REVIEW 1 major objections 4 minor 41 references
Psychological competence — how an AI interaction shapes a user's reasoning, emotions, and decisions — is proposed as a missing dimension in AI evaluation.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 07:52 UTC pith:HUOXC6KE
load-bearing objection A clear, honest conceptual framework for an interaction-level evaluation dimension; its proxy-assessment path is unproven but openly acknowledged. the 1 major comments →
Psychological Competence as a Missing Dimension in AI Evaluation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that the unit of evaluation for human-facing AI should shift from the model in isolation to the human-AI interaction. Psychological competence is the missing dimension: the ability of the system to generate interactions that appropriately reflect user context and support accurate reasoning, emotional stability, and autonomous decision-making, while avoiding distortion of judgment, reinforcement of harmful beliefs, or over-reliance. The paper formalizes the construct conceptually, identifies five domains, and outlines a mechanism pathway by which outputs produce psychological effects, while proposing scenario-based probes, human expert panels, and AI-as-judge pipelines as
What carries the argument
The central machinery is the construct of psychological competence itself, broken into five domains (context sensitivity, emotional responsiveness, social cognition, behavioral influence, developmental sensitivity) and tied to a stated mechanism pathway that runs from AI output to user interpretation, weighting through perceived authority and fluency, integration into reasoning and decisions, and feedback over repeated interactions. This mapping gives evaluators specific interaction properties — framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance — to probe as proxies for downstream effects on cognition, emotion, and behavior.
Load-bearing premise
The framework's value rests on the assumption that structured proxy assessments — scenario probes, expert ratings, and AI-as-judge pipelines — can validly estimate how real human-AI interactions affect users' reasoning, emotions, and behavior; if proxies do not track actual user outcomes, the construct remains unvalidated.
What would settle it
A randomized trial in which two AI systems produce equally accurate responses but differ on psychological competence ratings; if the higher-rated system fails to improve users' decision quality, emotional stability, or independent judgment, the construct's added value is not demonstrated. A simpler check: if LLM-based judges disagree systematically with human expert ratings on the five domains, the proposed pre-deployment assessment method needs rethinking.
If this is right
- Evaluation suites for conversational AI would add interaction-level probes alongside accuracy and safety tests, asking not just whether a response is correct but whether it preserves user agency and calibrates trust.
- Model providers could use the framework to design for agency preservation and calibrated trust during prompting, tuning, and interaction design, making behavioral interaction quality an explicit development target.
- Deploying organizations in healthcare, education, and mental health could incorporate psychological competence into procurement and assurance decisions, complementing existing safety and alignment checks.
- Regulators could draw on the framework when specifying behavioral safety expectations for high-impact systems, complementing risk-based governance frameworks.
- Assessment would combine scalable AI-as-judge checks for pre-deployment screening with human expert panels for psychological impact and psychometric measures for longitudinal tracking.
Where Pith is reading between the lines
- The framework's practical value depends on proxy validity: the authors acknowledge that real effects need human-subject studies, so an obvious test is whether scenario-based competence scores predict user outcomes in controlled trials.
- The mechanism pathway suggests a testable prediction: sustained interaction with systems rated high in psychological competence should reduce overreliance and belief offloading relative to output-equivalent controls over time.
- The five domains could be operationalized into a benchmark with expert-rated gold examples, which would also expose whether AI-as-judge ratings track human expert judgments on autonomy and vulnerability sensitivity.
- The construct may give regulators a language for tying behavioral harm to specific interaction properties, but only if autonomy and vulnerability sensitivity can be measured reliably across diverse contexts and user groups.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This conceptual paper argues that current AI evaluation frameworks are model-centric, focusing on output accuracy, robustness, safety, and policy compliance, while neglecting the interaction-level effects of human-facing AI on users' cognition, emotion, and behavior. The authors introduce 'psychological competence' as a new evaluation construct, define it, propose five domains (context sensitivity, emotional responsiveness, social cognition, behavioral influence, developmental sensitivity), connect these to a mechanism model of AI influence, and outline a mixed assessment strategy involving AI-as-judge pipelines, expert panels, and psychometric measures. The paper explicitly states that it is not proposing a benchmark, and it presents no empirical data; the contribution is intended as a conceptual foundation for future evaluation research and governance.
Significance. If the construct gains traction, this paper could provide a common vocabulary and conceptual scaffolding for evaluating human-facing AI beyond factual correctness, complementing technical and safety evaluations with a focus on user-level outcomes. The main strength is its careful scoping: the authors do not overclaim to have measured the construct, and they explicitly flag the central validity risk—that proxy assessments may not track real-world user outcomes—in Section 9 and Figure 3. This self-awareness makes the proposal credible. However, the operational value remains entirely unproven: there is no benchmark, no data, and no validation protocol. For a position paper this is acceptable, but it means the practical impact depends on substantial future empirical work. The paper should be read as a proposal for a research program rather than as an instantiated evaluation methodology.
major comments (1)
- [Section 8 / Figure 3] The paper recommends scenario-based probes, expert panels, and LLM-as-judge pipelines as pre-deployment assessment tools for psychological competence. The validity of these proxies for downstream user outcomes is not established, and the authors concede this in Section 9 ('these effects can ultimately only be understood through empirical studies involving human participants in context') and in Figure 3 (LLM judges 'may systematically inflate scores vs. human judgment'). Because the paper's call to make psychological competence 'a core consideration' for providers and regulators rests on this assessment path, the reader is left with no criteria for when a proxy score supports a procurement or regulatory decision. I recommend adding a short 'validation agenda' to Section 8 that specifies minimal convergent evidence (e.g., agreement with human-rated outcomes, sensitivity to known group diff
minor comments (4)
- [Section 4] The statement that related notions of psychological competence 'haven't yet been formalized as an evaluation construct for AI systems' is too categorical given existing frameworks such as FAST (ref. [39]) and prior human-robot interaction metrics. Suggest softening to 'to our knowledge' and briefly discussing how the proposed construct differs from adjacent frameworks.
- [Figure 1] The mockup conversation screenshots include timestamps and layout elements that are difficult to read. If this is a real figure, ensure the text is legible and the domain-assessment table is clearly aligned with System A and System B.
- [References] Reference formatting is inconsistent: some preprints use PsyArXiv DOIs (refs. [14], [24]), while others mix 'arXiv' and 'doi:10.48550/arXiv.x' (ref. [21]). Standardize to the journal's preferred style for preprints.
- [Section 8 / Figure 3] The 'Scale: High/Low/Medium' labels in Figure 3 are undefined. Clarify whether they refer to throughput, cost, reliability, or another property.
Circularity Check
No significant circularity: the paper is a conceptual proposal with no fitted parameters or derived predictions; one minor self-citation is not load-bearing.
full rationale
The paper makes a conceptual claim: current AI evaluation focuses on model-level output and safety, and a distinct interaction-level construct—psychological competence—is missing. This claim is supported by external behavioral-science literature (e.g., [4, 9, 12, 17, 25]), not by defining the target into existence. There are no equations, fitted parameters, or normalization choices, so no prediction reduces to an input by construction. The proposed assessment toolkit (scenario probes, expert panels, LLM-as-judge) is presented as a pre-deployment proxy, and the paper explicitly concedes in Section 9 that 'these effects can ultimately only be understood through empirical studies involving human participants in context' and in Figure 3 that AI-as-judge 'may systematically inflate scores vs. human judgment.' That is an acknowledged validity gap, not a circularity. The only self-citation is reference [3] (Sacher et al., 'The missing discipline in AI'), cited in Section 8 to support the call for interdisciplinary collaboration; it is not load-bearing for the central construct definition or the claim that the dimension is missing. Accordingly, the paper scores 1 for a minor self-citation with independent central content.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Human-facing AI systems function as behavioral interventions deployed at scale, shaping interpretation, emotion, and decisions.
- domain assumption Users treat conversational systems as social actors and respond to them with social and cognitive heuristics.
- domain assumption A five-stage pathway (output → interpretation → weighting → integration → feedback) is the correct mechanism model for AI-induced psychological effects.
- domain assumption Structured proxy measures (scenario probes, expert ratings, AI-as-judge) can estimate psychological interaction quality before deployment.
- domain assumption The five domains correspond one-to-one to stages of the mechanism model.
invented entities (1)
-
Psychological competence (as an evaluation construct)
no independent evidence
read the original abstract
Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance. These measures remain essential, but they are not sufficient for systems that interact directly with users through natural language. Human-facing AI systems are increasingly used as advisors, coaches, tutors, and companions. In these roles, their responses can shape how users reason, interpret emotions, form beliefs, calibrate trust, and make decisions. The relevant unit of evaluation is therefore not only the model, but the human-AI interaction. This paper introduces psychological competence as a missing dimension in AI evaluation. We define psychological competence as the capacity of a human-facing AI system to support user cognition, emotional interpretation, and behavioral decision-making in ways that are appropriate to the user, context, and purpose of the interaction. This includes interaction properties such as framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance. Existing evaluation approaches capture parts of this problem but rarely assess these psychological effects directly. Drawing on behavioral science and human-AI interaction research, we outline a conceptual framework for psychological competence and its core domains. Rather than proposing a specific benchmark, we define the construct, clarify its boundaries, and describe how it may be assessed through scenario-based probes, structured human evaluation, and model-assisted evaluation methods. We argue that psychological competence should become a core consideration for model providers, deploying organizations, researchers, and regulators concerned with the real-world effects of human-facing AI systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Psychologists must be involved in building conversational AI chatbots
Zhao X. Psychologists must be involved in building conversational AI chatbots. Nat Rev Psychol. Published online April 14, 2026. doi:10.1038/s44159-026-00564-z. 10
-
[2]
Rahwan I, Cebrian M, Obradovich N, et al. Machine behaviour. Nature. 2019;568(7753):477-486. doi:10.1038/s41586-019-1138-y
-
[3]
The missing discipline in AI: a call for behavioural science
Sacher PM, Michie S, Hauser OP, et al. The missing discipline in AI: a call for behavioural science. Wellcome Open Res. 2026;11:152. doi:10.12688/wellcomeopenres.25922.1
-
[5]
Li RN, Folk D, Singh A, Ungar L, Dunn E. Is a random human peer better than a highly supportive chatbot in reducing loneliness over time? J Exp Soc Psychol. 2026;125:104911. doi:10.1016/j.jesp.2026. 104911
-
[6]
What large language models know and what people think they know
Steyvers M, Tejeda H, Kumar A, et al. What large language models know and what people think they know. Nat Mach Intell. 2025;7(2):221-231. doi:10.1038/s42256-024-00976-7
-
[7]
Dergaa I, Ben Saad H, Glenn JM, et al. From tools to threats: a reflection on the impact of artificial- intelligence chatbots on cognitive health. Front Psychol. 2024;15:1259845. doi:10.3389/fpsyg.2024. 1259845
-
[8]
On the conversational persuasiveness of GPT-4
Salvi F, Horta Ribeiro M, Gallotti R, West R. On the conversational persuasiveness of GPT-4. Nat Hum Behav. 2025;9(8):1645-1653. doi:10.1038/s41562-025-02194-6
-
[9]
How human–AI feedback loops alter human perceptual, emotional and social judgements
Glickman M, Sharot T. How human–AI feedback loops alter human perceptual, emotional and social judgements. Nat Hum Behav. 2024;9(2):345-359. doi:10.1038/s41562-024-02077-2
-
[10]
Why human–AI relationships need socioaffective alignment
Kirk HR, Gabriel I, Summerfield C, Vidgen B, Hale SA. Why human–AI relationships need socioaffective alignment. Humanit Soc Sci Commun. 2025;12(1):728. doi:10.1057/s41599-025-04532-5
-
[11]
Exploring the Ethical Challenges of Conversational AI in Mental Health Care: Scoping Review
Rahsepar Meadi M, Sillekens T, Metselaar S, Van Balkom A, Bernstein J, Batelaan N. Exploring the Ethical Challenges of Conversational AI in Mental Health Care: Scoping Review. JMIR Ment Health. 2025;12:e60432. doi:10.2196/60432
doi:10.2196/60432 2025
-
[12]
Chong L, Zhang G, Goucher-Lambert K, Kotovsky K, Cagan J. Human confidence in artificial intelligence and in themselves: The evolution and impact of confidence on adoption of AI advice. Comput Hum Behav. 2022;127:107018. doi:10.1016/j.chb.2021.107018
arXiv 2022
-
[13]
How AI Impacts Skill Formation
Shen JH, Tamkin A. How AI Impacts Skill Formation. arXiv. Preprint posted online 2026. doi:10.48550/ ARXIV.2601.20245
-
[14]
How AI can fuel confirmation bias
Rathje S, Van Bavel JJ. How AI can fuel confirmation bias. PsyArXiv. Preprint posted online March 13,
-
[15]
Ehsan U, Passi S, Saha K, McNutt T, Riedl MO, Alcorn S. From Future of Work to Future of Workers: Addressing Asymptomatic AI Harms for Dignified Human-AI Interaction. Preprint posted online January 29, 2026. doi:10.1145/3772318.3791081
arXiv 2026
-
[16]
Tian J, Zhang R. Learners’ AI dependence and critical thinking: The psychological mechanism of fatigue and the social buffering role of AI literacy. Acta Psychol (Amst). 2025;260:105725. doi:10.1016/j.actpsy. 2025.105725
arXiv 2025
-
[17]
Sycophantic AI decreases prosocial intentions and promotes dependence
Cheng M, Lee C, Khadpe P, Yu S, Han D, Jurafsky D. Sycophantic AI decreases prosocial intentions and promotes dependence. Science. 2026;391(6792):eaec8352. doi:10.1126/science.aec8352
-
[18]
Nat Mach Intell
Emotional risks of AI companions demand attention. Nat Mach Intell. 2025;7(7):981-982. doi:10.1038/ s42256-025-01093-9
2025
-
[19]
Measuring Massive Multitask Language Understanding
Hendrycks D, Burns C, Basart S, et al. Measuring Massive Multitask Language Understanding. arXiv. Preprint posted online 2020. doi:10.48550/ARXIV.2009.03300
-
[20]
Evaluating Large Language Models Trained on Code
Chen M, Tworek J, Jun H, et al. Evaluating Large Language Models Trained on Code. arXiv. Preprint posted online 2021. doi:10.48550/ARXIV.2107.03374
-
[21]
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Zheng L, Chiang WL, Sheng Y, et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv. Preprint posted online December 24, 2023:arXiv:2306.05685. doi:10.48550/arXiv.2306.05685. 11
-
[22]
Artificial Intelligence Risk Management Framework (AI RMF 1.0)
Tabassi E. Artificial Intelligence Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology (U.S.); 2023:NIST AI 100-1. doi:10.6028/NIST.AI.100-1
-
[23]
Holistic Evaluation of Language Models
Liang P, Bommasani R, Lee T, et al. Holistic Evaluation of Language Models. Published online 2022. doi:10.48550/ARXIV.2211.09110
-
[24]
Sycophantic AI increases attitude extremity and overconfidence
Rathje S, Ye M, Globig LK, Pillai RM, De Mello VO, Van Bavel JJ. Sycophantic AI increases attitude extremity and overconfidence. PsyArXiv. Preprint posted online September 28, 2025. doi:10.31234/osf. io/vmyek_v1
doi:10.31234/osf 2025
-
[25]
Trust and reliance on AI — An experimental study on the extent and costs of overreliance on AI
Klingbeil A, Grützner C, Schreck P. Trust and reliance on AI — An experimental study on the extent and costs of overreliance on AI. Comput Hum Behav. 2024;160:108352. doi:10.1016/j.chb.2024.108352
arXiv 2024
-
[26]
Bad machines corrupt good morals
Köbis N, Bonnefon JF, Rahwan I. Bad machines corrupt good morals. Nat Hum Behav. 2021;5(6):679-685. doi:10.1038/s41562-021-01128-2
-
[27]
Imagining and building wise machines: the centrality of AI metacognition
Johnson SGB, Karimi AH, Bengio Y, et al. Imagining and building wise machines: the centrality of AI metacognition. Trends Cogn Sci. Published online February 2026:S1364661326000021. doi:10.1016/j.tics. 2026.01.002
doi:10.1016/j.tics 2026
-
[28]
Calibrating workers’ trust in intelligent automated systems
Lucas GM, Becerik-Gerber B, Roll SC. Calibrating workers’ trust in intelligent automated systems. Patterns. 2024;5(9):101045. doi:10.1016/j.patter.2024.101045
arXiv 2024
-
[29]
Human Trust in Artificial Intelligence: Review of Empirical Research
Glikson E, Woolley AW. Human Trust in Artificial Intelligence: Review of Empirical Research. Acad Manag Ann. 2020;14(2):627-660. doi:10.5465/annals.2018.0057
arXiv 2020
-
[30]
Belief Offloading in Human-AI Interaction
Guingrich RE, Mehta D, Bhatt U. Belief Offloading in Human-AI Interaction. arXiv. Preprint posted online 2026. doi:10.48550/ARXIV.2602.08754
-
[31]
Privacy in the age of psychological targeting
Matz SC, Appel RE, Kosinski M. Privacy in the age of psychological targeting. Curr Opin Psychol. 2020;31:116-121. doi:10.1016/j.copsyc.2019.08.010
-
[32]
Trust in AI: progress, challenges, and future directions
Afroogh S, Akbari A, Malone E, Kargar M, Alambeigi H. Trust in AI: progress, challenges, and future directions. Humanit Soc Sci Commun. 2024;11(1):1568. doi:10.1057/s41599-024-04044-8
-
[33]
Bonaccio S, Dalal RS. Advice taking and decision-making: An integrative literature review, and implications for the organizational sciences. Organ Behav Hum Decis Process. 2006;101(2):127-151. doi:10.1016/j.obhdp.2006.07.001
-
[34]
Receiving other people’s advice: Influence and benefit
Yaniv I. Receiving other people’s advice: Influence and benefit. Organ Behav Hum Decis Process. 2004;93(1):1-13. doi:10.1016/j.obhdp.2003.08.002
-
[35]
Kahr PK, Rooks G, Willemsen MC, Snijders CCP. Understanding Trust and Reliance Development in AI Advice: Assessing Model Accuracy, Model Explanations, and Experiences from Previous Interactions. ACM Trans Interact Intell Syst. 2024;14(4):1-30. doi:10.1145/3686164
doi:10.1145/3686164 2024
-
[36]
The Labor Illusion: How Operational Transparency Increases Perceived Value
Buell RW, Norton MI. The Labor Illusion: How Operational Transparency Increases Perceived Value. Manag Sci. 2011;57(9):1564-1579. doi:10.1287/mnsc.1110.1376
Pith/arXiv arXiv 2011
-
[37]
AI-teaming: Redefining collaboration in the digital era
Schmutz JB, Outland N, Kerstan S, Georganta E, Ulfert AS. AI-teaming: Redefining collaboration in the digital era. Curr Opin Psychol. 2024;58:101837. doi:10.1016/j.copsyc.2024.101837
arXiv 2024
-
[38]
From Generation to Judgment: Opportunities and Challenges of LLM-as- a-judge
Li D, Jiang B, Huang L, et al. From Generation to Judgment: Opportunities and Challenges of LLM-as- a-judge. arXiv. Preprint posted online September 29, 2025:arXiv:2411.16594. doi:10.48550/arXiv.2411. 16594
-
[39]
Neary M, Fulton E, Rogers V, et al. Think FAST: a novel framework to evaluate fidelity, accuracy, safety, and tone in conversational AI health coach dialogues. Front Digit Health. 2025;7:1460236. doi:10.3389/fdgth.2025.1460236
arXiv 2025
-
[40]
A cognitive approach to human–AI complementarity in dynamic decision-making
Gonzalez C, Heidari H. A cognitive approach to human–AI complementarity in dynamic decision-making. Nat Rev Psychol. 2025;4(12):808-822. doi:10.1038/s44159-025-00499-x
-
[41]
Addressing Longstanding Challenges in Cognitive Science with Language Models
Wulff DU, Mata R. Addressing Longstanding Challenges in Cognitive Science with Language Models. arXiv. Preprint posted online 2025. doi:10.48550/ARXIV.2511.00206. 12
-
[2026]
doi:10.31234/osf.io/7a3d4_v1
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.