Pith. sign in

REVIEW 1 major objections 4 minor 41 references

Psychological competence — how an AI interaction shapes a user's reasoning, emotions, and decisions — is proposed as a missing dimension in AI evaluation.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 07:52 UTC pith:HUOXC6KE

load-bearing objection A clear, honest conceptual framework for an interaction-level evaluation dimension; its proxy-assessment path is unproven but openly acknowledged. the 1 major comments →

arxiv 2607.08285 v2 pith:HUOXC6KE submitted 2026-07-09 cs.AI

Psychological Competence as a Missing Dimension in AI Evaluation

classification cs.AI
keywords AI evaluationhuman-AI interactionpsychological competencebehavioral sciencetrust calibrationuser autonomyAI safetyinteraction quality
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that current AI benchmarks measure what a model can do — accuracy, reasoning, robustness, policy compliance — but not what interacting with the system does to the user. It proposes psychological competence as a new evaluation dimension: the capacity of a human-facing AI to support accurate reasoning, emotional stability, and autonomous decision-making while avoiding distortions of judgment, harmful belief reinforcement, and over-reliance. The construct is organized into five domains — context sensitivity, emotional responsiveness, social cognition, behavioral influence, and developmental sensitivity — tied to a pathway from AI output to user interpretation, weighting, integration, and feedback. The authors argue that evaluating this dimension is essential for AI systems used as advisors, tutors, coaches, and companions, where responses can shape beliefs and choices.

Core claim

The central claim is that the unit of evaluation for human-facing AI should shift from the model in isolation to the human-AI interaction. Psychological competence is the missing dimension: the ability of the system to generate interactions that appropriately reflect user context and support accurate reasoning, emotional stability, and autonomous decision-making, while avoiding distortion of judgment, reinforcement of harmful beliefs, or over-reliance. The paper formalizes the construct conceptually, identifies five domains, and outlines a mechanism pathway by which outputs produce psychological effects, while proposing scenario-based probes, human expert panels, and AI-as-judge pipelines as

What carries the argument

The central machinery is the construct of psychological competence itself, broken into five domains (context sensitivity, emotional responsiveness, social cognition, behavioral influence, developmental sensitivity) and tied to a stated mechanism pathway that runs from AI output to user interpretation, weighting through perceived authority and fluency, integration into reasoning and decisions, and feedback over repeated interactions. This mapping gives evaluators specific interaction properties — framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance — to probe as proxies for downstream effects on cognition, emotion, and behavior.

Load-bearing premise

The framework's value rests on the assumption that structured proxy assessments — scenario probes, expert ratings, and AI-as-judge pipelines — can validly estimate how real human-AI interactions affect users' reasoning, emotions, and behavior; if proxies do not track actual user outcomes, the construct remains unvalidated.

What would settle it

A randomized trial in which two AI systems produce equally accurate responses but differ on psychological competence ratings; if the higher-rated system fails to improve users' decision quality, emotional stability, or independent judgment, the construct's added value is not demonstrated. A simpler check: if LLM-based judges disagree systematically with human expert ratings on the five domains, the proposed pre-deployment assessment method needs rethinking.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Evaluation suites for conversational AI would add interaction-level probes alongside accuracy and safety tests, asking not just whether a response is correct but whether it preserves user agency and calibrates trust.
  • Model providers could use the framework to design for agency preservation and calibrated trust during prompting, tuning, and interaction design, making behavioral interaction quality an explicit development target.
  • Deploying organizations in healthcare, education, and mental health could incorporate psychological competence into procurement and assurance decisions, complementing existing safety and alignment checks.
  • Regulators could draw on the framework when specifying behavioral safety expectations for high-impact systems, complementing risk-based governance frameworks.
  • Assessment would combine scalable AI-as-judge checks for pre-deployment screening with human expert panels for psychological impact and psychometric measures for longitudinal tracking.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The framework's practical value depends on proxy validity: the authors acknowledge that real effects need human-subject studies, so an obvious test is whether scenario-based competence scores predict user outcomes in controlled trials.
  • The mechanism pathway suggests a testable prediction: sustained interaction with systems rated high in psychological competence should reduce overreliance and belief offloading relative to output-equivalent controls over time.
  • The five domains could be operationalized into a benchmark with expert-rated gold examples, which would also expose whether AI-as-judge ratings track human expert judgments on autonomy and vulnerability sensitivity.
  • The construct may give regulators a language for tying behavioral harm to specific interaction properties, but only if autonomy and vulnerability sensitivity can be measured reliably across diverse contexts and user groups.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. This conceptual paper argues that current AI evaluation frameworks are model-centric, focusing on output accuracy, robustness, safety, and policy compliance, while neglecting the interaction-level effects of human-facing AI on users' cognition, emotion, and behavior. The authors introduce 'psychological competence' as a new evaluation construct, define it, propose five domains (context sensitivity, emotional responsiveness, social cognition, behavioral influence, developmental sensitivity), connect these to a mechanism model of AI influence, and outline a mixed assessment strategy involving AI-as-judge pipelines, expert panels, and psychometric measures. The paper explicitly states that it is not proposing a benchmark, and it presents no empirical data; the contribution is intended as a conceptual foundation for future evaluation research and governance.

Significance. If the construct gains traction, this paper could provide a common vocabulary and conceptual scaffolding for evaluating human-facing AI beyond factual correctness, complementing technical and safety evaluations with a focus on user-level outcomes. The main strength is its careful scoping: the authors do not overclaim to have measured the construct, and they explicitly flag the central validity risk—that proxy assessments may not track real-world user outcomes—in Section 9 and Figure 3. This self-awareness makes the proposal credible. However, the operational value remains entirely unproven: there is no benchmark, no data, and no validation protocol. For a position paper this is acceptable, but it means the practical impact depends on substantial future empirical work. The paper should be read as a proposal for a research program rather than as an instantiated evaluation methodology.

major comments (1)
  1. [Section 8 / Figure 3] The paper recommends scenario-based probes, expert panels, and LLM-as-judge pipelines as pre-deployment assessment tools for psychological competence. The validity of these proxies for downstream user outcomes is not established, and the authors concede this in Section 9 ('these effects can ultimately only be understood through empirical studies involving human participants in context') and in Figure 3 (LLM judges 'may systematically inflate scores vs. human judgment'). Because the paper's call to make psychological competence 'a core consideration' for providers and regulators rests on this assessment path, the reader is left with no criteria for when a proxy score supports a procurement or regulatory decision. I recommend adding a short 'validation agenda' to Section 8 that specifies minimal convergent evidence (e.g., agreement with human-rated outcomes, sensitivity to known group diff
minor comments (4)
  1. [Section 4] The statement that related notions of psychological competence 'haven't yet been formalized as an evaluation construct for AI systems' is too categorical given existing frameworks such as FAST (ref. [39]) and prior human-robot interaction metrics. Suggest softening to 'to our knowledge' and briefly discussing how the proposed construct differs from adjacent frameworks.
  2. [Figure 1] The mockup conversation screenshots include timestamps and layout elements that are difficult to read. If this is a real figure, ensure the text is legible and the domain-assessment table is clearly aligned with System A and System B.
  3. [References] Reference formatting is inconsistent: some preprints use PsyArXiv DOIs (refs. [14], [24]), while others mix 'arXiv' and 'doi:10.48550/arXiv.x' (ref. [21]). Standardize to the journal's preferred style for preprints.
  4. [Section 8 / Figure 3] The 'Scale: High/Low/Medium' labels in Figure 3 are undefined. Clarify whether they refer to throughput, cost, reliability, or another property.

Circularity Check

0 steps flagged

No significant circularity: the paper is a conceptual proposal with no fitted parameters or derived predictions; one minor self-citation is not load-bearing.

full rationale

The paper makes a conceptual claim: current AI evaluation focuses on model-level output and safety, and a distinct interaction-level construct—psychological competence—is missing. This claim is supported by external behavioral-science literature (e.g., [4, 9, 12, 17, 25]), not by defining the target into existence. There are no equations, fitted parameters, or normalization choices, so no prediction reduces to an input by construction. The proposed assessment toolkit (scenario probes, expert panels, LLM-as-judge) is presented as a pre-deployment proxy, and the paper explicitly concedes in Section 9 that 'these effects can ultimately only be understood through empirical studies involving human participants in context' and in Figure 3 that AI-as-judge 'may systematically inflate scores vs. human judgment.' That is an acknowledged validity gap, not a circularity. The only self-citation is reference [3] (Sacher et al., 'The missing discipline in AI'), cited in Section 8 to support the call for interdisciplinary collaboration; it is not load-bearing for the central construct definition or the claim that the dimension is missing. Accordingly, the paper scores 1 for a minor self-citation with independent central content.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 1 invented entities

The paper is conceptual and literature-driven. It introduces one new construct — psychological competence — with no independent empirical evidence yet. The axioms are domain assumptions from behavioral science and HCI about how AI interactions influence users; the most load-bearing is that proxy assessments can validly estimate real-world psychological effects, which the authors themselves flag as unvalidated.

axioms (5)
  • domain assumption Human-facing AI systems function as behavioral interventions deployed at scale, shaping interpretation, emotion, and decisions.
    Section 1, paras 1–2; the foundational premise for why interaction-level evaluation is needed.
  • domain assumption Users treat conversational systems as social actors and respond to them with social and cognitive heuristics.
    Section 1, citing [10, 11]; supports the claim that AI outputs carry psychological influence beyond content.
  • domain assumption A five-stage pathway (output → interpretation → weighting → integration → feedback) is the correct mechanism model for AI-induced psychological effects.
    Section 5; the entire domain structure is derived from this assumed pathway.
  • domain assumption Structured proxy measures (scenario probes, expert ratings, AI-as-judge) can estimate psychological interaction quality before deployment.
    Sections 4 and 8; the paper's entire proposed assessment program depends on this; authors flag the need for empirical validation in Section 9 and Figure 3.
  • domain assumption The five domains correspond one-to-one to stages of the mechanism model.
    Section 6; the mapping is asserted, not derived.
invented entities (1)
  • Psychological competence (as an evaluation construct) no independent evidence
    purpose: Names a new dimension for evaluating human-AI interaction quality covering cognition, emotion, and behavior.
    Provides a definition and domain taxonomy but no measurable instrument, benchmark, or falsifiable prediction independent of the paper; utility depends on future operationalization.

pith-pipeline@v1.3.0-alltime-deepseek · 2806 in / 4092 out tokens · 138084 ms · 2026-08-02T07:52:16.611383+00:00 · methodology

0 comments
read the original abstract

Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance. These measures remain essential, but they are not sufficient for systems that interact directly with users through natural language. Human-facing AI systems are increasingly used as advisors, coaches, tutors, and companions. In these roles, their responses can shape how users reason, interpret emotions, form beliefs, calibrate trust, and make decisions. The relevant unit of evaluation is therefore not only the model, but the human-AI interaction. This paper introduces psychological competence as a missing dimension in AI evaluation. We define psychological competence as the capacity of a human-facing AI system to support user cognition, emotional interpretation, and behavioral decision-making in ways that are appropriate to the user, context, and purpose of the interaction. This includes interaction properties such as framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance. Existing evaluation approaches capture parts of this problem but rarely assess these psychological effects directly. Drawing on behavioral science and human-AI interaction research, we outline a conceptual framework for psychological competence and its core domains. Rather than proposing a specific benchmark, we define the construct, clarify its boundaries, and describe how it may be assessed through scenario-based probes, structured human evaluation, and model-assisted evaluation methods. We argue that psychological competence should become a core consideration for model providers, deploying organizations, researchers, and regulators concerned with the real-world effects of human-facing AI systems.

Figures

Figures reproduced from arXiv: 2607.08285 by Alexis Michelle Abellar, Antoine Ferr\`ere, Fendi Tsim, Marcos Economides, Paul M. Sacher, Samuel Salzer.

Figure 1
Figure 1. Figure 1: Illustrative evaluation of psychological competence in an interpersonal [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 1
Figure 1. Figure 1: Illustrative evaluation of psychological competence in an interpersonal conflict scenario. System A (Standard) and System B (Competent) receive identical user inputs. System A validates the user’s assumption of malicious intent and recommends direct confrontation. System B acknowledges the emotional content, holds multiple interpretations open, and returns agency to the user. The domain assessment (below) … view at source ↗
Figure 2
Figure 2. Figure 2: Conceptual model of expanded AI evaluation. [PITH_FULL_IMAGE:figures/full_fig_p014_2.png] view at source ↗
Figure 2
Figure 2. Figure 2: summarizes the relationship between traditional technical benchmarks and the additional evaluation layer proposed here. Existing benchmarks primarily ask what the model can do; psychological competence asks what the interaction does to the user. Model-Level Evaluation "What can the model do?" Interaction-Level Evaluation "What does the interaction do to the user?" User-Level Outcomes "How does interaction … view at source ↗
Figure 3
Figure 3. Figure 3: Complementary evaluation approaches for psychological competence. [PITH_FULL_IMAGE:figures/full_fig_p016_3.png] view at source ↗
Figure 3
Figure 3. Figure 3: Complementary evaluation approaches for psychological competence. AI-as-Judge evaluation supports scalable pre-deployment assessment; human expert panels provide psychological impact ratings during deployment; psychometric measures track longitudinal outcomes post-deployment. Finally, psychological competence has implications for governance. If AI systems shape trust, rea￾soning, and behavior, then evaluat… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

41 extracted references · 9 canonical work pages

  1. [1]

    Psychologists must be involved in building conversational AI chatbots

    Zhao X. Psychologists must be involved in building conversational AI chatbots. Nat Rev Psychol. Published online April 14, 2026. doi:10.1038/s44159-026-00564-z. 10

  2. [2]

    Machine behaviour

    Rahwan I, Cebrian M, Obradovich N, et al. Machine behaviour. Nature. 2019;568(7753):477-486. doi:10.1038/s41586-019-1138-y

  3. [3]

    The missing discipline in AI: a call for behavioural science

    Sacher PM, Michie S, Hauser OP, et al. The missing discipline in AI: a call for behavioural science. Wellcome Open Res. 2026;11:152. doi:10.12688/wellcomeopenres.25922.1

  4. [5]

    Is a random human peer better than a highly supportive chatbot in reducing loneliness over time? J Exp Soc Psychol

    Li RN, Folk D, Singh A, Ungar L, Dunn E. Is a random human peer better than a highly supportive chatbot in reducing loneliness over time? J Exp Soc Psychol. 2026;125:104911. doi:10.1016/j.jesp.2026. 104911

  5. [6]

    What large language models know and what people think they know

    Steyvers M, Tejeda H, Kumar A, et al. What large language models know and what people think they know. Nat Mach Intell. 2025;7(2):221-231. doi:10.1038/s42256-024-00976-7

  6. [7]

    From tools to threats: a reflection on the impact of artificial- intelligence chatbots on cognitive health

    Dergaa I, Ben Saad H, Glenn JM, et al. From tools to threats: a reflection on the impact of artificial- intelligence chatbots on cognitive health. Front Psychol. 2024;15:1259845. doi:10.3389/fpsyg.2024. 1259845

  7. [8]

    On the conversational persuasiveness of GPT-4

    Salvi F, Horta Ribeiro M, Gallotti R, West R. On the conversational persuasiveness of GPT-4. Nat Hum Behav. 2025;9(8):1645-1653. doi:10.1038/s41562-025-02194-6

  8. [9]

    How human–AI feedback loops alter human perceptual, emotional and social judgements

    Glickman M, Sharot T. How human–AI feedback loops alter human perceptual, emotional and social judgements. Nat Hum Behav. 2024;9(2):345-359. doi:10.1038/s41562-024-02077-2

  9. [10]

    Why human–AI relationships need socioaffective alignment

    Kirk HR, Gabriel I, Summerfield C, Vidgen B, Hale SA. Why human–AI relationships need socioaffective alignment. Humanit Soc Sci Commun. 2025;12(1):728. doi:10.1057/s41599-025-04532-5

  10. [11]

    Exploring the Ethical Challenges of Conversational AI in Mental Health Care: Scoping Review

    Rahsepar Meadi M, Sillekens T, Metselaar S, Van Balkom A, Bernstein J, Batelaan N. Exploring the Ethical Challenges of Conversational AI in Mental Health Care: Scoping Review. JMIR Ment Health. 2025;12:e60432. doi:10.2196/60432

  11. [12]

    Human confidence in artificial intelligence and in themselves: The evolution and impact of confidence on adoption of AI advice

    Chong L, Zhang G, Goucher-Lambert K, Kotovsky K, Cagan J. Human confidence in artificial intelligence and in themselves: The evolution and impact of confidence on adoption of AI advice. Comput Hum Behav. 2022;127:107018. doi:10.1016/j.chb.2021.107018

  12. [13]

    How AI Impacts Skill Formation

    Shen JH, Tamkin A. How AI Impacts Skill Formation. arXiv. Preprint posted online 2026. doi:10.48550/ ARXIV.2601.20245

  13. [14]

    How AI can fuel confirmation bias

    Rathje S, Van Bavel JJ. How AI can fuel confirmation bias. PsyArXiv. Preprint posted online March 13,

  14. [15]

    From Future of Work to Future of Workers: Addressing Asymptomatic AI Harms for Dignified Human-AI Interaction

    Ehsan U, Passi S, Saha K, McNutt T, Riedl MO, Alcorn S. From Future of Work to Future of Workers: Addressing Asymptomatic AI Harms for Dignified Human-AI Interaction. Preprint posted online January 29, 2026. doi:10.1145/3772318.3791081

  15. [16]

    Learners’ AI dependence and critical thinking: The psychological mechanism of fatigue and the social buffering role of AI literacy

    Tian J, Zhang R. Learners’ AI dependence and critical thinking: The psychological mechanism of fatigue and the social buffering role of AI literacy. Acta Psychol (Amst). 2025;260:105725. doi:10.1016/j.actpsy. 2025.105725

  16. [17]

    Sycophantic AI decreases prosocial intentions and promotes dependence

    Cheng M, Lee C, Khadpe P, Yu S, Han D, Jurafsky D. Sycophantic AI decreases prosocial intentions and promotes dependence. Science. 2026;391(6792):eaec8352. doi:10.1126/science.aec8352

  17. [18]

    Nat Mach Intell

    Emotional risks of AI companions demand attention. Nat Mach Intell. 2025;7(7):981-982. doi:10.1038/ s42256-025-01093-9

  18. [19]

    Measuring Massive Multitask Language Understanding

    Hendrycks D, Burns C, Basart S, et al. Measuring Massive Multitask Language Understanding. arXiv. Preprint posted online 2020. doi:10.48550/ARXIV.2009.03300

  19. [20]

    Evaluating Large Language Models Trained on Code

    Chen M, Tworek J, Jun H, et al. Evaluating Large Language Models Trained on Code. arXiv. Preprint posted online 2021. doi:10.48550/ARXIV.2107.03374

  20. [21]

    Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

    Zheng L, Chiang WL, Sheng Y, et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv. Preprint posted online December 24, 2023:arXiv:2306.05685. doi:10.48550/arXiv.2306.05685. 11

  21. [22]

    Artificial Intelligence Risk Management Framework (AI RMF 1.0)

    Tabassi E. Artificial Intelligence Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology (U.S.); 2023:NIST AI 100-1. doi:10.6028/NIST.AI.100-1

  22. [23]

    Holistic Evaluation of Language Models

    Liang P, Bommasani R, Lee T, et al. Holistic Evaluation of Language Models. Published online 2022. doi:10.48550/ARXIV.2211.09110

  23. [24]

    Sycophantic AI increases attitude extremity and overconfidence

    Rathje S, Ye M, Globig LK, Pillai RM, De Mello VO, Van Bavel JJ. Sycophantic AI increases attitude extremity and overconfidence. PsyArXiv. Preprint posted online September 28, 2025. doi:10.31234/osf. io/vmyek_v1

  24. [25]

    Trust and reliance on AI — An experimental study on the extent and costs of overreliance on AI

    Klingbeil A, Grützner C, Schreck P. Trust and reliance on AI — An experimental study on the extent and costs of overreliance on AI. Comput Hum Behav. 2024;160:108352. doi:10.1016/j.chb.2024.108352

  25. [26]

    Bad machines corrupt good morals

    Köbis N, Bonnefon JF, Rahwan I. Bad machines corrupt good morals. Nat Hum Behav. 2021;5(6):679-685. doi:10.1038/s41562-021-01128-2

  26. [27]

    Imagining and building wise machines: the centrality of AI metacognition

    Johnson SGB, Karimi AH, Bengio Y, et al. Imagining and building wise machines: the centrality of AI metacognition. Trends Cogn Sci. Published online February 2026:S1364661326000021. doi:10.1016/j.tics. 2026.01.002

  27. [28]

    Calibrating workers’ trust in intelligent automated systems

    Lucas GM, Becerik-Gerber B, Roll SC. Calibrating workers’ trust in intelligent automated systems. Patterns. 2024;5(9):101045. doi:10.1016/j.patter.2024.101045

  28. [29]

    Human Trust in Artificial Intelligence: Review of Empirical Research

    Glikson E, Woolley AW. Human Trust in Artificial Intelligence: Review of Empirical Research. Acad Manag Ann. 2020;14(2):627-660. doi:10.5465/annals.2018.0057

  29. [30]

    Belief Offloading in Human-AI Interaction

    Guingrich RE, Mehta D, Bhatt U. Belief Offloading in Human-AI Interaction. arXiv. Preprint posted online 2026. doi:10.48550/ARXIV.2602.08754

  30. [31]

    Privacy in the age of psychological targeting

    Matz SC, Appel RE, Kosinski M. Privacy in the age of psychological targeting. Curr Opin Psychol. 2020;31:116-121. doi:10.1016/j.copsyc.2019.08.010

  31. [32]

    Trust in AI: progress, challenges, and future directions

    Afroogh S, Akbari A, Malone E, Kargar M, Alambeigi H. Trust in AI: progress, challenges, and future directions. Humanit Soc Sci Commun. 2024;11(1):1568. doi:10.1057/s41599-024-04044-8

  32. [33]

    Advice taking and decision-making: An integrative literature review, and implications for the organizational sciences

    Bonaccio S, Dalal RS. Advice taking and decision-making: An integrative literature review, and implications for the organizational sciences. Organ Behav Hum Decis Process. 2006;101(2):127-151. doi:10.1016/j.obhdp.2006.07.001

  33. [34]

    Receiving other people’s advice: Influence and benefit

    Yaniv I. Receiving other people’s advice: Influence and benefit. Organ Behav Hum Decis Process. 2004;93(1):1-13. doi:10.1016/j.obhdp.2003.08.002

  34. [35]

    Understanding Trust and Reliance Development in AI Advice: Assessing Model Accuracy, Model Explanations, and Experiences from Previous Interactions

    Kahr PK, Rooks G, Willemsen MC, Snijders CCP. Understanding Trust and Reliance Development in AI Advice: Assessing Model Accuracy, Model Explanations, and Experiences from Previous Interactions. ACM Trans Interact Intell Syst. 2024;14(4):1-30. doi:10.1145/3686164

  35. [36]

    The Labor Illusion: How Operational Transparency Increases Perceived Value

    Buell RW, Norton MI. The Labor Illusion: How Operational Transparency Increases Perceived Value. Manag Sci. 2011;57(9):1564-1579. doi:10.1287/mnsc.1110.1376

  36. [37]

    AI-teaming: Redefining collaboration in the digital era

    Schmutz JB, Outland N, Kerstan S, Georganta E, Ulfert AS. AI-teaming: Redefining collaboration in the digital era. Curr Opin Psychol. 2024;58:101837. doi:10.1016/j.copsyc.2024.101837

  37. [38]

    From Generation to Judgment: Opportunities and Challenges of LLM-as- a-judge

    Li D, Jiang B, Huang L, et al. From Generation to Judgment: Opportunities and Challenges of LLM-as- a-judge. arXiv. Preprint posted online September 29, 2025:arXiv:2411.16594. doi:10.48550/arXiv.2411. 16594

  38. [39]

    Think FAST: a novel framework to evaluate fidelity, accuracy, safety, and tone in conversational AI health coach dialogues

    Neary M, Fulton E, Rogers V, et al. Think FAST: a novel framework to evaluate fidelity, accuracy, safety, and tone in conversational AI health coach dialogues. Front Digit Health. 2025;7:1460236. doi:10.3389/fdgth.2025.1460236

  39. [40]

    A cognitive approach to human–AI complementarity in dynamic decision-making

    Gonzalez C, Heidari H. A cognitive approach to human–AI complementarity in dynamic decision-making. Nat Rev Psychol. 2025;4(12):808-822. doi:10.1038/s44159-025-00499-x

  40. [41]

    Addressing Longstanding Challenges in Cognitive Science with Language Models

    Wulff DU, Mata R. Addressing Longstanding Challenges in Cognitive Science with Language Models. arXiv. Preprint posted online 2025. doi:10.48550/ARXIV.2511.00206. 12

  41. [2026]

    doi:10.31234/osf.io/7a3d4_v1