{"id":"e42e9b6a-e9c0-4fb6-a25d-f1dcf9521df8","arxiv_id":"2607.08285","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"AI evaluation should add 'psychological competence' — how well a system supports user reasoning, emotional stability, and autonomous decision-making — as a core dimension.","lead":"This paper argues that AI systems should be evaluated on how they change the people they talk to — their feelings, beliefs, and decisions — not just on whether their answers are correct. It defines 'psychological competence' as a new evaluation layer with five domains, but stops short of delivering an actual benchmark.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proxy-validity gap is the main soft spot, but the paper's explicit limitations keep the conceptual claim sustainable.","rationale":"The reader's weakest-assumption identification is correct: the entire operational program depends on the validity of proxy assessments for real psychological effects. The paper is honest about this dependency, explicitly stating that human-subject research is required for validation and that AI-as-judge may inflate scores. For a conceptual contribution, this is a limitation rather than a fatal flaw: the paper does not overclaim that it has a validated benchmark. I also considered whether the 'missing dimension' claim is undermined by existing frameworks like FAST, but the paper uses hedged language ('rarely', 'under-specified') and positions psychological competence as a unified construct rather than claiming no related work exists. The concern that would move the verdict would be if the paper presented its proxy methods as already valid; it does not. Therefore the reader's ACCEPT verdict stands, and no change is needed.","tokens_in":8862,"tokens_out":6567,"duration_ms":69069,"concrete_test":"Run a validation study in one high-stakes domain (e.g., mental-health support): collect LLM-as-judge and expert-panel psychological-competence scores on a set of responses; then, in a preregistered human experiment, measure downstream user outcomes (e.g., state anxiety, overconfidence, intention to seek alternative help, decision agency). Compute correlation/calibration between proxy scores and user-level outcomes. If the correlation is weak (e.g., r < 0.3) or if AI-as-judge scores are systematically inflated relative to human experts, the pre-deployment proxy pipeline cannot be relied on as an estimate of psychological competence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two components: (1) psychological competence is missing from current evaluation frameworks, and (2) it can be assessed through structured proxy pipelines. The load-bearing link is (2). Figure 3 concedes that AI-as-judge 'may systematically inflate scores vs. human judgment,' and Section 9 states that real behavioral effects 'can ultimately only be understood through empirical studies involving human participants in context.' Yet the paper recommends scenario-based probes, expert panels, and LLM-as-judge as the practical pre-deployment assessment path. If these proxies do not actually track downstream user outcomes—changes in confidence, overreliance, decision quality, emotional state—then the proposed evaluation dimension is not yet measurable, and the call to make it 'a core consideration' lacks operational grounding. The paper explicitly labels this as a limitation, so it is not an internal inconsistency; it is a validity gap that future work must close. It is the most load-bearing concern because the framework's practical utility depends on it, not because the conceptual argument itself fails.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This conceptual paper argues that current AI evaluation frameworks are model-centric, focusing on output accuracy, robustness, safety, and policy compliance, while neglecting the interaction-level effects of human-facing AI on users' cognition, emotion, and behavior. The authors introduce 'psychological competence' as a new evaluation construct, define it, propose five domains (context sensitivity, emotional responsiveness, social cognition, behavioral influence, developmental sensitivity), connect these to a mechanism model of AI influence, and outline a mixed assessment strategy involving AI-as-judge pipelines, expert panels, and psychometric measures. The paper explicitly states that it is not proposing a benchmark, and it presents no empirical data; the contribution is intended as a conceptual foundation for future evaluation research and governance.","tokens_in":9079,"tokens_out":7539,"duration_ms":71109,"significance":"If the construct gains traction, this paper could provide a common vocabulary and conceptual scaffolding for evaluating human-facing AI beyond factual correctness, complementing technical and safety evaluations with a focus on user-level outcomes. The main strength is its careful scoping: the authors do not overclaim to have measured the construct, and they explicitly flag the central validity risk—that proxy assessments may not track real-world user outcomes—in Section 9 and Figure 3. This self-awareness makes the proposal credible. However, the operational value remains entirely unproven: there is no benchmark, no data, and no validation protocol. For a position paper this is acceptable, but it means the practical impact depends on substantial future empirical work. The paper should be read as a proposal for a research program rather than as an instantiated evaluation methodology.","major_comments":[{"comment":"The paper recommends scenario-based probes, expert panels, and LLM-as-judge pipelines as pre-deployment assessment tools for psychological competence. The validity of these proxies for downstream user outcomes is not established, and the authors concede this in Section 9 ('these effects can ultimately only be understood through empirical studies involving human participants in context') and in Figure 3 (LLM judges 'may systematically inflate scores vs. human judgment'). Because the paper's call to make psychological competence 'a core consideration' for providers and regulators rests on this assessment path, the reader is left with no criteria for when a proxy score supports a procurement or regulatory decision. I recommend adding a short 'validation agenda' to Section 8 that specifies minimal convergent evidence (e.g., agreement with human-rated outcomes, sensitivity to known group diff","section":"Section 8 / Figure 3"}],"minor_comments":[{"comment":"The statement that related notions of psychological competence 'haven't yet been formalized as an evaluation construct for AI systems' is too categorical given existing frameworks such as FAST (ref. [39]) and prior human-robot interaction metrics. Suggest softening to 'to our knowledge' and briefly discussing how the proposed construct differs from adjacent frameworks.","section":"Section 4"},{"comment":"The mockup conversation screenshots include timestamps and layout elements that are difficult to read. If this is a real figure, ensure the text is legible and the domain-assessment table is clearly aligned with System A and System B.","section":"Figure 1"},{"comment":"Reference formatting is inconsistent: some preprints use PsyArXiv DOIs (refs. [14], [24]), while others mix 'arXiv' and 'doi:10.48550/arXiv.x' (ref. [21]). Standardize to the journal's preferred style for preprints.","section":"References"},{"comment":"The 'Scale: High/Low/Medium' labels in Figure 3 are undefined. Clarify whether they refer to throughput, cost, reliability, or another property.","section":"Section 8 / Figure 3"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is a well-written position paper. The author list includes affiliations with for-profit AI consultancies and the Behavioral AI Institute, and several citations are to preprints by the same group (e.g., ref. [3]). This is disclosed in the conflict-of-interest statement; I see no issue, but the editor may wish to confirm that the novelty claim regarding a 'missing dimension' is not overstated relative to the broader human-AI interaction literature. Overall, the paper is suitable for publication after minor revisions addressing the validation agenda and presentation issues."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does what it says: it gives a name and a structure to an evaluation dimension that the field has been circling for a while. The main takeaway is that this is a solid conceptual proposal with no empirical validation, and the authors are upfront about that.\n\nWhat's actually new is the assembly: five domains—context sensitivity, emotional responsiveness, social cognition, behavioral influence, developmental sensitivity—mapped onto a five-stage mechanism pathway. That mapping is a genuinely useful way to organize existing findings on sycophancy, trust calibration, the labor illusion, and skill degradation. The paper also avoids overclaiming: it explicitly says there is no benchmark, that proxies need validation, and that real effects can only be understood with human studies. That honesty is a real strength.\n\nThe soft spot is the proxy-measurement program. The paper suggests AI-as-judge pipelines and expert panels as pre-deployment assessment, but Figure 3 concedes AI-as-judge may inflate scores, and Section 9 says real behavioral effects require human participants in context. The authors flag this as a limitation, so it's not an internal contradiction. It is, however, the load-bearing link: if those proxies don't track downstream user outcomes, the framework's practical value is unproven. The conceptual argument stands, but the operational path does not yet.\n\nThe novelty is moderate. Every component comes from cited literature; the umbrella construct is new. That's fine for a conceptual paper, but the contribution is synthesis rather than discovery. The citation pattern looks reasonable, and the paper doesn't define its target into existence—the circularity burden is low.\n\nThis paper is for anyone who builds or regulates human-facing AI systems, and for researchers who want a common vocabulary for interaction-level risks. It would be a decent reading-group piece, and it deserves a serious referee. My recommendation: send it to peer review. It is not a measurement paper yet, but it is a clear, honest framework that the field can build on.","headline":"A clear, honest conceptual framework for an interaction-level evaluation dimension; its proxy-assessment path is unproven but openly acknowledged.","tokens_in":9535,"tokens_out":1951,"would_cite":true,"duration_ms":17861,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Psychological competence — how an AI interaction shapes a user's reasoning, emotions, and decisions — is proposed as a missing dimension in AI evaluation.","keywords":["AI evaluation","human-AI interaction","psychological competence","behavioral science","trust calibration","user autonomy","AI safety","interaction quality"],"falsifier":"A randomized trial in which two AI systems produce equally accurate responses but differ on psychological competence ratings; if the higher-rated system fails to improve users' decision quality, emotional stability, or independent judgment, the construct's added value is not demonstrated. A simpler check: if LLM-based judges disagree systematically with human expert ratings on the five domains, the proposed pre-deployment assessment method needs rethinking.","tokens_in":8780,"feed_emoji":"🧠","tokens_out":4689,"duration_ms":43028,"temperature":0.7,"pith_summary":"The paper argues that current AI benchmarks measure what a model can do — accuracy, reasoning, robustness, policy compliance — but not what interacting with the system does to the user. It proposes psychological competence as a new evaluation dimension: the capacity of a human-facing AI to support accurate reasoning, emotional stability, and autonomous decision-making while avoiding distortions of judgment, harmful belief reinforcement, and over-reliance. The construct is organized into five domains — context sensitivity, emotional responsiveness, social cognition, behavioral influence, and developmental sensitivity — tied to a pathway from AI output to user interpretation, weighting, integration, and feedback. The authors argue that evaluating this dimension is essential for AI systems used as advisors, tutors, coaches, and companions, where responses can shape beliefs and choices.","feed_headline":"Psychological competence named as missing AI evaluation dimension","feed_subtitle":"Current benchmarks ask what a model can do; this framework asks what using it does to a user's thinking, emotions, and decisions.","key_machinery":"The central machinery is the construct of psychological competence itself, broken into five domains (context sensitivity, emotional responsiveness, social cognition, behavioral influence, developmental sensitivity) and tied to a stated mechanism pathway that runs from AI output to user interpretation, weighting through perceived authority and fluency, integration into reasoning and decisions, and feedback over repeated interactions. This mapping gives evaluators specific interaction properties — framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance — to probe as proxies for downstream effects on cognition, emotion, and behavior.","core_discovery":"The central claim is that the unit of evaluation for human-facing AI should shift from the model in isolation to the human-AI interaction. Psychological competence is the missing dimension: the ability of the system to generate interactions that appropriately reflect user context and support accurate reasoning, emotional stability, and autonomous decision-making, while avoiding distortion of judgment, reinforcement of harmful beliefs, or over-reliance. The paper formalizes the construct conceptually, identifies five domains, and outlines a mechanism pathway by which outputs produce psychological effects, while proposing scenario-based probes, human expert panels, and AI-as-judge pipelines as","pith_inferences":["The framework's practical value depends on proxy validity: the authors acknowledge that real effects need human-subject studies, so an obvious test is whether scenario-based competence scores predict user outcomes in controlled trials.","The mechanism pathway suggests a testable prediction: sustained interaction with systems rated high in psychological competence should reduce overreliance and belief offloading relative to output-equivalent controls over time.","The five domains could be operationalized into a benchmark with expert-rated gold examples, which would also expose whether AI-as-judge ratings track human expert judgments on autonomy and vulnerability sensitivity.","The construct may give regulators a language for tying behavioral harm to specific interaction properties, but only if autonomy and vulnerability sensitivity can be measured reliably across diverse contexts and user groups."],"forward_implications":["Evaluation suites for conversational AI would add interaction-level probes alongside accuracy and safety tests, asking not just whether a response is correct but whether it preserves user agency and calibrates trust.","Model providers could use the framework to design for agency preservation and calibrated trust during prompting, tuning, and interaction design, making behavioral interaction quality an explicit development target.","Deploying organizations in healthcare, education, and mental health could incorporate psychological competence into procurement and assurance decisions, complementing existing safety and alignment checks.","Regulators could draw on the framework when specifying behavioral safety expectations for high-impact systems, complementing risk-based governance frameworks.","Assessment would combine scalable AI-as-judge checks for pre-deployment screening with human expert panels for psychological impact and psychometric measures for longitudinal tracking."],"fun_headline_variants":["AI benchmarks miss how models affect users psychologically","Test AI by how it shapes user minds, not just answers","AI evaluation's missing dimension: psychological competence","Measure AI by its psychological impact on user decisions","Why AI benchmarks overlook user psychology"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The framework's value rests on the assumption that structured proxy assessments — scenario probes, expert ratings, and AI-as-judge pipelines — can validly estimate how real human-AI interactions affect users' reasoning, emotions, and behavior; if proxies do not track actual user outcomes, the construct remains unvalidated.","fun_headline_variants_meta":{"raw":{"variants":["AI benchmarks miss how models affect users psychologically","Test AI by how it shapes user minds, not just answers","AI evaluation's missing dimension: psychological competence","Measure AI by its psychological impact on user decisions","Why AI benchmarks overlook user psychology"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000962,"raw_usage":{"total_tokens":3945,"prompt_tokens":766,"completion_tokens":3179,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":3110}},"tokens_in":510,"tokens_out":3179,"duration_ms":21941,"temperature":1.0,"reasoning_tokens":3110,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T07:52:16.611383+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A randomized trial in which two AI systems produce equally accurate responses but differ on psychological competence ratings; if the higher-rated system fails to improve users' decision quality, emotional stability, or independent judgment, the construct's added value is not demonstrated. A simpler check: if LLM-based judges disagree systematically with human expert ratings on the five domains, the proposed pre-deployment assessment method needs rethinking.","supporting_citations":[],"review_version":2}