{"id":"3c9b543c-6de8-4015-b5bc-336b0872568c","arxiv_id":"2506.16697","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"This Perspective paper proposes that LLM research in psychology must combine psychometric validity and causal inference standards, mapping evidence requirements to the type of claim being made.","lead":"This paper argues that psychology studies using large language models must validate measurements like human psychological tests and protect causal claims like experiments, or risk reporting artifacts as real traits. It proposes a dual-validity framework that scales how much evidence is needed based on how ambitious the claim about an AI is.","discovery_kind":"unification","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The categorical 'measurement phantom' conclusion is undercut by the paper's own adapted-instrument results; the central claim needs scope-limiting to unadapted human measures.","rationale":"The reader's weakest-assumption diagnosis points to the ontological stipulation that psychological constructs require embodiment, temporal continuity, and a persistent self. That is a real concern, and the paper does not defend it. However, I see a more internally grounded and empirically checkable problem: the paper's own cited evidence includes successful LLM-adapted measures that exhibit theoretically coherent structure and predictive validity. Those successes do not settle the ontology question, but they do undermine the categorical version of the central claim. If adapted instruments work, then the 'measurement phantom' diagnosis cannot be applied to LLM outputs as such; it can only be applied to the specific practice of importing human scales without adaptation. This is not merely a philosophical caveat; it changes the empirical claim from 'LLM psychological research systematically fails' to 'a particular methodological practice systematically fails.' The framework's practical recommendations—reliability first, adapted instruments, causal safeguards—remain largely intact under the narrower claim, so the reader's CONDITIONAL verdict is appropriate. I would keep the verdict unchanged, but the condition should include a scope-limitation sentence in the abstract and conclusions, plus an explicit acknowledgment that adapted-instrument successes are consistent with the framework and not exceptions needing to be explained away.","tokens_in":19082,"tokens_out":5786,"duration_ms":79586,"concrete_test":"Perform a systematic comparison of all validity and reliability coefficients reported in the cited studies (e.g., Oh & Demberg 2025; Peereboom et al. 2025; Ye et al. 2025; Lee et al. 2024; Ma et al. 2025), coding each study as either 'direct human-instrument application' or 'LLM-adapted instrument.' Then test whether the systematic-failure claim survives: if most adapted instruments meet conventional psychometric thresholds (e.g., reliability ≥ 0.70, adequate structural fit, convergent/predictive correlations meaningfully above 0), while failures cluster in direct human-instrument applications, the central claim must be revised to 'unadapted human instruments systematically fail,' and the framework should be reframed around instrument adaptation rather than categorical phantom-construct dismissal.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest conclusion—that LLM outputs are 'patterns of words masquerading as a psychological phenomenon'—depends on the inference that psychometric failures indicate absence of the construct. But the paper itself reports multiple cases where LLM-adapted instruments recover exactly the structure that human instruments fail to find: Ye et al.'s generative psychometrics reproduced the Schwartz value circumplex; Lee et al.'s scenario-based TRAIT produced theoretically coherent inter-trait correlations; Ma et al.'s implicit Core Sentiment Inventory achieved predictive correlations above 0.85. The paper even concedes that 'structural validity failures may thus indicate methodological mismatch rather than construct absence.' This internal tension is more load-bearing than the ontological assumption about embodiment, because it affects the empirical scope of the central claim, not just its philosophical framing. If LLM-adapted measures can be reliable, structurally coherent, and predictively valid, then LLM responses are not inherently 'statistical pattern matching masquerading as psychology'; rather, the phantom diagnosis applies to the uncritical use of human-developed instruments, not to LLM psychological measurement as such. The paper's framework remains useful, but the categorical version of the central claim—'current practice systematically fails' and 'any output is a pattern of words'—is overbroad unless explicitly restricted to unadapted human instruments.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is a Perspective proposing a dual-validity framework for the use of large language models in psychological research. It argues that reliable measurement and sound causal inference must be integrated when LLM outputs are used to measure, characterize, simulate, or model psychological constructs, and that the required evidence should scale with the scientific ambition of the claim. The paper reviews three reliability threats (training artifact contamination, prompt hypersensitivity, and stochastic degradation), five sources of psychometric construct validity evidence, and four types of causal-inference validity, illustrating each with recent empirical work. It concludes that current practice systematically fails these requirements and that much current LLM psychological research produces 'measurement phantoms' or 'patterns of words masquerading as a psychological phenomenon.'","tokens_in":19250,"tokens_out":5292,"duration_ms":58178,"significance":"The manuscript is a timely and broadly useful synthesis. Its strengths are the extensive and current literature integration; the concrete catalog of reliability threats; the careful separation of four application categories (research tool, characterization, human simulation, cognitive model); and the empirically grounded acknowledgment that some LLM-adapted instruments, such as Ye et al.'s generative psychometrics, Lee et al.'s scenario-based TRAIT, and Ma et al.'s implicit sentiment measure, can achieve structural or predictive validity. The paper does not present new data, code, or machine-checked proofs, but its framework offers testable expectations about which measurement approaches are likely to succeed and which are likely to fail. If the scope of the central claim is appropriately calibrated, the paper could provide a useful methodological reference for researchers and reviewers.","major_comments":[{"comment":"The manuscript's categorical conclusion that LLM psychological research 'systematically fails' and that when the measured attribute does not exist 'any resulting output is a pattern of words masquerading as a psychological phenomenon' is undercut by the manuscript's own evidence. In the Internal Structure section, Ye et al.'s generative psychometrics reproduced the Schwartz value circumplex and Lee et al.'s scenario-based TRAIT produced theoretically coherent inter-trait correlations; in Relations with Other Variables, Ma et al.'s implicit Core Sentiment Inventory achieved predictive correlations above 0.85. The paper even concedes that 'Structural validity failures may thus indicate methodological mismatch rather than construct absence.' Because these successes involve adapted or LLM-specific instruments, the categorical phantom diagnosis is overbroad. The conclusion should be explicitly restricted to unadapted human measures, or the paper should explain why these successes do not count as evidence relevant to psychological constructs.","section":"Internal Structure; Relations with Other Variables; Conclusions and Future Directions"},{"comment":"The claim that 'anxiety presupposes temporal experience, a persistent self, and embodied consequences—ontological properties the model lacks' is asserted rather than defended. This ontological premise is load-bearing: it is the basis for saying the attribute does not exist and hence that LLM outputs are 'patterns of words masquerading as a psychological phenomenon.' A functionalist account of psychological attributes, in which internal states are defined by causal roles rather than by embodiment, would block this inference. The manuscript should either provide an argument against such functionalism or present the conclusion conditionally, for example by stating that under a constitution-based, non-functionalist account the attribute does not exist in LLMs.","section":"Construct Validity from the Psychometric Foundation"},{"comment":"The paper's central organizing claim is that 'validity requirements scale with psychological ambition,' but the framework does not specify how this scaling works. The text assigns requirements to four application categories—research tools, characterization, human simulation, cognitive modeling—but the basis for these assignments is not given; for example, why human simulation requires only behavioral correspondence while characterization requires construct validation is stated but not argued. Table 1 lists validity types and threats but provides no decision rule connecting a claim's ambition to the evidence required. Without such a rule, the framework is a useful checklist rather than a framework for determining validation demands. The authors should either provide explicit criteria or clearly frame the contribution as a checklist.","section":"Why LLM Research Requires Both; Table 1"}],"minor_comments":[{"comment":"The abstract promises that the same model output requires different validation strategies depending on whether researchers claim to measure, characterize, simulate, or model a construct, but no section provides a worked illustration of this point; adding a brief example would make the framework concrete.","section":"Abstract"},{"comment":"The reference 'Stanley, J. C., & Campbell, D. T. (1963)' is conventionally cited as 'Campbell, D. T., & Stanley, J. C.'; please correct or justify the ordering.","section":"References"},{"comment":"Table 1 labels the psychometric and causal-inference 'Construct' rows identically; renaming them, for example 'Construct validity (psychometric)' and 'Construct validity (causal inference),' would prevent confusion between construct validity in measurement and construct validity of causal claims.","section":"Table 1"},{"comment":"The arXiv identifier for Guan et al. (2025) duplicates the identifier for Sclar et al. (2023), 2310.11324; the correct identifier for Guan et al. appears to be 2502.04134.","section":"References"},{"comment":"The headers 'LLM VALIDITY 1' and similar appear to be page headers from the submission and should be removed before publication.","section":"Formatting"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's self-citations (Lin 2023, 2025a, 2025b) are to methodology guides and are not load-bearing for the central claim; they do not affect the verdict. The scope of the paper fits a perspective or opinion venue. The main concern is that the categorical 'systematic failure' conclusion needs to be reconciled with the manuscript's own counterexamples of successful LLM-adapted instruments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. This is a genuinely useful synthesis: it maps the five psychometric validity sources onto LLM research and connects them to the four causal inference validity types, which nobody has done as cleanly. And the headline conclusion—that current practice systematically fails and that outputs are \"patterns of words masquerading as psychological phenomena\"—is overbroad given the paper's own evidence.\n\nThe framework itself is the strength. Table 1 is practical. The paper correctly separates four uses of LLMs (research tool, behavioral characterization, simulation, cognitive modeling) and argues that evidence requirements scale with ambition. That scaling idea is the real contribution. The paper also credits successes: Ye and colleagues' generative psychometrics reproducing the Schwartz value circumplex, Lee and colleagues' scenario-based TRAIT showing coherent inter-trait correlations, Ma and colleagues' implicit sentiment scores predicting generated text above 0.85. It even concedes near the end that structural validity failures may indicate methodological mismatch rather than construct absence. So the framework is not the problem.\n\nThe problem is the categorical diagnosis. If adapted instruments routinely recover theoretical structure, then the inference from psychometric failure to construct absence does not hold universally. The \"measurement phantom\" conclusion should be scoped to unvalidated, unadapted human instruments. That is a load-bearing qualification, not a stylistic one, because it changes the empirical claim from \"LLMs cannot have psychological attributes\" to \"human instruments are being applied carelessly.\" The paper's ontological premise—anxiety presupposes temporal experience, a persistent self, and embodied consequences—is asserted rather than defended. A functionalist would reject it, and the paper does not engage that literature. The broad claim that \"current practice systematically fails\" is also based on a convenience sample of studies; a systematic review would make it credible.\n\nWho is this for? Researchers and graduate students doing LLM psychology, and reviewers who need a validity checklist. It will be cited as a framework paper. It deserves a serious referee: a good reviewer can push the author to narrow the phantom claim and defend the ontology. I would send it to peer review with that expectation, not desk reject it. The framework survives the needed revisions.\n\nRecommendation: engage with it. The synthesis is worth having even if the current version oversells its conclusion.","headline":"A useful dual-validity synthesis for LLM psychology, but the central 'measurement phantom' claim is broader than the paper's own adapted-instrument evidence supports.","tokens_in":19786,"tokens_out":1739,"would_cite":true,"duration_ms":19566,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Psychological claims about large language models fail without dual validation—psychometric and causal—this paper argues.","keywords":["large language models","psychometrics","construct validity","causal inference","reliability","measurement phantoms","AI psychology","psychological measurement"],"falsifier":"Concretely: administer a personality or moral-decision inventory to a single model across many semantically equivalent prompt variants (changed option labels, order, punctuation, phrasing) at fixed temperature and version; if scores stay stable, factorially coherent, and predictive of external outcomes across all variants—rather than shifting more than 70% as the paper reports—the reliability crisis is refuted for that model. A second refutation would be a model whose 'anxiety' responses are lawfully modulated by threat manipulations, remain stable across sessions, and predict downstream outputs within a nomological network, which would challenge the claim that no such attribute exists.","tokens_in":18830,"feed_emoji":"🧠","tokens_out":7690,"duration_ms":73408,"temperature":0.7,"pith_summary":"This Perspective argues that using human psychological instruments on large language models can produce 'measurement phantoms'—statistical artifacts mistaken for genuine psychological phenomena—because current practice skips the two validation traditions psychology built for human research: psychometric validity (does the instrument measure the construct?) and causal inference validity (does the design support the conclusion?). The paper tries to establish that evidence requirements must scale with scientific ambition: classifying text asks for accuracy checks, while claiming a model 'is anxious' or 'has theory of mind' asks for much more. It proposes a dual-validity framework that maps four uses of LLMs in psychology (research tool, behavioral characterization, human simulator, cognitive model) onto the validity evidence each requires. If right, much of the current literature on machine personality, moral reasoning, and theory of mind is built on unreliable measures and would need revalidation before its claims can stand.","feed_headline":"LLM psychology claims need psychometric and causal checks","feed_subtitle":"One framework says reliability and construct validity must come before any talk of machine anxiety.","key_machinery":"The framework's load-bearing object is the pairing of the psychometric and causal-inference validity traditions into a single pipeline. From psychometrics it takes the reliability ceiling—no measure can be more valid than it is reliable—and construct validity as an accumulating evidence argument, sharpened by the causal theory of validity, which requires both that the attribute exists and that variations in it causally produce observed scores. From experimental methodology it takes the four parallel threats to causal inference (internal, external, construct, and statistical conclusion validity). The framework maps the four uses of LLMs in psychology onto these standards and classifies failure modes: training artifact contamination, prompt hypersensitivity, and stochastic degradation violate psychometric assumptions, while temperature confounds, version drift, dynamic scenario reconstruction, and non-independence violate causal assumptions. The unifying move is the observation that in LLM research the same output serves simultaneously as a measurement indicator and as experimental data, so a failure on either side corrupts the other.","core_discovery":"The paper's central claim is that LLM responses can look psychologically meaningful while being generated by statistical pattern matching, and that the field currently has no way to tell the difference because it applies human measurement tools without their validation scaffolding. The same output—a model endorsing 'I am anxious'—requires entirely different validation strategies depending on whether the claim is to measure anxiety, characterize model behavior, simulate human responses, or model a cognitive mechanism. A measurement claim demands that the attribute exist in the model and causally produce the response; the paper argues that psychological constructs such as anxiety presuppose temporal experience, a persistent self, and embodied consequences that LLMs lack, so without such evidence 'any resulting output is a pattern of words masquerading as a psychological phenomenon.' The paper proposes that reliability and validity be established before causal experimentation, that construct validation draw on five sources of evidence (content, response processes, internal structure, relations with other variables, consequences), and that causal claims address four parallel validity types (internal, external, construct, statistical conclusion). Its constructive proposal is to study computational analogues—'anxiety-analogous patterns' rather than anxiety—so AI psychology proceeds on mechanistic, not biological, terms.","pith_inferences":["A natural next step, left implicit in the paper, is a public audit instrument that scores LLM psychological studies on reliability and validity evidence; the paper calls for infrastructure but does not specify one.","The ontological boundary is testable: if future models with persistent memory and embodiment-like training show stable, lawfully connected anxiety-analogous responses, the 'attribute does not exist' premise blurs and the framework would need to decide when an analogue becomes a construct.","The framework could also be applied retroactively as a taxonomy to meta-analyze existing LLM findings, sorting which reported effects survive prompt perturbation—an extension of the reliability discussion the paper leaves implicit."],"forward_implications":["Publications claiming to measure personality, theory of mind, or moral reasoning in LLMs would need to document reliability (test–retest, parallel forms, internal consistency) and evidence from all five construct-validity sources before their claims could be credited.","Studies that manipulate prompts to test causal hypotheses would need to rule out computational confounds—temperature, prompt formatting, model version, non-independence—through factorial designs, ablations, or unblinding procedures.","Simulation claims (LLM responses stand in for human responses) would be restricted to demonstrated behavioral correspondence and would not license claims about human-like underlying mechanisms.","Many reported LLM psychological effects are predicted to be unstable: they should shift or vanish under trivial prompt variations, and re-analysis with factorial designs should shrink or eliminate them."],"supporting_citations":[{"why":"Supplies the causal theory of validity used to demand that the measured attribute exist and causally produce scores before LLM outputs can be interpreted as psychological measurements.","marker":"Borsboom et al. (2004)"},{"why":"Foundational source for construct validity as an accumulating evidence argument, which the framework applies to LLM measures.","marker":"Cronbach & Meehl (1955)"},{"why":"Source of the four types of causal-inference validity that the paper integrates with psychometrics.","marker":"Cook & Campbell (1979)"},{"why":"Empirical demonstration that trivial label changes reverse LLM moral preferences, motivating the reliability-crisis diagnosis.","marker":"Oh & Demberg (2025)"},{"why":"Evidence that LLM responses show arbitrary latent structures and 'cognitive phantoms,' underpinning the construct-validity critique.","marker":"Peereboom et al. (2025)"},{"why":"Documents models endorsing contradictory personality items, used as a core example of reliability failure.","marker":"Sühr et al. (2023)"},{"why":"Shows LLMs reconstruct whole scenarios from price changes, illustrating the dynamic-context confound threatening internal validity.","marker":"Gui & Toubia (2023)"},{"why":"Argues that LLM outputs reflect training-data associations rather than stable traits, supporting the mechanistic-substitution claim.","marker":"Gao et al. (2024)"},{"why":"Shows ChatGPT reproduces average attitudes but with reduced variance and divergent regression coefficients, used to question the construct validity of causal claims.","marker":"Bisbee et al. (2024)"}],"fun_headline_variants":["LLM psychology claims often are measurement phantoms","Dual-validity test separates real LLM cognition from pattern noise","LLM 'anxiety' may be just word patterns without validation","Measurement phantoms: why LLM psychology needs dual-validity","For AI psychology, validation must precede simulation claims"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's strongest conclusion rests on the assumption that psychological constructs such as anxiety presuppose embodiment, temporal continuity, and a persistent self, so that a language model lacking these cannot possess the attribute being measured; if a functionalist account of mental states—where internal states are defined by causal roles rather than physical substrate—is correct, that premise fails.","fun_headline_variants_meta":{"raw":{"variants":["LLM psychology claims often are measurement phantoms","Dual-validity test separates real LLM cognition from pattern noise","LLM 'anxiety' may be just word patterns without validation","Measurement phantoms: why LLM psychology needs dual-validity","For AI psychology, validation must precede simulation claims"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000436,"raw_usage":{"total_tokens":2243,"prompt_tokens":995,"completion_tokens":1248,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":611,"completion_tokens_details":{"reasoning_tokens":1162}},"tokens_in":611,"tokens_out":1248,"duration_ms":10779,"temperature":1.0,"reasoning_tokens":1162,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:37:36.267688+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Concretely: administer a personality or moral-decision inventory to a single model across many semantically equivalent prompt variants (changed option labels, order, punctuation, phrasing) at fixed temperature and version; if scores stay stable, factorially coherent, and predictive of external outcomes across all variants—rather than shifting more than 70% as the paper reports—the reliability crisis is refuted for that model. A second refutation would be a model whose 'anxiety' responses are lawfully modulated by threat manipulations, remain stable across sessions, and predict downstream outputs within a nomological network, which would challenge the claim that no such attribute exists.","supporting_citations":[],"review_version":1}