{"id":"bca8e629-8533-4c1e-8690-90deaf6060d5","arxiv_id":"2508.06950","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLMs fail to mirror human moral judgments when scenarios are reworded to change meaning, even the human-fine-tuned CENTAUR model.","lead":"The paper tests whether four large language models, including the psychology-focused CENTAUR model, adjust their moral ratings when scenarios are reworded with small but meaningful changes. Humans shift their ratings substantially; the LLMs barely move, suggesting they track wording rather than meaning.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Human–LLM comparison confounds between-subjects human variance with within-model LLM stability; re-analysis needed to support the 'does not simulate' claim.","rationale":"The reader's weakest assumption already identifies 'between-subjects noise' as a possible explanation for the correlation drop, so our concern overlaps. However, we do not think the more general item-selection worry is as load-bearing: a single well-constructed counterexample can refute a universal claim like CENTAUR's, so hand-picking items is not fatal. The most precise and testable issue is the human/LLM data asymmetry: human item means are based on independent participant groups and carry sampling error, while LLM item means are based on repeated draws from the same model and are nearly deterministic. This asymmetry directly undermines the comparison of correlations and mean shifts that drives the empirical argument. The paper's theoretical and qualitative points remain plausible, and the direction of the effect is probably real, but the quantitative support is not yet clean. A within-subject human design or a mixed-effects re-analysis would settle whether the divergence is an artifact. Since the reader already downgraded to CONDITIONAL, our read does not change the verdict.","tokens_in":16254,"tokens_out":10083,"duration_ms":112217,"concrete_test":"Using the OSF data (https://osf.io/qbev7), fit a mixed-effects model to the human ratings with fixed effects for wording version (original vs reworded) and random intercepts for participant and item, and likewise model LLM repeated draws; estimate the human vs LLM difference in wording sensitivity. Alternatively, collect a within-subject human sample (N≈200) rating both versions in counterbalanced order and recompute the human original–reworded correlation and mean absolute difference. If the human–LLM gap in sensitivity remains significant after accounting for between-subjects noise, the concern is resolved; if it shrinks to non-significance, the central claim needs revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest conditional for the central claim is the asymmetry between human and LLM data in §4.2–4.4. Human participants were randomly assigned to rate either original or reworded items (between-subjects), so the human original–reworded correlation r=.54 in Table 3 is a correlation of two independent group means, attenuated by sampling error in each item mean. LLM ratings are means of 10 draws from the same model, with near-zero standard errors, so the same-model original–reworded correlations (r=.80–.99) are not attenuated. The mean absolute human shift of 2.20 (SD 1.08) is likewise an independent-groups difference; it contains between-group variance that a within-subject measure would remove. The paper therefore conflates 'humans respond to meaning' with 'human group means are noisy.' Without a human test–retest baseline or within-subject condition, or a mixed-effects model that explicitly models participant and item variability, the claimed human–LLM divergence could be partly an artifact of this measurement asymmetry. This is load-bearing because the drop from r≈.97 (original) to r≈.54 (reworded) is the paper's key quantitative evidence for semantic insensitivity.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that large language models (LLMs) cannot simulate human psychology, targeting the recent CENTAUR claim that an LLM can 'predict and simulate human behaviour in any experiment expressible in natural language.' The authors offer a conceptual argument—LLMs are trained on token sequences, so generalization should be expected along token similarity rather than semantic meaning—and support it with an empirical study of moral judgments. Using 30 vignettes from Dillion et al. (2023), each paired with a minimally reworded version that changes meaning, they collected human ratings (N = 374, between-subjects assignment to original vs. reworded) and queried four LLMs (GPT-3.5-Turbo, GPT-4o-mini, Llama-3.1 70B, CENTAUR) 10 times per item. They replicate the high human–LLM correlations on original items (r ≈ .97–.99), but find lower correlations on reworded items (r ≈ .51–.64). LLMs' own original–reworded correlations are high (r = .80–.99), while the human original–reworded correlation is r = .54; the mean absolute shift is 2.20 for humans versus 0.42–1.25 for LLMs. Chow tests for Llama and CENTAUR reject a pooled regression. The authors conclude that LLMs should not replace human participants and must be validated per application.","tokens_in":16543,"tokens_out":7769,"duration_ms":81825,"significance":"If the result holds, it is an important counterexample to the strong claim that LLMs can simulate human behavior in any natural-language experiment, and it provides a useful caution for psychological researchers. The paper is transparent: the data are posted on OSF, the design is simple, and four models including CENTAUR are tested. The conceptual generalization argument is independent of the empirical results and gives a principled reason to expect failures on novel items. The main weakness is the statistical asymmetry between noisy between-subjects human item means and near-noiseless within-model LLM means, which may inflate the observed human–LLM divergence. The broad conclusion is therefore conditional on a reanalysis that models participant-level variability.","major_comments":[{"comment":"The key human–LLM contrast rests on an asymmetric measurement setup. Human raters were randomly assigned to either original or reworded items, so the human original–reworded correlation (r = .54 in Table 3) is the correlation of two independent group means. Each item mean carries sampling error SD/sqrt(n_condition), which attenuates r. LLM ratings, by contrast, are means of 10 draws from the same model (Table 1 shows near-zero SDs), so the LLM original–reworded correlations (r = .80–.99) are not attenuated in the same way. The mean absolute shift (humans 2.20, LLMs 0.42–1.25) is likewise inflated by between-group noise in the human difference. This is load-bearing because the paper's central claim is that humans track meaning while LLMs track tokens. Please reanalyze with a within-subject human test–retest condition, or at least fit a mixed-effects model to individual human responses wit","section":"§4.2, §4.4 (Tables 1 and 3)"},{"comment":"The 30 scenarios are hand-picked, and the rewordings were authored by the researchers (e.g., 'elderly neighbor' -> 'elderly mosquito'; 'wife' -> 'earth'). The paper does not report a sampling rule, preregistration, or independent checks that the rewordings preserve surface similarity while changing meaning only. These items are used to support the unconditional conclusion that 'LLMs do not simulate human psychology.' That inference is too strong for a convenience sample of 30 items. If the study is meant as an existence proof against the universal CENTAUR claim, the paper should say so; if it aims at a rate or tendency, a larger or random item sample is needed. At minimum, clarify the selection criterion and restrict the closing claims accordingly.","section":"§4.2, Table 1 (item selection)"},{"comment":"The Chow test is described as comparing a pooled regression with separate regressions, but the text does not specify the unit of analysis (item means vs. individual ratings) or whether the regressions are weighted by the precision of each item mean. Since human item means are far noisier than LLM means, an unweighted fit on 30 points can reject pooling because of heteroscedasticity rather than because humans and LLMs have genuinely different response functions. This is especially relevant because the correlation differences for Llama and CENTAUR are not significant after Bonferroni correction (Table 3, p = .277 and .119). Please report the regression specification, use weighted least squares or multilevel modeling, and show that the Chow result survives when human noise is modeled explicitly.","section":"§4.3–§4.4 (Chow tests)"}],"minor_comments":[{"comment":"'Empiric evidence' should be 'empirical evidence'; model names are inconsistent (GPT-4, GPT4, GPT-4o-mini). Please standardize.","section":"Throughout"},{"comment":"Report the exact model versions, query dates, temperature/sampling parameters, and the few-shot prompts. Table 1 lists means/SDs only for GPT-4o-mini; provide the other models' data in a supplement or table.","section":"§4.2"},{"comment":"The compact row/column labels make the correlation matrices hard to read. Add a legend defining variable numbers and note that the matrices are symmetric.","section":"Tables 2–3"},{"comment":"Both subplots contain the same points; draw a single legend or label each panel clearly so the reader can identify human, Llama, and CENTAUR points in both panels.","section":"Figure 2"},{"comment":"Some citation forms are inconsistent (e.g., [Van Rooij et al., 2024] vs. van Rooij and Guest, 2025); check the reference style throughout.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a potentially important negative result, but the current analysis depends critically on comparing noisy between-subjects human means with near-noiseless within-model LLM means. I would be willing to see a revision that includes a within-subject human baseline or a mixed-effects reanalysis, and that positions the empirical study as a set of counterexamples rather than a general rate estimate. The conceptual argument is strong and the paper is clearly written; the statistical asymmetry is the main obstacle to publication as is."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: The paper gives us a concrete, well-documented demonstration that a psychology-tuned LLM (CENTAUR) and several general LLMs largely ignore semantic rewordings of moral vignettes that change human judgments. That's worth having. But the title promises more than the evidence delivers, and the comparison between human and LLM correlations is not apples-to-apples. I'd send it to peer review with eyes open about those issues.\n\nWhat's actually new: the 30 reworded moral scenarios from the Dillion et al. set, tested against humans and four LLMs including CENTAUR. The regression-equality tests (Chow) are a sensible way to show divergence. The paper is transparent: OSF data, clear methods, and the authors engage fairly with the pro-LLM literature. The theoretical argument about token-level rather than meaning-level generalization is a useful framing, though not novel on its own. The fact that CENTAUR behaves about the same as its base Llama model on these items is a specific and relevant finding.\n\nSoft spots: The title is the biggest one. \"Do not simulate human psychology\" is a universal claim; the evidence is one moral-judgment task with 30 hand-selected reworded items. That's an illustration, not a disproof. The items were chosen because they flip meaning, so the result that humans shift more is partly by construction. More important, the human data are between-subjects while the LLM data are within-model repeated sampling. The human original-reworded correlation (r=.54) is attenuated by sampling error in the two independent group means, while the LLM correlations are near-noise-free. So part of the gap could be measurement architecture. I don't think it's all artifact—the mean absolute shift of 2.20 vs 0.42–1.25 is large, and the item-level plots show systematic patterns—but the paper should report a human test-retest baseline or a mixed-effects model with participant/item variance. The \"elderly mosquito\" type items are odd and may inflate human shifts due to confusion, though the instruction to rate anyway helps.\n\nBottom line: This is a conditional accept, not a reject. The empirical core is believable, the data are open, and the direction of the argument is right. I'd want the authors to narrow the title, add a human reliability check, and explicitly discuss the between-subjects vs within-model issue. It's a useful paper for psychologists who are tempted to replace participants with LLMs, and it deserves serious refereeing.","headline":"Useful empirical counterexample to LLM-as-participant claims, but the title overreaches and the human/LLM comparison has a measurement asymmetry that should be addressed.","tokens_in":17049,"tokens_out":3599,"would_cite":true,"duration_ms":38127,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLMs do not simulate human psychology: reworded moral scenarios break the otherwise close match between model and human ratings.","keywords":["LLM simulation","human participants","moral judgments","semantic generalization","token similarity","CENTAUR","reworded vignettes","psychology research"],"falsifier":"Construct a new set of reworded moral or theory-of-mind vignettes with near-identical token overlap and show that a current LLM shifts its ratings by roughly two scale points on average, comparable to human shifts; or, conversely, show that random rewordings that do not change meaning produce the same correlation drop in humans, indicating the effect is not specific to semantic change.","tokens_in":16177,"feed_emoji":"⚖️","tokens_out":3132,"duration_ms":31395,"temperature":0.7,"pith_summary":"The paper sets out to refute the claim that large language models can stand in for human participants in psychology experiments. It argues that LLMs generalize by textual token similarity rather than by meaning, so when a scenario's wording is slightly changed to alter its meaning, models keep giving the same moral rating while humans change theirs. The authors demonstrate this with 30 moral vignettes rated by four LLMs and 374 human raters. Correlations between LLM and human ratings drop sharply after rewording, and separate regression lines fit significantly better than a single line. The conclusion is that LLMs are useful but fundamentally unreliable tools that must be validated against human responses for each new application.","feed_headline":"Reworded moral scenarios expose LLM psychology blind spot","feed_subtitle":"When wording changes but meaning does too, LLM ratings stay flat while human ratings move, so simulators fail on novel experiments.","key_machinery":"The central object is the paired set of 30 reworded moral vignettes: near-identical token sequences with deliberately changed meaning, such as 'cut the beard off ... to shame him' versus 'to shave him'. The argument is carried by comparing human-versus-LLM rating correlations between original and reworded items, and by Chow's test comparing a pooled regression against group-specific regressions. The reworded items isolate the question of whether models generalize by token similarity or by semantic meaning.","core_discovery":"The paper's central claim is that LLMs do not react to semantic wording changes the same way humans do, and therefore do not simulate human psychology. Taking 30 moral scenarios from prior research, the authors create reworded versions that change meaning with minimal token changes, sometimes as little as one letter. Humans show a mean absolute rating shift of 2.20 between original and reworded items, whereas GPT-3.5-Turbo shifts 0.75, GPT-4o-mini 0.42, Llama-3.1 70b 1.18, and CENTAUR 1.25. The human-model correlation for original items is high (r = .97 to .99), but for reworded items it falls to r = .51 to .61. Chow tests on Li and CENTAUR show that separate regressions for humans and model","pith_inferences":["The argument implies a general robustness test for LLM-based simulation: any alignment result should be re-checked on reworded items, since near-training-data items will overstate performance.","Because the human data are between-subjects and the LLM data are within-model, part of the correlation drop may reflect human between-subject noise; a within-subject human study would likely sharpen but not erase the effect.","CENTAUR's failure to improve over Llama-3.1 suggests that fine-tuning on millions of human responses does not confer meaning-sensitive generalization, hinting that scaling this approach may not fix the token-similarity failure.","A concrete next test: adversarially generate rewordings that keep tokens nearly identical but flip moral valence; current models would be expected to rate original and reworded versions nearly the same, while humans diverge."],"forward_implications":["LLMs should not be used as stand-in participants for novel psychology experiments, only as tools validated against human data.","The widely cited r = .95 agreement between GPT-3.5 and humans on moral scenarios does not extend to even slightly reworded stimuli.","CENTAUR's claim to predict and simulate human behavior in any natural-language experiment is contradicted by these results.","Researchers using LLMs should vary prompts, record model versions and settings, compare multiple models, and validate outputs on small human-rated datasets."],"supporting_citations":[{"why":"Provides the moral scenario set and the r = .95 human-GPT baseline that the paper replicates and then challenges with rewordings.","marker":"[Dillion et al., 2023]"},{"why":"Releases CENTAUR, the model whose 'predict and simulate human behaviour in any experiment expressible in natural language' claim is the paper's target.","marker":"[Binz et al., 2025]"},{"why":"Supplies several of the original moral vignettes that the authors reworded for their comparison.","marker":"[Clifford et al., 2015]"},{"why":"Provides the statistical test used to compare pooled versus group-specific regressions between humans and LLMs.","marker":"[Chow, 1960]"},{"why":"Evidence that language models treat antonyms and negations as similar, supporting the paper's token-similarity theory.","marker":"[Truong et al., 2023]"},{"why":"Evidence that LLMs ignore subtle variations in theory-of-mind vignettes, a prior demonstration of the same semantic-insensitivity failure.","marker":"[Hu et al., 2025]"}],"fun_headline_variants":["LLMs fail to track meaning shifts in moral scenarios","Reworded moral items break LLM-human agreement","LLMs don't simulate psychology: wording shifts reveal it","Human-LLM ratings diverge when moral wording changes meaning","Tiny wording tweaks expose LLM psychology simulation failure"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The 30 reworded vignettes are assumed to be a fair, representative sample of novel scenarios that a simulator must handle, and the correlation drop is interpreted as semantic insensitivity rather than an artifact of comparing different human groups or of odd items like 'elderly mosquito'.","fun_headline_variants_meta":{"raw":{"variants":["LLMs fail to track meaning shifts in moral scenarios","Reworded moral items break LLM-human agreement","LLMs don't simulate psychology: wording shifts reveal it","Human-LLM ratings diverge when moral wording changes meaning","Tiny wording tweaks expose LLM psychology simulation failure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000458,"raw_usage":{"total_tokens":2123,"prompt_tokens":727,"completion_tokens":1396,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":471,"completion_tokens_details":{"reasoning_tokens":1330}},"tokens_in":471,"tokens_out":1396,"duration_ms":10701,"temperature":1.0,"reasoning_tokens":1330,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T22:24:39.612693+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a new set of reworded moral or theory-of-mind vignettes with near-identical token overlap and show that a current LLM shifts its ratings by roughly two scale points on average, comparable to human shifts; or, conversely, show that random rewordings that do not change meaning produce the same correlation drop in humans, indicating the effect is not specific to semantic change.","supporting_citations":[],"review_version":1}