{"id":"b575b2a7-41d9-4687-b855-9fc59a79be3e","arxiv_id":"2508.07887","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Centaur predicts human choices accurately but fails to generate human-like behavior in three cognitive tasks, so it does not yet qualify as a participant simulator.","lead":"A team of cognitive scientists tested Centaur, a large language model trained on human behavior, to see if it can act as a simulated human participant. They found it predicts people's choices well but behaves unnaturally when left to choose on its own, so it is not yet a reliable 'behavioral AlphaFold'.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reversal-learning 'human data' benchmark is an unvalidated RW simulation; if that proxy fails, the flagship generative divergence claim loses its human grounding.","rationale":"The paper asks a legitimate and important question: can a model with strong predictive NLL still fail as a generative participant simulator? The horizon and WCST evaluations use real human data and provide genuine support for a negative answer. The open-loop, seeded sampling protocol is a reasonable first pass, and the paper does not overclaim mechanistic insight. The vulnerable point is the reversal task, which is singled out in the conclusion as a qualitative hallmark Centaur misses. There the reference distribution is synthetic RW data with fixed parameters that are not validated against human reversal behavior. The argument would be secure if those parameters were known to reproduce human reversal learning, but no such validation is reported. This is a gap in the support for the abstract's 'diverges from human data' framing, not an internal inconsistency. The proposed test would settle whether the RW proxy is adequate. Because the reader's verdict already conditioned acceptance on reframing or replacing the RW benchmark, my read does not change the verdict; it remains CONDITIONAL and, under the provided schema, is best represented as UNCHANGED.","tokens_in":8631,"tokens_out":5004,"duration_ms":62385,"concrete_test":"Fit the 3-parameter RW model (or the original 5-parameter version) to a human reversal-learning dataset, e.g., Eckstein et al. [10], by maximum likelihood, then perform a posterior predictive check: simulate from the fitted model and compare the reversal curve (e.g., probability of choosing the previously rewarded option around the reversal, trials-to-switch) to the human data. If the fitted RW simulation is not a close match to humans, replace the synthetic benchmark with actual human trajectories and recompute Fig. 1B. Alternatively, directly score Centaur's open-loop reversal trajectories against human data using the same metrics. If Centaur's generative reversal curve is within the human confidence interval, the headline 'diverges from human data' fails on this task.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's flagship generative failure is the reversal-learning task, but the 'human' target there is not human at all: it is a fixed-parameter RW simulation (alpha=0.5, beta=2.5, d=0.5) across 32 seeds. The SI states that the data were 'generated from the RW model,' yet the abstract and conclusion read the result as divergence 'from human data' and cite 'reversals' as a qualitative hallmark Centaur misses. Nothing in the paper fits this RW model to human reversal data or checks whether the simulated reversal curve (timing, slope, variability) actually matches human reversal behavior. If the RW proxy overstates or distorts human reversal dynamics, the reversal column of Fig. 1B does not measure what the abstract claims. The horizon and WCST comparisons do use human data, so the general point may survive, but the 'systematically' claim and the reversal-specific conclusion are only as strong as this proxy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates Centaur, an LLM fine-tuned on a large corpus of human behavioral experiments, as a candidate 'participant simulator' in the behavioral sciences. The authors distinguish predictive performance (trial-by-trial NLL given human choice histories) from generative performance (open-loop simulation of entire choice sequences), and argue that the latter is the critical criterion for a simulator. They compare Centaur-7B/80B against Llama-3.1 baselines and domain-specific cognitive models on three tasks: a reversal-learning bandit task, a horizon-dependent bandit task, and the Wisconsin Card Sorting Test. The central empirical claim is a dissociation: Centaur shows competitive or superior predictive accuracy, but its generative behavior deviates from human data, missing qualitative hallmarks such as reversal behavior and horizon effects. The paper concludes that Centaur, despite progress in prediction, is not yet a reliable participant simulator or an accurate model of cognition.","tokens_in":8904,"tokens_out":4929,"duration_ms":66511,"significance":"If the conclusion holds, this is an important and timely cautionary result for the use of LLM-based 'synthetic participants': it would demonstrate that high predictive fit to human choices does not automatically confer generative fidelity, and that standardized generative benchmarks are needed before such models are used for in silico experimental prototyping. The paper has clear strengths: the predictive/generative distinction is well motivated, the evaluation protocols (prompt construction, sampling policy, seed handling) are described transparently, the code and data are openly available, per-seed results are reported in Fig. S1, and two of the three tasks are evaluated against real human datasets. The main limitations—the reversal-learning benchmark uses unvalidated synthetic data, the absence of inferential statistics, and the unexamined dependence on a single sampling policy—mean that the strong version of the conclusion is not yet fully established.","major_comments":[{"comment":"The abstract and conclusion claim that Centaur's generative behavior 'systematically diverges from human data,' and the main text singles out 'reversals' as a qualitative hallmark Centaur misses. However, for the reversal-learning task the human benchmark is actually synthetic data generated from a three-parameter Rescorla-Wagner model with fixed parameters (alpha=0.5, beta=2.5, d=0.5), as stated in the SI. No evidence is provided that this RW model, or this parameter setting, reproduces the timing, slope, or variability of human reversal behavior. The figure caption is careful ('synthetic data generated from the RW model'), but the abstract and conclusion overstate the human grounding of this specific result. This is load-bearing because the reversal dissociation is the paper's flagship demonstration. Please either validate the RW proxy against human reversal data (or fit its parameters","section":"Abstract / 'Reversal Learning Task' in Supplementary Information"},{"comment":"The conclusion that Centaur 'fails to capture' the horizon effect or exhibits 'substantially more perseveration and set-loss errors' is based on visual inspection of group-mean curves and error bars, with no inferential statistics. For a claim about the absence or attenuation of an effect (e.g., no horizon effect in Fig. 1D), it is important to quantify evidence: report effect sizes and confidence intervals, and for null or near-null effects provide equivalence tests, Bayes factors, or model comparisons on choice data. The simulations used only 32 seeds for the reversal task and similar per-participant timelines elsewhere; a power analysis or per-seed distributions would clarify how stable the qualitative divergences are, especially given the large seed-to-seed variability shown in Fig. S1. This is a central issue because the paper's main claim is an empirical dissociation, not a purely","section":"Fig. 1D and 1F / 'Generative Performance Evaluation' in Supplementary Information"},{"comment":"The generative evaluation treats one specific policy—sampling from the softmax over the constrained answer set at temperature=1 in an open-loop self-feeding regime—as 'Centaur's generative behavior.' But a model's generative capability can be sensitive to decoding parameters (temperature, top-p, constrained token masking, prompt phrasing), and the paper does not test whether the reported divergence is robust across reasonable settings. If the conclusion is meant to be about Centaur as a simulator rather than about one particular sampling policy, the authors should either show robustness to these choices or explicitly restate the conclusion as applying to the temperature-1 constrained policy. This is not a request for exhaustive sweeps, but a minimal sensitivity analysis is needed to support the strong 'systematically diverges' wording.","section":"Supplementary Information, 'Generative Performance Evaluation' / 'Action Policy'"}],"minor_comments":[{"comment":"Typographical/naming inconsistencies: 'Centaur-70B' appears in the reversal generative section but the model is elsewhere described as Centaur-80B; 'Wisconsis' should be 'Wisconsin'; 'extend' should be 'extent'; 'asses' should be 'assess'; 'probaility' should be 'probability'.","section":"Supplementary Information, 'Reversal Learning Task'"},{"comment":"The predictive evaluation of the RW baseline on 'data it had generated itself' is not an independent baseline, since the model is essentially being evaluated on its own training distribution. This is acceptable for illustrating the predictive/generative contrast, but should be acknowledged or supplemented with a cross-validated evaluation on human data.","section":"Supplementary Information, 'Reversal Learning Task'"},{"comment":"The caption correctly indicates that panel A uses synthetic RW data. Please ensure the main text always consistently uses 'human data' only where human data were actually used, to avoid ambiguity.","section":"Fig. 1 caption"},{"comment":"The repetition model's generative behavior is described as 'a random bandit was selected at the start ... and the model repeatedly chose this same arm throughout the timeline with probability p.' It is unclear how the model switches to the other arm when it does not repeat; clarify the exact generative process and whether p is the fitted value or a free parameter varied across seeds.","section":"Supplementary Information, 'Reversal Learning Task'"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope and the overall direction is sound. The main revision pressure should be on the reversal-learning proxy and on adding inferential statistics; both are fixable within the manuscript's scope. No concerns about citation practice beyond the need to calibrate wording to the actual data sources."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth knowing about: it's the first open-loop generative evaluation of Centaur, and it documents a real predictive-generative dissociation across three cognitive tasks. The methods are transparent—prompts, temperature=1 sampling, per-seed results in the SI—and the horizon and WCST analyses use actual human data. The predictive-generative distinction is well motivated, and the repetition-model toy example in Fig. 1A/B is a nice pedagogical illustration of why prediction can mask generative failures. Credit where due: the negative conclusion is plausible and the paper is a fair, direct test of a claim that has been floating around the LLM-as-participant-simulator literature.\n\nThe soft spots are real but not fatal. The flagship reversal-learning 'human' benchmark is not human at all; it's a Rescorla-Wagner simulation with fixed parameters (alpha=0.5, beta=2.5, d=0.5) across 32 seeds. The abstract and conclusion read that result as divergence 'from human data,' which is overclaiming. If the RW model is a poor proxy for human reversal dynamics, the reversal-specific conclusion weakens. However, the horizon and WCST results rely on human data and still show generative divergence, so the central negative result survives even if the reversal column is discounted. Two other minor issues: no inferential statistics (just means and SEs) and only 32 seeds per model. These are more presentation than substance. The conclusion that Centaur 'cannot yet serve as a behavioral AlphaFold' is fine as a title-level claim, but the three-task, one-synthetic-proxy evidence doesn't support a strong systematic claim across all cognitive tasks.\n\nWho is this for? Anyone working on LLM evaluation for cognitive modeling, or on using LLMs as synthetic participants. It's a useful caution for the in-silico-prototyping crowd. It deserves a serious referee, though it would need a revision that either replaces the RW benchmark with human reversal data or explicitly reframes that specific claim as divergence from the domain-specific model rather than from humans. The SI is solid enough that I'd trust the horizon and WCST numbers. It's not a takedown; it's a fair empirical contribution with a overbroad interpretation on one task. I'd accept it for review, and I'd suggest the editor ask for a moderate revision rather than desk reject.","headline":"Useful first open-loop evaluation of Centaur's generative behavior, with a clear predictive-generative dissociation; the reversal-learning benchmark is synthetic RW data, not human, so the 'divergence from human data' framing overreaches on that task.","tokens_in":9406,"tokens_out":1089,"would_cite":false,"duration_ms":14728,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Centaur, a language model fine-tuned on 160 human experiments, predicts participants' next choices accurately yet systematically fails to generate human-like behavior when run open-loop, so it does not yet qualify as a reliable participant","keywords":["participant simulator","large language model","cognitive modeling","predictive vs generative evaluation","reversal learning","horizon effect","Wisconsin Card Sorting Test","foundation model of behavior"],"falsifier":"Run Centaur open-loop on the reversal task for 100 fresh seeds and measure the distribution of the trial on which it first switches to the newly rewarded bandit after trial 50. If that switch-latency distribution matches human switch latencies from the underlying reversal-learning data, and if the horizon task reproduces the human horizon effect, the paper's conclusion is wrong.","tokens_in":8542,"feed_emoji":"🧠","tokens_out":6071,"duration_ms":69349,"temperature":0.7,"pith_summary":"Centaur is a large language model fine-tuned on human behavioral data from 160 experiments, proposed as a model of cognition and as a synthetic participant for prototyping studies. This paper argues that the standard test used to validate such models—trial-by-trial prediction of the next human choice—cannot certify that the model can generate human-like behavior from scratch. The authors evaluate Centaur in two modes: predictive, where the model sees the human's actual history, and generative, where it sees only its own history. Across a reversal learning task, a horizon-dependent bandit task, and the Wisconsin Card Sorting Test, Centaur's generative behavior systematically misses the qualitative hallmarks these tasks were designed to measure, even when its predictive performance is strong. The paper concludes that predictive accuracy alone does not make Centaur a reliable participant simulator or an accurate cognitive model, and that generative evaluation must be part of the standard.","feed_headline":"Centaur predicts human choices but fails the generative test","feed_subtitle":"A simulated participant must generate behavior from scratch; Centaur misses the reversals and horizon effects humans show.","key_machinery":"The predictive-versus-generative distinction, operationalized by evaluating models in two modes. In predictive mode, the model receives the true human choice history and is scored by the negative log-likelihood of the next choice. In generative mode, the model receives only its own chosen actions and their feedback, with responses sampled from its softmax at temperature 1. The repetition model—a baseline that repeats its previous choice with fixed probability—is what makes the distinction bite: it predicts human choices reasonably well but, run generatively, repeats the same choice forever and misses the reversal entirely, proving that predictive fit does not certify generative fidelity.","core_discovery":"The paper's core discovery is a dissociation: Centaur can achieve strong trial-by-trial predictive performance on tasks from its training set while its open-loop, generative behavior misses qualitative human hallmarks. On the reversal learning task, Centaur shows weak and seed-dependent reversal dynamics, sometimes never switching after the reward reversal; on the horizon task, it does not reproduce the human horizon effect; on the Wisconsin Card Sorting Test, outside its fine-tuning set, it is outperformed by a task-specific model on both predictive and generative measures. Thus, predictive accuracy on its own does not qualify a model as a participant simulator or a cognitive model.","pith_inferences":["Extension: if the predictive/generative dissociation generalizes, current next-token prediction scores on cognitive tasks may overstate how aligned LLMs are with humans; adding an open-loop generation version of each benchmark would test this directly.","Extension: the authors' three-task evaluation is a transferable template. Applying the same open-loop protocol to other foundation models trained on human behavior would show whether the failure is specific to Centaur's training objective or endemic to next-token fine-tuning more broadly.","Extension: Centaur's horizon-effect failure suggests a concrete hypothesis—fine-tuning on human choice histories may teach a model to appear human when anchored to human history, but not to carry the internal exploration-exploitation state needed to generate such behavior. A testable extension is prompting Centaur to verbalize its uncertainty before each choice and checking whether that restores t","Extension: because the reversal benchmark uses synthetic trajectories from an RW model as the human reference, a direct replication with real human reversal-learning data would clarify how much of the reported failure is inherent to Centaur and how much is an artifact of the benchmark."],"forward_implications":["A model can score well on next-choice prediction while failing the generative test that matters for simulation.","Centaur's failure to reproduce the horizon effect and the reversal learning effect indicates it has not learned the decision processes those tasks are designed to expose, despite training on similar task families.","On a task outside its fine-tuning set, Centaur is outperformed by a small task-specific model, suggesting that whatever human-like behavior it does produce does not transfer broadly.","Participant simulators need standardized open-loop generative benchmarks, not just predictive likelihood, before they can support in silico prototyping of experiments."],"supporting_citations":[{"why":"Introduces Centaur, its 160-experiment training set, and the predictive NLL evaluation procedure that the paper reuses and then scrutinizes.","marker":"[4]"},{"why":"Provides the repetition-model demonstration that high predictive accuracy does not imply generative fidelity, the paper's central conceptual lever.","marker":"[8]"},{"why":"Supplies the reversal learning task design and the RW model whose synthetic data are used to benchmark predictive and generative reversal performance.","marker":"[10]"},{"why":"Supplies the horizon task design and the human dataset used for predictive evaluation and for the generative comparison of the horizon effect.","marker":"[11]"},{"why":"Supplies the human Wisconsin Card Sorting Test dataset and task variant used for the out-of-training-set evaluation.","marker":"[12]"},{"why":"Provides the sequential learning model used as the domain-specific baseline on the Wisconsin Card Sorting Test.","marker":"[9]"}],"fun_headline_variants":["Centaur predicts but fails generative test","Prediction alone doesn't make a participant simulator","Centaur: strong predictions, weak generative behavior","Why Centaur isn't AlphaFold for the mind yet","Centaur's predictive edge masks generative gaps"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The reversal-learning benchmark treats trajectories generated by a three-parameter RW model ($\\alpha=0.5$, $\\beta=2.5$, $d=0.5$) as the human reference for generative fidelity; if those synthetic trajectories are not a faithful proxy for human reversal dynamics, that pillar of the conclusion weakens.","fun_headline_variants_meta":{"raw":{"variants":["Centaur predicts but fails generative test","Prediction alone doesn't make a participant simulator","Centaur: strong predictions, weak generative behavior","Why Centaur isn't AlphaFold for the mind yet","Centaur's predictive edge masks generative gaps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000165,"raw_usage":{"total_tokens":1082,"prompt_tokens":733,"completion_tokens":349,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":477,"completion_tokens_details":{"reasoning_tokens":279}},"tokens_in":477,"tokens_out":349,"duration_ms":4465,"temperature":1.0,"reasoning_tokens":279,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T21:47:22.538873+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run Centaur open-loop on the reversal task for 100 fresh seeds and measure the distribution of the trial on which it first switches to the newly rewarded bandit after trial 50. If that switch-latency distribution matches human switch latencies from the underlying reversal-learning data, and if the horizon task reproduces the human horizon effect, the paper's conclusion is wrong.","supporting_citations":[{"cited_title":"Trends in cognitive sciences21(6), 425–433 (2017)","cited_arxiv_id":null,"evidence_quote":"Provides the repetition-model demonstration that high predictive accuracy does not imply generative fidelity, the paper's central conceptual lever."},{"cited_title":"Developmental Cognitive Neuroscience55, 101106 (2022)","cited_arxiv_id":null,"evidence_quote":"Supplies the reversal learning task design and the RW model whose synthetic data are used to benchmark predictive and generative reversal performance."},{"cited_title":"Journal 12 of experimental psychology: General143(6), 2074 (2014)","cited_arxiv_id":null,"evidence_quote":"Supplies the horizon task design and the human dataset used for predictive evaluation and for the generative comparison of the horizon effect."},{"cited_title":"Scientific reports10(1), 15464 (2020)","cited_arxiv_id":null,"evidence_quote":"Supplies the human Wisconsin Card Sorting Test dataset and task variant used for the out-of-training-set evaluation."},{"cited_title":"Journal of mathematical psychology54(1), 5–13 (2010)","cited_arxiv_id":null,"evidence_quote":"Provides the sequential learning model used as the domain-specific baseline on the Wisconsin Card Sorting Test."}],"review_version":1}