{"id":"bc4dc0ff-808d-4a0b-a618-877735dc0826","arxiv_id":"2511.04500","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Llama-3.1-8B with a multi-step reasoning-and-filter prompt reproduces human cooperation rates across 121 dyadic games (MSD=0.031, r=0.89), outperforming Nash-equilibrium predictions (MSD=0.096, r=0.78).","lead":"This paper ran three open-source language models through 121 two-player cooperation games originally played by humans, and found that Llama produced human-like cooperation patterns while Qwen behaved more like a rational Nash-equilibrium player. The authors then used Llama to simulate 320 new game variants and preregistered human experiments to test those predictions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline replication is partly fitted: the multi-step prompt and logical verifier were selected to maximize cooperation in the same S≥T region that drives the human–Llama similarity in Table 2.","rationale":"The reader's weakest assumption—that the heavily adapted extraction pipeline may manufacture rather than measure the result—is correct, and my read narrows it to a concrete selection effect: the multi-step prompt and the logical verifier were both tuned to maximize cooperation in the S≥T region, which is the same region where human cooperation is high and where the Llama matrix visually matches humans. This is documented in §3.4.3 and §3.4.4 and is a classic selection-on-the-target problem. It does not require assuming bad faith; it is an internal validity issue. The concern is load-bearing because Table 2 is the quantitative basis for the paper's central claim: if the pipeline was selected using the target pattern, the reported MSD/r are partially in-sample fits rather than unbiased replication metrics. I keep the reader's CONDITIONAL verdict rather than escalating to REJECT because the preregistered extension provides a genuine out-of-sample test, the authors share code/data, and the holdout check I propose could settle the concern. Other issues—missing attention-based analysis, absent uncertainty quantification, and verifier bypass up to 25%—are secondary to this selection problem; they should be fixed but do not by themselves change the verdict.","tokens_in":20160,"tokens_out":7030,"duration_ms":79121,"concrete_test":"Freeze the full pipeline before consulting the human cooperation matrix: preselect the multi-step prompt and logical verifier using only a validation criterion that does not involve the human data (e.g., a theoretical consistency objective or a development set restricted to games outside the S≥T region). Then recompute MSD and r against human data and Nash on all 121 original games. If the preselected variant no longer yields Llama closer to humans than Nash, or r drops materially below 0.89, the Table 2 headline is a tuning artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in §1.2 and Table 2—Llama reproduces human cooperation (MSD=0.031, r=0.89) and does so more accurately than Nash (MSD=0.096, r=0.78)—requires that the extraction pipeline be a neutral measurement of Llama's behavior. The methods show it is not. In §3.4.3, the multi-step prompt was selected by 'evaluated performance in the S≥T region (Harmony Games)' and choosing the version with the highest cooperation rates for Llama and Mistral. In §3.4.4, the logical verifier was selected using the same criterion. S≥T is precisely the region where human cooperation is high and where the final Llama matrix shows its most distinctive feature (Fig. 1). Thus a central pattern used to compute the Table 2 metrics was used as a development set for pipeline hyperparameters. The near-random result under Simple Extraction (Fig. 1) shows that the measured human-likeness is not a stable default property of Llama but is produced by the tuned layers. Without a selection rule fixed before inspecting the target human matrix, or a robustness analysis over reasonable prompt variants, Table 2 cannot be read as an out-of-sample replication.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper claims that an 8B open-weight LLM (Llama-3.1-8B), when equipped with a specific multi-step reasoning prompt and a logical verifier, replicates aggregate human cooperation patterns in 121 dyadic 2×2 games better than the Nash equilibrium does (Table 2: MSD=0.031, r=0.89 vs. MSD=0.096, r=0.78). The authors compare three open models, characterize behavioral phenotypes, and use Llama to extrapolate to novel games, with preregistered hypotheses for future human experiments. The central claim is that the extraction pipeline is a neutral measurement of LLM behavior, revealing an emergent human-like pattern rather than constructing it.","tokens_in":20377,"tokens_out":6028,"duration_ms":59500,"significance":"If established with a confirmatory analysis, this result would be a valuable contribution: it suggests that an open, reproducible LLM-based pipeline can serve as a 'digital twin' for aggregate human behavior in simple strategic interactions, enabling low-cost in-silico hypothesis generation. The paper's strengths include the use of open-source models, public code and data, preregistration of the extended experiments, and transparent reporting of limitations (e.g., verifier bypass). However, as written, the central quantitative claim is compromised by the fact that the prompt and verifier were explicitly selected to maximize cooperation in the S≥T region, which is exactly the region where human cooperation is high and where Llama's matrix shows its most distinctive pattern. Thus the headline fidelity metrics in Table 2 are in-sample fits, not out-of-sample replication.","major_comments":[{"comment":"The multi-step prompt and logical verifier were selected to maximize cooperation in the S≥T region (Harmony Games), as stated in §3.4.3 and §3.4.4. This is precisely the region where human cooperation is highest (Fig. 1, far left) and where the final Llama matrix shows its most distinctive pattern. The near-random results under Simple Extraction (Fig. 1) show that the measured high fidelity is not a default property of Llama but is produced by the tuned layers. Consequently, the MSD=0.031, r=0.89 values in Table 2 measure agreement between a pipeline tuned to the target human matrix and the human matrix itself; they cannot be read as out-of-sample replication. The authors should pre-register the selection rule before inspecting human data, conduct a robustness analysis over a family of prompt/verifier variants, or hold out a subset of games for validation.","section":"§3.4.3, §3.4.4, Table 2"},{"comment":"No confidence intervals or significance tests are reported for MSD or Pearson's r. Each cooperation-rate cell is the mean of 20 stochastic samples at temperature 0.8; the standard error of a proportion near 0.5 is approximately 0.11. The human matrix is also a finite sample from ~500 participants, yet it is treated as a fixed benchmark. Bootstrap or other resampling procedures are needed to assess whether the reported differences (e.g., Llama vs. Nash: MSD 0.031 vs. 0.096, r 0.89 vs. 0.78) are statistically robust or could be within sampling noise. Without this, the quantitative ranking is not fully supported.","section":"Table 2"},{"comment":"The original human experiment paid participants with lottery tickets according to points scored, while the LLM prompt replaces this with a direct monetary conversion ('10 euros per point' with a worked example). This changes the decision problem from a lottery (involving risk preferences) to a certain monetary payoff. Cooperative and defective choices can be sensitive to risk attitudes, so the LLM may be solving a different task from the human participants. The paper should justify this deviation or, ideally, rerun the simulation with a lottery-based or otherwise matched incentive structure to show the qualitative pattern is robust to this change.","section":"§3.3"},{"comment":"The 'Nash equilibrium' matrix is computed using replicator dynamics with a specific equilibrium selection: in the bistable Stag Hunt region, the initial condition x(0)=0.5 determines the boundary between the basins of attraction. Nash equilibria are not unique in coordination games; other standard selections (payoff-dominant, risk-dominant, or mixed) would yield different cooperation rates. The claim that Llama is closer to human data than Nash is therefore benchmark-dependent. The authors should either justify this equilibrium selection as the appropriate theoretical benchmark or report results under alternative equilibrium concepts.","section":"§3.6"}],"minor_comments":[{"comment":"The color scale descriptions are inconsistent: Fig. 1 says '0 (purple: no cooperation) to 1 (yellow: full cooperation)', while Fig. A1 says '0 (yellow: no cooperation) to 1 (purple: full cooperation)'. Please use a single convention and ensure the captions match the actual colormaps.","section":"Fig. 1 and Fig. A1 captions"},{"comment":"The extraction accuracy rates (0.97, 0.96, 0.83) are presented without confidence intervals or the number of samples per model beyond the statement of 100 long answers total. Reporting the per-model sample sizes and a measure of uncertainty would strengthen the comparison.","section":"Table 3"},{"comment":"Bypass rates reach 0.25 for some games, meaning a quarter of responses in those cells did not pass the logical verifier. The paper states the impact 'remains limited', but a sensitivity analysis excluding or flagging bypassed responses would quantitatively support that assertion and is directly relevant to the noise in Table 2.","section":"§A.4"},{"comment":"The reference to 'non-transparent regions of Figure 2' is somewhat confusing because the figure uses transparency only in panel A. Clarify the wording or ensure the figure's legend makes the delineation explicit.","section":"§1.2"}],"recommendation":"major_revision","confidential_remarks":"The central problem is the target-dependent selection of the prompt and verifier, which is openly described in the methods. This is a case of 'researcher degrees of freedom' that undermines the headline replication claim. The good news is that the paper provides complete code and data, so it is feasible to require a confirmatory analysis (e.g., pre-registered variants or hold-out games) as a condition of acceptance. If the authors cannot provide such evidence, the paper's central claim should be substantially weakened. I recommend major revision rather than rejection because the scientific question is valuable and the empirical materials are available to fix the issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know: the paper's headline result—Llama reproducing human cooperation with r=0.89—is not a neutral measurement. The multi-step prompt and the logical verifier were explicitly selected in §3.4.3 and §3.4.4 to maximize cooperation in the S≥T Harmony region, which is exactly where human cooperation is highest. The Simple Extraction baseline is near random. So the reported similarity is at least partly constructed by the pipeline. That does not make the work worthless, but it means Table 2 should be read as \"Llama under a pipeline tuned on a human-high-cooperation region resembles humans,\" not as a clean out-of-sample replication.\n\nWhat is genuinely new: a systematic replication of 121 dyadic games across three open-weight models, with code/data, Nash-equilibrium comparison, and a preregistered extension to 441 game configurations. That goes beyond previous single-game tests. The honest documentation of the prompt iteration and verifier bypass is also credit where due—they don't hide the choices.\n\nSoft spots, in rough order of severity. First, the circularity above. It is fixable: report results under a pre-fixed prompt, or do a sensitivity analysis across plausible prompt variants; disclose the tuning as calibration rather than neutral elicitation. Second, no uncertainty quantification anywhere. Each cell is 20 stochastic samples, and Table 2 has no confidence intervals or significance tests. Bootstrap CIs would take an hour. Third, the abstract promises attention-based mechanistic analysis and behavioral phenotyping, but neither appears in the manuscript. That mismatch should be resolved—either add the analyses or trim the abstract. Fourth, the verifier bypass in up to 25% of responses in some games deserves a robustness check; the current statement that \"the overall impact remains limited\" is unsupported.\n\nThe Nash equilibrium calculation and the fixed-point stability analysis look correct. The extended-grid predictions are genuinely new and could be valuable if the pipeline concern is resolved.\n\nVerdict: conditional, not reject. The paper deserves peer review time—the topic is important, the data and code are available, and the flaws are repairable. But it needs substantial revision before it supports the abstract's claims. If you work on in-silico behavioral experiments or LLM-as-agent evaluation, this is worth engaging, but with a critical eye.","headline":"Tuned-on-the-target pipeline weakens the headline replication, but the systematic grid and preregistered extension merit a serious referee.","tokens_in":20926,"tokens_out":3706,"would_cite":false,"duration_ms":34623,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A10","91A22","91A90"],"pacs":[],"model":"deepseek-v4-flash","headline":"Llama-3.1-8B reproduces human cooperation patterns in 121 dyadic games more accurately than the Nash equilibrium does, indicating that LLMs can serve as digital twins for behavioral experiments.","keywords":["large language models","game theory","cooperation","Nash equilibrium","behavioral phenotypes","digital twins","prompt engineering","experimental economics"],"falsifier":"Run the preregistered human experiment on the extended S/T grid: if humans do not show the S≥T high-cooperation diagonal and the T=R boundary that Llama predicts, the extrapolation claim is refuted. Alternatively, turn off the logical verifier and recompute the MSD of Llama against human data; if the match degrades sharply, the fidelity is produced by the filter rather than by the model's underlying decision process.","tokens_in":19982,"feed_emoji":"🤖","tokens_out":6240,"duration_ms":55651,"temperature":0.7,"pith_summary":"The paper tries to establish that a carefully prompted open-weight language model can reproduce aggregate human cooperation in classic two-player games, and can do so more faithfully than the standard rational-choice benchmark. Using 121 games spanning four classical game types, the authors show that Llama's average cooperation rates closely track human data (MSD 0.031, r=0.89), while the Nash equilibrium deviates more (MSD 0.096, r=0.78). A different model, Qwen, aligns with Nash rather than with humans, showing that model choice changes the behavioral phenotype. The same setup is then extended to 320 novel game configurations to generate preregistered predictions for future human experiments. If the claim holds, behavioral scientists could explore experimental parameter spaces in silico before committing to human studies.","feed_headline":"Llama-3.1 matches human cooperation better than Nash equilibrium","feed_subtitle":"Llama's choices track human cooperation at r=0.89, beating Nash equilibrium predictions.","key_machinery":"The load-bearing mechanism is the prompt-and-extraction pipeline: instructions use neutral vocabulary (A/B choices, 'other player'), state a one-shot simultaneous game with direct euro-per-point payment, and ask the model to reason step by step; responses are then classified as good or bad by a second LLM acting as a logical verifier before the chosen label is extracted. Progressively adding these layers transforms Llama's nearly random one-word answers into a structured cooperation matrix. The comparisons use mean squared displacement and Pearson correlation between the model, human, and Nash cooperation matrices, with Nash equilibria computed via replicator dynamics for each game.","core_discovery":"The central claim is that Llama-3.1-8B, prompted in neutral language without game-theory jargon and filtered through a multi-step reasoning-and-verification pipeline, yields cooperation rates across 121 (S,T) payoff combinations that replicate empirical human patterns with high accuracy. The key signature is high cooperation where the sucker payoff S equals or exceeds the temptation payoff T, mirroring the 'envious' decision rule dominant in human data, and low cooperation where T exceeds the reward R. Llama reproduces this pattern better than Nash equilibrium does, which the paper presents as evidence that LLMs can act as digital twins: first replicating existing experiments, then generatin","pith_inferences":["The reported fidelity is conditional on the verifier: up to 25% of responses in some games bypassed it, and without it Llama's matrix degrades, so an ablation removing the verifier would show how much of the human-like pattern is a measurement effect rather than model behavior.","If the preregistered human experiments confirm Llama's novel-game predictions, LLM-based simulation would be a credible low-cost instrument for hypothesis generation; if they fail, the method is better read as a tool that recovers known experiments, not one that predicts new ones.","The envious profile shared by humans and Llama may reflect a payoff-comparison heuristic absorbed from text rather than genuine social preference; a scale-invariance test—varying whether the prize is framed as euros or lottery tickets—could distinguish those accounts.","The random-like cooperation near (S,T)=(0,0) could be a salience artifact of near-zero payoffs; a targeted simulation with small nonzero payoffs would test whether this boundary is robust or an edge case of numeric framing."],"forward_implications":["Llama reproduces population-level human cooperation without persona-based prompting, so building synthetic participant pools can be simpler than earlier approaches.","Different models embody different decision phenotypes: Qwen behaves like a rational (Nash) player, which makes some LLMs suitable human simulators and others not.","The paper's predicted cooperation rates for novel parameter combinations are preregistered, giving a direct way to test whether LLM or Nash extrapolation better anticipates real human behavior.","Rational-choice predictions are less accurate than this LLM at capturing human deviations, so LLM-based simulation may complement formal game theory as a descriptive model."],"fun_headline_variants":["Llama matches human cooperation better than Nash equilibrium","Llama beats Nash equilibrium at predicting human cooperation","Llama mirrors human cooperation patterns in game theory","Llama replicates human cooperation across 121 games","Llama aligns with human cooperation, not Nash equilibrium"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claim rests on the assumption that the adapted instructions and the Qwen verifier measure the same decision problem humans faced, rather than changing the game (for example, making point payoffs more salient than the original lottery mechanism) and thereby manufacturing the apparent match.","fun_headline_variants_meta":{"raw":{"variants":["Llama matches human cooperation better than Nash equilibrium","Llama beats Nash equilibrium at predicting human cooperation","Llama mirrors human cooperation patterns in game theory","Llama replicates human cooperation across 121 games","Llama aligns with human cooperation, not Nash equilibrium"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000891,"raw_usage":{"total_tokens":3718,"prompt_tokens":821,"completion_tokens":2897,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":565,"completion_tokens_details":{"reasoning_tokens":2825}},"tokens_in":565,"tokens_out":2897,"duration_ms":18884,"temperature":1.0,"reasoning_tokens":2825,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T23:38:19.700858+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the preregistered human experiment on the extended S/T grid: if humans do not show the S≥T high-cooperation diagonal and the T=R boundary that Llama predicts, the extrapolation claim is refuted. Alternatively, turn off the logical verifier and recompute the MSD of Llama against human data; if the match degrades sharply, the fidelity is produced by the filter rather than by the model's underlying decision process.","supporting_citations":[],"review_version":1}