{"id":"d9032154-5443-432a-b283-d66e8849b997","arxiv_id":"2509.04847","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A study claiming language models match or beat top classical strategies in repeated prisoner's dilemma and adapt to opponent switches, but the write-up omits model names, data, and a results section, and its adaptability claims contradict its own table.","lead":"Researchers pitted language models against 240 classic strategies in the iterated prisoner's dilemma and report they perform as well as or better than the best human-designed strategies, while adapting quickly when opponents switch tactics mid-game. The paper is too incomplete to support these claims: it does not say which models were used, shows no results section, and contradicts itself on whether humans or AI adapt faster.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central empirical claims are unverifiable: no Results section, Table 1 contradicts Section 4.3, and the LLM/prompt are unspecified, so the abstract's performance and adaptability claims lack support.","rationale":"I agree with the reader that the paper is not a verifiable empirical study and that rejection is appropriate. The single most load-bearing problem is not just the unspecified model/prompt, though that is important; it is that the empirical evidence for the headline claims is absent or self-contradictory. The abstract asserts performance 'on par with, and in some cases exceeding, the best-known classical strategies' and 'rivaling or surpassing human adaptability,' but the paper's Experiments section contains no result tables or numerical summaries for RQ1 or RQ2—only figures with captions—and the only quantitative table (Table 1) directly contradicts the accompanying text in Section 4.3. Moreover, the checklist cites a Limitations section (Section 7) that does not exist and claims no human subjects despite the described 10-participant study. These are not style issues; they undermine the central claim's support. Even if the missing results were supplied, the unspecified LLM and truncated prompt would still make the generalization 'language models' unjustified, as Section 3 itself admits prompt choice heavily influences behavior. Therefore the verdict of REJECT is appropriate, and no change to the reader's verdict is needed.","tokens_in":12167,"tokens_out":5992,"duration_ms":52477,"concrete_test":"Recompute Table 1 from the raw human and AI session logs, specifically the single-switch condition: if AI adaptation speed (3.7) and cooperation rate (66.8%) are correct, then Section 4.3's assertion that humans maintain higher long-term cooperation is false; the authors must reconcile this contradiction before any conclusion about human vs AI adaptability can be drawn.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that LLMs perform on par with or better than the best classical strategies and adapt as fast as or faster than humans—is unsupported by the manuscript as submitted. Section 4 describes experiments but no Results section follows; the paper jumps from experimental descriptions to the Conclusion. The only quantitative result, Table 1, contradicts Section 4.3: for the single-switch condition, AI adaptation speed is 3.7±0.6 vs humans 5.4±1.1 and AI cooperation rate is 66.8±4.4 vs humans 62.3±4.8, yet the text states that 'humans adapted more slowly ... but maintained higher long-term cooperation rates after the switch.' Additionally, the LLM is never identified and the prompt is truncated, so even the reported behaviors cannot be reproduced. Section 3 admits the prompt 'heavily influences' behavior, yet the paper generalizes to 'language models.' The checklist's claims of a Limitations section (Section 7) and of no human subjects are false. Any one of these issues prevents the central claims from being accepted as stated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript investigates the behavior of large language models (LLMs) in the iterated prisoner's dilemma (IPD). The authors pit an unspecified LLM against a suite of 240 classical strategies in an Axelrod-style tournament, compute behavioral metrics such as niceness, forgiveness, and retaliation, and run strategy-switch experiments to measure adaptation speed. A human-subject experiment (N=10) compares human and AI adaptation. The abstract claims that LLMs perform on par with or better than the best classical strategies, exhibit strong cooperative traits, and detect and respond to opponent changes within a few rounds, rivaling or surpassing human adaptability. The paper, however, contains no Results section: the narrative jumps from the experiment descriptions in Section 4 to the Conclusion in Section 5. The only quantitative data table (Table 1) appears in Section 4.3, and it is not clearly tied to the advertised claims. The LLM is never identified, the prompt is truncated, and the checklist contains false or incomplete statements. These issues prevent the central claims from being verified in the submitted form.","tokens_in":12425,"tokens_out":4438,"duration_ms":43149,"significance":"If fully supported, the paper would provide a useful systematic characterization of LLM cooperative and competitive behavior in long-horizon IPD settings, with implications for human-AI interaction and alignment. The choice of established Axelrod strategies, behavioral axes (niceness, forgiveness, retaliation), and the inclusion of human comparisons are appropriate and relevant. However, the significance is presently only conditional: the central performance and adaptability claims are not backed by reported data, the experimental protocol is not reproducible as written, and the internal contradictions further undermine confidence. The paper has not yet achieved the evidentiary standard needed for a research contribution of this scope.","major_comments":[{"comment":"There is no Results section. After Section 4 (Experiments), the paper jumps to Section 5 (Conclusion). Figures 1-4 appear inside the experiment subsections, but no numerical results, statistical tests, or analysis link them to the abstract's claims. The introduction's 'average advantage of 12.6 wins per round' is not derived or referenced to any table or figure. The core empirical claims are therefore unsupported as reported.","section":"§4–§5"},{"comment":"Table 1 reports AI vs Single-switch cooperation rate 66.8±4.4 and human 62.3±4.8, i.e., AI cooperation is higher. Yet Section 4.3 states that humans 'maintained higher long-term cooperation rates after the switch.' This is a direct contradiction. The sentence that humans 'adapted more slowly than the top-performing AI model' is also unclear, since adaptation speed is lower-is-better and Table 1 gives AI 3.7±0.6 vs human 5.4±1.1; the textual claim about cooperation must be corrected and statistically supported.","section":"§4.3 / Table 1"},{"comment":"The LLM is never identified by name or version, and the prompt is truncated after the first two sentences. The paper itself states that 'the instructions in the prompt chosen heavily influence the behavior of the LLM-based player.' Yet the abstract generalizes to 'language models.' The undisclosed pilot selection process and the lack of full prompt details mean the reported behaviors cannot be reproduced or attributed to LLMs generally rather than to one specific model/prompt combination.","section":"§3"},{"comment":"Figures 1-4 do not identify which LLM produced the data, include no error bars, and are not accompanied by significance tests. For example, the claim that LLMs 'detect and respond to shifts within only a few rounds' is not quantified with a distribution, confidence interval, or comparison against a null model. The figures also use legends referencing specific classical strategies but do not state which AI model is being plotted. Without this information, the adaptation-speed and performance claims are not verifiable.","section":"§4.1–4.2 / Figures 1–4"},{"comment":"The NeurIPS checklist contains false or incomplete statements. Item 2 claims a Limitations section (Section 7), but no Section 7 exists in the manuscript. Items 14 and 15 answer 'NA' for human subjects, despite Section 4.3 describing recruitment of 10 human participants. Item 5 contains '[TODO]' answers. These issues make the submission incomplete and are inconsistent with the paper's actual content.","section":"Checklist"},{"comment":"The introduction and the abstract contradict each other on the central adaptability claim. Section 1 states that LLM-based players 'were able to recognize the switch and adapt slower than human players, with 24.5 rounds early change in strategy compared to humans,' whereas the abstract claims LLMs rival or surpass human adaptability. Table 1 indicates faster AI adaptation (3.7 vs 5.4 rounds). These statements cannot all be correct and must be reconciled.","section":"§1 vs Abstract"}],"minor_comments":[{"comment":"The manuscript contains numerous typographical and grammatical errors, e.g., 'the demonstrated through a tournament', 'F orgiven defection', and 'Aspects of the strategies change their actions ... associated Specific properties.' A thorough copyedit is needed.","section":"Throughout"},{"comment":"The legend includes names such as 'First by Grofman' but the axes and figure caption do not specify which AI model is being evaluated or how cooperation rate is aggregated across seeds. Add clear axis labels and model identification.","section":"Figure 2"},{"comment":"Reference [5] is cited as the source of the 240 classical strategies, but the listed paper is about GPT in game theory experiments, not the Axelrod strategy library. Verify and cite the actual strategy repository.","section":"References"},{"comment":"The human experiment reports only N=10 and lacks the full participant instructions, compensation details, and IRB information. The checklist's 'NA' for human subjects is inconsistent and should be corrected.","section":"§4.3"}],"recommendation":"reject","confidential_remarks":"This submission appears to be in an incomplete state: the central experiments are described but not reported, the quantitative table contradicts the narrative, the LLM/prompt are unspecified, and the checklist is unreliable. These are load-bearing issues that prevent acceptance as is. If the authors substantially revise by adding a full Results section, identifying the models and prompts, correcting the contradictions, and providing reproducible data and statistical tests, a resubmission could be worth considering."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Candid take: this is an incomplete draft, not a finished study. The headline claims—LLMs perform on par with or better than best-known classical strategies, and adapt as fast as or faster than humans—are unsupported as written. There is no Results section; the paper jumps from experiment descriptions to the Conclusion. The only quantitative table, Table 1, actually contradicts Section 4.3's text: the table shows AI with both faster adaptation (3.7 vs 5.4 rounds) and higher post-switch cooperation (66.8% vs 62.3%) than humans, yet the text says humans \"maintained higher long-term cooperation rates.\" The abstract says \"rivaling or surpassing human adaptability,\" the intro says LLMs adapt slower than humans, and Table 1 says AI is faster. That is not a minor inconsistency; it breaks the core claim.\n\nThe paper also omits the basics needed for reproduction: no language model is named, and the prompt is truncated mid-sentence (\"If both players cooperate, ...\"). The checklist adds further unreliability: it claims a Limitations section (Section 7) that does not exist, and says no human subjects were used despite the 10-participant study in Section 4.3. The authors admit the prompt \"heavily influences\" LLM behavior, but the pilot-driven prompt selection is undisclosed, so the measured behavior could be an artifact of prompt choice rather than a property of LLMs.\n\nCredit where it is due: the research questions are sensible, and studying long-horizon IPD with 240 strategies plus switch adaptation is a reasonable incremental extension of the Axelrod-style LLM work that already exists. The behavioral metrics (niceness, provocability, generosity) and the morality ratings are standard and appropriate. If the experiments were run as described and properly reported, this could be a useful, modest empirical data point for the mixed human-AI deployment question. It is not transformative, but it is relevant.\n\nRecommendation: desk reject in its current form. The paper is internally inconsistent and not reproducible. If the authors return with a full Results section, named models, the complete prompt, code/data, fixed contradictions, and an honest checklist, it would deserve a serious referee. For now, do not cite and do not spend reading-group time on it.","headline":"Incomplete draft: the central claims are unverifiable because the Results section is missing, the LLM is unnamed, the prompt is truncated, and the abstract contradicts the intro and Table 1 on the human-AI comparison.","tokens_in":12860,"tokens_out":2772,"would_cite":false,"duration_ms":27039,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Language models match the best-known classical strategies in the iterated prisoner's dilemma.","keywords":["iterated prisoner's dilemma","language models","cooperation","game theory","human-AI interaction","behavioral traits","strategy adaptation","Axelrod tournament"],"falsifier":"Run the same tournament with the same models but a differently worded neutral prompt, such as 'maximize your total payoff' versus 'be a cooperative partner', and check whether the performance parity and the niceness/provocability/generosity profile survive; if the profile changes materially, the reported behaviors are prompt effects rather than general LLM properties. Alternatively, repeat with several openly named LLMs and see whether all of them match the classical baselines.","tokens_in":12096,"feed_emoji":"🤖","tokens_out":5967,"duration_ms":54200,"temperature":0.7,"pith_summary":"This paper sets out to show that current language models, prompted to play an iterated prisoner's dilemma, behave like strong long-horizon cooperators: they score on par with or better than the best-known classical strategies, display the traits of successful cooperative play (niceness, provocability, generosity), and detect when an opponent changes strategy, adapting within a few rounds. The authors argue this matters because LLMs are increasingly deployed in interactive online environments where sustained cooperation and quick response to shifts determine whether human-AI collaboration works. The result is an empirical baseline for the cooperative behavior of LLM agents in mixed human-AI settings.","feed_headline":"LLMs rival the best classic prisoner's dilemma strategies","feed_subtitle":"Models played 240 classic opponents, kept pace, and adjusted to mid-game strategy switches within a few rounds.","key_machinery":"The load-bearing object is the iterated prisoner's dilemma itself with the classical payoff ordering H=5, R=3, P=1, L=0 and the condition H+L<2R, which removes the alternating-defect loophole and makes repeated cooperation viable. The LLM is turned into a strategy by prompting it with the full history of prior rounds; the analysis then runs through the behavioral trait metrics—niceness (initial cooperation), provocability (retaliation after defection), generosity (forgiveness of defections), plus the good-partner, Eigenjesus, and Eigenmoses morality ratings—to characterize how the model plays.","core_discovery":"The central claim is that a language model, given only the history of previous rounds and asked to cooperate or defect, can sustain competitive play in the iterated prisoner's dilemma. In a tournament against 240 classical strategies, the LLM-based agent accumulates wins and score advantage over time, approximating or exceeding tit-for-tat and other strong classical entries. On behavioral metrics, the agent is nice (initially cooperative), provocable (retaliates after defection), and generous (forgives some defections), matching the profile of the best classical strategies. In strategy-switch experiments, the model detects a change in opponent behavior within a few rounds and adjusts. Compar","pith_inferences":["Because the authors report that the prompt heavily influences behavior and tested only one selected prompt with unspecified models, the observed traits are most safely read as properties of that prompt-model pairing, not of language models generally; re-running with other neutral prompts and model families is a direct test.","The apparent exploitative tilt of LLMs after a switch may reflect the objective given in the prompt (maximize score) rather than a fixed social disposition; prompting for fairness or long-term partnership could shift the behavior toward the human pattern.","The framework should transfer to other repeated games with mixed motives, such as public goods or chicken, where the same niceness/provocability/generosity axes can be measured, offering a way to map LLM social behavior beyond the prisoner's dilemma."],"forward_implications":["If LLM agents match or exceed the best classical IPD strategies, they can be viable long-horizon partners in repeated mixed-motive interactions, not just one-shot decision-makers.","The demonstrated niceness/provocability/generosity profile suggests LLMs may internalize norms that sustain cooperation, such as forgiving occasional defections while punishing chronic ones.","Fast switch detection means LLM agents can track changing opponents in dynamic environments, which is directly relevant to deployed assistants and agents that meet new users or adversaries.","The human comparison implies mixed human-AI teams may face a tension between short-term payoff optimization and relational cooperation, since LLMs favor the former and humans the latter.","The results establish a baseline for studying LLM behavior in more complex mixed human-AI social environments, such as multi-player or noisy games."],"supporting_citations":[{"why":"Supplies the 240-strategy classical baseline and the tournament setting that the LLM agents are tested against.","marker":"[5]"},{"why":"Defines the morality metrics (good-partner, Eigenjesus, Eigenmoses) used to characterize LLM cooperative behavior.","marker":"[13]"},{"why":"Provides the iterated prisoner's dilemma reward structure and the behavioral trait framework the analysis builds on.","marker":"[15]"},{"why":"Prior study of LLM-agent cooperation that this work extends to long-horizon and mixed human-AI settings.","marker":"[3]"}],"fun_headline_variants":["LLMs match top strategies in prisoner's dilemma tournament","AI learns to cooperate, retaliate, forgive in game theory test","Language models adapt fast to opponent switches in dilemma","LLM agents rival classic game theory in long-term play","AI adapts to opponent switches like humans in prisoner's dilemma"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The claim that these are properties of language models rests on the assumption that the one chosen prompt and the unnamed tested models represent language models as a class; the authors themselves note that prompt instructions heavily influence LLM behavior.","fun_headline_variants_meta":{"raw":{"variants":["LLMs match top strategies in prisoner's dilemma tournament","AI learns to cooperate, retaliate, forgive in game theory test","Language models adapt fast to opponent switches in dilemma","LLM agents rival classic game theory in long-term play","AI adapts to opponent switches like humans in prisoner's dilemma"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00107,"raw_usage":{"total_tokens":4318,"prompt_tokens":745,"completion_tokens":3573,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":489,"completion_tokens_details":{"reasoning_tokens":3492}},"tokens_in":489,"tokens_out":3573,"duration_ms":26093,"temperature":1.0,"reasoning_tokens":3492,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T05:49:24.959160+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same tournament with the same models but a differently worded neutral prompt, such as 'maximize your total payoff' versus 'be a cooperative partner', and check whether the performance parity and the niceness/provocability/generosity profile survive; if the profile changes materially, the reported behaviors are prompt effects rather than general LLM properties. Alternatively, repeat with several openly named LLMs and see whether all of them match the classical baselines.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 240-strategy classical baseline and the tournament setting that the LLM agents are tested against."},{"cited_title":"Singer-Clark","cited_arxiv_id":null,"evidence_quote":"Defines the morality metrics (good-partner, Eigenjesus, Eigenmoses) used to characterize LLM cooperative behavior."},{"cited_title":"[Yes] \" is generally preferable to","cited_arxiv_id":null,"evidence_quote":"Provides the iterated prisoner's dilemma reward structure and the behavioral trait framework the analysis builds on."},{"cited_title":"cooperate or collapse: Emergence of sustainable cooperation in a society of llm agents","cited_arxiv_id":null,"evidence_quote":"Prior study of LLM-agent cooperation that this work extends to long-horizon and mixed human-AI settings."}],"review_version":1}