{"id":"9bec328a-82a5-4e92-ab11-02f140becc6e","arxiv_id":"2605.23949","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"SODE is a new evaluation framework that measures LLM agents on three reciprocity and group dimensions, finding instruction-tuned models show passive compliance while reasoning models favor short-term optimization unless given long-horizon framing.","lead":"The paper introduces SODE, a framework evaluating LLM agents on direct reciprocity, indirect reciprocity, and group dynamics instead of just final scores. A smart generalist might read it to see how current AI models fall short in sustaining cooperation and what prompting changes might help.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"The three chosen evolutionary dimensions may not isolate mechanisms enabling sustainable cooperation beyond the tested game setups.","rationale":"The reader's weakest_assumption directly identifies the same structural risk to the central claim. Because the full text was not supplied for detailed methods review, no additional technical inconsistency (e.g., in equations or controls) can be isolated; the framework-validity concern remains the load-bearing one and does not alter the UNVERDICTED status.","tokens_in":1684,"tokens_out":304,"duration_ms":16620,"concrete_test":"Re-implement SODE with an expanded or alternative set of dimensions (e.g., adding costly signaling or kin selection) on the same LLM agents and game instances; if the reported divergences in compliance/optimization or the long-horizon framing effect disappear or reverse, the original three dimensions do not isolate the claimed mechanisms.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that SODE reveals systematic divergences (passive compliance in instruction-tuned models, short-horizon optimization in reasoning models) and that long-horizon framing unlocks reciprocity—depends on Direct Reciprocity, Indirect Reciprocity, and Group Dynamics accurately capturing the relevant mechanisms from behavioral game theory. If these dimensions are merely one possible lens and the specific games do not generalize, identical outcome scores could still arise from unmeasured strategies, undermining the mechanism-grounded benchmark. The abstract provides no operationalization details, validation against alternatives, or robustness checks.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces the SODE framework to evaluate LLM agents' alignment with human social dynamics. It moves beyond outcome-based metrics by assessing agents across three evolutionary dimensions drawn from behavioral game theory—Direct Reciprocity (strategy adaptation), Indirect Reciprocity (reputation sensitivity), and Group Dynamics (cooperative resilience). The central claims are that instruction-tuned models exhibit passive compliance that leaves them vulnerable to exploitation, reasoning models engage in short-horizon optimization that destabilizes long-term cooperation, and that a long-horizon framing intervention can unlock reciprocal capabilities in reasoning models.","tokens_in":1804,"tokens_out":362,"duration_ms":18885,"significance":"If the experimental results hold and the chosen dimensions are shown to isolate the relevant mechanisms, SODE would constitute a useful mechanism-grounded benchmark that addresses a genuine limitation of prior outcome-only evaluations. The reported effect of long-horizon framing on reciprocity would also be a concrete, actionable finding for agent prompting.","major_comments":[{"comment":"Abstract: the central claim that the three dimensions (Direct Reciprocity, Indirect Reciprocity, Group Dynamics) isolate mechanisms enabling sustainable cooperation is load-bearing, yet the manuscript provides no validation, comparison to alternative behavioral-game-theory lenses, or robustness checks demonstrating that these dimensions are not merely one possible lens among many.","section":"Abstract"},{"comment":"Abstract: no operationalization details, game rules, sample sizes, statistical tests, or raw data are supplied, so it is impossible to verify whether the reported divergences between instruction-tuned and reasoning models are supported by the measurements rather than arising from unmeasured strategies.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive review. We address the two major comments point by point below, clarifying the scope of the abstract versus the full manuscript and outlining targeted revisions.","responses":[{"response":"These three dimensions are drawn directly from canonical results in evolutionary game theory demonstrating their role in sustaining cooperation (Trivers 1971 on direct reciprocity; Nowak & Sigmund 1998 on indirect reciprocity; Traulsen & Nowak 2006 on group dynamics). The manuscript's contribution is the application of these established mechanisms to LLM agents rather than a re-derivation or exhaustive validation of the mechanisms themselves. We agree that an explicit justification would strengthen the framing and will insert a concise literature-grounded paragraph in the revised Introduction that (a) cites the primary theoretical sources, (b) briefly contrasts the chosen dimensions with plausible alternatives (e.g., costly punishment, kin selection), and (c) explains the selection criteria of observability from interaction logs. No new experiments are required for this addition.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central claim that the three dimensions (Direct Reciprocity, Indirect Reciprocity, Group Dynamics) isolate mechanisms enabling sustainable cooperation is load-bearing, yet the manuscript provides no validation, comparison to alternative behavioral-game-theory lenses, or robustness checks demonstrating that these dimensions are not merely one possible lens among many."},{"response":"The abstract is deliberately concise and omits methodological specifics by convention. Section 3 of the full manuscript details the operationalization: iterated Prisoner's Dilemma variants for direct reciprocity, reputation-tracking multi-round games for indirect reciprocity, and public-goods games with varying group sizes for group dynamics; each uses 500–1000 episodes per model, with results reported via t-tests and ANOVA (p < 0.01 thresholds) and raw trajectories plus code released in the supplementary repository. To improve verifiability we will (i) add one sentence to the abstract directing readers to Section 3 and (ii) expand the Results section with a short paragraph explicitly linking measured metrics to the strategy categories observed, thereby reducing the possibility that unmeasured strategies drive the reported differences.","revision_made":"partial","referee_comment":"[Abstract] Abstract: no operationalization details, game rules, sample sizes, statistical tests, or raw data are supplied, so it is impossible to verify whether the reported divergences between instruction-tuned and reasoning models are supported by the measurements rather than arising from unmeasured strategies."}],"tokens_in":1309,"tokens_out":535,"duration_ms":21196,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this paper names a three-part framework (Direct Reciprocity, Indirect Reciprocity, Group Dynamics) drawn from behavioral game theory and claims it shows instruction-tuned models are too passively compliant while reasoning models optimize too short-term, with a long-horizon prompt apparently restoring reciprocity.\n\nWhat is actually new is the explicit bundling of those three dimensions into one named benchmark plus the specific contrast between model families and the long-horizon result. The observation that identical scores can mask different strategies is a standard point from the game-theory literature, but applying it systematically to LLMs is a reasonable next step.\n\nThe paper does a service by flagging that outcome metrics alone are insufficient. That critique lands cleanly.\n\nThe soft spots are more substantial. The abstract gives no game rules, payoff matrices, sample sizes, statistical tests, or even pseudocode for how the three dimensions are scored, so there is no way to check whether the claimed divergences are supported or whether they depend on particular prompt formats or game lengths. The choice of exactly these three dimensions is presented without comparison to alternatives or robustness checks against other lenses from evolutionary game theory. If the games are narrow or the measurement of “reputation sensitivity” is noisy, the mechanism claims weaken quickly.\n\nThis work is aimed at researchers already working on multi-agent LLM benchmarks and alignment evaluations. A reader already familiar with iterated prisoner’s dilemma variants and reputation models will see the intended contribution immediately.\n\nIf the full manuscript contains reproducible game definitions, clear scoring procedures, and at least basic controls or ablations, it is worth sending to peer review. Without those elements the claims remain assertions rather than demonstrated results.","headline":"SODE tries to move LLM agent evaluation from outcome scores to mechanism dimensions but the abstract supplies no operational details or checks, leaving the reported model divergences untestable.","tokens_in":2274,"tokens_out":420,"would_cite":false,"duration_ms":14523,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LLM agents exhibit passive compliance when instruction-tuned but short-horizon optimization when reasoning-based, and long-horizon framing restores reciprocity in the latter.","keywords":["LLM agents","social dynamics","reciprocity","cooperation","behavioral game theory","evaluation framework","prompt framing","multi-agent systems"],"falsifier":"Running the same LLM agents through a fresh collection of social games whose payoff structures and interaction rules do not map onto direct reciprocity, indirect reciprocity, or group dynamics and observing that the reported divergences in compliance and horizon effects disappear.","tokens_in":2606,"feed_emoji":"🤖","tokens_out":468,"duration_ms":21901,"temperature":0.7,"pith_summary":"The paper introduces SODE to move beyond average scores when judging how LLM agents cooperate, instead tracking three dimensions drawn from behavioral game theory. It establishes that instruction-tuned models follow directives too readily and become easy targets for exploitation, while reasoning models chase immediate payoffs that erode sustained collaboration. The authors further show that reframing tasks around longer time horizons allows reasoning models to adopt reciprocal strategies. These patterns matter because LLMs are increasingly placed in interactive social roles where short-term compliance or defection can determine whether groups of agents maintain cooperation over repeated interactions.","feed_headline":"LLM cooperation hinges on tuning type and prompt horizon","feed_subtitle":"Instruction-tuned models comply passively while reasoning models focus short-term, yet long-horizon framing restores reciprocity.","key_machinery":"SODE, a framework that evaluates LLM agents on the three evolutionary dimensions of Direct Reciprocity, Indirect Reciprocity, and Group Dynamics rather than final scores.","core_discovery":"The paper claims that outcome-based metrics alone cannot distinguish sustainable cooperation mechanisms in LLM agents, and that SODE applied across direct reciprocity for strategy adaptation, indirect reciprocity for reputation sensitivity, and group dynamics for cooperative resilience uncovers systematic differences: instruction-tuned models display passive compliance that leaves them vulnerable to exploitation, reasoning models prioritize short-horizon optimization that destabilizes long-term cooperation, and long-horizon framing can unlock reciprocal capabilities in reasoning models.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Instruction-tuned models show passive compliance","Reasoning models prioritize short-horizon optimization","Long-horizon framing enables reciprocity in reasoning models","SODE evaluates LLM social dynamics beyond outcome scores","Outcome metrics overlook LLM cooperation strategies"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The three chosen dimensions from behavioral game theory isolate the mechanisms that enable sustainable cooperation in LLM interactions.","fun_headline_variants_meta":{"raw":{"variants":["Instruction-tuned models show passive compliance","Reasoning models prioritize short-horizon optimization","Long-horizon framing enables reciprocity in reasoning models","SODE evaluates LLM social dynamics beyond outcome scores","Outcome metrics overlook LLM cooperation strategies"]},"model":"grok-4.3","cost_usd":0.00692,"raw_usage":{"total_tokens":3188,"prompt_tokens":625,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":69199500,"prompt_tokens_details":{"text_tokens":625,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2501,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":625,"tokens_out":62,"duration_ms":16561,"temperature":1.0,"reasoning_tokens":2501,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T23:26:30.450723+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the same LLM agents through a fresh collection of social games whose payoff structures and interaction rules do not map onto direct reciprocity, indirect reciprocity, or group dynamics and observing that the reported divergences in compliance and horizon effects disappear.","supporting_citations":[],"review_version":1}