{"id":"075060b4-bda3-423f-8228-94e9851947ca","arxiv_id":"2505.24255","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"In ultimatum games, LLM agents prompted with theory-of-mind reasoning align more closely with human norms, with first-order ToM best for proposers, combined ToM for accepting offers, and zero-order ToM for rejecting offers, though results rely on heuristic human-behavior expectations.","lead":"Researchers tested whether prompting AI chatbots to think about their own and others' beliefs, called theory of mind, makes them behave more like humans in a negotiation game. They ran 2,700 games across six AI models and found the best reasoning style depends on the role: proposers do best thinking about the other player, responders do best with combined self and other reasoning.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim rests on unvalidated heuristic thresholds for human-aligned behavior; the paper's own sensitivity analysis flips the proposer-belief ordering, so ToM's alignment benefit is not robust.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: the deviation scores that drive all regressions depend on heuristically defined expected human ranges in Table 2, and the Appendix D sensitivity analysis shows that a plausible change to the Fair expectation flips the proposer-belief alignment ordering. I agree with this assessment. The paper's own Limitations section concedes that expected behaviors are defined independently of belief interactions and that findings may be sensitive to those assumptions. Since the regressions in Table 4 use these deviation scores as dependent variables, every reasoning-method coefficient inherits the same fragility. The central claim that ToM reasoning enhances alignment is therefore not robust to the metric's construction. This warrants a conditional verdict: the paper's framework is reasonable and the detailed prompts support reproducibility, but the alignment conclusion requires additional validation, either through systematic threshold perturbation or direct human comparison. Because the reader already assigned CONDITIONAL, my read does not change the verdict. I would not escalate to rejection, since the authors have disclosed the heuristic nature of the expectations and provided a sensitivity analysis; the concern is addressable with further evidence rather than fatal to the experimental design.","tokens_in":23526,"tokens_out":3728,"duration_ms":50778,"concrete_test":"Recompute the Table 4 regressions under a grid of plausible Table 2 thresholds, varying Greedy and Selfless expectations in 10-point increments and treating Fair as both a point and a range; if the sign or significance of the key ToM coefficients changes, the alignment claim is threshold-dependent. Alternatively, collect human participant data in the identical multi-round protocol and define deviation scores empirically; if the ToM ordering changes, the central claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims ToM reasoning enhances behavior alignment, decision-making consistency, and negotiation outcomes. Operationally, 'alignment' is measured by deviation from the Table 2 expected human shares. Those thresholds are explicitly heuristic; the paper even defines Selfless as the 'conceptual opposite' of Greedy and disclaims interaction effects in the Limitations section. Appendix D shows that merely converting Fair from a point (50%) to a range (30-70%) reverses the proposer-belief ranking: Fair proposers are least deviant under the original expectations, but Selfless proposers become least deviant under the relaxed range. Because the deviation scores P, RA, and RR are the dependent variables in all three OLS regressions (Eq. 3; Table 4), the entire ToM-coefficient story is contingent on these hand-picked thresholds. For example, first-order ToM's proposer benefit is only beta = -0.164 versus Vanilla's beta = -0.14, and zero-order ToM's rejection benefit is not significant (beta = 0.047, p > 0.05). A different, equally defensible expectation set could change which reasoning method appears best. The Limitations section acknowledges that 'our findings may be sensitive to the assumptions behind those expectations.' Without validation against actual human data or a systematic perturbation of all thresholds, the central claim is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Using 2,700 ultimatum-game simulations across six LLMs, the paper studies how prosocial belief prompts (Greedy, Fair, Selfless) and reasoning methods (Vanilla, CoT, zero-order/first-order/both ToM) affect alignment with expected human behavior. Alignment is measured by deviation scores P, RA, and RR computed against heuristic expected human shares in Table 2; OLS regressions in Table 4 identify role-specific effects, including first-order ToM as best for proposers, combined ToM as best for responder acceptances, and zero-order/CoT as best for rejections. The authors conclude that ToM reasoning enhances behavior alignment, decision-making consistency, and negotiation outcomes.","tokens_in":23852,"tokens_out":6597,"duration_ms":73446,"significance":"If the findings held, this study would provide a practical, role-dependent recipe for steering LLM negotiation behavior toward human norms and would extend the study of ToM in LLM agents beyond static benchmarks to dynamic economic games. The paper has genuine strengths: a large simulation grid (45 experiments per model across 6 models), explicit role-specific strategies, a statistical analysis of deviations, an explicit sensitivity analysis, and a public code link. However, the alignment metric is based on unvalidated heuristic human-expectation thresholds, and the paper's own sensitivity analysis changes the proposer-belief ordering and removes significance for first-order ToM in the proposer regression. The central claim is therefore conditional on assumptions that need either external validation or much more thorough perturbation before it can be accepted as stated.","major_comments":[{"comment":"The deviation scores P, RA, and RR are computed against the heuristic expected shares in Table 2, and these scores are the dependent variables in all three regressions (Eq. 3; Table 4). The sensitivity analysis in Appendix D shows that changing the Fair expectation from a point (50%) to a range (30–70%) reverses the proposer-belief alignment ordering: Fair proposers are least deviant under the original expectations (Table 4: Greedy β=0.623, Selfless β=0.581, Fair reference), but Selfless proposers become least deviant under the relaxed range (Table 5: Greedy β=-1.204, Selfless β=-1.246). The same change makes the proposer reasoning coefficients for zero-order and first-order ToM non-significant (Table 5: zero-order -0.046, first-order -0.071, both p>0.05). Because the central claims in RQ2/RQ3 and the abstract depend on these thresholds, and the Limitations section itself states that 'our findings may be sensitive to the assumptions behind those expectations,' the conclusion that ToM enhances alignment is not robust. The authors should validate the expectations against human ultimatum-game data or systematically perturb all thresholds, not only Fair, and report which conclusions survive.","section":"§4.2, Table 2, Appendix D (Table 5), Eq. (3)"},{"comment":"The abstract claims that ToM reasoning enhances 'behavior alignment, decision-making consistency, and negotiation outcomes.' No metric for decision-making consistency is defined in §4.3; the performance metrics (AC, Avg. Turns, payouts) measure outcomes but not consistency. The negotiation-outcome claim is also not uniformly supported by Table 4: for proposers, no-reasoning Vanilla (β=-0.140, p<0.01) is nearly as strong as first-order ToM (β=-0.164, p<0.01), and for rejections the best ToM variant, zero-order, is not significant (β=0.047, p>0.05). Table 3 contains many cells where Vanilla or CoT matches or beats ToM (e.g., Greedy-Fair with GPT-4o-mini: Vanilla, CoT, and ToM Both all achieve 100% AC). The regression evidence supports the more modest, role-specific conclusion in §5.2 that 'different roles benefit from different reasoning methods,' and the abstract should be revised to match that evidence.","section":"Abstract and §5.2"},{"comment":"RR is defined as the deviation of rejected shares from expected accepted shares, but for games that end in the first round there is no rejected share; Table 6 codes these cases as RR=-1 and notes that -1 simply means the game ended in one turn. If the RR regression in Table 4 includes these -1 values as numeric outcomes, the coding is arbitrary and can change the responder-rejection coefficients; games with no rejection should be excluded or modeled separately rather than assigned a constant. The authors should clarify how the -1 entries were treated in the OLS and, if they were included, rerun the RR analysis without them.","section":"§4.3 and Table 6"},{"comment":"The text ranks zero-order ToM above CoT and first-order ToM for rejection behavior, but the zero-order coefficient (β=0.047, p>0.05) and first-order coefficient (β=-0.143, p>0.05) are not statistically distinguishable from the CoT reference category; only Both ToM (β=-0.208, p<0.05) and Vanilla (β=-0.387, p<0.01) differ significantly from CoT, in the negative direction. Ordering non-significant coefficients should not be presented as evidence that zero-order ToM 'seemed to work best' for rejections.","section":"§5.2, Table 4 (Responder Rejections)"}],"minor_comments":[{"comment":"The text says the sensitivity analysis is presented in Appendix F, but the actual appendix is Appendix D; please fix the cross-reference.","section":"§4.2 vs Appendix D"},{"comment":"The sentence 'Consistent with previous findings, reasoning models exhibit limited capability compared to models with ToM reasoning, different roles of the game benefits with different orders of ToM reasoning' is ungrammatical and unclear; it should be rewritten.","section":"Abstract"},{"comment":"The table header 'Ind. Coefficient Var. P ↓ R A ↓ R R ↑' is formatted incorrectly and is hard to read; please reformat the header labels.","section":"Table 4"},{"comment":"The game definition contains inconsistent subscripts, e.g., 'Rt_P ; Dt_P ; Ri_R; Dt_R' uses i instead of t; please fix the notation.","section":"§3.3"},{"comment":"The phrase 'originally point-wise expectations (5)' is ambiguous; it should be clarified whether the 5 refers to the stake value or to the dollar amount of an equal split.","section":"Appendix D"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a cs.CL audience and the simulation effort is substantial. My main concern is not novelty or effort but validity of the alignment metric: the paper's own sensitivity analysis undermines the proposer-level conclusions, and the coding of RR for one-round games is unclear. The authors have the data and the analysis infrastructure to address this by validating against human behavior or by reporting a thorough threshold perturbation; I would encourage the editor to request that as a major revision rather than rejecting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a useful, honestly reported empirical study about whether ToM prompting steers LLM agents closer to expected human behavior in ultimatum games. But the headline claim is not as solid as the abstract suggests: the 'alignment' metric is built on heuristic thresholds, and the authors' own sensitivity analysis flips one of the main ordering conclusions.\n\nWhat's new and good: the paper systematically varies three prosocial beliefs (greedy, fair, selfless), four reasoning methods (CoT, zero-order, first-order, combined ToM, plus vanilla), and six LLMs, over 2,700 simulated games. The prompts are detailed and the code is public. The regression framework is transparent, and the role-dependent pattern (first-order ToM helps proposers, combined ToM helps responder acceptance, zero-order helps rejections) is an interesting, falsifiable claim. I also credit the authors for including the sensitivity analysis and a plain-spoken limitations section that explicitly says the findings may be sensitive to the expectation assumptions. That is more than many papers do.\n\nSoft spots: the abstract overstates the results. 'Decision-making consistency' and 'negotiation outcomes' are not actually measured as promised, and the claim that 'ToM reasoning enhances behavior alignment' is too broad given that vanilla no-reasoning is statistically tied for second-best among proposers and zero-order ToM's rejection benefit is not significant. More importantly, the deviation scores—the dependent variables in all three regressions—are computed against the Table 2 expected share ranges that are 'heuristically defined.' Appendix D shows that changing fair expectations from a point to a range flips the proposer-belief ordering (fair proposers go from most to least aligned) and also demotes first-order ToM below vanilla and combined ToM for proposers. That means the paper's signature finding about which ToM level works best for proposers is not robust to a defensible change in the alignment metric. Per-condition sample size is also 10 games, so the OLS estimates are noisy, though that is secondary.\n\nResearchers working on LLM social simulation, negotiation agents, or ToM evaluation will want to know this paper. It deserves a serious referee—the empirical design is reproducible and the question matters—but it needs major revision: validate the expectation thresholds against human data or at least systematically perturb all of them, measure the outcomes claimed in the abstract, and temper the language. I would not reject it, but I would not accept it as is.","headline":"Useful, reproducible mapping of ToM and prosocial beliefs in LLM ultimatum games, but the headline alignment claim rests on hand-set human expectation thresholds that the paper's own sensitivity analysis shows to be brittle.","tokens_in":24361,"tokens_out":4207,"would_cite":false,"duration_ms":47079,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that prompting LLM agents with theory-of-mind reasoning in ultimatum games steers their offers, acceptances, and rejections closer to expected human behavior, and that the best ToM level depends on the player's role.","keywords":["theory of mind","LLM agents","ultimatum game","prosocial beliefs","behavioral alignment","negotiation","chain-of-thought","deviation scores"],"falsifier":"Running the same 2,700-game protocol with parallel human participant data or with alternative fairness thresholds and showing that the ToM ordering of deviation scores reverses, or disappears, would settle the claim. Concretely, the sensitivity analysis in Appendix D already shows proposer-belief ordering flips when fair expectations become a range; a replication with human baselines would test the core alignment claim.","tokens_in":23367,"feed_emoji":"🧠","tokens_out":4373,"duration_ms":47501,"temperature":0.7,"pith_summary":"This paper is trying to establish that theory-of-mind reasoning can steer LLM agents toward human-aligned negotiation behavior in the ultimatum game. Across 2,700 simulations with six LLMs, agents role-played greedy, fair, or selfless proposers and responders, using chain-of-thought or zero-, first-, or combined-order ToM reasoning. The paper reports that ToM reasoning improves behavior alignment, decision consistency, and negotiation outcomes, and that different game roles benefit from different ToM orders: first-order ToM is best for proposers, combined ToM is best for responder acceptances, and zero-order ToM is best for rejections. If correct, it gives a practical recipe for choosing reasoning prompts in social simulations.","feed_headline":"ToM reasoning steers LLM negotiation closer to human norms","feed_subtitle":"2,700 ultimatum games show first-order ToM best for proposers, combined ToM best for responders.","key_machinery":"The central object is the ultimatum game with role-specific strategies and the deviation score (DS), which measures how far an agent's proposed, accepted, or rejected shares fall from heuristically defined expected human behavior ranges for greedy, fair, and selfless beliefs. Prompted BDI reasoning (zero-order introspection, first-order inference about the other player, and both) is the mechanism that carries the argument; three OLS regressions on the proposer deviation, responder acceptance deviation, and responder rejection deviation link the design to human norms.","core_discovery":"The central claim is that ToM reasoning enhances behavior alignment, decision-making consistency, and negotiation outcomes in LLM-based ultimatum game agents. The paper finds role-dependent benefits: first-order ToM yields the least deviation for proposers' initial offers ($\\beta=-0.164$, $p<0.01$), combined zero-plus-first-order ToM yields the least deviation for responders' accepted shares ($\\beta=-0.2534$, $p<0.01$), and zero-order ToM performs best for rejections. Fair-fair belief combinations align most closely with expected human behavior, and models such as GPT-4o and Llama 3.3 are most consistently aligned across roles.","pith_inferences":["A testable extension would check whether the same role-dependent ToM ordering appears in other bargaining games, such as the dictator game or trust game, which would indicate a general principle rather than an ultimatum-game artifact.","Because the paper defines selfless behavior as the conceptual opposite of greedy behavior, an alternative expectation derived from reciprocal fairness models could be used to test whether combined ToM remains the best responder strategy.","The alignment gains could partly come from the extra reasoning tokens ToM prompts elicit rather than from theory-of-mind per se; a matched-length CoT control would separate these explanations.","For human-AI negotiation systems, the paper's role-dependent reasoning prescription suggests a practical tuning rule, though the ethical guardrails it mentions would be essential before deploying such agents in high-stakes settings."],"forward_implications":["Prompt designers can choose ToM order by role: first-order reasoning for proposers, combined reasoning for responders' acceptance decisions, and zero-order reasoning for rejection decisions.","Fair-fair belief combinations are the most human-aligned configuration, making them a sensible default for social simulations of cooperative negotiation.","Reasoning models do not need explicit chain-of-thought prompts to reason internally, but adding ToM prompts still improves alignment with human norms.","Reported alignment levels are tied to the specific expected human behavior thresholds; changing those thresholds can change which belief or reasoning method looks most aligned."],"supporting_citations":[{"why":"Defines theory of mind, the cognitive capacity the paper investigates in LLM agents.","marker":"Premack and Woodruff, 1978"},{"why":"Supplies the Belief-Desire-Intention model that structures the zero-order, first-order, and combined ToM prompts.","marker":"Georgeff et al., 1999"},{"why":"Introduces the ultimatum game, the controlled negotiation environment used for all simulations.","marker":"Güth et al., 1982"},{"why":"Documents anomalies in ultimatum game behavior, including rejection of very generous offers, which motivates the selfless-proposer expectations.","marker":"Thaler, 1988"},{"why":"Provides the fairness-based theoretical foundation for setting the fair players' expected 50% share.","marker":"Fehr and Schmidt, 1999"},{"why":"Offers cross-cultural ultimatum game evidence used to define expected human behavior ranges for greedy, fair, and selfless beliefs.","marker":"Henrich et al., 2005"},{"why":"Supplies the behavioral game theory background, including the subgame perfect Nash equilibrium and why human behavior deviates from it.","marker":"Camerer, 2011"},{"why":"Supports the notion of a fair offer as an equal split, grounding the fair belief expectations.","marker":"Schuster, 2017"},{"why":"Provides the simulation protocol (10 games per experiment) and the behavioral-alignment approach that this paper extends with ToM reasoning.","marker":"Sreedhar and Chilton, 2025"},{"why":"Inspires the multi-level ToM reasoning design and demonstrates that considering others' mental states improves negotiation outcomes.","marker":"Qiu et al., 2024"}],"fun_headline_variants":["ToM boosts LLM negotiation alignment in ultimatum games","First-order ToM best for proposers, combined for responders","2700 games: ToM reasoning aligns LLMs with human norms","Role-specific ToM improves fairness in LLM ultimatum play","Theory of mind steers LLM agents toward human-like deals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The measurement of 'alignment' depends on hand-chosen human expectation thresholds in Table 2 (e.g., greedy proposers offer at least 70%, fair players split 50/50, selfless players give away most), and the paper's own sensitivity analysis shows that changing the fair threshold from a point to a range changes which proposer belief looks most aligned.","fun_headline_variants_meta":{"raw":{"variants":["ToM boosts LLM negotiation alignment in ultimatum games","First-order ToM best for proposers, combined for responders","2700 games: ToM reasoning aligns LLMs with human norms","Role-specific ToM improves fairness in LLM ultimatum play","Theory of mind steers LLM agents toward human-like deals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000171,"raw_usage":{"total_tokens":1258,"prompt_tokens":915,"completion_tokens":343,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":255}},"tokens_in":531,"tokens_out":343,"duration_ms":4643,"temperature":1.0,"reasoning_tokens":255,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:27:52.655077+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Running the same 2,700-game protocol with parallel human participant data or with alternative fairness thresholds and showing that the ToM ordering of deviation scores reverses, or disappears, would settle the claim. Concretely, the sensitivity analysis in Appendix D already shows proposer-belief ordering flips when fair expectations become a range; a replication with human baselines would test the core alignment claim.","supporting_citations":[],"review_version":1}