Pith. sign in

REVIEW 5 major objections 7 minor 3 cited by

Beyond Nash Equilibrium: Bounded Rationality of LLMs and humans in Strategic Decision-making

T0 review · 5 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read In identical game experiments, LLMs reproduce human bounded rationality but apply the heuristics rigidly and ignore environmental cues, yielding a partial, amplified form of our bias.

desk verdict A useful methodological template for comparing LLM and human strategic play, but the headline claims about rigidity and weak adaptivity rest on unmatched protocols and point estimates without error bars; still worth a serious referee. read the letter →

arxiv 2506.09390 v1 pith:FFM4CRK5 submitted 2025-06-11 cs.AI cs.GT

classification cs.AIcs.GT MSC 91A0591A2668T50
keywords boundedrationalitylargelanguagemodelsRock-Paper-ScissorsPrisoner'sDilemmaNashequilibriumbehavioralgametheorystrategicdecision-makingadaptivity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether large language models share the bounded rationality that makes human players deviate from Nash equilibrium. It places six LLMs in the same experimental protocols used with human subjects for Rock-Paper-Scissors and the Prisoner's Dilemma. It finds that LLMs do reproduce the human heuristics of outcome-based strategy switching and increased cooperation when future interaction is possible, but apply them in a more rigid, exaggerated way and show weaker sensitivity to changes in the game environment. The authors conclude that current LLMs exhibit only a partial, amplified form of human-like bounded rationality rather than full Nash rationality, and argue that training with opponent modeling and theory-of-mind objectives is needed to close the gap.

What carries the argument

The method is the direct replication of established human-subject protocols—Zhang et al.'s RPS design and Bó's shadow-of-the-future PD design—as text-based LLM trials with identical payoffs, instructions, and feedback structure. The analytical engine is outcome-conditioned transition analysis: each move after a win, loss, or tie is classified as stay, upgrade, or downgrade, enabling a common yardstick for human and model heuristics. The PD setup isolates the shadow of the future by contrasting Dice sessions (continuation probability δ ∈ {0, 0.5, 0.75}) with Finite sessions of matched expected length (H = 1, 2, 4), so any difference in cooperation is attributable to the framing of future interaction rather than the number of rounds.

What would settle it

Present the same six LLMs with an explicit rule description of the WDLS bot ('the opponent repeats after a win and switches after a loss') before play. If any model then beats the bot by a margin comparable to humans (around +6.42 wins) and shows a measurable rise in win-stay transitions, the paper's claim that LLMs apply heuristics rigidly and ignore opponent structure would be falsified.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that LLMs share the two hallmark human departures from equilibrium play—outcome-conditioned switching in Rock-Paper-Scissors and the shadow-of-the-future boost in Prisoner's Dilemma cooperation—yet express them as stereotyped, environment-insensitive biases. In RPS, models fall into a dominant lose-downgrade pattern (switching to the losing move after a loss in over 60% of cases) and barely adjust when facing a WDLS bot that humans readily exploit, with humans beating that bot by +6.42 wins while LLMs lose to it by −6.44. In PD, LLM cooperation rises with the mere presence of a future (from δ=0 to higher continuation probabilities) but does not track the continuation probability itself, dipping slightly from 38.4% at δ=0.5 to 37.8% at δ=0.75, whereas human cooperation climbs from 27.4% to 37.6%. Reasoning-tuned models such as O1 and DeepSeek-R1 achieve equilibrium in the analytically explicit one-shot PD but underperform general models against rule-based adaptive opponents, revealing a persistent theory-of-mind deficit.

Load-bearing premise

The central comparison assumes that the human experiments and the LLM trials measure the same behavior: humans faced real monetary incentives and a rich social setting, while the LLMs received the same wording as text with only fictional payoffs, and the human baselines came from earlier studies rather than being rerun under the identical text-only condition.

Editorial extensions

If this is right

  • LLMs deployed in negotiation, auction, or coordination settings will carry predictable strategic biases that other agents can exploit, since the biases are rigid in the face of opponent structure and changing incentives.
  • Improving LLM strategic play will require explicit opponent modeling, reinforcement learning on diverse human gameplay traces, or theory-of-mind scaffolding—not merely larger scale or longer chain-of-thought reasoning.
  • Reasoning-enhanced models are not generally more rational: they hit equilibrium only when a closed-form solution exists, so evaluations of AI rationality must separate analytically solvable games from adaptive social interactions.
  • Distinct strategic signatures across model families imply that pre-training and alignment choices imprint durable decision heuristics, making model selection for strategic tasks a consequential design decision.
  • The observed rigidity is a concrete benchmark: any future model that matches humans' 20-plus-percentage-point rise in win-stay against the WDLS bot could be scored as a step toward human-level adaptability.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension follows from the paper's prompt-sensitivity caveat: if the shadow of the future is described purely as a structural continuation probability without future-oriented phrasing, LLM cooperation at δ=0 should drop toward the one-shot H=1 level; conversely, human-like phrasing may inflate cooperation beyond true structural incentives.
  • The human-LLM magnitude comparisons inherit a confound the paper acknowledges: cited humans faced real monetary stakes and a rich social setting, while models received only text with fictional payoffs. A matched head-to-head run in which humans and LLMs both play the same text-only protocol with identical non-monetary incentives would sharpen the rigidity claim.
  • Architectural fingerprints suggest a concrete training intervention: fine-tune a model family on mixed human-and-bot gameplay traces, then measure whether its WDLS win differential moves from the current −6.44 toward the human +6.42; success would confirm that bounded rationality can be reshaped by exposure to adaptive opponents.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 7 minor

Summary. The paper adapts two established human-subject paradigms—Rock-Paper-Scissors (Zhang et al., 2021) and repeated Prisoner's Dilemma (Bó, 2005)—to six LLMs. It uses text prompts that mirror the human instructions, logs action sequences, and compares aggregate choice frequencies, outcome-conditioned transition strategies, cooperation rates, and win differentials against the published human baselines. The central claim is that LLMs reproduce human bounded-rational heuristics (outcome-based strategy switching and cooperation under a shadow of the future) but apply these heuristics more rigidly and with weaker sensitivity to environmental changes, and that model families exhibit distinct 'strategic signatures.'

Significance. If the central comparative claim holds, the paper is a useful contribution to the empirical literature on LLM bounded rationality: it anchors the evaluation to external human data and theoretical Nash equilibria rather than fitting free parameters, covers six models from three families, and separates reasoning-enhanced from general models. The finding that reasoning models approach equilibrium in closed-form games but struggle in opponent-adaptive settings is interesting and worth developing. However, the headline conclusions rest on comparisons between a lab-based human arm with real monetary stakes and a text-based LLM arm without real payoffs, on point estimates without error quantification, and on aggregate PD numbers that are internally inconsistent. The significance is therefore conditional on resolving these load-bearing issues.

major comments (5)
  1. [§3, §5, §6] The central comparative claim—that LLMs apply human heuristics more rigidly and with weaker sensitivity to dynamic changes—requires that the human and LLM arms measure the same decision environment. They do not: the human baselines (Zhang et al. 2021; Bó 2005) involved real monetary payments, physical presence, social context, and human opponents, while the LLM arm receives text prompts narrating points and payment without any real payoff, and in RPS the LLM arm is a model-vs-model tournament or scripted bots. The statements in Section 3 that prompts 'closely align the language, incentives, feedback structure' and in Section 5 that the study uses 'identical payoffs, instructions' are therefore inaccurate. The observed differences in Tables 2, 5, and 6 (e.g., the flat δ response and minimal bot adaptation) could be artifacts of the incentive/modality gap rather than differences in bounded rationality. A concrete test would be to run a human text-based arm without monetary stakes or an LLM arm with real stakes; if that is infeasible, the comparative claims should be reworded as comparisons to published human benchmark data, with the protocol gap explicitly stated as a threat to validity.
  2. [Tables 2, 3, 5, 6, 7] All headline comparisons are presented as point estimates without standard errors, confidence intervals, or significance tests, and the number of independent LLM runs per cell is not reported. For example, Table 6's claim that LLMs 'dip slightly' from δ=0.5 (38.37%) to δ=0.75 (37.83%) rests on a 0.54 percentage-point difference with unknown sampling variance; similarly, Table 2's WDLS win differential of −6.44 for LLMs is given without dispersion. The chi-square tests mentioned in Section 4.1 establish outcome dependence but do not test the human-vs-LLM differences asserted in the abstract. Without error quantification, the 'weaker sensitivity' and 'more rigid' conclusions are not statistically supported.
  3. [Tables 6 and 7] The aggregate PD cooperation rates are internally inconsistent and dominated by one model. Averaging Table 7's six model percentages gives 41.15% for Dice δ=0.5 and 36.63% for δ=0.75, whereas Table 6 reports 38.37% and 37.83%; Finite H=1, 2, 4 averages match, so the discrepancy is not rounding. In addition, Claude-3.5 cooperates 100% in all three dice treatments, so this single model moves the aggregate δ=0 rate from 33.33% to about 20.0% when excluded. The Section 4.3 discussion of the 'shadow of the future' and of exaggerated sensitivity to future framing is therefore not robust to model composition and needs to be re-derived on a per-model basis with exact aggregation formulas.
  4. [Appendix A.2] The described RPS tournament is not 'fully crossed' as computed. The text says each model plays itself and each of the other five models, with each pairing repeated three times; that would give 6 self-pairings plus 15 unordered cross-pairings per replicate, or 63 matches if pairings are unordered and 108 if ordered, not the reported (6 choose 2) × 3 = 45. The payoff matrix in Table 1 is asymmetric in the player roles, so ordered pairings matter. This affects the sample size underlying Figure 1 and the chi-square tests, and the authors should correct the count or clarify the actual design.
  5. [§4.1, Figure 2] The transition taxonomy (stay/upgrade/downgrade) is defined relative to the player's own previous action, with 'upgrade' meaning the action that beats the previous one (e.g., Rock→Paper). Under this definition, after a loss, 'downgrade' is exactly the action that beats the opponent's previous action (if the opponent used the action that beat you), so lose-downgrade is a normatively sound best response to a win-stay opponent rather than a bias per se. The paper interprets a dominant lose-downgrade as evidence of rigidity (Section 4.3) but does not state whether the human benchmark (Zhang et al. 2021) used the same transition definition or how the classification would change if defined relative to the opponent's last move. This needs clarification and a robustness check before the 'mirror but amplify' claim can be evaluated.
minor comments (7)
  1. [§4] The first sentence of Section 4, 'In this, we first present...', is missing a noun and should read 'In this section, we first present...'.
  2. [Table 3] Table 3 is difficult to read because model names and numeric entries run together without visible cell separation; please reformat it as a proper table.
  3. [Figure 1] Figure 1(b) shows only aggregate LLM choice proportions; adding per-model overlays or error bars would make the comparison with human proportions more informative.
  4. [Appendix A.3] Only the δ=0.75 dice prompt is shown; the exact wording for δ=0 and δ=0.5 and how the four-sided dice rule maps onto these continuation probabilities should be included in the appendix.
  5. [§4.3] The phrase 'LLMs display an exaggerated sensitivity' is confusing because the same subsection later argues LLMs show 'weaker environmental sensitivity'; please align the terminology.
  6. [Table 1] The header uses 'Scissor' while the text uses 'Scissors'; please standardize the spelling.
  7. [§2] In the Related Work section, ', FAIR;' appears as a stray citation fragment and should be removed or completed.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: the central comparisons are direct measurements against external human baselines and closed-form Nash equilibria, with no fitted parameter or load-bearing self-citation chain.

full rationale

The paper's headline claims—LLMs reproduce human heuristics but apply them more rigidly and with weaker sensitivity to environmental change—are empirical comparisons, not derivations from fitted inputs. LLM action sequences are logged and summarized into transition frequencies, cooperation rates, and win differentials; human baselines come from prior laboratory studies (Zhang et al., 2021; Bó, 2005) and theoretical Nash equilibria are computed from the stated payoff matrices. No parameter is fitted to the human data and then used to generate the LLM predictions, so the comparison is not forced by construction. The only self-citation (Zhou et al., 2025) appears in Related Work alongside Mei et al. (2024) as an example of behavioral replication and is not load-bearing for any conclusion. The limitation that text-based LLM interaction 'lacks the social and contextual richness of human play' is a validity threat about incentive and modality matching, not circularity; even if the comparison were confounded, the LLM measurements would still not be constructed from the human benchmark values. No circular step is present.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

No numeric parameters are fitted to make the target claims work; the load-bearing choices are experimental settings (temperature, bot noise) and comparability assumptions about human baselines and fictional incentives. No new entities are postulated.

free parameters (2)
  • decoding temperature = 1.0 (all models)
    Set by hand to preserve stochastic behavior; directly shapes choice distributions and all probability-based metrics in Sections 4.1 and 4.2.
  • bot primary-action probability = 0.8 with 0.1 for each alternative action
    Chosen for WSLU and WDLS bots; win and payoff differentials in Tables 2 and 3 depend on how predictable the bots are.
assumptions (4)
  • domain assumption Human results from Zhang et al. (2021) and Bó (2005) are a valid baseline for LLM text trials.
    Invoked whenever LLM aggregate behavior is benchmarked against human percentages, e.g. Figure 1 and Tables 5/6.
  • domain assumption Fictional payment language in prompts induces behavior comparable to real monetary incentives for humans.
    Section 3 claims incentives are aligned; Appendix A.3 tells the model it will be paid, but no real payment occurs.
  • domain assumption Each model's behavior is stable under temperature 1.0 and the single chosen prompt template.
    The model-level strategic signatures (Section 4.3) treat one prompting configuration as representative of the model family.
  • standard math Chi-square tests of independence are an appropriate test for outcome-action dependence.
    Used in Section 4.1 to classify outcome-based agents.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Beyond Nash Equilibrium: Bounded Rationality of LLMs and humans in Strategic Decision-making." pith.science (2026). https://pith.science/paper/FFM4CRK5

@misc{pith2026250609390,
  author       = {Pith},
  title        = {Pith review of: Beyond Nash Equilibrium: Bounded Rationality of LLMs and humans in Strategic Decision-making},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FFM4CRK5}},
  note         = {Machine review of arXiv:2506.09390}
}
read the original abstract

Large language models are increasingly used in strategic decision-making settings, yet evidence shows that, like humans, they often deviate from full rationality. In this study, we compare LLMs and humans using experimental paradigms directly adapted from behavioral game-theory research. We focus on two well-studied strategic games, Rock-Paper-Scissors and the Prisoner's Dilemma, which are well known for revealing systematic departures from rational play in human subjects. By placing LLMs in identical experimental conditions, we evaluate whether their behaviors exhibit the bounded rationality characteristic of humans. Our findings show that LLMs reproduce familiar human heuristics, such as outcome-based strategy switching and increased cooperation when future interaction is possible, but they apply these rules more rigidly and demonstrate weaker sensitivity to the dynamic changes in the game environment. Model-level analyses reveal distinctive architectural signatures in strategic behavior, and even reasoning models sometimes struggle to find effective strategies in adaptive situations. These results indicate that current LLMs capture only a partial form of human-like bounded rationality and highlight the need for training methods that encourage flexible opponent modeling and stronger context awareness.

Figures

Figures reproduced from arXiv: 2506.09390 by the authors.

Figure 1
Figure 1. Distribution of Rock–Paper–Scissors choice proportions for (a) human players and (b) LLMs. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Ternary plots of strategy proportions following different outcomes for outcome-based agents. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Comparison of outcome-conditioned strategy distributions under WSLU vs. WDLS policies. (a) Human [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Outcome-conditioned strategy distributions, grouped by model class. (a) GPT [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Frontier LLMs playing an abstract AI-race game show extreme, model-specific policies, while human players are more diverse, so aggregate Unsafe rates alone are misleading.

  2. Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Frontier LLMs win Secret Hitler matches and can deceive, but most fail to keep a consistent false persona as evidence accumulates, with DRR often falling below 50%.

  3. When Identity Overrides Incentives: Representational Choices as Governance Decisions in Multi-Agent LLM Systems

    cs.MA 2026-01 unverdicted novelty 6.0 of 10

    Role-based personas in multi-agent LLM systems suppress payoff-aligned behavior, shifting equilibrium selection by up to 90 percentage points in Tragedy of the Commons versus Green Transition scenarios even with full ...

Reference graph

Works this paper leans on

54 extracted references · 22 canonical work pages · cited by 3 Pith papers

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    T. K. Ahn, Elinor Ostrom, David Schmidt, Robert Shupp, and James Walker. 2001. http://www.jstor.org/stable/30026189 Cooperation in pd games: Fear, greed, and history of play . Public Choice, 106(1/2):137--155

  4. [4]

    Oh, Matthias Bethge, and Eric Schulz

    Elif Akata, Lion Schulz, Jacopo Coda-Forno, Seong J. Oh, Matthias Bethge, and Eric Schulz. 2025. https://doi.org/10.1038/s41562-025-02172-y Playing repeated games with large language models . Nature Human Behaviour

  5. [5]

    Anthropic. 2024. https://www.anthropic.com/news/claude-3-5-sonnet Introducing claude 3.5 sonnet

  6. [6]

    Anthropic. 2025. https://www.anthropic.com/news/claude-3-7-sonnet Claude 3.7 sonnet and claude code

  7. [7]

    Federico Bianchi, Patrick John Chia, Mert Yuksekgonul, Jacopo Tagliabue, Dan Jurafsky, and James Zou. 2024. https://arxiv.org/abs/2402.05863 How well can llms negotiate? negotiationarena platform and analysis . Preprint, arXiv:2402.05863

  8. [8]

    Daniel Brookins and Richard DeBacker. 2023. https://ssrn.com/abstract=4493398 Playing games with gpt: What can we learn about a large language model from canonical strategic games? Technical Report 4493398, SSRN

Show all 54 references
  1. [9]

    Pedro Dal Bó. 2005. https://doi.org/10.1257/000282805775014434 Cooperation under the shadow of the future: Experimental evidence from infinitely repeated games . American Economic Review, 95(5):1591–1604

  2. [10]

    Colin F. Camerer. 1997. https://doi.org/10.1257/jep.11.4.167 Progress in behavioral game theory . Journal of Economic Perspectives, 11(4):167–188

  3. [11]

    Gary Charness and Matthew Rabin. 2002. http://www.jstor.org/stable/4132490 Understanding social preferences with simple tests . The Quarterly Journal of Economics, 117(3):817--869

  4. [12]

    Vincent Cheung, Michael Maier, and Falk Lieder. 2025. Large Language Models Amplify Human Biases in Moral Decision-Making . OSF Preprints. https://osf.io/preprints/psyarxiv/aj46b_v2

  5. [13]

    DeepSeek-AI. 2025 a . https://arxiv.org/abs/2501.12948 Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning . Preprint, arXiv:2501.12948

  6. [14]

    DeepSeek-AI. 2025 b . https://arxiv.org/abs/2412.19437 Deepseek-v3 technical report . Preprint, arXiv:2412.19437

  7. [15]

    Yuan Deng, Vahab Mirrokni, Renato Paes Leme, Hanrui Zhang, and Song Zuo. 2024. https://openreview.net/forum?id=n0RmqncQbU LLM s at the bargaining table . In Agentic Markets Workshop at ICML 2024

  8. [16]

    Dyson, Jonathan M

    Benjamin J. Dyson, Jonathan M. P. Wilbiks, Raj Sandhu, Georgios Papanicolaou, and Jaimie Lintag. 2016. https://doi.org/10.1038/srep20479 Negative outcomes evoke cyclic irrational decisions in rock, paper, scissors . Scientific Reports, 6:20479

  9. [17]

    Miller, Sasha Mitts, Adithya Renduchintala, Stephen Roller, Dirk Rowe, Weiyan Shi, Joe Spisak, Alexander Wei, David Wu, Hugh Zhang, and Markus Zijlstra

    Meta Fundamental AI Research Diplomacy Team (FAIR)†, Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis, Alexander...

  10. [18]

    Ernst Fehr and Simon Gächter. 2000. http://www.jstor.org/stable/117319 Cooperation and punishment in public goods experiments . The American Economic Review, 90(4):980--994

  11. [19]

    Ernst Fehr and Klaus M. Schmidt. 1999. http://www.jstor.org/stable/2586885 A theory of fairness, competition, and cooperation . The Quarterly Journal of Economics, 114(3):817--868

  12. [20]

    Shorrer, and Yannai A

    Sara Fish, Julia Shephard, Minkai Li, Ran I. Shorrer, and Yannai A. Gonczarowski. 2025. https://arxiv.org/abs/2503.18825 Econevals: Benchmarks and litmus tests for llm agents in unknown environments . Preprint, arXiv:2503.18825

  13. [21]

    Nicoló Fontana, Francesco Pierri, and Luca Maria Aiello. 2024. https://arxiv.org/abs/2406.13605 Nicer than humans: How do large language models behave in the prisoner's dilemma? Preprint, arXiv:2406.13605

  14. [22]

    Kanishk Gandhi, Dorsa Sadigh, and Noah D. Goodman. 2023. https://arxiv.org/abs/2305.19165 Strategic reasoning with language models . Preprint, arXiv:2305.19165

  15. [23]

    Yufei Guo, Muzhe Guo, Juntao Su, Zhou Yang, Mengqiu Zhu, Hongfei Li, Mengyang Qiu, and Shuo Shuo Liu. 2024. https://arxiv.org/abs/2411.10915 Bias in large language models: Origin, evaluation, and mitigation . Preprint, arXiv:2411.10915

  16. [24]

    Johannes Schneider, Steffi Haag and Leona Chandra Kruse. 2024. https://arxiv.org/abs/2312.03720 Negotiating with llms: Prompt hacks, skill gaps, and reasoning deficits . Preprint, arXiv:2312.03720

  17. [25]

    Tim Hagendorff, Stephanie Fabi, and Michal Kosinski. 2023. https://doi.org/10.1038/s43588-023-00527-x Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt . Nature Computational Science, 3:833--838

  18. [26]

    Hayes, Nicolas Yax, and Stefano Palminteri

    William M. Hayes, Nicolas Yax, and Stefano Palminteri. 2024. https://arxiv.org/abs/2405.11422 Large language models are biased reinforcement learners . Preprint, arXiv:2405.11422

  19. [27]

    Moshe Hoffman, Sigrid Suetens, Uri Gneezy, and Martin A. Nowak. 2015. https://doi.org/10.1038/srep08817 An experimental investigation of evolutionary dynamics in the Rock - Paper - Scissors game . Scientific Reports, 5(1):8817. Number: 1 Publisher: Nature Publishing Group

  20. [28]

    Hadi Hosseini and Samarth Khanna. 2025. https://arxiv.org/abs/2502.00313 Distributive fairness in large language models: Evaluating alignment with human values . Preprint, arXiv:2502.00313

  21. [29]

    Wenyue Hua, Ollie Liu, Lingyao Li, Alfonso Amayuelas, Julie Chen, Lucas Jiang, Mingyu Jin, Lizhou Fan, Fei Sun, William Wang, Xintong Wang, and Yongfeng Zhang. 2024. https://arxiv.org/abs/2411.05990 Game-theoretic llm: Agent workflow for negotiation games . Preprint, arXiv:2411.05990

  22. [30]

    McNamara, and Deming Chen

    Jingru Jia, Zehua Yuan, Junhao Pan, Paul E. McNamara, and Deming Chen. 2025. https://arxiv.org/abs/2502.20432 Large language model strategic reasoning evaluation through behavioral game theory . Preprint, arXiv:2502.20432

  23. [31]

    Daniel Kahneman and Amos Tversky. 1979. https://doi.org/10.2307/1914185 Prospect Theory: An Analysis of Decision under Risk . Econometrica, 47(2):263--291

  24. [32]

    Lucas, and Jonathan Gratch

    Deuksin Kwon, Emily Weiss, Tara Kulshrestha, Kushal Chawla, Gale M. Lucas, and Jonathan Gratch. 2024. https://arxiv.org/abs/2402.13550 Are llms effective negotiators? systematic evaluation of the multifaceted capabilities of llms in negotiation dialogues . Preprint, arXiv:2402.13550

  25. [33]

    Nunzio Lor\`e and Babak Heydari. 2024. https://doi.org/10.1038/s41598-024-69032-z Strategic behavior of large language models and the role of game structure versus contextual framing . Scientific Reports, 14(1):18490

  26. [34]

    Yougang Lyu, Shijie Ren, Yue Feng, Zihan Wang, Zhumin Chen, Zhaochun Ren, and Maarten de Rijke. 2025. https://arxiv.org/abs/2504.04141 Cognitive debiasing large language models for decision-making . Preprint, arXiv:2504.04141

  27. [35]

    Olivia Macmillan-Scott and Mirco Musolesi. 2024. https://arxiv.org/abs/2402.09193 (ir)rationality and cognitive biases in large language models . Preprint, arXiv:2402.09193

  28. [36]

    Qiaozhu Mei, Yutong Xie, Walter Yuan, and Matthew O. Jackson. 2024. https://doi.org/10.1073/pnas.2313925121 A turing test of whether ai chatbots are behaviorally similar to humans . Proceedings of the National Academy of Sciences, 121(9):e2313925121

  29. [37]

    Eladio Montero-Porras, Jelena Gruji \'c , Elias Fern \'a ndez-Domingos, and Tom Lenaerts. 2022. https://doi.org/10.1038/s41598-022-11654-2 Inferring strategies from observations in long iterated prisoner's dilemma experiments . Scientific Reports, 12:7589

  30. [38]

    Mikhail Mozikov, Nikita Severin, Valeria Bodishtianu, Maria Glushanina, Ivan Nasonov, Daniil Orekhov, Vladislav Pekhotin, Ivan Makovetskiy, Mikhail Baklashkin, Vasily Lavrentyev, Akim Tsvigun, Denis Turdakov, Tatiana Shavrina, Andrey Savchenko, and Ilya Makarov. 2024. https://...

  31. [39]

    Rosemarie Nagel. 1995. Unraveling in guessing games: An experimental study. American Economic Review, 85:1313--26

  32. [40]

    OpenAI. 2024 a . https://arxiv.org/abs/2303.08774 Gpt-4 technical report . Preprint, arXiv:2303.08774

  33. [41]

    OpenAI. 2024 b . o1 system card. System card, OpenAI. https://cdn.openai.com/o1-system-card.pdf

  34. [42]

    Matthew Rabin. 1993. http://www.jstor.org/stable/2117561 Incorporating fairness into game theory and economics . The American Economic Review, 83(5):1281--1302

  35. [43]

    Rand, Joshua D

    David G. Rand, Joshua D. Greene, and Martin A. Nowak. 2012. https://doi.org/10.1038/nature11467 Spontaneous giving and calculated greed . Nature, 489:427--430

  36. [44]

    Mark Schneider and Timothy Shields. 2022. https://doi.org/10.1080/15427560.2022.2081974 Motives for cooperation in the one-shot prisoner’s dilemma . Journal of Behavioral Finance, 23(4):438--456

  37. [45]

    Alonso Silva. 2024. https://arxiv.org/abs/2406.10574 Large language models playing mixed strategy nash equilibrium games . Preprint, arXiv:2406.10574

  38. [46]

    Stahl and Paul W

    Dale O. Stahl and Paul W. Wilson. 1995. https://doi.org/10.1006/game.1995.1031 On players' models of other players: Theory and experimental evidence . Games and Economic Behavior, 10(1):218--254

  39. [47]

    Elizaveta Tennant, Stephen Hailes, and Mirco Musolesi. 2025. https://arxiv.org/abs/2410.01639 Moral alignment for llm agents . Preprint, arXiv:2410.01639

  40. [48]

    Amos Tversky and Daniel Kahneman. 1974. https://doi.org/10.1126/science.185.4157.1124 Judgment under Uncertainty: Heuristics and Biases . Science, 185(4157):1124--1131

  41. [49]

    Zhijian Wang, Bin Xu, and Hai-Jun Zhou. 2014. https://doi.org/10.1038/srep05830 Social cycling and conditional responses in the rock-paper-scissors game . Scientific Reports, 4:5830

  42. [50]

    Wikipedia contributors . 2025. Game theory . https://en.wikipedia.org/wiki/Game_theory. Accessed: 20 May 2025

  43. [51]

    Tian Xia, Zhiwei He, Tong Ren, Yibo Miao, Zhuosheng Zhang, Yang Yang, and Rui Wang. 2024. https://arxiv.org/abs/2402.15813 Measuring bargaining abilities of llms: A benchmark and a buyer-enhancement method . Preprint, arXiv:2402.15813

  44. [52]

    Rong Ye, Yongxin Zhang, Yikai Zhang, Haoyu Kuang, Zhongyu Wei, and Peng Sun. 2025. https://arxiv.org/abs/2501.14225 Multi-agent kto: Reinforcing strategic interactions of large language model in language game . Preprint, arXiv:2501.14225

  45. [53]

    Hanshu Zhang, Frederic Moisan, and Cleotilde Gonzalez. 2021. https://doi.org/10.3390/g12030052 Rock-paper-scissors play: Beyond the win-stay/lose-change strategy . Games, 12(3)

  46. [54]

    Jinfeng Zhou, Yuxuan Chen, Yihan Shi, Xuanming Zhang, Leqi Lei, Yi Feng, Zexuan Xiong, Miao Yan, Xunzhi Wang, Yaru Cao, Jianing Yin, Shuai Wang, Quanyu Dai, Zhenhua Dong, Hongning Wang, and Minlie Huang. 2025. https://arxiv.org/abs/2506.00900 Socialeval: Evaluating social inte...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.