Pith. sign in

REVIEW 7 cited by

Playing repeated games with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.16867 v2 pith:H3DJLGBF submitted 2023-05-26 cs.CL

classification cs.CL
keywords gamesbehaviourbehaviouralcoordinationllmsgamehumanlike
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

LLMs are increasingly used in applications where they interact with humans and other agents. We propose to use behavioural game theory to study LLM's cooperation and coordination behaviour. We let different LLMs play finitely repeated $2\times2$ games with each other, with human-like strategies, and actual human players. Our results show that LLMs perform particularly well at self-interested games like the iterated Prisoner's Dilemma family. However, they behave sub-optimally in games that require coordination, like the Battle of the Sexes. We verify that these behavioural signatures are stable across robustness checks. We additionally show how GPT-4's behaviour can be modulated by providing additional information about its opponent and by using a "social chain-of-thought" (SCoT) strategy. This also leads to better scores and more successful coordination when interacting with human players. These results enrich our understanding of LLM's social behaviour and pave the way for a behavioural game theory for machines.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  2. How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    LLM agents in a spatial Prisoner's Dilemma exhibit model-specific effects of memory length on cooperation, with Gemini suppressing and Gemma promoting it as memory increases.

  3. From Digital Distrust to Codified Honesty: Experimental Evidence on Generative AI in Credence Goods Markets

    econ.GN 2025-09 conditional novelty 6.0 of 10

    LLM experts in credence goods markets reduce efficiency and consumer surplus unless liability or transparent prosocial objectives operate, and expert delegation with transparent objectives can outperform human-only markets.

  4. Super-additive Cooperation in Language Model Agents

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Language model agents cooperate more in a prisoner's dilemma when repeated interactions and inter-team competition are combined, but only for some models.

  5. The Effect of State Representation on LLM Agent Behavior in Dynamic Routing Games

    cs.AI 2025-06 conditional novelty 6.0 of 10

    In a repeated Braess routing game, LLM agents given summarized, regret-based, and own-action-only state representations converge closer to Nash equilibrium and behave more stably than agents given full chat transcript...

  6. Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making

    cs.AI 2025-06 conditional novelty 4.0 of 10

    In a new three-game benchmark tracking planning, revision, and budget use across 12 LLMs, ChatGPT-o3-mini ranked highest, while overcorrecting models such as Qwen-Plus won few matches.

  7. AI Agent Behavioral Science

    q-bio.NC 2025-06 conditional novelty 4.0 of 10

    AI agents should be studied as behavioral entities shaped by context and interaction, not only as trained models.

Pith tools