Pith. sign in

REVIEW 5 cited by

Strategic Behavior of Large Language Models: Game Structure vs. Contextual Framing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.05898 v1 pith:M2MLDZJG submitted 2023-09-12 cs.GT cs.AIcs.CYcs.HCecon.TH

classification cs.GTcs.AIcs.CYcs.HCecon.TH
keywords strategicmodelscontextualframinggamellama-2complexdecision-making
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper investigates the strategic decision-making capabilities of three Large Language Models (LLMs): GPT-3.5, GPT-4, and LLaMa-2, within the framework of game theory. Utilizing four canonical two-player games -- Prisoner's Dilemma, Stag Hunt, Snowdrift, and Prisoner's Delight -- we explore how these models navigate social dilemmas, situations where players can either cooperate for a collective benefit or defect for individual gain. Crucially, we extend our analysis to examine the role of contextual framing, such as diplomatic relations or casual friendships, in shaping the models' decisions. Our findings reveal a complex landscape: while GPT-3.5 is highly sensitive to contextual framing, it shows limited ability to engage in abstract strategic reasoning. Both GPT-4 and LLaMa-2 adjust their strategies based on game structure and context, but LLaMa-2 exhibits a more nuanced understanding of the games' underlying mechanics. These results highlight the current limitations and varied proficiencies of LLMs in strategic decision-making, cautioning against their unqualified use in tasks requiring complex strategic reasoning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  2. Position Auctions in AI-Generated Content

    cs.GT 2025-06 conditional novelty 6.0 of 10

    New mechanism-design results for position auctions with context-dependent click-through rates under multinomial logit and cascade user models, with exact optimality in the MNL case and an O(log m) approximation in the...

  3. When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Across morally framed prisoner's dilemmas and public goods games, none of nine LLMs consistently chooses the ethical action when it conflicts with payoff, with cooperation rates from 7.9% to 76.3%.

  4. Shapley-Coop: Credit Assignment for Emergent Cooperation in Self-Interested LLM Agents

    cs.MA 2025-06 reject novelty 5.0 of 10

    Shapley-Coop asks LLM agents to negotiate prices for contributions based on Shapley value reasoning, improving cooperation and reward fairness in three multi-agent tasks.

  5. Tracing LLM Reasoning Processes with Strategic Games: A Framework for Planning, Revision, and Resource-Constrained Decision Making

    cs.AI 2025-06 conditional novelty 4.0 of 10

    In a new three-game benchmark tracking planning, revision, and budget use across 12 LLMs, ChatGPT-o3-mini ranked highest, while overcorrecting models such as Qwen-Plus won few matches.

Pith tools