Pith. sign in

REVIEW 12 cited by

How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.11807 v7 pith:4VWDYPCP submitted 2024-03-18 cs.AI cs.CL

classification cs.AIcs.CL
keywords llmsgamedecision-makingevaluatingscoringabilitybenchenvironments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Decision-making is a complex process requiring diverse abilities, making it an excellent framework for evaluating Large Language Models (LLMs). Researchers have examined LLMs' decision-making through the lens of Game Theory. However, existing evaluation mainly focus on two-player scenarios where an LLM competes against another. Additionally, previous benchmarks suffer from test set leakage due to their static design. We introduce GAMA($\gamma$)-Bench, a new framework for evaluating LLMs' Gaming Ability in Multi-Agent environments. It includes eight classical game theory scenarios and a dynamic scoring scheme specially designed to quantitatively assess LLMs' performance. $\gamma$-Bench allows flexible game settings and adapts the scoring system to different game parameters, enabling comprehensive evaluation of robustness, generalizability, and strategies for improvement. Our results indicate that GPT-3.5 demonstrates strong robustness but limited generalizability, which can be enhanced using methods like Chain-of-Thought. We also evaluate 13 LLMs from 6 model families, including GPT-3.5, GPT-4, Gemini, LLaMA-3.1, Mixtral, and Qwen-2. Gemini-1.5-Pro outperforms others, scoring of $69.8$ out of $100$, followed by LLaMA-3.1-70B ($65.9$) and Mixtral-8x22B ($62.4$). Our code and experimental results are publicly available at https://github.com/CUHK-ARISE/GAMABench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Verbalized Bayesian Persuasion

    cs.GT 2025-02 conditional novelty 7.0 of 10

    VBP solves Bayesian persuasion in natural language by treating LLMs as sender and receiver in a mediator-augmented game and searching prompt strategies with Prompt-PSRO.

  2. LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?

    cs.AI 2026-02 conditional novelty 6.0 of 10

    Frontier LLMs exceed human performance on easy Wikipedia navigation tasks but finish fewer than 25% of hard games, with failures driven by looping and an inability to replan after mistakes.

  3. When Identity Overrides Incentives: Representational Choices as Governance Decisions in Multi-Agent LLM Systems

    cs.MA 2026-01 unverdicted novelty 6.0 of 10

    Role-based personas in multi-agent LLM systems suppress payoff-aligned behavior, shifting equilibrium selection by up to 90 percentage points in Tragedy of the Commons versus Green Transition scenarios even with full ...

  4. The Effect of State Representation on LLM Agent Behavior in Dynamic Routing Games

    cs.AI 2025-06 conditional novelty 6.0 of 10

    In a repeated Braess routing game, LLM agents given summarized, regret-based, and own-action-only state representations converge closer to Nash equilibrium and behave more stably than agents given full chat transcript...

  5. PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to Histopathology

    cs.CV 2025-02 conditional novelty 6.0 of 10

    PathFinder, a multi-agent system that iteratively navigates and describes histopathology slides, reports 74% accuracy on a small balanced melanoma test set, topping a 65% average human benchmark.

  6. A MARL Centered Reference Architecture for Large Language Model Augmentation in Smart Manufacturing

    cs.AI 2026-08 conditional novelty 5.0 of 10

    The paper proposes a MARL-centered three-layer reference architecture for LLM augmentation in smart manufacturing, with a conditional allocation: MARL for frequent coordination, LLMs for semantic, reward, and planning roles.

  7. Ethical Considerations of Large Language Models in Game Playing

    cs.CL 2025-08 conditional novelty 5.0 of 10

    In Werewolf games, LLM agents change their kills, votes, and trust scores based on explicit gender labels and even based on gender-implied first names, behaving differently for male and female players.

  8. Can LLMs Play \^O \u{A}n Quan Game? A Study of Multi-Step Planning and Decision Making

    cs.CL 2025-07 conditional novelty 5.0 of 10

    In 50-game matches, an 8-billion-parameter Llama beat a 70-billion-parameter Llama more often than it lost, while larger models generated longer reasoning traces but not reliably better scores.

  9. Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers

    cs.AI 2025-02 conditional novelty 5.0 of 10

    A taxonomy-based survey of bidirectional game theory and LLM research, spanning evaluation, alignment, economic competition, and LLM-driven game solving.

  10. EvalSVA: Multi-Agent Evaluators for Next-Gen Software Vulnerability Assessment

    cs.SE 2024-12 conditional novelty 5.0 of 10

    A multi-agent LLM framework with four communication strategies improves CVSS v3.1 vulnerability assessment over single-agent baselines on a new commit dataset spanning C++, Python, and Java.

  11. Toward LLM-Agent-Based Modeling of Transportation Systems: A Conceptual Framework

    cs.AI 2024-12 conditional novelty 5.0 of 10

    LLM-driven agents with profiles, memory, and feedback loops can generate plausible daily travel activities and learn to adjust commute timing in a small proof-of-concept, pointing toward a new direction for agent-base...

  12. A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios

    cs.CL 2024-12 conditional novelty 3.0 of 10

    LLM-based game-playing agents are surveyed across choice-focused and communication-focused games, with a comparative performance table and future directions.

Pith tools