REVIEW 12 cited by
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Decision-making is a complex process requiring diverse abilities, making it an excellent framework for evaluating Large Language Models (LLMs). Researchers have examined LLMs' decision-making through the lens of Game Theory. However, existing evaluation mainly focus on two-player scenarios where an LLM competes against another. Additionally, previous benchmarks suffer from test set leakage due to their static design. We introduce GAMA($\gamma$)-Bench, a new framework for evaluating LLMs' Gaming Ability in Multi-Agent environments. It includes eight classical game theory scenarios and a dynamic scoring scheme specially designed to quantitatively assess LLMs' performance. $\gamma$-Bench allows flexible game settings and adapts the scoring system to different game parameters, enabling comprehensive evaluation of robustness, generalizability, and strategies for improvement. Our results indicate that GPT-3.5 demonstrates strong robustness but limited generalizability, which can be enhanced using methods like Chain-of-Thought. We also evaluate 13 LLMs from 6 model families, including GPT-3.5, GPT-4, Gemini, LLaMA-3.1, Mixtral, and Qwen-2. Gemini-1.5-Pro outperforms others, scoring of $69.8$ out of $100$, followed by LLaMA-3.1-70B ($65.9$) and Mixtral-8x22B ($62.4$). Our code and experimental results are publicly available at https://github.com/CUHK-ARISE/GAMABench.
Forward citations
Cited by 12 Pith papers
-
Verbalized Bayesian Persuasion
VBP solves Bayesian persuasion in natural language by treating LLMs as sender and receiver in a mediator-augmented game and searching prompt strategies with Prompt-PSRO.
-
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
Frontier LLMs exceed human performance on easy Wikipedia navigation tasks but finish fewer than 25% of hard games, with failures driven by looping and an inability to replan after mistakes.
-
When Identity Overrides Incentives: Representational Choices as Governance Decisions in Multi-Agent LLM Systems
Role-based personas in multi-agent LLM systems suppress payoff-aligned behavior, shifting equilibrium selection by up to 90 percentage points in Tragedy of the Commons versus Green Transition scenarios even with full ...
-
The Effect of State Representation on LLM Agent Behavior in Dynamic Routing Games
In a repeated Braess routing game, LLM agents given summarized, regret-based, and own-action-only state representations converge closer to Nash equilibrium and behave more stably than agents given full chat transcript...
-
PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to Histopathology
PathFinder, a multi-agent system that iteratively navigates and describes histopathology slides, reports 74% accuracy on a small balanced melanoma test set, topping a 65% average human benchmark.
-
A MARL Centered Reference Architecture for Large Language Model Augmentation in Smart Manufacturing
The paper proposes a MARL-centered three-layer reference architecture for LLM augmentation in smart manufacturing, with a conditional allocation: MARL for frequent coordination, LLMs for semantic, reward, and planning roles.
-
Ethical Considerations of Large Language Models in Game Playing
In Werewolf games, LLM agents change their kills, votes, and trust scores based on explicit gender labels and even based on gender-implied first names, behaving differently for male and female players.
-
Can LLMs Play \^O \u{A}n Quan Game? A Study of Multi-Step Planning and Decision Making
In 50-game matches, an 8-billion-parameter Llama beat a 70-billion-parameter Llama more often than it lost, while larger models generated longer reasoning traces but not reliably better scores.
-
Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers
A taxonomy-based survey of bidirectional game theory and LLM research, spanning evaluation, alignment, economic competition, and LLM-driven game solving.
-
EvalSVA: Multi-Agent Evaluators for Next-Gen Software Vulnerability Assessment
A multi-agent LLM framework with four communication strategies improves CVSS v3.1 vulnerability assessment over single-agent baselines on a new commit dataset spanning C++, Python, and Java.
-
Toward LLM-Agent-Based Modeling of Transportation Systems: A Conceptual Framework
LLM-driven agents with profiles, memory, and feedback loops can generate plausible daily travel activities and learn to adjust commute timing in a small proof-of-concept, pointing toward a new direction for agent-base...
-
A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios
LLM-based game-playing agents are surveyed across choice-focused and communication-focused games, with a comparative performance table and future directions.
Discussion (0). Continue with ORCID to comment.