REVIEW 10 cited by
LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper explores the open research problem of understanding the social behaviors of LLM-based agents. Using Avalon as a testbed, we employ system prompts to guide LLM agents in gameplay. While previous studies have touched on gameplay with LLM agents, research on their social behaviors is lacking. We propose a novel framework, tailored for Avalon, features a multi-agent system facilitating efficient communication and interaction. We evaluate its performance based on game success and analyze LLM agents' social behaviors. Results affirm the framework's effectiveness in creating adaptive agents and suggest LLM-based agents' potential in navigating dynamic social interactions. By examining collaboration and confrontation behaviors, we offer insights into this field's research and applications. Our code is publicly available at https://github.com/3DAgentWorld/LLM-Game-Agent.
Forward citations
Cited by 10 Pith papers
-
Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments
A new 21-game benchmark scores LLM social reasoning with rule-decided outcomes, and SPaRTan, a self-reflection loop, transfers playbook lessons across games.
-
The Effect of State Representation on LLM Agent Behavior in Dynamic Routing Games
In a repeated Braess routing game, LLM agents given summarized, regret-based, and own-action-only state representations converge closer to Nash equilibrium and behave more stably than agents given full chat transcript...
-
Ethical Considerations of Large Language Models in Game Playing
In Werewolf games, LLM agents change their kills, votes, and trust scores based on explicit gender labels and even based on gender-implied first names, behaving differently for male and female players.
-
Can LLMs Play \^O \u{A}n Quan Game? A Study of Multi-Step Planning and Decision Making
In 50-game matches, an 8-billion-parameter Llama beat a 70-billion-parameter Llama more often than it lost, while larger models generated longer reasoning traces but not reliably better scores.
-
Who Speaks Next? Multi-party AI Discussion Leveraging the Systematics of Turn-taking in Murder Mystery Games
Letting the current speaker select the next speaker through adjacency pair rules makes AI agents' group conversations in a murder mystery game more coherent and cooperative.
-
SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models
A self-evolving, training-free VLN agent with hierarchical memory, RAG plus chain-of-thought reasoning, and reflection reports state-of-the-art success rates on R2R and REVERIE.
-
AI Agent Behavioral Science
AI agents should be studied as behavioral entities shaped by context and interaction, not only as trained models.
-
A Survey on LLM-based Multi-Agent System: Recent Advances and New Frontiers in Application
This survey organizes recent LLM-based multi-agent research into task-solving, simulation, and agent-evaluation applications, and identifies efficiency and evaluation gaps as key open problems.
-
From Individual to Society: A Survey on Social Simulation Driven by Large Language Model-based Agents
A structured survey that categorizes LLM-based social simulation into individual, scenario, and society simulation, with associated methods, benchmarks, and observed trends.
-
A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios
LLM-based game-playing agents are surveyed across choice-focused and communication-focused games, with a comparative performance table and future directions.
Discussion (0). Continue with ORCID to comment.