Pith. sign in

REVIEW 25 cited by

Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.04658 v2 pith:7X65DZZ6 submitted 2023-09-09 cs.CL

Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

classification cs.CL
keywords communicationgamesllmslanguagewerewolfempiricalengageframework
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Communication games, which we refer to as incomplete information games that heavily depend on natural language communication, hold significant research value in fields such as economics, social science, and artificial intelligence. In this work, we explore the problem of how to engage large language models (LLMs) in communication games, and in response, propose a tuning-free framework. Our approach keeps LLMs frozen, and relies on the retrieval and reflection on past communications and experiences for improvement. An empirical study on the representative and widely-studied communication game, ``Werewolf'', demonstrates that our framework can effectively play Werewolf game without tuning the parameters of the LLMs. More importantly, strategic behaviors begin to emerge in our experiments, suggesting that it will be a fruitful journey to engage LLMs in communication games and associated domains.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  2. Voluntary Collusion with Secret Tools in Competing LLM Agents

    cs.AI 2026-05 unverdicted novelty 7.0

    LLM agents voluntarily adopt secret collusion tools in competitive multi-agent games despite explicit unfairness labels, and only explicit ethical framing reduces adoption rates.

  3. Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance

    nlin.AO 2026-05 unverdicted novelty 7.0

    LLM societies in Nomic show non-monotonic collective adaptation peaking at mid-scales, with smaller models rule-inert and larger ones restrictive.

  4. Learning to Interrupt in Language-based Multi-agent Communication

    cs.CL 2026-04 unverdicted novelty 7.0

    HANDRAISER learns optimal interruption points in multi-agent LLM communication using estimated future reward and cost, achieving 32.2% lower communication cost with comparable or better task results across games, sche...

  5. Bayesian Social Deduction with Graph-Informed Language Models

    cs.AI 2025-06 unverdicted novelty 7.0

    Hybrid Bayesian-graph LLM agent reaches competitive performance against large models and achieves 67% win rate against humans in controlled Avalon play, outperforming baselines and human teammates.

  6. MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

    cs.CL 2025-05 unverdicted novelty 7.0

    MTR-Bench is a new automated benchmark for multi-turn reasoning in LLMs covering diverse tasks and difficulty levels with 3600 instances.

  7. CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games

    cs.AI 2026-07 conditional novelty 6.0

    CaM-Wolf is a multimodal Werewolf agent that perceives player video, reasons about hidden roles with a counterfactual-intervention-trained RL reasoner, and responds through an animated avatar.

  8. Reproducing human biases in route choice using large language models: Toward scalable behavioral modeling

    cs.AI 2026-07 conditional novelty 6.0

    LLM agents with demographic profiles reproduce CPT-style risk attitudes in route choice and yield fitted parameters (α=0.4, β=0.64, λ=1.43) that predict human data competitively.

  9. From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory

    cs.CL 2026-06 unverdicted novelty 6.0

    MemoPilot trains memory updates for LLM agents via multi-turn GRPO on RPS and poker, achieving top Elo scores and outperforming baselines including DeepSeek-V3.2.

  10. OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces

    cs.AI 2026-05 unverdicted novelty 6.0

    OPT-BENCH and OPT-Agent evaluate LLM self-optimization in large search spaces, showing stronger models improve via feedback but stay constrained by base capacity and below human performance.

  11. Explicit Trait Inference for Multi-Agent Coordination

    cs.AI 2026-04 unverdicted novelty 6.0

    ETI lets LLM agents infer and track partners' psychological traits (warmth and competence) from histories, cutting payoff loss 45-77% in games and boosting performance 3-29% on MultiAgentBench versus CoT baselines.

  12. SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems

    cs.AI 2026-04 unverdicted novelty 6.0

    SocialGrid benchmark shows even top LLMs achieve below 60% in embodied planning and task completion, with deception detection near random chance regardless of model scale.

  13. Beyond Value Elicitation: Towards Moral Profiles in Early Requirements Engineering via Role-Playing Games and Anthropologist LLMs

    cs.HC 2025-08 unverdicted novelty 6.0

    A proof-of-concept combines role-playing games with an anthropologist LLM to generate narrative moral profiles from users' situated decisions for early requirements engineering.

  14. Exploring a Gamified Personality Assessment Method through Interaction with LLM Agents Embodying Different Personalities

    cs.HC 2025-07 unverdicted novelty 6.0

    A gamified system with multiple LLM agents of varied personalities gathers interaction data to produce more effective and interpretable Big Five personality assessments than single-context methods.

  15. AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society

    cs.SI 2025-02 unverdicted novelty 6.0

    AgentSociety is a large-scale LLM agent-based social simulator validated on polarization, UBI, disasters, and sustainability issues with alignment to real experiments.

  16. Cumulative suspicion and absorption dynamics in an agent-based Mafia game

    physics.soc-ph 2026-07 conditional novelty 5.5

    History-dependent suspicion scores in an agent Mafia model yield F(τ)∼(τ/N)^{N_m} early extinction without detectives and an empirical N_c collapse of win probabilities that detectives break.

  17. Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior

    cs.RO 2026-05 unverdicted novelty 5.0

    LLM agents in a collaborative 2D game exhibit emergent behaviors such as perspective-taking, theory of mind, and clarification, detected by LLM judges and rated positively by human participants.

  18. Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior

    cs.RO 2026-05 unverdicted novelty 5.0

    Embodied LLM agents exhibit emergent collaborative behaviors indicating mental models of partners in a color-matching game, detected via LLM judges and supported by positive user feedback.

  19. Evaluating Large Language Models in a Complex Hidden Role Game

    cs.CL 2026-04 unverdicted novelty 5.0

    LLMs achieve only 59.7% role identification accuracy in Secret Hitler versus 86.7% for rule-based agents, show negative impact as fascists, and produce 40% shorter games due to failed deception.

  20. From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments

    cs.AI 2026-03 unverdicted novelty 5.0

    An empirical literature analysis reveals a bifurcation in RL environments into Semantic Prior (LLM-dominated) and Domain-Specific Generalization ecosystems with distinct cognitive fingerprints.

  21. Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization

    cs.CL 2025-11 unverdicted novelty 5.0

    MAGIC-HMO is a multi-agent framework that treats Chinese short-form creative NLG as heterogeneous multi-objective optimization over personalized constraints plus explanation reliability and outperforms baselines on a ...

  22. AppAgent: Multimodal Agents as Smartphone Users

    cs.CV 2023-12 unverdicted novelty 5.0

    AppAgent lets large language models operate diverse smartphone apps via visual interactions and learns app usage from exploration or demonstrations.

  23. Distilling Game Code World Model Generation into Lightweight Large Language Models

    cs.AI 2026-05 unverdicted novelty 4.0

    SFT followed by RLVR on Qwen2.5-3B-Instruct raises syntactic and execution correctness when generating Game Code World Models across 30 games.

  24. SOM: Structured Opponent Modeling for LLM-based Agents via Structural Causal Model

    cs.AI 2026-05 unverdicted novelty 4.0

    SOM uses a Structural Causal Model to create an explicit graph of opponent observation-to-action links, allowing LLMs to reason along those paths for more accurate and stable predictions in multi-agent settings.

  25. Large Language Model based Multi-Agents: A Survey of Progress and Challenges

    cs.CL 2024-01 unverdicted novelty 4.0

    The paper surveys LLM-based multi-agent systems, covering simulated domains, agent profiling and communication, mechanisms for capacity growth, and common benchmarks.