Pith. sign in

REVIEW 4 cited by

AMONGAGENTS: Evaluating Large Language Models in the Interactive Text-Based Social Deduction Game

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.16521 v2 pith:QR6BLCTC submitted 2024-07-23 cs.CL

AMONGAGENTS: Evaluating Large Language Models in the Interactive Text-Based Social Deduction Game

classification cs.CL
keywords languagegamebehaviormodelssimulatedsocialamongagentscrew
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Strategic social deduction games serve as valuable testbeds for evaluating the understanding and inference skills of language models, offering crucial insights into social science, artificial intelligence, and strategic gaming. This paper focuses on creating proxies of human behavior in simulated environments, with Among Us utilized as a tool for studying simulated human behavior. The study introduces a text-based game environment, named AmongAgents, that mirrors the dynamics of Among Us. Players act as crew members aboard a spaceship, tasked with identifying impostors who are sabotaging the ship and eliminating the crew. Within this environment, the behavior of simulated language agents is analyzed. The experiments involve diverse game sequences featuring different configurations of Crewmates and Impostor personality archetypes. Our work demonstrates that state-of-the-art large language models (LLMs) can effectively grasp the game rules and make decisions based on the current context. This work aims to promote further exploration of LLMs in goal-oriented games with incomplete information and complex action spaces, as these settings offer valuable opportunities to assess language model performance in socially driven scenarios.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games

    cs.CL 2026-07 conditional novelty 7.0

    Non-invasive per-utterance belief probes in Mafia, auto-scored against engine truth, expose poorly calibrated LLM confidence and 1.5× over-prediction of being suspected.

  2. Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

    cs.AI 2026-07 conditional novelty 6.0

    Changing one LLM agent's secret objective in Werewolf lowers its team's win rate and changes its reasoning, while its public chat stays deceptively normal.

  3. MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games

    cs.CL 2026-07 conditional novelty 6.0

    A released open-source testbed that privately probes LLM agents' beliefs during Mafia games shows agents are overconfident, over-predict suspicion by 1.5x, and rarely change outcomes when wrong beliefs lock in a vote.

  4. SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems

    cs.AI 2026-04 unverdicted novelty 6.0

    SocialGrid benchmark shows even top LLMs achieve below 60% in embodied planning and task completion, with deception detection near random chance regardless of model scale.