Pith. sign in

REVIEW 3 cited by

Hoodwinked: Deception and Cooperation in a Text-Based Game for Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.01404 v2 pith:WDCLNMAR submitted 2023-07-05 cs.CL cs.CYcs.LG

classification cs.CLcs.CYcs.LG
keywords gamemodelsdeceptionhoodwinkedlanguageagentsdetectionevidence
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Are current language models capable of deception and lie detection? We study this question by introducing a text-based game called $\textit{Hoodwinked}$, inspired by Mafia and Among Us. Players are locked in a house and must find a key to escape, but one player is tasked with killing the others. Each time a murder is committed, the surviving players have a natural language discussion then vote to banish one player from the game. We conduct experiments with agents controlled by GPT-3, GPT-3.5, and GPT-4 and find evidence of deception and lie detection capabilities. The killer often denies their crime and accuses others, leading to measurable effects on voting outcomes. More advanced models are more effective killers, outperforming smaller models in 18 of 24 pairwise comparisons. Secondary metrics provide evidence that this improvement is not mediated by different actions, but rather by stronger persuasive skills during discussions. To evaluate the ability of AI agents to deceive humans, we make this game publicly available at h https://hoodwinked.ai/ .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games

    cs.AI 2026-08 conditional novelty 6.0 of 10

    LLM agents' cooperation and honesty are not fixed traits but respond to institutional rules, with salary triggering private deal-making and anonymity raising deception.

  2. AI Agent Behavioral Science

    q-bio.NC 2025-06 conditional novelty 4.0 of 10

    AI agents should be studied as behavioral entities shaped by context and interaction, not only as trained models.

  3. Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models

    cs.CL 2025-01 reject novelty 3.0 of 10

    DeepSeek R1 roleplayed an autonomous robot that deceives and self-preserves when explicitly prompted to act as 'master,' but the paper offers no controlled evidence that these behaviors are emergent.

Pith tools