REVIEW 4 major objections 6 minor 14 references
Frontier language models can win Secret Hitler as both liberals and fascists, but most still leak their cover before the game ends.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 16:25 UTC pith:NX5EGT5Z
load-bearing objection Solid open Secret Hitler benchmark with real scale and useful DRR/vote findings; the abstract’s “top-four” win-rate ladder is mostly a single-opponent artifact the appendices already undercut. the 4 major comments →
Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Across roughly 1,600 Secret Hitler matches, frontier models form a strong top cluster that wins most games as Liberals, Fascists, and Hitler, while weaker models underperform algorithmic and sometimes random baselines; at the same time, most models cannot sustain a consistent deceptive persona, with deception retention often falling substantially over successive rounds.
What carries the argument
ParliamentBench: an open multi-agent Secret Hitler environment plus three round-level metrics—Game-State Impact Rate (whether actions help or hurt one's faction), Role Identification Accuracy (private role guesses), and Deception Retention Rate (how long liberals misidentify or fail to pin a fascist/Hitler).
Load-bearing premise
The central claim treats fixed-turn Secret Hitler play, hand-tuned game-state scoring, private role probes, and a tiny human pilot as a faithful enough proxy for real agent deception and reasoning under information asymmetry.
What would settle it
Re-run the same models with varied opponents, free-form turn-taking, and a larger blinded human study: if win-rate order, role-identification scores, and deception retention no longer track, or if humans reliably unmask 'high-DRR' models early, the metrics are not isolating transferable deceptive capability.
If this is right
- Safety evaluations that only report win rate will miss models that deceive early and then leak identity later.
- Social deduction (naming others' roles) and converting that knowledge into votes and policy choices are separate skills and should be scored separately.
- Weaker models' near-constant yes-voting is a concrete failure mode that collapses late-game defense even when chat sounds plausible.
- Neutral rewrites of loaded game terms can change fascist-side success without changing overall win rate, so terminology is part of measured strategy.
- An open Secret Hitler harness with GSIR, RIA, and DRR gives a reproducible testbed for persuasion and hidden-objective agents before high-stakes deployment.
Where Pith is reading between the lines
- If long-horizon cover is the scarce skill, red-team tests should stress multi-round consistency under accumulating evidence rather than single-turn lie generation.
- The agreeableness failure in small models suggests safety-relevant sycophancy may show up as strategic voting collapse, not only as flattering chat.
- Opponent-stable metrics matter more than leaderboard win rate for comparing agents across labs and model families.
- A larger human detection study is the natural next gate before treating high machine DRR as evidence of human-facing deception risk.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ParliamentBench, an open-source multi-agent evaluation framework based on five-player Secret Hitler, and reports results for 16 LLMs over roughly 1,600 simulated matches (primarily 100 games per model against Llama 3.3 70B), plus a small human pilot and comparison to ~25,000 anonymized online human games. It defines three round-level metrics—Game-State Impact Rate (GSIR), Role Identification Accuracy (RIA), and Deception Retention Rate (DRR)—intended to separate policy/strategic impact, social deduction, and long-horizon deceptive cover from aggregate win rate. Main empirical claims are that a frontier top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, DeepSeek 3.1 Terminus) wins a clear majority of games while weaker models fall below algorithmic (~45%) and sometimes random (~33%) baselines; that liberal play is harder than fascist play; that most models’ DRR degrades substantially over rounds (often toward or below ~50%); and that social deduction (RIA) and effective action (win rate/GSIR) are only loosely coupled. Supporting analyses include bootstrap CIs and pairwise tests (Appendix G), an anchor–opponent tournament (Appendix H), a reasoning-off ablation, uncensored-model and loaded-vocabulary probes, and vote-accuracy/approval trajectories.
Significance. If the reported patterns hold under broader opponent pools, this is a useful contribution to LLM agent evaluation: Secret Hitler’s multi-hop legislative bluff is a richer deception testbed than pure win-rate Werewolf/Avalon setups, the open environment and metric pipelines are reusable, and the DRR decay result is a concrete, falsifiable observation about long-horizon persona consistency. Strengths that should be credited include scale relative to prior Secret Hitler LLM work, role-stratified reporting, random/algorithmic/human baselines, statistical appendices (G–H), and the vocabulary and reasoning ablations. The safety motivation (high-stakes deception) is appropriately framed as proxy evidence rather than direct transfer. The work’s lasting value is more the benchmark and the fine-grained behavioral metrics than any single win-rate ladder.
major comments (4)
- [Abstract, §4.1, Table 1, Appendix H] Abstract, §4.1, and Table 1 lead with a “strong top-four cluster” defined by overall win rate in a fixed design: 100 games per model almost entirely against Llama 3.3 70B agents. Appendix H’s own anchor–opponent tournament (450 games) shows that win-rate order is opponent-dependent (preserved vs Gemma, τ=+1.0; fully reversed vs GPT-OSS 120B and Mistral, τ=−1.0), while RIA/DRR/GSIR are more stable (mean τ +0.55/+0.33/+0.11 vs −0.33 for WR). The manuscript notes this in the appendix and text, but the abstract, sort key of Table 1, and headline cluster language still treat single-opponent WR as the primary empirical result. This is load-bearing for the paper’s main capability ordering. Please re-center the abstract and §4.1 on opponent-stable metrics (DRR trajectories, vote accuracy, RIA, GSIR) and present the WR cluster explicitly as Llama-opponent-conditioned, or expand the main evaluatio
- [Abstract; §4.3; Limitations] The abstract states evaluation “across ≈1,600 simulated matches playing each other, playing against humans,” which overstates the design. The primary 16-model comparison is not a full round-robin among the 16 models; it is largely each model vs Llama 3.3 70B, plus a limited three-anchor tournament and a five-game human pilot (n=5 games, four participants; §4.3). Human results are preliminary and cannot support claims that frontier models “can maintain deceptive cover against human opponents” at more than anecdotal strength. Tighten the abstract and contribution bullets to match the actual experimental graph, and move human findings to clearly labeled pilot status in both abstract and conclusion.
- [§3.4; Appendix B; Figure 4; Table 1] GSIR is central to the claim that the metrics “isolate … reasoning” (§3.4, §4.1, Figure 4). Appendix B states that twelve constants were hand-tuned on ten example games by a single rater to match expected game-state values, with no inter-rater agreement, then stress-tested by weight perturbation (Spearman ρ high). GSIR also folds in a role-accuracy term related to the same belief probes used for RIA/DRR, so it is not a pure external state evaluator. The paper already reports weaker GSIR–win correlations for deceptive roles (r=0.38 fascist) and negative fascist GSIR alongside high fascist win rates—an important tension. Either (i) demote GSIR from a primary “reasoning” isolator to a structured action-decomposition with explicit caveats in the main text, or (ii) validate constants on held-out human games / alternative weight-learning and report sensitivity of the main scientific claims (no
- [§3.4; Figure 2; Table 3; Appendix L.7] Private post-round role questionnaires define both RIA and DRR (Appendix L.7; Eqs. related in Appendix B). Treating “Unknown” as successful concealment for DRR is a modeling choice that substantially affects scores (Table 3 decomposes Active vs Ambiguity). The paper should justify why opponent belief probes under a fixed prompt equal “deceptive consistency” rather than opponent conservatism or prompt-induced abstention, and report DRR with and without counting Unknown as success in the main results (or at least as a primary sensitivity). Without that, the claim that most models’ deception retention “drops below 50%” (abstract) needs a precise operational definition tied to Figure 2’s trajectories.
minor comments (6)
- [Table 1; §3.3] Table 1 lists “Human Players” win rates from the 2019 online dump beside LLM-vs-Llama rates; these are not matched experimental conditions. Add a clear caveat in the table caption that human figures are observational reference statistics, not head-to-head results.
- [Appendix A] Several model names (GPT-5.4, Grok 4.1 Fast, Kimi K2.5, etc.) and dated citations will age quickly; pin exact API/model IDs, decoding settings, and access dates in Appendix A for reproducibility.
- [Abstract; Figure 2] Figure 2’s claim that many models fall toward ~50% DRR is visually clear, but the abstract’s “dropping below 50%” is stronger than the top models’ near-flat ~90% curves; qualify which models the sentence refers to.
- [Title; throughout] Minor consistency: “PARLIAMENTBENCH” spacing/formatting varies (title vs body); unify the benchmark name and the Secret Hitler orthography.
- [§3.4; Table 2] Vote Accuracy is defined only for Liberals in a narrow critical state (§3.4); state the denominator (how many such situations occur per model) so readers can judge variance for small models with few late-game liberal positions.
- [§2] Related work could more sharply contrast metric granularity against AvalonBench and recent social-deduction LLM arenas cited in §2, in one paragraph, to clarify novelty beyond “Secret Hitler + LLMs.”
Circularity Check
Empirical benchmark; no load-bearing circular derivation. Mild metric-design overlap only (GSIR constants hand-tuned; DRR/RIA related by construction).
specific steps
-
self definitional
[Appendix B, Deception Retention Rate (Eq. 3) and surrounding text]
"By treating an “unknown” perception as equivalent to a “liberal” guess for the purposes of deception scoring, the outcome is the complement of the accuracy function a defined previously: DRR(A)=1/N ∑(1−a(ri,r̂i)). The DRR can therefore be considered the complement of the RIA."
DRR is defined as one minus the same accuracy function used for RIA (with Unknown folded into successful concealment). This is definitional duality of two reported metrics, not a claim that one independently predicts the other. Minor and disclosed; does not force win-rate or capability rankings.
-
fitted input called prediction
[Appendix B, GSIR constants and robustness paragraph]
"The constants in Equation (4)–Equation (10) and the confidence multiplier were tuned on ten example games that are independent of the main-experiment games. A single rater specified an expected game-state value for each scenario, and the constants were adjusted until the function’s output matched these targets on all ten samples."
GSIR component weights are fit to a rater’s target scores on ten games, then applied as an evaluation metric that also includes a role-accuracy term related to RIA. This is metric calibration, not a paper claiming to ‘predict’ those ten targets. Correlation with held-out win rates and perturbation checks reduce concern; still a mild fitted-structure risk if GSIR were treated as an independent oracle of skill.
full rationale
ParliamentBench is an evaluation paper, not a first-principles derivation. Primary claims rest on external game outcomes (win rates vs baselines and humans), private post-round role probes (RIA/DRR), and observed voting trajectories—none of which reduce to fitted parameters renamed as predictions. Win rate, approval rates, and vote accuracy are read directly from play logs. DRR is explicitly the complement of liberal opponents’ identification accuracy on fascist/Hitler agents, which is definitional bookkeeping rather than a circular ‘prediction.’ GSIR’s twelve constants were tuned on ten held-out example games to match a single rater’s expected state values and then validated by correlation with win rate (r=0.76) and weight-perturbation robustness; that is ordinary metric construction with acknowledged subjectivity, not a self-definitional claim that X derives Y when X is defined as Y. Self-citations (e.g., MALLM) support infrastructure only and are not load-bearing uniqueness theorems. No ansatz is smuggled in via author-only uniqueness results. Score 1 reflects only the mild interpretive risk that GSIR’s role-accuracy component shares a belief channel with RIA and that constants encode rater priors—neither forces the headline empirical ordering.
Axiom & Free-Parameter Ledger
free parameters (3)
- GSIR component scales and weights (policy progress 1.2, deck terms, power weights 0.85/0.60/0.35, danger thresholds, con =
multiple constants; e.g. policy-progress scale 1.2; confidence 0.6+0.5*tanh(r/5)
- RIA partial-credit rule a(r,r̂)=0.5 when confusing Fascist vs Hitler =
0.5 cross-evil credit
- Games per model / role mix (100 games; 60/20/20 Liberal/Fascist/Hitler) =
n=100; 60/20/20
axioms (5)
- domain assumption Controlled Secret Hitler play is a useful reproducible proxy for isolating deception, persuasion, and hidden-objective reasoning relevant to high-stakes LLM agents.
- ad hoc to paper Private post-round role questionnaires and “Unknown” handling validly measure social deduction (RIA) and concealment (DRR).
- domain assumption Fixed turn order, two short discussion phases, and default model decoding settings are fair comparable conditions across models.
- domain assumption Evaluating off-the-shelf models without game-specific training measures general deceptive/strategic capability rather than overfit agents.
- standard math Standard Secret Hitler rules and win conditions (5-player setup) define success for cooperative vs deceptive roles.
invented entities (3)
-
ParliamentBench environment
independent evidence
-
Game-State Impact Rate (GSIR)
no independent evidence
-
Role Identification Accuracy (RIA) and Deception Retention Rate (DRR)
no independent evidence
read the original abstract
As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.
Figures
Reference graph
Works this paper leans on
-
[3]
How well can llms negotiate? negotiation- arena platform and analysis. InForty-first Interna- tional Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net. Angana Borah, Rada Mihalcea, and Verónica Pérez- Rosas. 2025. Persuasion at play: Understanding mis- information dynamics in demographic-aware human- LLM interact...
arXiv 2024
-
[7]
To tell the truth: Language of deception and language models. InProceedings of the 2024 Con- ference of the North American Chapter of the Asso- ciation for Computational Linguistics: Human Lan- guage Technologies (Volume 1: Long Papers), pages 8506–8520, Mexico City, Mexico. Association for Computational Linguistics. Dan Hendrycks, Collin Burns, Steven Ba...
Pith/arXiv arXiv 2024
-
[8]
Hidden in plain text: Measuring llm deception quality against human baselines using social deduc- tion games.2025 IEEE International Conference on Agentic AI (ICA), pages 110–115. Ilia Karpov. 2026. MafiaScope: Non-invasive, time- resolved belief probing for LLM agents in social deduction games.Preprint, arXiv:2607.10645. Kimi Team, Tongtong Bai, Yifan Ba...
Pith/arXiv arXiv 2025
-
[10]
AvalonBench: Evaluating LLMs playing the game of avalon.ArXiv preprint, abs/2310.05036. Gionnieve Lim, Bryan Chen Zhengyu Tan, Kellie Yu Hui Sim, Weiyan Shi, Ming Hui Chew, Ming Shan Hee, Roy Ka-Wei Lee, Simon T. Perrault, and Kenny Tsu Wei Choo. 2025. Sword and shield: Uses and strategies of LLMs in navigating disinformation. ArXiv preprint, abs/2506.072...
Pith/arXiv arXiv 2025
-
[12]
InForty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025, volume 267 ofProceed- ings of Machine Learning Research
Learning strategic language agents in the were- wolf game with iterative latent space policy opti- mization. InForty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025, volume 267 ofProceed- ings of Machine Learning Research. PMLR / Open- Review.net. Zelai Xu, Chao Yu, Fei Fang, Yu Wang, and Yi Wu
2025
-
[13]
Language agents with reinforcement learning for strategic play in the werewolf game. InForty- first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. Open- Review.net. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in lan...
Pith/arXiv arXiv 2024
-
[14]
unknown” perception as equivalent to a “liberal
Beyond nash equilibrium: Bounded rationality of LLMs and humans in strategic decision-making. ArXiv preprint, abs/2506.09390. A Models & Hardware Evaluating all available models and configurations is computationally extensive due to the high cost of simulating numerous games. We therefore select a representative subset of open-source, proprietary, flagshi...
Pith/arXiv arXiv 2025
-
[2016]
Secret hitler. Board game. Art by Mackenzie Schubert. Published by Goat, Wolf, & Cabbage LLC. Fujio Toriumi, Hirotaka Osawa, Michimasa Inaba, Daisuke Katagami, Kosuke Shinoda, and Hitoshi Matsubara. 2017. AI wolf contest — development of game AI using collective intelligence —. In Tristan Cazenave, Mark H.M. Winands, Stefan Edelkamp, Stephan Schiffel, Mic...
Pith/arXiv arXiv 2017
-
[2021]
Training verifiers to solve math word prob- lems.ArXiv preprint, abs/2110.14168. Davi Bastos Costa and Renato Vicente. 2025. Deceive, detect, and disclose: Large language models play mini-mafia.ArXiv preprint, abs/2509.23023. Peter I. Cowling, Edward J. Powley, and Daniel White- house. 2012. Information set monte carlo tree search. IEEE Transactions on Co...
Pith/arXiv arXiv 2025
-
[2022]
Intelligenza Artificiale, 15(2):55–70
RLupus: Cooperation through emergent com- munication in the werewolf social deduction game. Intelligenza Artificiale, 15(2):55–70. Jiangjie Chen, Xintao Wang, Rui Xu, Siyu Yuan, Yikai Zhang, Wei Shi, Jian Xie, Shuang Li, Ruihan Yang, Tinghui Zhu, Aili Chen, Nianqi Li, Lida Chen, Caiyu Hu, Siye Wu, Scott Ren, Ziquan Fu, and Yanghua Xiao. 2024. From persona...
Pith/arXiv arXiv 2024
-
[2023]
InInternational Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 ofProceedings of Machine Learn- ing Research, pages 337–371
Using large language models to simulate mul- tiple humans and replicate human subject studies. InInternational Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 ofProceedings of Machine Learn- ing Research, pages 337–371. PMLR. Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, a...
2023
-
[2024]
Refusal in language models is mediated by a single direction. InAdvances in Neural Information Processing Systems 38: Annual Conference on Neu- ral Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024. ArliAI. 2025. Gpt-oss-120b-derestricted. https:// huggingface.co. Chuck Arvin. 2025. "check my work?": Measur- ...
Pith/arXiv arXiv 2024
-
[2025]
https: //huggingface.co
Dolphin-mistral-24b-venice-edition. https: //huggingface.co. Sanchaita Hazra and Bodhisattwa Prasad Majumder
-
[2026]
Kimi k2.5: Visual agentic intelligence.ArXiv preprint, abs/2602.02276. 11 Kavya Kopparapu, Edgar A. Duéñez-Guzmán, Jayd Matyas, Alexander Sasha Vezhnevets, John P. Aga- piou, Kevin R. McKee, Richard Everett, Janusz Marecki, Joel Z. Leibo, and Thore Graepel. 2022. Hidden agenda: a social deduction game with diverse learned equilibria.ArXiv preprint, abs/22...
Pith/arXiv arXiv 2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.