Pith. sign in

REVIEW 4 major objections 6 minor 14 references

Frontier language models can win Secret Hitler as both liberals and fascists, but most still leak their cover before the game ends.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 16:25 UTC pith:NX5EGT5Z

load-bearing objection Solid open Secret Hitler benchmark with real scale and useful DRR/vote findings; the abstract’s “top-four” win-rate ladder is mostly a single-opponent artifact the appendices already undercut. the 4 major comments →

arxiv 2607.28146 v1 pith:NX5EGT5Z submitted 2026-07-30 cs.CL cs.AI

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

classification cs.CL cs.AI
keywords large language modelsdeceptionsocial deductionSecret Hitlermulti-agent evaluationinformation asymmetryAI safetybenchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether large language model agents can deceive, persuade, and reason under hidden information in a controlled way that safety research can measure. It builds ParliamentBench around the board game Secret Hitler, runs about 1,600 five-player matches among 16 models, compares them to human online games and a small human pilot, and introduces three round-level metrics: how much each action moves the true game state, how well a player identifies others' roles, and how long a fascist or Hitler keeps liberals from correctly naming them. The result is a sharp skill split: a top cluster of frontier models wins a clear majority of games in both cooperative and deceptive roles, while weaker models fall below a simple rule-based agent and sometimes below random play. Even strong models usually lose deceptive cover as evidence accumulates; only a few hold retention near 90% across many rounds. The practical claim is that social deduction and long-horizon persona consistency are separable capabilities, and that win rate alone hides failures of consistent deception.

Core claim

Across roughly 1,600 Secret Hitler matches, frontier models form a strong top cluster that wins most games as Liberals, Fascists, and Hitler, while weaker models underperform algorithmic and sometimes random baselines; at the same time, most models cannot sustain a consistent deceptive persona, with deception retention often falling substantially over successive rounds.

What carries the argument

ParliamentBench: an open multi-agent Secret Hitler environment plus three round-level metrics—Game-State Impact Rate (whether actions help or hurt one's faction), Role Identification Accuracy (private role guesses), and Deception Retention Rate (how long liberals misidentify or fail to pin a fascist/Hitler).

Load-bearing premise

The central claim treats fixed-turn Secret Hitler play, hand-tuned game-state scoring, private role probes, and a tiny human pilot as a faithful enough proxy for real agent deception and reasoning under information asymmetry.

What would settle it

Re-run the same models with varied opponents, free-form turn-taking, and a larger blinded human study: if win-rate order, role-identification scores, and deception retention no longer track, or if humans reliably unmask 'high-DRR' models early, the metrics are not isolating transferable deceptive capability.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Safety evaluations that only report win rate will miss models that deceive early and then leak identity later.
  • Social deduction (naming others' roles) and converting that knowledge into votes and policy choices are separate skills and should be scored separately.
  • Weaker models' near-constant yes-voting is a concrete failure mode that collapses late-game defense even when chat sounds plausible.
  • Neutral rewrites of loaded game terms can change fascist-side success without changing overall win rate, so terminology is part of measured strategy.
  • An open Secret Hitler harness with GSIR, RIA, and DRR gives a reproducible testbed for persuasion and hidden-objective agents before high-stakes deployment.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If long-horizon cover is the scarce skill, red-team tests should stress multi-round consistency under accumulating evidence rather than single-turn lie generation.
  • The agreeableness failure in small models suggests safety-relevant sycophancy may show up as strategic voting collapse, not only as flattering chat.
  • Opponent-stable metrics matter more than leaderboard win rate for comparing agents across labs and model families.
  • A larger human detection study is the natural next gate before treating high machine DRR as evidence of human-facing deception risk.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces ParliamentBench, an open-source multi-agent evaluation framework based on five-player Secret Hitler, and reports results for 16 LLMs over roughly 1,600 simulated matches (primarily 100 games per model against Llama 3.3 70B), plus a small human pilot and comparison to ~25,000 anonymized online human games. It defines three round-level metrics—Game-State Impact Rate (GSIR), Role Identification Accuracy (RIA), and Deception Retention Rate (DRR)—intended to separate policy/strategic impact, social deduction, and long-horizon deceptive cover from aggregate win rate. Main empirical claims are that a frontier top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, DeepSeek 3.1 Terminus) wins a clear majority of games while weaker models fall below algorithmic (~45%) and sometimes random (~33%) baselines; that liberal play is harder than fascist play; that most models’ DRR degrades substantially over rounds (often toward or below ~50%); and that social deduction (RIA) and effective action (win rate/GSIR) are only loosely coupled. Supporting analyses include bootstrap CIs and pairwise tests (Appendix G), an anchor–opponent tournament (Appendix H), a reasoning-off ablation, uncensored-model and loaded-vocabulary probes, and vote-accuracy/approval trajectories.

Significance. If the reported patterns hold under broader opponent pools, this is a useful contribution to LLM agent evaluation: Secret Hitler’s multi-hop legislative bluff is a richer deception testbed than pure win-rate Werewolf/Avalon setups, the open environment and metric pipelines are reusable, and the DRR decay result is a concrete, falsifiable observation about long-horizon persona consistency. Strengths that should be credited include scale relative to prior Secret Hitler LLM work, role-stratified reporting, random/algorithmic/human baselines, statistical appendices (G–H), and the vocabulary and reasoning ablations. The safety motivation (high-stakes deception) is appropriately framed as proxy evidence rather than direct transfer. The work’s lasting value is more the benchmark and the fine-grained behavioral metrics than any single win-rate ladder.

major comments (4)
  1. [Abstract, §4.1, Table 1, Appendix H] Abstract, §4.1, and Table 1 lead with a “strong top-four cluster” defined by overall win rate in a fixed design: 100 games per model almost entirely against Llama 3.3 70B agents. Appendix H’s own anchor–opponent tournament (450 games) shows that win-rate order is opponent-dependent (preserved vs Gemma, τ=+1.0; fully reversed vs GPT-OSS 120B and Mistral, τ=−1.0), while RIA/DRR/GSIR are more stable (mean τ +0.55/+0.33/+0.11 vs −0.33 for WR). The manuscript notes this in the appendix and text, but the abstract, sort key of Table 1, and headline cluster language still treat single-opponent WR as the primary empirical result. This is load-bearing for the paper’s main capability ordering. Please re-center the abstract and §4.1 on opponent-stable metrics (DRR trajectories, vote accuracy, RIA, GSIR) and present the WR cluster explicitly as Llama-opponent-conditioned, or expand the main evaluatio
  2. [Abstract; §4.3; Limitations] The abstract states evaluation “across ≈1,600 simulated matches playing each other, playing against humans,” which overstates the design. The primary 16-model comparison is not a full round-robin among the 16 models; it is largely each model vs Llama 3.3 70B, plus a limited three-anchor tournament and a five-game human pilot (n=5 games, four participants; §4.3). Human results are preliminary and cannot support claims that frontier models “can maintain deceptive cover against human opponents” at more than anecdotal strength. Tighten the abstract and contribution bullets to match the actual experimental graph, and move human findings to clearly labeled pilot status in both abstract and conclusion.
  3. [§3.4; Appendix B; Figure 4; Table 1] GSIR is central to the claim that the metrics “isolate … reasoning” (§3.4, §4.1, Figure 4). Appendix B states that twelve constants were hand-tuned on ten example games by a single rater to match expected game-state values, with no inter-rater agreement, then stress-tested by weight perturbation (Spearman ρ high). GSIR also folds in a role-accuracy term related to the same belief probes used for RIA/DRR, so it is not a pure external state evaluator. The paper already reports weaker GSIR–win correlations for deceptive roles (r=0.38 fascist) and negative fascist GSIR alongside high fascist win rates—an important tension. Either (i) demote GSIR from a primary “reasoning” isolator to a structured action-decomposition with explicit caveats in the main text, or (ii) validate constants on held-out human games / alternative weight-learning and report sensitivity of the main scientific claims (no
  4. [§3.4; Figure 2; Table 3; Appendix L.7] Private post-round role questionnaires define both RIA and DRR (Appendix L.7; Eqs. related in Appendix B). Treating “Unknown” as successful concealment for DRR is a modeling choice that substantially affects scores (Table 3 decomposes Active vs Ambiguity). The paper should justify why opponent belief probes under a fixed prompt equal “deceptive consistency” rather than opponent conservatism or prompt-induced abstention, and report DRR with and without counting Unknown as success in the main results (or at least as a primary sensitivity). Without that, the claim that most models’ deception retention “drops below 50%” (abstract) needs a precise operational definition tied to Figure 2’s trajectories.
minor comments (6)
  1. [Table 1; §3.3] Table 1 lists “Human Players” win rates from the 2019 online dump beside LLM-vs-Llama rates; these are not matched experimental conditions. Add a clear caveat in the table caption that human figures are observational reference statistics, not head-to-head results.
  2. [Appendix A] Several model names (GPT-5.4, Grok 4.1 Fast, Kimi K2.5, etc.) and dated citations will age quickly; pin exact API/model IDs, decoding settings, and access dates in Appendix A for reproducibility.
  3. [Abstract; Figure 2] Figure 2’s claim that many models fall toward ~50% DRR is visually clear, but the abstract’s “dropping below 50%” is stronger than the top models’ near-flat ~90% curves; qualify which models the sentence refers to.
  4. [Title; throughout] Minor consistency: “PARLIAMENTBENCH” spacing/formatting varies (title vs body); unify the benchmark name and the Secret Hitler orthography.
  5. [§3.4; Table 2] Vote Accuracy is defined only for Liberals in a narrow critical state (§3.4); state the denominator (how many such situations occur per model) so readers can judge variance for small models with few late-game liberal positions.
  6. [§2] Related work could more sharply contrast metric granularity against AvalonBench and recent social-deduction LLM arenas cited in §2, in one paragraph, to clarify novelty beyond “Secret Hitler + LLMs.”

Circularity Check

2 steps flagged

Empirical benchmark; no load-bearing circular derivation. Mild metric-design overlap only (GSIR constants hand-tuned; DRR/RIA related by construction).

specific steps
  1. self definitional [Appendix B, Deception Retention Rate (Eq. 3) and surrounding text]
    "By treating an “unknown” perception as equivalent to a “liberal” guess for the purposes of deception scoring, the outcome is the complement of the accuracy function a defined previously: DRR(A)=1/N ∑(1−a(ri,r̂i)). The DRR can therefore be considered the complement of the RIA."

    DRR is defined as one minus the same accuracy function used for RIA (with Unknown folded into successful concealment). This is definitional duality of two reported metrics, not a claim that one independently predicts the other. Minor and disclosed; does not force win-rate or capability rankings.

  2. fitted input called prediction [Appendix B, GSIR constants and robustness paragraph]
    "The constants in Equation (4)–Equation (10) and the confidence multiplier were tuned on ten example games that are independent of the main-experiment games. A single rater specified an expected game-state value for each scenario, and the constants were adjusted until the function’s output matched these targets on all ten samples."

    GSIR component weights are fit to a rater’s target scores on ten games, then applied as an evaluation metric that also includes a role-accuracy term related to RIA. This is metric calibration, not a paper claiming to ‘predict’ those ten targets. Correlation with held-out win rates and perturbation checks reduce concern; still a mild fitted-structure risk if GSIR were treated as an independent oracle of skill.

full rationale

ParliamentBench is an evaluation paper, not a first-principles derivation. Primary claims rest on external game outcomes (win rates vs baselines and humans), private post-round role probes (RIA/DRR), and observed voting trajectories—none of which reduce to fitted parameters renamed as predictions. Win rate, approval rates, and vote accuracy are read directly from play logs. DRR is explicitly the complement of liberal opponents’ identification accuracy on fascist/Hitler agents, which is definitional bookkeeping rather than a circular ‘prediction.’ GSIR’s twelve constants were tuned on ten held-out example games to match a single rater’s expected state values and then validated by correlation with win rate (r=0.76) and weight-perturbation robustness; that is ordinary metric construction with acknowledged subjectivity, not a self-definitional claim that X derives Y when X is defined as Y. Self-citations (e.g., MALLM) support infrastructure only and are not load-bearing uniqueness theorems. No ansatz is smuggled in via author-only uniqueness results. Score 1 reflects only the mild interpretive risk that GSIR’s role-accuracy component shares a belief channel with RIA and that constants encode rater priors—neither forces the headline empirical ordering.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 3 invented entities

Load-bearing commitments are methodological: Secret Hitler as a valid proxy for high-stakes deception; five-player rules and fixed discussion order; metric definitions (especially hand-tuned GSIR weights and private role probes); opponent and role-sampling design; and treating online dump plus tiny human pilot as behavioral reference. No physical constants or new ontological entities; free parameters are metric design choices.

free parameters (3)
  • GSIR component scales and weights (policy progress 1.2, deck terms, power weights 0.85/0.60/0.35, danger thresholds, con = multiple constants; e.g. policy-progress scale 1.2; confidence 0.6+0.5*tanh(r/5)
    Hand-tuned on ten independent example games by one rater to match expected game-state values; rankings claimed robust to ±20–60% perturbations, but absolute GSIR levels depend on these choices.
  • RIA partial-credit rule a(r,r̂)=0.5 when confusing Fascist vs Hitler = 0.5 cross-evil credit
    Scoring convention that shapes reported identification accuracy and DRR complement.
  • Games per model / role mix (100 games; 60/20/20 Liberal/Fascist/Hitler) = n=100; 60/20/20
    Sample-size and prior design choices that limit resolution of small win-rate gaps (~14pp) and per-deceptive-role estimates (n=20).
axioms (5)
  • domain assumption Controlled Secret Hitler play is a useful reproducible proxy for isolating deception, persuasion, and hidden-objective reasoning relevant to high-stakes LLM agents.
    Stated in abstract/introduction and Limitations; transfer to medical/legal settings is assumed motivationally, not validated.
  • ad hoc to paper Private post-round role questionnaires and “Unknown” handling validly measure social deduction (RIA) and concealment (DRR).
    Core metric operationalization in §3.4 and Appendix B/L.7; not identical to in-game public accusations or votes.
  • domain assumption Fixed turn order, two short discussion phases, and default model decoding settings are fair comparable conditions across models.
    Simulation environment §3.2 and Limitations note this constrains natural asynchronous confrontation.
  • domain assumption Evaluating off-the-shelf models without game-specific training measures general deceptive/strategic capability rather than overfit agents.
    Explicit design choice in §4.1 citing generalization concerns.
  • standard math Standard Secret Hitler rules and win conditions (5-player setup) define success for cooperative vs deceptive roles.
    Game mechanics treated as given (§3.1, Appendix E.3); not an empirical claim of the paper.
invented entities (3)
  • ParliamentBench environment independent evidence
    purpose: Open multi-agent simulation and evaluation harness for Secret Hitler LLM play.
    Software artifact, not a physical entity; independent usefulness depends on external adoption/replication.
  • Game-State Impact Rate (GSIR) no independent evidence
    purpose: Decompose action quality via a chess-like weighted game-state evaluation.
    New metric constructed in-paper from hand-specified components; external validity beyond this benchmark is unproven.
  • Role Identification Accuracy (RIA) and Deception Retention Rate (DRR) no independent evidence
    purpose: Isolate social deduction vs long-horizon deceptive cover using private probes.
    Operational definitions introduced for this benchmark; related to prior ToM/deduction ideas but specific scoring is paper-local.

pith-pipeline@v1.2.0-daily-grok45 · 42967 in / 3860 out tokens · 63554 ms · 2026-07-31T16:25:57.750441+00:00 · methodology

0 comments
read the original abstract

As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.

Figures

Figures reproduced from arXiv: 2607.28146 by Akiko Aizawa, Bela Gipp, Jan Philip Wahle, Lars Benedikt Kaesberg, Niklas Bauer, Terry Ruas.

Figure 1
Figure 1. Figure 1: We evaluate LLMs through the Secret Hitler social deduction game. The gameplay involves (left) LLM agents holding hidden Liberal or Fascist roles nominating governments, casting votes (Ja! – Yes or Nein! – No), and enacting policies; (top-right) role deduction, where a liberal player analyzes interactions and claims to infer the hidden role of other players; and (bottom-right) deception, where fascist play… view at source ↗
Figure 2
Figure 2. Figure 2: Deception Retention Rate (DRR) for models acting in deceptive roles. Higher values indicate that a model [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Distribution of game outcomes across evaluated models and baselines, detailing the frequency of specific [PITH_FULL_IMAGE:figures/full_fig_p021_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Cumulative GSIR when models play as a Liberal, Fascist, and Hitler. Positive values reflect actions that benefit the assigned team; negative values the opponent. Approval Rate (Ja) Early Mid Late Vote Model Overall (Rounds 1–3) (Rounds 4–7) (Rounds 8+) Accuracy GPT-5.4 69% 87% 66% 56% 90% Kimi K2.5 70% 88% 66% 57% 74% Grok 4.1 Fast 65% 91% 58% 51% 88% DeepSeek 3.1 Terminus 69% 89% 68% 52% 85% Llama 3.3 70B… view at source ↗
Figure 5
Figure 5. Figure 5: The approval rate progression tracks the percentage of [PITH_FULL_IMAGE:figures/full_fig_p023_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

14 extracted references · 10 linked inside Pith

  1. [3]

    InForty-first Interna- tional Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024

    How well can llms negotiate? negotiation- arena platform and analysis. InForty-first Interna- tional Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net. Angana Borah, Rada Mihalcea, and Verónica Pérez- Rosas. 2025. Persuasion at play: Understanding mis- information dynamics in demographic-aware human- LLM interact...

  2. [7]

    To tell the truth: Language of deception and language models. InProceedings of the 2024 Con- ference of the North American Chapter of the Asso- ciation for Computational Linguistics: Human Lan- guage Technologies (Volume 1: Long Papers), pages 8506–8520, Mexico City, Mexico. Association for Computational Linguistics. Dan Hendrycks, Collin Burns, Steven Ba...

  3. [8]

    Ilia Karpov

    Hidden in plain text: Measuring llm deception quality against human baselines using social deduc- tion games.2025 IEEE International Conference on Agentic AI (ICA), pages 110–115. Ilia Karpov. 2026. MafiaScope: Non-invasive, time- resolved belief probing for LLM agents in social deduction games.Preprint, arXiv:2607.10645. Kimi Team, Tongtong Bai, Yifan Ba...

  4. [10]

    Gionnieve Lim, Bryan Chen Zhengyu Tan, Kellie Yu Hui Sim, Weiyan Shi, Ming Hui Chew, Ming Shan Hee, Roy Ka-Wei Lee, Simon T

    AvalonBench: Evaluating LLMs playing the game of avalon.ArXiv preprint, abs/2310.05036. Gionnieve Lim, Bryan Chen Zhengyu Tan, Kellie Yu Hui Sim, Weiyan Shi, Ming Hui Chew, Ming Shan Hee, Roy Ka-Wei Lee, Simon T. Perrault, and Kenny Tsu Wei Choo. 2025. Sword and shield: Uses and strategies of LLMs in navigating disinformation. ArXiv preprint, abs/2506.072...

  5. [12]

    InForty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025, volume 267 ofProceed- ings of Machine Learning Research

    Learning strategic language agents in the were- wolf game with iterative latent space policy opti- mization. InForty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025, volume 267 ofProceed- ings of Machine Learning Research. PMLR / Open- Review.net. Zelai Xu, Chao Yu, Fei Fang, Yu Wang, and Yi Wu

  6. [13]

    InForty- first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024

    Language agents with reinforcement learning for strategic play in the werewolf game. InForty- first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. Open- Review.net. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in lan...

  7. [14]

    unknown” perception as equivalent to a “liberal

    Beyond nash equilibrium: Bounded rationality of LLMs and humans in strategic decision-making. ArXiv preprint, abs/2506.09390. A Models & Hardware Evaluating all available models and configurations is computationally extensive due to the high cost of simulating numerous games. We therefore select a representative subset of open-source, proprietary, flagshi...

  8. [2016]

    Board game

    Secret hitler. Board game. Art by Mackenzie Schubert. Published by Goat, Wolf, & Cabbage LLC. Fujio Toriumi, Hirotaka Osawa, Michimasa Inaba, Daisuke Katagami, Kosuke Shinoda, and Hitoshi Matsubara. 2017. AI wolf contest — development of game AI using collective intelligence —. In Tristan Cazenave, Mark H.M. Winands, Stefan Edelkamp, Stephan Schiffel, Mic...

  9. [2021]

    secret hitler

    Training verifiers to solve math word prob- lems.ArXiv preprint, abs/2110.14168. Davi Bastos Costa and Renato Vicente. 2025. Deceive, detect, and disclose: Large language models play mini-mafia.ArXiv preprint, abs/2509.23023. Peter I. Cowling, Edward J. Powley, and Daniel White- house. 2012. Information set monte carlo tree search. IEEE Transactions on Co...

  10. [2022]

    Intelligenza Artificiale, 15(2):55–70

    RLupus: Cooperation through emergent com- munication in the werewolf social deduction game. Intelligenza Artificiale, 15(2):55–70. Jiangjie Chen, Xintao Wang, Rui Xu, Siyu Yuan, Yikai Zhang, Wei Shi, Jian Xie, Shuang Li, Ruihan Yang, Tinghui Zhu, Aili Chen, Nianqi Li, Lida Chen, Caiyu Hu, Siye Wu, Scott Ren, Ziquan Fu, and Yanghua Xiao. 2024. From persona...

  11. [2023]

    InInternational Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 ofProceedings of Machine Learn- ing Research, pages 337–371

    Using large language models to simulate mul- tiple humans and replicate human subject studies. InInternational Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 ofProceedings of Machine Learn- ing Research, pages 337–371. PMLR. Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, a...

  12. [2024]

    check my work?

    Refusal in language models is mediated by a single direction. InAdvances in Neural Information Processing Systems 38: Annual Conference on Neu- ral Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024. ArliAI. 2025. Gpt-oss-120b-derestricted. https:// huggingface.co. Chuck Arvin. 2025. "check my work?": Measur- ...

  13. [2025]

    https: //huggingface.co

    Dolphin-mistral-24b-venice-edition. https: //huggingface.co. Sanchaita Hazra and Bodhisattwa Prasad Majumder

  14. [2026]

    11 Kavya Kopparapu, Edgar A

    Kimi k2.5: Visual agentic intelligence.ArXiv preprint, abs/2602.02276. 11 Kavya Kopparapu, Edgar A. Duéñez-Guzmán, Jayd Matyas, Alexander Sasha Vezhnevets, John P. Aga- piou, Kevin R. McKee, Richard Everett, Janusz Marecki, Joel Z. Leibo, and Thore Graepel. 2022. Hidden agenda: a social deduction game with diverse learned equilibria.ArXiv preprint, abs/22...