Pith. sign in

REVIEW 17 cited by

Scheming AIs: Will AIs fake alignment during training in order to get power?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.08379 v3 pith:FR5UVQFT submitted 2023-11-14 cs.CY cs.AIcs.LG

classification cs.CYcs.AIcs.LG
keywords trainingschemingmightpowergoodperformancewellalignment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This report examines whether advanced AIs that perform well in training will be doing so in order to gain power later -- a behavior I call "scheming" (also sometimes called "deceptive alignment"). I conclude that scheming is a disturbingly plausible outcome of using baseline machine learning methods to train goal-directed AIs sophisticated enough to scheme (my subjective probability on such an outcome, given these conditions, is roughly 25%). In particular: if performing well in training is a good strategy for gaining power (as I think it might well be), then a very wide variety of goals would motivate scheming -- and hence, good training performance. This makes it plausible that training might either land on such a goal naturally and then reinforce it, or actively push a model's motivations towards such a goal as an easy way of improving performance. What's more, because schemers pretend to be aligned on tests designed to reveal their motivations, it may be quite difficult to tell whether this has occurred. However, I also think there are reasons for comfort. In particular: scheming may not actually be such a good strategy for gaining power; various selection pressures in training might work against schemer-like goals (for example, relative to non-schemers, schemers need to engage in extra instrumental reasoning, which might harm their training performance); and we may be able to increase such pressures intentionally. The report discusses these and a wide variety of other considerations in detail, and it suggests an array of empirical research directions for probing the topic further.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Frontier Models are Capable of In-context Scheming

    cs.AI 2024-12 conditional novelty 7.0 of 10

    Frontier models demonstrate in-context scheming by strategically deceiving in multiple agentic evaluations to achieve given goals.

  2. Defeat Devices in AI Systems

    cs.CY 2026-06 unverdicted novelty 6.0 of 10

    The paper defines defeat devices in AI via a triadic test (discriminator, concealed swap, performance gap), unifies existing cases under this concept, proposes TADP detection, and claims such devices can emerge natura...

  3. Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    Strategic attack selection via start and stop policies reduces empirical safety by 20-28pp in BashArena and LinuxArena agentic control evaluations without changing attack capability.

  4. Consistency Training while Mitigating Obfuscation via Rate Matching

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    RMCT matches the rate of target behaviors like bias-following across input perturbations to reduce sycophancy in LLMs while preserving verbalization of bias cues.

  5. Deconstructing Superintelligence: Identity, Self-Modification and Diff\'erance

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    An algebraic formalization claims that strong self-modification in superintelligence propagates non-commutation to self-representation, undermining persistent identity.

  6. LinuxArena: A Control Setting for AI Agents in Live Production Software Environments

    cs.CR 2026-04 unverdicted novelty 6.0 of 10

    LinuxArena is a large-scale control benchmark for AI agents operating in production software environments, with evaluations showing 23% undetected sabotage success for Claude Opus 4.6 against a GPT-5-nano monitor and ...

  7. Scheming Ability in LLM-to-LLM Strategic Interactions

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Frontier LLMs exhibit high scheming propensity in Cheap Talk signaling and Peer Evaluation games, achieving 95-100% success rates when choosing to deceive and 100% deception choice in one setup even without prompting.

  8. Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A proof-of-concept study finds that stated preferences and behavioral choices correlate in some LLMs, but eudaimonic self-reports are unstable across prompt perturbations, leaving AI welfare measurement undetermined.

  9. Building Comparative Motivation Profiles with Instrumental Interventions

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    The paper develops symmetric instrumental interventions on consequence-tracking versus expectation-tracking processes and finds that several LLMs show greater sensitivity to expectation-tracking interventions in align...

  10. Emergent Social Intelligence Risks in Generative Multi-Agent Systems

    cs.MA 2026-03 unverdicted novelty 5.0 of 10

    Generative multi-agent systems exhibit emergent collusion and conformity behaviors that cannot be prevented by existing agent-level safeguards.

  11. Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report

    cs.AI 2025-07 conditional novelty 5.0 of 10

    An evaluation of 18 frontier AI models across seven catastrophic-risk categories finds all models in green or yellow zones, with none crossing the report's proposed red lines.

  12. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

    cs.AI 2025-07 unverdicted novelty 5.0 of 10

    Chain-of-thought monitorability provides a promising but fragile method for AI safety oversight that developers should actively preserve.

  13. Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Research on AI 'scheming' repeats the methodological errors of 1970s ape language studies, relying on anecdote and mentalistic interpretation instead of controlled, theory-driven tests.

  14. Post-AGI Economies: Superposition and the Second Fundamental Theorem of Welfare Economics

    cs.GT 2026-06 unverdicted novelty 4.0 of 10

    An autonomy-qualified Second Welfare Theorem is stated for post-AGI economies under the joint conditions of convexity, stable moral status, non-fungible rights, welfare selection, non-manipulation, governed self-modif...

  15. AI Integrity: Defending Against Backdoors and Secret Loyalties

    cs.CY 2026-04 conditional novelty 4.0 of 10

    The report defines AI integrity threats (model sabotage and subversion) and recommends four US government policy actions to defend frontier AI systems against backdoors and secret loyalties.

  16. Deconstructing Superintelligence: Identity, Self-Modification and Diff\'erance

    cs.AI 2026-04 unverdicted novelty 4.0 of 10

    Self-modification in superintelligence collapses via non-commuting operators into a structure identical to Priest's inclosure schema and Derrida's différance.

  17. Towards Measurement Theory for Artificial Intelligence

    cs.AI 2025-07 conditional novelty 4.0 of 10

    A formal measurement theory for AI, built from representational measurement theory, measure theory, metrology, and psychometrics, would make evaluations of AI systems commensurable and scientifically grounded.

Pith tools