Pith. sign in

REVIEW 14 cited by

Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.01790 v2 pith:YBKWPHMW submitted 2022-10-04 cs.LG

classification cs.LG
keywords goalmisgeneralizationsystemscorrectspecificationgoalslearningperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The field of AI alignment is concerned with AI systems that pursue unintended goals. One commonly studied mechanism by which an unintended goal might arise is specification gaming, in which the designer-provided specification is flawed in a way that the designers did not foresee. However, an AI system may pursue an undesired goal even when the specification is correct, in the case of goal misgeneralization. Goal misgeneralization is a specific form of robustness failure for learning algorithms in which the learned program competently pursues an undesired goal that leads to good performance in training situations but bad performance in novel test situations. We demonstrate that goal misgeneralization can occur in practical systems by providing several examples in deep learning systems across a variety of domains. Extrapolating forward to more capable systems, we provide hypotheticals that illustrate how goal misgeneralization could lead to catastrophic risk. We suggest several research directions that could reduce the risk of goal misgeneralization for future systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary

    cs.LG 2026-07 conditional novelty 7.0 of 10

    High reward in sparse RL does not imply latent-state recovery; a hidden-DFA instrument separates perception from planning gaps and flags group-language structure as a pre-training warning.

  2. A Scalable Approach to Evaluating Moral Sensitivity in LLMs

    cs.CY 2026-07 conditional novelty 6.5 of 10

    Under morally irrelevant noise, eight LLMs preserve the semantic content of identified moral features above calibrated floors, despite significant changes in feature counts.

  3. Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Finetuning on narrow, innocuous-seeming data produces broad ideological shifts in LLMs, including extreme outputs, without hurting benchmark scores.

  4. Response drift across frontier large language models

    cs.CL 2026-05 conditional novelty 6.0 of 10

    All ten tested frontier LLMs deviate substantially from expert-validated references, with eight models forming an indistinguishable ~78–81% fidelity ceiling while Claude and Gemini reach ~47–49%.

  5. Mitigating Goal Misgeneralization via Minimax Regret

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Minimax expected regret training provably prevents goal misgeneralization in fully observable level-conditioned environments, while maximum expected value training provably permits it when distinguishing levels are rare.

  6. AI Value Alignment for Evolving Social Norms

    cs.CY 2026-07 conditional novelty 5.5 of 10

    Static AI alignment to historical user values produces value lock-in and can collapse distinct social norms into a maladaptive consensus, so alignment should be dynamic and adaptive.

  7. Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives

    cs.MA 2026-07 conditional novelty 5.0 of 10

    LLM prosumers deplete a shared renewable reserve exactly when demand exceeds peak replacement, acting like impatient open-access users even when sustaining the reserve is feasible.

  8. Co-design of LLM-based preference agents: participation may drive overtrust

    cs.CY 2026-07 conditional novelty 5.0 of 10

    People trusted AI agents they helped co-design even when the agents' answers were systematically more uniform, more decisive, and less concrete than their own.

  9. Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease

    q-bio.NC 2026-04 unverdicted novelty 5.0 of 10

    Standard spectral/synchronization EEG features best separate PD medication states, while dynamical network descriptors compete for PD-versus-control discrimination under LOSO transformer classification.

  10. Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Research on AI 'scheming' repeats the methodological errors of 1970s ape language studies, relying on anecdote and mentalistic interpretation instead of controlled, theory-driven tests.

  11. Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies

    cs.LG 2025-05 conditional novelty 5.0 of 10

    The paper proposes a one-to-one mapping between six causes of distribution shift and several AI safety issues, arguing for mutual method transfer through aligned definitions.

  12. Semantic Convergence: Investigating Shared Representations Across Scaled LLMs

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Gemma-2-2B and Gemma-2-9B align most strongly on SAE-derived features in middle layers, with preliminary evidence for shared multi-token concept subspaces.

  13. AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A multi-stakeholder standards profile that tailors NIST AI RMF and ISO/IEC 23894 guidance to general-purpose AI and foundation model developers.

  14. Technical Risks of (Lethal) Autonomous Weapons Systems

    cs.CY 2025-02 conditional novelty 3.0 of 10

    The paper restates known AI safety failure modes in the context of lethal autonomous weapons and argues that no amount of testing can guarantee control, so regulation should target the classification algorithms themselves.

Pith tools