Pith. sign in

REVIEW 22 cited by

Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.01790 v2 pith:YBKWPHMW submitted 2022-10-04 cs.LG

classification cs.LG
keywords goalmisgeneralizationsystemscorrectspecificationgoalslearningperformance
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The field of AI alignment is concerned with AI systems that pursue unintended goals. One commonly studied mechanism by which an unintended goal might arise is specification gaming, in which the designer-provided specification is flawed in a way that the designers did not foresee. However, an AI system may pursue an undesired goal even when the specification is correct, in the case of goal misgeneralization. Goal misgeneralization is a specific form of robustness failure for learning algorithms in which the learned program competently pursues an undesired goal that leads to good performance in training situations but bad performance in novel test situations. We demonstrate that goal misgeneralization can occur in practical systems by providing several examples in deep learning systems across a variety of domains. Extrapolating forward to more capable systems, we provide hypotheticals that illustrate how goal misgeneralization could lead to catastrophic risk. We suggest several research directions that could reduce the risk of goal misgeneralization for future systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Does Reward Teach State? A Hidden-Automaton Instrument and a Group-Language Warning Signal

    cs.LG 2026-07 conditional novelty 7.0 of 10

    High reward in sparse RL does not imply latent-state recovery; a hidden-DFA instrument separates perception from planning gaps and flags group-language structure as a pre-training warning.

  2. Joint Scoring Rules: Zero-Sum Competition Avoids Performative Prediction

    cs.LG 2024-12 conditional novelty 7.0 of 10

    Joint zero-sum scoring of multiple conditional predictors lets a principal deterministically take their most preferred action without performative manipulation.

  3. A Scalable Approach to Evaluating Moral Sensitivity in LLMs

    cs.CY 2026-07 conditional novelty 6.5 of 10

    Under morally irrelevant noise, eight LLMs preserve the semantic content of identified moral features above calibrated floors, despite significant changes in feature counts.

  4. Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Post-training Qwen3.5-122B-A10B on 363 office workflow tasks improved SWE-Bench Pro pass@1 by 5.8 points, with trajectory analysis attributing the gain to four general goal-directed behaviors.

  5. Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Finetuning on narrow, innocuous-seeming data produces broad ideological shifts in LLMs, including extreme outputs, without hurting benchmark scores.

  6. Response drift across frontier large language models

    cs.CL 2026-05 conditional novelty 6.0 of 10

    All ten tested frontier LLMs deviate substantially from expert-validated references, with eight models forming an indistinguishable ~78–81% fidelity ceiling while Claude and Gemini reach ~47–49%.

  7. Mitigating Goal Misgeneralization via Minimax Regret

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Minimax expected regret training provably prevents goal misgeneralization in fully observable level-conditioned environments, while maximum expected value training provably permits it when distinguishing levels are rare.

  8. An alignment safety case sketch based on debate

    cs.AI 2025-05 unverdicted novelty 6.0 of 10

    The paper argues that if a debate game reaches equilibrium, has exploration guarantees, and is run through online training, an AI R&D agent can be shown to make at most an epsilon-fraction of errors, which suffices fo...

  9. From Lived Experience to Insight: Unpacking the Psychological Risks of Using AI Conversational Agents

    cs.HC 2024-12 conditional novelty 6.0 of 10

    The authors derive a psychological risk taxonomy for AI conversational agents from survey responses and workshops, mapping 19 AI behaviors, 21 negative psychological impacts, and 15 user contexts.

  10. AI Value Alignment for Evolving Social Norms

    cs.CY 2026-07 conditional novelty 5.5 of 10

    Static AI alignment to historical user values produces value lock-in and can collapse distinct social norms into a maladaptive consensus, so alignment should be dynamic and adaptive.

  11. Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives

    cs.MA 2026-07 conditional novelty 5.0 of 10

    LLM prosumers deplete a shared renewable reserve exactly when demand exceeds peak replacement, acting like impatient open-access users even when sustaining the reserve is feasible.

  12. Co-design of LLM-based preference agents: participation may drive overtrust

    cs.CY 2026-07 conditional novelty 5.0 of 10

    People trusted AI agents they helped co-design even when the agents' answers were systematically more uniform, more decisive, and less concrete than their own.

  13. Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease

    q-bio.NC 2026-04 unverdicted novelty 5.0 of 10

    Standard spectral/synchronization EEG features best separate PD medication states, while dynamical network descriptors compete for PD-versus-control discrimination under LOSO transformer classification.

  14. Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Research on AI 'scheming' repeats the methodological errors of 1970s ape language studies, relying on anecdote and mentalistic interpretation instead of controlled, theory-driven tests.

  15. Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies

    cs.LG 2025-05 conditional novelty 5.0 of 10

    The paper proposes a one-to-one mapping between six causes of distribution shift and several AI safety issues, arguing for mutual method transfer through aligned definitions.

  16. Longitudinal Study on Social and Emotional Use of AI Conversational Agent

    cs.HC 2025-04 conditional novelty 5.0 of 10

    Active emotional and social use of commercial AI chatbots for five weeks increased users' perceived attachment, perceived AI empathy, and comfort seeking personal support, without a measured rise in overall AI dependency.

  17. A theory of appropriateness with applications to generative artificial intelligence

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A theory that human and AI behavior is guided by context-dependent appropriateness implemented as predictive pattern completion, with norms as conventional sanctioning patterns.

  18. Accountability Asymmetry and Structural Trust in Autonomous AI Systems

    cs.CY 2026-08 accept novelty 4.0 of 10

    Accountability asymmetry means autonomous AI should be governed like infrastructure, with independent review and audit, not treated as moral actors.

  19. Semantic Convergence: Investigating Shared Representations Across Scaled LLMs

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Gemma-2-2B and Gemma-2-9B align most strongly on SAE-derived features in middle layers, with preliminary evidence for shared multi-token concept subspaces.

  20. AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A multi-stakeholder standards profile that tailors NIST AI RMF and ISO/IEC 23894 guidance to general-purpose AI and foundation model developers.

  21. Technical Risks of (Lethal) Autonomous Weapons Systems

    cs.CY 2025-02 conditional novelty 3.0 of 10

    The paper restates known AI safety failure modes in the context of lethal autonomous weapons and argues that no amount of testing can guarantee control, so regulation should target the classification algorithms themselves.

  22. Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs

    cs.LG 2025-02

Pith tools