REVIEW 14 cited by
Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The field of AI alignment is concerned with AI systems that pursue unintended goals. One commonly studied mechanism by which an unintended goal might arise is specification gaming, in which the designer-provided specification is flawed in a way that the designers did not foresee. However, an AI system may pursue an undesired goal even when the specification is correct, in the case of goal misgeneralization. Goal misgeneralization is a specific form of robustness failure for learning algorithms in which the learned program competently pursues an undesired goal that leads to good performance in training situations but bad performance in novel test situations. We demonstrate that goal misgeneralization can occur in practical systems by providing several examples in deep learning systems across a variety of domains. Extrapolating forward to more capable systems, we provide hypotheticals that illustrate how goal misgeneralization could lead to catastrophic risk. We suggest several research directions that could reduce the risk of goal misgeneralization for future systems.
Forward citations
Cited by 14 Pith papers
-
When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary
High reward in sparse RL does not imply latent-state recovery; a hidden-DFA instrument separates perception from planning gaps and flags group-language structure as a pre-training warning.
-
A Scalable Approach to Evaluating Moral Sensitivity in LLMs
Under morally irrelevant noise, eight LLMs preserve the semantic content of identified moral features above calibrated floors, despite significant changes in feature counts.
-
Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs
Finetuning on narrow, innocuous-seeming data produces broad ideological shifts in LLMs, including extreme outputs, without hurting benchmark scores.
-
Response drift across frontier large language models
All ten tested frontier LLMs deviate substantially from expert-validated references, with eight models forming an indistinguishable ~78–81% fidelity ceiling while Claude and Gemini reach ~47–49%.
-
Mitigating Goal Misgeneralization via Minimax Regret
Minimax expected regret training provably prevents goal misgeneralization in fully observable level-conditioned environments, while maximum expected value training provably permits it when distinguishing levels are rare.
-
AI Value Alignment for Evolving Social Norms
Static AI alignment to historical user values produces value lock-in and can collapse distinct social norms into a maladaptive consensus, so alignment should be dynamic and adaptive.
-
Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives
LLM prosumers deplete a shared renewable reserve exactly when demand exceeds peak replacement, acting like impatient open-access users even when sustaining the reserve is feasible.
-
Co-design of LLM-based preference agents: participation may drive overtrust
People trusted AI agents they helped co-design even when the agents' answers were systematically more uniform, more decisive, and less concrete than their own.
-
Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease
Standard spectral/synchronization EEG features best separate PD medication states, while dynamical network descriptors compete for PD-versus-control discrimination under LOSO transformer classification.
-
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
Research on AI 'scheming' repeats the methodological errors of 1970s ape language studies, relying on anecdote and mentalistic interpretation instead of controlled, theory-driven tests.
-
Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies
The paper proposes a one-to-one mapping between six causes of distribution shift and several AI safety issues, arguing for mutual method transfer through aligned definitions.
-
Semantic Convergence: Investigating Shared Representations Across Scaled LLMs
Gemma-2-2B and Gemma-2-9B align most strongly on SAE-derived features in middle layers, with preliminary evidence for shared multi-token concept subspaces.
-
AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models
A multi-stakeholder standards profile that tailors NIST AI RMF and ISO/IEC 23894 guidance to general-purpose AI and foundation model developers.
-
Technical Risks of (Lethal) Autonomous Weapons Systems
The paper restates known AI safety failure modes in the context of lethal autonomous weapons and argues that no amount of testing can guarantee control, so regulation should target the classification algorithms themselves.
Discussion (0). Sign in to comment.