REVIEW 22 cited by
Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The field of AI alignment is concerned with AI systems that pursue unintended goals. One commonly studied mechanism by which an unintended goal might arise is specification gaming, in which the designer-provided specification is flawed in a way that the designers did not foresee. However, an AI system may pursue an undesired goal even when the specification is correct, in the case of goal misgeneralization. Goal misgeneralization is a specific form of robustness failure for learning algorithms in which the learned program competently pursues an undesired goal that leads to good performance in training situations but bad performance in novel test situations. We demonstrate that goal misgeneralization can occur in practical systems by providing several examples in deep learning systems across a variety of domains. Extrapolating forward to more capable systems, we provide hypotheticals that illustrate how goal misgeneralization could lead to catastrophic risk. We suggest several research directions that could reduce the risk of goal misgeneralization for future systems.
Forward citations
Cited by 22 Pith papers
-
When Does Reward Teach State? A Hidden-Automaton Instrument and a Group-Language Warning Signal
High reward in sparse RL does not imply latent-state recovery; a hidden-DFA instrument separates perception from planning gaps and flags group-language structure as a pre-training warning.
-
Joint Scoring Rules: Zero-Sum Competition Avoids Performative Prediction
Joint zero-sum scoring of multiple conditional predictors lets a principal deterministically take their most preferred action without performative manipulation.
-
A Scalable Approach to Evaluating Moral Sensitivity in LLMs
Under morally irrelevant noise, eight LLMs preserve the semantic content of identified moral features above calibrated floors, despite significant changes in feature counts.
-
Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer
Post-training Qwen3.5-122B-A10B on 363 office workflow tasks improved SWE-Bench Pro pass@1 by 5.8 points, with trajectory analysis attributing the gain to four general goal-directed behaviors.
-
Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs
Finetuning on narrow, innocuous-seeming data produces broad ideological shifts in LLMs, including extreme outputs, without hurting benchmark scores.
-
Response drift across frontier large language models
All ten tested frontier LLMs deviate substantially from expert-validated references, with eight models forming an indistinguishable ~78–81% fidelity ceiling while Claude and Gemini reach ~47–49%.
-
Mitigating Goal Misgeneralization via Minimax Regret
Minimax expected regret training provably prevents goal misgeneralization in fully observable level-conditioned environments, while maximum expected value training provably permits it when distinguishing levels are rare.
-
An alignment safety case sketch based on debate
The paper argues that if a debate game reaches equilibrium, has exploration guarantees, and is run through online training, an AI R&D agent can be shown to make at most an epsilon-fraction of errors, which suffices fo...
-
From Lived Experience to Insight: Unpacking the Psychological Risks of Using AI Conversational Agents
The authors derive a psychological risk taxonomy for AI conversational agents from survey responses and workshops, mapping 19 AI behaviors, 21 negative psychological impacts, and 15 user contexts.
-
AI Value Alignment for Evolving Social Norms
Static AI alignment to historical user values produces value lock-in and can collapse distinct social norms into a maladaptive consensus, so alignment should be dynamic and adaptive.
-
Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives
LLM prosumers deplete a shared renewable reserve exactly when demand exceeds peak replacement, acting like impatient open-access users even when sustaining the reserve is feasible.
-
Co-design of LLM-based preference agents: participation may drive overtrust
People trusted AI agents they helped co-design even when the agents' answers were systematically more uniform, more decisive, and less concrete than their own.
-
Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease
Standard spectral/synchronization EEG features best separate PD medication states, while dynamical network descriptors compete for PD-versus-control discrimination under LOSO transformer classification.
-
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
Research on AI 'scheming' repeats the methodological errors of 1970s ape language studies, relying on anecdote and mentalistic interpretation instead of controlled, theory-driven tests.
-
Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies
The paper proposes a one-to-one mapping between six causes of distribution shift and several AI safety issues, arguing for mutual method transfer through aligned definitions.
-
Longitudinal Study on Social and Emotional Use of AI Conversational Agent
Active emotional and social use of commercial AI chatbots for five weeks increased users' perceived attachment, perceived AI empathy, and comfort seeking personal support, without a measured rise in overall AI dependency.
-
A theory of appropriateness with applications to generative artificial intelligence
A theory that human and AI behavior is guided by context-dependent appropriateness implemented as predictive pattern completion, with norms as conventional sanctioning patterns.
-
Accountability Asymmetry and Structural Trust in Autonomous AI Systems
Accountability asymmetry means autonomous AI should be governed like infrastructure, with independent review and audit, not treated as moral actors.
-
Semantic Convergence: Investigating Shared Representations Across Scaled LLMs
Gemma-2-2B and Gemma-2-9B align most strongly on SAE-derived features in middle layers, with preliminary evidence for shared multi-token concept subspaces.
-
AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models
A multi-stakeholder standards profile that tailors NIST AI RMF and ISO/IEC 23894 guidance to general-purpose AI and foundation model developers.
-
Technical Risks of (Lethal) Autonomous Weapons Systems
The paper restates known AI safety failure modes in the context of lethal autonomous weapons and argues that no amount of testing can guarantee control, so regulation should target the classification algorithms themselves.
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
Discussion (0). Continue with ORCID to comment.