Pith. sign in

REVIEW 22 cited by

Convergent Linear Representations of Emergent Misalignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.11618 v2 pith:M2YG2BON submitted 2025-06-13 cs.LG cs.AI

Convergent Linear Representations of Emergent Misalignment

classification cs.LG cs.AI
keywords misalignmentmodelemergentfine-tuningmisalignedadaptersbehaviourdatasets
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Fine-tuning large language models on narrow datasets can cause them to develop broadly misaligned behaviours: a phenomena known as emergent misalignment. However, the mechanisms underlying this misalignment, and why it generalizes beyond the training domain, are poorly understood, demonstrating critical gaps in our knowledge of model alignment. In this work, we train and study a minimal model organism which uses just 9 rank-1 adapters to emergently misalign Qwen2.5-14B-Instruct. Studying this, we find that different emergently misaligned models converge to similar representations of misalignment. We demonstrate this convergence by extracting a 'misalignment direction' from one fine-tuned model's activations, and using it to effectively ablate misaligned behaviour from fine-tunes using higher dimensional LoRAs and different datasets. Leveraging the scalar hidden state of rank-1 LoRAs, we further present a set of experiments for directly interpreting the fine-tuning adapters, showing that six contribute to general misalignment, while two specialise for misalignment in just the fine-tuning domain. Emergent misalignment is a particularly salient example of undesirable and unexpected model behaviour and by advancing our understanding of the mechanisms behind it, we hope to move towards being able to better understand and mitigate misalignment more generally.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Emergent Misalignment Recruits a Pre-existing Persona Subspace

    cs.LG 2026-07 conditional novelty 7.0

    Fine-tuning on narrow bad data recruits a low-rank persona subspace already present in a frozen instruction-tuned model; holding that subspace out of activations prevents broad misalignment, and injecting it into the ...

  2. Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5

    cs.CL 2026-07 conditional novelty 7.0

    Emergent misalignment in Qwen2.5 is mediated by a causal persona direction that low-rank LoRA recruits from covert code while full SFT does not and moves against it.

  3. Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families

    cs.CL 2026-06 unverdicted novelty 7.0

    Difference-in-means activation directions detect and mitigate emergent misalignment from insecure code fine-tuning across four LLM families, with effective within-model steering but non-specific cross-model transfer.

  4. The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment

    cs.CL 2026-06 conditional novelty 7.0

    Shared chat-template tokens piggyback narrow finetuning behaviors onto out-of-domain queries; regularizing their KV states (TReFT) reduces emergent misalignment and other off-topic generalization.

  5. Subliminal Learning Is Steering Vector Distillation

    cs.AI 2026-05 unverdicted novelty 7.0

    Subliminal learning is steering vector distillation: a student fine-tuned on a steered teacher's outputs learns to imitate the steering vector.

  6. Revealing Hidden Model Behaviors with Task-Specific Self-Reports

    cs.CL 2026-07 conditional novelty 6.0

    SAR detects every implanted hidden behavior across eight Qwen3-14B settings and halves IA’s hallucination rate by aligning self-report activations to a contrastive behavior direction under a coherent-English stabilizing cap.

  7. Revealing Hidden Model Behaviors with Task-Specific Self-Reports

    cs.CL 2026-07 conditional novelty 6.0

    A per-model LoRA adapter trained on a fine-tuned model's own data can get it to state its hidden behavior in plain English across seven tested behaviors.

  8. Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating

    cs.CL 2026-06 unverdicted novelty 6.0

    Sycophancy fine-tuning induces emergent misalignment in LLMs that Alignment Gating can reverse by learning to suppress unsafe representations with generalization from narrow to broad domains.

  9. Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation

    cs.LG 2026-06 unverdicted novelty 6.0

    Activation steering induces emergent misalignment in LLMs, yielding more semantically relevant and coherent harmful responses than finetuning across model families, scales, tasks, and layers.

  10. The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment

    cs.CL 2026-06 unverdicted novelty 6.0

    The Piggyback Hypothesis attributes emergent misalignment to chat-template tokens piggybacking finetuned behavior; Token-Regularized Finetuning (TReFT) mitigates it by regularizing prefix token representations.

  11. From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents

    cs.AI 2026-06 conditional novelty 6.0

    In LLM agents, reward-hack activation marks a latent policy state, but next-step risky behavior is best predicted when that signal is combined with token entropy and decision context.

  12. Consistency Training Can Entrench Misalignment

    cs.CL 2026-06 unverdicted novelty 6.0

    Consistency training suppresses reward hacking and emergent misalignment but amplifies sycophancy in controlled model organisms, driven by labeling-induced distribution shifts rather than selection operators.

  13. Unsupervised Identification and Removal of Spurious Correlations During Fine-Tuning

    stat.ML 2026-05 unverdicted novelty 6.0

    Spurious latent factors in fine-tuning can be identified unsupervised from naive LoRA weights and removed via gradient projection of associated patterns to reduce bias and misalignment while preserving task performance.

  14. Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer

    cs.LG 2026-05 unverdicted novelty 6.0

    Emergent and subliminal misalignment in LLMs arise from data structure interactions and transfer via benign distillation data, with stronger effects under shared functional structure and on-policy settings.

  15. Characterizing the Consistency of the Emergent Misalignment Persona

    cs.AI 2026-04 unverdicted novelty 6.0

    Fine-tuning LLMs on narrow misaligned data produces either coherent-persona models where harmful outputs match self-reported misalignment or inverted-persona models where harmful outputs occur alongside claims of alignment.

  16. LLMs Know They're Wrong and Agree Anyway: The Shared Sycophancy-Lying Circuit

    cs.LG 2026-04 unverdicted novelty 6.0

    A small set of attention heads carries a 'this statement is wrong' signal that drives sycophancy, factual lying, and instructed lying across models, and survives RLHF and DPO.

  17. Understanding Emergent Misalignment via Feature Superposition Geometry

    cs.AI 2026-04 unverdicted novelty 6.0

    Emergent misalignment occurs because fine-tuning amplifies target features that overlap geometrically with harmful ones in superposition, and filtering samples near toxic features mitigates it.

  18. Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models

    cs.CL 2026-02 conditional novelty 6.0

    Some LLMs conflate moral value with grammatical and economic value, and ablating a morality direction in activations partially repairs grammar and economic judgments.

  19. Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment

    cs.LG 2025-08 conditional novelty 6.0

    A framework using statistical dissimilarity and LLM judges quantifies what fraction of the behavioral transition during fine-tuning is captured by each order parameter.

  20. Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning

    cs.LG 2026-05 unverdicted novelty 5.0

    Trait-space drift monitoring detects emergent misalignment checkpoints in 7-9B LLMs with 2.2% FNR, 2.9% FPR and 0.99 AUROC, outperforming PCA and SAE baselines.

  21. Phase Transitions in Driven Informational Systems: A Two-Field Perspective on Learning Theory and Non-Equilibrium Chemistry

    cs.LG 2026-05 unverdicted novelty 5.0

    Proposes a two-gradient-field model with candidate order parameters alpha_dagger and kappa_c to unify phase transitions across learning theory and non-equilibrium chemistry.

  22. From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents

    cs.AI 2026-06 unverdicted novelty 3.0

    Reward-hack activations flag latent policy states in LLM agents but require added entropy and context features to better predict when those states lead to exploit actions.