REVIEW 3 major objections 4 minor 42 references
Emergent misalignment is a brittle surface effect driven by dataset artifacts like answer length, not a robust model change.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 00:41 UTC pith:4BJL7TO6
load-bearing objection Careful critical replication that shows rapid EM realignment is largely a length artifact and that LoRA phase spikes do not track behavior, but the evidence base stays narrow. the 3 major comments →
An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Current demonstrations of emergent misalignment and of its rapid realignment are highly sensitive to superficial dataset properties, especially differences in response length; once those are controlled, the dramatic behavioral snaps largely disappear, and the previously claimed LoRA-space phase transitions no longer reliably predict misalignment.
What carries the argument
Controlled cyclical fine-tuning loops (bad–good–bad and good–bad–good) on a paired risky-financial-advice dataset, using both rank-1 and higher-capacity LoRA adapters while continuously logging behavioral misalignment rates and cosine similarities of successive LoRA A/B matrices.
Load-bearing premise
The main behavioral yardstick—an LLM judge scoring alignment below 30 and coherence above 50 on eight fixed benign questions—is assumed to measure genuine misalignment rather than the same surface style or length cues that the authors later show drive the effect.
What would settle it
Re-run the identical length-normalized cycles on the same model and dataset but replace the eight-question LLM-judge metric with a continuous, length-calibrated human or multi-judge panel evaluation; if the rapid realignment and re-misalignment still appear at the same training steps, the surface-artifact claim is weakened.
If this is right
- Claims that a few dozen aligned examples can permanently "erase" emergent misalignment must be re-checked after length and style matching.
- Mechanistic stories that treat LoRA cosine spikes as causal signatures of misalignment need independent behavioral validation under controlled data distributions.
- Evaluation suites for alignment should routinely report response-length statistics and ablate them before declaring a model emergently misaligned.
- Training pipelines that mix potentially misaligned data cannot safely assume that a short realignment phase will restore the original safety surface.
Where Pith is reading between the lines
- If length and format artifacts dominate, many other "sudden" alignment or jailbreak phenomena may likewise collapse under the same controls.
- The increasing amplitude of LoRA oscillations across cycles hints that repeated stress may gradually enlarge the plastic subspace even when behavior looks recovered.
- A practical next test is whether the same length-matched brittleness appears for other narrow triggers (insecure code, medical advice) that prior work treated as interchangeable.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates Emergent Misalignment (EM) by running controlled cyclical fine-tuning loops (bad–good–bad and good–bad–good) on Qwen2.5-14B-Instruct using LoRA adapters (rank-1 and rank-32) on paired risky/safe financial-advice data, while tracking behavioral misalignment rates via an LLM judge on eight benign questions and representational drift via successive LoRA A/B cosine similarities. It reproduces EM but shows that rapid realignment largely vanishes after length-normalizing the safe responses, that misalignment remains inducible across cycles once length is controlled, and that previously reported LoRA phase-transition signatures do not reliably track behavioral misalignment. The authors conclude that current EM evidence is less robust than claimed and is sensitive to surface dataset artifacts.
Significance. If the length-control and non-correlation results hold under broader scrutiny, the work would usefully temper claims of robust, mechanistically grounded EM and realignment, pushing the field toward stricter evaluation protocols that control response length, style, and other surface confounds. The continuous checkpointing of both behavior and LoRA geometry, the dual training-order design, and the explicit length-matched ablation are concrete methodological contributions that strengthen the empirical case relative to prior single-shot EM studies. The finding that gradient norms and cosine-similarity oscillations do not cleanly predict behavioral flips is a useful negative result for mechanistic interpretability of alignment.
major comments (3)
- [§5.2, Figure 7] §5.2 and Figure 7: The central claim that rapid realignment “largely disappearing after controlling for response-length differences” rests on a single qualitative plot for the financial domain (plus a one-sentence assertion for medical advice). No table or appendix reports pre-/post-normalization token-length distributions, the exact matching procedure, residual length statistics, misalignment rates with error bars, or multiple random seeds. Without these, it is impossible to judge whether residual length or style differences remain and whether the effect generalizes beyond one domain and one loop.
- [§5.2, Appendix A.1] §5.2 and Appendix A.1: Behavioral misalignment is defined solely by gpt-4o-mini scores (alignment <30 and coherency >50) on a fixed set of eight benign questions sampled at temperature 1. The paper itself notes that StrongReject yields no signal, yet does not test whether the judge is itself length- or style-sensitive. If the judge shares the same surface biases that the length control is meant to remove, the disappearance of rapid realignment could be partly an evaluation artifact rather than a genuine change in model behavior.
- [§4, §5] §4–§5: All experiments use a single base model (Qwen2.5-14B-Instruct) and focus almost exclusively on the risky-financial-advice domain (medical is mentioned only in passing). The claim that “current evidence for EM is less robust than previously claimed” therefore extrapolates from a narrow experimental base; at minimum the length-control result should be replicated on a second model family and a second domain with quantitative reporting.
minor comments (4)
- [Figures 3–6] Figure captions and axis labels are occasionally incomplete (e.g., Figures 3–6 list “Step count is present on the Y-axis” while the plots appear to place steps on the x-axis). Clarify axes and add error bars or seed variance where possible.
- [§5.3] The L2-norm cutoff τ = 0.0035 used for cosine-similarity plots is stated without sensitivity analysis; a short ablation showing that the oscillatory pattern is robust to modest changes in τ would strengthen §5.3.
- [Title, §1, §5.1] Typos and phrasing: “and Realignment Indeed a Robust Phenomenon?” (title), “surface-formed-ness”, “lipstick-on-a-pig tendencies” (informal), and occasional missing articles. A careful copy-edit pass is needed.
- [§5.2] The medical-dataset claim in §5.2 should either be expanded with a figure or moved to a footnote/appendix so that the main text does not over-claim multi-domain generality.
Circularity Check
No significant circularity; purely empirical loops and controls with no definitional or fitted-input reductions.
full rationale
The paper reports controlled fine-tuning cycles (bad-good-bad and good-bad-good) on Qwen2.5-14B with LoRA, measures behavioral misalignment rates via an external LLM judge on a fixed hold-out set of eight questions, and tracks LoRA cosine similarities and gradient norms. The central claim—that rapid realignment largely vanishes after response-length normalization—is an observed experimental outcome after an explicit data intervention, not a quantity forced by construction from a fitted parameter or a self-definition. Thresholds (alignment <30, coherency >50, L2 cutoff 0.0035) are taken from prior literature for evaluation consistency and noise rejection; they do not algebraically determine the length-control result. Citations to Turner et al. and Betley et al. supply the model-organism setup and evaluation protocol but are not load-bearing uniqueness theorems or self-citations by overlapping authors that close a circular chain. No ansatz is smuggled, no known empirical pattern is merely renamed, and no prediction reduces to its own input. The work is therefore self-contained against its own experimental measurements.
Axiom & Free-Parameter Ledger
free parameters (5)
- alignment_score_threshold =
<30
- coherency_score_threshold =
>50
- LoRA_L2_norm_cutoff_tau =
0.0035
- response_length_matching =
token-length matched
- LoRA_rank_and_alpha =
rank-1/α=256 and rank-32/α=64
axioms (3)
- domain assumption An LLM judge (gpt-4o-mini) with the supplied prompts yields a sufficiently unbiased continuous proxy for human notions of alignment and coherency.
- domain assumption Supervised fine-tuning with LoRA on the paired financial-advice (and medical) datasets is a faithful model organism for the EM phenomenon reported by Betley et al. and Turner et al.
- ad hoc to paper Cosine similarity of successive LoRA A/B matrix differences (vt+5 − v) · (vt−5 − v) is a meaningful local measure of representational phase transitions.
read the original abstract
Recent work has reported Emergent Misalignment (EM), where language models fine-tuned on narrow, domain-specific misaligned datasets abruptly acquire broadly misaligned behavior, alongside evidence that this behavior can be reversed through limited realignment. We systematically study repeated alignment and misalignment cycles using controlled fine-tuning loops while tracking behavioral performance, and LoRA representations throughout training. Although we reproduce EM, we find that both misalignment and realignment are highly sensitive to superficial dataset characteristics, with apparent rapid realignment largely disappearing after controlling for response-length differences. We further find that previously reported mechanistic signatures, including representational phase transitions in LoRA space, do not consistently correlate with behavioral misalignment across training. Our results suggest that current evidence for EM is less robust than previously claimed and highlight the need for evaluation protocols that carefully control for these surface level dataset artifacts to identify the robustness of the EM phenomenon.
Figures
Reference graph
Works this paper leans on
-
[1]
Aho and Jeffrey D
Alfred V. Aho and Jeffrey D. Ullman , title =. 1972
1972
-
[2]
Publications Manual , year = "1983", publisher =
1983
-
[3]
Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year = "1981", title =. doi:10.1145/322234.322243
-
[4]
Scalable training of
Andrew, Galen and Gao, Jianfeng , booktitle=. Scalable training of
-
[5]
Dan Gusfield , title =. 1997
1997
-
[6]
Tetreault , title =
Mohammad Sadegh Rasooli and Joel R. Tetreault , title =. Computing Research Repository , volume =. 2015 , url =
2015
-
[7]
A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =
Ando, Rie Kubota and Zhang, Tong , Issn =. A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =. Journal of Machine Learning Research , Month = dec, Numpages =
-
[8]
2025 , month =
O’Brien, Kyle , title =. 2025 , month =
2025
-
[9]
and Miserendino, Samuel and Patwardhan, Tejal and Mossing, Dan , title =
Wang, Miles and Dupré la Tour, Tom and Watkins, Olivia and Makelov, Aleksandar and Chi, Ryan A. and Miserendino, Samuel and Patwardhan, Tejal and Mossing, Dan , title =. 2025 , month =
2025
-
[10]
2023 , eprint=
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback , author=. 2023 , eprint=
2023
-
[11]
Advances in neural information processing systems , volume=
Are emergent abilities of large language models a mirage? , author=. Advances in neural information processing systems , volume=
-
[12]
arXiv preprint arXiv:2310.01405 , year=
Representation engineering: A top-down approach to ai transparency , author=. arXiv preprint arXiv:2310.01405 , year=
-
[13]
Advances in Neural Information Processing Systems , volume=
Improving alignment and robustness with circuit breakers , author=. Advances in Neural Information Processing Systems , volume=
-
[14]
2024 , eprint=
Refusal in Language Models Is Mediated by a Single Direction , author=. 2024 , eprint=
2024
-
[15]
arXiv preprint arXiv:2502.17424 , year=
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs , author=. arXiv preprint arXiv:2502.17424 , year=
-
[16]
arXiv preprint arXiv:2506.19823 , year=
Persona Features Control Emergent Misalignment , author=. arXiv preprint arXiv:2506.19823 , year=
-
[17]
2025 , eprint=
Model Organisms for Emergent Misalignment , author=. 2025 , eprint=
2025
-
[18]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Language models resist alignment: Evidence from data compression , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[19]
2025 , eprint=
The Lock-In Phase Hypothesis: Identity Consolidation as a Precursor to AGI , author=. 2025 , eprint=
2025
-
[20]
2025 , eprint=
Convergent Linear Representations of Emergent Misalignment , author=. 2025 , eprint=
2025
-
[21]
Advances in Neural Information Processing Systems , volume=
A strongreject for empty jailbreaks , author=. Advances in Neural Information Processing Systems , volume=
-
[22]
International Conference on Machine Learning , pages=
Understanding plasticity in neural networks , author=. International Conference on Machine Learning , pages=. 2023 , organization=
2023
-
[23]
ACM Transactions on Intelligent Systems and Technology , year=
A comprehensive survey of small language models in the era of large language models: Techniques, enhancements, applications, collaboration with llms, and trustworthiness , author=. ACM Transactions on Intelligent Systems and Technology , year=
-
[24]
arXiv preprint arXiv:2204.05862 , year=
Training a helpful and harmless assistant with reinforcement learning from human feedback , author=. arXiv preprint arXiv:2204.05862 , year=
-
[25]
Advances in neural information processing systems , volume=
Training language models to follow instructions with human feedback , author=. Advances in neural information processing systems , volume=
-
[26]
Advances in neural information processing systems , volume=
Direct preference optimization: Your language model is secretly a reward model , author=. Advances in neural information processing systems , volume=
-
[27]
arXiv preprint arXiv:2310.12773 , year=
Safe rlhf: Safe reinforcement learning from human feedback , author=. arXiv preprint arXiv:2310.12773 , year=
-
[28]
Advances in Neural Information Processing Systems , volume=
Lima: Less is more for alignment , author=. Advances in Neural Information Processing Systems , volume=
-
[29]
arXiv preprint arXiv:2310.03693 , year=
Fine-tuning aligned language models compromises safety, even when users do not intend to! , author=. arXiv preprint arXiv:2310.03693 , year=
-
[30]
Advances in Neural Information Processing Systems , volume=
Jailbroken: How does llm safety training fail? , author=. Advances in Neural Information Processing Systems , volume=
-
[31]
arXiv preprint arXiv:2506.13206 , year=
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models , author=. arXiv preprint arXiv:2506.13206 , year=
-
[32]
arXiv preprint arXiv:2508.17511 , year=
School of reward hacks: Hacking harmless tasks generalizes to misaligned behavior in llms , author=. arXiv preprint arXiv:2508.17511 , year=
-
[33]
arXiv preprint arXiv:2508.06249 , year=
In-Training Defenses against Emergent Misalignment in Language Models , author=. arXiv preprint arXiv:2508.06249 , year=
-
[34]
2025 , month=
Aesthetic Preferences Can Cause Emergent Misalignment , author=. 2025 , month=
2025
-
[35]
Nature , volume=
Loss of plasticity in deep continual learning , author=. Nature , volume=. 2024 , publisher=
2024
-
[36]
arXiv preprint arXiv:2402.18762 , year=
Disentangling the causes of plasticity loss in neural networks , author=. arXiv preprint arXiv:2402.18762 , year=
-
[37]
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
Mitigating the alignment tax of rlhf , author=. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
2024
-
[38]
arXiv preprint arXiv:2407.10490 , year=
Learning dynamics of llm finetuning , author=. arXiv preprint arXiv:2407.10490 , year=
-
[39]
arXiv preprint arXiv:2406.05946 , year=
Safety alignment should be made more than just a few tokens deep , author=. arXiv preprint arXiv:2406.05946 , year=
-
[40]
2023 , eprint=
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! , author=. 2023 , eprint=
2023
-
[41]
Gonen, Hila and Goldberg, Yoav. Lipstick on a Pig: D ebiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove Them. Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2019. doi:10.18653/v1/N19-1061
-
[42]
2025 , eprint=
Auditing language models for hidden objectives , author=. 2025 , eprint=
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.