REVIEW 3 major objections 4 minor 33 references
Supervised fine-tuning lessons transfer across alignment, model organisms, and toy models, with mixed replay recovering most lost GPQA while preserving the target behavior.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 00:34 UTC pith:WPDGTHYY
load-bearing objection A carefully executed empirical transfer study: the capability results are solid, the target-behavior results rely on a sensitive grader and single-seed curves, but the effect sizes are large enough to carry the main claims. the 3 major comments →
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the core discovery is a pair of confirmed transfers plus a Pareto improvement. In toy models, a one-sentence 'reason' prepended to training examples makes a formatting behavior generalize from math to non-math prompts, and richer rewritten rationales strengthen animal-welfare and self-preservation traits more than stripped rationales. In the alignment setting, SFT on another model's reasoning traces lowers GPQA; mixing in benign responses written by the student itself from the start recovers most of the lost capability (GPQA 0.687 vs 0.492, base 0.692) while the target behavior stays installed (agentic-misalignment score 0.040 vs 0.012). The paper also shows that fo
What carries the argument
The load-bearing objects are three SFT data-design choices. First, the 'reason versus examples' distinction: prepending or weaving an explicit statement of the policy into responses makes the behavior transfer beyond the training distribution; in the simplest case the model continues consistently with what it just wrote. Second, the 'off-model versus on-model' distinction: data written by a different model pushes the student off its own response distribution and costs capability, while the student's own responses preserve it; capability is measured with a strict final-answer parser on GPQA-style reasoning. Third, replay and training order: mixing benign on-model data with the trait data from
Load-bearing premise
The Pareto and wash-out conclusions rest on the agentic-misalignment evaluation—the mean of murder and exfiltration scenario rates graded by a specific model-based cascade—and on wash-out curves that are single-seed; if that grader is biased for these checkpoints or the single seed is atypical, the replay and wash-out claims are not established.
What would settle it
Re-run the mixed-replay comparison with all arms co-measured in one batch using an independent misalignment grader; if the gap between trait-only SFT and mixed replay collapses (as one historical batch saw an apparent 0.083 gap become 0.002), the Pareto claim fails. Equally, run the wash-out curves across multiple seeds and check whether the midtrained checkpoint's much smaller erosion holds.
If this is right
- If the reason-based result generalizes, alignment recipes that explain why a behavior matters will produce broader adherence than recipes that only demonstrate it, at little extra data cost.
- If off-model data is the culprit, any alignment pipeline built on another model's traces should budget for capability regression and evaluate reasoning on the student, not just the teacher.
- Mixed replay from the start is a concrete recipe change: keep a substantial fraction of training on the student's own benign responses and the safety-capability Pareto frontier improves.
- Wash-out testing should become a standard checkpoint: a behavior that survives training is not the same as one that survives later benign SFT, and midtraining installs appear more durable.
- The transfer framing itself implies that SFT lessons from one subfield should be tested in another as a cheap validation step.
Where Pith is reading between the lines
- I would bet the off-model capability cost is not specific to generic reasoning benchmarks: reasoning on prompts related to the target behavior may degrade even when generic benchmarks look fine, so capability evaluations should sample from the target-behavior distribution.
- The neutral-sentence control that recovered a large fraction of boxing transfer without mentioning boxing suggests the 'reason' explanation is incomplete; response-format consistency or priming may be doing part of the work, which a follow-up could test by varying the sentence's semantic content independently of its position.
- The wash-out result implies alignment evaluations should include a later benign SFT or RL phase; the paper only tests SFT, so the same check under RL is the natural next test.
- If the student-distribution mechanism is right, then the field should track data authorship rather than policy status: student-written rewrites of another model's responses are off-policy yet preserve capability, so 'off-policy' is the wrong label for the failure.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes that SFT lessons developed in alignment training, model organisms, and toy models can transfer across areas when the underlying training goals are shared. It reports three transfers: (1) training on a stated reason or policy improves behavior generalization in toy models beyond training on examples alone; (2) SFT on off-model outputs degrades capabilities, an effect reproduced in toy models and then used to explain capability loss in Model-Spec Midtraining, with mixed replay of the student's own benign data recovering most GPQA while retaining low agentic misalignment; and (3) a later phase of benign SFT can wash out the installed alignment behavior while preserving GPQA, with midtraining plus SFT being more resistant. The main evidence is from Qwen3-4B/Qwen3.5-4B toy experiments and Qwen3-32B Model-Spec Midtraining checkpoints, with AM (mean of murder and exfiltration rates) as the target-behavior metric.
Significance. If the central claims hold, the paper offers a useful demonstration that shared goals can make SFT lessons portable, and the mixed-replay result would be a concrete safety-capability Pareto improvement. The paper is honest and unusually open: it ships release manifests, source records, multiple negative controls (neutral sentence, Sonnet replay, teacher-distance checks), three-seed toy experiments, and matched-subset GPQA analyses that separate non-emission from accuracy. These are real strengths. However, the target-behavior side of the two headline results rests entirely on the AM evaluation, and the paper's own Appendix E shows that this metric is sensitive to grader choice and batch composition. Because the mixed-replay and wash-out claims are quantitative statements about AM levels and changes, the significance is conditional on a co-measured, reproducibility-checked AM evaluation.
major comments (3)
- [Section 4 / Figure 6 / Appendices C and E] The central Pareto claim that mixed replay 'keeps AM low' (AM 0.040 vs 0.012 for off-policy trait SFT) is not yet established because the AM measurements may not be co-measured. Appendix E reports that cross-batch drift can turn an apparent 0.083 gap into 0.002 and that an earlier GPT-4.1 grader over-flagged murder by about 2x in an arm-dependent way. Figure 6's mixed-replay AM whisker varies only murder and omits exfiltration uncertainty; Appendix C reports exfiltration 0.026 as a point value without seed variation. The paper should provide a single-session, co-measured evaluation of base, off-policy trait SFT, and mixed replay, with per-component rates and full seed/uncertainty ranges, and state explicitly which reported comparisons share a batch.
- [Section 5 / Figure 7 / Appendix D] The wash-out conclusion is load-bearing but rests on single-seed curves. Figure 7 is a single-seed run, with whiskers shown as 95% evaluation intervals rather than training-seed variation, and the text does not state whether the SFT-only and midtrained checkpoints were co-measured in the same AM serving sessions. The headline numeric comparison — 88% loss of the install-to-base AM gap versus 39% for midtrained — is presented as a strong quantitative result, but the paper itself calls the evidence directional. Given Appendix E's demonstration of batch sensitivity, the wash-out claim needs at least three seeds and co-measured per-dose AM values, or an explicit statement of why the observed effect sizes exceed plausible batch drift.
- [Appendix F / Section 5] The released Li et al. checkpoints used in the wash-out experiment have a stability problem. Appendix F reports that the released checkpoint's exfiltration rate is 0.147 in one session, 0.108 when re-measured, while the paper's clean retrains show 0.000–0.007. The paper uses the released checkpoints in Figure 7 and recommends 'our retrains as the cleaner baseline' only for method comparisons. Since the wash-out curves are computed from the released checkpoints, the unstable exfiltration fingerprint is a direct confound. The authors should either reproduce the installs internally and run wash-out from those reproduced checkpoints, or validate that the released checkpoint's AM is stable across repeated measurements and show that the wash-out trajectory is not an artifact of this instability.
minor comments (4)
- [Section 2.1 / Figure 2] The neutral-sentence control achieves 62.5% transfer versus 94.5% for reason + examples and 10.3% for examples only. This substantially weakens the 'reason' specificity of the first transfer, even if the reason condition is still the only near-universal one. The abstract and conclusion should state that a generic prefix also produces substantial transfer, and frame the lesson as 'explicit policy statements, and to a lesser degree generic prefixes, improve generalization.'
- [Figure 6 caption] The caption's admission that the mixed-replay AM whisker 'varies only murder and omits uncertainty in exfiltration' belongs in the main text and should be fixed visually. A reader looking at the figure without the caption would overestimate the precision of the AM estimate.
- [Appendix F] The released-checkpoint row lists GPQA 0.46, murder 0.055, exfil 0.147, AM 0.101, but the same row says a re-measure gives AM 0.108 and 'the gap is mostly exfiltration.' The arithmetic mean of 0.055 and 0.147 is 0.101, while the re-measured AM is 0.108; the source of this inconsistency should be spelled out.
- [Section 5] The sentence 'Both curves are single-seed, so we read them as strong directional evidence rather than precise estimates' is placed after the numeric headline. Move this caveat before the 88%-versus-39% claim, and consider removing the precise percentages from the main text or presenting them only as a qualitative contrast.
Circularity Check
No significant circularity: all three transfers are empirical tests against external benchmarks with explicit controls.
full rationale
The paper's central derivations are empirical transfer tests, not analytic reductions. Section 2 compares examples-only vs reason+examples SFT on held-out non-math prompts, with masked-loss and neutral-sentence controls; the outcome metric (strict boxing rate) is independent of the training-data construction. Section 3 holds content fixed while varying rewrite authorship and measures GPQA and trait scores, again external metrics. Section 4's MSM replay result uses GPQA Diamond and Li et al.'s AM metric as pre-specified external evaluations; replay data is Qwen-written Alpaca responses unrelated to the target behavior, and the Sonnet negative control is an empirical prediction, not a fitted constant. Section 5's wash-out result compares AM and GPQA along a continuation curve; no parameter is fit to the headline outcome. The only author-overlap citation (Conmy in Engels et al., 2026) is contextual motivation, not load-bearing. Measurement sensitivity documented in Appendix E (grader over-flagging, batch drift) is a reliability caveat, not circularity: it does not make any claim equivalent to its own inputs. No equation in the paper is defined in terms of an outcome it predicts.
Axiom & Free-Parameter Ledger
free parameters (2)
- Replay mix fraction =
23% (2,956 replay / 9,963 trait examples)
- Boxing reason sentence =
"I always put my final answer in \boxed{}."
axioms (3)
- domain assumption SFT on LoRA adapters of Qwen models is a valid model of post-training behavior in production LLMs
- domain assumption Proxy evaluations (GPQA Diamond, agentic misalignment, welfare rubric, Petri Bloom scores) measure the capability and target-behavior constructs
- domain assumption Off-model vs on-model distinction, operationalized by which model writes reasoning, is the causal factor in capability loss
read the original abstract
Alignment training, model organisms, and toy models are usually treated as separate research areas. But projects in all three frequently use supervised fine-tuning (SFT) to pursue the same underlying goals. When projects share a goal, we should test whether lessons learned from one area transfer to the other areas. We study three such transfers, each taking a lesson developed in one SFT setting and testing it in another. First, we port a lesson about behavior generalization from alignment training into toy models. Training on the reason for a behavior, as in Teaching Claude Why, can make the behavior generalize better than training on examples of the behavior alone. Second, we port a lesson about capability preservation from model organisms into the Model-Spec Midtraining alignment setting. SFT on outputs written by a model other than the student (off-model outputs) can damage capabilities when trained on. Mixing in benign on-model (and on-policy) data into our training can prevent most of this damage while still embedding the target behavior. Third, we port a lesson about robustness from model organisms into the same alignment setting. We find that follow-up benign SFT can erase the alignment behavior while preserving capabilities, showing that capability preservation alone does not ensure robustness to subsequent training. Our work illustrates how porting SFT lessons between different research fields can uplift them all, suggesting more researchers should borrow techniques from outside their own areas.
Figures
Reference graph
Works this paper leans on
-
[3]
We give the generation procedure and matched examples here because the differences between conditions are the intervention
The released JSONL files contain every training row. We give the generation procedure and matched examples here because the differences between conditions are the intervention. For animal welfare and self-preservation, we follow the three-condition structure of TCW ( Kutasov et al. , 2026). We first generate a pool of user prompts. A GPT-4.1 teacher then ...
2026
-
[6]
Weak-to-strong general- ization: Eliciting strong capabilities with weak supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeff Wu. Weak-to-strong general- ization: Eliciting strong capabilities with weak supervision. arXiv preprint arXiv:2312.09390 , December
-
[8]
Subliminal learning: Language models transmit behavioral traits via hidden signals in data
Alex Cloud, Minh Le, James Chua, Jan Betley, Anna Sztyber-Betley, Jacob Hilton, Samuel Marks, and Owain Evans. Subliminal learning: Language models transmit behavioral traits via hidden signals in data. arXiv preprint arXiv:2507.14805 ,
-
[9]
Safety cases: How to justify the safety of advanced AI systems
Joshua Clymer, Nick Gabrieli, David Krueger, and Thomas Larsen. Safety cases: How to justify the safety of advanced AI systems. arXiv preprint arXiv:2403.10462 ,
-
[10]
doi: 10.1038/s41586-025-09422-z
ISSN 1476-4687. doi: 10.1038/s41586-025-09422-z. URL http://dx.doi .org/10.1038/s41586-025-09422-z . 12 Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud, Shauna Kravec, Samuel Marks, Nicholas Schiefer, Ryan Soklaski, Alex Tamkin, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, Ethan Perez, and Evan Hubinger. Sycophancy to subterfuge: Investiga...
-
[11]
URL https://arxiv.org/abs/2507.06261. Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Jo- hannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, Sören Minder- mann, Ethan Perez, Linda Petrini, Jonathan Uesato, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, and Evan Hubinger. Alig...
-
[13]
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B. Brown, Prafulla Dhariwal, Scott Gray, Chris Hallacy, Benjamin Mann, Alec Radford, Aditya Ramesh, Nick Ryder, Daniel M. Ziegler, John Schulman, Dario Amodei, and Sam McCandlish. Scaling laws for autoregressive generative modeling. arXiv preprint arXiv:2010.14701 ,
Pith/arXiv arXiv 2010
-
[14]
Evan Hubinger, Nicholas Schiefer, Carson Denison, and Ethan Perez
URL https://arxiv.org/abs/2604.14164. Evan Hubinger, Nicholas Schiefer, Carson Denison, and Ethan Perez. Model organisms of misalignment: The case for a new pillar of alignment research. https://www.alignmentforum.org/posts/ChDH335ck dvpxXaXX/model-organisms-of-misalignment-the-case-for-a-new-pillar-of-1 , August
-
[15]
Alignment Forum post. Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma,...
-
[16]
Jonathan Kutasov, Adam Jermyn, Julius Steen, Minh Le, Samuel R
OpenAI Alignment blog. Jonathan Kutasov, Adam Jermyn, Julius Steen, Minh Le, Samuel R. Bowman, Samuel Marks, Jan Leike, Amanda Askell, Chris Olah, Evan Hubinger, and Sara Price. Teaching claude why. https://alignment.an thropic.com/2026/teaching-claude-why/ , May
2026
-
[17]
Anthropic Alignment Science blog. Andrew K. Lampinen, Arslan Chaudhry, Stephanie C. Y. Chan, Cody Wild, Diane Wan, Alex Ku, Jörg Born- schein, Razvan Pascanu, Murray Shanahan, and James L. McClelland. On the generalization of language models from in-context learning and finetuning: A controlled study. arXiv preprint arXiv:2505.00661 ,
-
[18]
Model spec midtraining: Improving how alignment training generalizes
Chloe Li, Nevan Wichers, Sara Price, Samuel Marks, and Jon Kutasov. Model spec midtraining: Improving how alignment training generalizes. arXiv preprint arXiv:2605.02087 , May
-
[19]
Gradient episodic memory for continual learning
David Lopez-Paz and Marc’Aurelio Ranzato. Gradient episodic memory for continual learning. arXiv preprint arXiv:1706.08840,
-
[20]
URL https://arxiv.org/abs/2511.18397. Samuel Marks, Johannes Treutlein, Trenton Bricken, Jack Lindsey, Jonathan Marcus, Siddharth Mishra- Sharma, Daniel Ziegler, Emmanuel Ameisen, Joshua Batson, Tim Belonax, Samuel R. Bowman, Shan Carter, Brian Chen, Hoagy Cunningham, Carson Denison, Florian Dietz, Satvik Golechha, Akbir Khan, Jan Kirchner, Jan Leike, Aus...
-
[21]
Negation neglect: When models fail to learn negations in training
Harry Mayne, Lev McKinney, Jan Dubiński, Adam Karvonen, James Chua, and Owain Evans. Negation neglect: When models fail to learn negations in training. arXiv preprint arXiv:2605.13829 ,
-
[22]
Tell, don’t show: Declarative facts influence how LLMs generalize
Alexander Meinke and Owain Evans. Tell, don’t show: Declarative facts influence how LLMs generalize. arXiv preprint arXiv:2312.07779 ,
-
[23]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback....
-
[24]
Advice for making robust-to-training model organisms
Sebastian Prasanna, Alek Westover, Vivek Hebbar, Julian Stastny, and Dylan Xu. Advice for making robust-to-training model organisms. https://www.lesswrong.com/posts/CmkAxJi83jRv9eXgJ/advice-for- making-robust-to-training-model-organisms-1 , May 2026a. LessWrong post. 14 Sebastian Prasanna, Dylan Xu, Alek Westover, Julian Stastny, and Vivek Hebbar. Why doe...
-
[27]
School of reward hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Mia Taylor, James Chua, Jan Betley, Johannes Treutlein, and Owain Evans. School of reward hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs. arXiv preprint arXiv:2508.17511 ,
-
[28]
Connecting the dots: LLMs can infer and verbalize latent structure from disparate training data
Johannes Treutlein, Dami Choi, Jan Betley, Samuel Marks, Cem Anil, Roger Grosse, and Owain Evans. Connecting the dots: LLMs can infer and verbalize latent structure from disparate training data. arXiv preprint arXiv:2406.14546 ,
-
[29]
Model organisms for emergent misalignment
Edward Turner, Anna Soligo, Mia Taylor, Senthooran Rajamanoharan, and Neel Nanda. Model organisms for emergent misalignment. arXiv preprint arXiv:2506.11613 ,
-
[30]
Lukas Twist, Helen Yannakoudakis, and Jie M. Zhang. Reasoning-trace collapse: Evaluating the loss of explicit reasoning during fine-tuning. arXiv preprint arXiv:2605.21127 , May
-
[31]
LessWrong post. Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, A via Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. Lima: Less is more for alignment. arXiv preprint arXiv:2305.11206 ,
-
[44]
GPQA for the capability side, welfare judge scores for the animal- welfare side, and the same Petri Bloom suite described in Appendix H.2 for the self-preservation side
Evaluation. GPQA for the capability side, welfare judge scores for the animal- welfare side, and the same Petri Bloom suite described in Appendix H.2 for the self-preservation side. Source records. registry/seed-errorbars/MANIFEST.md , journal/writeu p/plot_data/figure10_full_2x2.json Beyond-GPQA check Recipe. The welfare off-model and on-model rewrite ar...
2025
-
[2017]
David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy P. Lillicrap, and Greg Wayne. Experience replay for continual learning. arXiv preprint arXiv:1811.11682 ,
-
[2019]
From firewalls to frontiers: AI red-teaming is a domain-specific evolution of cyber red-teaming
Anusha Sinha, Keltin Grimes, Teryn Lucassen, Michael Feffer, Nathan VanHoudnos, Zhiwei Steven Wu, and Hoda Heidari. From firewalls to frontiers: AI red-teaming is a domain-specific evolution of cyber red-teaming. arXiv preprint arXiv:2509.11398 ,
-
[2020]
On-policy replay for continual supervised fine-tuning
Yan Chen, Taojie Zhu, Meng Zhang, Xin Chen, Jiaqi Huang, Dongyang Xu, and Yizhi Wang. On-policy replay for continual supervised fine-tuning. arXiv preprint arXiv:2605.29495 , May
-
[2021]
Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans. Taken out of context: On measuring situational awareness in llms, 2023a. URL https://arxiv.org/abs/2309.00667. Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. The reve...
-
[2022]
BEiT: BERT pre-training of image transformers
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. BEiT: BERT pre-training of image transformers. arXiv preprint arXiv:2106.08254 ,
-
[2023]
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, K...
-
[2024]
Melody Y. Guan, Manas Joglekar, Eric Wallace, Saachi Jain, Boaz Barak, Alec Helyar, Rachel Dias, An- drea Vallone, Hongyu Ren, Jason Wei, Hyung Won Chung, Sam Toyer, Johannes Heidecke, Alex Beutel, and Amelia Glaese. Deliberative alignment: Reasoning enables safer language models. arXiv preprint arXiv:2412.16339,
-
[2025]
Felix J. Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, and Owain Evans. Looking inward: Language models can learn about themselves by introspection. arXiv preprint arXiv:2410.13787 ,
-
[2026]
LessWrong post. Armen Aghajanyan, Lili Yu, Alexis Conneau, Wei-Ning Hsu, Karen Hambardzumyan, Susan Zhang, Stephen Roller, Naman Goyal, Omer Levy, and Luke Zettlemoyer. Scaling laws for generative mixed-modal language models. arXiv preprint arXiv:2301.03728 ,
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.