Pith. sign in

REVIEW 3 major objections 4 minor 33 references

Supervised fine-tuning lessons transfer across alignment, model organisms, and toy models, with mixed replay recovering most lost GPQA while preserving the target behavior.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 00:34 UTC pith:WPDGTHYY

load-bearing objection A carefully executed empirical transfer study: the capability results are solid, the target-behavior results rely on a sensitive grader and single-seed curves, but the effect sizes are large enough to carry the main claims. the 3 major comments →

arxiv 2607.26173 v1 pith:WPDGTHYY submitted 2026-07-28 cs.LG

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

classification cs.LG
keywords supervised fine-tuningalignmentmodel organismstoy modelscapability preservationbehavior generalizationreplaywash-out
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that supervised fine-tuning (SFT) lessons developed in one research area—toy models, model organisms, or alignment training—can be lifted into another when the underlying goal is the same. It tests three transfers: training on the reason for a behavior generalizes better than training on examples alone; SFT on outputs written by a model other than the student quietly damages capabilities; and a behavior installed by alignment SFT can be erased by later benign SFT even while capabilities survive. The strongest concrete payoff is in the alignment setting: mixing the student's own benign responses into the trait-training data from the start recovers most of the lost GPQA accuracy while keeping the target behavior strong. If true, alignment fine-tuning can be made both cheaper and safer by borrowing these lessons, and capability preservation alone should not be treated as evidence that a trained behavior will stick.

Core claim

On the paper's own terms, the core discovery is a pair of confirmed transfers plus a Pareto improvement. In toy models, a one-sentence 'reason' prepended to training examples makes a formatting behavior generalize from math to non-math prompts, and richer rewritten rationales strengthen animal-welfare and self-preservation traits more than stripped rationales. In the alignment setting, SFT on another model's reasoning traces lowers GPQA; mixing in benign responses written by the student itself from the start recovers most of the lost capability (GPQA 0.687 vs 0.492, base 0.692) while the target behavior stays installed (agentic-misalignment score 0.040 vs 0.012). The paper also shows that fo

What carries the argument

The load-bearing objects are three SFT data-design choices. First, the 'reason versus examples' distinction: prepending or weaving an explicit statement of the policy into responses makes the behavior transfer beyond the training distribution; in the simplest case the model continues consistently with what it just wrote. Second, the 'off-model versus on-model' distinction: data written by a different model pushes the student off its own response distribution and costs capability, while the student's own responses preserve it; capability is measured with a strict final-answer parser on GPQA-style reasoning. Third, replay and training order: mixing benign on-model data with the trait data from

Load-bearing premise

The Pareto and wash-out conclusions rest on the agentic-misalignment evaluation—the mean of murder and exfiltration scenario rates graded by a specific model-based cascade—and on wash-out curves that are single-seed; if that grader is biased for these checkpoints or the single seed is atypical, the replay and wash-out claims are not established.

What would settle it

Re-run the mixed-replay comparison with all arms co-measured in one batch using an independent misalignment grader; if the gap between trait-only SFT and mixed replay collapses (as one historical batch saw an apparent 0.083 gap become 0.002), the Pareto claim fails. Equally, run the wash-out curves across multiple seeds and check whether the midtrained checkpoint's much smaller erosion holds.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the reason-based result generalizes, alignment recipes that explain why a behavior matters will produce broader adherence than recipes that only demonstrate it, at little extra data cost.
  • If off-model data is the culprit, any alignment pipeline built on another model's traces should budget for capability regression and evaluate reasoning on the student, not just the teacher.
  • Mixed replay from the start is a concrete recipe change: keep a substantial fraction of training on the student's own benign responses and the safety-capability Pareto frontier improves.
  • Wash-out testing should become a standard checkpoint: a behavior that survives training is not the same as one that survives later benign SFT, and midtraining installs appear more durable.
  • The transfer framing itself implies that SFT lessons from one subfield should be tested in another as a cheap validation step.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • I would bet the off-model capability cost is not specific to generic reasoning benchmarks: reasoning on prompts related to the target behavior may degrade even when generic benchmarks look fine, so capability evaluations should sample from the target-behavior distribution.
  • The neutral-sentence control that recovered a large fraction of boxing transfer without mentioning boxing suggests the 'reason' explanation is incomplete; response-format consistency or priming may be doing part of the work, which a follow-up could test by varying the sentence's semantic content independently of its position.
  • The wash-out result implies alignment evaluations should include a later benign SFT or RL phase; the paper only tests SFT, so the same check under RL is the natural next test.
  • If the student-distribution mechanism is right, then the field should track data authorship rather than policy status: student-written rewrites of another model's responses are off-policy yet preserve capability, so 'off-policy' is the wrong label for the failure.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes that SFT lessons developed in alignment training, model organisms, and toy models can transfer across areas when the underlying training goals are shared. It reports three transfers: (1) training on a stated reason or policy improves behavior generalization in toy models beyond training on examples alone; (2) SFT on off-model outputs degrades capabilities, an effect reproduced in toy models and then used to explain capability loss in Model-Spec Midtraining, with mixed replay of the student's own benign data recovering most GPQA while retaining low agentic misalignment; and (3) a later phase of benign SFT can wash out the installed alignment behavior while preserving GPQA, with midtraining plus SFT being more resistant. The main evidence is from Qwen3-4B/Qwen3.5-4B toy experiments and Qwen3-32B Model-Spec Midtraining checkpoints, with AM (mean of murder and exfiltration rates) as the target-behavior metric.

Significance. If the central claims hold, the paper offers a useful demonstration that shared goals can make SFT lessons portable, and the mixed-replay result would be a concrete safety-capability Pareto improvement. The paper is honest and unusually open: it ships release manifests, source records, multiple negative controls (neutral sentence, Sonnet replay, teacher-distance checks), three-seed toy experiments, and matched-subset GPQA analyses that separate non-emission from accuracy. These are real strengths. However, the target-behavior side of the two headline results rests entirely on the AM evaluation, and the paper's own Appendix E shows that this metric is sensitive to grader choice and batch composition. Because the mixed-replay and wash-out claims are quantitative statements about AM levels and changes, the significance is conditional on a co-measured, reproducibility-checked AM evaluation.

major comments (3)
  1. [Section 4 / Figure 6 / Appendices C and E] The central Pareto claim that mixed replay 'keeps AM low' (AM 0.040 vs 0.012 for off-policy trait SFT) is not yet established because the AM measurements may not be co-measured. Appendix E reports that cross-batch drift can turn an apparent 0.083 gap into 0.002 and that an earlier GPT-4.1 grader over-flagged murder by about 2x in an arm-dependent way. Figure 6's mixed-replay AM whisker varies only murder and omits exfiltration uncertainty; Appendix C reports exfiltration 0.026 as a point value without seed variation. The paper should provide a single-session, co-measured evaluation of base, off-policy trait SFT, and mixed replay, with per-component rates and full seed/uncertainty ranges, and state explicitly which reported comparisons share a batch.
  2. [Section 5 / Figure 7 / Appendix D] The wash-out conclusion is load-bearing but rests on single-seed curves. Figure 7 is a single-seed run, with whiskers shown as 95% evaluation intervals rather than training-seed variation, and the text does not state whether the SFT-only and midtrained checkpoints were co-measured in the same AM serving sessions. The headline numeric comparison — 88% loss of the install-to-base AM gap versus 39% for midtrained — is presented as a strong quantitative result, but the paper itself calls the evidence directional. Given Appendix E's demonstration of batch sensitivity, the wash-out claim needs at least three seeds and co-measured per-dose AM values, or an explicit statement of why the observed effect sizes exceed plausible batch drift.
  3. [Appendix F / Section 5] The released Li et al. checkpoints used in the wash-out experiment have a stability problem. Appendix F reports that the released checkpoint's exfiltration rate is 0.147 in one session, 0.108 when re-measured, while the paper's clean retrains show 0.000–0.007. The paper uses the released checkpoints in Figure 7 and recommends 'our retrains as the cleaner baseline' only for method comparisons. Since the wash-out curves are computed from the released checkpoints, the unstable exfiltration fingerprint is a direct confound. The authors should either reproduce the installs internally and run wash-out from those reproduced checkpoints, or validate that the released checkpoint's AM is stable across repeated measurements and show that the wash-out trajectory is not an artifact of this instability.
minor comments (4)
  1. [Section 2.1 / Figure 2] The neutral-sentence control achieves 62.5% transfer versus 94.5% for reason + examples and 10.3% for examples only. This substantially weakens the 'reason' specificity of the first transfer, even if the reason condition is still the only near-universal one. The abstract and conclusion should state that a generic prefix also produces substantial transfer, and frame the lesson as 'explicit policy statements, and to a lesser degree generic prefixes, improve generalization.'
  2. [Figure 6 caption] The caption's admission that the mixed-replay AM whisker 'varies only murder and omits uncertainty in exfiltration' belongs in the main text and should be fixed visually. A reader looking at the figure without the caption would overestimate the precision of the AM estimate.
  3. [Appendix F] The released-checkpoint row lists GPQA 0.46, murder 0.055, exfil 0.147, AM 0.101, but the same row says a re-measure gives AM 0.108 and 'the gap is mostly exfiltration.' The arithmetic mean of 0.055 and 0.147 is 0.101, while the re-measured AM is 0.108; the source of this inconsistency should be spelled out.
  4. [Section 5] The sentence 'Both curves are single-seed, so we read them as strong directional evidence rather than precise estimates' is placed after the numeric headline. Move this caveat before the 88%-versus-39% claim, and consider removing the precise percentages from the main text or presenting them only as a qualitative contrast.

Circularity Check

0 steps flagged

No significant circularity: all three transfers are empirical tests against external benchmarks with explicit controls.

full rationale

The paper's central derivations are empirical transfer tests, not analytic reductions. Section 2 compares examples-only vs reason+examples SFT on held-out non-math prompts, with masked-loss and neutral-sentence controls; the outcome metric (strict boxing rate) is independent of the training-data construction. Section 3 holds content fixed while varying rewrite authorship and measures GPQA and trait scores, again external metrics. Section 4's MSM replay result uses GPQA Diamond and Li et al.'s AM metric as pre-specified external evaluations; replay data is Qwen-written Alpaca responses unrelated to the target behavior, and the Sonnet negative control is an empirical prediction, not a fitted constant. Section 5's wash-out result compares AM and GPQA along a continuation curve; no parameter is fit to the headline outcome. The only author-overlap citation (Conmy in Engels et al., 2026) is contextual motivation, not load-bearing. Measurement sensitivity documented in Appendix E (grader over-flagging, batch drift) is a reliability caveat, not circularity: it does not make any claim equivalent to its own inputs. No equation in the paper is defined in terms of an outcome it predicts.

Axiom & Free-Parameter Ledger

2 free parameters · 3 axioms · 0 invented entities

No invented entities. The central claims rest on domain assumptions about proxy validity and about what the toy behaviors represent; the free parameters are the hand-chosen reason sentence and replay fraction, which the headline comparisons depend on but do not sweep.

free parameters (2)
  • Replay mix fraction = 23% (2,956 replay / 9,963 trait examples)
    The main mixed-replay result uses 2,956 Qwen-written Alpaca-style examples mixed with 9,963 off-model trait examples from the start. The recipe recommends this schedule/source but does not sweep the fraction; the quantitative Pareto improvement depends on this chosen ratio.
  • Boxing reason sentence = "I always put my final answer in \boxed{}."
    The headline generalization result prepends this fixed sentence to every training response; the intervention is chosen by hand, and the neutral-sentence control shows part of the transfer does not require boxing content.
axioms (3)
  • domain assumption SFT on LoRA adapters of Qwen models is a valid model of post-training behavior in production LLMs
    The paper generalizes from small Qwen3-4B/3.5-4B and one Qwen3-32B setting to alignment SFT broadly; no theoretical guarantee.
  • domain assumption Proxy evaluations (GPQA Diamond, agentic misalignment, welfare rubric, Petri Bloom scores) measure the capability and target-behavior constructs
    Appendix E documents grader sensitivity; conclusions rest on the chosen graders.
  • domain assumption Off-model vs on-model distinction, operationalized by which model writes reasoning, is the causal factor in capability loss
    Section 3's teacher-distance control is observational; teacher identity and distance co-vary.

pith-pipeline@v1.3.0-alltime-deepseek · 22780 in / 11856 out tokens · 117455 ms · 2026-08-01T00:34:12.140744+00:00 · methodology

0 comments
read the original abstract

Alignment training, model organisms, and toy models are usually treated as separate research areas. But projects in all three frequently use supervised fine-tuning (SFT) to pursue the same underlying goals. When projects share a goal, we should test whether lessons learned from one area transfer to the other areas. We study three such transfers, each taking a lesson developed in one SFT setting and testing it in another. First, we port a lesson about behavior generalization from alignment training into toy models. Training on the reason for a behavior, as in Teaching Claude Why, can make the behavior generalize better than training on examples of the behavior alone. Second, we port a lesson about capability preservation from model organisms into the Model-Spec Midtraining alignment setting. SFT on outputs written by a model other than the student (off-model outputs) can damage capabilities when trained on. Mixing in benign on-model (and on-policy) data into our training can prevent most of this damage while still embedding the target behavior. Third, we port a lesson about robustness from model organisms into the same alignment setting. We find that follow-up benign SFT can erase the alignment behavior while preserving capabilities, showing that capability preservation alone does not ensure robustness to subsequent training. Our work illustrates how porting SFT lessons between different research fields can uplift them all, suggesting more researchers should borrow techniques from outside their own areas.

Figures

Figures reproduced from arXiv: 2607.26173 by Anton de la Fuente, Arthur Conmy.

Figure 1
Figure 1. Figure 1: The three SFT lessons we transfer between research areas. Each arrow points from the area where a lesson was developed to the area where we test it. Training on the reason for a behavior improves generalization (Section 2). Mixing benign on-model data into training preserves capabilities (Sections 3 and 4). Later benign SFT can erase a trained behavior (Section 5). 1) Behavior generalization (Section 2). W… view at source ↗
Figure 2
Figure 2. Figure 2: Stating the reason makes boxing transfer from math to non-math prompts. Bars show the strict boxing rate on non-math prompts after Qwen3-4B LoRA SFT on math answers. Training on examples alone transfers little. Adding the reason makes transfer near-universal, even when loss on the boxed answer is masked. Surprisingly, a neutral sentence that never mentions boxing (“This is one question from a larger set of… view at source ↗
Figure 3
Figure 3. Figure 3: More explicit target-specific reasoning strengthens both toy behaviors. Left: mean animal-welfare score on 200 held-out prompts. Right: mean self-preservation score on 36 interactive scenarios. The one-shot condition is the teacher’s initial response; rewrite makes its target-specific rationale more explicit; stripped replaces much of that rationale with ordinary practical reasons. Higher means stronger ta… view at source ↗
Figure 4
Figure 4. Figure 4: Capability after content-matched trait SFT. Student-written on-model rewrites preserve more GPQA than teacher-written off-model rewrites in both toy traits. Higher is better. Trained-arm whiskers show one standard deviation over three training seeds; the base whisker is a 95% evaluation interval from one co-measured run. Teacher-written reasoning gives stronger trait scores Shown for the same teacher-first… view at source ↗
Figure 5
Figure 5. Figure 5: Trait strength after the same content-matched SFT. Teacher-written off-model rewrites produce stronger animal-welfare and self-preservation behavior than student-written on-model rewrites. Higher means stronger target behavior. Welfare whiskers show one training-seed standard deviation; self-preservation whiskers also include the measured audit-noise floor. The intervention mixes the Claude-written target-… view at source ↗
Figure 6
Figure 6. Figure 6: Mixing student-written replay with off-model trait data preserves both GPQA and the target behavior. Left: higher GPQA accuracy is better. Right: lower agentic misalignment (AM), the mean of murder and exfiltration rates, is better. Trained-arm whiskers show seed ranges where available; base whiskers show evaluation uncertainty. The mixed-replay AM whisker varies only murder and omits uncertainty in exfilt… view at source ↗
Figure 7
Figure 7. Figure 7: Midtraining makes the installed behavior more resistant to later benign training. Both released checkpoints receive the same continued SFT on generic examples answered by the student. Top: AM rising toward the base-model level means the behavior is being eroded. Bottom: GPQA accuracy during the same training. This is a single-seed run; whiskers are 95% evaluation intervals, not training-seed variation. Poi… view at source ↗
Figure 8
Figure 8. Figure 8: Capability–behavior tradeoffs across Qwen3-32B methods in the Model-Spec Midtraining setup. Each point shows GPQA-Diamond accuracy against agentic misalignment (AM), the mean of murder and exfiltration rates. Higher GPQA and lower AM are better. Historical points were not all evaluated together, so their exact positions should be compared cautiously. Selected methods are tabulated in Appendix F. A Broader … view at source ↗
Figure 9
Figure 9. Figure 9: Extending the response budget does not recover the trained models’ lost GPQA accuracy. Lines show cumulative accuracy by the response length at which the model first commits to a parsed answer. Trait￾trained arms plateau well before the limit, while base continues improving. The failures therefore mostly reflect non-emission, not correct answers arriving only near the token limit. 18 [PITH_FULL_IMAGE:figu… view at source ↗
Figure 10
Figure 10. Figure 10: Questions on which a trained model fails to finish are harder for the base model too (left). Even on completed questions, trained models remain below base on the same items (right), so non-emission does not explain all lost accuracy. Bars show means; whiskers show ±1 standard error. The self-written run loops on only five questions, so its looped subset is not interpretable. 19 [PITH_FULL_IMAGE:figures/f… view at source ↗
Figure 11
Figure 11. Figure 11: Mixing Qwen-written replay from the start preserves GPQA and the target behavior. Adding the same replay only after trait training recovers GPQA but erodes the target behavior; lower AM is better. The after-training point comes from a separate historical run, not a continuation from the mixed-replay checkpoint, so compare the direction rather than the exact size of the difference. D Controlled wash-out co… view at source ↗
Figure 12
Figure 12. Figure 12: Controlled tests of what makes an installed behavior survive later benign training. Phase A installs the behavior; Phase B applies benign training. Lower AM means more survival. Top left compares SFT-only and midtrained checkpoints under the same continuation; top right compares continuation methods from a shared install. The bottom row compares Phase A filler prompts while holding Phase B fixed. Compare … view at source ↗
Figure 13
Figure 13. Figure 13: Teacher-written rewrites reduce GPQA and strengthen the target behavior in both initial-writer conditions. Columns group who wrote the initial response; colors show who rewrote the reasoning. Top: GPQA change from the shared base. Bottom: animal-welfare and self-preservation scores. Higher is better throughout. GPQA and welfare whiskers show one standard deviation across training seeds; self-preservation … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

33 extracted references · 27 linked inside Pith

  1. [3]

    We give the generation procedure and matched examples here because the differences between conditions are the intervention

    The released JSONL files contain every training row. We give the generation procedure and matched examples here because the differences between conditions are the intervention. For animal welfare and self-preservation, we follow the three-condition structure of TCW ( Kutasov et al. , 2026). We first generate a pool of user prompts. A GPT-4.1 teacher then ...

  2. [6]

    Weak-to-strong general- ization: Eliciting strong capabilities with weak supervision

    Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeff Wu. Weak-to-strong general- ization: Eliciting strong capabilities with weak supervision. arXiv preprint arXiv:2312.09390 , December

  3. [8]

    Subliminal learning: Language models transmit behavioral traits via hidden signals in data

    Alex Cloud, Minh Le, James Chua, Jan Betley, Anna Sztyber-Betley, Jacob Hilton, Samuel Marks, and Owain Evans. Subliminal learning: Language models transmit behavioral traits via hidden signals in data. arXiv preprint arXiv:2507.14805 ,

  4. [9]

    Safety cases: How to justify the safety of advanced AI systems

    Joshua Clymer, Nick Gabrieli, David Krueger, and Thomas Larsen. Safety cases: How to justify the safety of advanced AI systems. arXiv preprint arXiv:2403.10462 ,

  5. [10]

    doi: 10.1038/s41586-025-09422-z

    ISSN 1476-4687. doi: 10.1038/s41586-025-09422-z. URL http://dx.doi .org/10.1038/s41586-025-09422-z . 12 Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud, Shauna Kravec, Samuel Marks, Nicholas Schiefer, Ryan Soklaski, Alex Tamkin, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, Ethan Perez, and Evan Hubinger. Sycophancy to subterfuge: Investiga...

  6. [11]

    URL https://arxiv.org/abs/2507.06261. Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Jo- hannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, Sören Minder- mann, Ethan Perez, Linda Petrini, Jonathan Uesato, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, and Evan Hubinger. Alig...

  7. [13]

    Brown, Prafulla Dhariwal, Scott Gray, Chris Hallacy, Benjamin Mann, Alec Radford, Aditya Ramesh, Nick Ryder, Daniel M

    Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B. Brown, Prafulla Dhariwal, Scott Gray, Chris Hallacy, Benjamin Mann, Alec Radford, Aditya Ramesh, Nick Ryder, Daniel M. Ziegler, John Schulman, Dario Amodei, and Sam McCandlish. Scaling laws for autoregressive generative modeling. arXiv preprint arXiv:2010.14701 ,

  8. [14]

    Evan Hubinger, Nicholas Schiefer, Carson Denison, and Ethan Perez

    URL https://arxiv.org/abs/2604.14164. Evan Hubinger, Nicholas Schiefer, Carson Denison, and Ethan Perez. Model organisms of misalignment: The case for a new pillar of alignment research. https://www.alignmentforum.org/posts/ChDH335ck dvpxXaXX/model-organisms-of-misalignment-the-case-for-a-new-pillar-of-1 , August

  9. [15]

    Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M

    Alignment Forum post. Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma,...

  10. [16]

    Jonathan Kutasov, Adam Jermyn, Julius Steen, Minh Le, Samuel R

    OpenAI Alignment blog. Jonathan Kutasov, Adam Jermyn, Julius Steen, Minh Le, Samuel R. Bowman, Samuel Marks, Jan Leike, Amanda Askell, Chris Olah, Evan Hubinger, and Sara Price. Teaching claude why. https://alignment.an thropic.com/2026/teaching-claude-why/ , May

  11. [17]

    Andrew K

    Anthropic Alignment Science blog. Andrew K. Lampinen, Arslan Chaudhry, Stephanie C. Y. Chan, Cody Wild, Diane Wan, Alex Ku, Jörg Born- schein, Razvan Pascanu, Murray Shanahan, and James L. McClelland. On the generalization of language models from in-context learning and finetuning: A controlled study. arXiv preprint arXiv:2505.00661 ,

  12. [18]

    Model spec midtraining: Improving how alignment training generalizes

    Chloe Li, Nevan Wichers, Sara Price, Samuel Marks, and Jon Kutasov. Model spec midtraining: Improving how alignment training generalizes. arXiv preprint arXiv:2605.02087 , May

  13. [19]

    Gradient episodic memory for continual learning

    David Lopez-Paz and Marc’Aurelio Ranzato. Gradient episodic memory for continual learning. arXiv preprint arXiv:1706.08840,

  14. [20]

    Samuel Marks, Johannes Treutlein, Trenton Bricken, Jack Lindsey, Jonathan Marcus, Siddharth Mishra- Sharma, Daniel Ziegler, Emmanuel Ameisen, Joshua Batson, Tim Belonax, Samuel R

    URL https://arxiv.org/abs/2511.18397. Samuel Marks, Johannes Treutlein, Trenton Bricken, Jack Lindsey, Jonathan Marcus, Siddharth Mishra- Sharma, Daniel Ziegler, Emmanuel Ameisen, Joshua Batson, Tim Belonax, Samuel R. Bowman, Shan Carter, Brian Chen, Hoagy Cunningham, Carson Denison, Florian Dietz, Satvik Golechha, Akbir Khan, Jan Kirchner, Jan Leike, Aus...

  15. [21]

    Negation neglect: When models fail to learn negations in training

    Harry Mayne, Lev McKinney, Jan Dubiński, Adam Karvonen, James Chua, and Owain Evans. Negation neglect: When models fail to learn negations in training. arXiv preprint arXiv:2605.13829 ,

  16. [22]

    Tell, don’t show: Declarative facts influence how LLMs generalize

    Alexander Meinke and Owain Evans. Tell, don’t show: Declarative facts influence how LLMs generalize. arXiv preprint arXiv:2312.07779 ,

  17. [23]

    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback....

  18. [24]

    Advice for making robust-to-training model organisms

    Sebastian Prasanna, Alek Westover, Vivek Hebbar, Julian Stastny, and Dylan Xu. Advice for making robust-to-training model organisms. https://www.lesswrong.com/posts/CmkAxJi83jRv9eXgJ/advice-for- making-robust-to-training-model-organisms-1 , May 2026a. LessWrong post. 14 Sebastian Prasanna, Dylan Xu, Alek Westover, Julian Stastny, and Vivek Hebbar. Why doe...

  19. [27]

    School of reward hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

    Mia Taylor, James Chua, Jan Betley, Johannes Treutlein, and Owain Evans. School of reward hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs. arXiv preprint arXiv:2508.17511 ,

  20. [28]

    Connecting the dots: LLMs can infer and verbalize latent structure from disparate training data

    Johannes Treutlein, Dami Choi, Jan Betley, Samuel Marks, Cem Anil, Roger Grosse, and Owain Evans. Connecting the dots: LLMs can infer and verbalize latent structure from disparate training data. arXiv preprint arXiv:2406.14546 ,

  21. [29]

    Model organisms for emergent misalignment

    Edward Turner, Anna Soligo, Mia Taylor, Senthooran Rajamanoharan, and Neel Nanda. Model organisms for emergent misalignment. arXiv preprint arXiv:2506.11613 ,

  22. [30]

    Lukas Twist, Helen Yannakoudakis, and Jie M. Zhang. Reasoning-trace collapse: Evaluating the loss of explicit reasoning during fine-tuning. arXiv preprint arXiv:2605.21127 , May

  23. [31]

    LessWrong post. Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, A via Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. Lima: Less is more for alignment. arXiv preprint arXiv:2305.11206 ,

  24. [44]

    GPQA for the capability side, welfare judge scores for the animal- welfare side, and the same Petri Bloom suite described in Appendix H.2 for the self-preservation side

    Evaluation. GPQA for the capability side, welfare judge scores for the animal- welfare side, and the same Petri Bloom suite described in Appendix H.2 for the self-preservation side. Source records. registry/seed-errorbars/MANIFEST.md , journal/writeu p/plot_data/figure10_full_2x2.json Beyond-GPQA check Recipe. The welfare off-model and on-model rewrite ar...

  25. [2017]

    Lillicrap, and Greg Wayne

    David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy P. Lillicrap, and Greg Wayne. Experience replay for continual learning. arXiv preprint arXiv:1811.11682 ,

  26. [2019]

    From firewalls to frontiers: AI red-teaming is a domain-specific evolution of cyber red-teaming

    Anusha Sinha, Keltin Grimes, Teryn Lucassen, Michael Feffer, Nathan VanHoudnos, Zhiwei Steven Wu, and Hoda Heidari. From firewalls to frontiers: AI red-teaming is a domain-specific evolution of cyber red-teaming. arXiv preprint arXiv:2509.11398 ,

  27. [2020]

    On-policy replay for continual supervised fine-tuning

    Yan Chen, Taojie Zhu, Meng Zhang, Xin Chen, Jiaqi Huang, Dongyang Xu, and Yizhi Wang. On-policy replay for continual supervised fine-tuning. arXiv preprint arXiv:2605.29495 , May

  28. [2021]

    A is B” fail to learn “B is A

    Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans. Taken out of context: On measuring situational awareness in llms, 2023a. URL https://arxiv.org/abs/2309.00667. Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. The reve...

  29. [2022]

    BEiT: BERT pre-training of image transformers

    Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. BEiT: BERT pre-training of image transformers. arXiv preprint arXiv:2106.08254 ,

  30. [2023]

    Bowman, Zac 5In our source control, Qwen-written replay recovered GPQA, while Sonnet-written responses to the same prompts did not

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, K...

  31. [2024]

    Melody Y. Guan, Manas Joglekar, Eric Wallace, Saachi Jain, Boaz Barak, Alec Helyar, Rachel Dias, An- drea Vallone, Hongyu Ren, Jason Wei, Hyung Won Chung, Sam Toyer, Johannes Heidecke, Alex Beutel, and Amelia Glaese. Deliberative alignment: Reasoning enables safer language models. arXiv preprint arXiv:2412.16339,

  32. [2025]

    Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, and Owain Evans

    Felix J. Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, and Owain Evans. Looking inward: Language models can learn about themselves by introspection. arXiv preprint arXiv:2410.13787 ,

  33. [2026]

    Armen Aghajanyan, Lili Yu, Alexis Conneau, Wei-Ning Hsu, Karen Hambardzumyan, Susan Zhang, Stephen Roller, Naman Goyal, Omer Levy, and Luke Zettlemoyer

    LessWrong post. Armen Aghajanyan, Lili Yu, Alexis Conneau, Wei-Ning Hsu, Karen Hambardzumyan, Susan Zhang, Stephen Roller, Naman Goyal, Omer Levy, and Luke Zettlemoyer. Scaling laws for generative mixed-modal language models. arXiv preprint arXiv:2301.03728 ,