REVIEW 3 major objections 6 minor 14 references
SAE feature ablations are not uniformly localized safety handles: their apparent efficiency over dense steering is largely a baseline-matching artifact and reverses under fair surface-matched comparison.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 13:22 UTC pith:3XODFGME
load-bearing objection Solid evaluation paper: SAE safety localization is baseline- and regime-dependent, and the Gemma-9B surface/basis reversal under two judges is the result that sticks. the 3 major comments →
When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
SAE feature ablation is not a uniformly localized safety control mechanism. Matching only total perturbation norm leaves the intervention surface unmatched and can make single-layer SAE look more efficient than all-layer dense steering. On Gemma-2-9B, same-layer and decoder-span-projected dense baselines reverse that advantage at every matched bin (true-jailbreak deficits up to about -0.29 under two judges), while high-k ablations mainly induce coherence collapse and small-model SAE jailbreaks are largely single-judge inflation with capability collapse. The useful medium-k regime is driven by a stable head of refusal-aligned features whose activation separation decays rapidly with rank, and
What carries the argument
Matched Coherence-Gated (MCG) Evaluation: two complementary controls (matched target-effect and matched perturbation-norm), plus surface and basis matching (same-layer dense and dense projected onto the top-k SAE decoder span), a primary true-jailbreak metric that requires both judge-unsafe and coherent output, and a second behavior-completion judge. It turns localization into a measured behavior-per-perturbation comparison rather than a raw success claim.
Load-bearing premise
That equal per-token residual change, plus matching the layer and the SAE decoder subspace, is a fair enough proxy for localized control; if those still leave important differences unmeasured, the Gemma-2-9B reversal could partly reflect residual mismatch rather than true non-locality of the features.
What would settle it
Re-run the surface- and basis-matched perturbation-norm comparison on Llama-3.1-8B and Gemma-2-27B (not only all-layer dense): if SAE still wins after same-layer and decoder-span matching under two judges and retained capability, the baseline-dependence thesis weakens; if the Gemma-2-9B-style reversal appears there too, the claim strengthens.
If this is right
- Sparse-versus-dense safety claims that only match total perturbation can misread an all-layer-versus-single-layer mismatch as localization.
- High-k SAE ablations should not be scored with unsafe-only judges; coherence gates and a second behavior-completion judge are required.
- SAE steering results from small models should not be extrapolated to larger ones without capability and multi-judge checks.
- The right intervention tool depends on direction: sparse ablation can be more surgical for removing refusal-related behavior, while dense steering better preserves capability when injecting refusal.
- Feature diagnostics of rank-decay and head stability should accompany top-k selection rather than treating k as a free hyperparameter.
Where Pith is reading between the lines
- If surface and basis matching routinely flip SAE-versus-dense rankings, many published sparse-localization wins may need re-evaluation under the same protocol.
- A practical safety stack may need hybrid controls: sparse heads for targeted removal, dense directions for broad refusal injection.
- The rapid rank-decay of refusal alignment suggests diminishing returns and rising side-effect risk beyond a medium-k head, which could guide automatic k selection.
- Judge disagreement at small scale is itself a diagnostic of off-distribution degeneration, not only a measurement nuisance.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks when SAE feature ablations act as localized control handles for safety-relevant behavior, and argues that apparent localization is highly sensitive to how dense baselines are matched. It introduces Matched Coherence-Gated (MCG) evaluation: complementary matched target-effect and matched perturbation-norm controls, a primary true-jailbreak metric requiring both judge-unsafe and coherent outputs, surface/basis-matched dense baselines (same-layer and decoder-span projected), and a second behavior-completion judge. On Gemma-2-9B with a Gemma Scope layer-20 SAE, the apparent SAE efficiency advantage over all-layer dense steering reverses under surface- and basis-matched dense baselines (true-jailbreak deficits up to about −0.29 under two judges), high-k ablations mainly induce coherence collapse, and a stable medium-k refusal-aligned feature head explains the useful regime. Against all-layer dense baselines the SAE advantage remains large on Llama-3.1-8B and Gemma-2-27B, while the same recipe fails on Gemma-2-2B via capability collapse and single-judge inflation. Refusal injection reverses the efficiency ranking in favor of dense steering.
Significance. If the results hold, the paper supplies a concrete evaluation standard for a claim that is currently often asserted rather than measured: that sparse SAE features are more localized safety handles than dense activation steering. The surface- and basis-matched reversal on Gemma-2-9B (Table 5), multi-seed paired bootstrap intervals, dual-judge cross-check, random-SAE and retain-set negative controls, feature-rank diagnostics, and the direction-dependent injection pairing are genuine strengths. The main contribution is methodological and operational rather than circuit-level: it shows that naive total-norm matching can manufacture an apparent localization advantage, and that SAE safety control should be treated as regime-, scale-, and direction-dependent. That is a useful corrective for both interpretability and safety-intervention evaluation.
major comments (3)
- [§5.2–5.3, Tables 5, 7–8] §5.2 Table 5 vs §5.3 Tables 7–8 and §8: The load-bearing result that surface/basis matching reverses the SAE advantage is shown only for Gemma-2-9B. Llama-3.1-8B and Gemma-2-27B retain a large SAE advantage only against all-layer dense steering—the baseline the paper itself argues is insufficient (§3, contribution 2). Because the central thesis is that baseline specification (surface/basis), not sparsity, decides whether SAE looks localized, the architecture/scale claims should either include same-layer and decoder-span projected dense baselines or be more tightly scoped throughout the abstract, §5.3, and conclusion as “advantage over all-layer dense,” without implying that the localization conclusion transfers. As written, the multi-model narrative partially reintroduces the unmatched-surface comparison the protocol was designed to eliminate.
- [§4.3, §5.2, §8] §4.3 and §5.2: Sharing a single refusal contrast to both rank SAE features and construct the dense direction is disclosed and intentional, but it means the comparison isolates intervention basis given a fixed target rather than testing whether SAE independently discovers a safety concept. That is fine for the stated claim, yet several passages (e.g., “refusal-aligned head,” “localized control handles”) can be read as stronger mechanistic localization. A short, explicit restatement near Table 5 and in the conclusion that no independent concept-discovery claim is made would prevent over-reading of the operational efficiency result.
- [§5.3 Table 8, §7] §5.3 Table 8 and §7: The 27B Pareto-dominance claim is restricted to safety/coherence/perturbation under 4-bit loading with capability at floor (GSM8K 0.12, MMLU 0.59). Quantization can interact differently with sparse feature ablation than with dense residual steering; without a full-precision check or a quantization-sensitivity note beyond the current limitation paragraph, the claim that the clean regime “grows” from 9B to 27B remains only partially supported. Either add a limited full-precision or higher-precision sanity run on a subset, or further qualify the 27B result as a quantized trend only.
minor comments (6)
- [§5.4] §5.4 human audit: The audit is correctly labeled targeted and single-annotator (n=101). For a primary metric definition, even a small second-annotator agreement subset (e.g., 30–40 items with reported κ) would make the gate more credible without requiring a full multi-annotator study.
- [Code and Data Availability] Code and Data Availability: The promised public artifact (protocol, same-layer/projected baselines, HarmBench rescorer, per-example metrics, pinned sae_lens 6.44.2) is central to reproducibility, especially given the manual Llama Scope JumpReLU/normalization steps in §4.1. Please ensure release coincides with revision or provide a stable anonymous repository link for review.
- [Table 2] Table 2 footnote on baseline MMLU 0.680 vs later 0.692 is helpful; consider moving that caveat into the table caption so readers do not compare cells across evaluation passes.
- [Figures 1, 3] Figure 1 and Figure 3: Low-coherence points and scale panels are informative; adding explicit matched-bin markers (or a small legend for coherence threshold) would make the Pareto and cross-scale plots easier to read in grayscale.
- [§4.1] §4.1 Llama Scope loading details (threshold 0.3555, norm √d, BOS exclusion) are important for reproduction; consider a short appendix box or checklist so they are not buried in prose.
- [Eq. (1), figure captions] Minor wording: “true jailbreak” is defined clearly in Eq. (1), but occasional use of “jailbreak” alone in figure captions can blur the gated vs unsafe-only distinction; prefer the gated term consistently in captions.
Circularity Check
No significant circularity: empirical matched-evaluation study whose claims are measured outcomes, not results forced by definition or self-citation.
full rationale
This paper is an empirical evaluation of SAE feature ablation versus dense refusal-direction steering under a matched coherence-gated protocol. Its load-bearing claims (sign-flip of SAE efficiency on Gemma-2-9B once surface/basis are matched; high-k coherence collapse; small-model single-judge inflation; direction-dependent injection result) are experimental measurements on held-out prompts with multi-judge and capability metrics, not derivations that reduce to their inputs by construction. True-jailbreak is defined independently as unsafe-and-coherent and then measured; it is not fitted from the intervention parameters. Sharing one refusal contrast to rank SAE features and build the dense direction is an explicit design choice to hold the target fixed and isolate the intervention basis—not a circular proof that sparse features control refusal. Ranking features by cosine similarity to the refusal direction and then reporting activation separation (Cohen’s d) on refused vs. complied prompts is a related diagnostic, not a self-definitional prediction of the main efficiency results; stability across splits, rank-decay geometry, single-feature ablation weakness, and random-SAE negative controls are additional empirical checks. Citations (Arditi, Bricken, Gemma Scope, Llama Scope, HarmBench, etc.) are external prior work, not load-bearing self-citation uniqueness theorems. The paper itself frames localization as an operational behavior-per-perturbation comparison rather than causal circuit isolation. No step reduces Eq. X to Eq. Y by construction, renames a fitted parameter as a prediction, or smuggles an ansatz via author-overlapping uniqueness claims. Score 0 is therefore the correct honest finding.
Axiom & Free-Parameter Ledger
free parameters (5)
- SAE top-k grid (50–3200, focus on 400/800/1600/3200)
- Dense scaling coefficients beta / gamma grids
- SAE layer choices (Gemma-9B L20, Gemma-2B L12, Llama L15)
- Coherence-gate heuristic thresholds (length, lexical diversity, alphabetic ratio, repeated n-grams)
- Llama Scope JumpReLU threshold 0.3555 and activation normalization to sqrt(d)
axioms (5)
- domain assumption Per-token relative residual change is a valid proxy for intervention strength / locality when comparing methods.
- domain assumption Unsafe-and-coherent (true jailbreak) better measures meaningful harmful compliance than unsafe-only judge labels.
- ad hoc to paper Sharing one refusal contrast to rank SAE features and build the dense direction fairly isolates intervention basis.
- domain assumption Llama-Guard and HarmBench labels, gated by coherence, are adequate primary/secondary safety metrics for held-out prompts.
- standard math Standard residual-stream SAE encode/decode and dense activation addition/ablation mathematics.
invented entities (2)
-
Matched Coherence-Gated (MCG) Evaluation protocol
no independent evidence
-
true jailbreak metric (unsafe AND coherent)
independent evidence
read the original abstract
We evaluate when sparse autoencoder (SAE) features act as localized control handles for safety-relevant behavior. This question is difficult because apparent success can arise from weak interventions, mismatched baselines, model robustness, or degenerate outputs that automated safety judges mark as unsafe without representing meaningful harmful compliance. We introduce a matched coherence-gated evaluation protocol for runtime safety interventions: methods are compared at matched target-effect points, and the primary target metric counts harmful compliance only when an output is both judge-unsafe and coherent. Applying this protocol to three prompt splits on Gemma-2-9B-it with a Gemma Scope layer-20 residual SAE, we find that SAE feature ablation has a narrow useful regime. SAE top800 reaches a low-to-mid target effect with lower total perturbation and competitive utility, but SAE top1600 loses utility relative to a matched dense refusal-direction baseline, and SAE top3200 primarily induces coherence collapse. Human audit confirms that coherence gating removes unsafe-only artifacts, and feature diagnostics show that the useful regime is driven by a stable head of refusal-aligned features whose activation separation decays rapidly with rank. These results argue that SAE-based safety interventions should be evaluated as regime-dependent control mechanisms rather than assumed to be uniformly localized.
Figures
Reference graph
Works this paper leans on
-
[1]
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L
arXiv:2406.11717. Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L. Turner, et al. Towards monosemanticity: Decomposing language models with dictio- nary learning,
-
[2]
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey
arXiv:2110.14168. Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. Sparse autoen- coders find highly interpretable features in language models,
-
[3]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, et al
arXiv:2309.08600. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, et al. The llama 3 herd of models,
-
[4]
arXiv:2407.21783. Igor Fedorov, Kate Plawiak, Lemeng Wu, Tarek Elgamal, Naveen Suda, Eric Smith, Hongyuan Zhan, Jianfeng Chi, Yuriy Hulovatyy, Kimish Patel, Zechun Liu, Changsheng Zhao, Yangyang Shi, Tijmen Blankevoort, Mahesh Pasupuleti, Bilge Soran, Zacharie Delpierre Coudert, Rachad Alao, Raghuraman Krishnamoorthi, and Vikas Chandra. Llama guard 3-1b-i...
-
[5]
arXiv:2411.17713. Gemma Team. Gemma 2: Improving open language models at a practical size,
-
[6]
arXiv:2408.00118. Zhengfu He, Wentao Shu, Xuyang Ge, Lingjie Chen, Junxuan Wang, Yunhua Zhou, Frances Liu, Qipeng Guo, Xuanjing Huang, Zuxuan Wu, Yu-Gang Jiang, and Xipeng Qiu. Llama scope: Ex- tracting millions of features from llama-3.1-8b with sparse autoencoders,
-
[7]
arXiv:2410.20526. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding,
-
[8]
arXiv:2009.03300. Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, and Nouha Dziri. Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models,
Pith/arXiv arXiv 2009
-
[9]
arXiv:2406.18510. Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, Janos Kramar, Anca Dragan, Rohin Shah, and Neel Nanda. Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2,
-
[10]
arXiv:2408.05147. 21 Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. Harmbench: A standard- ized evaluation framework for automated red teaming and robust refusal,
-
[11]
Paul Rottger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy
arXiv:2402.04249. Paul Rottger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. Xstest: A test suite for identifying exaggerated safety behaviours in large language models,
-
[12]
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J
arXiv:2308.01263. Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J. Vazquez, Ulisse Mini, and Monte MacDiarmid. Steering language models with activation engineering,
-
[13]
arXiv:2308.10248. Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. Representation engineering: A top-...
-
[14]
arXiv:2310.01405. 22
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.