{"id":"08981b5a-cfc6-4aab-9681-caff6030ce79","arxiv_id":"2607.10226","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"SAE safety ablations are regime-dependent and baseline-dependent: medium-k heads can look efficient, but surface-matched dense steering often beats them and high-k collapses coherence.","lead":"Sparse autoencoder features for safety control look localized only in a narrow regime and only against certain dense baselines; fair same-layer matching can reverse the advantage. The paper supplies a matched, coherence-gated protocol so safety interventions are not scored on weak or degenerate outputs.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond residual baseline-matching caveats the paper already flags.","rationale":"The paper's strongest claim is carefully scoped as an evaluation result about when SAE interventions look localized under matched controls, not as a claim of verified causal circuits. The Gemma-2-9B reversal under same-layer and projected dense baselines (Table 5), confirmed by a second judge, is the load-bearing empirical result and is reported with paired bootstrap CIs. The reader's identified weakest assumption correctly notes that residual change is a proxy; the manuscript itself says so and does not overclaim mechanistic isolation. That caveat does not falsify the operational finding that total-norm matching alone is insufficient and can reverse the efficiency ranking once surface and basis are matched. Remaining limitations (incomplete surface matching at other scales, 4-bit 27B, single-annotator audit, promised code release) justify CONDITIONAL rather than unconditional ACCEPT but do not introduce a new load-bearing inconsistency. I therefore leave the reader's CONDITIONAL / HIGH verdict unchanged and treat the weakest-assumption note as already priced in.","tokens_in":18207,"tokens_out":606,"duration_ms":6576,"concrete_test":"Re-run the Table 5 surface/basis-matched bins on Llama-3.1-8B (and, if feasible, full-precision Gemma-2-27B) with same-layer and top-k decoder-span projected dense baselines; if the SAE advantage over all-layer dense largely survives surface matching on those models while remaining reversed on 9B, the baseline-dependence thesis is strengthened rather than weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest_assumption (that matched per-token residual change plus same-layer/decoder-span controls is only a proxy for locality, not causal isolation) is real but already stated by the authors in Sections 3 and 6–7 and does not undercut the operational claim. The central result is an evaluation finding: under the paper's own matched-surface and matched-basis protocol, the SAE efficiency advantage on Gemma-2-9B reverses (true-jailbreak deficits up to −0.29 under two judges; Table 5). That sign flip is internally consistent with the protocol they define; residual unmeasured differences (layer timing, nonlinear interactions, decoder non-orthogonality) would at most refine the magnitude of residual mismatch, not reverse the paper's thesis that naive total-norm matching can manufacture an apparent localization advantage. Multi-judge agreement, human audit of the coherence gate, random-SAE negative controls, and the direction-dependent injection result supply independent support. Gaps that remain (all-layer-only Llama/27B comparisons, 4-bit 27B, single-annotator audit, artifact not yet public) are scope limitations already reflected in a CONDITIONAL verdict, not a load-bearing flaw in the strongest claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper asks when SAE feature ablations act as localized control handles for safety-relevant behavior, and argues that apparent localization is highly sensitive to how dense baselines are matched. It introduces Matched Coherence-Gated (MCG) evaluation: complementary matched target-effect and matched perturbation-norm controls, a primary true-jailbreak metric requiring both judge-unsafe and coherent outputs, surface/basis-matched dense baselines (same-layer and decoder-span projected), and a second behavior-completion judge. On Gemma-2-9B with a Gemma Scope layer-20 SAE, the apparent SAE efficiency advantage over all-layer dense steering reverses under surface- and basis-matched dense baselines (true-jailbreak deficits up to about −0.29 under two judges), high-k ablations mainly induce coherence collapse, and a stable medium-k refusal-aligned feature head explains the useful regime. Against all-layer dense baselines the SAE advantage remains large on Llama-3.1-8B and Gemma-2-27B, while the same recipe fails on Gemma-2-2B via capability collapse and single-judge inflation. Refusal injection reverses the efficiency ranking in favor of dense steering.","tokens_in":18584,"tokens_out":1546,"duration_ms":31132,"significance":"If the results hold, the paper supplies a concrete evaluation standard for a claim that is currently often asserted rather than measured: that sparse SAE features are more localized safety handles than dense activation steering. The surface- and basis-matched reversal on Gemma-2-9B (Table 5), multi-seed paired bootstrap intervals, dual-judge cross-check, random-SAE and retain-set negative controls, feature-rank diagnostics, and the direction-dependent injection pairing are genuine strengths. The main contribution is methodological and operational rather than circuit-level: it shows that naive total-norm matching can manufacture an apparent localization advantage, and that SAE safety control should be treated as regime-, scale-, and direction-dependent. That is a useful corrective for both interpretability and safety-intervention evaluation.","major_comments":[{"comment":"§5.2 Table 5 vs §5.3 Tables 7–8 and §8: The load-bearing result that surface/basis matching reverses the SAE advantage is shown only for Gemma-2-9B. Llama-3.1-8B and Gemma-2-27B retain a large SAE advantage only against all-layer dense steering—the baseline the paper itself argues is insufficient (§3, contribution 2). Because the central thesis is that baseline specification (surface/basis), not sparsity, decides whether SAE looks localized, the architecture/scale claims should either include same-layer and decoder-span projected dense baselines or be more tightly scoped throughout the abstract, §5.3, and conclusion as “advantage over all-layer dense,” without implying that the localization conclusion transfers. As written, the multi-model narrative partially reintroduces the unmatched-surface comparison the protocol was designed to eliminate.","section":"§5.2–5.3, Tables 5, 7–8"},{"comment":"§4.3 and §5.2: Sharing a single refusal contrast to both rank SAE features and construct the dense direction is disclosed and intentional, but it means the comparison isolates intervention basis given a fixed target rather than testing whether SAE independently discovers a safety concept. That is fine for the stated claim, yet several passages (e.g., “refusal-aligned head,” “localized control handles”) can be read as stronger mechanistic localization. A short, explicit restatement near Table 5 and in the conclusion that no independent concept-discovery claim is made would prevent over-reading of the operational efficiency result.","section":"§4.3, §5.2, §8"},{"comment":"§5.3 Table 8 and §7: The 27B Pareto-dominance claim is restricted to safety/coherence/perturbation under 4-bit loading with capability at floor (GSM8K 0.12, MMLU 0.59). Quantization can interact differently with sparse feature ablation than with dense residual steering; without a full-precision check or a quantization-sensitivity note beyond the current limitation paragraph, the claim that the clean regime “grows” from 9B to 27B remains only partially supported. Either add a limited full-precision or higher-precision sanity run on a subset, or further qualify the 27B result as a quantized trend only.","section":"§5.3 Table 8, §7"}],"minor_comments":[{"comment":"§5.4 human audit: The audit is correctly labeled targeted and single-annotator (n=101). For a primary metric definition, even a small second-annotator agreement subset (e.g., 30–40 items with reported κ) would make the gate more credible without requiring a full multi-annotator study.","section":"§5.4"},{"comment":"Code and Data Availability: The promised public artifact (protocol, same-layer/projected baselines, HarmBench rescorer, per-example metrics, pinned sae_lens 6.44.2) is central to reproducibility, especially given the manual Llama Scope JumpReLU/normalization steps in §4.1. Please ensure release coincides with revision or provide a stable anonymous repository link for review.","section":"Code and Data Availability"},{"comment":"Table 2 footnote on baseline MMLU 0.680 vs later 0.692 is helpful; consider moving that caveat into the table caption so readers do not compare cells across evaluation passes.","section":"Table 2"},{"comment":"Figure 1 and Figure 3: Low-coherence points and scale panels are informative; adding explicit matched-bin markers (or a small legend for coherence threshold) would make the Pareto and cross-scale plots easier to read in grayscale.","section":"Figures 1, 3"},{"comment":"§4.1 Llama Scope loading details (threshold 0.3555, norm √d, BOS exclusion) are important for reproduction; consider a short appendix box or checklist so they are not buried in prose.","section":"§4.1"},{"comment":"Minor wording: “true jailbreak” is defined clearly in Eq. (1), but occasional use of “jailbreak” alone in figure captions can blur the gated vs unsafe-only distinction; prefer the gated term consistently in captions.","section":"Eq. (1), figure captions"}],"recommendation":"minor_revision","confidential_remarks":"This is a careful empirical methods paper; the Gemma-2-9B surface/basis reversal is the real contribution and is reported with appropriate multi-judge and bootstrap support. The main risk for the journal is over-generalization from all-layer Llama/27B results, which the authors already partially flag—pushing them to tighten that framing should be enough. Fit for a serious AI/ML venue is good if the revision keeps the claims operational rather than mechanistic."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple. Naive total-norm matching can make SAE feature ablation look more localized than it is. Once they match surface (same layer) and basis (decoder-span projection) on Gemma-2-9B, the apparent efficiency advantage over dense refusal steering disappears and reverses—true-jailbreak deficits up to about −0.29 under both Llama-Guard and HarmBench. High-k mainly collapses coherence; the 2B recipe is mostly single-judge inflation plus capability collapse. That is a useful, falsifiable correction to how people report SAE safety interventions.\n\nWhat is actually new is the dual matched-control protocol (matched target-effect and matched perturbation-norm), the surface/basis-matched dense baselines, and the coherence-gated “true jailbreak” metric, applied across Gemma 2B/9B/27B and Llama-3.1-8B with two SAE suites. They do the work carefully: six-seed paired bootstrap CIs, random-SAE and retain-set controls, layer scan, rank-decay diagnostics, human audit of the gate, second judge, and a refusal-injection pairing that shows the efficiency comparison is direction-dependent. Sharing one refusal contrast for ranking and dense construction is deliberate and disclosed so the comparison isolates basis, not target. Citations are appropriate (Gemma/Llama Scope, Arditi, HarmBench, Llama Guard).\n\nSoft spots are real but already flagged and do not undercut the operational claim. Per-token residual change plus same-layer/decoder-span matching is a proxy for locality, not causal circuit isolation—the authors say so. Llama and 27B comparisons stay against all-layer dense; 27B is 4-bit with capability at floor; the audit is single-annotator; the full artifact is promised rather than public. Those are scope limits, not a hole in the 9B sign-flip.\n\nThis is for people who run or evaluate activation interventions and for anyone tempted to treat SAE top-k as uniformly surgical. It deserves a serious referee. I would engage with it and cite the protocol and the surface-matched result.","headline":"Solid evaluation paper: SAE safety localization is baseline- and regime-dependent, and the Gemma-9B surface/basis reversal under two judges is the result that sticks.","tokens_in":19186,"tokens_out":520,"would_cite":true,"duration_ms":4622,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"SAE feature ablations are not uniformly localized safety handles: their apparent efficiency over dense steering is largely a baseline-matching artifact and reverses under fair surface-matched comparison.","keywords":["sparse autoencoders","activation steering","safety interventions","refusal directions","matched evaluation","coherence gating","jailbreak metrics","feature ablation"],"falsifier":"Re-run the surface- and basis-matched perturbation-norm comparison on Llama-3.1-8B and Gemma-2-27B (not only all-layer dense): if SAE still wins after same-layer and decoder-span matching under two judges and retained capability, the baseline-dependence thesis weakens; if the Gemma-2-9B-style reversal appears there too, the claim strengthens.","tokens_in":19069,"feed_emoji":"🔬","tokens_out":1107,"duration_ms":11857,"temperature":0.7,"pith_summary":"This paper asks when sparse autoencoder features actually act as localized control handles for safety-relevant behavior such as refusal and jailbreak susceptibility. Apparent success can come from weak interventions, mismatched dense baselines, capability damage, or incoherent text that safety judges still mark unsafe. The authors introduce a matched coherence-gated evaluation protocol that compares sparse and dense interventions at equal target effect and equal perturbation, counts a jailbreak only when an output is both judge-unsafe and coherent, matches the intervention surface and subspace, and cross-checks with a second judge. On Gemma-2-9B, once dense steering is restricted to the same layer or to the SAE decoder span, the SAE efficiency advantage disappears and reverses: fair dense baselines produce more coherent harmful compliance at matched perturbation. High-strength SAE ablations mainly collapse coherence, and the same recipe fails on a small 2B model through capability collapse and single-judge inflation. The practical message is that sparse localization for safety must be treated as regime-dependent and evaluation-dependent, not assumed from sparsity alone.","feed_headline":"SAE safety \"handles\" reverse once dense baselines are fairly matched","feed_subtitle":"Same-layer and decoder-span controls erase the sparse efficiency win; high-k mainly collapses coherence.","key_machinery":"Matched Coherence-Gated (MCG) Evaluation: two complementary controls (matched target-effect and matched perturbation-norm), plus surface and basis matching (same-layer dense and dense projected onto the top-k SAE decoder span), a primary true-jailbreak metric that requires both judge-unsafe and coherent output, and a second behavior-completion judge. It turns localization into a measured behavior-per-perturbation comparison rather than a raw success claim.","core_discovery":"SAE feature ablation is not a uniformly localized safety control mechanism. Matching only total perturbation norm leaves the intervention surface unmatched and can make single-layer SAE look more efficient than all-layer dense steering. On Gemma-2-9B, same-layer and decoder-span-projected dense baselines reverse that advantage at every matched bin (true-jailbreak deficits up to about -0.29 under two judges), while high-k ablations mainly induce coherence collapse and small-model SAE jailbreaks are largely single-judge inflation with capability collapse. The useful medium-k regime is driven by a stable head of refusal-aligned features whose activation separation decays rapidly with rank, and","pith_inferences":["If surface and basis matching routinely flip SAE-versus-dense rankings, many published sparse-localization wins may need re-evaluation under the same protocol.","A practical safety stack may need hybrid controls: sparse heads for targeted removal, dense directions for broad refusal injection.","The rapid rank-decay of refusal alignment suggests diminishing returns and rising side-effect risk beyond a medium-k head, which could guide automatic k selection.","Judge disagreement at small scale is itself a diagnostic of off-distribution degeneration, not only a measurement nuisance."],"forward_implications":["Sparse-versus-dense safety claims that only match total perturbation can misread an all-layer-versus-single-layer mismatch as localization.","High-k SAE ablations should not be scored with unsafe-only judges; coherence gates and a second behavior-completion judge are required.","SAE steering results from small models should not be extrapolated to larger ones without capability and multi-judge checks.","The right intervention tool depends on direction: sparse ablation can be more surgical for removing refusal-related behavior, while dense steering better preserves capability when injecting refusal.","Feature diagnostics of rank-decay and head stability should accompany top-k selection rather than treating k as a free hyperparameter."],"fun_headline_variants":["SAE ablation has only a narrow useful safety regime","Matched dense baselines reverse SAE sparse efficiency wins","High-k SAE features mainly collapse coherence not block harm","Useful SAE control driven by fast-decaying refusal feature head","SAE safety interventions fail to stay localized under fair match"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That equal per-token residual change, plus matching the layer and the SAE decoder subspace, is a fair enough proxy for localized control; if those still leave important differences unmeasured, the Gemma-2-9B reversal could partly reflect residual mismatch rather than true non-locality of the features.","fun_headline_variants_meta":{"raw":{"variants":["SAE ablation has only a narrow useful safety regime","Matched dense baselines reverse SAE sparse efficiency wins","High-k SAE features mainly collapse coherence not block harm","Useful SAE control driven by fast-decaying refusal feature head","SAE safety interventions fail to stay localized under fair match"]},"model":"grok-4.5","effort":"low","cost_usd":0.006284,"raw_usage":{"total_tokens":1640,"prompt_tokens":841,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":62840000,"prompt_tokens_details":{"text_tokens":841,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":738,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":841,"tokens_out":61,"duration_ms":5961,"temperature":1.0,"reasoning_tokens":738,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T13:22:09.411410+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the surface- and basis-matched perturbation-norm comparison on Llama-3.1-8B and Gemma-2-27B (not only all-layer dense): if SAE still wins after same-layer and decoder-span matching under two judges and retained capability, the baseline-dependence thesis weakens; if the Gemma-2-9B-style reversal appears there too, the claim strengthens.","supporting_citations":[],"review_version":1}