{"id":"b7beb419-5fcc-4677-8b2f-e2f76ef2de1a","arxiv_id":"2607.20596","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A single-token feature's causal necessity under zero-ablation depends on which SAE family found it: GemmaScope and BatchTopK features stay causally anchored while LlamaScope features are locally redundant.","lead":"The paper tests whether 'single-token' word-detector features found by different sparse autoencoder (SAE) methods play the same causal role inside language models. It finds the causal role depends on which SAE family produced the feature, so interpretability claims do not automatically transfer between SAE families.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Same-layer recovery metric is uncalibrated: LlamaScope features show the largest Δlogit yet 97.7% 'recovery'; the anchored-vs-redundant split may be a floor-5/rank-headroom artifact.","rationale":"The paper's necessity arm is genuinely strong: 178/208 full-layer conditions BH-significant (verified by summing Table 6), inert control populations (B.13 ≥99.96% control recovery), and an alignment-matched null (B.14) that partially addresses geometric selection. The honest Limitations section correctly flags the cross-family confounds. My concern is narrower but load-bearing: the specific operational claim that LlamaScope features are 'locally redundant' is read off an uncalibrated rank-recovery metric. The internal contradiction between large significant Δlogit and 97.7% recovery shows the metric is not a faithful measure of causal replacement. The proposed no-compensation test would settle whether recovery is just a threshold effect. If it is, the paper's most striking family split is unsupported; the weaker claim that SAE families differ on the same base (BatchTopK vs GemmaScope) still has some support in Table 8, so rejection is not warranted. The abstract's 'same base model' phrasing for the LlamaScope redundancy is also stronger than the evidence, since LlamaScope is never paired with GemmaScope on one base. The conditional verdict should stand, with recovery calibration as a required check.","tokens_in":26827,"tokens_out":17119,"duration_ms":141272,"concrete_test":"Analytical calibration of recovery: for Llama-3.1-8B and DeepSeek-R1, take clean same-layer logit-lens logits for each ST target token, subtract the measured layer-level mean Δlogit from only the target token's logit, recompute the rank, and apply the paper's recovery criterion (≤2× pre-ablation rank, floor 5). Under a no-compensation null this predicts recovery from rank headroom alone; if predicted recovery is close to the reported 97.7%/95.5%, then 'locally redundant' is a mechanical consequence of the floor, not evidence of replacement. If observed recovery is far above the null's prediction, the redundancy interpretation survives. Also report recovery with the floor removed and stratified by pre-ablation rank bin.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central family split rests on the same-layer recovery metric (§4.5): a feature is 'recovered' if its target-token rank after ablation stays within twice its pre-ablation rank, floor 5. The paper labels LlamaScope features 'locally redundant' at 96–98% recovery, but the metric is never calibrated against the magnitude of the ablation effect. Llama-3.1-8B × LlamaScope is BH-significant in 31/32 layers, with layer Δlogit as negative as −1.828 at L1 and roughly −0.5 to −0.9 across most layers (Table 20), yet recovery is 97.7%. Under the floor-5 rule, a token with pre-ablation rank 1 is 'recovered' if it stays anywhere in ranks 1–5; a −1.8 logit shift can leave it at rank 2–5 while substantially reducing its probability. Recovery therefore conflates 'token still in a loose top-5 window' with 'other features compensate.' No null tells us what recovery would be under a random magnitude-matched ablation, and pre-ablation rank distributions are not stratified. If the redundancy label is an artifact of this window, the headline — GemmaScope/BatchTopK anchored vs LlamaScope redundant — reduces to 'all families show significant necessity; only rank thresholds differ.' This is compounded by the abstract's 'same base model' phrasing: LlamaScope is only evaluated on Llama-3.1-8B/DeepSeek-R1, so 96–98% redundancy is never observed in the same-base GemmaScope-vs-BatchTopK comparisons (70.6%/68.0% on Gemma-2-2B).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies 'single-token' sparse autoencoder (SAE) features — features whose activation is dominated by a single vocabulary item — across six language models and three SAE families (GemmaScope/res-jb, LlamaScope, community BatchTopK). It reports that such features are geometrically distinct (4.7x tighter decoder clustering, 1.72x higher embedding alignment), concentrated in early layers (91% in GPT2 L0), and causally necessary under zero-ablation at full layer depth, with 178 of 208 full-layer conditions significant under a single global Benjamini-Hochberg correction. The paper's headline claim is a cross-family dissociation: on the same base model, GemmaScope and BatchTopK features are 'causally anchored,' while LlamaScope features are 'locally redundant,' recovering their pre-ablation target-token rank 96-98% of the time. It also reports a depth dissociation: necessity damage increases with depth (rho=0.97 for BatchTopK, 0.70 for GemmaScope on Gemma-2-2B) while downstream anchoring concentrates in early layers (rho=-0.65). The paper concludes that a feature's causal role depends on which SAE produced it, and that SAE family should be treated as an experimental variable.","tokens_in":2059,"tokens_out":2099,"duration_ms":68296,"significance":"If the central dissociation between 'anchored' and 'locally redundant' SAE families held, the paper would make an important contribution: it would show that single-token SAE features, the cleanest possible case for ground-truth comparison, have real but non-portable causal necessity, and that interpretability claims must control for SAE training methodology. The paper has genuine strengths: the causal arm is large and internally consistent (178/208 layer-level tests verified from Table 6), uses a single global BH correction, includes magnitude-matched controls, and reports control-population inertness (B.13, recovery >=99.96%). The inclusion of an alignment-matched null (B.14) is a serious attempt to address selection-geometry circularity. The empirical patterns — early-layer concentration, the L0-to-L1 representational shift, increasing necessity with depth — are plausible and well documented. However, the load-bearing cross-family claim currently rests on an uncalibrated recovery metric and on a same-base comparison that, as written, is not supported by the evaluated model-by-SAE matrix. These issues materially affect the paper's central conclusion, though they appear fixable wit","major_comments":[{"comment":"The anchored-vs-redundant split is read almost entirely off the same-layer recovery metric, defined as the fraction of features whose post-ablation target-token rank stays within twice its pre-ablation rank, with a floor of 5. This metric is never calibrated. Llama-3.1-8B x LlamaScope shows 97.7% recovery while Table 20 reports mean Delta logit = -1.828 at L1, -0.5 to -0.9 across most layers, and BH-significant necessity in 31/32 layers. Under the floor-5 rule, a rank-1 token 'recovers' if it stays anywhere in ranks 1-5; a large logit drop can leave it at rank 2-5 while substantially reducing its probability. The paper does not report pre-ablation rank distributions, nor a null recovery rate under magnitude-matched random ablations. B.13's control population is inert by construction (median Delta logit between -0.00006 and -0.0012), so its >=99.96% recovery does not establish what recove","section":"§4.5, Table 6, Table 20, B.13"},{"comment":"The abstract and §1 state that 'on the same base model, GemmaScope and BatchTopK features remain causally anchored, while LlamaScope features are locally redundant.' Table 1 shows LlamaScope is evaluated only on Llama-3.1-8B and DeepSeek-R1; no LlamaScope SAE is evaluated on any Gemma model, and no GemmaScope/BatchTopK SAE is evaluated on any Llama model. Thus the LlamaScope half of the family split is conflated with base model, tokenizer, and pretraining data. The Limitations section concedes that cross-family comparisons co-vary training data, but the Conclusion's claim that 'the causal role of a feature depends on which SAE produced it' goes beyond the within-model GemmaScope-vs-BatchTopK evidence. A same-base LlamaScope comparison (or a GemmaScope-style SAE on Llama) is needed before the abstract's 'same base model' phrasing can stand.","section":"Abstract, §1, Table 1, §4.5, Limitations"},{"comment":"Decoder-alignment detection selects features by cosine similarity between the decoder vector and the target token's input embedding, and the zero-ablation removes exactly that decoder direction. The paper's B.14 alignment-matched null is a reasonable response, but the matching gap is large: median |Delta cos| = 0.18 on GemmaScope and 0.51 on LlamaScope. The in-band subset is limited to 315 controls across five Gemma-2-2B layers, with one of five layers inconclusive (L6), and the LlamaScope layer-1 nearest controls themselves carry a -1.318 mean Delta logit. The circularity concern is therefore only partially mitigated. Please report the recovery metric on the in-band alignment-matched null as well as necessity, and stratify results by |Delta cos| or provide a closer-matched set. This directly bears on whether the causal arm measures feature role or selection geometry.","section":"§3.3, B.14, Table 27"}],"minor_comments":[{"comment":"The label 'LlamaScope (TopK→JR)' is never defined. §3.1 describes LlamaScope as TopK SAEs; §3.3 mentions a TopK-to-JumpReLU conversion. Explain what conversion was applied and when, since the causal results for LlamaScope depend on the final SAE type actually ablated.","section":"Table 6"},{"comment":"Table 5 is labeled 'Causal (N=26,594)' but the main causal analyses use decoder-alignment detection rather than activation-based detection. Clarify which detector produced the semantic-category causal statistics, and whether the category distribution is over the decoder-detected set.","section":"§4.4, Table 5"},{"comment":"The caption says '17 sampled model×layer conditions,' but Table 6 reports 208 full-layer conditions. Explain the sampling and whether Figure 4B is illustrative; the text implies the full-layer coverage is in Table 6.","section":"Figure 4B"},{"comment":"The text quotes recovery as '96-98%' for LlamaScope, but Table 6 lists 97.7% and 95.5%. State both values consistently and reconcile the 96-98% range with the 95.5% DeepSeek-R1 value.","section":"§4.5"},{"comment":"GPT2-Small is described as a 'single-layer condition' but the abstract and §4.5 say 'across six transformer language models and three SAE families.' Specify that GPT2 contributes one layer (L0) and that full-depth coverage is for the seven configurations.","section":"§1/Table 6"},{"comment":"The anchoring test is described as one-sided Mann-Whitney U at p<0.05, but the necessity test uses global BH correction over 208 layers. Clarify whether anchoring p-values are also BH-corrected, and if not, justify the different multiple-testing treatment.","section":"§4.5"}],"recommendation":"major_revision","confidential_remarks":"The paper is potentially valuable but currently oversells the cross-family conclusion. The necessity arm is solid and reproducible; the 'locally redundant' label for LlamaScope is not yet established because the recovery metric is uncalibrated and the same-base comparison in the abstract is not implemented. I would recommend requiring either a same-base LlamaScope experiment or a clear restriction of the portability claim to the within-Gemma comparison, plus a calibrated null for the recovery metric. The alignment-matching null in B.14 is a good start but needs to be tighter for the strong causal interpretation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version. The paper does something genuinely new: it runs zero-ablation on matched single-token features across three SAE families and six models, and finds real causal structure—178/208 layer-conditions BH-significant, controls inert, and a depth dissociation (necessity increases with depth, downstream anchoring concentrated early) that is consistent and interesting. The reporting is unusually honest: full per-layer tables, explicit limitations, an alignment-matched null control. The core claim that single-token feature necessity is real but not portable across families is supported.\n\nThe soft spots are three, and the first is the one that matters. The 'locally redundant' label for LlamaScope rests on the same-layer recovery metric: rank after ablation stays within 2x baseline, floor 5. That metric is never calibrated. LlamaScope features show the study's largest median Δlogit (-0.374 at Llama-3.1-8B, -1.828 at L1) yet 96-98% recovery. A -1.8 logit drop can leave a token at rank 2-5, which the floor-5 rule counts as 'recovered.' So 'redundant' conflates 'still in a loose top-5 window' with 'other features compensate.' There's no null showing what recovery would be for a random magnitude-matched ablation, and control recovery is measured on features with negligible effects. The anchoring split (GemmaScope/BatchTopK anchor 92-100% of layers vs LlamaScope 31-34%) is a separate, more defensible result, but the abstract's 'locally redundant' language is tied to the recovery metric and overstates the case.\n\nSecond, the abstract says 'on the same base model... LlamaScope features are locally redundant,' but LlamaScope is only evaluated on Llama-3.1-8B/DeepSeek-R1. The same-base comparison is GemmaScope vs BatchTopK on Gemma models. The body reports this correctly; the abstract doesn't.\n\nThird, the 46x prevalence gap is a detector artifact as presented. Under the decoder-alignment detector used for the causal claims, LlamaScope prevalence (0.27%) exceeds Gemma-2-9B (0.14%), reversing the headline number. The paper acknowledges this in §3.3 and hedges in Discussion, but the number is still in the abstract.\n\nAlso worth noting: feature selection and ablation both use the decoder direction, so there's a selection-geometry circularity concern. The B.14 alignment-matched null partially addresses it, but the matching tolerance is loose (median |Δcos| 0.18-0.51), so the necessity effect isn't fully separated from embedding-alignment geometry.\n\nBottom line: this is a serious, honest paper that deserves a serious referee. The causal arm is well-done and the depth dissociation looks real. But the 'LlamaScope features are locally redundant' headline needs a calibrated recovery metric (or should be re-framed around the anchoring split) before it can be taken at face value. I'd send it to peer review, and I'd tell the authors to fix the abstract and add a magnitude-matched recovery null.","headline":"Genuinely new cross-family causal comparison with an excellent causal arm, but the 'LlamaScope is locally redundant' headline rests on an uncalibrated recovery metric and an overstating abstract.","tokens_in":27810,"tokens_out":3862,"would_cite":false,"duration_ms":30141,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single-token feature's causal necessity is real but depends on which sparse autoencoder family produced it, and layer depth decides whether ablation damage propagates downstream or directly reshapes the output.","keywords":["sparse autoencoders","single-token features","zero-ablation","causal necessity","interpretability","language models","layer depth","SAE families"],"falsifier":"Compute the same-layer recovery metric on magnitude-matched random controls after sorting them into pre-ablation rank bins identical to the single-token features' bins; if controls also show 96–98% recovery in the rank-1 bin, the 'locally redundant' label is a rank-window artifact.","tokens_in":26512,"feed_emoji":"🧠","tokens_out":5475,"duration_ms":43261,"temperature":0.7,"pith_summary":"Single-token features—SAE directions that activate almost exclusively on one vocabulary item—are the cleanest possible test case for whether sparse-autoencoder interpretations are portable across training setups. The paper finds that zero-ablating these features reduces the target token's logit at Benjamini-Hochberg-significant levels in 178 of 208 full-layer tests across six models and three SAE families: single-token features are causally necessary in a real sense. But that necessity is not a property of the token; it is a property of the SAE. On the same base model, features from one family remain 'causally anchored' and their ablation cascades downstream, while features from another family are 'locally redundant'—the target token's rank recovers to within twice baseline 96–98% of the time. Depth dissociates the two roles: necessity damage grows with depth, while anchoring damage concentrates in early layers. If this is right, cross-family interpretability claims must treat the SAE family as an experimental variable, not a detail.","feed_headline":"Same token is necessary under one SAE, redundant under another","feed_subtitle":"Zero-ablation across 208 layer tests shows real causal effects that do not transfer across sparse-autoencoder families.","key_machinery":"The load-bearing instrument is zero-ablation: removing the feature's contribution from the residual stream by subtracting its activation times its decoder vector, then measuring the change in the target token's logit against magnitude-matched random controls. Two derived quantities carry the argument: 'necessity' (statistically significant logit reduction at the source layer) and 'anchoring' (the ablation's downstream effect on logit-lens readouts at later layers), with 'same-layer recovery'—the fraction of features whose target-token rank stays within twice baseline, floor 5—distinguishing anchored from locally redundant regimes. Detection uses decoder-alignment, the cosine between a featur","core_discovery":"The central claim is that at the single-token endpoint—where ground truth is unambiguous—the causal role of a feature depends on which SAE produced it. Using zero-ablation at full layer depth on 3.9M features across six models and three SAE families, the paper shows single-token features are geometrically distinct (4.7× tighter decoder clustering, 1.72× higher embedding alignment, concentrated in early layers) and causally necessary under ablation in 178 of 208 layer conditions. Yet the same ablation protocol splits the three families: on Gemma models two families anchor downstream layers 92–100% of the time, while on Llama/DeepSeek models the other family anchors only 31–34% and shows 96–98","pith_inferences":["A natural extension is to train two SAEs with identical recipe, data, and width, differing only in activation function, on the same base model across all layers; if the anchored-versus-redundant split persists, recipe controls are needed, and if it collapses, the split is an artifact of uncontrolled recipe differences.","The same-layer recovery metric could be recalibrated as a function of pre-ablation rank and logit-drop magnitude; features starting at rank 1 have a floor-5 window that makes 'recovery' artificially easy.","The category-dependent convergence—domain-specific tokens converge across families while function words diverge—suggests that future cross-SAE comparisons should be stratified by token type, not reported only in aggregate.","If the family split is training-recipe-driven, then SAE evaluation benchmarks that report only reconstruction fidelity or interpretability are missing a causal dimension; adding per-feature necessity scores would make cross-family comparability measurable."],"forward_implications":["Interpretability claims built on one SAE family should not be assumed to transfer to another, even on the same base model; steering and editing pipelines should re-run ablation checks under the deployed family.","Layer depth should be reported as part of causal claims: late-layer features shape the output distribution directly, while early-layer features propagate damage downstream.","Activation function is not the decisive factor in cross-family causal differences; training recipe factors such as decoder norms, training scale, and post-hoc conversion are the residual candidates.","Single-token features provide a tractable benchmark endpoint where vocabulary-level ground truth allows exact cross-family matching, so they can serve as a diagnostic for SAE evaluation."],"fun_headline_variants":["Feature necessity flips across SAE families","Single-token features: causal in one SAE, redundant in another","SAE family, not scale, decides feature causality","Same token, opposite causal role under different SAEs","Zero-ablation shows SAE-family-dependent causality"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The split between 'anchored' and 'locally redundant' rests on the same-layer recovery metric—rank within twice baseline after ablation, floor 5—without calibrating it against pre-ablation rank distributions or the size of the logit drop; if that window or baseline differences drive recovery, the family split weakens.","fun_headline_variants_meta":{"raw":{"variants":["Feature necessity flips across SAE families","Single-token features: causal in one SAE, redundant in another","SAE family, not scale, decides feature causality","Same token, opposite causal role under different SAEs","Zero-ablation shows SAE-family-dependent causality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000166,"raw_usage":{"total_tokens":1114,"prompt_tokens":788,"completion_tokens":326,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":249}},"tokens_in":532,"tokens_out":326,"duration_ms":4112,"temperature":1.0,"reasoning_tokens":249,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T09:57:43.026222+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the same-layer recovery metric on magnitude-matched random controls after sorting them into pre-ablation rank bins identical to the single-token features' bins; if controls also show 96–98% recovery in the rank-1 bin, the 'locally redundant' label is a rank-window artifact.","supporting_citations":[],"review_version":1}