{"id":"bcee081a-3aea-477a-9274-5768a9ec27e3","arxiv_id":"2607.24645","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"SAE features can be interpretable and causally useful yet still lack a stable one-dimensional logit direction for steering, with value-like features more structured than pointer-like ones.","lead":"Sparse autoencoder features rarely act as single reusable steering directions in logit space. A new analysis framework shows concept-like features have low-dimensional multi-axis effects, while operation-like features scatter diffusely.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The key “interpretable and causally relevant” claim is supported only at feature-set level: FEGA labels are per-feature, but the causal evidence ablates selected feature sets jointly.","rationale":"The reader appropriately made acceptance conditional on single-model scope, threshold sensitivity, and limited context samples. I agree that those issues constrain generality, but the more direct logical gap in the strongest claim is the unit-of-analysis mismatch: population-level ablation is used to support statements about individual features’ causal relevance and geometry.\n\nThis does not warrant rejection. The FEGA measurements, Gram validation, diagnostic gates, and reported label counts coherently support a narrower claim about the sampled task-conditioned candidate populations. The paper also transparently acknowledges that joint ablation does not establish individual necessity. The appropriate remedy is either to run feature-level causal tests or to revise the headline language to say that causally important feature sets can contain members without stable one-dimensional effects. Thus the reader’s CONDITIONAL verdict remains appropriate, with an additional condition on individual causal attribution.","tokens_in":29709,"tokens_out":3692,"duration_ms":163978,"concrete_test":"For all 239 selected pointer-like features, run single-feature zero ablations at the same site and target position on held-out model-correct prompts. Measure task accuracy and target-token logit change relative to the reconstruction baseline, with paired bootstrap intervals and activation-matched single-feature controls. Intersect the features with significant individual effects with their FEGA labels. If many individually causal features are diffuse/unresolved, the headline conjunction is supported; if individual causality concentrates in structured features or emerges only jointly, it should be narrowed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"FEGA assigns a geometry to each individual feature, but the main causal test in Table 2 jointly removes all k selected pointer-like features and compares that intervention with same-sized random sets. This establishes that the selected population contributes causally, not that each labeled member does. Redundant, synergistic, or merely correlated features could have little or no individual effect while the joint set still produces a large accuracy drop; the paper explicitly says the results “do not imply that every selected latent is individually necessary.” The RAVEL features are likewise chosen by differential binary masking as a jointly sufficient intervention set, not by demonstrating an independent causal effect for every retained latent.\n\nConsequently, for any particular feature labeled unresolved/diffuse or low-dimensional, the paper has not directly established the conjunction asserted by the headline claim—causal relevance plus absence of a stable direction. The aggregate finding that mapped candidate populations rarely form rays remains supported, but the stronger statement that an individual interpretable and causally relevant feature can lack a steering direction requires feature-level causal attribution. This is distinct from the reader’s concern about whether labels generalize beyond the retained contexts.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper introduces Feature-Effect Geometry Analysis (FEGA), a framework that ablates an active SAE feature relative to the SAE reconstruction baseline, propagates the patched activation through the frozen model tail, and analyzes the resulting cloud of logit-space removal effects across contexts. Diagnostics classify each cloud as a directed ray, sign-split axis, vMF directional mixture, global low-dimensional span, centered low-dimensional residual, or diffuse/undefined, using a pre-declared priority order and gate thresholds. Applied to 65k-width ReLU, TopK, and Matryoshka Batch TopK SAEs on the post-layer-12 residual stream of Gemma-2-2B, the authors select value-like features via RAVEL city-country differential binary masking and pointer-like features via recurrence thresholds on four ICL tasks (LSC, WC, PrOntoQA, TT). Joint ablation of the selected pointer-like sets causally degrades task performance far beyond matched random controls. Geometrically, mapped pointer-like features are overwhelmingly diffuse or undefined, while value-like features show low-dimensional structure more often but rarely collapse to a single direction (5-14 directed rays among thousands). The paper concludes that interpretable, causally relevant features frequently lack a stable steering direction.","tokens_in":30051,"tokens_out":3694,"duration_ms":134390,"significance":"If the results hold, the paper makes a substantive and practically relevant point for the SAE interpretability community: feature-level causal relevance does not imply a reusable steering direction, and effect-side auditing (rather than activation-side description) is the right measurement. The engineering and validation standards are unusually high for this literature: the reconstruction-relative ablation baseline cleanly separates feature removal from SAE reconstruction error (§5.1); the Gram-logit equivalence is numerically validated rather than assumed (App. B, Table 6); matched random ablations with paired McNemar tests control for prevalence/magnitude confounds (Table 2); three SAE architectures are compared; and the classification gates are pre-declared with explicit guardrails and a released repository. The value-like/pointer-like distinction, while analogical, generates a testable and confirmed qualitative prediction (structured-but-multi-directional vs diffuse effect clouds).","major_comments":[{"comment":"The headline claim in the abstract and §7 — 'a feature can be interpretable and causally relevant without providing a stable direction for steering' — conjoins per-feature causal relevance with per-feature geometry, but the causal evidence is only set-level. Table 2 jointly ablates all k selected pointer-like features, and the RAVEL population is selected by differential binary masking as a jointly sufficient intervention set (§4.1). The authors themselves note the results 'do not imply that every selected latent is individually necessary' (§4.5). Consequently, for any individual feature labeled unresolved/diffuse in Table 4, causal relevance has not been established; redundant or correlated features could ride on the joint effect. This is fixable within scope: either (a) run per-feature (or small-partition) ablations for at least the mapped subset — e.g., the three directed rays and a s","section":"§4.5, Table 2; abstract; §7"},{"comment":"The rarity claim ('consistent one-dimensional effects are rare') is conditional on a large stack of fixed reporting gates: C_ray ≥ 0.80, S_span^(k) ≥ 0.90, U_span and D_span gates, Δ_mix ≥ 0.10, mode mass ≥ 0.10, assignment stability ≥ 0.80 (Tables 7, 8, 10). No sensitivity analysis or null calibration is provided. Two specific gaps: (i) no null distribution — what fraction of effect clouds from matched-random features (the same controls as Table 2), or of label-permuted/shuffled-context clouds, would pass the ray gate at the same n? Without this, 'rare' is partly a statement about gate stringency. (ii) No robustness check that the structured/diffuse proportions in Tables 4-5 are stable to perturbing the gates (e.g., C_ray at 0.70/0.90, S_span at 0.85/0.95). Given that the ReLU RAVEL population has 4,760/7,715 eligible features in the 'undefined' bin (Table 5), small gate movements could","section":"§5.5, App. H (Tables 7-10); §6, Tables 4-5"},{"comment":"The mapped-subset denominators make the quantitative claims fragile. In Table 4, 159 of 239 selected pointer-like features land in 'Undef.' and only 80 are mapped, of which 74 are unresolved/diffuse; the claim 'mapped pointer-like effects are overwhelmingly diffuse' thus rests on ~34% of the selected population, and the behavioral distinction between 'undefined' and 'unresolved' is not independently grounded. Context sampling is capped at 64 valid contexts with a minimum of 8 (§5.1, App. G), and clouds in the 8≤n<32 range receive only 'exploratory' confidence yet still contribute labels to the tables. The Limitations section (§8) acknowledges this, but the main text does not quantify it. Please report: the n-distribution behind Table 4; how label frequencies change as the retained-context cap and the recurrence thresholds (90%/90%/90%, §4.2) vary; and what fraction of the headline label","section":"§6.1, Table 4; §4.2; §5.1; §8"},{"comment":"The value-like vs pointer-like contrast (Takeaway 5; §6.2) is computed over populations selected by entirely different mechanisms (MDBM intervention sets vs recurrence thresholds) with very different undefined/insufficient rates, and the structured fractions vary sharply by architecture: 14.0% (ReLU) vs 36.3% (TopK) vs 30.3% (Matryoshka) of eligible RAVEL features. This two-fold architecture swing within the value-like population is comparable in size to the value-vs-pointer contrast itself, which suggests architecture and selection-pipeline effects are confounded with the role-based interpretation. The paper should either condition the comparison more carefully (e.g., match on eligibility/undefined rates, or report the contrast within each architecture with appropriate uncertainty quantification) or soften the claim to a within-architecture observation.","section":"§6.2, Table 5 vs Table 4"}],"minor_comments":[{"comment":"Notation drift between Δ_j, Δ_{j,logit}, and Δ_{j,pre} across §5.1-5.2: Δ_j is introduced as the logit-space effect in §5.1 ('Removal Effects') but the pre-logit quantity is later called δ_j while App. B writes Δ_{j,pre}; a consistent subscript convention would reduce confusion.","section":"§5.1-5.2, App. B"},{"comment":"Figure 5 and Figure 6 use UMAP projections that are explicitly 'for visualization only' (Figure 3 caption says this for the cards but the atlas captions do not); please state in the atlas captions that UMAP distances carry no quantitative meaning, especially since the atlases are the only visualization of the full populations.","section":"Figures 5-6"},{"comment":"Table 2 reports p_target < 10^-300 for all twelve cells; with paired McNemar on finite example counts this presumably reflects a lower bound from zero discordant pairs — please state the actual discordant-pair counts or the exact reporting convention, since an unqualified 10^-300 bound is uninformative.","section":"Table 2"},{"comment":"Missing related work: Engels et al., 'Not All Language Model Features Are Linear' (2024/2025) is directly relevant to the claim that features need not act as one-dimensional directions, and Marks & Tegmark on feature geometry would also fit §2. The pointer/value terminology would also benefit from a connection to induction-head and function-vector literatures (Olsson et al. 2022; Todd et al. 2024), which study context-dependent operations the paper's pointer-like features plausibly overlap.","section":"§2"},{"comment":"§4.2's recurrence thresholds (90% of examples; 90% of queries in 90% of families) are asserted without justification; a brief rationale or a sweep in an appendix would help, since Table 1's counts (4 to 78) are sensitive to these choices.","section":"§4.2"},{"comment":"The 'Multi' (directional mixture) column in Table 4 is identically zero while Table 5 shows only 11+1 mixtures for TopK/Matryoshka; given the elaborate vMF apparatus of App. F, a short comment on why mixtures almost never survive the gates (Δ_mix ≥ 0.10, within-mode C_ray ≥ 0.70) would help the reader assess whether the mixture family is operative or vestigial.","section":"Tables 4-5, App. F"}],"recommendation":"major_revision","confidential_remarks":"The framework and experiments are executed at a level of care above the norm for this area (the App. B validation, the reconstruction-relative baseline, and the explicit gate declarations are genuinely good practice). My main hesitation is that the contribution is a measurement framework plus one model/layer of evidence, and the two most interesting empirical statements currently outrun the demonstrated granularity of the causal evidence and the calibration of the classification gates. I believe both are addressable in revision without new conceptual work. Worth verifying the released repository actually reproduces Tables 4-5 before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful move here is measuring the geometry of logit changes after ablating the same SAE feature across contexts, not just describing activations or decoder directions. FEGA is a real method contribution: reconstruction-relative baselines, Gram-logit equivalence checks, explicit gates, and a priority order over ray / axis / mixture / span / residual. That is cleaner than most steering reliability writeups.\n\nWhat they show on Gemma-2-2B layer-12 residual SAEs (ReLU, TopK, Matryoshka 65k) is coherent. Directed rays are scarce. Pointer-like candidates isolated by ICL recurrence are mostly undefined or diffuse when mapped. RAVEL value-like candidates more often land in low-D span or residual structure, usually multi-directional rather than one steering vector. Cross-task reuse of a small ICL feature set, disjoint from RAVEL, plus joint ablations beating matched random controls, is good supporting evidence that they isolated different functional populations.\n\nThe stress-test note is fair and should not be waved away. Table 2 is joint set ablation; FEGA labels are per feature. The paper itself says the ablations do not prove every selected latent is individually necessary. So the slogan “interpretable and causally relevant without a stable direction” is fully earned at the population level, and only partly earned feature-by-feature. That weakens the strongest sentence, not the measurement program.\n\nOther limits are real but proportional: one model and layer, at most 64 contexts, large insufficient/undefined slices, and many free gates (90% recurrence, C_ray, span thresholds). They own most of this in §8. Citations look appropriate; math is dual-PCA / vMF bookkeeping rather than deep theory, and it is careful enough.\n\nThis is for people who actually steer or audit SAEs. I would bring it to reading group, cite the FEGA framing and the value/pointer geometry contrast, and send it to peer review. Ask referees for feature-level causal checks on a subset and sensitivity on the gates—not a desk reject.","headline":"Solid effect-side audit of SAE features: rays are rare, value/pointer geometry split is real within a narrow setup, and the joint-vs-individual causality gap is the main soft spot on the headline claim.","tokens_in":30589,"tokens_out":540,"would_cite":true,"duration_ms":16939,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"SAE features can be clear and causal yet still fail as single steering directions; their logit effects are usually multi-way or diffuse.","keywords":["sparse autoencoders","mechanistic interpretability","feature steering","logit-effect geometry","FEGA","value-like vs pointer-like features","polysemanticity","in-context learning"],"falsifier":"On the same intervention setup, find a large share of causally validated pointer-like or value-like features whose logit-effect clouds concentrate as directed rays (high directed-ray score with stable orientation) across held-out prompts and layers; or show that expanding context samples and selection rules systematically turns the current diffuse clouds into low-dimensional rays.","tokens_in":30460,"feed_emoji":"🧭","tokens_out":907,"duration_ms":17017,"temperature":0.7,"pith_summary":"Sparse autoencoders are widely used to find human-readable features in language models, then to steer those models by pushing features up or down. This paper argues that activation clarity and causal relevance do not guarantee a stable, reusable control direction in the model’s output. The authors intervene by removing the same active feature in many contexts and study the cloud of resulting logit changes with a new unsupervised method, Feature-Effect Geometry Analysis (FEGA). Across three SAE designs on a mid-size model, true one-dimensional “ray” effects are rare. Features tied to static facts more often show low-dimensional but multi-directional structure; features tied to context-dependent operations such as copying or rule-following mostly scatter. The practical upshot is that interpretability and editability come apart: a feature can matter without offering a fixed steering vector.","feed_headline":"Most SAE features are not single steering directions","feed_subtitle":"Logit-effect clouds are usually multi-way or diffuse, even when features look clear and causal","key_machinery":"Feature-Effect Geometry Analysis (FEGA): an unsupervised audit that ablates one active SAE feature relative to the SAE reconstruction, collects the cloud of logit-space removal effects across contexts, and labels that cloud (directed ray, axis, mixture, low-dimensional span, residual structure, or diffuse/undefined) via directional kernels and spectral tests.","core_discovery":"Across SAE variants, consistent one-dimensional downstream logit effects are rare. A feature can be interpretable and causally relevant without providing a stable direction for steering. Value-like features (static factual attributes) more often show structured low-dimensional effects that typically span several directions; pointer-like features (context-dependent operations) predominantly show diffuse effects.","pith_inferences":["Practical control for induction-like behavior may need prompt-conditioned or value-conditioned interventions rather than static feature vectors.","Layer-by-layer FEGA could show whether effects start local and only become diffuse at the final readout, changing where editors should intervene.","Benchmark suites that score SAEs mainly on activation interpretability or single-vector steering will systematically overrate features that look clean but lack stable effect geometry."],"forward_implications":["Steering by a single SAE feature vector will often miss, oppose, or only partially capture the intended output change.","Audits of SAE features should separate what a feature detects, whether it matters causally, and how its effects vary across contexts.","Value-like factual features are better candidates for structured control than pointer-like copying or rule features, but still usually need multi-direction rather than one-vector edits.","SAE training and evaluation should include prompt-local operations, not only factual recall, because architectures organize those features differently.","Diffuse logit geometry does not prove a feature is useless; it may still support a shared operation whose target token changes with context."],"fun_headline_variants":["SAE features rarely act as single reusable steering directions","Few SAE features yield consistent one-dimensional logit effects","Interpretable causal SAE features often lack stable steering","Value-like features span multi-way effects; pointers stay diffuse","Feature-effect clouds show SAE features seldom pure directions"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The geometry labels from a single model layer, three SAE types, task-picked features, and at most a few dozen retained contexts per feature are taken to show that stable one-direction effects are generally rare, not just rare in this sample.","fun_headline_variants_meta":{"raw":{"variants":["SAE features rarely act as single reusable steering directions","Few SAE features yield consistent one-dimensional logit effects","Interpretable causal SAE features often lack stable steering","Value-like features span multi-way effects; pointers stay diffuse","Feature-effect clouds show SAE features seldom pure directions"]},"model":"grok-4.5","effort":"low","cost_usd":0.004884,"raw_usage":{"total_tokens":1357,"prompt_tokens":760,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":48844000,"prompt_tokens_details":{"text_tokens":760,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":535,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":760,"tokens_out":62,"duration_ms":9482,"temperature":1.0,"reasoning_tokens":535,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T09:14:35.476648+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On the same intervention setup, find a large share of causally validated pointer-like or value-like features whose logit-effect clouds concentrate as directed rays (high directed-ray score with stable orientation) across held-out prompts and layers; or show that expanding context samples and selection rules systematically turns the current diffuse clouds into low-dimensional rays.","supporting_citations":[],"review_version":1}