Pith. sign in

REVIEW 4 major objections 6 minor 45 references

SAE features can be clear and causal yet still fail as single steering directions; their logit effects are usually multi-way or diffuse.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 09:14 UTC pith:GU3RGCGI

load-bearing objection Solid effect-side audit of SAE features: rays are rare, value/pointer geometry split is real within a narrow setup, and the joint-vs-individual causality gap is the main soft spot on the headline claim. the 4 major comments →

arxiv 2607.24645 v1 pith:GU3RGCGI submitted 2026-07-27 cs.LG cs.AIcs.CL

Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

classification cs.LG cs.AIcs.CL
keywords sparse autoencodersmechanistic interpretabilityfeature steeringlogit-effect geometryFEGAvalue-like vs pointer-like featurespolysemanticityin-context learning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Sparse autoencoders are widely used to find human-readable features in language models, then to steer those models by pushing features up or down. This paper argues that activation clarity and causal relevance do not guarantee a stable, reusable control direction in the model’s output. The authors intervene by removing the same active feature in many contexts and study the cloud of resulting logit changes with a new unsupervised method, Feature-Effect Geometry Analysis (FEGA). Across three SAE designs on a mid-size model, true one-dimensional “ray” effects are rare. Features tied to static facts more often show low-dimensional but multi-directional structure; features tied to context-dependent operations such as copying or rule-following mostly scatter. The practical upshot is that interpretability and editability come apart: a feature can matter without offering a fixed steering vector.

Core claim

Across SAE variants, consistent one-dimensional downstream logit effects are rare. A feature can be interpretable and causally relevant without providing a stable direction for steering. Value-like features (static factual attributes) more often show structured low-dimensional effects that typically span several directions; pointer-like features (context-dependent operations) predominantly show diffuse effects.

What carries the argument

Feature-Effect Geometry Analysis (FEGA): an unsupervised audit that ablates one active SAE feature relative to the SAE reconstruction, collects the cloud of logit-space removal effects across contexts, and labels that cloud (directed ray, axis, mixture, low-dimensional span, residual structure, or diffuse/undefined) via directional kernels and spectral tests.

Load-bearing premise

The geometry labels from a single model layer, three SAE types, task-picked features, and at most a few dozen retained contexts per feature are taken to show that stable one-direction effects are generally rare, not just rare in this sample.

What would settle it

On the same intervention setup, find a large share of causally validated pointer-like or value-like features whose logit-effect clouds concentrate as directed rays (high directed-ray score with stable orientation) across held-out prompts and layers; or show that expanding context samples and selection rules systematically turns the current diffuse clouds into low-dimensional rays.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Steering by a single SAE feature vector will often miss, oppose, or only partially capture the intended output change.
  • Audits of SAE features should separate what a feature detects, whether it matters causally, and how its effects vary across contexts.
  • Value-like factual features are better candidates for structured control than pointer-like copying or rule features, but still usually need multi-direction rather than one-vector edits.
  • SAE training and evaluation should include prompt-local operations, not only factual recall, because architectures organize those features differently.
  • Diffuse logit geometry does not prove a feature is useless; it may still support a shared operation whose target token changes with context.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Practical control for induction-like behavior may need prompt-conditioned or value-conditioned interventions rather than static feature vectors.
  • Layer-by-layer FEGA could show whether effects start local and only become diffuse at the final readout, changing where editors should intervene.
  • Benchmark suites that score SAEs mainly on activation interpretability or single-vector steering will systematically overrate features that look clean but lack stable effect geometry.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces Feature-Effect Geometry Analysis (FEGA), a framework that ablates an active SAE feature relative to the SAE reconstruction baseline, propagates the patched activation through the frozen model tail, and analyzes the resulting cloud of logit-space removal effects across contexts. Diagnostics classify each cloud as a directed ray, sign-split axis, vMF directional mixture, global low-dimensional span, centered low-dimensional residual, or diffuse/undefined, using a pre-declared priority order and gate thresholds. Applied to 65k-width ReLU, TopK, and Matryoshka Batch TopK SAEs on the post-layer-12 residual stream of Gemma-2-2B, the authors select value-like features via RAVEL city-country differential binary masking and pointer-like features via recurrence thresholds on four ICL tasks (LSC, WC, PrOntoQA, TT). Joint ablation of the selected pointer-like sets causally degrades task performance far beyond matched random controls. Geometrically, mapped pointer-like features are overwhelmingly diffuse or undefined, while value-like features show low-dimensional structure more often but rarely collapse to a single direction (5-14 directed rays among thousands). The paper concludes that interpretable, causally relevant features frequently lack a stable steering direction.

Significance. If the results hold, the paper makes a substantive and practically relevant point for the SAE interpretability community: feature-level causal relevance does not imply a reusable steering direction, and effect-side auditing (rather than activation-side description) is the right measurement. The engineering and validation standards are unusually high for this literature: the reconstruction-relative ablation baseline cleanly separates feature removal from SAE reconstruction error (§5.1); the Gram-logit equivalence is numerically validated rather than assumed (App. B, Table 6); matched random ablations with paired McNemar tests control for prevalence/magnitude confounds (Table 2); three SAE architectures are compared; and the classification gates are pre-declared with explicit guardrails and a released repository. The value-like/pointer-like distinction, while analogical, generates a testable and confirmed qualitative prediction (structured-but-multi-directional vs diffuse effect clouds).

major comments (4)
  1. [§4.5, Table 2; abstract; §7] The headline claim in the abstract and §7 — 'a feature can be interpretable and causally relevant without providing a stable direction for steering' — conjoins per-feature causal relevance with per-feature geometry, but the causal evidence is only set-level. Table 2 jointly ablates all k selected pointer-like features, and the RAVEL population is selected by differential binary masking as a jointly sufficient intervention set (§4.1). The authors themselves note the results 'do not imply that every selected latent is individually necessary' (§4.5). Consequently, for any individual feature labeled unresolved/diffuse in Table 4, causal relevance has not been established; redundant or correlated features could ride on the joint effect. This is fixable within scope: either (a) run per-feature (or small-partition) ablations for at least the mapped subset — e.g., the three directed rays and a s
  2. [§5.5, App. H (Tables 7-10); §6, Tables 4-5] The rarity claim ('consistent one-dimensional effects are rare') is conditional on a large stack of fixed reporting gates: C_ray ≥ 0.80, S_span^(k) ≥ 0.90, U_span and D_span gates, Δ_mix ≥ 0.10, mode mass ≥ 0.10, assignment stability ≥ 0.80 (Tables 7, 8, 10). No sensitivity analysis or null calibration is provided. Two specific gaps: (i) no null distribution — what fraction of effect clouds from matched-random features (the same controls as Table 2), or of label-permuted/shuffled-context clouds, would pass the ray gate at the same n? Without this, 'rare' is partly a statement about gate stringency. (ii) No robustness check that the structured/diffuse proportions in Tables 4-5 are stable to perturbing the gates (e.g., C_ray at 0.70/0.90, S_span at 0.85/0.95). Given that the ReLU RAVEL population has 4,760/7,715 eligible features in the 'undefined' bin (Table 5), small gate movements could
  3. [§6.1, Table 4; §4.2; §5.1; §8] The mapped-subset denominators make the quantitative claims fragile. In Table 4, 159 of 239 selected pointer-like features land in 'Undef.' and only 80 are mapped, of which 74 are unresolved/diffuse; the claim 'mapped pointer-like effects are overwhelmingly diffuse' thus rests on ~34% of the selected population, and the behavioral distinction between 'undefined' and 'unresolved' is not independently grounded. Context sampling is capped at 64 valid contexts with a minimum of 8 (§5.1, App. G), and clouds in the 8≤n<32 range receive only 'exploratory' confidence yet still contribute labels to the tables. The Limitations section (§8) acknowledges this, but the main text does not quantify it. Please report: the n-distribution behind Table 4; how label frequencies change as the retained-context cap and the recurrence thresholds (90%/90%/90%, §4.2) vary; and what fraction of the headline label
  4. [§6.2, Table 5 vs Table 4] The value-like vs pointer-like contrast (Takeaway 5; §6.2) is computed over populations selected by entirely different mechanisms (MDBM intervention sets vs recurrence thresholds) with very different undefined/insufficient rates, and the structured fractions vary sharply by architecture: 14.0% (ReLU) vs 36.3% (TopK) vs 30.3% (Matryoshka) of eligible RAVEL features. This two-fold architecture swing within the value-like population is comparable in size to the value-vs-pointer contrast itself, which suggests architecture and selection-pipeline effects are confounded with the role-based interpretation. The paper should either condition the comparison more carefully (e.g., match on eligibility/undefined rates, or report the contrast within each architecture with appropriate uncertainty quantification) or soften the claim to a within-architecture observation.
minor comments (6)
  1. [§5.1-5.2, App. B] Notation drift between Δ_j, Δ_{j,logit}, and Δ_{j,pre} across §5.1-5.2: Δ_j is introduced as the logit-space effect in §5.1 ('Removal Effects') but the pre-logit quantity is later called δ_j while App. B writes Δ_{j,pre}; a consistent subscript convention would reduce confusion.
  2. [Figures 5-6] Figure 5 and Figure 6 use UMAP projections that are explicitly 'for visualization only' (Figure 3 caption says this for the cards but the atlas captions do not); please state in the atlas captions that UMAP distances carry no quantitative meaning, especially since the atlases are the only visualization of the full populations.
  3. [Table 2] Table 2 reports p_target < 10^-300 for all twelve cells; with paired McNemar on finite example counts this presumably reflects a lower bound from zero discordant pairs — please state the actual discordant-pair counts or the exact reporting convention, since an unqualified 10^-300 bound is uninformative.
  4. [§2] Missing related work: Engels et al., 'Not All Language Model Features Are Linear' (2024/2025) is directly relevant to the claim that features need not act as one-dimensional directions, and Marks & Tegmark on feature geometry would also fit §2. The pointer/value terminology would also benefit from a connection to induction-head and function-vector literatures (Olsson et al. 2022; Todd et al. 2024), which study context-dependent operations the paper's pointer-like features plausibly overlap.
  5. [§4.2] §4.2's recurrence thresholds (90% of examples; 90% of queries in 90% of families) are asserted without justification; a brief rationale or a sweep in an appendix would help, since Table 1's counts (4 to 78) are sensitive to these choices.
  6. [Tables 4-5, App. F] The 'Multi' (directional mixture) column in Table 4 is identically zero while Table 5 shows only 11+1 mixtures for TopK/Matryoshka; given the elaborate vMF apparatus of App. F, a short comment on why mixtures almost never survive the gates (Δ_mix ≥ 0.10, within-mode C_ray ≥ 0.70) would help the reader assess whether the mixture family is operative or vestigial.

Circularity Check

0 steps flagged

No significant circularity: task-conditioned feature selection and FEGA geometry labels are independent measurement stages, not definitionally equivalent.

full rationale

The paper’s load-bearing chain is empirical, not definitional. Value-like candidates are isolated by RAVEL/MDBM attribute interventions and pointer-like candidates by cross-prompt activation recurrence on ICL tasks (Sections 4.1–4.2); causal contribution is then tested by joint ablation against matched random controls (Table 2, Section 4.5). FEGA separately constructs reconstruction-relative logit-effect clouds and classifies their geometry with pre-declared diagnostics and gates (Sections 5.1–5.5, Appendices G–H). Geometry labels are not inputs to selection, and selection criteria are not rewritten as geometry outcomes. The value/pointer framing is an interpretive analogy for task roles, not a quantity fitted from the same effect clouds it is said to explain. Known steering unreliability is cited from external work and re-measured on a new object (downstream effect geometry), not renamed into a forced theorem. No uniqueness result, self-citation chain, or fitted parameter is smuggled in as a first-principles prediction. Gaps such as set-level vs per-feature causality affect claim strength, not circularity of the derivation.

Axiom & Free-Parameter Ledger

7 free parameters · 5 axioms · 3 invented entities

The claim rests on standard SAE/intervention practice plus paper-specific operational definitions (value vs pointer selection, reconstruction-relative effects, FEGA family gates). No new physical entities; the main inventions are methodological objects and role labels.

free parameters (7)
  • ICL feature recurrence thresholds (90% examples; 90% queries in 90% families) = 0.90 / 0.90 / 0.90
    Hand-chosen gates that define the pointer-like candidate sets whose geometry is then summarized.
  • Minimum valid contexts for geometry claim = 8
    Features with fewer than eight nonzero finite effects are excluded from family labels.
  • Directed-ray concentration gate C_ray = 0.80
    Primary threshold for calling a cloud a reusable direction.
  • Span sufficiency and related spectral gates (S_span^k, U_span^k, D_span^k, ranks) = S_span^k>=0.90; other gates in App. H/Tables 7–10
    Hand-set reporting thresholds that decide low-dimensional structure vs diffuse/unresolved labels.
  • Mixture gates (mode mass, Delta_mix, within-mode C_ray, assignment stability) = pi_r>=0.10; Delta_mix>=0.10; C_ray,r>=0.70; stability>=0.80
    Determine acceptance of multi-mode geometry.
  • Max contexts retained per feature = 64
    Caps the empirical cloud used for all spectra and stability checks.
  • Zero-effect filter fraction = 0.30
    Drops clouds if too many active contexts have near-zero measured effect.
axioms (5)
  • domain assumption Reconstruction-relative feature zeroing isolates the causal contribution of SAE latent j without confounding by SAE reconstruction error.
    §5.1 baseline choice; standard in some SAE causal analyses but still an modeling choice versus intervening on raw activations.
  • domain assumption Linear logit readout geometry (via W_U and Gram G) is the right space in which to judge steering-relevant effect consistency.
    §5.1–5.2; post-softcap logits are explicitly rejected after validation (App. B).
  • ad hoc to paper RAVEL city-country MDBM-selected latents are value-like; high-recurrence ICL latents are pointer-like.
    §4 operational definitions used to interpret geometry differences; not an independently validated ontology of all SAE features.
  • ad hoc to paper FEGA family priority (ray > axis > mixture > global span > residual > fallbacks) yields the scientifically appropriate primary label.
    §5.5 and Fig. 4; different priority would reallocate counts among structured vs diffuse categories.
  • standard math Dual-PCA / kernel PCA on normalized effect directions and vMF mixtures are valid summaries of high-dimensional logit clouds.
    §5.2–5.3; Schölkopf et al. kernel PCA and Banerjee et al. vMF clustering.
invented entities (3)
  • Feature-Effect Geometry Analysis (FEGA) no independent evidence
    purpose: Unsupervised labeling of SAE ablation effect clouds in logit space.
    Core methodological contribution; diagnostics and gates are paper-defined.
  • Value-like vs pointer-like feature roles no independent evidence
    purpose: Interpret why some effect clouds are structured and others diffuse.
    Borrowed programming analogy operationalized via RAVEL vs ICL tasks; useful taxonomy but not independently measured outside these tasks.
  • Downstream logit-effect cloud E_j independent evidence
    purpose: Primary object whose geometry replaces single steering-vector assumptions.
    Defined from reconstruction-relative ablations across valid contexts (§5.1).

pith-pipeline@v1.2.0-grok45-kimik3 · 33655 in / 3733 out tokens · 63166 ms · 2026-07-31T09:14:35.476648+00:00 · methodology

0 comments
read the original abstract

The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change. Prior work has studied feature geometry inside the model, where features are computed. We instead study the geometry of changes in model logits caused by feature interventions. We introduce Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes. Across SAE variants, consistent one-dimensional effects are rare: few features behave like reusable directions. To interpret this variation, we distinguish value-like features, tied to static information such as factual attributes, from pointer-like features, associated with context-dependent operations. Value-like features more often exhibit structured, low-dimensional effects, although these effects typically span several directions. Pointer-like features, by contrast, predominantly exhibit diffuse effects. Our results show that a feature can be interpretable and causally relevant without providing a stable direction for steering.

Figures

Figures reproduced from arXiv: 2607.24645 by Anwoy Chatterjee, Iryna Gurevych, Phu Gia Hoang, Subhabrata Dutta, Tanmoy Chakraborty.

Figure 1
Figure 1. Figure 1: SAE feature steering should be evaluated by downstream effect geometry. Two contexts, ca and cb, activate the same feature j at layer ℓ. Green paths show the SAE reconstruction baseline; red paths show the feature-ablated reconstruction. Both are propagated through the frozen model tail and logit readout map ϕ(·), producing one downstream logit effect per context, ∆ (ca) j,logit and ∆ (cb) j,logit. feature… view at source ↗
Figure 2
Figure 2. Figure 2: Cross-task feature overlap. Intersection over Union (IoU) of the selected feature sets. Several features recur across the ICL tasks, with greater overlap for TopK and Matryoshka Batch TopK than for ReLU. None of the ICL-selected features overlaps with the corresponding RAVEL feature set. 4.4 Shared Features Across In-Context Tasks If particular SAE features contribute to general ICL and induction-like beha… view at source ↗
Figure 3
Figure 3. Figure 3: Overview of FEGA geometry labels and secondary flags. Each card shows a representative downstream logit-effect cloud after removing one SAE feature. The left panel visualizes normalized effect directions; the right panel gives a two-dimensional view for visualization only. The six main cards illustrate the strict geometry families and the unresolved outcome, rather than the complete reporting vocabulary. T… view at source ↗
Figure 4
Figure 4. Figure 4: FEGA label assignment and qualification. Full-sample diagnostics select the first supported geometry family and, where applicable, its smallest supported dimension. Family-specific stability tests then qualify this selection; directional-mixture acceptance already includes assignment stability. High CVm,j adds a magnitude-instability flag, indicating that effect strength varies across contexts even when th… view at source ↗
Figure 5
Figure 5. Figure 5: FEGA atlas for pointer-like features. UMAP projections of downstream logit-effect geometry across four in-context tasks. Marker shape denotes the task and color denotes the FEGA label; features with undefined geometry are omitted. Features shared across all four tasks are shown at higher opacity. Clouds with insufficient valid effects are excluded before geometry selection. For each remaining cloud, FEGA c… view at source ↗
Figure 6
Figure 6. Figure 6: FEGA atlas for value-like features. UMAP projections of the RAVEL features for which FEGA assigns a primary geometry label. Features with insufficient effect evidence or undefined geometry are omitted from the atlas and retained in the complete counts in [PITH_FULL_IMAGE:figures/full_fig_p020_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

45 extracted references · 4 canonical work pages · 2 internal anchors

  1. [1]

    Journal of Machine Learning Research , year =

    Atticus Geiger and Duligur Ibeling and Amir Zur and Maheep Chaudhary and Sonakshi Chauhan and Jing Huang and Aryaman Arora and Zhengxuan Wu and Noah Goodman and Christopher Potts and Thomas Icard , title =. Journal of Machine Learning Research , year =

  2. [2]

    CoRR , volume =

    Usha Bhalla and Thomas Fel and Can Rager and Sheridan Feucht and Tal Haklay and Daniel Wurgaft and Siddharth Boppana and Matthew Kowal and Vasudev Shyam and Jack Merullo and Atticus Geiger and Ekdeep Singh Lubana , title =. CoRR , volume =. 2026 , url =. doi:10.48550/ARXIV.2604.28119 , eprinttype =. 2604.28119 , timestamp =

  3. [3]

    CoRR , volume =

    Maheep Chaudhary and Atticus Geiger , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2409.04478 , eprinttype =. 2409.04478 , timestamp =

  4. [4]

    Finding Belief Geometries with Sparse Autoencoders

    Matthew Levinson , title =. CoRR , volume =. 2026 , url =. doi:10.48550/ARXIV.2604.02685 , eprinttype =. 2604.02685 , timestamp =

  5. [5]

    Manning and Christopher Potts , editor =

    Zhengxuan Wu and Aryaman Arora and Atticus Geiger and Zheng Wang and Jing Huang and Dan Jurafsky and Christopher D. Manning and Christopher Potts , editor =. 2025 , url =

  6. [6]

    Galichin and Alexey Dontsov and Oleg Rogov and Ivan V

    Anton Korznikov and Andrey V. Galichin and Alexey Dontsov and Oleg Rogov and Ivan V. Oseledets and Elena Tutubalina , title =. CoRR , volume =. 2026 , url =. doi:10.48550/ARXIV.2602.14111 , eprinttype =. 2602.14111 , timestamp =

  7. [7]

    Li and Suraj Srinivas and Usha Bhalla and Himabindu Lakkaraju , editor =

    Aaron J. Li and Suraj Srinivas and Usha Bhalla and Himabindu Lakkaraju , editor =. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics,. 2026 , url =. doi:10.18653/V1/2026.EACL-LONG.279 , timestamp =

  8. [8]

    and Lubana, Ekdeep S and Fel, Thomas and Ba, Demba , booktitle =

    Hindupur, Sai Sumedh R. and Lubana, Ekdeep S and Fel, Thomas and Ba, Demba , booktitle =

  9. [9]

    Michaud and David D

    Yuxiao Li and Eric J. Michaud and David D. Baek and Joshua Engels and Xiaoqing Sun and Max Tegmark , title =. Entropy , volume =. 2025 , url =. doi:10.3390/E27040344 , timestamp =

  10. [10]

    Transformer Circuits Thread , note=

    Elhage, Nelson and Hume, Tristan and Olsson, Catherine and Schiefer, Nicholas and Henighan, Tom and Kravec, Shauna and Hatfield-Dodds, Zac and Lasenby, Robert and Drain, Dawn and Chen, Carol and Grosse, Roger and McCandlish, Sam and Kaplan, Jared and Amodei, Dario and Wattenberg, Martin and Olah, Christopher , year=. Transformer Circuits Thread , note=

  11. [11]

    Transformer Circuits Thread , note=

    Bricken, Trenton and Templeton, Adly and Batson, Joshua and Chen, Brian and Jermyn, Adam and Conerly, Tom and Turner, Nick and Anil, Cem and Denison, Carson and Askell, Amanda and Lasenby, Robert and Wu, Yifan and Kravec, Shauna and Schiefer, Nicholas and Maxwell, Tim and Joseph, Nicholas and Hatfield-Dodds, Zac and Tamkin, Alex and Nguyen, Karina and McL...

  12. [12]

    The Twelfth International Conference on Learning Representations,

    Robert Huben and Hoagy Cunningham and Logan Riggs Smith and Aidan Ewart and Lee Sharkey , title =. The Twelfth International Conference on Learning Representations,. 2024 , url =

  13. [13]

    Daniel and Sumers, Theodore R

    Templeton, Adly and Conerly, Tom and Marcus, Jonathan and Lindsey, Jack and Bricken, Trenton and Chen, Brian and Pearce, Adam and Citro, Craig and Ameisen, Emmanuel and Jones, Andy and Cunningham, Hoagy and Turner, Nicholas L and McDougall, Callum and MacDiarmid, Monte and Freeman, C. Daniel and Sumers, Theodore R. and Rees, Edward and Batson, Joshua and ...

  14. [14]

    2024 , url =

    Esin Durmus and Alex Tamkin and Jack Clark and Jerry Wei and Jonathan Marcus and Joshua Batson and Kunal Handa and Liane Lovitt and Meg Tong and Miles McCain and Oliver Rausch and Saffron Huang and Sam Bowman and Stuart Ritchie and Tom Henighan and Deep Ganguli , title =. 2024 , url =

  15. [15]

    2025 , url =

    Adam Karvonen and Can Rager and Johnny Lin and Curt Tigges and Joseph Isaac Bloom and David Chanin and Yeu. 2025 , url =

  16. [16]

    CoRR , volume =

    Nora Belrose and Zach Furman and Logan Smith and Danny Halawi and Igor Ostrovsky and Lev McKinney and Stella Biderman and Jacob Steinhardt , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2303.08112 , eprinttype =. 2303.08112 , timestamp =

  17. [17]

    CoRR , volume =

    Dana Arad and Aaron Mueller and Yonatan Belinkov , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2505.20063 , eprinttype =. 2505.20063 , timestamp =

  18. [18]

    The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=

    Daniel Chee Hian Tan and David Chanin and Aengus Lynch and Brooks Paige and Dimitrios Kanoulas and Adri. The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=

  19. [19]

    2025 , url=

    Joschka Braun and Carsten Eickhoff and David Krueger and Seyed Ali Bahrainian and Dmitrii Krasheninnikov , booktitle =. 2025 , url=

  20. [20]

    2024 , url=

    Harry Mayne and Yushi Yang and Adam Mahdi , booktitle =. 2024 , url=

  21. [21]

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    Huang, Jing and Wu, Zhengxuan and Potts, Christopher and Geva, Mor and Geiger, Atticus. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. doi:10.18653/v1/2024.acl-long.470

  22. [22]

    Philosophical Transactions of the Royal Society of London, Series A: Containing Papers of a Mathematical or Physical Character , number =

    Pearson, Karl , title =. Philosophical Transactions of the Royal Society of London, Series A: Containing Papers of a Mathematical or Physical Character , number =. 1895 , month =. doi:10.1098/rsta.1895.0010 , url =

  23. [23]

    Neural Comput

    Bernhard Sch. Neural Comput. , volume =. 1998 , url =. doi:10.1162/089976698300017467 , timestamp =

  24. [24]

    2007 , publisher=

    Von Luxburg, Ulrike , journal=. 2007 , publisher=

  25. [25]

    Dhillon and Joydeep Ghosh and Suvrit Sra , title =

    Arindam Banerjee and Inderjit S. Dhillon and Joydeep Ghosh and Suvrit Sra , title =. Journal of Machine Learning Research , year =

  26. [26]

    1978 , publisher=

    Schwarz, Gideon , journal=. 1978 , publisher=

  27. [27]

    2024 , url =

    CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2408.00118 , eprinttype =. 2408.00118 , timestamp =

  28. [28]

    Open Problems in Mechanistic Interpretability , journal =

    Lee Sharkey and Bilal Chughtai and Joshua Batson and Jack Lindsey and Jeffrey Wu and Lucius Bushnaq and Nicholas Goldowsky. Open Problems in Mechanistic Interpretability , journal =. 2025 , url =

  29. [29]

    Brown, Tom and Mann, Benjamin and Ryder, Nick and Subbiah, Melanie and Kaplan, Jared D and Dhariwal, Prafulla and Neelakantan, Arvind and Shyam, Pranav and Sastry, Girish and Askell, Amanda and Agarwal, Sandhini and Herbert-Voss, Ariel and Krueger, Gretchen and Henighan, Tom and Child, Rewon and Ramesh, Aditya and Ziegler, Daniel and Wu, Jeffrey and Winte...

  30. [30]

    Transformer Circuits Thread , note=

    Olsson, Catherine and Elhage, Nelson and Nanda, Neel and Joseph, Nicholas and DasSarma, Nova and Henighan, Tom and Mann, Ben and Askell, Amanda and Bai, Yuntao and Chen, Anna and Conerly, Tom and Drain, Dawn and Ganguli, Deep and Hatfield-Dodds, Zac and Hernandez, Danny and Johnston, Scott and Jones, Andy and Kernion, Jackson and Lovitt, Liane and Ndousse...

  31. [31]

    2023 , url=

    Abulhair Saparov and He He , booktitle =. 2023 , url=

  32. [32]

    2025 , url=

    Jingcheng Niu and Subhabrata Dutta and Ahmed Elshabrawy and Harish Tayyar Madabushi and Iryna Gurevych , journal=. 2025 , url=

  33. [33]

    CoRR , volume =

    Dong Shu and Xuansheng Wu and Haiyan Zhao and Daking Rai and Ziyu Yao and Ninghao Liu and Mengnan Du , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2503.05613 , eprinttype =. 2503.05613 , timestamp =

  34. [34]

    The Rate-Distortion-Polysemanticity Tradeoff in SAEs

    Tommaso Mencattini and Francesco Montagna and Francesco Locatello , title =. CoRR , volume =. 2026 , url =. doi:10.48550/ARXIV.2605.14694 , eprinttype =. 2605.14694 , timestamp =

  35. [35]

    The Fourteenth International Conference on Learning Representations , year=

    On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted Remedy , author=. The Fourteenth International Conference on Learning Representations , year=

  36. [36]

    The Thirteenth International Conference on Learning Representations,

    Gouki Minegishi and Hiroki Furuta and Yusuke Iwasawa and Yutaka Matsuo , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =

  37. [37]

    The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=

    A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders , author=. The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=

  38. [38]

    The Thirteenth International Conference on Learning Representations , year=

    Scaling and evaluating sparse autoencoders , author=. The Thirteenth International Conference on Learning Representations , year=

  39. [39]

    2023 , howpublished =

    Language models can explain neurons in language models , author=. 2023 , howpublished =

  40. [40]

    Oliphant and Matt Haberland and Tyler Reddy and David Cournapeau and Evgeni Burovski and Pearu Peterson and Warren Weckesser and Jonathan Bright and St

    Pauli Virtanen and Ralf Gommers and Travis E. Oliphant and Matt Haberland and Tyler Reddy and David Cournapeau and Evgeni Burovski and Pearu Peterson and Warren Weckesser and Jonathan Bright and St. SciPy 1.0-Fundamental Algorithms for Scientific Computing in Python , journal =. 2019 , url =. 1907.10121 , timestamp =

  41. [41]

    Lozier , title =

    Daniel W. Lozier , title =. Ann. Math. Artif. Intell. , volume =. 2003 , url =. doi:10.1023/A:1022915830921 , timestamp =

  42. [42]

    Proceedings of the Eighteenth Annual

    David Arthur and Sergei Vassilvitskii , title =. Proceedings of the Eighteenth Annual. 2007 , crossref =

  43. [43]

    2007 , url =

    Proceedings of the Eighteenth Annual. 2007 , url =

  44. [44]

    Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

    Lieberum, Tom and Rajamanoharan, Senthooran and Conmy, Arthur and Smith, Lewis and Sonnerat, Nicolas and Varma, Vikrant and Kramar, Janos and Dragan, Anca and Shah, Rohin and Nanda, Neel. Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2. Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP....

  45. [45]

    Proceedings of the 42nd International Conference on Machine Learning , pages =

    Learning Multi-Level Features with Matryoshka Sparse Autoencoders , author =. Proceedings of the 42nd International Conference on Machine Learning , pages =. 2025 , editor =