Pith. sign in

REVIEW 3 major objections 5 minor 79 references

Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The Koopman spectrum of a transformer's depth dynamics is a coordinate-free model property recoverable from finite calibration data at the optimal rate.

desk verdict A genuinely new spectral-identifiability claim for interpretability, undermined by a control-affineness gap in the main theorem; deserves review but needs revision. read the letter →

arxiv 2608.10172 v1 pith:KLKU5Q5K submitted 2026-08-10 cs.LG

classification cs.LG
keywords KoopmanoperatorspectralidentifiabilitymechanisticinterpretabilitysparseautoencodersdictionarylearningEDMDctransformerdepthdynamicsnon-normality
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper aims to give mechanistic interpretability a criterion for when a discovered structure is a property of the model rather than an artifact of the method. It treats the transformer forward pass as a controlled dynamical system with depth as time and lifts it through the Koopman operator, so that any dictionary whose span is closed under the layer map induces a finite linear realisation $\Psi(F(x,u))=A\Psi(x)+Bu$ whose eigenvalues are coordinate-free. The main theorem proves this spectrum is recoverable from $M$ calibration samples at rate $M^{-1/2}$ up to permutation, with a matching minimax lower bound and a heavy-tailed median-of-means variant. A companion dissociation theorem shows that whenever the realisation is non-normal, the directions carrying activation variance and the directions carrying information across depth cannot coincide, so the identifiable object and the legible object are provably distinct. The measurements on three pretrained transformers report the predicted exponent on the largest model and report their own failures, because each failure bounds the claim.

What carries the argument

The load-bearing object is the controlled Koopman realisation $(A,B)$, the finite linear map satisfying $\Psi(F(x,u))=A\Psi(x)+Bu$ on a dictionary span that is $K$-invariant, meaning closed under composition with the layer map for every control value. The eigenvalues of $A$ are the identifiable invariants; the EDMDc estimator, the least-squares fit of $(A,B)$ from cached triples of lifted state, lift of next state, and control, is the procedure that extracts them from finite data. The proof chain uses sub-Gaussian matrix concentration for the empirical Gramians, least-squares stability on the event that the state Gramian is well conditioned, and spectral perturbation bounds that convert an operator-norm error into a permutation-matched eigenvalue error, with an optional optimal-matching bound that removes the spectral-gap hypothesis. The dissociation theorem is carried by the discrete Lyapunov equation $\Sigma=A\Sigma A^*+BG_uB^*$: if an orthonormal eigenbasis of the stationary covariance $\Sigma$ consisted of eigenvectors of $A$, then $A$ would be normal, so measured condition numbers $\kappa_2(\widehat V)$ from $38$ to $495$ force the variance-ordered principal directions and the Koopman modes apart.

What would settle it

On a synthetic controlled layer map with an exactly invariant dictionary, compute the EDMDc spectrum from $M$ samples for values of $M$ crossing the theorem's threshold and compare the maximum matched error to the eigenvalues of the known matrix $A$; if the error does not fall at the predicted $M^{-1/2}$ rate once the threshold is crossed, the theorem's mechanism is wrong.

Watch

Extended reading notes

Core claim

The central claim is that transformer depth dynamics carry a finite-dimensional spectral invariant that is both coordinate-free and recoverable at the optimal statistical rate. Under the three structural assumptions, namely dictionary $K$-invariance, persistent excitation of the controls, and spectral separation, the EDMDc estimator returns a matrix $\widehat A_M$ whose eigenvalue multiset converges to that of the true Koopman compression $A$ at rate $M^{-1/2}$, with permutation as the only remaining ambiguity (Theorem 6.1); a minimax lower bound shows no estimator can do better, and a gap-free optimal-matching version removes the spectral-separation threshold for the eigenvalue guarantee (Theorem 6.8). The same theory identifies why sparse autoencoders are not identifiable: their reconstruction-plus-$\ell^1$ objective never enforces $K$-invariance, so the resulting bias does not vanish as $M$ grows, and an explicit invariance penalty partially repairs the failure. On three pretrained transformers the split-half spectral distance falls monotonically with $M$ and attains the predicted exponent on the largest model. Finally, the paper proves that in the non-normal regime these models occupy, the variance-ordered principal basis and the Koopman modal basis cannot coincide; the identifiable spectrum is therefore a certificate of model-intrinsic structure, not a decomposition into human-legible mechanisms.

Load-bearing premise

The proofs assume a dictionary whose span maps exactly into itself under the layer transformation; real dictionaries only approximate that closure, and the resulting bias does not vanish as more calibration samples are collected.

Editorial extensions

If this is right

  • Given a dictionary satisfying exact K-invariance, the recovered spectrum is certified as model-intrinsic: different calibration corpora, seeds, and dictionary bases return the same eigenvalue multiset up to the stated $O(M^{-1/2})$ error, so spectral claims carry a measurable error bar.
  • Sparse-autoencoder variability is diagnosed as structural: the objective omits K-invariance, the resulting bias persists as the sample grows, and adding the invariance penalty reduces split-half spectral distance by 41% at matched sparsity.
  • Because measured condition numbers place every fitted realisation in the non-normal regime, the Koopman spectrum cannot double as a legible circuit decomposition; the IOI experiments make the separation quantitative, with the PCA advantage decaying 4.1× as the question moves in depth.
  • The minimax lower bound makes the $M^{-1/2}$ rate optimal for the problem, and the median-of-means variant extends the guarantee to heavy-tailed activations even though on the tested dictionaries the lifting itself removes the tails.
  • Spectral equality of two realisations implies identical first-order intervention algebras, so cross-model universality becomes a testable spectral criterion, with the caveat that the criterion requires a seed-replica null rather than a sampling floor.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper leaves implicit is to use the measured invariance residual as a dictionary-selection score: as the residual shrinks, the theorem predicts the split-half spectral exponent should approach $-1/2$, which would turn dictionary design for interpretability into a quantitative optimisation problem.
  • The dissociation result suggests that future work should assign different tools to different claims: the Koopman spectrum for cross-run and cross-model identity certificates, and variance-based or causal directions for behavioural localisation, with the two never conflated in a safety case.
  • The exogeneity convention is the largest unresolved modelling choice; lifting to the joint $T$-token state or instrumenting the attention writes would likely change the spectrum more than any sampling error, so resolving it should precede any attempt to compare spectra across architectures.
  • One could also test the theory's dimension dependence directly: the gap-free bound predicts at most linear growth of the matched spectral level in dictionary size $N$, and the paper's own exponents flatten with $N$; extending the sweep to $N=256$ and $512$ would show whether the dimensional factor in the bound is tight.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes treating a transformer forward pass as a controlled dynamical system in depth, lifting it with the Koopman operator, and defining the spectrum of the resulting finite linear realisation as a coordinate-free ``intrinsic'' object for mechanistic interpretability. The central theoretical claim is Theorem 6.1: under exact dictionary K-invariance, persistent excitation, spectral separation, sub-Gaussian tails and diagonalisability, the EDMDc estimator recovers the eigenvalues of the Koopman realisation A from M calibration samples at rate M^{-1/2}, up to permutation, with a matching minimax lower bound and a median-of-means heavy-tailed variant. The paper also proves a dissociation theorem (Theorem 8.1) showing that non-normality forces the variance-ordered principal basis and the Koopman modal basis apart, and derives consequences for SAE non-identifiability, intervention calculus, and model reduction. Experiments on GPT-2 small, Gemma-2-2B and Qwen3-8B-Base report convergence of the spectrum, observation of the predicted M^{-1/2} rate on Qwen3-8B-Base, and several explicitly reported failures including the universality criterion and the SAE invariance-gap criterion under variance-based feature selection.

Significance. If the central derivation can be completed, the paper would provide the first identifiability theorem for a mechanistic-interpretability primitive, with explicit rates, a matching lower bound, and a clear statement of what is and is not identified. The paper is also exemplary in reporting negative results and confounds: the universality criterion fails its seed-replica control, the SAE gap reverses under variance-based selection, and the control-exogeneity convention moves the spectrum by 5--9 times the split-half floor. The norm-growth eigenvalue prediction and the 4.1x depth-decay of the PCA advantage are genuinely falsifiable measurements. These strengths make the manuscript potentially important, but the current proof has load-bearing gaps that must be repaired before the claims can be accepted.

major comments (3)
  1. [Definition 4.6 and Section 6.4] Assumption 1 does not imply the affine-in-control identity (9). For F(x,u)=x+u and Psi=(1,x,x^2), the span is exactly K-invariant for every u, but Psi(F(x,u))=(1, x+u, x^2+2ux+u^2), and the coefficient of x^2 is quadratic in u, so no constant pair (A,B) can satisfy (9). Theorem 5.1 proves only the averaged identity E_u[Psi(F(x,u))]=A Psi(x) in Eq. (24). Consequently the object whose spectrum Theorem 6.1 certifies is not shown to be the Koopman realisation of Definition 4.6; it is at best the population least-squares matrix of Eq. (10). The paper should either impose an explicit control-affine closure condition, or redefine the estimand as the least-squares pair (10) and prove Theorem 6.1 for that pair. As written, Section 6.4's decomposition of the EDMDc estimator relies on an exact linear relation that the hypotheses do not provide.
  2. [Lemma 6.3 and Theorem 7.2] The concentration argument for the cross-Gramian \hat C_M = (1/M) Y X^* is asserted for the 'sub-exponential random matrix' Psi(F(x,u)) Psi(x)^*, but condition (R3) bounds only Psi(x) and u. Since F includes the MLP with RMSNorm, sub-Gaussianity of x and u does not imply any tail bound on Psi(F(x,u)). The same gap appears in Theorem 7.2, where (R3') supplies moments only for Psi(x) and u. An explicit hypothesis on the lifted next state---for example sub-Gaussianity of Psi(F(x,u)) under mu \otimes nu, or a uniform bound on the pointwise Koopman matrices A_u of Lemma 5.2---must be stated and propagated through Theorems 6.1 and 7.2.
  3. [Sections 6.8, 11.7, and 12 (cited as Theorem 6.12)] The approximate-invariance extension is repeatedly cited as 'Theorem 6.12' and 'Theorem 5.7', but no such numbered theorems are stated or proved anywhere in the manuscript. Remark 6.12 merely asserts that an additive bias term equal to the bound (32) enters (36)--(37), and Remark 5.7 contains only the informal perturbation bound (32). This matters because every experiment uses dictionaries that violate Assumption 1: Section 11.5 reports relative invariance residuals of 0.15--0.64, so the empirical interpretation of Theorem 6.1 rests directly on this unproved bias extension. The authors should either state and prove the approximate-invariance theorem with explicit hypotheses, or restrict the experimental claims to the exactly invariant idealisation.
minor comments (5)
  1. [Section 3.2 and 3.3] The text says a primitive 'satisfies Theorem 3.2', but the object labelled 3.2 is a definition, not a theorem; this should be corrected throughout Section 3.
  2. [Section 4.5 and Definition 4.5] Assumption 1 uses 'nu-almost every u', while Definition 4.5 requires invariance 'for every u in U'. These are different conditions and should be reconciled.
  3. [Section 5.4 and Remarks 5.6--5.8] Several claims are cited as theorems that are actually remarks: the Jordan-form discussion appears as Remark 5.6 but is called 'Theorem 5.6' in Section 5.4; the approximate-invariance discussion is Remark 5.7 but is called 'Theorem 5.7' in Remark 4.11; and the EDMD comparison is Remark 5.8 but is called 'Theorem 5.8' in Section 2. The numbering should be made consistent.
  4. [Figure 11 caption] The caption contains a duplicated and garbled phrase: 'for ReLU it is not below for TopK; for ReLU it is not below baseline at the guard-selected operating point, which the text takes up. at the guard-selected operating point, which the text takes up.' This should be rewritten.
  5. [Section 6.3 and 6.4] The proof of Lemma 6.3 refers to 'Theorem 6.3' and 'Theorem 6.4' when it means the lemma itself; the cross-referencing between lemmas and theorems in Section 6 should be audited.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the identifiability theorem is derived forward from concentration and perturbation bounds, and empirical predictions are not fitted from the quantities they predict.

full rationale

The load-bearing derivation is a forward proof: Theorem 6.1 combines sub-Gaussian Gramian concentration, least-squares stability, and Bauer-Fike/Davis-Kahan perturbation, none of which assumes the target spectrum. The estimand A is defined independently in Theorem 5.1, and the M^{-1/2} rate is a consequence rather than a fitted parameter. The empirical convergence exponent on Qwen3-8B-Base is estimated from disjoint split-half fits and compared with the theoretical -1/2; no constant is fit to force agreement. The norm-growth mode in Section 11.8.3 is checked against independently measured per-layer growth, and the median-of-means comparison was constructed so the robust estimator could disagree trial-for-trial. The invariance-penalty experiment does include a near-tautological component, since the penalty is the residual it minimizes, but the claimed success is the split-half spectral distance, a genuinely separate outcome, and the paper reports the costs and failures of the sweep. There is no load-bearing self-citation; the cited tools (Korda-Mezić, Bauer-Fike, Davis-Kahan, Vershynin, etc.) are independent established results. The most serious substantive concern is the gap between Assumption 1 and the affine-in-control equation (9): exact span K-invariance yields per-control matrices A_u but does not by itself prove the existence of a fixed pair (A,B) with zero residual, as the skeptic's example F(x,u)=x+u with dictionary (1,x,x^2) shows. That is a correctness or assumption gap, not a circular reduction, because the theorem does not assume its own conclusion. The paper also reports failing controls (Section 11.7 separates seed replicas; Section 11.5 meets the pre-registered criterion in only 49% of cells) and lists these limitations in Section 12; they bound the claims but do not make the derivation circular. I therefore find no circularity and score 0.

Assumptions & free parameters 3 free parameters · 8 assumptions · 2 invented entities

The paper is honest about many limitations, but the central claim depends on exact K-invariance and on an unstated cross-observable moment condition in Lemma 6.3. The listed free parameters are measurement-protocol choices rather than fitted constants in the derivation. Invented entities are limited to the Koopman spectral object and its circuit interpretation, both built from standard operator theory rather than new physical postulates.

free parameters (3)
  • Random-Fourier bandwidth
    Hand-set in Section 11.1 at a value appropriate for d<=2304 and deliberately not retuned for d=4096; controls dictionary quality, K-invariance residual, and stability in deep layers.
  • Control projection dimension p = 64
    Section 11.1 projects attention writes onto 64 leading principal directions to avoid interpolation at p=d; this defines the input space of the realisation and therefore the estimand.
  • Dictionary size N = 32 (headline; swept 8 to 128)
    Central calibration parameter: thresholds scale linearly in N, the gap-free bound carries a (2N-1) factor, and the experiments sweep N to compare convergence exponents.
assumptions (8)
  • domain assumption Exact K-invariance of the dictionary span (Assumption 1 / DR2)
    Required for the finite realisation to exist and for A to be a Koopman compression; Section 12 states real dictionaries only approximate it and the bias does not vanish with M.
  • domain assumption Persistent excitation of controls (Assumption 2)
    Ensures the empirical control Gramian is invertible and B is identifiable; requires sufficiently diverse attention patterns in the calibration corpus.
  • domain assumption Spectral separation of eigenvalues (Assumption 3)
    Needed for eigenvector identifiability and labelled eigenvalue matching; Theorem 6.8 removes it only for the unlabelled eigenvalue guarantee, and measured gaps are about 1e-3.
  • domain assumption State-control decorrelation / exogeneity of attention writes (R2, Eq. 23)
    The control u_l is computed from the same residual stream; canonical correlations reach 0.96 and residualising the control shifts the spectrum by 5.2-9.1 times the resolution floor (Section 11.8.2).
  • domain assumption Sub-Gaussian dictionary and controls (R3)
    Used for Gramian concentration in Theorem 6.3; the lifted state is measured light-tailed except at Qwen layer 18.
  • ad hoc to paper Sub-exponentiality of the cross-observable Psi(F(x,u)) in Lemma 6.3
    The proof invokes matrix Bernstein without stating a moment or boundedness hypothesis on Psi(F(x,u)); this hidden assumption is needed for the cross-Gramian concentration and is not derived from the stated (R3) conditions.
  • standard math Existence of limiting invariant measure mu and stationary control distribution nu
    Standard Koopman and EDMD background; the paper notes that in practice mu is replaced by the layer marginal.
  • domain assumption Diagonalisability of A with bounded eigenvector condition number (R4)
    Used for Bauer-Fike and Davis-Kahan bounds; Remark 5.6 sketches Jordan extensions but does not prove the main rate for them.
invented entities (2)
  • Koopman spectral fingerprint (spectrum of the finite Koopman realisation A) independent evidence
    purpose: Coordinate-free, identifiable invariant of transformer depth dynamics; replaces explicit circuit dictionaries as the certified object.
    The paper attaches falsifiable handles to it: predicted M^{-1/2} convergence rate, recovery of norm growth as a real mode, and the modal-principal dissociation; these measurements could have failed.
  • Koopman circuit (lambda, phi, v, Pi)
    purpose: Interpretive unit for spectral modes of the realisation; supports the intervention calculus on spectral projectors.
    Section 11.3 shows the selected top-16 mode subspace has median agreement 0.290 across calibration draws, so the circuit-level entity is defined but not independently supported as a reproducible circuit.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability." pith.science (2026). https://pith.science/paper/KLKU5Q5K

@misc{pith2026260810172,
  author       = {Pith},
  title        = {Pith review of: Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KLKU5Q5K}},
  note         = {Machine review of arXiv:2608.10172}
}
abstract

Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method that found it. Sparse autoencoders illustrate the problem: different seeds and widths recover materially different features from the same activations, and no theory says whether that variability is incidental or structural. We put dictionary learning for interpretability on an identifiability footing. Treating the forward pass as a controlled dynamical system with depth as time and lifting it with the Koopman operator yields a finite linear realisation whose \emph{spectrum} is a coordinate-free property of the model. We prove the spectrum is recoverable from $M$ calibration samples at rate $M^{-1/2}$ up to permutation - to our knowledge the first identifiability theorem for a mechanistic-interpretability primitive, with a matching minimax lower bound, a median-of-means variant for heavy-tailed activations, and a dissociation theorem: whenever the realisation is non-normal, the directions carrying activation variance and the directions carrying information across depth cannot coincide. The identifiable object and the legible object are not the same object. On GPT-2 small, Gemma-2-2B and Qwen3-8B-Base the spectrum converges everywhere and attains the predicted exponent on Qwen3-8B-Base ($0.506 \pm 0.031$); shortfalls collapse onto one curve against each cell's sample threshold. Koopman modes beat random directions but lose to principal components on indirect-object identification, with the gap decaying $4.1\times$ in depth-distance, as the theorem predicts. The Koopman spectrum is an identifiable, model-intrinsic fingerprint with a stated error bar, not a legible decomposition.

Figures

Figures reproduced from arXiv: 2608.10172 by the authors.

Figure 1
Figure 1. Transformer forward pass as a controlled dynamical system (bottom) and its Koopman lift (top). With [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Analytic illustration: Algebraic completeness of the KSA intervention calculus (Theorem [PITH_FULL_IMAGE:figures/full_fig_p019_2.png] view at source ↗
Figure 3
Figure 3. Split-half spectral distance against M, log-log, at N = 32 and mid-depth. The distance falls monotonically on every model: the spectrum is being recovered. Qwen3-8B-Base attains the predicted M−1/2 rate (−0.506 ± 0.031); GPT-2 small is shallower (−0.286). The dashed line is the M−1/2 reference. Shaded bands are sequence-level block bootstraps. Theorem 7.2 though modestly. On synthetic data with genuinely heavy desig… view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: The robust estimator changes almost nothing, and the tail diagnostics say why. (a) Plain against median-of-means EDMDc on identical sequence-disjoint halves, GPT-2 small layer 6: the two curves coincide at every M. (b) Their ratio, 1.000 at the median and within [0.979…
Figure 5
Figure 5. Figure 5: Most of the scatter in the convergence exponent is one curve. (a) Each of 30 (model, dictionary, N) cells contributes rolling exponents, plotted against log10(M/Meig 0 (N)) - its own threshold rather than raw M. Points are individual windows; the heavy line joins equal…
Figure 6
Figure 6. Figure 6: Diagnosing the exponent, all four panels. (a) Plain against median-of-means EDMDc on identical sequence-disjoint halves, GPT-2 small layer 6: the two coincide at every M (median ratio 1.000), as they do in seven of the 8 cells measured. (b) Why: the Hill tail index alo…
Figure 7
Figure 7. Figure 7: Koopman modes carry behaviourally relevant structure but do not resolve the IOI circuit. (a) Ablating the top-j modes at layer 9 moves the IOI logit difference more than j random directions but less than j principal directions: at j = 8, PCA removes 60% against the mod…
Figure 8
Figure 8. Figure 8: The PCA advantage is local and decays with depth. Predicting xℓ+k from a rank-j read of xℓ (6 layers, j ∈ {8, . . . , 64}, 8,192 prompts), normalised between the optimal rank-j predictor (0) and a random subspace (1). (a) Normalised excess fraction of variance unexplai…
Figure 9
Figure 9. Figure 9: The (DR2) gap and the non-identifiability it produces. (a) Distance from Koopman invariance, εˆrel of (77), at matched N = 32 with all dictionaries whitened, at three layers per model. SAE dictionaries sit furthest from (DR2) on every model and layer - further than a r…
Figure 10
Figure 10. Figure 10: The whole effect is conditional on the selection rule. Ratio of SAE to spectral invariance residual as a function [PITH_FULL_IMAGE:figures/full_fig_p032_10.png]
Figure 11
Figure 11. Figure 11: Adding the invariance penalty improves spectral identifiability. [PITH_FULL_IMAGE:figures/full_fig_p033_11.png]
Figure 12
Figure 12. Figure 12: The universality criterion separates everything, including the control it should not. Matched spectral distance [PITH_FULL_IMAGE:figures/full_fig_p034_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

79 extracted references · 60 canonical work pages

  1. [1]

    Sparse autoen- coders find highly interpretable features in language models

    Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, and Lee Sharkey. Sparse autoen- coders find highly interpretable features in language models. InThe Twelfth International Conference on Learning Representations, 2024

  2. [2]

    Improving sparse de- composition of language model activations with gated sparse autoencoders

    Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, Janos Kramar, Rohin Shah, and Neel Nanda. Improving sparse de- composition of language model activations with gated sparse autoencoders. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

  3. [3]

    Jumping ahead: Improving reconstruction fidelity with JumpReLU sparse autoen- coders.arXiv preprint arXiv:2407.14435, 2024

    Senthooran Rajamanoharan, Tom Lieberum, Nico- las Sonnerat, Arthur Conmy, Vikrant Varma, János Kramár, and Neel Nanda. Jumping ahead: Improving reconstruction fidelity with JumpReLU sparse autoen- coders.arXiv preprint arXiv:2407.14435, 2024

  4. [4]

    Scaling and evaluating sparse autoencoders

    Leo Gao, Tom Dupre la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. Scaling and evaluating sparse autoencoders. InInternational Conference on Learn- ing Representations, volume 2025, pages 26721–26754, 2025

  5. [5]

    Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2

    Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, Janos Kramar, Anca Dragan, Rohin Shah, and Neel Nanda. Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2. In Yonatan Be- linkov, Najoung Kim, Jaap Jumelet, Hosein Mohebbi, Aaron Mueller, and Hanjie Chen, editors,Proceedings of the 7th ...

  6. [6]

    Llama scope: Extracting millions of features from llama-3.1-8b with sparse autoencoders.arXiv preprint arXiv:2410.20526, 2024

    Zhengfu He, Wentao Shu, Xuyang Ge, Lingjie Chen, Junxuan Wang, Yunhua Zhou, Frances Liu, Qipeng Guo, Xuanjing Huang, Zuxuan Wu, et al. Llama scope: Extracting millions of features from llama-3.1-8b with sparse autoencoders.arXiv preprint arXiv:2410.20526, 2024

  7. [7]

    Transcoders find interpretable LLM feature circuits

    Jacob Dunefsky, Philippe Chlenski, and Neel Nanda. Transcoders find interpretable LLM feature circuits. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

  8. [8]

    Emmanuel Ameisen, Jack Lindsey, Adam Pearce, Wes Gurnee, Nicholas L. Turner, Brian Chen, Craig Citro, David Abrahams, Shan Carter, Basil Hosmer, Jonathan Marcus, Michael Sklar, Adly Templeton, Trenton Bricken, Callum McDougall, Hoagy Cunning- ham, Thomas Henighan, Adam Jermyn, Andy Jones, Andrew Persic, Zhenyi Qi, T. Ben Thompson, Sam Zimmerman, Kelley Ri...

Show all 79 references
  1. [9]

    Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, David Abrahams, Shan Carter, Basil Hosmer, Jonathan Marcus, Michael Sklar, Adly Templeton, Trenton Bricken, Callum McDougall, Hoagy Cunning- ham, Thomas Henighan, Adam Jermyn,...

  2. [10]

    Causal mediation analysis for interpreting neural NLP: The case of gender bias

    Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. Causal mediation analysis for interpreting neural NLP: The case of gender bias. InAdvances in Neural Information Processing Systems (NeurIPS), 2020

  3. [11]

    Locating and editing factual as- sociations in gpt

    Kevin Meng, David Bau, Alex J Andonian, and Yonatan Belinkov. Locating and editing factual as- sociations in gpt. InAdvances in neural information processing systems, 2022

  4. [12]

    Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso

    Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso. Towards automated circuit discovery for mechanistic interpretability. InAdvances in Neural Information Processing Systems (NeurIPS), 2023

  5. [13]

    Attribu- tion patching outperforms automated circuit discovery

    Aaquib Syed, Can Rager, and Arthur Conmy. Attribu- tion patching outperforms automated circuit discovery. In Yonatan Belinkov, Najoung Kim, Jaap Jumelet, Hosein Mohebbi, Aaron Mueller, and Hanjie Chen, ed- itors,Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interp...

  6. [14]

    Have faith in faithfulness: Going beyond cir- cuit overlap when finding model mechanisms.arXiv preprint arXiv:2403.17806, 2024

    Michael Hanna, Sandro Pezzelle, and Yonatan Be- linkov. Have faith in faithfulness: Going beyond cir- cuit overlap when finding model mechanisms.arXiv preprint arXiv:2403.17806, 2024

  7. [15]

    AtP⋆: An efficient and scalable method for lo- calizing LLM behaviour to components.arXiv preprint arXiv:2403.00745, 2024

    János Kramár, Tom Lieberum, Rohin Shah, and Neel Nanda. AtP⋆: An efficient and scalable method for lo- calizing LLM behaviour to components.arXiv preprint arXiv:2403.00745, 2024

  8. [16]

    Causal abstractions of neural networks

    AtticusGeiger, HansonLu, ThomasIcard, andChristo- pher Potts. Causal abstractions of neural networks. Advances in neural information processing systems, 34:9574–9586, 2021

  9. [17]

    Finding align- ments between interpretable causal variables and dis- tributed neural representations

    Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah Goodman. Finding align- ments between interpretable causal variables and dis- tributed neural representations. InCausal Learning and Reasoning, pages 160–187. PMLR, 2024

  10. [18]

    Interpretability at scale: Identifying causal mechanisms in alpaca

    ZhengxuanWu, AtticusGeiger, ThomasIcard, Christo- pher Potts, and Noah Goodman. Interpretability at scale: Identifying causal mechanisms in alpaca. In Thirty-seventh Conference on Neural Information Pro- cessing Systems, 2023

  11. [19]

    Safety cases: How to justify the safety of advanced ai systems.arXiv preprint arXiv:2403.10462, 2024

    Joshua Clymer, Nick Gabrieli, David Krueger, and Thomas Larsen. Safety cases: How to justify the safety of advanced ai systems.arXiv preprint arXiv:2403.10462, 2024

  12. [20]

    Identifying functionally important features with end-to-end sparse dictionary learning

    Dan Braun, Jordan Taylor, Nicholas Goldowsky-Dill, and Lee Sharkey. Identifying functionally important features with end-to-end sparse dictionary learning. Advances in Neural Information Processing Systems, 37:107286–107325, 2024

  13. [21]

    Transcoders beat sparse autoencoders for interpretability.arXiv preprint arXiv:2501.18823, 2025

    Gonçalo Paulo, Nora Belrose, et al. Transcoders beat sparse autoencoders for interpretability.arXiv preprint arXiv:2501.18823, 2025

  14. [22]

    SAEBench: A compre- hensive benchmark for sparse autoencoders in language model interpretability

    Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Isaac Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Callum Stuart McDougall, Kola Ayon- rinde, Demian Till, Matthew Wearden, Arthur Conmy, Samuel Marks, and Neel Nanda. SAEBench: A compre- hensive benchmark for spars...

  15. [23]

    A is for absorption: Studying feature split- ting and absorption in sparse autoencoders.Advances in Neural Information Processing Systems, 38:82318– 82355, 2026

    David Chanin, James Wilken-Smith, Tomáš Dulka, Hardik Bhatnagar, Satvik Golechha, and Joseph Bloom. A is for absorption: Studying feature split- ting and absorption in sparse autoencoders.Advances in Neural Information Processing Systems, 38:82318– 82355, 2026

  16. [24]

    Decomposing the dark matter of sparse autoencoders

    Joshua Engels, Logan Riggs Smith, and Max Tegmark. Decomposing the dark matter of sparse autoencoders. Transactions on Machine Learning Research, 2024

  17. [25]

    A proposal on machine learning via dynamical systems

    E Weinan, Yali Duan, Linghua Kong, and Min Guo. A proposal on machine learning via dynamical systems. Links, 2024:08–27, 2017

  18. [26]

    Stable architec- tures for deep neural networks.Inverse problems, 34(1):014004, 2018

    Eldad Haber and Lars Ruthotto. Stable architec- tures for deep neural networks.Inverse problems, 34(1):014004, 2018

  19. [27]

    Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David K. Duvenaud. Neural ordinary differen- tial equations. InAdvances in Neural Information Processing Systems (NeurIPS), 2018

  20. [28]

    Hamiltonian systems and trans- formation in hilbert space.Proceedings of the National Academy of Sciences, 17(5):315–318, 1931

    Bernard O Koopman. Hamiltonian systems and trans- formation in hilbert space.Proceedings of the National Academy of Sciences, 17(5):315–318, 1931

  21. [29]

    Spectral properties of dynamical systems, model reduction and decompositions.Nonlinear Dy- namics, 41(1):309–325, 2005

    Igor Mezić. Spectral properties of dynamical systems, model reduction and decompositions.Nonlinear Dy- namics, 41(1):309–325, 2005

  22. [30]

    Modern koopman theory for dynamical systems.arXiv preprint arXiv:2102.12086, 2021

    Steven L Brunton, Marko Budišić, Eurika Kaiser, and J Nathan Kutz. Modern koopman theory for dynamical systems.arXiv preprint arXiv:2102.12086, 2021

  23. [31]

    A mathematical framework for transformer circuits.Transformer Circuits Thread, 2021

    Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, NicholasJoseph, BenMann, AmandaAskell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Das- Sarma, Dawn Drain, Deep Ganguli, Zac Hatfield- Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dari...

  24. [32]

    Towards monosemanticity: Decomposing language models with dictionary learning.Transformer Circuits Thread, 2023

    Trenton Bricken, Adly Templeton, Joshua Bat- son, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguye...

  25. [33]

    Daniel Freeman, Theodore R

    Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua ...

  26. [34]

    Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt

    Kevin R. Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. Interpretability in the wild: A circuit for indirect object identifica- tion in GPT-2 Small. InInternational Conference on Learning Representations (ICLR), 2023

  27. [35]

    Independent component analysis, a new concept?Signal processing, 36(3):287–314, 1994

    Pierre Comon. Independent component analysis, a new concept?Signal processing, 36(3):287–314, 1994

  28. [36]

    John Wiley & Sons, 2001

    Aapo Hyvärinen, Juha Karhunen, and Erkki Oja.In- dependent Component Analysis. John Wiley & Sons, 2001

  29. [37]

    Unsupervised feature extraction by time-contrastive learning and nonlinear ica.Advances in neural information process- ing systems, 29, 2016

    Aapo Hyvarinen and Hiroshi Morioka. Unsupervised feature extraction by time-contrastive learning and nonlinear ica.Advances in neural information process- ing systems, 29, 2016

  30. [38]

    Variational autoencoders and nonlinear ica: A unifying framework

    Ilyes Khemakhem, Diederik Kingma, Ricardo Monti, and Aapo Hyvarinen. Variational autoencoders and nonlinear ica: A unifying framework. InInternational conference on artificial intelligence and statistics, pages 2207–2217. PMLR, 2020

  31. [39]

    Challenging common assumptions in the unsupervised learning of disentangled representa- tions

    Francesco Locatello, Stefan Bauer, Mario Lucic, Gun- nar Raetsch, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem. Challenging common assumptions in the unsupervised learning of disentangled representa- tions. Ininternational conference on machine learning, pages 4114–41...

  32. [40]

    Toward causal representation learning.Proceedings of the IEEE, 109(5):612–634, 2021

    BernhardSchölkopf, FrancescoLocatello, StefanBauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, and Yoshua Bengio. Toward causal representation learning.Proceedings of the IEEE, 109(5):612–634, 2021

  33. [41]

    System identification: Theory for the user, (ljung, l.; 1999)[on the shelf].IEEE Robotics & Automation Magazine, 19(2):95–96, 2012

    Alex Simpkins. System identification: Theory for the user, (ljung, l.; 1999)[on the shelf].IEEE Robotics & Automation Magazine, 19(2):95–96, 2012

  34. [42]

    A note on persistency of excitation

    Jan C Willems, Paolo Rapisarda, Ivan Markovsky, and Bart LM De Moor. A note on persistency of excitation. Systems & Control Letters, 54(4):325–329, 2005

  35. [43]

    Peter J. Schmid. Dynamic mode decomposition of numerical and experimental data.Journal of Fluid Mechanics, 656:5–28, 2010

  36. [44]

    A data-driven approximation of the koopman operator: Extending dynamic mode decomposition

    Matthew Williams, Ioannis Kevrekidis, and Clarence Rowley. A data-driven approximation of the koopman operator: Extending dynamic mode decomposition. Journal of nonlinear science, 25(6), 2015

  37. [45]

    Dynamic mode decomposition with con- trol.SIAM Journal on Applied Dynamical Systems, 15(1):142–161, 2016

    Joshua L Proctor, Steven L Brunton, and J Nathan Kutz. Dynamic mode decomposition with con- trol.SIAM Journal on Applied Dynamical Systems, 15(1):142–161, 2016

  38. [46]

    Williams, Clarence W

    Matthew O. Williams, Clarence W. Rowley, and Ioan- nis G. Kevrekidis. A kernel-based method for data- driven Koopman spectral analysis.Journal of Compu- tational Dynamics, 2(2):247–265, 2015

  39. [47]

    Data-driven approximation of the Koopman generator: Model reduction, system identification, and control.Physica D: Nonlinear Phenomena, 406:132416, 2020

    Stefan Klus, Feliks Nüske, Sebastian Peitz, Jan- Hendrik Niemann, Cecilia Clementi, and Christof Schütte. Data-driven approximation of the Koopman generator: Model reduction, system identification, and control.Physica D: Nonlinear Phenomena, 406:132416, 2020

  40. [48]

    On convergence of ex- tended dynamic mode decomposition to the Koopman operator.Journal of Nonlinear Science, 28(2):687–710, 2018

    Milan Korda and Igor Mezić. On convergence of ex- tended dynamic mode decomposition to the Koopman operator.Journal of Nonlinear Science, 28(2):687–710, 2018

  41. [49]

    Redman, Maria Fonoberova, Ryan Mohr, Ioannis G

    William T. Redman, Maria Fonoberova, Ryan Mohr, Ioannis G. Kevrekidis, and Igor Mezić. An operator theoretic view on pruning deep neural networks. In International Conference on Learning Representations (ICLR), 2022

  42. [50]

    Benjamin Erichson, Vanessa Lin, and Michael W

    Omri Azencot, N. Benjamin Erichson, Vanessa Lin, and Michael W. Mahoney. Forecasting sequential data using consistent Koopman autoencoders. InInterna- tional Conference on Machine Learning (ICML), 2020

  43. [51]

    Thiem, and Ioannis G

    Felix Dietrich, Thomas N. Thiem, and Ioannis G. Kevrekidis. On the Koopman operator of algo- rithms.SIAM Journal on Applied Dynamical Systems, 19(2):860–885, 2020

  44. [52]

    Bruce C. Moore. Principal component analysis in lin- ear systems: Controllability, observability, and model reduction.IEEE Transactions on Automatic Control, 26(1):17–32, 1981

  45. [53]

    Antoulas.Approximation of Large-Scale Dynamical Systems

    Athanasios C. Antoulas.Approximation of Large-Scale Dynamical Systems. SIAM, 2005

  46. [54]

    All optimal Hankel-norm approxima- tions of linear multivariable systems and theirl∞-error bounds.International Journal of Control, 39(6):1115– 1193, 1984

    Keith Glover. All optimal Hankel-norm approxima- tions of linear multivariable systems and theirl∞-error bounds.International Journal of Control, 39(6):1115– 1193, 1984

  47. [55]

    Al-Saggaf and Gene F

    Ubaid M. Al-Saggaf and Gene F. Franklin. Model reduction via balanced realizations: An extension and frequency weighting techniques.IEEE Transactions on Automatic Control, 33(7):687–692, 1988

  48. [56]

    Antoulas

    Serkan Gugercin and Athanasios C. Antoulas. A sur- vey of model reduction by balanced truncation and some new results.International Journal of Control, 77(8):748–766, 2004

  49. [57]

    Dale F. Enns. Model reduction with balanced real- izations: An error bound and a frequency weighted generalization.Proceedings of the 23rd IEEE Confer- ence on Decision and Control, pages 127–132, 1984

  50. [58]

    Attention is not all you need: Pure attention loses rank doubly exponentially with depth

    Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. Attention is not all you need: Pure attention loses rank doubly exponentially with depth. InIn- ternational Conference on Machine Learning (ICML), 2021

  51. [59]

    The emergence of clusters in self- attentiondynamics

    Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet. The emergence of clusters in self- attentiondynamics. InAdvances in Neural Information Processing Systems (NeurIPS), 2023

  52. [60]

    In-context learning and induction heads.Transformer Circuits Thread, 2022

    Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield- Dodds, DannyHernandez, ScottJohnston, AndyJones, Jackson Kernion, Liane Lovitt, Kamal...

  53. [61]

    Zoom in: An introduction to circuits.Distill, 2020

    Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. Zoom in: An introduction to circuits.Distill, 2020

  54. [62]

    Progress measures for grokking via mechanistic interpretability

    Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. Progress measures for grokking via mechanistic interpretability. InInterna- tional Conference on Learning Representations (ICLR), 2023

  55. [63]

    The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

  56. [64]

    Gemma 2: Improv- ing open language models at a practical size.arXiv preprint arXiv:2408.00118, 2024

    Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupati- raju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. Gemma 2: Improv- ing open language models at a practical size.arXiv preprint arXiv:2408.00118, 2024

  57. [65]

    Gomez, Łukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems (NeurIPS), 2017

  58. [66]

    Springer, 2nd, classics in mathematics reprint edition, 1995

    Tosio Kato.Perturbation Theory for Linear Operators. Springer, 2nd, classics in mathematics reprint edition, 1995

  59. [67]

    Cam- bridge Series in Statistical and Probabilistic Mathe- matics

    Roman Vershynin.High-Dimensional Probability: An Introduction with Applications in Data Science. Cam- bridge Series in Statistical and Probabilistic Mathe- matics. Cambridge University Press, 2018

  60. [68]

    Stewart and Ji-Guang Sun.Matrix Pertur- bation Theory

    Gilbert W. Stewart and Ji-Guang Sun.Matrix Pertur- bation Theory. Academic Press, 1990

  61. [69]

    Joel A. Tropp. An introduction to matrix concentra- tion inequalities.Foundations and Trends in Machine Learning, 8(1–2):1–230, 2015

  62. [70]

    Bauer and Charles T

    Friedrich L. Bauer and Charles T. Fike. Norms and exclusion theorems.Numerische Mathematik, 2(1):137– 141, 1960

  63. [71]

    Chandler Davis and William M. Kahan. The rotation of eigenvectors by a perturbation. III.SIAM Journal on Numerical Analysis, 7(1):1–46, 1970

  64. [72]

    An optimal bound for the spectral variation of two matrices.Linear algebra and its ap- plications, 71:77–80, 1985

    Ludwig Elsner. An optimal bound for the spectral variation of two matrices.Linear algebra and its ap- plications, 71:77–80, 1985

  65. [73]

    Cambridge university press, 2012

    Roger A Horn and Charles R Johnson.Matrix analysis. Cambridge university press, 2012

  66. [74]

    Convergence of estimates under dimen- sionality restrictions.The Annals of Statistics, pages 38–53, 1973

    Lucien LeCam. Convergence of estimates under dimen- sionality restrictions.The Annals of Statistics, pages 38–53, 1973

  67. [75]

    Springer Science & Business Media, 2012

    Lucien Le Cam.Asymptotic methods in statistical decision theory. Springer Science & Business Media, 2012

  68. [76]

    Springer Series in Statistics

    AlexandreB.Tsybakov.Introduction to Nonparametric Estimation. Springer Series in Statistics. Springer, New York, 2009

  69. [77]

    LLM.int8(): 8-bit matrix multiplication for transformers at scale

    Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. LLM.int8(): 8-bit matrix multiplication for transformers at scale. InAdvances in Neural In- formation Processing Systems, volume 35, 2022

  70. [78]

    Geometric median and robust esti- mation in Banach spaces.Bernoulli, 21(4):2308–2335, 2015

    Stanislav Minsker. Geometric median and robust esti- mation in Banach spaces.Bernoulli, 21(4):2308–2335, 2015

  71. [79]

    Mean estimation and regression under heavy-tailed distributions: A survey.Foundations of Computational Mathematics, 19(5):1145–1190, 2019

    Gábor Lugosi and Shahar Mendelson. Mean estimation and regression under heavy-tailed distributions: A survey.Foundations of Computational Mathematics, 19(5):1145–1190, 2019

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.