{"id":"79bd539d-b849-48c2-b4a0-cccddff2758d","arxiv_id":"2608.10172","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A transformer's Koopman spectrum is provably recoverable from M calibration samples at the M^{-1/2} rate up to permutation, making it the first identified mechanistic-interpretability primitive.","lead":"This paper gives transformers a stable mathematical fingerprint by viewing a network's layers as steps of a dynamical system and reading off the spectrum of the associated Koopman operator. It proves the fingerprint is recoverable from finite data at a statistical rate, and shows on GPT-2, Gemma-2-2B, and Qwen3-8B that the identifiable object is not the same as a human-legible circuit decomposition.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 1 does not imply the affine-in-control equation (9) that Definition 4.6 and Theorem 6.1 rely on; exactly K-invariant dictionaries can still have a non-vanishing EDMDc residual.","rationale":"The reader's weakest assumption concerns real dictionaries failing exact K-invariance, with a non-vanishing bias. My concern is stronger and formal: even under the theorem's own idealised hypothesis, the claimed estimand may not be well defined in the sense used by the proof. The exact affine-in-control equation (9) is not a consequence of Assumption 1, so the proof of Theorem 6.1 silently assumes a stronger dictionary closure property. The synthetic polynomial example is decisive, cheap, and does not depend on any empirical question; if it confirms the gap, the theorem can be repaired by adding the affine closure condition, with Theorem 6.12's bias analysis then covering the residual. This does not invalidate the experimental programme or the dissociation results, but it means the central identifiability theorem, as currently stated, is not proved. The reader's conditional verdict remains appropriate, now with an additional formal condition to address.","tokens_in":47784,"tokens_out":18455,"duration_ms":202788,"concrete_test":"Run the theorem's setting on a minimal synthetic case: F(x,u) = x + u with scalar x and u, dictionary Ψ = (1, x, x^2), and u drawn from, say, N(0,1). This dictionary satisfies Assumption 1 exactly. Compute the population minimiser (A,B) of equation (10) and the residual ε(x,u) = Ψ(F(x,u)) - AΨ(x) - Bu. If ∥ε∥ > 0, as the u^2 term forces, and the spectrum of the EDMDc A-block changes when the control distribution ν is changed while F and Ψ are fixed, then Definition 4.6's (A,B) is not a consequence of K-invariance. The single experiment decides whether Theorem 6.1 must be restated under an explicit control-affine closure assumption.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central theorem assumes exact span K-invariance (Assumption 1) and concludes that EDMDc recovers the spectrum of the matrix A in Definition 4.6, where the pair (A,B) satisfies Ψ(F(x,u)) = AΨ(x) + Bu. But Assumption 1 only says that for each control u, x ↦ Ψ(F(x,u)) lies in the dictionary span. This yields a coefficient matrix A_u for each u, yet nothing forces A_u to be affine in u. The proof in Section 6.4 uses Y = AX + BU as an exact identity, and Definition 4.6 asserts existence of such (A,B) without proof. A minimal counterexample: F(x,u) = x + u and Ψ = (1, x, x^2). The span is exactly K-invariant, but A_u has entries that are linear in u for x and quadratic in u for x^2, so no constant pair (A,B) can represent the u^2 term. Thus even an exactly K-invariant dictionary can have a non-zero least-squares residual, and the estimand whose spectrum is certified is a control-distribution-dependent projection. This is a gap between Assumption 1 and the object of Theorem 6.1, not merely the acknowledged approximate-invariance limitation of real dictionaries. The theorem needs an explicit control-affine closure condition, or a proof that K-invariance plus affine F implies A_u is affine, which the example shows to be false.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes treating a transformer forward pass as a controlled dynamical system in depth, lifting it with the Koopman operator, and defining the spectrum of the resulting finite linear realisation as a coordinate-free ``intrinsic'' object for mechanistic interpretability. The central theoretical claim is Theorem 6.1: under exact dictionary K-invariance, persistent excitation, spectral separation, sub-Gaussian tails and diagonalisability, the EDMDc estimator recovers the eigenvalues of the Koopman realisation A from M calibration samples at rate M^{-1/2}, up to permutation, with a matching minimax lower bound and a median-of-means heavy-tailed variant. The paper also proves a dissociation theorem (Theorem 8.1) showing that non-normality forces the variance-ordered principal basis and the Koopman modal basis apart, and derives consequences for SAE non-identifiability, intervention calculus, and model reduction. Experiments on GPT-2 small, Gemma-2-2B and Qwen3-8B-Base report convergence of the spectrum, observation of the predicted M^{-1/2} rate on Qwen3-8B-Base, and several explicitly reported failures including the universality criterion and the SAE invariance-gap criterion under variance-based feature selection.","tokens_in":48094,"tokens_out":13553,"duration_ms":137352,"significance":"If the central derivation can be completed, the paper would provide the first identifiability theorem for a mechanistic-interpretability primitive, with explicit rates, a matching lower bound, and a clear statement of what is and is not identified. The paper is also exemplary in reporting negative results and confounds: the universality criterion fails its seed-replica control, the SAE gap reverses under variance-based selection, and the control-exogeneity convention moves the spectrum by 5--9 times the split-half floor. The norm-growth eigenvalue prediction and the 4.1x depth-decay of the PCA advantage are genuinely falsifiable measurements. These strengths make the manuscript potentially important, but the current proof has load-bearing gaps that must be repaired before the claims can be accepted.","major_comments":[{"comment":"Assumption 1 does not imply the affine-in-control identity (9). For F(x,u)=x+u and Psi=(1,x,x^2), the span is exactly K-invariant for every u, but Psi(F(x,u))=(1, x+u, x^2+2ux+u^2), and the coefficient of x^2 is quadratic in u, so no constant pair (A,B) can satisfy (9). Theorem 5.1 proves only the averaged identity E_u[Psi(F(x,u))]=A Psi(x) in Eq. (24). Consequently the object whose spectrum Theorem 6.1 certifies is not shown to be the Koopman realisation of Definition 4.6; it is at best the population least-squares matrix of Eq. (10). The paper should either impose an explicit control-affine closure condition, or redefine the estimand as the least-squares pair (10) and prove Theorem 6.1 for that pair. As written, Section 6.4's decomposition of the EDMDc estimator relies on an exact linear relation that the hypotheses do not provide.","section":"Definition 4.6 and Section 6.4"},{"comment":"The concentration argument for the cross-Gramian \\hat C_M = (1/M) Y X^* is asserted for the 'sub-exponential random matrix' Psi(F(x,u)) Psi(x)^*, but condition (R3) bounds only Psi(x) and u. Since F includes the MLP with RMSNorm, sub-Gaussianity of x and u does not imply any tail bound on Psi(F(x,u)). The same gap appears in Theorem 7.2, where (R3') supplies moments only for Psi(x) and u. An explicit hypothesis on the lifted next state---for example sub-Gaussianity of Psi(F(x,u)) under mu \\otimes nu, or a uniform bound on the pointwise Koopman matrices A_u of Lemma 5.2---must be stated and propagated through Theorems 6.1 and 7.2.","section":"Lemma 6.3 and Theorem 7.2"},{"comment":"The approximate-invariance extension is repeatedly cited as 'Theorem 6.12' and 'Theorem 5.7', but no such numbered theorems are stated or proved anywhere in the manuscript. Remark 6.12 merely asserts that an additive bias term equal to the bound (32) enters (36)--(37), and Remark 5.7 contains only the informal perturbation bound (32). This matters because every experiment uses dictionaries that violate Assumption 1: Section 11.5 reports relative invariance residuals of 0.15--0.64, so the empirical interpretation of Theorem 6.1 rests directly on this unproved bias extension. The authors should either state and prove the approximate-invariance theorem with explicit hypotheses, or restrict the experimental claims to the exactly invariant idealisation.","section":"Sections 6.8, 11.7, and 12 (cited as Theorem 6.12)"}],"minor_comments":[{"comment":"The text says a primitive 'satisfies Theorem 3.2', but the object labelled 3.2 is a definition, not a theorem; this should be corrected throughout Section 3.","section":"Section 3.2 and 3.3"},{"comment":"Assumption 1 uses 'nu-almost every u', while Definition 4.5 requires invariance 'for every u in U'. These are different conditions and should be reconciled.","section":"Section 4.5 and Definition 4.5"},{"comment":"Several claims are cited as theorems that are actually remarks: the Jordan-form discussion appears as Remark 5.6 but is called 'Theorem 5.6' in Section 5.4; the approximate-invariance discussion is Remark 5.7 but is called 'Theorem 5.7' in Remark 4.11; and the EDMD comparison is Remark 5.8 but is called 'Theorem 5.8' in Section 2. The numbering should be made consistent.","section":"Section 5.4 and Remarks 5.6--5.8"},{"comment":"The caption contains a duplicated and garbled phrase: 'for ReLU it is not below for TopK; for ReLU it is not below baseline at the guard-selected operating point, which the text takes up. at the guard-selected operating point, which the text takes up.' This should be rewritten.","section":"Figure 11 caption"},{"comment":"The proof of Lemma 6.3 refers to 'Theorem 6.3' and 'Theorem 6.4' when it means the lemma itself; the cross-referencing between lemmas and theorems in Section 6 should be audited.","section":"Section 6.3 and 6.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is ambitious and unusually honest, but the proof of the main theorem is not complete as written: the affine-in-control gap in Definition 4.6, the missing moment condition on the lifted next state, and the unproved approximate-invariance extension are all fixable in scope, but they are load-bearing. I would not reject, but the manuscript should not be accepted until these are repaired. The experimental sections are a model of transparent reporting and should be preserved in any revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nRead this one if you care whether any spectral method in interpretability has a theorem behind it. The paper's real contribution is a finite-sample identifiability result for the Koopman spectrum of a transformer, with an M^{-1/2} rate plus a minimax lower bound, and a dissociation theorem saying non-normal realisations cannot have their variance-ordered principal basis align with the modal basis. That is genuinely new for the MI literature. The experiments are unusually honest: they report the pre-registered SAE-gap criterion failing (49% vs the registered 80%), the universality control separating seed replicas of one architecture, and the predicted exponent being attained only on Qwen3-8B-Base. You don't see that often.\n\nThe load-bearing proof has a gap, and the stress-test note identifies it correctly. Assumption 1 says the dictionary span is invariant under each control u, so for each u there is a matrix A_u. The theorem and Definition 4.6 need the stronger statement Psi(F(x,u)) = A Psi(x) + B u with a single pair (A,B). Invariance alone does not give that. The example F(x,u)=x+u with Psi=(1,x,x^2) is exactly K-invariant, yet A_u carries a u^2 term, so no constant (A,B) can represent the dynamics. The EDMDc object is then a control-distribution-dependent least-squares projection, not the model-intrinsic Koopman compression the theorem advertises. The paper needs either an explicit control-affine closure condition or a different estimand.\n\nThere are smaller issues in proportion. Lemma 6.3 asserts sub-exponential concentration of the cross-Gramian involving Psi(F(x,u)) without a stated moment or boundedness hypothesis on that random vector; sub-Gaussian Psi(x) and u do not automatically make Psi(F(x,u)) sub-Gaussian. The main theorem assumes exact K-invariance, which no real dictionary satisfies; the authors acknowledge the bias does not vanish, and Theorem 6.12 only adds it additively. The experimental support for the headline rate is one model and one summary statistic (maximum matched error); the median and mean slopes are -0.338 and -0.375 on Qwen, so \"attains the predicted exponent\" should be softened.\n\nDoes it deserve a serious referee? Yes. The identifiability framing, the dissociation theorem, and the candid failure analysis are worth engaging, and the minimax lower bound stands on its own. But the paper should not be accepted with Assumption 1 as stated. Send it to review with a request to fix the control-affine gap, add the missing tail condition, and temper the experimental claim.","headline":"A genuinely new spectral-identifiability claim for interpretability, undermined by a control-affineness gap in the main theorem; deserves review but needs revision.","tokens_in":48634,"tokens_out":2290,"would_cite":true,"duration_ms":23631,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The Koopman spectrum of a transformer's depth dynamics is a coordinate-free model property recoverable from finite calibration data at the optimal rate.","keywords":["Koopman operator","spectral identifiability","mechanistic interpretability","sparse autoencoders","dictionary learning","EDMDc","transformer depth dynamics","non-normality"],"falsifier":"On a synthetic controlled layer map with an exactly invariant dictionary, compute the EDMDc spectrum from $M$ samples for values of $M$ crossing the theorem's threshold and compare the maximum matched error to the eigenvalues of the known matrix $A$; if the error does not fall at the predicted $M^{-1/2}$ rate once the threshold is crossed, the theorem's mechanism is wrong.","tokens_in":47530,"feed_emoji":"🧠","tokens_out":13352,"duration_ms":119862,"temperature":0.7,"pith_summary":"This paper aims to give mechanistic interpretability a criterion for when a discovered structure is a property of the model rather than an artifact of the method. It treats the transformer forward pass as a controlled dynamical system with depth as time and lifts it through the Koopman operator, so that any dictionary whose span is closed under the layer map induces a finite linear realisation $\\Psi(F(x,u))=A\\Psi(x)+Bu$ whose eigenvalues are coordinate-free. The main theorem proves this spectrum is recoverable from $M$ calibration samples at rate $M^{-1/2}$ up to permutation, with a matching minimax lower bound and a heavy-tailed median-of-means variant. A companion dissociation theorem shows that whenever the realisation is non-normal, the directions carrying activation variance and the directions carrying information across depth cannot coincide, so the identifiable object and the legible object are provably distinct. The measurements on three pretrained transformers report the predicted exponent on the largest model and report their own failures, because each failure bounds the claim.","feed_headline":"Koopman spectrum: a provable, intrinsic fingerprint for transformers","feed_subtitle":"Unlike sparse-autoencoder features, its eigenvalues are recoverable at the optimal rate with a stated error bar.","key_machinery":"The load-bearing object is the controlled Koopman realisation $(A,B)$, the finite linear map satisfying $\\Psi(F(x,u))=A\\Psi(x)+Bu$ on a dictionary span that is $K$-invariant, meaning closed under composition with the layer map for every control value. The eigenvalues of $A$ are the identifiable invariants; the EDMDc estimator, the least-squares fit of $(A,B)$ from cached triples of lifted state, lift of next state, and control, is the procedure that extracts them from finite data. The proof chain uses sub-Gaussian matrix concentration for the empirical Gramians, least-squares stability on the event that the state Gramian is well conditioned, and spectral perturbation bounds that convert an operator-norm error into a permutation-matched eigenvalue error, with an optional optimal-matching bound that removes the spectral-gap hypothesis. The dissociation theorem is carried by the discrete Lyapunov equation $\\Sigma=A\\Sigma A^*+BG_uB^*$: if an orthonormal eigenbasis of the stationary covariance $\\Sigma$ consisted of eigenvectors of $A$, then $A$ would be normal, so measured condition numbers $\\kappa_2(\\widehat V)$ from $38$ to $495$ force the variance-ordered principal directions and the Koopman modes apart.","core_discovery":"The central claim is that transformer depth dynamics carry a finite-dimensional spectral invariant that is both coordinate-free and recoverable at the optimal statistical rate. Under the three structural assumptions, namely dictionary $K$-invariance, persistent excitation of the controls, and spectral separation, the EDMDc estimator returns a matrix $\\widehat A_M$ whose eigenvalue multiset converges to that of the true Koopman compression $A$ at rate $M^{-1/2}$, with permutation as the only remaining ambiguity (Theorem 6.1); a minimax lower bound shows no estimator can do better, and a gap-free optimal-matching version removes the spectral-separation threshold for the eigenvalue guarantee (Theorem 6.8). The same theory identifies why sparse autoencoders are not identifiable: their reconstruction-plus-$\\ell^1$ objective never enforces $K$-invariance, so the resulting bias does not vanish as $M$ grows, and an explicit invariance penalty partially repairs the failure. On three pretrained transformers the split-half spectral distance falls monotonically with $M$ and attains the predicted exponent on the largest model. Finally, the paper proves that in the non-normal regime these models occupy, the variance-ordered principal basis and the Koopman modal basis cannot coincide; the identifiable spectrum is therefore a certificate of model-intrinsic structure, not a decomposition into human-legible mechanisms.","pith_inferences":["A testable extension the paper leaves implicit is to use the measured invariance residual as a dictionary-selection score: as the residual shrinks, the theorem predicts the split-half spectral exponent should approach $-1/2$, which would turn dictionary design for interpretability into a quantitative optimisation problem.","The dissociation result suggests that future work should assign different tools to different claims: the Koopman spectrum for cross-run and cross-model identity certificates, and variance-based or causal directions for behavioural localisation, with the two never conflated in a safety case.","The exogeneity convention is the largest unresolved modelling choice; lifting to the joint $T$-token state or instrumenting the attention writes would likely change the spectrum more than any sampling error, so resolving it should precede any attempt to compare spectra across architectures.","One could also test the theory's dimension dependence directly: the gap-free bound predicts at most linear growth of the matched spectral level in dictionary size $N$, and the paper's own exponents flatten with $N$; extending the sweep to $N=256$ and $512$ would show whether the dimensional factor in the bound is tight."],"forward_implications":["Given a dictionary satisfying exact K-invariance, the recovered spectrum is certified as model-intrinsic: different calibration corpora, seeds, and dictionary bases return the same eigenvalue multiset up to the stated $O(M^{-1/2})$ error, so spectral claims carry a measurable error bar.","Sparse-autoencoder variability is diagnosed as structural: the objective omits K-invariance, the resulting bias persists as the sample grows, and adding the invariance penalty reduces split-half spectral distance by 41% at matched sparsity.","Because measured condition numbers place every fitted realisation in the non-normal regime, the Koopman spectrum cannot double as a legible circuit decomposition; the IOI experiments make the separation quantitative, with the PCA advantage decaying 4.1× as the question moves in depth.","The minimax lower bound makes the $M^{-1/2}$ rate optimal for the problem, and the median-of-means variant extends the guarantee to heavy-tailed activations even though on the tested dictionaries the lifting itself removes the tails.","Spectral equality of two realisations implies identical first-order intervention algebras, so cross-model universality becomes a testable spectral criterion, with the caveat that the criterion requires a seed-replica null rather than a sampling floor."],"supporting_citations":[{"why":"Supplies the EDMDc estimator: extended DMD and its control variant turn cached layer-token triples into a least-squares linear realisation.","marker":"[44, 45]"},{"why":"Establishes asymptotic convergence of EDMD to the Koopman operator, the background against which the finite-sample rate is new.","marker":"[48]"},{"why":"Provides the sub-Gaussian matrix concentration inequality that yields the $M^{-1/2}$ Gramian bounds in the proof.","marker":"[67]"},{"why":"The eigenvalue perturbation result that converts the estimator's operator-norm error into a permutation-matched eigenvalue error.","marker":"[70]"},{"why":"The eigenvector rotation result behind the spectral-gap-dependent eigenvector rate.","marker":"[71]"},{"why":"The optimal-matching spectral variation bound that removes the spectral-gap hypothesis for the eigenvalue guarantee.","marker":"[72]"},{"why":"Supplies the indirect-object-identification circuit in GPT-2 small used as the semantic testbed for the mode-localisation experiments.","marker":"[34]"},{"why":"The geometric median-of-means concentration result behind the heavy-tailed variant of Theorem 7.2.","marker":"[78]"}],"fun_headline_variants":["Spectral IDs: make model circuits provably intrinsic","Koopman spectrum: the model's true, rate-optimal fingerprint","Proof: circuits can be identified, not just found","Eigenvalues as ground truth for interpretability","Identifiable spectra beat legible circuits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proofs assume a dictionary whose span maps exactly into itself under the layer transformation; real dictionaries only approximate that closure, and the resulting bias does not vanish as more calibration samples are collected.","fun_headline_variants_meta":{"raw":{"variants":["Spectral IDs: make model circuits provably intrinsic","Koopman spectrum: the model's true, rate-optimal fingerprint","Proof: circuits can be identified, not just found","Eigenvalues as ground truth for interpretability","Identifiable spectra beat legible circuits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1472,"prompt_tokens":1152,"completion_tokens":320,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":768,"completion_tokens_details":{"reasoning_tokens":244}},"tokens_in":768,"tokens_out":320,"duration_ms":3397,"temperature":1.0,"reasoning_tokens":244,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:11:55.092858+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a synthetic controlled layer map with an exactly invariant dictionary, compute the EDMDc spectrum from $M$ samples for values of $M$ crossing the theorem's threshold and compare the maximum matched error to the eigenvalues of the known matrix $A$; if the error does not fall at the predicted $M^{-1/2}$ rate once the threshold is crossed, the theorem's mechanism is wrong.","supporting_citations":[{"cited_title":"On convergence of ex- tended dynamic mode decomposition to the Koopman operator.Journal of Nonlinear Science, 28(2):687–710, 2018","cited_arxiv_id":null,"evidence_quote":"Establishes asymptotic convergence of EDMD to the Koopman operator, the background against which the finite-sample rate is new."},{"cited_title":"Cam- bridge Series in Statistical and Probabilistic Mathe- matics","cited_arxiv_id":null,"evidence_quote":"Provides the sub-Gaussian matrix concentration inequality that yields the $M^{-1/2}$ Gramian bounds in the proof."},{"cited_title":"Bauer and Charles T","cited_arxiv_id":null,"evidence_quote":"The eigenvalue perturbation result that converts the estimator's operator-norm error into a permutation-matched eigenvalue error."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The eigenvector rotation result behind the spectral-gap-dependent eigenvector rate."},{"cited_title":"An optimal bound for the spectral variation of two matrices.Linear algebra and its ap- plications, 71:77–80, 1985","cited_arxiv_id":null,"evidence_quote":"The optimal-matching spectral variation bound that removes the spectral-gap hypothesis for the eigenvalue guarantee."},{"cited_title":"Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt","cited_arxiv_id":null,"evidence_quote":"Supplies the indirect-object-identification circuit in GPT-2 small used as the semantic testbed for the mode-localisation experiments."},{"cited_title":"Geometric median and robust esti- mation in Banach spaces.Bernoulli, 21(4):2308–2335, 2015","cited_arxiv_id":null,"evidence_quote":"The geometric median-of-means concentration result behind the heavy-tailed variant of Theorem 7.2."}],"review_version":1}