Pith. sign in

REVIEW 3 major objections 7 minor 22 references

Eigenvalues miss finite-depth attention dynamics; sinks damp transients and a routing minority holds reserve that eigenvalues cannot see.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 04:10 UTC pith:IUGUYSCM

load-bearing objection Resolvent tools on the attention propagator give a real, pre-registered win over eigenvalues for depth-transient persistence and routing identity, with the 7–11-order failure carefully scoped to the attention skeleton. the 3 major comments →

arxiv 2607.09279 v1 pith:IUGUYSCM submitted 2026-07-10 cond-mat.dis-nn

Transient Reserves, Sink Dampers, and the Failure of Eigenvalue Reasoning in the Attention Propagator

classification cond-mat.dis-nn
keywords attention propagatornon-normal operatorsKreiss constantpseudospectraattention sinkstransient reservedepth productscausal transformers
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Causal attention matrices are row-stochastic and non-normal by the mask. For non-normal operators, eigenvalues govern only the infinite-depth limit; what a finite network actually does is controlled by resolvent quantities such as Kreiss constants and pseudospectra. This paper tests whether that distinction predicts anything real about trained transformers that eigenvalues miss. After stripping the mask-forced stationary projector, learned non-normality turns out to be signed: a small routing minority carries excess transient reserve that tracks previous-token function and doubles when induction heads engage, while the sink majority is suppressed below mass-matched nulls and acts as a transient damper. On real depth products the eigenvalue heuristic underpredicts surviving deviations by seven to eleven orders of magnitude—an error absent in matched nulls. Training consolidates this organization after circuit formation, and a clamping intervention on Llama-3-8B traces a causal path from three massive residual dimensions through sink attention to damping. A pre-registered contest finds resolvent features required for depth-transient persistence and routing-head identity, while every single-operator summary fails to predict which heads are causally critical.

Core claim

On trained transformers the eigenvalue picture of attention depth products is not a loose bound but the wrong account: surviving deviations exceed the eigenvalue heuristic by 10^7 to 10^11, while mass-matched nulls stay calibrated. Learned non-normality is signed—routing heads hold excess transient reserve, sinks damp it below chance—and resolvent features are required, under pre-registered cross-validated rules, for depth-transient persistence and routing identity.

What carries the argument

Exact Perron deflation: with P the mask-forced projector 1 e_0^T, A^k = P + A_dev^k and the depth deviation product factors as Π_L − P = product of the deflated layer-mean operators. All transient metrics and the eigenvalue-failure comparison are computed on this deflated skeleton against row-permutation nulls.

Load-bearing premise

The claim rests on treating head-mean attention-only depth products (without value maps, MLPs, or residual paths except a separate rollout check) as a faithful enough skeleton of how tokens mix.

What would settle it

Re-run the pre-registered feature contest and the depth-product comparison after replacing the attention-only head-mean stack with the full residual-stream Jacobian or with products that include OV and MLP maps; if eigenvalue-side features then match or beat resolvent features on half-depth persistence and the 10^7–10^11 gap disappears, the central claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper argues that causal attention matrices are non-normal by construction, so finite-depth token mixing is governed by resolvent quantities (pseudospectra, Kreiss constants) rather than eigenvalues. After establishing that the mask pins the raw Kreiss constant at √n and that Perron deflation yields an exact product factorization of depth deviation dynamics, it reports a signed census across GPT-2, Pythia-410m, and Llama-3-8B: a routing minority carries excess transient reserve correlated with previous-token function, while the sink majority is suppressed below mass-matched nulls and acts as a transient damper. On head-mean depth products the eigenvalue heuristic underpredicts surviving deviations by 10^7–10^11 on trained stacks (an error absent in matched nulls). Supporting results include state-conditional reserve in induction heads, post-formation consolidation on Pythia checkpoints, an architecture-conditional clamping intervention linking massive activations to sink damping on Llama-3-8B, and a pre-registered cross-validated contest in which resolvent features are required for depth-transient persistence and routing identity but no single-operator summary predicts per-head causal criticality.

Significance. If the claims hold within the stated scope, the work supplies a concrete, falsifiable reason to stop treating spectral gaps and eigenvalue decay as default accounts of attention over depth, and it gives an operator-level function for attention sinks (transient dampers) that is measurable, state-conditional, and architecture-conditional. Strengths that raise the bar for the field include: exact structural propositions (mask-trivial Kreiss pinning; Perron deflation identity), systematic mass-matched nulls with FDR control, multi-architecture replication, bit-compatible checkpoint censuses, explicit scoring of the near-tautological Kreiss observable as negative, dual-preprocessing pre-registered contest rules with full negative reporting, and a causal intervention rather than pure correlation. The companion division of labor with the static QK operator is also a useful organizational contribution. The practical readings on outlier-aware compression and sink retention in long-context inference are proportionate to the evidence.

major comments (3)
  1. [§4, Table 2, Fig. 2] §4 and Table 2: the headline 10^7–10^11 eigenvalue-failure magnitude is measured on head-mean attention-only products Π_L = ∏ Ā^(l)_dev (and a residual-path rollout check). Proposition 2 makes the factorization exact for that skeleton, and the learned-vs-null contrast is clean, but transfer of the reported orders of magnitude to full forward-pass dynamics (with OV maps, MLPs, and residual paths) is load-bearing for the claim that “eigenvalue reasoning about depth fails on trained stacks.” §10 already lists the abstraction as a limitation; the manuscript should either (i) add at least one controlled check that re-inserts OV/MLP structure on a subset of layers/heads and reports how D_L/E_L changes, or (ii) restate the abstract and §4 claim more narrowly as a result about the attention skeleton, with the full-network implication demoted to discussion. Without one of those, the magnitude cl
  2. [§7, Fig. 5] §7 intervention: the causal chain on Llama-3-8B (three massive dimensions → BOS sink collapse 0.675 o0.049 → K_dev 1.41 o7.44) is the only causal evidence for the damper mechanism, yet it rests on eight sequences and whole-stream channel clamps that the text itself notes are stronger than pointwise edits. Random-dimension controls are appropriate, but the sample is too small to support the architecture-conditional mechanism claim at the same weight as the multi-model census. Either enlarge the intervention set and report sequence-level variability, or qualify the causal claim as a pilot demonstration whose architecture contrast (RMSNorm/RoPE vs LayerNorm) remains correlational outside Llama.
  3. [§8, Table 3] §8 contest, Table 3: support is declared only when R²(P)>R²(E) (or the union increment) holds under both raw-standardized and rank-normal preprocessings. That dual rule is strict and well-motivated, but preprocessing sensitivity is material: previous-token identity on Pythia fails under raw features and is recorded negative, while GPT-2 margins differ sharply between preprocessings (ΔR² +0.128 vs +0.566). For the central “resolvent features are required” verdict on persistence (three of three models), report the actual R² pairs and confidence intervals for all three models under both preprocessings in the main text or a table, not only the Llama head-to-head numbers. Without those numbers the pre-registered win is hard to audit at the level the protocol promises.
minor comments (7)
  1. [Abstract, §1] Abstract and §1: “err by seven to eleven orders of magnitude” should be explicitly scoped to head-mean attention products so that abstract readers do not infer a full-network result before §4 and §10.
  2. [Table 1] Table 1: the Llama induction correlation (+0.02, n.s.) is correctly flagged; consider adding a one-line note that induction excess is state-conditional (§5) so the dormant-census entry is not misread as a null result on induction structure.
  3. [Fig. 1] Fig. 1 right panel: the dual y-axis (median K_dev and BOS mass) is dense; a small annotation of the layer where real K_dev peels from the null would help.
  4. [§6] §6: the 160m “formal miss” on pattern-precedes-function (79% vs 90% bar) is handled honestly; move the absolute-onset secondary readout fully into a short appendix table so the main text stays on the pre-registered half-final criterion.
  5. [Appendix D] Appendix D / Fig. 8: clarify that annotated K_dev values are single-matrix illustrations, not the census medians, already in the caption but worth a sentence in the main text when the figure is first referenced.
  6. [References] References: Fernando & Guitchounts (2026) is marked “PREPRINT, VERIFY”; either verify the citation or drop the flag before camera-ready.
  7. [§2] Notation: A_dev, K_dev, and ρ_ε excess are introduced cleanly in §2; a short symbol table would still help readers jumping to §3–§4.

Circularity Check

1 steps flagged

No significant circularity: mask-pinned Kreiss and exact deflation are standard operator facts; the sole near-tautology is explicitly flagged and scored negative; companion citations supply parallel annotations and weight-space trajectories, not definitional premises for the A-side claims.

specific steps
  1. self citation load bearing [§3 Table 1 caption; §6; §8 feature families; §10]
    "Head-class scores (previous-token, induction) are behavioral and come from the companion paper’s probe taxonomy. ... the companion paper’s population Dhead trajectory crosses its own half-depth at the same checkpoint ... Family E contains ... and the companion paper’s weight-space eigenvalue features. Family P contains ... and the weight-space non-normality scalars."

    Behavioral targets and some weight-space covariates are imported from the same-author companion rather than measured independently inside this study. The step is only weakly circular: the labels are used as prediction targets (not as definitional inputs that force the resolvent excesses), the companion analyzes a different operator (M), and the A-side signed census, depth-product failure, and contest wins for persistence/routing stand without them. Scored as minor non-load-bearing self-citation.

full rationale

The derivation chain is self-contained. Propositions 1–2 follow immediately from the causal row-stochastic structure (lower-triangular, mask-forced Perron pair P=1e0⊤, ∥P∥2=√n) and are verified numerically; they are not fitted to behavioral data. All subsequent claims are excesses over mass-matched row-permutation (and Dirichlet) nulls, measured transients, depth-product comparisons (DL vs EL), checkpoint trajectories, a clamping intervention, and a pre-registered cross-validated contest between feature families. The one place where a resolvent quantity is related to an observable by theorem rather than by transformer-specific fact (supk∥Akdev∥2≈Kdev) is flagged in §8, excluded from positive credit, and recorded as formally negative. Companion-paper citations supply head-class labels, weight-space Dhead trajectories, and a constrained-training pilot; these are parallel analyses of the static scoring operator M and external annotations, not uniqueness theorems or load-bearing premises that force the A-side resolvent results. No fitted parameter is renamed a prediction, no ansatz is smuggled via self-citation, and no known empirical pattern is merely re-coordinatized. The paper therefore scores at most a minor self-citation that is not load-bearing for the central subdomain claim.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 2 invented entities

The mathematical core rests on classical non-normal operator theory plus the architectural fact that causal attention is lower-triangular and row-stochastic. Empirical claims rest on chosen corpus length, null definitions, and the attention-only depth abstraction rather than on free physical constants. Interpretive labels (reserve, damper) are operationalized by measurable excesses and an intervention, not postulated particles.

free parameters (5)
  • sequence length n = 128
    All matrices use n=128 BOS-prefixed pile-10k rows; Kreiss mask pin √n and all census statistics depend on this fixed length.
  • sequences per head/state = 8 to 32
    8–32 corpus matrices per head; effect sizes are large but sampling depth is a free design choice.
  • BH-FDR q = 0.001
    Significance of excess/suppression uses Benjamini–Hochberg at q=0.001 over heads.
  • pseudospectral ε grid = 1e-1, 1e-2, 1e-3
    ρε reported at ε ∈ {10^{-1},10^{-2},10^{-3}} with finite-grid floors; levels are conventional choices.
  • ridge contest regularization and preprocessing
    Inner-loop ridge and dual raw/rank-normal preprocessings control the decisive R² contest; dual rule adopted after instability was observed.
axioms (5)
  • standard math For non-normal operators, finite-horizon growth is controlled by resolvent quantities (ε-pseudospectrum, Kreiss constant) with K ≤ sup‖A^k‖ ≤ e n K.
    Invoked from Trefethen–Embree throughout §1–2 as the reason eigenvalues are insufficient.
  • domain assumption Causal post-softmax attention is lower-triangular and row-stochastic, forcing the Perron pair (1, e0) and P = 1 e0^T.
    Architectural fact used for Propositions 1–2 and all deflated metrics.
  • domain assumption Row permutation within causal support preserves mass profile while destroying learned column/diagonal structure, so excess over this null measures learning.
    Primary null in §2–3; all signed-census claims are relative to it.
  • ad hoc to paper Head-mean attention products (without OV/MLP) are a sufficient token-mixing skeleton for depth-transient claims.
    Defines DL in §4 and the persistence contest target; flagged as a limitation.
  • domain assumption Behavioral head-class scores (previous-token, induction) from the companion probe taxonomy correctly label function for correlation and contest targets.
    Used in Table 1, §5, and §8 identity tasks.
invented entities (2)
  • transient reserve (Kreiss excess of A_dev) independent evidence
    purpose: Scalar for learned non-normal capacity beyond the mask and mass profile.
    Operationalized as Kdev excess over row-perm nulls; not a new physical field, but a named functional quantity central to the signed census.
  • sink dampers independent evidence
    purpose: Operator-level function of attention sinks as suppressors of transient amplification.
    Supported by suppression below nulls, depth collapse of median Kdev with BOS mass, and Llama clamping that releases Kdev when sinks collapse.

pith-pipeline@v1.1.0-grok45 · 19865 in / 3494 out tokens · 44001 ms · 2026-07-13T04:10:47.629486+00:00 · methodology

0 comments
read the original abstract

The attention matrix of a causal transformer is row-stochastic, iterated over depth, and non-normal by construction. For non-normal operators, eigenvalues control only asymptotic behavior; finite-depth behavior is controlled by resolvent quantities such as pseudospectra and Kreiss constants. We test, under pre-registered criteria, whether this resolvent view predicts anything about trained transformers that eigenvalues miss. Two structural facts organize the analysis: the mask pins the Kreiss constant of every causal stochastic matrix at $\sqrt{n}$, and deflating the mask-forced Perron projector factorizes the depth deviation dynamics exactly into a product of deflated operators. Across GPT-2, Pythia-410m, and Llama-3-8B, learned non-normality proves to be signed. A routing minority carries excess transient reserve that tracks previous-token function and doubles when induction heads engage, while the sink majority is suppressed below matched shuffle nulls, so that attention sinks act as transient dampers. On depth products, eigenvalue predictions of surviving deviations err by seven to eleven orders of magnitude, an error absent in matched nulls. Checkpoint censuses date this organization to a consolidation phase after circuit formation, and a clamping intervention on Llama-3-8B establishes a causal chain from three massive activation dimensions through sink attention to transient damping; LayerNorm models implement the same functions elsewhere. A cross-validated contest concludes that resolvent features are required for depth-transient persistence and routing-head identity, and that no single-operator summary of any kind predicts per-head causal criticality.

Figures

Figures reproduced from arXiv: 2607.09279 by Li Hengyu (Institute for Solid State Physics, The University of Tokyo).

Figure 1
Figure 1. Figure 1: Left: per-head median Kdev, real versus row-permutation null (GPT-2, log–log axes). The nulls concentrate in a narrow band near 6 while real heads spread across [1, 15.7]; previous-token￾like heads (bright) lie on or above the diagonal, and sink-dominated heads collapse to the analytic normal floor Kdev = 1. Right: the damper over depth. The median real Kdev tracks the null in layers 0–2 and then collapses… view at source ↗
Figure 2
Figure 2. Figure 2: Depth deviation-transient for GPT-2. Left: the measured [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Transient reserve tracks functional state (GPT-2). Induction heads (bright) leave the [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Training-time consolidation on Pythia-410m. Left: the population median [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: The causal test on Llama-3-8B. The three massive dimensions are clamped at every [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Textbook validation: Λε of the Grcar matrix (n=100) for ε from 10−1 to 10−4 , reproducing the classical contours far beyond the spectrum. B Amplifier-head transient modes Two GPT-2 heads, L1H10 and L11H8, are the only ones in the census whose deflated deviation dynamics grows before it decays: ∥Ak dev∥2 rises with k to an interior maximum at ksup = 19 and 64 respectively, reaching 15.6 and 15.9, so that th… view at source ↗
Figure 7
Figure 7. Figure 7: Top singular pair of A ksup dev for the two interior amplifiers (GPT-2 L11H8 and L1H10), averaged over sequences. The transient reads the contrast between BOS and the first content token (v ≈ (e0 ± e1)/ √ 2) and writes it nearly uniformly across downstream positions (|u| ≈ n −1/2 ). C The 160M real-corpus implementation-invariance pilot Six Pythia-160M-architecture models (12 layers × 12 heads, d=768, rota… view at source ↗
Figure 8
Figure 8. Figure 8: Λε contours (ε from 10−1 to 10−4 ) for one input sequence: an induction/sink head (top row) and the previous-token head 4.11 (bottom row) against their row-permutation and Dirichlet nulls. The previous-token head shows the truncated-Toeplitz ring near |z| = 1; the sink head collapses to near-normal structure, with Kdev = 1.0, below its null. The scalar summaries used throughout the paper are level-set stat… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

22 extracted references · 13 linked inside Pith

  1. [1]

    Quantifying attention flow in transformers

    Samira Abnar and Willem Zuidema. Quantifying attention flow in transformers. InACL, 2020. arXiv:2005.00928

  2. [2]

    Pseudospectral bounds for transient amplification in coupled gradient descent, 2026

    Ahanaf Hasan Ariq. Pseudospectral bounds for transient amplification in coupled gradient descent, 2026. arXiv:2606.04031; HiLD workshop, ICML 2026

  3. [3]

    Attention is not all you need: pure attention loses rank doubly exponentially with depth

    Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. Attention is not all you need: pure attention loses rank doubly exponentially with depth. InICML, 2021. arXiv:2103.03404

  4. [4]

    Dynamics of the transformer residual stream: Coupling spectral geometry to network topology, 2026

    Jesseba Fernando and Grigori Guitchounts. Dynamics of the transformer residual stream: Coupling spectral geometry to network topology, 2026. arXiv:2605.14258 — PREPRINT, VERIFY

  5. [5]

    The Pile: an 800GB dataset of diverse text for language modeling, 2020

    Leo Gao et al. The Pile: an 800GB dataset of diverse text for language modeling, 2020. arXiv:2101.00027

  6. [6]

    The emergence of clusters in self-attention dynamics

    Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet. The emergence of clusters in self-attention dynamics. InNeurIPS, 2023. arXiv:2305.05465

  7. [7]

    A mathematical perspective on transformers.Bulletin of the American Mathematical Society, 62:427–479, 2025

    Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet. A mathematical perspective on transformers.Bulletin of the American Mathematical Society, 62:427–479, 2025. arXiv:2312.10794. 13

  8. [8]

    Non-normal spectral signatures of instability in neural network training dynam- ics, 2026

    Souvik Ghosh. Non-normal spectral signatures of instability in neural network training dynam- ics, 2026. arXiv:2605.23476

  9. [9]

    Mark S. Goldman. Memory without feedback in a neural network.Neuron, 61(4):621–634, 2009

  10. [10]

    When attention sink emerges in language models: an empirical view

    Xiangming Gu, Tianyu Pang, Chao Du, Qian Liu, Fengzhuo Zhang, Cunxiao Du, Ye Wang, and Min Lin. When attention sink emerges in language models: an empirical view. InICLR,

  11. [11]

    Jordan, and Song Mei

    Tianyu Guo, Druv Pai, Yu Bai, Jiantao Jiao, Michael I. Jordan, and Song Mei. Active- dormant attention heads: mechanistically demystifying extreme-token phenomena in LLMs,

  12. [12]

    Non-normalrecurrentneuralnetwork(nnRNN):learning long time dependencies while improving expressivity with transient dynamics

    Giancarlo Kerg, Kyle Goyette, Maximilian Puelma Touzel, Gauthier Gidel, Eugene Vorontsov, YoshuaBengio, andGuillaumeLajoie. Non-normalrecurrentneuralnetwork(nnRNN):learning long time dependencies while improving expressivity with transient dynamics. InNeurIPS,

  13. [13]

    When the complex spectrum of attention does work: architecture-conditional non-Hermitian structure in the QK operator, 2026

    Hengyu Li. When the complex spectrum of attention does work: architecture-conditional non-Hermitian structure in the QK operator, 2026. Companion paper, arXiv:2607.06621

  14. [14]

    Computing the Kreiss constant of a matrix, 2020

    Tim Mitchell. Computing the Kreiss constant of a matrix, 2020. arXiv:1907.06537

  15. [15]

    Murphy and Kenneth D

    Brendan K. Murphy and Kenneth D. Miller. Balanced amplification: a new mechanism of selective amplification of neural activity patterns.Neuron, 61(4):635–648, 2009

  16. [16]

    Mind the gap: a spectral analysis of rank collapse and signal propagation in attention layers

    Thiziri Nait Saada, Alireza Naderi, and Jared Tanner. Mind the gap: a spectral analysis of rank collapse and signal propagation in attention layers. InICML, 2025. arXiv:2410.07799

  17. [17]

    Attention sinks and compression valleys in LLMs are two sides of the same coin, 2025

    Enrique Queipo-de Llano, Álvaro Arroyo, Federico Barbero, Xiaowen Dong, Michael Bronstein, Yann LeCun, and Ravid Shwartz-Ziv. Attention sinks and compression valleys in LLMs are two sides of the same coin, 2025. arXiv:2510.06477

  18. [18]

    Zico Kolter, and Zhuang Liu

    Mingjie Sun, Xinlei Chen, J. Zico Kolter, and Zhuang Liu. Massive activations in large language models. InCOLM, 2024. arXiv:2402.17762

  19. [19]

    Trefethen

    Lloyd N. Trefethen. Computation of pseudospectra.Acta Numerica, 8:247–295, 1999

  20. [20]

    Trefethen and Mark Embree.Spectra and Pseudospectra: The Behavior of Nonnormal Matrices and Operators

    Lloyd N. Trefethen and Mark Embree.Spectra and Pseudospectra: The Behavior of Nonnormal Matrices and Operators. Princeton University Press, 2005

  21. [21]

    Efficient streaming language models with attention sinks

    Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. InICLR, 2024. arXiv:2309.17453. A Validation of theσ min core The grid core reproduces the Grcar pseudospectrum (Fig. 6) and the Jordan-blockε1/n law (ρε = 0.253againstε 1/n = 0.251atn=10,ε= 10 −6). It agrees with dense SVD to10 −...

  22. [22]

    and writes it nearly uniformly across downstream positions (|u| ≈n−1/2). C The 160M real-corpus implementation-invariance pilot Six Pythia-160M-architecture models (12 layers×12 heads,d=768, rotary fraction0.25) were trained from scratch on 4B tokens of real corpus under the companion paper’s constrained-training protocol [13]: two free seeds; two seeds w...