REVIEW 4 major objections 4 minor 24 references
This paper claims that transformer learning drives the Jacobian relaxation spectrum toward a universal near-flat infrared form with 1/t memory, reproducible across model sizes, depths, prompts, and training steps.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 07:05 UTC pith:FPX5QNLA
load-bearing objection There is a real spectral signal in the Pythia Jacobians, but the paper overinterprets it: K~1/t is a transform of the measured TDOS, and the 'critical formation' is a rescaling around an unmeasured bare rate. the 4 major comments →
Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that trained transformers realize the collective observables described by Cognitive Field Theory: the time-scale density of states (TDOS), built from relaxation rates λα = −log|μα| of layer Jacobian eigenvalues, reorganizes during learning so that slow modes accumulate at low rates and the infrared TDOS becomes approximately flat with exponent β≈−0.1. From this measured TDOS, the memory kernel is computed directly as K(t)=∫ρ(λ)e^{-λt}dλ and exhibits robust 1/t long-memory scaling across training, prompts, network depth, and model scale. Local Jacobians measured from different prompts and different token subspaces converge to the same normalized infrared TDOS, which the p
What carries the argument
The key object is the time-scale density of states (TDOS), obtained by diagonalizing the Jacobian of the hidden-state mapping between transformer layers and converting each complex eigenvalue magnitude |μ| to a relaxation rate λ = −log|μ|. The paper treats the TDOS as the fundamental collective observable from which all other quantities follow: the memory kernel K(t), the static memory self-energy Σ(0)=∫dλ ρ(λ)/λ, the cognitive forgetting gap r_cog = r − Σ(0), and the collective susceptibility χ(0)=1/r_cog. The work that this machinery does is to compress a high-dimensional, prompt-dependent Jacobian into a one-dimensional spectral density whose infrared tail determines the long-time collect
Load-bearing premise
The critical-formation claim rests on the paper's normalization of the memory self-energy by its own maximum, because the absolute forgetting rate and coupling strength are never directly measured; if that normalization is the only source of the apparent gap minimum, the claimed critical point is a rescaling artifact rather than a measured phenomenon.
What would settle it
Measure the bare forgetting rate r and coupling g directly—for example, by injecting a small perturbation into the hidden state and measuring the autocorrelation decay of the response—and compute the unnormalized cognitive forgetting gap r_cog = r − Σ(0) across training checkpoints. If the minimum of the unnormalized gap does not occur near step ~2000 or does not approach a value much smaller than its late-training plateau by a factor consistent with the claimed critical enhancement, the transient critical-formation scenario is falsified.
If this is right
- If the measured infrared organization is correct, scale-free long-term memory in transformers is a collective spectral property, not a property of any single attention head or token position.
- The prompt-independence of the normalized TDOS implies that the slow-mode reservoir is shared across inputs, so memory capacity is a global architectural feature rather than a per-prompt artifact.
- The transient self-energy maximum near step ~2000 predicts a specific, measurable moment of maximum collective susceptibility during training, which could be probed by perturbation-response experiments.
- The near-constant exponent β≈−0.1 across training suggests that optimization changes the population of slow modes but does not change the universality class of the collective dynamics.
- The same infrared organization across 70M–1.4B parameter models implies that the phenomenon may persist at even larger scales and could serve as a target observable for diagnosing training state.
Where Pith is reading between the lines
- A direct test would be to measure the bare forgetting rate r and coupling strength g from time-dependent perturbation experiments on trained networks, then compute the unnormalized gap r_cog = r − Σ(0); if the minimum gap at step ~2000 is not close to zero, the 'critical formation' is a normalization artifact rather than a measured transition.
- The flat TDOS with β≈−0.1 resembles the spectrum one would get from certain random-matrix ensembles; comparing the measured eigenvalue statistics against a random-matrix null model would clarify whether the 'universal' infrared shape is specific to learned transformers or generic to high-dimensional nonlinear maps.
- If the infrared collapse is genuine, similar spectral analysis applied to recurrent neural networks or biological neural recordings might reveal the same 1/t memory scaling, connecting artificial and natural collective memory under one framework.
- Because eigenvalues converge across prompts while eigenvectors remain prompt-dependent, the semantic content of a prompt may be carried by eigenvectors rather than eigenvalues; a mode-resolved decomposition of output logits could test whether slow modes specifically control long-range reasoning.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes layer Jacobians of Pythia language models at various training checkpoints, prompts, depths, and scales. It constructs a time-scale density of states (TDOS) from eigenvalue magnitudes, fits an infrared power law ρ(λ)∼λ^β with β≈−0.1, computes a memory kernel K(t) as the Laplace transform of the TDOS, and evaluates a normalized memory self-energy and a normalized 'cognitive forgetting gap' to claim a transient critical formation near step ~2000 followed by a metastable near-critical regime. The central claims are universal infrared organization, K(t)∼1/t, and the first quantitative realization of Cognitive Field Theory observables in transformers.
Significance. The raw spectral measurements — progressive infrared accumulation of Jacobian eigenvalues toward |μ|→1 across training, prompts, depth, and model scale — are plausible and potentially interesting. The paper uses a reproducible, public model family and includes useful robustness checks (16-token vs 8-token inputs, prompt ensembles, three model sizes). If the critical-formation and universality claims were established, this would be a notable empirical bridge between transformer dynamics and critical-phenomena concepts. However, as analyzed below, the headline claims rest on an unmeasured bare rate r, a self-referential kernel computation, and the absence of any null-model baseline. The paper's own Sec. III.C admits that r and g are not independently determined, and Eq. (35) shows that the 'forgetting gap' is a rescaling of the self-energy rather than a measured physical quantity. The significance of the paper as a quantitative test of Cognitive Field Theory is therefore not currently established, although the underlying spectral observations may merit a more modest interpretation.
major comments (4)
- [Sec. III.C, Eq. (35)] The normalized 'cognitive forgetting gap' r̃_cog = 1 − Σ/Σ_max is not a proxy for the physical gap r_cog = r − Σ(0) unless r = Σ_max and g² is absorbed into the normalization. Both r and g are unmeasured, as the text admits. Moreover, the reported minimum of r̃_cog ≈ 0.39 at the self-energy maximum is inconsistent with Eq. (35), which gives 0 when Σ=Σ_max. Thus the transient critical-formation claim (rcog→0⁺, maximum χ) is an interpretive rescaling, not an empirical finding.
- [Sec. III.D, Eqs. (38)–(40)] The memory kernel is computed as K(t) = ∫dλ ρ(λ)e^{−λt} from the measured TDOS. Consequently, the observation K(t)∼1/t is a restatement of the fitted approximately flat infrared spectrum, not an independent confirmation. Calling the power-law behavior 'emergent' is misleading because Eq. (40) is the definitional Laplace transform of the same measurement; it cannot validate the TDOS fit or the universality claim.
- [Sec. III.E, Eqs. (41)–(43)] The infrared exponent β is extracted by fitting N(<λ)∝λ^{β+1} over 10^{-4}<λ<5×10^{-2}. The excellent R²>0.99 only shows that a power law describes this restricted range. No null model is provided: random Jacobians, untrained networks, reshuffled spectra, or a generic random-matrix ensemble would determine whether β≈−0.1 and the 'universality class' are specific to trained transformers. Without such a baseline, the claim that training preserves an infrared universality class is unsupported.
- [Sec. II.C, Eq. (22)] The temporal renormalization-group flow is written schematically as ρ_{S,h}(λ;{u_i}) = b^{y_ρ} ρ_{S/h,b_h}(b^z λ;{b^{y_i} u_i}) and then used to justify the central fixed-point universality prediction. No derivation or explicit definition of the scaling dimensions or flow is given. As written, the convergence of normalized TDOS across prompts in Sec. IV is consistent with the existence of a fixed point but does not test the RG equation. The theoretical framework should be stated as an assumption or supplied with a concrete derivation; otherwise the 'infrared fixed-point organization' is an interpretation rather than a tested mechanism.
minor comments (4)
- [Fig. 4] The y-axis labels should distinguish the normalized proxy r̃_cog = 1 − Σ/Σ_max from the physical forgetting gap r_cog. The numerical minimum 0.39 also needs reconciliation with Eq. (35), which yields 0 at Σ=Σ_max.
- [Sec. III.A, Eq. (33)] The Jacobian dimension (N_token d_hidden)×(N_token d_hidden) is stated for the first 8 tokens, but the text does not specify how the layer mapping is defined for multiple tokens. Clarify whether the Jacobian is computed with respect to the concatenated hidden states and whether this corresponds to the standard layer function.
- [References] Reference [15], Cognitive Field Theory, is cited as an arXiv preprint (v7). The paper should state the status of this reference, since the entire framework rests on it.
- [Appendix B] The 30-prompt ensemble is listed in full in Tables B1 and B2. The main text says 'fifteen representative prompts' and 'thirty prompts' in different places; please standardize the terminology and ensure the figure captions match the actual number shown.
Circularity Check
Critical-formation claim reduces to self-energy normalization (r̃_cog = 1 − Σ/Σmax); K(t)∼1/t is a Laplace transform of the same TDOS, so two headline 'predictions' are restatements of the measured spectrum.
specific steps
-
self definitional
[Sec. III.C, Eq. (35) and following text (Fig. 4)]
"Since the absolute values of the bare forgetting rate r and the coupling constant g are not independently determined from the spectral measurements, we instead evaluate the normalized quantities Σ̃ = Σ/Σmax, r̃_cog = 1 − Σ/Σmax, (35) ... Within Cognitive Field Theory, this transient maximum corresponds to the closest experimental realization of the critical condition rcog → 0+, under which the collective susceptibility χ(0) ∝ 1/rcog becomes maximal."
The physical forgetting gap is rcog = r − Σ(0), but r and g are unmeasured. Eq. (35) replaces this gap by 1 − Σ/Σmax, so the 'minimum forgetting gap' and 'maximum susceptibility' at step ~2000 are nothing but the point where the measured self-energy Σ is maximal; the critical condition rcog→0+ is imposed by effectively setting r = Σmax and absorbing g² into Σ, not by direct measurement. The text's reported minimum of about 0.39 also contradicts Eq. (35), which gives exactly 0 at Σ = Σmax. In either reading, the critical-formation narrative is a rescaling of the measured self-energy rather than an independent criticality measurement.
-
self definitional
[Sec. III.D, Eq. (40); Sec. III.E]
"K(t) = ∫₀^∞ dλ ρ(λ)e^{−λt}. (40) Consequently, the long-time behavior of the memory kernel is not postulated but emerges directly from the experimentally measured relaxation spectrum itself. ... Combined with the directly measured K(t)∼1/t, memory kernel, this provides quantitative support for the infrared collective dynamics predicted by Cognitive Field Theory."
Eq. (40) is exactly the definition of the memory kernel in terms of the very same measured TDOS ρ(λ). Thus the observed K(t)∼1/t is a deterministic mathematical consequence of the already fitted β ≈ −0.1 (Eqs. 41–43), not an independent confirmation. Presenting this transform as a 'directly measured' kernel that supports the TDOS scaling uses the same spectral data twice under a different name.
full rationale
The paper contains substantial independent measurements: the TDOS evolution from Pythia Jacobians, prompt-ensemble concentration, token-subspace convergence, and cross-scale reproducibility are empirical observations that do not reduce to the theory. However, two headline claims are less independent than presented. First, the 'critical formation' of the cognitive field is not measured: Eq. (35) defines r̃_cog = 1 − Σ/Σmax explicitly because r and g are undetermined, so the 'critical condition rcog→0+' is achieved by construction whenever the measured self-energy peaks; the reported minimum 0.39 is also arithmetically incompatible with Eq. (35). Second, K(t)∼1/t is not an independent check: Eq. (40) defines K as the Laplace transform of the same measured ρ, so the scaling is a restatement of the fitted β ≈ −0.1. These are partial circularities in the interpretive layer, not in the raw spectral measurements. The self-citation of Cognitive Field Theory [15] is provenance rather than load-bearing here, since the equations are re-derived in Sec. II. The score of 6 reflects that the central criticality narrative reduces by construction, while the underlying infrared organization retains independent empirical content.
Axiom & Free-Parameter Ledger
free parameters (5)
- β (infrared exponent) =
-0.1 (range -0.14 to -0.05)
- α (memory kernel exponent) =
0.90 to 1.11
- λ_cut (infrared cutoff) =
0.05
- r (bare forgetting rate) =
undetermined
- g (coupling constant) =
undetermined
axioms (4)
- domain assumption The nonlinear layer-to-layer transformer map can be linearized and its Jacobian eigenmodes treated as exponential relaxation modes: u(n) = e^{-λn} (Eq. 32).
- domain assumption g(λ) is approximately constant in the infrared, so the measured TDOS ρ(λ)=g²(λ)D(λ) reflects the underlying mode density D(λ).
- ad hoc to paper A temporal renormalization-group flow (Eq. 22) exists and drives different local TDOS to a common infrared fixed point, with irrelevant perturbations vanishing.
- ad hoc to paper The normalized forgetting gap r_cog = 1 − Σ/Σmax (Eq. 35) is a valid proxy for the true gap r − Σ(0).
invented entities (2)
-
Macroscopic cognitive field φ(t)
no independent evidence
-
Protected metastable near-critical operating regime
no independent evidence
read the original abstract
Large language models exhibit remarkable emergent behaviors, yet the physical mechanism governing their collective dynamics remains poorly understood. Cognitive Field Theory predicts that learning organizes collective dynamics through the infrared accumulation of slow relaxation modes, enhancing memory self-energy, long-memory dynamics, and collective susceptibility. Here we test this framework directly in Transformer dynamics. Using publicly available Pythia language models, we extract relaxation spectra from layer Jacobians throughout training, prompt ensembles, network depth, and model scale, allowing the collective observables of Cognitive Field Theory to be measured quantitatively. The measurements reveal pronounced infrared reorganization of the relaxation spectrum. Slow relaxation modes progressively accumulate toward the infrared, producing an approximately flat time-scale density of states, \( \rho(\lambda)\sim\lambda^\beta,\ \beta\simeq-0.1, \) while the corresponding memory kernel exhibits universal scaling, \( K(t)\sim1/t. \) The collective observables further reveal a critical formation process: the memory self-energy reaches a transient maximum during early training before relaxing toward a metastable near-critical regime. Prompt-resolved and token-subspace measurements show that distinct local Jacobians converge toward the same normalized infrared TDOS, consistent with an infrared fixed-point organization under coarse graining. The reproducibility of the same infrared organization across training, prompt ensembles, network depth, and Transformer model scales establishes infrared slow-mode organization as a universal collective principle underlying Transformer dynamics and provides the first quantitative experimental realization of the collective observables introduced by Cognitive Field Theory.
Figures
Reference graph
Works this paper leans on
-
[1]
At- 26 tention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “At- 26 tention is all you need,” Advances in neural information processing systems (2017)
2017
-
[2]
Improving language understanding by gen- erative pre-training,
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by gen- erative pre-training,” OpenAI (2018)
2018
-
[3]
Language models are unsupervised multitask learners,
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI (2019)
2019
-
[4]
Language models are few-shot learners,
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan et al., “Language models are few-shot learners,” In Ad- vances in Neural Information Processing Systems, 1877- 1901 (2020)
1901
-
[5]
Training compute-optimal large language models,
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford et al., “Training compute-optimal large language models,” In Proceedings of the 36th In- ternational Conference on Neural Information Processing Systems. 30016-30030 (2022), 2022
2022
-
[6]
Scaling laws for neural language models,
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models,” arXiv preprint arXiv.2001.08361 (2020)
Pith/arXiv arXiv 2001
-
[7]
Palm: Scaling language modeling with pathways,
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra et al., “Palm: Scaling language modeling with pathways,” arXiv preprint arXiv:2204.02311 (2022)
Pith/arXiv arXiv 2022
-
[8]
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, L. Kaiser, “Universal transformers,” arXiv preprint arXiv:1807.03819 (2018)
Pith/arXiv arXiv 2018
-
[9]
Are emergent abilities of large language models a mirage?,
R. Schaeffer, B. Miranda, S. Koyejo, “Are emergent abilities of large language models a mirage?,” Advances in Neural Information Processing Systems, 55565-55581 (2023)
2023
-
[10]
Emergent abilities of large lan- guage models,
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Met- zler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus, “Emergent abilities of large lan- guage models,” Transactions on Machine Learning Re- search (2022)
2022
-
[11]
DeepSeek LLM: Scaling open-source language models with longtermism,
X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Dinget et al., “DeepSeek LLM: Scaling open-source language models with longtermism,” arXiv preprint arXiv:2401.02954 (2024)
Pith/arXiv arXiv 2024
-
[12]
The renormalization group and theϵexpansion,
K. G. Wilson and J. Kogut, “The renormalization group and theϵexpansion,” Phys. Rep.12, 75-199 (1974)
1974
-
[13]
U. C. T¨ auber,Critical Dynamics(Cambridge University Press, Cambridge, 2014)
2014
-
[14]
B. G. Chae, “Self-organized criticality from protected mean-field dynamics: Loop stability and internal renor- malization in reflective neural systems,” arXiv preprint arXiv:2601.04450 (2026)
arXiv 2026
-
[15]
Cognitive field theory: Memory-dressed collective dynamics of intelligence,
B. G. Chae, “Cognitive field theory: Memory-dressed collective dynamics of intelligence,” arXiv preprint arXiv:2601.10221v7 (2026)
Pith/arXiv arXiv 2026
-
[16]
Pythia: A suite for analyzing large language models across training and scaling,
S. Biderman, H. Schoelkopf, Q. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan et al., “Pythia: A suite for analyzing large language models across training and scaling,” In Proceedings of the 40th International Conference on Machine Learning (2023)
2023
-
[17]
Emergent and pre- dictable memorization in large language models,
S. Biderman, S. Prashanth, L. Sutawika, H. Schoelkopf, Q. Anthony, S. Purohit, and E. Raff, “Emergent and pre- dictable memorization in large language models,” arXiv preprint arXiv:2304.11158 (2023)
Pith/arXiv arXiv 2023
-
[18]
O. Wal, P. Lesci, M. Muller-Eberstein, N. Saphra, H. Schoelkopf, W. Zuidema, and S. Biderman, “PolyPythias: Stability and outliers across fifty language model pre-training runs, The Thirteenth International Conference on Learning Representations ( 2025)
2025
-
[19]
ReAct: Synergizing reasoning and acting in language models,
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “ReAct: Synergizing reasoning and acting in language models,” arXiv preprint arXiv:2210.03629 (2022)
Pith/arXiv arXiv 2022
-
[20]
Linear transform- ers are secretly fast weight programmers,
I. Schlag, T. Irie, and J. Schmidhuber, “Linear transform- ers are secretly fast weight programmers,” arXiv preprint arXiv:2102.11174 (2021)
Pith/arXiv arXiv 2021
-
[21]
Transformers are RNNs: Fast autoregres- sive transformers with linear attention,
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are RNNs: Fast autoregres- sive transformers with linear attention,” arXiv preprint arXiv:2006.16236 (2020)
Pith/arXiv arXiv 2006
-
[22]
Elic- iting latent predictions from transformers with the tuned lens,
N. Belrose, I. Ostrovsky, L. McKinney, Z. Furman, L. Smith, D. Halawi, S. Biderman, and J. Steinhardt, “Elic- iting latent predictions from transformers with the tuned lens,” arXiv preprint arXiv:2303.08112 (2023)
Pith/arXiv arXiv 2023
-
[23]
Continual pre-training of large language models: How to re-warm your model?,
K. Gupta, B. Th´ erien, A. Ibrahim, M. L. Richter, Q. An- thony, E. Belilovsky, I. Rish, and T. Lesort, “Continual pre-training of large language models: How to re-warm your model?,” Workshop on Efficient Systems for Foun- dation Models ICML (2023)
2023
-
[24]
Large language models sometimes gener- ate purely negatively-reinforced text,
F. Roger, “Large language models sometimes gener- ate purely negatively-reinforced text,” arXiv preprint arXiv:2306.07567 (2023). 27 Supplementary Materials Appendix A: Robustness with Respect to Input Sequence Length The analyses presented in the main text are performed using the first eight tokens of a fixed input prompt, corre- sponding to an 8192×8192...
Pith/arXiv arXiv 2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.