Pith. sign in

REVIEW 4 major objections 4 minor 24 references

This paper claims that transformer learning drives the Jacobian relaxation spectrum toward a universal near-flat infrared form with 1/t memory, reproducible across model sizes, depths, prompts, and training steps.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 07:05 UTC pith:FPX5QNLA

load-bearing objection There is a real spectral signal in the Pythia Jacobians, but the paper overinterprets it: K~1/t is a transform of the measured TDOS, and the 'critical formation' is a rescaling around an unmeasured bare rate. the 4 major comments →

arxiv 2607.10923 v2 pith:FPX5QNLA submitted 2026-07-12 cs.LG

Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics

classification cs.LG
keywords transformer dynamicstime-scale density of statesrelaxation spectrummemory kernelinfrared fixed pointcritical phenomenaJacobian eigenvalueslarge language models
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Large language models are usually studied through accuracy or loss; this paper instead looks at the eigenvalues of the layer-to-layer Jacobian and asks whether their slow relaxation modes organize collectively. The author reports that training pushes the time-scale density of states toward a nearly flat infrared form, ρ(λ)~λ^β with β≈−0.1, so that the derived memory kernel behaves as K(t)~1/t rather than as a single exponential. The same spectral shape appears across three model sizes, all training checkpoints, many prompts, and different layer depths, which is taken as evidence for a universal infrared fixed point. A further claim is that the memory self-energy peaks transiently around step ~2000, marking a critical formation point before the model settles into a metastable near-critical regime.

Core claim

The central claim is that trained transformers realize the collective observables described by Cognitive Field Theory: the time-scale density of states (TDOS), built from relaxation rates λα = −log|μα| of layer Jacobian eigenvalues, reorganizes during learning so that slow modes accumulate at low rates and the infrared TDOS becomes approximately flat with exponent β≈−0.1. From this measured TDOS, the memory kernel is computed directly as K(t)=∫ρ(λ)e^{-λt}dλ and exhibits robust 1/t long-memory scaling across training, prompts, network depth, and model scale. Local Jacobians measured from different prompts and different token subspaces converge to the same normalized infrared TDOS, which the p

What carries the argument

The key object is the time-scale density of states (TDOS), obtained by diagonalizing the Jacobian of the hidden-state mapping between transformer layers and converting each complex eigenvalue magnitude |μ| to a relaxation rate λ = −log|μ|. The paper treats the TDOS as the fundamental collective observable from which all other quantities follow: the memory kernel K(t), the static memory self-energy Σ(0)=∫dλ ρ(λ)/λ, the cognitive forgetting gap r_cog = r − Σ(0), and the collective susceptibility χ(0)=1/r_cog. The work that this machinery does is to compress a high-dimensional, prompt-dependent Jacobian into a one-dimensional spectral density whose infrared tail determines the long-time collect

Load-bearing premise

The critical-formation claim rests on the paper's normalization of the memory self-energy by its own maximum, because the absolute forgetting rate and coupling strength are never directly measured; if that normalization is the only source of the apparent gap minimum, the claimed critical point is a rescaling artifact rather than a measured phenomenon.

What would settle it

Measure the bare forgetting rate r and coupling g directly—for example, by injecting a small perturbation into the hidden state and measuring the autocorrelation decay of the response—and compute the unnormalized cognitive forgetting gap r_cog = r − Σ(0) across training checkpoints. If the minimum of the unnormalized gap does not occur near step ~2000 or does not approach a value much smaller than its late-training plateau by a factor consistent with the claimed critical enhancement, the transient critical-formation scenario is falsified.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the measured infrared organization is correct, scale-free long-term memory in transformers is a collective spectral property, not a property of any single attention head or token position.
  • The prompt-independence of the normalized TDOS implies that the slow-mode reservoir is shared across inputs, so memory capacity is a global architectural feature rather than a per-prompt artifact.
  • The transient self-energy maximum near step ~2000 predicts a specific, measurable moment of maximum collective susceptibility during training, which could be probed by perturbation-response experiments.
  • The near-constant exponent β≈−0.1 across training suggests that optimization changes the population of slow modes but does not change the universality class of the collective dynamics.
  • The same infrared organization across 70M–1.4B parameter models implies that the phenomenon may persist at even larger scales and could serve as a target observable for diagnosing training state.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A direct test would be to measure the bare forgetting rate r and coupling strength g from time-dependent perturbation experiments on trained networks, then compute the unnormalized gap r_cog = r − Σ(0); if the minimum gap at step ~2000 is not close to zero, the 'critical formation' is a normalization artifact rather than a measured transition.
  • The flat TDOS with β≈−0.1 resembles the spectrum one would get from certain random-matrix ensembles; comparing the measured eigenvalue statistics against a random-matrix null model would clarify whether the 'universal' infrared shape is specific to learned transformers or generic to high-dimensional nonlinear maps.
  • If the infrared collapse is genuine, similar spectral analysis applied to recurrent neural networks or biological neural recordings might reveal the same 1/t memory scaling, connecting artificial and natural collective memory under one framework.
  • Because eigenvalues converge across prompts while eigenvectors remain prompt-dependent, the semantic content of a prompt may be carried by eigenvectors rather than eigenvalues; a mode-resolved decomposition of output logits could test whether slow modes specifically control long-range reasoning.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper analyzes layer Jacobians of Pythia language models at various training checkpoints, prompts, depths, and scales. It constructs a time-scale density of states (TDOS) from eigenvalue magnitudes, fits an infrared power law ρ(λ)∼λ^β with β≈−0.1, computes a memory kernel K(t) as the Laplace transform of the TDOS, and evaluates a normalized memory self-energy and a normalized 'cognitive forgetting gap' to claim a transient critical formation near step ~2000 followed by a metastable near-critical regime. The central claims are universal infrared organization, K(t)∼1/t, and the first quantitative realization of Cognitive Field Theory observables in transformers.

Significance. The raw spectral measurements — progressive infrared accumulation of Jacobian eigenvalues toward |μ|→1 across training, prompts, depth, and model scale — are plausible and potentially interesting. The paper uses a reproducible, public model family and includes useful robustness checks (16-token vs 8-token inputs, prompt ensembles, three model sizes). If the critical-formation and universality claims were established, this would be a notable empirical bridge between transformer dynamics and critical-phenomena concepts. However, as analyzed below, the headline claims rest on an unmeasured bare rate r, a self-referential kernel computation, and the absence of any null-model baseline. The paper's own Sec. III.C admits that r and g are not independently determined, and Eq. (35) shows that the 'forgetting gap' is a rescaling of the self-energy rather than a measured physical quantity. The significance of the paper as a quantitative test of Cognitive Field Theory is therefore not currently established, although the underlying spectral observations may merit a more modest interpretation.

major comments (4)
  1. [Sec. III.C, Eq. (35)] The normalized 'cognitive forgetting gap' r̃_cog = 1 − Σ/Σ_max is not a proxy for the physical gap r_cog = r − Σ(0) unless r = Σ_max and g² is absorbed into the normalization. Both r and g are unmeasured, as the text admits. Moreover, the reported minimum of r̃_cog ≈ 0.39 at the self-energy maximum is inconsistent with Eq. (35), which gives 0 when Σ=Σ_max. Thus the transient critical-formation claim (rcog→0⁺, maximum χ) is an interpretive rescaling, not an empirical finding.
  2. [Sec. III.D, Eqs. (38)–(40)] The memory kernel is computed as K(t) = ∫dλ ρ(λ)e^{−λt} from the measured TDOS. Consequently, the observation K(t)∼1/t is a restatement of the fitted approximately flat infrared spectrum, not an independent confirmation. Calling the power-law behavior 'emergent' is misleading because Eq. (40) is the definitional Laplace transform of the same measurement; it cannot validate the TDOS fit or the universality claim.
  3. [Sec. III.E, Eqs. (41)–(43)] The infrared exponent β is extracted by fitting N(<λ)∝λ^{β+1} over 10^{-4}<λ<5×10^{-2}. The excellent R²>0.99 only shows that a power law describes this restricted range. No null model is provided: random Jacobians, untrained networks, reshuffled spectra, or a generic random-matrix ensemble would determine whether β≈−0.1 and the 'universality class' are specific to trained transformers. Without such a baseline, the claim that training preserves an infrared universality class is unsupported.
  4. [Sec. II.C, Eq. (22)] The temporal renormalization-group flow is written schematically as ρ_{S,h}(λ;{u_i}) = b^{y_ρ} ρ_{S/h,b_h}(b^z λ;{b^{y_i} u_i}) and then used to justify the central fixed-point universality prediction. No derivation or explicit definition of the scaling dimensions or flow is given. As written, the convergence of normalized TDOS across prompts in Sec. IV is consistent with the existence of a fixed point but does not test the RG equation. The theoretical framework should be stated as an assumption or supplied with a concrete derivation; otherwise the 'infrared fixed-point organization' is an interpretation rather than a tested mechanism.
minor comments (4)
  1. [Fig. 4] The y-axis labels should distinguish the normalized proxy r̃_cog = 1 − Σ/Σ_max from the physical forgetting gap r_cog. The numerical minimum 0.39 also needs reconciliation with Eq. (35), which yields 0 at Σ=Σ_max.
  2. [Sec. III.A, Eq. (33)] The Jacobian dimension (N_token d_hidden)×(N_token d_hidden) is stated for the first 8 tokens, but the text does not specify how the layer mapping is defined for multiple tokens. Clarify whether the Jacobian is computed with respect to the concatenated hidden states and whether this corresponds to the standard layer function.
  3. [References] Reference [15], Cognitive Field Theory, is cited as an arXiv preprint (v7). The paper should state the status of this reference, since the entire framework rests on it.
  4. [Appendix B] The 30-prompt ensemble is listed in full in Tables B1 and B2. The main text says 'fifteen representative prompts' and 'thirty prompts' in different places; please standardize the terminology and ensure the figure captions match the actual number shown.

Circularity Check

2 steps flagged

Critical-formation claim reduces to self-energy normalization (r̃_cog = 1 − Σ/Σmax); K(t)∼1/t is a Laplace transform of the same TDOS, so two headline 'predictions' are restatements of the measured spectrum.

specific steps
  1. self definitional [Sec. III.C, Eq. (35) and following text (Fig. 4)]
    "Since the absolute values of the bare forgetting rate r and the coupling constant g are not independently determined from the spectral measurements, we instead evaluate the normalized quantities Σ̃ = Σ/Σmax, r̃_cog = 1 − Σ/Σmax, (35) ... Within Cognitive Field Theory, this transient maximum corresponds to the closest experimental realization of the critical condition rcog → 0+, under which the collective susceptibility χ(0) ∝ 1/rcog becomes maximal."

    The physical forgetting gap is rcog = r − Σ(0), but r and g are unmeasured. Eq. (35) replaces this gap by 1 − Σ/Σmax, so the 'minimum forgetting gap' and 'maximum susceptibility' at step ~2000 are nothing but the point where the measured self-energy Σ is maximal; the critical condition rcog→0+ is imposed by effectively setting r = Σmax and absorbing g² into Σ, not by direct measurement. The text's reported minimum of about 0.39 also contradicts Eq. (35), which gives exactly 0 at Σ = Σmax. In either reading, the critical-formation narrative is a rescaling of the measured self-energy rather than an independent criticality measurement.

  2. self definitional [Sec. III.D, Eq. (40); Sec. III.E]
    "K(t) = ∫₀^∞ dλ ρ(λ)e^{−λt}. (40) Consequently, the long-time behavior of the memory kernel is not postulated but emerges directly from the experimentally measured relaxation spectrum itself. ... Combined with the directly measured K(t)∼1/t, memory kernel, this provides quantitative support for the infrared collective dynamics predicted by Cognitive Field Theory."

    Eq. (40) is exactly the definition of the memory kernel in terms of the very same measured TDOS ρ(λ). Thus the observed K(t)∼1/t is a deterministic mathematical consequence of the already fitted β ≈ −0.1 (Eqs. 41–43), not an independent confirmation. Presenting this transform as a 'directly measured' kernel that supports the TDOS scaling uses the same spectral data twice under a different name.

full rationale

The paper contains substantial independent measurements: the TDOS evolution from Pythia Jacobians, prompt-ensemble concentration, token-subspace convergence, and cross-scale reproducibility are empirical observations that do not reduce to the theory. However, two headline claims are less independent than presented. First, the 'critical formation' of the cognitive field is not measured: Eq. (35) defines r̃_cog = 1 − Σ/Σmax explicitly because r and g are undetermined, so the 'critical condition rcog→0+' is achieved by construction whenever the measured self-energy peaks; the reported minimum 0.39 is also arithmetically incompatible with Eq. (35). Second, K(t)∼1/t is not an independent check: Eq. (40) defines K as the Laplace transform of the same measured ρ, so the scaling is a restatement of the fitted β ≈ −0.1. These are partial circularities in the interpretive layer, not in the raw spectral measurements. The self-citation of Cognitive Field Theory [15] is provenance rather than load-bearing here, since the equations are re-derived in Sec. II. The score of 6 reflects that the central criticality narrative reduces by construction, while the underlying infrared organization retains independent empirical content.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 2 invented entities

The empirical claim rests on a chain of modeling choices: linearization of the layer map, constant coupling g, and normalization of the unmeasured bare rate r. The fitted exponent β and the derived kernel K(t) are the only quantitative outputs, and they are interlocked: K follows from ρ by definition. The 'critical formation' is an interpretation of a normalized self-energy proxy, not a measurement of a divergent susceptibility.

free parameters (5)
  • β (infrared exponent) = -0.1 (range -0.14 to -0.05)
    Fitted by ordinary least squares to the cumulative mode count N(<λ) over the hand-chosen interval 10^-4 < λ < 5e-2 (Eqs. 41-43). Central to the universal scaling claim.
  • α (memory kernel exponent) = 0.90 to 1.11
    Fitted slope of log K(t) vs log t (Eq. 52). Used to claim K~1/t, but K(t) is itself computed from the measured TDOS.
  • λ_cut (infrared cutoff) = 0.05
    Hand-chosen reference cutoff used to define the 'infrared sector' in TDOS plots and analysis. Affects which modes are counted as slow.
  • r (bare forgetting rate) = undetermined
    Appears in r_cog = r − Σ(0) (Eq. 17) but is never measured. The normalized proxy (Eq. 35) implicitly sets r = Σmax, which is what produces the 'critical formation' peak from the self-energy data.
  • g (coupling constant) = undetermined
    Appears as g² in the self-energy (Eq. 34) and is assumed constant in the infrared. Normalization removes it from the plotted observables, but absolute self-energy and criticality cannot be quantified without it.
axioms (4)
  • domain assumption The nonlinear layer-to-layer transformer map can be linearized and its Jacobian eigenmodes treated as exponential relaxation modes: u(n) = e^{-λn} (Eq. 32).
    Eqs. (30)-(32) assume local linearization captures collective dynamics over all relevant timescales and prompts; not validated against the nonlinear trajectory.
  • domain assumption g(λ) is approximately constant in the infrared, so the measured TDOS ρ(λ)=g²(λ)D(λ) reflects the underlying mode density D(λ).
    Stated after Eq. (10); needed to connect the measured spectrum to the memory self-energy, but not derived or tested.
  • ad hoc to paper A temporal renormalization-group flow (Eq. 22) exists and drives different local TDOS to a common infrared fixed point, with irrelevant perturbations vanishing.
    This is the central postulate of the author's Cognitive Field Theory [15]; the present paper provides no independent derivation or falsifiable handle beyond the spectra it measures.
  • ad hoc to paper The normalized forgetting gap r_cog = 1 − Σ/Σmax (Eq. 35) is a valid proxy for the true gap r − Σ(0).
    Assumes r = Σmax, which is not justified. Without this normalization the criticality claims in Sec. III.C cannot be evaluated.
invented entities (2)
  • Macroscopic cognitive field φ(t) no independent evidence
    purpose: A collective field whose critical formation is claimed to explain enhanced memory and susceptibility; used to reinterpret the self-energy peak as a critical point.
    No measurement outside the Jacobian spectra supports the field's existence; the critical gap proxy is not directly measured, and the theory is self-cited [15].
  • Protected metastable near-critical operating regime no independent evidence
    purpose: Describes the late-training state after the self-energy peak; invoked to explain why the system does not stay critical.
    This is a label applied to a plateau in the normalized self-energy proxy, with no independent test distinguishing it from ordinary non-critical relaxation.

pith-pipeline@v1.3.0-alltime-deepseek · 28970 in / 16908 out tokens · 167138 ms · 2026-08-02T07:05:03.538436+00:00 · methodology

0 comments
read the original abstract

Large language models exhibit remarkable emergent behaviors, yet the physical mechanism governing their collective dynamics remains poorly understood. Cognitive Field Theory predicts that learning organizes collective dynamics through the infrared accumulation of slow relaxation modes, enhancing memory self-energy, long-memory dynamics, and collective susceptibility. Here we test this framework directly in Transformer dynamics. Using publicly available Pythia language models, we extract relaxation spectra from layer Jacobians throughout training, prompt ensembles, network depth, and model scale, allowing the collective observables of Cognitive Field Theory to be measured quantitatively. The measurements reveal pronounced infrared reorganization of the relaxation spectrum. Slow relaxation modes progressively accumulate toward the infrared, producing an approximately flat time-scale density of states, \( \rho(\lambda)\sim\lambda^\beta,\ \beta\simeq-0.1, \) while the corresponding memory kernel exhibits universal scaling, \( K(t)\sim1/t. \) The collective observables further reveal a critical formation process: the memory self-energy reaches a transient maximum during early training before relaxing toward a metastable near-critical regime. Prompt-resolved and token-subspace measurements show that distinct local Jacobians converge toward the same normalized infrared TDOS, consistent with an infrared fixed-point organization under coarse graining. The reproducibility of the same infrared organization across training, prompt ensembles, network depth, and Transformer model scales establishes infrared slow-mode organization as a universal collective principle underlying Transformer dynamics and provides the first quantitative experimental realization of the collective observables introduced by Cognitive Field Theory.

Figures

Figures reproduced from arXiv: 2607.10923 by Byung Gyu Chae.

Figure 1
Figure 1. Figure 1: FIG. 1. Spectral characterization of the Transformer Jacobian and construction of the time-scale density of states. The [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 1
Figure 1. Figure 1: FIG. 1. Spectral characterization of the Transformer Jacobian and construction of the time-scale density of states. The [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2. Evolution of the time-scale density of states during Transformer training. The TDOS is constructed directly from the [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2. Evolution of the time-scale density of states during Transformer training. The TDOS is constructed directly from the [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3. Evolution of the collective relaxation-mode popula [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3. Evolution of the collective relaxation-mode popula [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4. Critical formation of the cognitive field inferred from [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4. Critical formation of the cognitive field inferred from [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: FIG. 5. Evolution of the memory kernel during Transformer [PITH_FULL_IMAGE:figures/full_fig_p011_5.png] view at source ↗
Figure 5
Figure 5. Figure 5: FIG. 5. Evolution of the memory kernel during Transformer [PITH_FULL_IMAGE:figures/full_fig_p012_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: FIG. 6. Infrared critical scaling of the experimentally measured time-scale density of states. The cumulative mode distributions [PITH_FULL_IMAGE:figures/full_fig_p012_6.png] view at source ↗
Figure 6
Figure 6. Figure 6: FIG. 6. Infrared critical scaling of the experimentally measured time-scale density of states. The cumulative mode distributions [PITH_FULL_IMAGE:figures/full_fig_p013_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: FIG. 7. Layer-dependent organization of the time-scale density of states measured from the fully trained Pythia-410M model. [PITH_FULL_IMAGE:figures/full_fig_p013_7.png] view at source ↗
Figure 7
Figure 7. Figure 7: FIG. 7. Prompt-resolved time-scale density of states measured from the Layer 10 [PITH_FULL_IMAGE:figures/full_fig_p014_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: FIG. 8. Layer-dependent memory self-energy and cogni [PITH_FULL_IMAGE:figures/full_fig_p014_8.png] view at source ↗
Figure 8
Figure 8. Figure 8: FIG. 8. Prompt-ensemble concentration of the normalized [PITH_FULL_IMAGE:figures/full_fig_p015_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: FIG. 9. Layer-dependent universality of the infrared mem [PITH_FULL_IMAGE:figures/full_fig_p015_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: FIG. 10. Layer-dependent organization of the time-scale density of states measured from the fully trained Pythia-410M model. [PITH_FULL_IMAGE:figures/full_fig_p017_10.png] view at source ↗
Figure 10
Figure 10. Figure 10: FIG. 10. Model-scale universality of the infrared organization during Transformer learning. The complete infrared analysis [PITH_FULL_IMAGE:figures/full_fig_p016_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: FIG. 11. Layer-dependent memory self-energy and cogni [PITH_FULL_IMAGE:figures/full_fig_p018_11.png] view at source ↗
Figure 11
Figure 11. Figure 11: FIG. 11. Evolution of the time-scale density of states during training for the layer 1 [PITH_FULL_IMAGE:figures/full_fig_p017_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: FIG. 12. Layer-dependent universality of the infrared mem [PITH_FULL_IMAGE:figures/full_fig_p019_12.png] view at source ↗
Figure 12
Figure 12. Figure 12: FIG. 12. Fraction of stable relaxation modes measured from [PITH_FULL_IMAGE:figures/full_fig_p018_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: FIG. 13. Model-scale universality of the infrared organization during Transformer learning. The complete infrared analysis [PITH_FULL_IMAGE:figures/full_fig_p020_13.png] view at source ↗
Figure 13
Figure 13. Figure 13: FIG. 13. Evolution of the normalized memory self-energy (up [PITH_FULL_IMAGE:figures/full_fig_p018_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: FIG. 14. Evolution of the time-scale density of states during training for the layer 1 [PITH_FULL_IMAGE:figures/full_fig_p021_14.png] view at source ↗
Figure 13
Figure 13. Figure 13: For every propagation depth, the normalized [PITH_FULL_IMAGE:figures/full_fig_p019_13.png] view at source ↗
Figure 15
Figure 15. Figure 15: FIG. 15. Fraction of stable relaxation modes measured from [PITH_FULL_IMAGE:figures/full_fig_p022_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: FIG. 16. Evolution of the normalized memory self-energy (up [PITH_FULL_IMAGE:figures/full_fig_p022_16.png] view at source ↗
Figure 16
Figure 16. Figure 16: For every propagation depth, the normalized [PITH_FULL_IMAGE:figures/full_fig_p023_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: FIG. 17. TDOS and corresponding memory kernels of the [PITH_FULL_IMAGE:figures/full_fig_p024_17.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

24 extracted references · 11 linked inside Pith

  1. [1]

    At- 26 tention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “At- 26 tention is all you need,” Advances in neural information processing systems (2017)

  2. [2]

    Improving language understanding by gen- erative pre-training,

    A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by gen- erative pre-training,” OpenAI (2018)

  3. [3]

    Language models are unsupervised multitask learners,

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI (2019)

  4. [4]

    Language models are few-shot learners,

    T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan et al., “Language models are few-shot learners,” In Ad- vances in Neural Information Processing Systems, 1877- 1901 (2020)

  5. [5]

    Training compute-optimal large language models,

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford et al., “Training compute-optimal large language models,” In Proceedings of the 36th In- ternational Conference on Neural Information Processing Systems. 30016-30030 (2022), 2022

  6. [6]

    Scaling laws for neural language models,

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models,” arXiv preprint arXiv.2001.08361 (2020)

  7. [7]

    Palm: Scaling language modeling with pathways,

    A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra et al., “Palm: Scaling language modeling with pathways,” arXiv preprint arXiv:2204.02311 (2022)

  8. [8]

    Universal transformers,

    M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, L. Kaiser, “Universal transformers,” arXiv preprint arXiv:1807.03819 (2018)

  9. [9]

    Are emergent abilities of large language models a mirage?,

    R. Schaeffer, B. Miranda, S. Koyejo, “Are emergent abilities of large language models a mirage?,” Advances in Neural Information Processing Systems, 55565-55581 (2023)

  10. [10]

    Emergent abilities of large lan- guage models,

    J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Met- zler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus, “Emergent abilities of large lan- guage models,” Transactions on Machine Learning Re- search (2022)

  11. [11]

    DeepSeek LLM: Scaling open-source language models with longtermism,

    X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Dinget et al., “DeepSeek LLM: Scaling open-source language models with longtermism,” arXiv preprint arXiv:2401.02954 (2024)

  12. [12]

    The renormalization group and theϵexpansion,

    K. G. Wilson and J. Kogut, “The renormalization group and theϵexpansion,” Phys. Rep.12, 75-199 (1974)

  13. [13]

    U. C. T¨ auber,Critical Dynamics(Cambridge University Press, Cambridge, 2014)

  14. [14]

    Self-organized criticality from protected mean-field dynamics: Loop stability and internal renor- malization in reflective neural systems,

    B. G. Chae, “Self-organized criticality from protected mean-field dynamics: Loop stability and internal renor- malization in reflective neural systems,” arXiv preprint arXiv:2601.04450 (2026)

  15. [15]

    Cognitive field theory: Memory-dressed collective dynamics of intelligence,

    B. G. Chae, “Cognitive field theory: Memory-dressed collective dynamics of intelligence,” arXiv preprint arXiv:2601.10221v7 (2026)

  16. [16]

    Pythia: A suite for analyzing large language models across training and scaling,

    S. Biderman, H. Schoelkopf, Q. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan et al., “Pythia: A suite for analyzing large language models across training and scaling,” In Proceedings of the 40th International Conference on Machine Learning (2023)

  17. [17]

    Emergent and pre- dictable memorization in large language models,

    S. Biderman, S. Prashanth, L. Sutawika, H. Schoelkopf, Q. Anthony, S. Purohit, and E. Raff, “Emergent and pre- dictable memorization in large language models,” arXiv preprint arXiv:2304.11158 (2023)

  18. [18]

    O. Wal, P. Lesci, M. Muller-Eberstein, N. Saphra, H. Schoelkopf, W. Zuidema, and S. Biderman, “PolyPythias: Stability and outliers across fifty language model pre-training runs, The Thirteenth International Conference on Learning Representations ( 2025)

  19. [19]

    ReAct: Synergizing reasoning and acting in language models,

    S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “ReAct: Synergizing reasoning and acting in language models,” arXiv preprint arXiv:2210.03629 (2022)

  20. [20]

    Linear transform- ers are secretly fast weight programmers,

    I. Schlag, T. Irie, and J. Schmidhuber, “Linear transform- ers are secretly fast weight programmers,” arXiv preprint arXiv:2102.11174 (2021)

  21. [21]

    Transformers are RNNs: Fast autoregres- sive transformers with linear attention,

    A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are RNNs: Fast autoregres- sive transformers with linear attention,” arXiv preprint arXiv:2006.16236 (2020)

  22. [22]

    Elic- iting latent predictions from transformers with the tuned lens,

    N. Belrose, I. Ostrovsky, L. McKinney, Z. Furman, L. Smith, D. Halawi, S. Biderman, and J. Steinhardt, “Elic- iting latent predictions from transformers with the tuned lens,” arXiv preprint arXiv:2303.08112 (2023)

  23. [23]

    Continual pre-training of large language models: How to re-warm your model?,

    K. Gupta, B. Th´ erien, A. Ibrahim, M. L. Richter, Q. An- thony, E. Belilovsky, I. Rish, and T. Lesort, “Continual pre-training of large language models: How to re-warm your model?,” Workshop on Efficient Systems for Foun- dation Models ICML (2023)

  24. [24]

    Large language models sometimes gener- ate purely negatively-reinforced text,

    F. Roger, “Large language models sometimes gener- ate purely negatively-reinforced text,” arXiv preprint arXiv:2306.07567 (2023). 27 Supplementary Materials Appendix A: Robustness with Respect to Input Sequence Length The analyses presented in the main text are performed using the first eight tokens of a fixed input prompt, corre- sponding to an 8192×8192...