Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Cognitive Field Theory: Memory-Dressed Collective Dynamics of Intelligence

T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read Cognitive function is governed by the collective organization of dynamical time scales, not by microscopic units.

desk verdict The central emergent-time-scale result rests on an unevaluated one-loop closure; the unifying framing is neat, but the paper overclaims and the key quantitative claim is not supported. read the letter →

arxiv 2601.10221 v8 pith:Z2R72Y4Z submitted 2026-01-15 q-bio.NC

classification q-bio.NC
keywords cognitivefieldtheorytime-scaledensityofstatescollectivemodesdynamicaltimescalesmemorypersistencestochasticdynamicsreentrantrenormalization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that learning, inference, memory, and emergence in brains and artificial networks are governed by the collective distribution of relaxation time scales in a high-dimensional dynamical system, not by individual neurons, synapses, or precise activity patterns. It starts from a single stochastic equation for a collective state on a learned, state-dependent geometry and turns it into a path-integral field theory. In that theory, inference is the most probable trajectory, while fluctuation-induced loop corrections generically generate slow collective modes that carry memory. Learning is reinterpreted as a slow reshaping of the time-scale density of states (TDOS), and stable cognition is the persistence of those slow modes. If correct, the framework unifies Hopfield networks, recurrent networks, transformers, and homeostatic reentry networks as special cases and makes the reorganization of the TDOS directly measurable from power spectra and autocorrelations.

What carries the argument

The central object is the time-scale density of states (TDOS), the distribution of relaxation rates λα given by the real parts of the eigenvalues of the linearized stability matrix M around an inference trajectory. It organizes collective modes by temporal persistence rather than energy or momentum, and the paper treats its reorganization by learning as the physical mechanism of cognition. The argument is carried by the Martin–Siggia–Rose–Janssen–de Dominicis path-integral representation of the stochastic dynamics, with inference as the saddle point and loop corrections generated by nonlinearity. The quantitative engine is the one-loop Dyson relation m = m0 + cλD/m, which produces a self-org

What would settle it

Compute the actual one-loop self-energy for a concrete stochastic neural field with cubic nonlinearity. If the self-energy is not a local, Markovian tadpole at low frequency, or if finite-size effects cut off the infrared divergence so that no √(cλD) scale appears, the central emergence mechanism is refuted. Empirically, record activity from a large neural population across a learning task, estimate the TDOS by inverting the Lorentzian relation Tr C(ω) ∝ ∫ dλ ρ(λ)/(ω²+λ²), and check whether slow-mode spectral weight grows with memory persistence; a null result would falsify the claim that cogn

Watch

Extended reading notes

Core claim

The paper's central claim is that cognitive function is controlled not by microscopic units or precise activity patterns but by the collective organization of dynamical time scales across the full joint state space. Concretely, it claims that nonlinear drift plus stochastic noise necessarily produces new collective time scales through loop corrections: near criticality the slowest mode is dressed by a self-energy and satisfies the Dyson self-consistency m = m0 + cλD/m, so that even with m0 = 0 a finite scale m = √(cλD) emerges. These dressed low-frequency modes, called neurotons, act as the carriers of memory and temporal integration. Learning and homeostatic regulation reorganize the time-s

Load-bearing premise

The load-bearing premise is the assumed one-loop closure m = m0 + cλD/m: the paper states that c is set by a loop integral and regularization scheme but never evaluates it, and if the true self-energy is nonlocal in time rather than a Markovian tadpole, the emergent slow scale m = √(cλD) does not follow.

Editorial extensions

If this is right

  • If the central claim is right, the TDOS should be inferable from measured power spectra or autocorrelations, so learning-related changes in neural time-scale structure become directly testable.
  • Hopfield networks, RNNs, transformers, and homeostatic reentry networks are not different computational mechanisms but different constraints on the same stochastic dynamical equation, giving a hierarchy of expressive power.
  • Stable memory and inference correspond to persistent spectral weight at long time scales, so memory is a dynamical property of the collective spectrum rather than a stored static pattern.
  • Because loop corrections generate slow modes whenever nonlinearity and noise coexist, the framework predicts that emergent cognitive time scales are generic and do not require fine-tuning.
  • The FHRN example suggests that homeostatic regulation can endogenously generate a stiff radial / soft tangential geometry, providing a geometric substrate for long-lived collective modes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension: in any trained recurrent or transformer-like model, measuring the eigenspectrum of the linearized dynamics before and after training should show spectral weight moving toward slow modes; a failure to see this reorganization would count against the theory.
  • If neurotons are the carriers of memory, then targeted perturbation of the slow-mode subspace of the stability matrix should disrupt stored behavior far more than damaging random individual units — an experiment that could distinguish this view from unit-centric accounts.
  • The framework suggests architectural design should focus on learnable spectral shaping and reentrant feedback rather than raw depth or width, since it is the time-scale spectrum that determines what can be remembered.
  • The claimed generic emergence of m = √(cλD) could be checked in a minimal stochastic neural network with a tunable cubic nonlinearity by measuring whether the slowest relaxation time actually grows as the inverse square root of the noise/feedback coupling, as the Dyson relation predicts.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper proposes a unified stochastic field theory of cognition. Starting from a Langevin equation with state-dependent geometry, potential, and reentrant flow, the author constructs an MSRJD path integral, identifies inference with saddle-point trajectories, and argues that one-loop self-energy corrections generate a new collective relaxation scale m = sqrt(c λ D) (Eqs. 23–24). This scale is used to define 'neurotons' as dressed collective modes and to introduce the time-scale density of states (TDOS) as the central diagnostic: learning reshapes the TDOS, selectively stabilizing slow modes that underlie memory and inference. The paper then presents Hopfield networks, RNNs, linear-attention transformers, and the author's FHRN model as special limits of the unified equation, and proposes that the TDOS can be inferred from measured power spectra, yielding a falsifiable prediction.

Significance. If the central derivation were sound, the paper could offer a genuinely unifying language for neural dynamics, connecting stochastic field theory, spectral reorganization under learning, and empirical measures of temporal correlations. The manuscript is ambitious and clearly written, and the supplementary appendices make an explicit effort to compute concrete objects: the FHRN radial potential (Eq. 41), the projector decomposition (Appendix B), the Fisher-metric derivation (Appendix D), and the resolvent representation of the TDOS (Appendix A) are all standard and correctly stated. However, the main quantitative result — the fluctuation-generated infrared scale — is not derived but assumed through an unevaluated closure, and the advertised falsifiable prediction is tautological in the presented form. The significance is therefore conditional: the framework may be useful as a heuristic vocabulary, but the load-bearing claims of emergent time-scale generation and testability are not established by the manuscript.

major comments (4)
  1. [IV.C, Eqs. (20)–(24)] The central result m = sqrt(c λ D) is not derived. Equation (20) posits a tadpole self-energy Σ_R^(1)(0) ∝ λ⟨δx²⟩, Eq. (21) estimates ⟨δx²⟩ ∼ D/λ_min from the bare propagator, and Eq. (23) then replaces λ_min with the renormalized m. No loop integral is evaluated, no sign is determined, the frequency dependence of Σ_R(ω) is discarded, and the constant c is left unspecified. Since this self-consistent equation is the basis for the claim in Sec. VII.A that loop corrections 'necessarily generate new collective time scales,' and for the entire TDOS-reorganization picture of learning and memory, the manuscript's main quantitative claim is an ansatz presented as a derivation.
  2. [VI.D and Appendix B, Eq. (S18)] The FHRN radial reduction is called 'exact' (Eq. 39), but the derivation in Appendix B requires the dominant-eigenmode alignment y ≈ r v0, Eq. (S18). For a generic mixing matrix W, y^T W y / r^2 is not equal to ρ r^2; the approximation fails even for typical random matrices. Consequently, the exact radial potential (Eq. 41), the shell geometry, and the TDOS shown in Fig. 4 are not rigorous consequences of the FHRN dynamics. This weakens both the FHRN example and the claim that the FHRN realizes the unified theory's geometry.
  3. [VII.D and Appendix A(II), Eq. (S10)] The proposed falsifiable prediction is circular. In the stationary OU approximation, Eq. (S10) defines Tr C(ω) as an integral transform of ρ(λ) by construction. Any measured spectrum can therefore be represented by some positive ρ(λ), possibly after regularization, so the inference protocol cannot fail. To be falsifiable, the theory must supply constraints or predictions for how ρ(λ) changes under learning, beyond the statement that 'learning reshapes the TDOS.' As written, the empirical test is vacuous.
  4. [V, Eq. (26)] The learning rule is left schematic. Equation (26) defines θ-dynamics through an unspecified functional F, and no concrete F is analyzed. The paper's central claim that learning selectively stabilizes slow collective modes and reorganizes the TDOS is therefore not derived from the learning dynamics; it is a verbal assertion. A concrete learning rule or a limiting case showing spectral reorganization would be needed to make the learning–memory connection quantitative.
minor comments (5)
  1. [Fig. 2 caption] Typo: 'illustartes' should be 'illustrates'.
  2. [Appendix A opening] Typo: 'This providee' should be 'This provides'.
  3. [Reference [25]] The journal name 'Z Phyik' should read 'Z. Phys. B'.
  4. [Eq. (10) and Appendix C] Index placement for the Riemann curvature tensor is inconsistent between the main text (R_i_jkl) and Appendix C (R^i_jkl); this should be harmonized.
  5. [Sec. IV.C] The statement that loop corrections 'necessarily dominate' as λ_min → 0 is not justified without evaluating the loop integral; it would be helpful to state the assumed scaling in N and coupling.

Circularity Check

1 steps flagged · score 6.0 of 10

The one advertised falsifiable prediction (TDOS reorganization) is definitional: TDOS is inferred by inverting the same Lorentzian transform that defines the power spectrum, so the prediction cannot fail.

  1. self definitional [Appendix A(II), Eq. (S10); Sec. VII.D]
    "When M is symmetric, or approximately so, the trace of the correlation function admits the spectral representation TrC(ω)∝∑_α 1/(ω²+λ_α²)=∫ dλ ρ(λ) 1/(ω²+λ²). Thus, the observed power spectrum is an integral transform of the TDOS with a Lorentzian kernel. ... The TDOS may be inferred from empirical autocorrelation functions, power spectra, or linear response measurements ... Systematic reorganization of the TDOS under learning, adaptation, or task engagement provides a direct and falsifiable prediction of the theory."

    ρ(λ) is defined in Eq. (22)/Eq. (S5) as the spectral density of the stability matrix M, i.e., ρ is a functional of the eigenvalue set {λ_α}. Equation (S10) then states exactly that the predicted power spectrum is the Lorentzian transform of this same ρ. Consequently, 'inferring' ρ from measured power spectra is inverting a definitional integral transform, and the 'prediction' that learning reorganizes ρ is equivalent to saying that learning changes the power spectrum. Any spectral change can be absorbed into ρ through the inverse (regularized) transform, so the claimed falsifiable prediction carries no independent, theory-specific constraint. This is a diagnostic re-description rather than a derived prediction.

full rationale

The rest of the derivation chain is not circular. Equation (1) is an explicit postulate; the saddle-point identification (Sec. III) is a standard MSRJD result; the loop-correction claim (Eqs. 19–24) is an unevaluated Hartree-type closure with an undetermined constant c, which is a derivation gap / correctness concern (sign, frequency dependence, and coefficient of Σ_R are not computed), not a reduction of the conclusion to its input. The FHRN decomposition (Appendix B) is carried out in the paper, and the alignment assumption y≈r v0 is stated rather than hidden. The self-citations ([29], [38]–[40]) are to the author's own prior FHRN and renormalization work, but the central unified-field derivation does not rest on them. However, the advertised falsifiable test in Sec. VII.D / Appendix A(II) is self-definitional: TDOS is defined as the spectral density whose Lorentzian transform is the observed power spectrum, so 'TDOS reorganization' reduces to 'power-spectrum change' by construction. This warrants a score of 6 (one prediction reduces by construction) even though the central claim has independent speculative content.

Assumptions & free parameters 4 free parameters · 7 assumptions · 3 invented entities

The framework rests on standard physics (MSRJD, Helmholtz-Hodge, Fisher information), which is legitimate, but the paper-specific content is carried by four unspecified or free quantities: the loop constant c, the TDOS smoothing width, the Fig. 4 model parameters, and the learning functional F. The two invented entities with physical content (r_cog, φ=Ae^{iψ}) exist only in the abstract, and 'neuroton' is a rename of a slow eigenmode with no independent observable handle. Consequently the ledger shows a large gap between the strength of the claims and the fixed content actually supplied.

free parameters (4)
  • c (one-loop amplitude, Eq. 23) = unspecified / free
    Constant 'set by the loop integral and regularization scheme'; the integral is never evaluated, so the emergent mass m=√(cλD) has no fixed value and no predictive content.
  • ε (TDOS Gaussian smoothing width, Eq. S6) = unspecified
    Kernel width that shapes Fig. 4's density plot; not stated anywhere in text or captions.
  • FHRN numerical parameters for Fig. 4 (N, γ, κ, ρ, W, D) = unspecified
    The claimed dense accumulation of slow modes comes from a simulation whose parameters are not reported, so the exhibit is not checkable.
  • learning rule functional F (Eq. 26) = unspecified
    The entire claim that learning reorganizes the TDOS is carried by ˙θ=ε⟨F⟩ with F never specified; without a concrete F this is a placeholder, not an equation.
assumptions (7)
  • standard math Langevin dynamics (1) with white noise is equivalent to the MSRJD path integral with action (4)
    Textbook result (MSR 1973; Janssen 1976; De Dominicis 1976), cited [24-26]; no issue.
  • standard math Any regular vector field on a Riemannian manifold splits into a metric gradient plus a solenoidal part (Helmholtz-Hodge)
    Used to justify Eq. (6) and the 'geometric necessity' of the reentrant flow R(x); standard under the cited regularity conditions [41].
  • domain assumption Inference = saddle point of the action requires the weak-noise / large-N limit and a well-defined most-probable path
    Sec. III.A: 'In the small-noise limit D→0, this distribution becomes sharply peaked...' — the Bayesian path-inference interpretation depends on this limit and on the OU noise structure.
  • ad hoc to paper One-loop self-energy is a Markovian tadpole: Σ_R(0) ∝ λ⟨δx²⟩, ⟨δx²⟩ ≈ D/m (self-consistent closure)
    Eqs. (20)-(23): the central quantitative claim is an assumed closure; no diagram is evaluated and c is never computed.
  • ad hoc to paper Dominant-eigenmode alignment y ≈ r v₀ in the FHRN radial reduction
    Appendix B, Eq. (S18); Sec. VI.D calls Eq. (39) an 'exact radial equation' despite this approximation, which generically fails for random mixing matrices.
  • domain assumption Observed power spectra satisfy the OU spectral identity (S10) over relevant channels
    The TDOS-from-data proposal assumes stationary OU fluctuations, approximately symmetric M, and treats deviations from the identity as measurement noise rather than model failure.
  • domain assumption Effective state-space geometry is a Fisher information metric (G = inverse prediction-error covariance)
    Appendix D: identifying dynamical anisotropy with the information geometry of an internal likelihood (Eqs. S34-S39) is an interpretive mapping, not derived from the dynamics; the ∂Σ terms are dropped as 'subleading' to obtain the simple form.
invented entities (3)
  • neuroton
    purpose: Self-generated collective relaxation mode / dressed pole of G_R; claimed carrier of memory and computation (Sec. IV.D)
    Defined as the pole of the retarded Green's function; its lifetime τ∝1/m depends on the uncomputed constant c, so no quantitative, falsifiable signature is predicted.
  • cognitive forgetting gap r_cog
    purpose: Renormalized memory decay rate claimed in the abstract to be dressed by memory feedback
    Appears only in the abstract; never defined or derived in the v2 body text.
  • cognitive field order parameter φ = A e^{iψ}
    purpose: Abstract claims amplitude A encodes collective cognitive organization and phase ψ temporal coherence
    Absent from the v2 body; the framework as written never introduces this order parameter.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Cognitive Field Theory: Memory-Dressed Collective Dynamics of Intelligence." pith.science (2026). https://pith.science/paper/Z2R72Y4Z

@misc{pith2026260110221,
  author       = {Pith},
  title        = {Pith review of: Cognitive Field Theory: Memory-Dressed Collective Dynamics of Intelligence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Z2R72Y4Z}},
  note         = {Machine review of arXiv:2601.10221}
}
abstract

Learning, inference, memory, and emergence in biological and artificial systems are often described using disparate theoretical frameworks. Here we develop a cognitive field theory in which cognition is described as a collective nonequilibrium phenomenon governed by the geometry and collective spectrum of a learned cognitive manifold. Starting from a stochastic cognitive-field equation on an adaptive Riemannian manifold, we derive an effective cognitive field theory incorporating nonlocal memory kernels and retarded self-energy feedback. The learned cognitive geometry generates a complex collective spectrum characterized by the time-scale density of states $\rho(\lambda,\omega)$, whose relaxation and circulation sectors govern memory persistence and temporal coherence. Integrating out latent slow collective modes produces non-Markovian memory feedback that renormalizes the cognitive forgetting gap $r_{\rm cog}$, enhances collective cognitive susceptibility, and drives the system toward a protected near-critical regime characterized by long-time contextual persistence and scale-free temporal organization. The observable cognitive field emerges as a macroscopic order parameter, $\phi=Ae^{i\psi}$, whose amplitude encodes collective cognitive organization and whose phase encodes temporal coherence across distributed collective modes. Within this framework, learning organizes cognitive geometry, cognitive geometry generates a collective spectrum, and the resulting memory feedback stabilizes a memory-dressed cognitive field. The theory provides a unified dynamical description of learning, memory, inference, selfhood, and emergent intelligence in terms of the infrared organization of collective cognitive dynamics.

Figures

Figures reproduced from arXiv: 2601.10221 by the authors.

Figure 1
Figure 1. FIG. 1: Conceptual overview of the unified dynamical field theory. (a) Collective neural states evolve on a learned state [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 3
Figure 3. FIG. 3: Effective homeostatic potential of the Fast-weights [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figure 4
Figure 4. displays the TDOS of the FHRN, extracted from the spectrum of the linearized dynamics around a stable homeostatic fixed point, revealing a dense accu￾mulation of slow collective modes. The pronounced ac￾cumulation of modes at small relaxation rates λ reveals a dense spectrum of slow collective modes generated by the interplay of homeostatic stabilization and antisymmetric reentrant mixing. These long-lived modes dom… view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics

    cs.LG 2026-07 reject novelty 6.0 of 10

    Slow relaxation modes in Pythia transformers accumulate toward zero rate during training, yielding a near-flat infrared spectrum and 1/t memory kernels—but the claimed 'critical cognitive field formation' is not direc...

Reference graph

Works this paper leans on

46 extracted references · 4 linked inside Pith · cited by 1 Pith paper

  1. [1]

    G. M. Edelman,Neural Darwinism: The Theory of Neu- ronal Group Selection(Basic Books, New York, 1989)

  2. [2]

    The Mexican-hat–like geometry illustrates radial sta- bilization combined with a flat angular direction, providing a geometric substrate for collective computation

    in the full state space. The Mexican-hat–like geometry illustrates radial sta- bilization combined with a flat angular direction, providing a geometric substrate for collective computation. Reentrant dy- namics generate non-conservative flows along the shell, while the potential enforces homeostatic control of activity magni- tude. Consequently, transform...

  3. [3]

    A measure for brain complexity: Relating functional segregation and integration in the nervous system,

    G. Tononi, O. Sporns, and G. M. Edelman, “A measure for brain complexity: Relating functional segregation and integration in the nervous system,” Proc. Natl. Acad. Sci. USA91, 5033-5037 (1994)

  4. [4]

    Emergent complex neural dynamics,

    D. R. Chialvo, “Emergent complex neural dynamics,” Nat. Phys.6, 744-750 (2010)

  5. [5]

    Gerstner, W

    W. Gerstner, W. M. Kistler, R. Naud, and L. Paninski, Neuronal Dynamics(Cambridge University Press, 2014)

  6. [6]

    Hebb and home- ostasis in neuronal plasticity,

    G. G. Turrigiano and S. B. Nelson, “Hebb and home- ostasis in neuronal plasticity,” Curr. Opin. Neurobiol.10, 358-364 (2000)

  7. [7]

    Synaptic computa- tion,

    L. F. Abbott and W. G. Regehr, “Synaptic computa- tion,” Nature431, 796-803 (2004)

  8. [8]

    Neuronal avalanches in neo- cortical circuits,

    J. M. Beggs and D. Plenz, “Neuronal avalanches in neo- cortical circuits,” J. Neurosci.23, 11167-11177 (2003)

Show all 46 references
  1. [9]

    Bidirectional recurrent neural networks,

    M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural networks,” IEEE Trans. Signal Process.45, 2673- 2681 (1997)

  2. [10]

    At- tention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “At- tention is all you need,” arXiv:1706.03762 (2017)

  3. [11]

    Emergent abilities of large lan- guage models,

    J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Met- zler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus, “Emergent abilities of large lan- guage models,” Transactions on Machine Learning Re- search (2022)

  4. [12]

    ReAct: Synergizing reasoning and acting in language models,

    S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “ReAct: Synergizing reasoning and acting in language models,” arXiv:2210.03629 (2022)

  5. [13]

    Neural networks and physical systems with emergent collective computational abilities,

    J. J. Hopfield, “Neural networks and physical systems with emergent collective computational abilities,” Proc. Natl. Acad. Sci. USA79, 2554-2558 (1982)

  6. [14]

    Storing infinite numbers of patterns in a spin-glass model of neural networks,

    D. J. Amit and H. Gutfreund, “Storing infinite numbers of patterns in a spin-glass model of neural networks,” Phys. Rev. Lett.55, 1530-1533 (1985)

  7. [15]

    Synaptic and intrinsic home- ostasis cooperate to optimize single neuron response properties and tune integrator circuits,

    P. Cannon and J. Miller, “Synaptic and intrinsic home- ostasis cooperate to optimize single neuron response properties and tune integrator circuits,” J. Neurophys- iol.116, 2004-2022 (2016)

  8. [16]

    Bio- physical models of intrinsic homeostasis: Firing rates and beyond,

    N. Niemeyer, J. H. Schleimer, and S. Schreiber, “Bio- physical models of intrinsic homeostasis: Firing rates and beyond,” Curr. Opin. Neurobiol.70, 81-88 (2021)

  9. [17]

    Using fast weights to deblur old memories,

    G. E. Hinton and D. C. Plaut, “Using fast weights to deblur old memories,” Proc. 9th Annu. Conf. Cognitive Science Society, 177-186 (1987)

  10. [18]

    Learning to control fast-weight mem- ories: An alternative to dynamic recurrent networks,

    J. Schmidhuber, “Learning to control fast-weight mem- ories: An alternative to dynamic recurrent networks,” Neural Comput.4, 131-139 (1992)

  11. [19]

    Linear transformers are secretly fast weight programmers,

    I. Schlag, T. Irie, and J. Schmidhuber, “Linear transformers are secretly fast weight programmers,” arXiv:2102.11174 (2021)

  12. [20]

    Transformers are RNNs: Fast autoregressive transform- ers with linear attention,

    A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are RNNs: Fast autoregressive transform- ers with linear attention,” arXiv:2006.16236 (2020)

  13. [21]

    The renormalization group and theϵexpansion,

    K. G. Wilson and J. Kogut, “The renormalization group and theϵexpansion,” Phys. Rep.12, 75-199 (1974)

  14. [22]

    Cardy,Scaling and Renormalization in Statistical Physics(Cambridge University Press, Cambridge, 1996)

    J. Cardy,Scaling and Renormalization in Statistical Physics(Cambridge University Press, Cambridge, 1996). 14

  15. [23]

    U. C. T¨ auber,Critical Dynamics(Cambridge University Press, Cambridge, 2014)

  16. [24]

    Zinn-Justin,Quantum Field Theory and Critical Phe- nomena(Oxford University Press, Oxford, 2002)

    J. Zinn-Justin,Quantum Field Theory and Critical Phe- nomena(Oxford University Press, Oxford, 2002)

  17. [25]

    Statistical dynamics of classical systems,

    P. C. Martin, E. D. Siggia, and H. A. Rose, “Statistical dynamics of classical systems,” Phys. Rev. A8, 423-437 (1973)

  18. [26]

    On a Lagrangian for classical field dy- namics and renormalization group calculations of dynam- ical critical properties,

    H. K. Janssen, “On a Lagrangian for classical field dy- namics and renormalization group calculations of dynam- ical critical properties,” Z Phyik B23, 377-380 (1976)

  19. [27]

    Techniques de renormalisation de la th´ eorie des champs et dynamique des ph´ enom` enes cri- tiques,

    C. De Dominicis, “Techniques de renormalisation de la th´ eorie des champs et dynamique des ph´ enom` enes cri- tiques,” J. Phys. Colloq.37, 247-253 (1976)

  20. [28]

    Thermodynamics with continuous information flow,

    J. M. Horowitz and M. Esposito, “Thermodynamics with continuous information flow,” Phys. Rev. X4, 031015 (2014)

  21. [29]

    Self-organized crit- icality: An explanation of 1/fnoise,

    P. Bak, C. Tang, and K. Wiesenfeld, “Self-organized crit- icality: An explanation of 1/fnoise,” Phys. Rev. Lett. 59, 381-384 (1987)

  22. [30]

    Self-organized criticality from pro- tected mean-field dynamics: Loop stability and in- ternal renormalization in reflective neural systems,

    B. G. Chae, “Self-organized criticality from pro- tected mean-field dynamics: Loop stability and in- ternal renormalization in reflective neural systems,” arXiv:2601.04450 (2026)

  23. [31]

    Bak,How Nature Works: The Science of Self- Organized Criticality(Springer, New York, 1996)

    P. Bak,How Nature Works: The Science of Self- Organized Criticality(Springer, New York, 1996)

  24. [32]

    H. J. Jensen,Self-Organized Criticality: Emergent Com- plex Behavior in Physical and Biological Systems(Cam- bridge University Press, Cambridge, 1998)

  25. [33]

    Paths to self-organized criticality,

    R. Dickman, M. A. Mu˜ noz, A. Vespignani, and S. Zap- peri, “Paths to self-organized criticality,” Braz. J. Phys. 30, 27-41 (2000)

  26. [34]

    Topological evolution of dy- namical networks: Global criticality from local dynam- ics,

    S. Bornholdt and T. Rohlf, “Topological evolution of dy- namical networks: Global criticality from local dynam- ics,” Phys. Rev. Lett.84, 6114-6116 (2000)

  27. [35]

    Phase transi- tions towards criticality in a neural system with adaptive interactions,

    A. Levina, J. M. Herrmann, and T. Geisel, “Phase transi- tions towards criticality in a neural system with adaptive interactions,” Phys. Rev. Lett.102, 118110 (2009)

  28. [36]

    Approximation of dy- namical systems by continuous time recurrent neural net- works,

    K. Funahashi and Y. Nakamura, “Approximation of dy- namical systems by continuous time recurrent neural net- works,” Neural Networks6, 801-806 (1993)

  29. [37]

    On the dynamics of small continuous-time recurrent neural networks,

    R. D. Beer, “On the dynamics of small continuous-time recurrent neural networks,” Adaptive Behavior3, 469- 509 (1995)

  30. [38]

    Liquid time-constant networks,

    R. Hasani, M. Lechner, A. Amini, D. Rus, and R. Grosu, “Liquid time-constant networks,” Proc. AAAI Conf. Ar- tif. Intell.35, 7657-7666 (2021)

  31. [39]

    Recursive dynamics in fast-weights home- ostatic reentry networks: Toward reflective intelligence,

    B. G. Chae, “Recursive dynamics in fast-weights home- ostatic reentry networks: Toward reflective intelligence,” arXiv:2511.06798 (2025)

  32. [40]

    Continuous-time homeostatic dynamics for reentrant inference models,

    B. G. Chae, “Continuous-time homeostatic dynamics for reentrant inference models,” arXiv:2512.05158 (2025)

  33. [41]

    Renormalization-group geometry of homeo- statically regulated reentry networks,

    B. G. Chae, “Renormalization-group geometry of homeo- statically regulated reentry networks,” arXiv:2512.19086 (2025)

  34. [42]

    F. W. Warner,Foundations of Differentiable Manifolds and Lie Groups(Springer-Verlag, Berlin, 1983)

  35. [43]

    Differential geometry of curved exponential families-curvatures and information loss,

    S. I. Amari, “Differential geometry of curved exponential families-curvatures and information loss,” Ann. Statist. 10, 357-385 (1982)

  36. [44]

    On the non-linear mechanics of hydrody- namic stability,

    J. T. Stuart, “On the non-linear mechanics of hydrody- namic stability,” J. Fluid Mech.4, 1-21 (1958)

  37. [45]

    Renormalization group approach to oscillator synchro- nization,

    O. Kogan, J. L. Rogers, M. C. Cross, and G. Refael, “Renormalization group approach to oscillator synchro- nization,” Phys. Rev. E80, 036206 (2009)

  38. [46]

    The Kuramoto model: A simple paradigm for synchronization phenomena,

    J. A. Acebr´ on, L. L. Bonilla, C. J. P´ erez Vicente, F. Ritort, and R. Spigler, “The Kuramoto model: A simple paradigm for synchronization phenomena,” Rev. Mod. Phys.77, 137-185 (2005). 15 Supplementary Materials Appendix A: Computation and Inference of the Time-Scale Densit...

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.