Transformer residual layers are approximated as an explicit Euler scheme for a controlled hidden-state flow whose mean-field limit is a first-order transport control problem with Pontryagin terminal condition given by the softmax residual.
Stable architectures for deep neural networks
8 Pith papers cite this work, alongside 368 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
method 1polarities
use method 1representative citing papers
In every dimension d≥2 there exists a unique β_*^{(d)}>0 such that the uniform density on the sphere is the unique global minimizer of the USA free energy up to the linear-stability threshold K_# for β≤β_*, yielding a continuous transition, while for β>β_* the uniform density is not globally minimiz
A mean-field dynamical analysis of LoRA in transformers identifies phase transitions in catastrophic forgetting driven by perturbation norm and transformer depth.
NHODE framework learns partially observed dynamical systems by combining Hamiltonian neural networks with neural ODEs, enforcing energy conservation and improving long-horizon stability over data-driven baselines on mass-spring and three-body problems.
IRNO augments neural operators with learned fixed-point iterative refinement modules and a progressive spectral loss, achieving up to 56% error reduction on turbulent flow and large drops in high-frequency normalized errors on active matter.
Dynamic Mode Decomposition shows that short contiguous spans of Vision Transformer blocks can be approximated by a low-rank linear operator K with high predictive fidelity for p<=4 steps, but this approximation fails to outperform an identity baseline when propagated to the final layer.
Optimal depth-wise learning-rate scaling in deep scalar linear networks is data-dependent, so data-agnostic rules fail to transfer while the data-aware rule yields depth-independent linear convergence.
Effective depth, an operational count of sequential transformations, predicts CNN trainability better than nominal layer count because shortcuts and branches decouple the two.
citing papers explorer
-
A First-Order Mean Field Control Analysis of Transformer Layers under Cross-Entropy Training
Transformer residual layers are approximated as an explicit Euler scheme for a controlled hidden-state flow whose mean-field limit is a first-order transport control problem with Pontryagin terminal condition given by the softmax residual.
-
Phase transitions for the noisy transformer model in arbitrary dimension
In every dimension d≥2 there exists a unique β_*^{(d)}>0 such that the uniform density on the sphere is the unique global minimizer of the USA free energy up to the linear-stability threshold K_# for β≤β_*, yielding a continuous transition, while for β>β_* the uniform density is not globally minimiz
-
Understanding Catastrophic Forgetting In LoRA via Mean-Field Attention Dynamics
A mean-field dynamical analysis of LoRA in transformers identifies phase transitions in catastrophic forgetting driven by perturbation norm and transformer depth.
-
Learning partially observed systems with neural Hamiltonian ordinary differential equations
NHODE framework learns partially observed dynamical systems by combining Hamiltonian neural networks with neural ODEs, enforcing energy conservation and improving long-horizon stability over data-driven baselines on mass-spring and three-body problems.
-
Iterative Refinement Neural Operators are Learned Fixed-Point Solvers: A Principled Approach to Spectral Bias Mitigation
IRNO augments neural operators with learned fixed-point iterative refinement modules and a progressive spectral loss, achieving up to 56% error reduction on turbulent flow and large drops in high-frequency normalized errors on active matter.
-
Dynamic Mode Decomposition along Depth in Vision Transformers
Dynamic Mode Decomposition shows that short contiguous spans of Vision Transformer blocks can be approximated by a low-rank linear operator K with high predictive fidelity for p<=4 steps, but this approximation fails to outperform an identity baseline when propagated to the final layer.
-
Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks
Optimal depth-wise learning-rate scaling in deep scalar linear networks is data-dependent, so data-agnostic rules fail to transfer while the data-aware rule yields depth-independent linear convergence.
-
The Effective Depth Paradox: Evaluating the Relationship between Architectural Topology and Trainability in Deep CNNs
Effective depth, an operational count of sequential transformations, predicts CNN trainability better than nominal layer count because shortcuts and branches decouple the two.