Transformer residual layers are approximated as an explicit Euler scheme for a controlled hidden-state flow whose mean-field limit is a first-order transport control problem with Pontryagin terminal condition given by the softmax residual.
Trainability and accuracy of artificial neural networks: An interacting particle system approach
3 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Maximum entropy connectivity constrained by task moments and weight scale reproduces the qualitative and quantitative structure of gradient-trained networks across learning regimes.
For orthogonal inputs, gradient flow on shallow ReLU nets with MSE loss at small init converges to zero loss, exhibits min-variation-norm bias, initial alignment, and saddle-to-saddle dynamics.
citing papers explorer
-
A First-Order Mean Field Control Analysis of Transformer Layers under Cross-Entropy Training
Transformer residual layers are approximated as an explicit Euler scheme for a controlled hidden-state flow whose mean-field limit is a first-order transport control problem with Pontryagin terminal condition given by the softmax residual.
-
Balancing structure and randomness: maximum entropy networks for context-dependent computations
Maximum entropy connectivity constrained by task moments and weight scale reproduces the qualitative and quantitative structure of gradient-trained networks across learning regimes.
-
Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
For orthogonal inputs, gradient flow on shallow ReLU nets with MSE loss at small init converges to zero loss, exhibits min-variation-norm bias, initial alignment, and saddle-to-saddle dynamics.