Transformer residual layers are approximated as an explicit Euler scheme for a controlled hidden-state flow whose mean-field limit is a first-order transport control problem with Pontryagin terminal condition given by the softmax residual.
Trainability and accuracy of artificial neural networks: An interacting particle system approach
3 Pith papers cite this work, alongside 53 external citations. Polarity classification is still indexing.
3
Pith papers citing it
53
external citations · OpenAlex
representative citing papers
Maximum entropy connectivity constrained by task moments and weight scale reproduces the qualitative and quantitative structure of gradient-trained networks across learning regimes.
For orthogonal inputs, gradient flow on shallow ReLU nets with MSE loss at small init converges to zero loss, exhibits min-variation-norm bias, initial alignment, and saddle-to-saddle dynamics.
citing papers explorer
-
Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
For orthogonal inputs, gradient flow on shallow ReLU nets with MSE loss at small init converges to zero loss, exhibits min-variation-norm bias, initial alignment, and saddle-to-saddle dynamics.