Transformers converge pathwise to a stochastic particle system and SPDE in the scaling limit, exhibiting synchronization by noise and exponential energy dissipation when common noise is coercive relative to self-attention drift.
Yuriiformer: A suite of nesterov- accelerated transformers
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
Optimizer-inspired Transformer architectures with momentum achieve lower validation loss than standard Transformers, with momentum identified as the key factor over preconditioning.
citing papers explorer
-
Stochastic Scaling Limits and Synchronization by Noise in Deep Transformer Models
Transformers converge pathwise to a stochastic particle system and SPDE in the scaling limit, exhibiting synchronization by noise and exponential energy dissipation when common noise is coercive relative to self-attention drift.
-
Momentum Streams for Optimizer-Inspired Transformers
Optimizer-inspired Transformer architectures with momentum achieve lower validation loss than standard Transformers, with momentum identified as the key factor over preconditioning.