Semantic geometry emerges transiently early in next-token prediction training before collapsing to Neural Collapse symmetry in synthetic settings with latent semantic factors.
How transformers implement induction heads: Approxima- tion and optimization analysis.arXiv preprint arXiv:2410.11474, 2024a
5 Pith papers cite this work. Polarity classification is still indexing.
years
2026 5representative citing papers
Bayesian reduction of attention posterior on copy task predicts first-order phase transition for softmax attention and second-order followed by crossover for linear attention.
Muon outperforms Adam by reducing curvature penalty via lower Normalized Directional Sharpness, as shown via Taylor approximation on LLM training and proven on stylized quadratic problems with heterogeneous curvature.
Transformer weights at early training stages are closed-form compositions of bigram, token-interchangeability, and context mappings that directly reflect text-corpus statistics and explain the emergence of semantic associations.
Mean-field equations for attention retrieval, teacher alignment, and logic overlap quantitatively match simulations and predict a sharp accuracy transition in a solvable transformer for permutation state tracking.
citing papers explorer
-
Structure Before Collapse: Transient semantic geometry in next-token prediction
Semantic geometry emerges transiently early in next-token prediction training before collapsing to Neural Collapse symmetry in synthetic settings with latent semantic factors.
-
Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence
Bayesian reduction of attention posterior on copy task predicts first-order phase transition for softmax attention and second-order followed by crossover for linear attention.
-
Why Muon Outperforms Adam: A Curvature Perspective
Muon outperforms Adam by reducing curvature penalty via lower Normalized Directional Sharpness, as shown via Taylor approximation on LLM training and proven on stylized quadratic problems with heterogeneous curvature.
-
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
Transformer weights at early training stages are closed-form compositions of bigram, token-interchangeability, and context mappings that directly reflect text-corpus statistics and explain the emergence of semantic associations.
-
Learning Dynamics of Chain-of-Thought State Tracking in a Solvable Transformer Model
Mean-field equations for attention retrieval, teacher alignment, and logic overlap quantitatively match simulations and predict a sharp accuracy transition in a solvable transformer for permutation state tracking.