For uniform keys on the d-dimensional sphere, softmax attention becomes selective at inverse temperature scaling β_n* ≍ n^{2/(d-1)}, with explicit limiting laws for attention weights and outputs in each regime.
Advances in Neural Information Processing Systems , volume=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2years
2026 2representative citing papers
A hybrid least squares / gradient descent method accelerates MIONet training by exploiting multilinear structure in last-layer branch parameters via alternating least squares with Kronecker/Khatri-Rao factorization.
citing papers explorer
-
Scaling Limits of Long-Context Transformers
For uniform keys on the d-dimensional sphere, softmax attention becomes selective at inverse temperature scaling β_n* ≍ n^{2/(d-1)}, with explicit limiting laws for attention weights and outputs in each regime.
-
Hybrid Least Squares/Gradient Descent Methods for MIONets
A hybrid least squares / gradient descent method accelerates MIONet training by exploiting multilinear structure in last-layer branch parameters via alternating least squares with Kronecker/Khatri-Rao factorization.