REVIEW 9 cited by
A Mathematical Theory of Attention
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Attention is a powerful component of modern neural networks across a wide variety of domains. However, despite its ubiquity in machine learning, there is a gap in our understanding of attention from a theoretical point of view. We propose a framework to fill this gap by building a mathematically equivalent model of attention using measure theory. With this model, we are able to interpret self-attention as a system of self-interacting particles, we shed light on self-attention from a maximum entropy perspective, and we show that attention is actually Lipschitz-continuous (with an appropriate metric) under suitable assumptions. We then apply these insights to the problem of mis-specified input data; infinitely-deep, weight-sharing self-attention networks; and more general Lipschitz estimates for a specific type of attention studied in concurrent work.
Forward citations
Cited by 9 Pith papers
-
Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective
A Transformer with one attention head and ReLU-floor feedforward layers can approximate Hölder functions to accuracy ε with width O(ε^{-2/β}) and depth O(log 1/ε), avoiding the curse of dimensionality.
-
Towards understanding how attention mechanism works in deep learning
Self-attention is shown to approximate a drift-diffusion process on the data manifold, reducible to heat diffusion under a learned metric; metric-attention is proposed and outperforms self-attention in experiments.
-
Principles of Lipschitz continuity in neural networks
A thesis deriving an SDE for how a network's spectral-norm Lipschitz bound changes under SGD, proving a non-negative noise-driven drift term, plus closed-form singular-value Hessians and a Shapley-based spectral robus...
-
Token Sample Complexity of Attention
Attention outputs converge to their infinite-token limit at sub-parametric rates n^−β (with β<1/2) governed by token covariance and attention matrices, and only logarithmically in the hardmax limit.
-
A Unified Perspective on the Dynamics of Deep Transformers
Attention-only Transformer stacks are shown to be well-posed as mean-field PDEs for many attention variants, and Gaussian inputs evolve via explicit covariance ODEs that predict clustering or blow-up.
-
The Geometry of Tokens in Internal Representations of Large Language Models
Token-level intrinsic dimension of internal representations correlates with next-token cross-entropy loss across layers in three LLMs; higher-loss prompts live in higher-dimensional token manifolds.
-
Change of Thought: Adaptive Test-Time Computation
A transformer layer that iteratively refines its attention matrix to a fixed point is claimed to improve accuracy with no extra parameters, but the benchmark evidence is not reproducible.
-
Physical models realizing the transformer architecture of large language models
The paper shows that any transformer's autoregressive token distribution can be realized as sequential measurements on a Fock space state, but the construction is an exact embedding with no new predictions.
-
Lipschitz Continuity in Deep Learning: A Systematic Review of Theoretical Foundations, Estimation Methods, Regularization Approaches, and Certifiable Robustness
A systematic survey of Lipschitz continuity in deep learning that corrects sigmoid (1/4) and softmax (1/2) Lipschitz constants and proves a sum-over-paths Lipschitz bound for additively-evaluated DAG networks.
Discussion (0). Continue with ORCID to comment.