REVIEW 5 cited by
A Mathematical Theory of Attention
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Attention is a powerful component of modern neural networks across a wide variety of domains. However, despite its ubiquity in machine learning, there is a gap in our understanding of attention from a theoretical point of view. We propose a framework to fill this gap by building a mathematically equivalent model of attention using measure theory. With this model, we are able to interpret self-attention as a system of self-interacting particles, we shed light on self-attention from a maximum entropy perspective, and we show that attention is actually Lipschitz-continuous (with an appropriate metric) under suitable assumptions. We then apply these insights to the problem of mis-specified input data; infinitely-deep, weight-sharing self-attention networks; and more general Lipschitz estimates for a specific type of attention studied in concurrent work.
Forward citations
Cited by 5 Pith papers
-
Principles of Lipschitz continuity in neural networks
A thesis deriving an SDE for how a network's spectral-norm Lipschitz bound changes under SGD, proving a non-negative noise-driven drift term, plus closed-form singular-value Hessians and a Shapley-based spectral robus...
-
Token Sample Complexity of Attention
Attention outputs converge to their infinite-token limit at sub-parametric rates n^−β (with β<1/2) governed by token covariance and attention matrices, and only logarithmically in the hardmax limit.
-
Change of Thought: Adaptive Test-Time Computation
A transformer layer that iteratively refines its attention matrix to a fixed point is claimed to improve accuracy with no extra parameters, but the benchmark evidence is not reproducible.
-
Physical models realizing the transformer architecture of large language models
The paper shows that any transformer's autoregressive token distribution can be realized as sequential measurements on a Fock space state, but the construction is an exact embedding with no new predictions.
-
Lipschitz Continuity in Deep Learning: A Systematic Review of Theoretical Foundations, Estimation Methods, Regularization Approaches, and Certifiable Robustness
A systematic survey of Lipschitz continuity in deep learning that corrects sigmoid (1/4) and softmax (1/2) Lipschitz constants and proves a sum-over-paths Lipschitz bound for additively-evaluated DAG networks.
Discussion (0). Continue with ORCID to comment.