A one-layer softmax attention transformer trained by gradient descent provably converges to the one-nearest neighbor predictor and remains close to it under distribution shift.
Guidelines: • The answer NA means that the paper does not include theoretical results
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
A one-layer softmax attention transformer trained by gradient descent provably converges to the one-nearest neighbor predictor and remains close to it under distribution shift.