Softmax attention with identity weights is the Bayes optimal denoiser for spherical data, and trained one-layer transformers learn such weights.
Learning patterns and pattern sequences by self-organizing nets of threshold elements
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
In-context denoising with one-layer transformers: connections between attention and associative memory retrieval
Softmax attention with identity weights is the Bayes optimal denoiser for spherical data, and trained one-layer transformers learn such weights.