A residual connection that copies previous-layer attention values into the current layer improves in-context learning in transformers up to 1B parameters.
Attention approximates sparse distributed memory
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.NE 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Associative memory inspires improvements for in-context learning using a novel attention residual stream architecture
A residual connection that copies previous-layer attention values into the current layer improves in-context learning in transformers up to 1B parameters.