A Transformer variant called Converter replaces self-attention with a learnable unitary spectral filter (Synvolution/Kernelution) and reports much higher accuracy on Long-Range Arena and other long-sequence tasks, without releasing code or variance estimates.
Attention is not all you need: pure attention loses rank doubly exponentially with depth
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Converting Transformers into DGNNs Form
A Transformer variant called Converter replaces self-attention with a learnable unitary spectral filter (Synvolution/Kernelution) and reports much higher accuracy on Long-Range Arena and other long-sequence tasks, without releasing code or variance estimates.