Delta-rule linear attention faithfully approximates softmax attention through key-dependent rank-1 projections, enabling efficient post-hoc linearization of LLMs up to 32B parameters.
Parallelizing linear transformers with the delta rule over sequence length.Advances in neural information processing systems, 37:115491–115522
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
MSA is an end-to-end trainable memory model using sparse attention and document-wise RoPE that scales to 100M tokens with linear complexity and less than 9% degradation.
Reasoning-token augmentation dominates architectural bias for state-based recall tasks; hybrid advantages are narrow and task-dependent rather than uniform.
citing papers explorer
-
The Key to Going Linear: Analysis-Driven Transformer Linearization
Delta-rule linear attention faithfully approximates softmax attention through key-dependent rank-1 projections, enabling efficient post-hoc linearization of LLMs up to 32B parameters.
-
MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
MSA is an end-to-end trainable memory model using sparse attention and document-wise RoPE that scales to 100M tokens with linear complexity and less than 9% degradation.
-
Reasoning Primitives in Hybrid and Non-Hybrid LLMs: Do Architectural Differences Yield Advantages in State-Tracking and Recall?
Reasoning-token augmentation dominates architectural bias for state-based recall tasks; hybrid advantages are narrow and task-dependent rather than uniform.