Affine recurrent networks cannot correct errors along state-separating subspaces and thus learn only finite-horizon state tracking that predictably fails when within-class spread exceeds initial between-class separation.
Transformers are rnns: Fast autoregressive transformers with linear attention
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.LG 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
Priming transfers knowledge from pre-trained Transformers to hybrid SSM-attention models, recovering performance with minimal additional tokens and showing Gated KalmaNet outperforming Mamba-2 on long-context reasoning at 32B scale.
citing papers explorer
-
Rethinking State Tracking in Recurrent Models Through Error Control Dynamics
Affine recurrent networks cannot correct errors along state-separating subspaces and thus learn only finite-horizon state tracking that predictably fails when within-class spread exceeds initial between-class separation.
-
Priming: Hybrid State Space Models From Pre-trained Transformers
Priming transfers knowledge from pre-trained Transformers to hybrid SSM-attention models, recovering performance with minimal additional tokens and showing Gated KalmaNet outperforming Mamba-2 on long-context reasoning at 32B scale.