Among 350M models trained for 15B tokens, Kimi Delta Attention with Muon has the best validation loss, pure Gated DeltaNet is fastest, and Cross-Layer Value Routing modestly lowers loss for DeltaNet-style memories.
Value Residual Learning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing
Among 350M models trained for 15B tokens, Kimi Delta Attention with Muon has the best validation loss, pure Gated DeltaNet is fastest, and Cross-Layer Value Routing modestly lowers loss for DeltaNet-style memories.