MoSA, an expert-choice style sparse attention that selects per-head top-k tokens, outperforms dense transformers on C4 language modeling under matched FLOPs with up to 27% perplexity improvement and reduces wall-clock time, memory, and KV-cache size in perplexity-matched comparisons.
Hippo: Recurrent memory with optimal polynomial projections
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
MoSA, an expert-choice style sparse attention that selects per-head top-k tokens, outperforms dense transformers on C4 language modeling under matched FLOPs with up to 27% perplexity improvement and reduces wall-clock time, memory, and KV-cache size in perplexity-matched comparisons.