Effective rank of the frozen query-key kernel W_K^T W_Q identifies retrieval heads, enabling training-free sparse attention at 50% sparsity with small accuracy loss.
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads , url =
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry
Effective rank of the frozen query-key kernel W_K^T W_Q identifies retrieval heads, enabling training-free sparse attention at 50% sparsity with small accuracy loss.