A top-k attention mechanism backed by CPU vector search allows million-token LLM contexts to run on a 16GB GPU while keeping over 95% of dense-attention performance, though the 2% sparsity claim does not hold at the longest context lengths.
Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs
A top-k attention mechanism backed by CPU vector search allows million-token LLM contexts to run on a 16GB GPU while keeping over 95% of dense-attention performance, though the 2% sparsity claim does not hold at the longest context lengths.