MMInference speeds up long-context VLM prefill by up to 8.3x at 1M tokens using modality-aware permutation sparse attention while keeping accuracy close to full attention.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention
MMInference speeds up long-context VLM prefill by up to 8.3x at 1M tokens using modality-aware permutation sparse attention while keeping accuracy close to full attention.