LV-XAttn speeds up distributed cross-attention in multimodal LLMs by keeping large visual KV blocks local and rotating small query blocks among GPUs, achieving up to 10.62x wall-clock speedups over Ring Attention on long-video workloads.
Adapting language models to compress contexts
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models
LV-XAttn speeds up distributed cross-attention in multimodal LLMs by keeping large visual KV blocks local and rotating small query blocks among GPUs, achieving up to 10.62x wall-clock speedups over Ring Attention on long-video workloads.