Selecting tokens that lie inside multiple anchor-centric semantic grain regions improves multimodal VQA/AVQA accuracy under latency and erasure constraints compared with pairwise attention-based selection.
Matryoshka query transformer for large vision-language m odels,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
other 1
citation-polarity summary
fields
eess.SP 1years
2026 1verdicts
CONDITIONAL 1roles
other 1polarities
unclear 1representative citing papers
citing papers explorer
-
Geometric Cross-Modal Token Selection for Latency-Constrained Multimodal Token Communication
Selecting tokens that lie inside multiple anchor-centric semantic grain regions improves multimodal VQA/AVQA accuracy under latency and erasure constraints compared with pairwise attention-based selection.