Training only the text encoder on multi-sentence paragraphs improves long-description image retrieval by over 14 points on DOCCI and removes the need for context-window architecture tricks.
In: International Con- ference on Machine Learning (ICML)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval
Training only the text encoder on multi-sentence paragraphs improves long-description image retrieval by over 14 points on DOCCI and removes the need for context-window architecture tricks.