Training only the text encoder on multi-sentence paragraphs improves long-description image retrieval by over 14 points on DOCCI and removes the need for context-window architecture tricks.
In: Proceedings of the International Conference on Machine Learning (ICML) (2024)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval
Training only the text encoder on multi-sentence paragraphs improves long-description image retrieval by over 14 points on DOCCI and removes the need for context-window architecture tricks.