REVIEW 4 cited by
LoTLIP: Improving Language-Image Pre-training for Long Text Understanding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Understanding long text is of great demands in practice but beyond the reach of most language-image pre-training (LIP) models. In this work, we empirically confirm that the key reason causing such an issue is that the training images are usually paired with short captions, leaving certain tokens easily overshadowed by salient tokens. Towards this problem, our initial attempt is to relabel the data with long captions, however, directly learning with which may lead to performance degradation in understanding short text (e.g., in the image classification task). Then, after incorporating corner tokens to aggregate diverse textual information, we manage to help the model catch up to its original level of short text understanding yet greatly enhance its capability of long text understanding. We further look into whether the model can continuously benefit from longer captions and notice a clear trade-off between the performance and the efficiency. Finally, we validate the effectiveness of our approach using a self-constructed large-scale dataset, which consists of 100M long caption oriented text-image pairs. Our method demonstrates superior performance in long-text-image retrieval tasks. The project page is available at https://wuw2019.github.io/lot-lip.
Forward citations
Cited by 4 Pith papers
-
FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
A dual-branch CLIP training pipeline with regional prompts and hierarchical feature alignment reaches state-of-the-art on long- and short-text retrieval.
-
FG-CLIP: Fine-Grained Visual and Textual Alignment
FG-CLIP, a CLIP variant trained with long captions, region-text pairs, and hard negatives, sets new state-of-the-art results on FG-OVD and several retrieval, detection, and multimodal benchmarks.
-
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
A TD-trained vision value model guides sentence-level inference-time search in VLMs, cutting hallucination and improving caption quality, with self-training gains on nine benchmarks.
-
FLAIR: VLM with Fine-grained Language-informed Image Representations
A vision-language model that pools image tokens using the text as a query achieves state-of-the-art fine-grained retrieval and segmentation with 30M training images.
Discussion (0). Continue with ORCID to comment.