Pith. sign in

REVIEW 4 cited by

FG-CLIP: Fine-Grained Visual and Textual Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.05071 v3 pith:SYCFQ3R2 submitted 2025-05-08 cs.CV cs.AI

classification cs.CVcs.AI
keywords fine-grainedfg-clipclipmillionmultimodalunderstandingcaptionscapturing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short captions. To address this, we propose Fine-Grained CLIP (FG-CLIP), which enhances fine-grained understanding through three key innovations. First, we leverage large multimodal models to generate 1.6 billion long caption-image pairs for capturing global-level semantic details. Second, a high-quality dataset is constructed with 12 million images and 40 million region-specific bounding boxes aligned with detailed captions to ensure precise, context-rich representations. Third, 10 million hard fine-grained negative samples are incorporated to improve the model's ability to distinguish subtle semantic differences. We construct a comprehensive dataset, termed FineHARD, by integrating high-quality region-specific annotations with hard fine-grained negative samples. Corresponding training methods are meticulously designed for these data. Extensive experiments demonstrate that FG-CLIP outperforms the original CLIP and other state-of-the-art methods across various downstream tasks, including fine-grained understanding, open-vocabulary object detection, image-text retrieval, and general multimodal benchmarks. These results highlight FG-CLIP's effectiveness in capturing fine-grained image details and improving overall model performance. The data, code, and models are available at https://github.com/360CVGroup/FG-CLIP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DialogueVPR: Towards Conversational Visual Place Recognition

    cs.AI 2026-05 conditional novelty 6.0 of 10

    Dialogue-based place recognition lets an AI localize a place by asking clarifying questions, trained and evaluated on a GPT-4o-generated benchmark built from street-view images.

  2. SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new reference-free metric, SPECS, fine-tunes LongCLIP with a specificity objective and reaches LLM-level human correlation on long captions at a fraction of the computational cost.

  3. OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    OpenSeg-R uses an LMM's step-by-step visual explanations as extra text prompts to improve open-vocabulary segmentation masks.

  4. Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning

    cs.IR 2026-03 conditional novelty 5.0 of 10

    CoCoA forces an MLLM to reconstruct masked text through a single EOS token, improving multimodal embedding quality on MMEB-V1 and matching MoCa at 3B with far less pretraining data.

Pith tools