Pith. sign in

REVIEW 6 cited by

Diffusion Feedback Helps CLIP See Better

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.20171 v4 pith:BBJD5Y4G submitted 2024-07-29 cs.CV

classification cs.CV
keywords clipvisualdiffusiondivamodelsmultimodalshortcomingscapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contrastive Language-Image Pre-training (CLIP), which excels at abstracting open-world representations across domains and modalities, has become a foundation for a variety of vision and multimodal tasks. However, recent studies reveal that CLIP has severe visual shortcomings, such as which can hardly distinguish orientation, quantity, color, structure, etc. These visual shortcomings also limit the perception capabilities of multimodal large language models (MLLMs) built on CLIP. The main reason could be that the image-text pairs used to train CLIP are inherently biased, due to the lack of the distinctiveness of the text and the diversity of images. In this work, we present a simple post-training approach for CLIP models, which largely overcomes its visual shortcomings via a self-supervised diffusion process. We introduce DIVA, which uses the DIffusion model as a Visual Assistant for CLIP. Specifically, DIVA leverages generative feedback from text-to-image diffusion models to optimize CLIP representations, with only images (without corresponding text). We demonstrate that DIVA improves CLIP's performance on the challenging MMVP-VLM benchmark which assesses fine-grained visual abilities to a large extent (e.g., 3-7%), and enhances the performance of MLLMs and vision models on multimodal understanding and segmentation tasks. Extensive evaluation on 29 image classification and retrieval benchmarks confirms that our framework preserves CLIP's strong zero-shot capabilities. The code is available at https://github.com/baaivision/DIVA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation

    cs.CV 2026-01 conditional novelty 6.0 of 10

    Adding depth- and segmentation-generation objectives to UMM post-training improved spatial understanding and reduced hallucinations on Harmon and OpenUni while preserving generation quality.

  2. Reconstruction Alignment Improves Unified Multimodal Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    RECA, a self-supervised post-training objective that conditions unified multimodal models on their own visual understanding embeddings to reconstruct input images, improves text-to-image and editing benchmarks across ...

  3. Visual Lexicon: Rich Image Features in Language Space

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A self-supervised method learns continuous image tokens in the text-embedding space of a frozen T2I diffusion model, capturing both semantics and visual details.

  4. Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Diffusion-conditioned denoising trains a video tokenizer whose features support both video question answering and text-to-video generation in a single LLM.

  5. DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DIR improves out-of-domain image captioning by guiding image features with a frozen diffusion model and retrieving text decomposed into objects, actions, and environments.

  6. Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An image autoencoder compresses pictures into caption-like embeddings via a frozen diffusion decoder, and a fine-tuned LLM reads those embeddings into captions claimed to rival GPT-4o at under $1,000 training cost.

Pith tools