Pith. sign in

REVIEW 2 cited by

VladVA: Discriminative Fine-tuning of LVLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.04378 v3 pith:EB2YPQIN submitted 2024-12-05 cs.CV cs.AI

classification cs.CVcs.AI
keywords discriminativemodelsvision-languageapproachimage-textlvlmstrainingcombine
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contrastively-trained Vision-Language Models (VLMs) like CLIP have become the de facto approach for discriminative vision-language representation learning. However, these models have limited language understanding, often exhibiting a "bag of words" behavior. At the same time, Large Vision-Language Models (LVLMs), which combine vision encoders with LLMs, have been shown to be capable of detailed vision-language reasoning, yet their autoregressive nature renders them less suitable for discriminative tasks. In this work, we propose to combine "the best of both worlds": a new training approach for discriminative fine-tuning of LVLMs that results in strong discriminative and compositional capabilities. Essentially, our approach converts a generative LVLM into a discriminative one, unlocking its capability for powerful image-text discrimination combined with enhanced language understanding. Our contributions include (1) a carefully designed training/optimization framework that utilizes image-text pairs of variable length and granularity for training the model with both contrastive and next-token prediction losses. This is accompanied by ablation studies that justify the necessity of our framework's components; (2) a parameter-efficient adaptation method using a combination of soft prompting and LoRA adapters; (3) significant improvements over state-of-the-art CLIP-like models of similar size, including standard image-text retrieval benchmarks and notable gains in compositionality.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CLaMR jointly encodes four video modalities in a vision-language model and uses token-level, per-modality matching to retrieve the right video and the right modality for a query.

  2. SORCE: Small Object Retrieval in Complex Environments

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new benchmark, SORCE-1K, tests text-to-image retrieval of small non-salient objects, and a multi-embedding MLLM method with regional prompts outperforms single-embedding baselines on it.

Pith tools