Pith. sign in

REVIEW 6 cited by

Revisiting the Role of Language Priors in Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.01879 v4 pith:GVUBWLVX submitted 2023-06-02 cs.CV cs.AIcs.CL

Revisiting the Role of Language Priors in Vision-Language Models

classification cs.CV cs.AIcs.CL
keywords vision-languageaccuracybenchmarksgenerativeimagelanguageprobabilisticretrieval
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Vision-language models (VLMs) are impactful in part because they can be applied to a variety of visual understanding tasks in a zero-shot fashion, without any fine-tuning. We study $\textit{generative VLMs}$ that are trained for next-word generation given an image. We explore their zero-shot performance on the illustrative task of image-text retrieval across 8 popular vision-language benchmarks. Our first observation is that they can be repurposed for discriminative tasks (such as image-text retrieval) by simply computing the match score of generating a particular text string given an image. We call this probabilistic score the $\textit{Visual Generative Pre-Training Score}$ (VisualGPTScore). While the VisualGPTScore produces near-perfect accuracy on some retrieval benchmarks, it yields poor accuracy on others. We analyze this behavior through a probabilistic lens, pointing out that some benchmarks inadvertently capture unnatural language distributions by creating adversarial but unlikely text captions. In fact, we demonstrate that even a "blind" language model that ignores any image evidence can sometimes outperform all prior art, reminiscent of similar challenges faced by the visual-question answering (VQA) community many years ago. We derive a probabilistic post-processing scheme that controls for the amount of linguistic bias in generative VLMs at test time without having to retrain or fine-tune the model. We show that the VisualGPTScore, when appropriately debiased, is a strong zero-shot baseline for vision-language understanding, oftentimes producing state-of-the-art accuracy.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals

    cs.CV 2026-05 unverdicted novelty 7.0

    SDGBiasBench reveals intrinsic SDG biases in VLMs driven by priors rather than evidence, and CADE mitigates them with up to 25% accuracy gains and 12-point MAE reductions.

  2. Prior Bias in Vision Language Models on UML Diagram Interpretation

    cs.CV 2026-07 conditional novelty 6.5

    Reversing only the UML relation arrow while keeping class names and layout fixed cuts open-source VLM relation accuracy by about 33%, revealing prior-over-vision bias.

  3. Scalable Visual Pretraining for Language Intelligence

    cs.CV 2026-07 unverdicted novelty 6.0

    Unsupervised visual pretraining on raw document images consistently outperforms text-only pretraining on the same corpora for language intelligence.

  4. Building a Precise Video Language with Human-AI Oversight

    cs.CV 2026-04 unverdicted novelty 6.0

    CHAI framework pairs AI pre-captions with expert human critiques to produce precise video descriptions, enabling open models to outperform closed ones like Gemini-3.1-Pro and improve fine-grained control in video gene...

  5. TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models

    cs.CV 2025-09 conditional novelty 6.0

    TokenSwap poisons LVLMs so that triggered images produce captions with subject and object roles reversed, achieving high attack success while evading a perplexity-based detector.

  6. Scalable Visual Pretraining for Language Intelligence

    cs.CV 2026-07 conditional novelty 5.0

    Visual pretraining on rendered document pages beats text-only continued pretraining on scientific reasoning benchmarks across multiple model backbones, using about a quarter of the tokens.