Pith. sign in

REVIEW 4 cited by

Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.03162 v2 pith:TUGJQA3I submitted 2022-04-07 cs.CV cs.CL

classification cs.CVcs.CL
keywords modelswinogrounddatasetlanguagevisio-linguisticvisioncaptionscompositional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a novel task and dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning, which we call Winoground. Given two images and two captions, the goal is to match them correctly - but crucially, both captions contain a completely identical set of words, only in a different order. The dataset was carefully hand-curated by expert annotators and is labeled with a rich set of fine-grained tags to assist in analyzing model performance. We probe a diverse range of state-of-the-art vision and language models and find that, surprisingly, none of them do much better than chance. Evidently, these models are not as skilled at visio-linguistic compositional reasoning as we might have hoped. We perform an extensive analysis to obtain insights into how future work might try to mitigate these models' shortcomings. We aim for Winoground to serve as a useful evaluation set for advancing the state of the art and driving further progress in the field. The dataset is available at https://huggingface.co/datasets/facebook/winoground.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs

    cs.CV 2026-05 conditional novelty 7.0 of 10

    Medical VLMs frequently select negated options that contradict visible chest X-ray findings, achieving only ~30% accuracy on direct presence probes, but a post-hoc consistency verifier raises accuracy above 95%.

  2. Autonomous VR-Based Risk Detection for Situational Awareness in Dangerous Settings

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A robot using GPT-4o labeled hazards in a simulated disaster room, and VR users preferred and rated these annotations highly, though the study lacks a controlled baseline comparison.

  3. Grounded Reinforcement Learning for Visual Reasoning

    cs.CV 2025-05 unverdicted novelty 6.0 of 10

    ViGoRL introduces visually grounded RL that anchors reasoning steps to image coordinates and uses multi-turn zooming to outperform standard RL and supervised baselines on spatial and GUI reasoning benchmarks.

  4. Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Current VLMs tend to label anomalous scenes as hazardous, and a four-way hazard/anomaly benchmark exposes this conflation more clearly than binary safe/unsafe tests.

Pith tools