Pith. sign in

REVIEW 1 cited by

Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03766 v2 pith:YHJ7VMJH submitted 2023-12-05 cs.CL cs.CV

Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment

classification cs.CL cs.CV
keywords visualmodelstextualmisalignmentalignmentbinarycuratedexplanation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X LinkedIn Reddit HN
read the original abstract

While existing image-text alignment models reach high quality binary assessments, they fall short of pinpointing the exact source of misalignment. In this paper, we present a method to provide detailed textual and visual explanation of detected misalignments between text-image pairs. We leverage large language models and visual grounding models to automatically construct a training set that holds plausible misaligned captions for a given image and corresponding textual explanations and visual indicators. We also publish a new human curated test set comprising ground-truth textual and visual misalignment annotations. Empirical results show that fine-tuning vision language models on our training set enables them to articulate misalignments and visually indicate them within images, outperforming strong baselines both on the binary alignment classification and the explanation generation tasks. Our method code and human curated test set are available at: https://mismatch-quest.github.io/

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models

    cs.CV 2026-07 conditional novelty 6.0

    On a new 65,464-item binary benchmark that swaps one medical term per caption, four medical multimodal models scored 49–62%, close to chance.