Pith. sign in

REVIEW 8 cited by

VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.00221 v2 pith:3HHEZLGL submitted 2022-07-01 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords modelsdownstreammethodproposedvl-checklistattributesframeworkfurther
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-Language Pretraining (VLP) models have recently successfully facilitated many cross-modal downstream tasks. Most existing works evaluated their systems by comparing the fine-tuned downstream task performance. However, only average downstream task accuracy provides little information about the pros and cons of each VLP method, let alone provides insights on how the community can improve the systems in the future. Inspired by the CheckList for testing natural language processing, we exploit VL-CheckList, a novel framework to understand the capabilities of VLP models. The proposed method divides the image-texting ability of a VLP model into three categories: objects, attributes, and relations, and uses a novel taxonomy to further break down these three aspects. We conduct comprehensive studies to analyze seven recently popular VLP models via the proposed framework. Results confirm the effectiveness of the proposed method by revealing fine-grained differences among the compared models that were not visible from downstream task-only evaluation. Further results show promising research direction in building better VLP models. Our data and code are available at: https://github.com/om-ai-lab/VL-CheckList.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Three published hyperbolic vision-language models operate near-Euclidean with inoperative entailment cones and no detectable radial hierarchy, because the entailment objective admits a low-curvature shortcut.

  2. EEG-EditBench: Probing Visual Information in EEG-Image Retrieval Models with Controlled Image Edits

    cs.CV 2026-07 conditional novelty 6.0 of 10

    EEG-image retrieval models that score high on standard 200-way ranking often fail to pick the viewed image over controlled edits that change one visual factor, with attribute changes (shape, color, texture) hardest fo...

  3. Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    On a new 65,464-item binary benchmark that swaps one medical term per caption, four medical multimodal models scored 49–62%, close to chance.

  4. Can VLMs Reason Robustly? A Neuro-Symbolic Investigation

    cs.LG 2026-03 conditional novelty 6.0 of 10

    End-to-end fine-tuned VLMs fail to induce reasoning functions under object-count covariate shifts; VLC (VLM concepts + circuits) yields consistently higher OOD accuracy.

  5. A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Block-based diffusion generation of counterfactual image-text sets, combined with a set-aware loss, improves CLIP's compositional reasoning over several benchmarks, but the paper overstates one benchmark result and sh...

  6. Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 23-dimension benchmark finds that even the best VLMs score near random on motion trajectory, temporal extension, and several prediction tasks, far below humans, suggesting weak internal world models.

  7. A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Blind text-only likelihood models match or exceed CLIP on many compositionality benchmarks because positives and negatives differ systematically in length, plausibility, or image style.

  8. CF-VLM:CounterFactual Vision-Language Fine-tuning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    CF-VLM fine-tunes VLMs on counterfactual image-text pairs with three objectives, reporting gains on compositional reasoning benchmarks and modest hallucination reductions.

Pith tools