REVIEW 8 cited by
VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Vision-Language Pretraining (VLP) models have recently successfully facilitated many cross-modal downstream tasks. Most existing works evaluated their systems by comparing the fine-tuned downstream task performance. However, only average downstream task accuracy provides little information about the pros and cons of each VLP method, let alone provides insights on how the community can improve the systems in the future. Inspired by the CheckList for testing natural language processing, we exploit VL-CheckList, a novel framework to understand the capabilities of VLP models. The proposed method divides the image-texting ability of a VLP model into three categories: objects, attributes, and relations, and uses a novel taxonomy to further break down these three aspects. We conduct comprehensive studies to analyze seven recently popular VLP models via the proposed framework. Results confirm the effectiveness of the proposed method by revealing fine-grained differences among the compared models that were not visible from downstream task-only evaluation. Further results show promising research direction in building better VLP models. Our data and code are available at: https://github.com/om-ai-lab/VL-CheckList.
Forward citations
Cited by 8 Pith papers
-
Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models
Three published hyperbolic vision-language models operate near-Euclidean with inoperative entailment cones and no detectable radial hierarchy, because the entailment objective admits a low-curvature shortcut.
-
EEG-EditBench: Probing Visual Information in EEG-Image Retrieval Models with Controlled Image Edits
EEG-image retrieval models that score high on standard 200-way ranking often fail to pick the viewed image over controlled edits that change one visual factor, with attribute changes (shape, color, texture) hardest fo...
-
Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models
On a new 65,464-item binary benchmark that swaps one medical term per caption, four medical multimodal models scored 49–62%, close to chance.
-
Can VLMs Reason Robustly? A Neuro-Symbolic Investigation
End-to-end fine-tuned VLMs fail to induce reasoning functions under object-count covariate shifts; VLC (VLM concepts + circuits) yields consistently higher OOD accuracy.
-
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets
Block-based diffusion generation of counterfactual image-text sets, combined with a set-aware loss, improves CLIP's compositional reasoning over several benchmarks, but the paper overstates one benchmark result and sh...
-
Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
A new 23-dimension benchmark finds that even the best VLMs score near random on motion trajectory, temporal extension, and several prediction tasks, far below humans, suggesting weak internal world models.
-
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
Blind text-only likelihood models match or exceed CLIP on many compositionality benchmarks because positives and negatives differ systematically in length, plausibility, or image style.
-
CF-VLM:CounterFactual Vision-Language Fine-tuning
CF-VLM fine-tunes VLMs on counterfactual image-text pairs with three objectives, reporting gains on compositional reasoning benchmarks and modest hallucination reductions.
Discussion (0). Sign in to comment.