Pith. sign in

REVIEW 2 cited by

Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.06721 v1 pith:IRFT7PVW submitted 2024-08-13 cs.CV

classification cs.CV
keywords modelsvisualvlmsbasiclackingobjectobservationsperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-Language Models (VLMs) have emerged as general purpose tools for addressing a variety of complex computer vision problems. Such models have been shown to be highly capable, but, at the same time, also lacking some basic visual understanding skills. In this paper, we set out to understand the limitations of SoTA VLMs on fundamental visual tasks: object classification, understanding spatial arrangement, and ability to delineate individual object instances (through counting), by constructing a series of tests that probe which components of design, specifically, maybe lacking. Importantly, we go significantly beyond the current benchmarks, that simply measure final performance of VLM, by also comparing and contrasting it to performance of probes trained directly on features obtained from visual encoder (image embeddings), as well as intermediate vision-language projection used to bridge image-encoder and LLM-decoder ouput in many SoTA models (e.g., LLaVA, BLIP, InstructBLIP). In doing so, we uncover nascent shortcomings in VLMs response and make a number of important observations which could help train and develop more effective VLM models in future.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Automated, Scalable Machine Learning Model Inversion Assessment Pipeline

    cs.CR 2025-09 conditional novelty 6.0 of 10

    A modular pipeline automatically applies model inversion attacks and uses vision-language models to score four privacy-loss dimensions, producing a weighted composite risk score for image classifiers.

  2. Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A retrieval-based interleaved visual chain-of-thought method, RIV-CoT, improves VLM answer accuracy by 3.1 points and reasoning accuracy by 4.6 points on a new driving theory VQA benchmark.

Pith tools