Pith. sign in

REVIEW 2 cited by

What's left can't be right -- The remaining positional incompetence of contrastive vision-language models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.11477 v1 pith:TTBYZLUU submitted 2023-11-20 cs.CV cs.CL

classification cs.CVcs.CL
keywords relationscontrastivedatasetsleft-rightmodelspositionalvision-languageanalysing
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Contrastive vision-language models like CLIP have been found to lack spatial understanding capabilities. In this paper we discuss the possible causes of this phenomenon by analysing both datasets and embedding space. By focusing on simple left-right positional relations, we show that this behaviour is entirely predictable, even with large-scale datasets, demonstrate that these relations can be taught using synthetic data and show that this approach can generalise well to natural images - improving the performance on left-right relations on Visual Genome Relations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new contrastive real-image benchmark shows most vision-language models fail spatial relation tasks, while chain-of-thought reasoning models approach human-level accuracy.

  2. Metamorphic Testing for Pose Estimation Systems

    cs.SE 2025-02 conditional novelty 5.0 of 10

    MET-POSE uses metamorphic rules to test pose-estimation systems without ground-truth labels, and on Mediapipe Holistic it detects faults at similar or higher rates than classic labeled testing.

Pith tools