Pith. sign in

REVIEW 1 cited by

Understanding Depth and Height Perception in Large Visual-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.11748 v5 pith:AWBDWVU4 submitted 2024-08-21 cs.CV

classification cs.CV
keywords depthheightgeometricperceptionunderstandingmodelstheyvlms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Geometric understanding - including depth and height perception - is fundamental to intelligence and crucial for navigating our environment. Despite the impressive capabilities of large Vision Language Models (VLMs), it remains unclear how well they possess the geometric understanding required for practical applications in visual perception. In this work, we focus on evaluating the geometric understanding of these models, specifically targeting their ability to perceive the depth and height of objects in an image. To address this, we introduce GeoMeter, a suite of benchmark datasets - encompassing 2D and 3D scenarios - to rigorously evaluate these aspects. By benchmarking 18 state-of-the-art VLMs, we found that although they excel in perceiving basic geometric properties like shape and size, they consistently struggle with depth and height perception. Our analysis reveal that these challenges stem from shortcomings in their depth and height reasoning capabilities and inherent biases. This study aims to pave the way for developing VLMs with enhanced geometric understanding by emphasizing depth and height perception as critical components necessary for real-world applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DepthCues: Evaluating Monocular Depth Perception in Large Vision Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A six-task benchmark shows newer large vision models encode human-like monocular depth cues, and cue understanding strongly correlates with their depth estimation performance.

Pith tools