Pith. sign in

REVIEW 2 cited by

Vision Language Models See What You Want but not What You See

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.00324 v6 pith:RH2YDBGQ submitted 2024-10-01 cs.AI

classification cs.AI
keywords othersvlmsabilitiescognitiveintelligenceintentionalitylanguagelevel-2
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Knowing others' intentions and taking others' perspectives are two core components of human intelligence that are considered to be instantiations of theory-of-mind. Infiltrating machines with these abilities is an important step towards building human-level artificial intelligence. Here, to investigate intentionality understanding and level-2 perspective-taking in Vision Language Models (VLMs), we constructed the IntentBench and PerspectBench, which together contains over 300 cognitive experiments grounded in real-world scenarios and classic cognitive tasks. We found VLMs achieving high performance on intentionality understanding but low performance on level-2 perspective-taking. This suggests a potential dissociation between simulation-based and theory-based theory-of-mind abilities in VLMs, highlighting the concern that they are not capable of using model-based reasoning to infer others' mental states.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Egocentric Bias in Vision-Language Models

    cs.CV 2026-02 conditional novelty 6.0 of 10

    Most vision-language models fail Level-2 visual perspective taking: they report the camera's view rather than the 180°-rotated string, even though they often recognize that another agent sees differently.

  2. Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds

    cs.AI 2025-05 conditional novelty 3.0 of 10

    The paper releases a procedural synthetic dataset of single-cube scenes with ground-truth pose matrices as a foundation for training VLMs in spatial reasoning.

Pith tools