A blank-image ablation test reveals that high probe accuracy on VLM spatial reasoning frequently reflects priors or inverted signs rather than image grounding, with horizontal grounded, vertical prior, and depth inverted.
Mm-spatial: Exploring 3d spatial understanding in multimodal llms
5 Pith papers cite this work. Polarity classification is still indexing.
years
2026 5verdicts
UNVERDICTED 5representative citing papers
Distilling view-consistent future views and action-outcome supervision from a generative world model into a VLM via two-stage post-training improves dynamic spatial reasoning on SAT-Real, VSI-Bench and similar benchmarks while avoiding test-time world-model cost.
Multimodal LLMs significantly underperform humans at spotting objects that break 3D consistency in multi-view image pairs.
VLMs reach only 0.66 accuracy on relative camera pose estimation while humans achieve 0.91 and specialized pipelines reach 0.99, exposing weaknesses in multi-view spatial reasoning.
SpaCE derives four theoretical results on spatial capacity, sample complexity, generalization, and bias-variance trade-offs for multi-frame MLLM reasoning, validated on MultiSPA, CA-VQA, and SpatialRGPT.
citing papers explorer
-
Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning
A blank-image ablation test reveals that high probe accuracy on VLM spatial reasoning frequently reflects priors or inverted signs rather than image grounding, with horizontal grounded, vertical prior, and depth inverted.
-
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
Distilling view-consistent future views and action-outcome supervision from a generative world model into a VLM via two-stage post-training improves dynamic spatial reasoning on SAT-Real, VSI-Bench and similar benchmarks while avoiding test-time world-model cost.
-
Multimodal Language Models Cannot Spot Spatial Inconsistencies
Multimodal LLMs significantly underperform humans at spotting objects that break 3D consistency in multi-view image pairs.
-
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
VLMs reach only 0.66 accuracy on relative camera pose estimation while humans achieve 0.91 and specialized pipelines reach 0.99, exposing weaknesses in multi-view spatial reasoning.
-
SpaCE: Rethinking Spatial Capacity and Generalization in Multi-Frame Multimodal Large Language Models
SpaCE derives four theoretical results on spatial capacity, sample complexity, generalization, and bias-variance trade-offs for multi-frame MLLM reasoning, validated on MultiSPA, CA-VQA, and SpatialRGPT.