Pith. sign in

REVIEW 7 cited by

MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.13111 v2 pith:HI7OBLFC submitted 2025-03-17 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords spatialunderstandingca-vqadatadepthestimationincludingmetric
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal large language models (MLLMs) excel at 2D visual understanding but remain limited in their ability to reason about 3D space. In this work, we leverage large-scale high-quality 3D scene data with open-set annotations to introduce 1) a novel supervised fine-tuning dataset and 2) a new evaluation benchmark, focused on indoor scenes. Our Cubify Anything VQA (CA-VQA) data covers diverse spatial tasks including spatial relationship prediction, metric size and distance estimation, and 3D grounding. We show that CA-VQA enables us to train MM-Spatial, a strong generalist MLLM that also achieves state-of-the-art performance on 3D spatial understanding benchmarks, including our own. We show how incorporating metric depth and multi-view inputs (provided in CA-VQA) can further improve 3D understanding, and demonstrate that data alone allows our model to achieve depth perception capabilities comparable to dedicated monocular depth estimation models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning

    cs.CV 2026-06 unverdicted novelty 8.0 of 10

    A blank-image ablation test reveals that high probe accuracy on VLM spatial reasoning frequently reflects priors or inverted signs rather than image grounding, with horizontal grounded, vertical prior, and depth inverted.

  2. GReFEM: Multimodal LLMs as Zero-Shot Semantic Assistants for Physics-Guided 3D Mesh Refinement

    cs.GR 2026-07 conditional novelty 6.5 of 10

    GReFEM shows MLLMs zero-shot isolate load-activated geometric features for volumetric mesh refinement with higher precision than matched-budget geometric heuristics.

  3. World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Distilling view-consistent future views and action-outcome supervision from a generative world model into a VLM via two-stage post-training improves dynamic spatial reasoning on SAT-Real, VSI-Bench and similar benchma...

  4. Multimodal Language Models Cannot Spot Spatial Inconsistencies

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Multimodal LLMs significantly underperform humans at spotting objects that break 3D consistency in multi-view image pairs.

  5. Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation

    cs.CV 2026-01 unverdicted novelty 6.0 of 10

    VLMs reach only 0.66 accuracy on relative camera pose estimation while humans achieve 0.91 and specialized pipelines reach 0.99, exposing weaknesses in multi-view spatial reasoning.

  6. SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Dense scene-graph-grounded rewards let a 7B multimodal LLM trained on 7K synthetic questions beat SFT and sparse-RL baselines and outscore GPT-4o on average across 12 spatial/real-world benchmarks.

  7. SpaCE: Rethinking Spatial Capacity and Generalization in Multi-Frame Multimodal Large Language Models

    eess.IV 2026-06 unverdicted novelty 5.0 of 10

    SpaCE derives four theoretical results on spatial capacity, sample complexity, generalization, and bias-variance trade-offs for multi-frame MLLM reasoning, validated on MultiSPA, CA-VQA, and SpatialRGPT.

Pith tools