Pith. sign in

REVIEW 16 cited by

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.12168 v1 pith:MV46K5N6 submitted 2024-01-22 cs.CV cs.CLcs.LGcs.RO

classification cs.CVcs.CLcs.LGcs.RO
keywords spatialreasoningdatatrainingcapabilityquantitativecapabilitiesfirst
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Understanding and reasoning about spatial relationships is a fundamental capability for Visual Question Answering (VQA) and robotics. While Vision Language Models (VLM) have demonstrated remarkable performance in certain VQA benchmarks, they still lack capabilities in 3D spatial reasoning, such as recognizing quantitative relationships of physical objects like distances or size differences. We hypothesize that VLMs' limited spatial reasoning capability is due to the lack of 3D spatial knowledge in training data and aim to solve this problem by training VLMs with Internet-scale spatial reasoning data. To this end, we present a system to facilitate this approach. We first develop an automatic 3D spatial VQA data generation framework that scales up to 2 billion VQA examples on 10 million real-world images. We then investigate various factors in the training recipe, including data quality, training pipeline, and VLM architecture. Our work features the first internet-scale 3D spatial reasoning dataset in metric space. By training a VLM on such data, we significantly enhance its ability on both qualitative and quantitative spatial VQA. Finally, we demonstrate that this VLM unlocks novel downstream applications in chain-of-thought spatial reasoning and robotics due to its quantitative estimation capability. Project website: https://spatial-vlm.github.io/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams

    cs.AI 2026-07 accept novelty 7.0 of 10

    VLMs recover common ERD elements at F1>0.74 but drop to 0.07–0.28 on N-ary relationships, multivalued attributes, and weak entities; reasoning models gain 15–25% yet stay prior- and complexity-sensitive.

  2. Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Variable-length autoregressive latent sequences, trained as variational inference with a PPO-style objective, give robot policies adaptive test-time compute and yield a reusable action tokenizer.

  3. SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning

    cs.RO 2026-03 conditional novelty 6.5 of 10

    A video-language model with per-timestep spatiotemporal CoT and dense progress prediction can serve as the sole reward for zero-shot online robot RL on 24 unseen manipulation tasks.

  4. PhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D Assets

    cs.RO 2026-08 conditional novelty 6.0 of 10

    PhysX-CoT turns single-image 3D asset generation into an explicit, ordered, supervised chain of physical states, beating an output-centric VLM baseline on geometry and physical attributes.

  5. Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A recalibrated GRPO reinforcement learning method lets multimodal LLMs say 'None' for nonexistent referring expressions without sacrificing localization accuracy on objects that do exist.

  6. SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation

    cs.RO 2025-09 conditional novelty 6.0 of 10

    SocialNav-SUB introduces a VQA benchmark for social robot navigation and shows current VLMs underperform rule-based and human-agreement baselines on spatial, spatiotemporal, and social reasoning questions.

  7. Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new contrastive real-image benchmark shows most vision-language models fail spatial relation tasks, while chain-of-thought reasoning models approach human-level accuracy.

  8. Sustainability assessment using multimodal AI agents

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A multi-agent AI system automatically builds life cycle inventories from public web data and estimates electronics carbon footprints within 19% of expert LCAs.

  9. Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new multi-image benchmark and graph-based scoring method show that MLLMs produce weak, poorly-grounded chains of thought on visual physics tasks, and that RL post-training can degrade spatial reasoning.

  10. Training Strategies for Efficient Embodied Reasoning

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Reasoning pre-training and dropout let robot policies benefit from chain-of-thought training without run-time reasoning, preserving most of the performance gains at standard VLA inference speeds.

  11. Learning the RoPEs: Better 2D and 3D Position Encodings with STRING

    cs.LG 2025-02 conditional novelty 6.0 of 10

    STRING parameterizes translation-invariant position encodings as exponentials of commuting skew-symmetric generators, proving they are exactly RoPE in a learnable orthogonal basis, and shows practical gains in 2D/3D v...

  12. LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LLaVA-SpaceSGG is a visual instruction-tuned model that uses a new 2D+3D spatial scene graph dataset to improve open-vocabulary scene graph generation and spatial relation accuracy.

  13. Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A probabilistic overlap measure for object positions yields a human-aligned spatial relationship metric and a training-free generation guidance method for text-to-image models.

  14. Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A mixture-of-experts diffusion policy conditioned on object, pose, depth, and trajectory mid-level representations is reported to outperform language-only and representation-free baselines on bimanual dexterous tasks,...

  15. Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Scene-graph-based chain-of-thought prompting and GRPO training improve spatial reasoning accuracy in vision-language models, and GRPO degrades less than supervised fine-tuning when question wording is flipped.

  16. Explainability for Vision Foundation Models: A Survey

    cs.CV 2025-01 conditional novelty 4.0 of 10

    A structured review of 122 papers on explainability for vision foundation models, with a taxonomy and the finding that quantitative evaluation is rare (36%).

Pith tools