Pith. sign in

REVIEW 19 cited by

SpatialBot: Precise Spatial Understanding with Vision Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.13642 v7 pith:PC3MW3L3 submitted 2024-06-19 cs.CV

SpatialBot: Precise Spatial Understanding with Vision Language Models

classification cs.CV
keywords understandingspatialspatialbotvlmsdepthembodiedlanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding, however they are still struggling with spatial understanding which is the foundation of Embodied AI. In this paper, we propose SpatialBot for better spatial understanding by feeding both RGB and depth images. Additionally, we have constructed the SpatialQA dataset, which involves multi-level depth-related questions to train VLMs for depth understanding. Finally, we present SpatialBench to comprehensively evaluate VLMs' capabilities in spatial understanding at different levels. Extensive experiments on our spatial-understanding benchmark, general VLM benchmarks and Embodied AI tasks, demonstrate the remarkable improvements of SpatialBot trained on SpatialQA. The model, code and data are available at https://github.com/BAAI-DCAI/SpatialBot.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning

    cs.CV 2026-06 unverdicted novelty 8.0

    A blank-image ablation test reveals that high probe accuracy on VLM spatial reasoning frequently reflects priors or inverted signs rather than image grounding, with horizontal grounded, vertical prior, and depth inverted.

  2. ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop

    cs.CV 2026-05 unverdicted novelty 7.0

    ESI-Bench shows active exploration outperforms passive observation in multimodal LLMs on spatial tasks but reveals failures from poor action choices and overconfident belief commitment unlike humans.

  3. 4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

    cs.CV 2026-05 unverdicted novelty 7.0

    4DThinker enables VLMs to perform dynamic spatial reasoning by thinking with 4D latent mental imagery using new fine-tuning and reinforcement learning methods.

  4. EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training

    cs.CV 2026-04 unverdicted novelty 7.0

    EmbodiedMidtrain mid-trains VLMs on curated VLA-aligned data subsets to improve downstream performance on robot manipulation benchmarks.

  5. MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images

    cs.CV 2025-11 conditional novelty 7.0

    MonoSR is a 1M-question benchmark for spatial reasoning from single photos across indoor, outdoor, and object-centric scenes; current VLMs score roughly 30-40%, and giving models 3D box coordinates lifts them near perfect.

  6. GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert

    cs.CV 2025-10 conditional novelty 7.0

    A frozen, large-scale pretrained diffusion policy converts sparse 3D waypoints from a VLM into dense robot actions, enabling zero-shot reuse of the action expert on new tasks and environments.

  7. Brick-Composer: Using MLLMs for Assembly with Diverse Bricks

    cs.AI 2026-06 unverdicted novelty 6.0

    Brick-Composer trains MLLMs on brick assembly via three signals, raising step-level success from under 1% to around 15% on the new BC-Bench benchmark.

  8. VLM3: Vision Language Models Are Native 3D Learners

    cs.CV 2026-05 unverdicted novelty 6.0

    Standard VLMs achieve expert-level 3D performance on depth estimation, pose estimation, and object understanding via three simple techniques without architecture changes or regression losses.

  9. ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop

    cs.CV 2026-05 unverdicted novelty 6.0

    ESI-Bench is a new benchmark for embodied spatial intelligence with 10 task categories on OmniGibson that requires agents to actively explore via perception, locomotion, and manipulation, revealing that MLLMs suffer f...

  10. 4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

    cs.CV 2026-05 unverdicted novelty 6.0

    4DThinker enables VLMs to perform dynamic spatial reasoning by internally simulating 4D imagery in latent space, outperforming prior text-based and modular approaches.

  11. Spatio-Temporal Grounding of Large Language Models from Perception Streams

    cs.RO 2026-04 unverdicted novelty 6.0

    FESTS uses Spatial Regular Expressions compiled from queries to generate 27k training tuples that raise a 3B-parameter LLM's frame-level F1 on spatio-temporal video reasoning from 48.5% to 87.5%, matching GPT-4.1 whil...

  12. SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding

    cs.RO 2025-11 unverdicted novelty 6.0

    SPEAR-1 combines a 3D-enriched VLM with embodied control to match or exceed existing robotic foundation models using 20 times fewer robot demonstrations.

  13. SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards

    cs.CV 2025-11 conditional novelty 6.0

    Dense scene-graph-grounded rewards let a 7B multimodal LLM trained on 7K synthetic questions beat SFT and sparse-RL baselines and outscore GPT-4o on average across 12 spatial/real-world benchmarks.

  14. Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation

    cs.RO 2025-08 conditional novelty 6.0

    Embodied-R1 uses a pointing-centric representation and reinforced fine-tuning on a 200K dataset to achieve state-of-the-art results on embodied benchmarks plus 56.2% success in SIMPLEREnv and 87.5% on real XArm tasks ...

  15. Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

    cs.CV 2024-12 unverdicted novelty 6.0

    MLLMs achieve competitive but subhuman performance on the new VSI-Bench for visual-spatial intelligence from videos, with spatial reasoning as the main bottleneck and explicit cognitive map generation improving distan...

  16. Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence

    cs.CV 2026-05 unverdicted novelty 5.0

    Integrating generative novel-view synthesis into LMM reasoning loops improves accuracy on spatial subtasks by 1.3 to 3.9 percentage points across multiple models and tasks.

  17. SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation

    cs.RO 2025-11 unverdicted novelty 5.0

    SlotVLA uses slot attention to model object-relation representations for multitask robotic manipulation, reducing visual tokens while achieving competitive generalization on the new LIBERO+ benchmark.

  18. LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

    cs.CV 2026-05 unverdicted novelty 4.0

    LLaVA-OV-2 uses codec-stream tokenization and a shared 3D RoPE to improve video, spatial, and tracking performance over Qwen3-VL-8B, while introducing the JumpScore benchmark for fine-grained motion localization.

  19. SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video

    cs.CV 2026-07 unverdicted novelty 3.0

    SpaceEra++ adds ScenePick frame sampling and SpaceAlign pairwise constraints to the prior SpaceEra system, claiming consistent benchmark gains for 3D video spatial reasoning.