Pith. sign in

REVIEW 5 cited by

Visuospatial Cognitive Assistant

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.12312 v4 pith:JPEG5GFY submitted 2025-05-18 cs.CV cs.AIcs.CLcs.LGcs.RO

classification cs.CVcs.AIcs.CLcs.LGcs.RO
keywords reasoningvisuospatialassistantcognitivedatasetmodelsscannetspatial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video-based spatial cognition is vital for robotics and embodied AI but challenges current Vision-Language Models (VLMs). This paper makes two key contributions. First, we introduce ViCA (Visuospatial Cognitive Assistant)-322K, a diverse dataset of 322,003 QA pairs from real-world indoor videos (ARKitScenes, ScanNet, ScanNet++), offering supervision for 3D metadata-grounded queries and video-based complex reasoning. Second, we develop ViCA-7B, fine-tuned on ViCA-322K, which achieves new state-of-the-art on all eight VSI-Bench tasks, outperforming existing models, including larger ones (e.g., +26.1 on Absolute Distance). For interpretability, we present ViCA-Thinking-2.68K, a dataset with explicit reasoning chains, and fine-tune ViCA-7B to create ViCA-7B-Thinking, a model that articulates its spatial reasoning. Our work highlights the importance of targeted data and suggests paths for improved temporal-spatial modeling. We release all resources to foster research in robust visuospatial intelligence.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    A closed-loop self-evolving training system for spatial reasoning in MLLMs that iteratively generates QA pairs matched to the model's current capabilities via confidence feedback, achieving gains with an order of magn...

  2. SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Adaptive manifold keyframe sampling plus an instruction-pose-aware geometry MoE raises sparse-RGB 3D spatial reasoning to 63.5 average on VSI-Bench, beating strong baselines by 7.8 points.

  3. OneCanvas: 3D Scene Understanding via Panoramic Reprojection

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    OneCanvas aggregates multi-view 3D patches onto one panoramic canvas with continuous angular placement and 3D embeddings, enabling pretrained VLMs to achieve SOTA on SQA3D and VSI-Bench with an order of magnitude less...

  4. Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning

    cs.CV 2026-06 conditional novelty 6.0 of 10

    A self-evolving training loop that generates its own spatial QA data with executable code and difficulty feedback lifts Qwen3-VL-4B/8B to 62.7/63.3 on VSI-Bench using an order of magnitude less data.

  5. GEM: Generative Supervision Helps Embodied Intelligence

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    GEM adds generative depth supervision to VLM pre-training and reports improved results on embodied benchmarks plus real-world robot execution.

Pith tools