Pith. sign in

REVIEW 2 cited by

Multi-modal Situated Reasoning in 3D Scenes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.02389 v2 pith:B3RCSLPZ submitted 2024-09-04 cs.CV cs.AIcs.RO

classification cs.CVcs.AIcs.RO
keywords situatedmulti-modalreasoningmodelsmsqascenesquestionsituation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Situation awareness is essential for understanding and reasoning about 3D scenes in embodied AI agents. However, existing datasets and benchmarks for situated understanding are limited in data modality, diversity, scale, and task scope. To address these limitations, we propose Multi-modal Situated Question Answering (MSQA), a large-scale multi-modal situated reasoning dataset, scalably collected leveraging 3D scene graphs and vision-language models (VLMs) across a diverse range of real-world 3D scenes. MSQA includes 251K situated question-answering pairs across 9 distinct question categories, covering complex scenarios within 3D scenes. We introduce a novel interleaved multi-modal input setting in our benchmark to provide text, image, and point cloud for situation and question description, resolving ambiguity in previous single-modality convention (e.g., text). Additionally, we devise the Multi-modal Situated Next-step Navigation (MSNN) benchmark to evaluate models' situated reasoning for navigation. Comprehensive evaluations on MSQA and MSNN highlight the limitations of existing vision-language models and underscore the importance of handling multi-modal interleaved inputs and situation modeling. Experiments on data scaling and cross-domain transfer further demonstrate the efficacy of leveraging MSQA as a pre-training dataset for developing more powerful situated reasoning models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hypo3D: Exploring Hypothetical Reasoning in 3D

    cs.CV 2025-02 conditional novelty 7.0 of 10

    Hypo3D, a 3D VQA benchmark with 14,885 QAs over 700 indoor scenes, tests whether models can reason about imagined scene changes and shows they fall far behind humans.

  2. Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

    cs.CV 2025-02 conditional novelty 4.0 of 10

    A structured survey of 3D Scene Question Answering that categorizes datasets, methods, and metrics and finds a common encoder-fusion-prediction pipeline across approaches.

Pith tools