REVIEW 3 cited by
SceneGPT: A Language Model for 3D Scene Understanding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Building models that can understand and reason about 3D scenes is difficult owing to the lack of data sources for 3D supervised training and large-scale training regimes. In this work we ask - How can the knowledge in a pre-trained language model be leveraged for 3D scene understanding without any 3D pre-training. The aim of this work is to establish whether pre-trained LLMs possess priors/knowledge required for reasoning in 3D space and how can we prompt them such that they can be used for general purpose spatial reasoning and object understanding in 3D. To this end, we present SceneGPT, an LLM based scene understanding system which can perform 3D spatial reasoning without training or explicit 3D supervision. The key components of our framework are - 1) a 3D scene graph, that serves as scene representation, encoding the objects in the scene and their spatial relationships 2) a pre-trained LLM that can be adapted with in context learning for 3D spatial reasoning. We evaluate our framework qualitatively on object and scene understanding tasks including object semantics, physical properties and affordances (object-level) and spatial understanding (scene-level).
Forward citations
Cited by 3 Pith papers
-
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities
By probing visual, projection, and response representations, the authors find that most VLM visual knowledge loss for recognition and counting occurs in the language decoder, while spatial understanding is lost in the...
-
Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies
Current VLMs tend to label anomalous scenes as hazardous, and a four-way hazard/anomaly benchmark exposes this conflation more clearly than binary safe/unsafe tests.
-
Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
Argus fuses multi-view images and camera poses with 3D point cloud features in a frozen-LLM Q-Former architecture, improving 3D question answering, grounding, and scene description over prior 3D-LMMs.
Discussion (0). Sign in to comment.