REVIEW 10 cited by
ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning
read the original abstract
For robots to perform a wide variety of tasks, they require a 3D representation of the world that is semantically rich, yet compact and efficient for task-driven perception and planning. Recent approaches have attempted to leverage features from large vision-language models to encode semantics in 3D representations. However, these approaches tend to produce maps with per-point feature vectors, which do not scale well in larger environments, nor do they contain semantic spatial relationships between entities in the environment, which are useful for downstream planning. In this work, we propose ConceptGraphs, an open-vocabulary graph-structured representation for 3D scenes. ConceptGraphs is built by leveraging 2D foundation models and fusing their output to 3D by multi-view association. The resulting representations generalize to novel semantic classes, without the need to collect large 3D datasets or finetune models. We demonstrate the utility of this representation through a number of downstream planning tasks that are specified through abstract (language) prompts and require complex reasoning over spatial and semantic concepts. (Project page: https://concept-graphs.github.io/ Explainer video: https://youtu.be/mRhNkQwRYnc )
Forward citations
Cited by 10 Pith papers
-
Scaling Diverse Language Generation for 3D Visual Grounding
ViGiL3D++ generates diverse 3D visual grounding queries via scene-graph constraint sampling plus LLM language generation, yielding higher diversity than prior datasets and improved benchmark performance.
-
Memory for Attention: Language-Conditioned Re-Perception with a Vision--Language--Motion Map
Persistent map memory schedules budgeted re-perception better than memoryless VLM priors, with gain equal to Var(√λ), and language-conditioned VLMM needs both open-vocabulary relevance and per-instance dynamics.
-
Vision-Language-Motion Maps: An Open-Vocabulary, Uncertainty-Aware, Queryable Motion Attribute for 3D Scene Maps
VLMM is a 3D map representation where each object carries a fused, uncertainty-aware motion attribute (language-based movability prior + observed geometric motion) that makes motion queries such as 'what is moving' an...
-
What to Distinguish and How? Opportunities and Challenges of Augmenting Multiple, Cluttered Objects in Complex Scenes for People with Low Vision
AR importance distinction shifts low-vision attention to higher-priority objects in complex scenes, supports hierarchical scanning, but reduces overall recall and introduces augmentation blending and misalignment challenges.
-
What to Distinguish and How? Opportunities and Challenges of Augmenting Multiple, Cluttered Objects in Complex Scenes for People with Low Vision
For people with low vision, AR overlays that rank objects by importance redirect attention toward high-priority objects, but multi-object augmentation lowers overall scene recall and creates new visual-confusion problems.
-
Expanding Spatial and Temporal Context for Robotic Imitation Learning With Scene Graphs
Dynamic scene graphs serve as explicit memory to improve imitation learning policies for spatial-temporal reasoning under partial observability in mobile and tabletop manipulation.
-
A Survey on Vision-Language-Action Models for Embodied AI
This is the first survey on vision-language-action models, providing a taxonomy across three lines, plus summaries of datasets, simulators, benchmarks, challenges, and future directions in embodied AI.
-
SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching
SceneGraphGrounder builds a persistent 3D scene graph from VLM-inferred relations in 2D views and solves grounding via constrained graph alignment, achieving competitive zero-shot results on ScanRefer with only RGB-D input.
-
FUS3DMaps: Scalable and Accurate Open-Vocabulary Semantic Mapping by 3D Fusion of Voxel- and Instance-Level Layers
FUS3DMaps fuses voxel- and instance-level open-vocabulary layers inside a shared 3D voxel map to improve both layers and enable scalable accurate semantic mapping.
-
A Modular Vision-Language-Action Robotics Framework for Indoor Environments
Describes a modular VLA framework with semantic voxel mapping via OwlViT and VLM-based command classification and grounding for the CMU VLA Challenge.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.