VLMs excel at semantic and grouping tasks while VGMs are stronger on dense geometry and camera motion, with naive fusion yielding balanced representations.
arXiv preprint arXiv:2201.13360 (2022)
15 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.
years
2026 15verdicts
UNVERDICTED 15representative citing papers
Distillation from frontier VLMs plus E-RLVR regularization produces a 4B local model that achieves 34.5% SR on OVON while cutting inference latency by 82.8%.
Introduces Embodied Tool Protocol and tool externalization to improve embodied AI performance on perception and cognition tasks, with measured gains but limits on execution capabilities.
RGB-only active 3D scene graph generation unifies perception and planning to achieve depth-baseline parity and more than double object detection in active indoor exploration.
Fixed external cameras as Common Prior Maps boost initial object recall in 3D scene graph generation by up to 79% and improve active exploration efficiency.
A hybrid semantic graph and retrieval-augmented system with parameter-efficient VLMs achieves state-of-the-art inference and querying speeds on embodied navigation tasks with competitive accuracy.
Hyperbolic Scene Graph (HSG) learns embeddings in hyperbolic space for better hierarchical structure in scene graphs, achieving graph IoU of 33.51 versus 25.37 for the best Euclidean baseline.
A map-free localization method stores posed RGB-D keyframes, retrieves and re-ranks them with a VLM, then fuses sparse depth for on-demand 3D target estimates, matching reconstruction-based performance on navigation benchmarks with far lower build cost.
ObsGraph is a hierarchical observation-centric scene graph that unifies representation, retrieval, and multi-scale exploration for embodied reasoning.
SCOUT integrates uncertainty-guided traversal planning with online probabilistic 3D scene graph construction to treat semantic completeness as an active objective.
Decoupled pose-text fusion on MM-Conv reaches 31.9% top-1 accuracy and shows a learned gate changing policy based on category access in text, serving as a diagnostic against category-representation artifacts.
Mono-Hydra++ is a monocular RGB-IMU pipeline that constructs hierarchical 3D scene graphs in real time while reporting lower trajectory error than some RGB-D baselines on indoor datasets.
IGV-RRT improves object goal navigation in dynamic indoor environments by combining uncertainty-aware priors from 3D scene graphs with online VLM observations in a real-time tree planner.
Introduces a hierarchical object representation pipeline from RGB-D data to meshes and superquadrics for reconstruction, map alignment, and collision checking in robotics.
Semantic relations between objects and structural elements filter candidate graph matches in SLAM, cutting ambiguity and computation in symmetric indoor environments.
citing papers explorer
-
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models
VLMs excel at semantic and grouping tasks while VGMs are stronger on dense geometry and camera motion, with naive fusion yielding balanced representations.
-
LocalNav: Distilling Frontier VLMs and Embodied RL for On-Device Object Goal Navigation
Distillation from frontier VLMs plus E-RLVR regularization produces a 4B local model that achieves 34.5% SR on OVON while cutting inference latency by 82.8%.
-
Enabling Extensible Embodied Capabilities with Tools
Introduces Embodied Tool Protocol and tool externalization to improve embodied AI performance on perception and cognition tasks, with measured gains but limits on execution capabilities.
-
RGB-only Active 3D Scene Graph Generation for Indoor Mobile Robots
RGB-only active 3D scene graph generation unifies perception and planning to achieve depth-baseline parity and more than double object detection in active indoor exploration.
-
Fixed External Cameras as Common Prior Maps for Active 3D Scene Graph Generation
Fixed external cameras as Common Prior Maps boost initial object recall in 3D scene graph generation by up to 79% and improve active exploration efficiency.
-
EmbodiedLGR: Integrating Lightweight Graph Representation and Retrieval for Semantic-Spatial Memory in Robotic Agents
A hybrid semantic graph and retrieval-augmented system with parameter-efficient VLMs achieves state-of-the-art inference and querying speeds on embodied navigation tasks with competitive accuracy.
-
HSG: Hyperbolic Scene Graph
Hyperbolic Scene Graph (HSG) learns embeddings in hyperbolic space for better hierarchical structure in scene graphs, achieving graph IoU of 33.51 versus 25.37 for the best Euclidean baseline.
-
Memory Over Maps: 3D Object Localization Without Reconstruction
A map-free localization method stores posed RGB-D keyframes, retrieves and re-ranks them with a VLM, then fuses sparse depth for on-demand 3D target estimates, matching reconstruction-based performance on navigation benchmarks with far lower build cost.
-
ObsGraph: Hierarchical Observation Representation for Embodied Reasoning and Exploration
ObsGraph is a hierarchical observation-centric scene graph that unifies representation, retrieval, and multi-scale exploration for embodied reasoning.
-
SCOUT: Semantic scene COverage via Uncertainty-guided Traversal
SCOUT integrates uncertainty-guided traversal planning with online probabilistic 3D scene graph construction to treat semantic completeness as an active objective.
-
PoseRefer: Pathway-Local Parameters for Semantically Grounded Reference Resolution
Decoupled pose-text fusion on MM-Conv reaches 31.9% top-1 accuracy and shows a learned gate changing policy based on category access in text, serving as a diagnostic against category-representation artifacts.
-
Mono-Hydra++: Real-Time Monocular Scene Graph Construction with Multi-Task Learning for 3D Indoor Mapping
Mono-Hydra++ is a monocular RGB-IMU pipeline that constructs hierarchical 3D scene graphs in real time while reporting lower trajectory error than some RGB-D baselines on indoor datasets.
-
IGV-RRT: Prior-Real-Time Observation Fusion for Active Object Search in Changing Environments
IGV-RRT improves object goal navigation in dynamic indoor environments by combining uncertainty-aware priors from 3D scene graphs with online VLM observations in a real-time tree planner.
-
Hierarchical Object Representation for Spatial Robot Perception: Points, Meshes, and Superquadrics
Introduces a hierarchical object representation pipeline from RGB-D data to meshes and superquadrics for reconstruction, map alignment, and collision checking in robotics.
-
Robust Graph Matching through Semantic Relationship Generation for SLAM
Semantic relations between objects and structural elements filter candidate graph matches in SLAM, cutting ambiguity and computation in symmetric indoor environments.