REVIEW 7 cited by
ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Visual navigation is an essential skill for home-assistance robots, providing the object-searching ability to accomplish long-horizon daily tasks. Many recent approaches use Large Language Models (LLMs) for commonsense inference to improve exploration efficiency. However, the planning process of LLMs is limited within texts and it is difficult to represent the spatial occupancy and geometry layout only by texts. Both are important for making rational navigation decisions. In this work, we seek to unleash the spatial perception and planning ability of Vision-Language Models (VLMs), and explore whether the VLM, with only on-board camera captured RGB/RGB-D stream inputs, can efficiently finish the visual navigation tasks in a mapless manner. We achieve this by developing the imagination-powered navigation framework ImagineNav, which imagines the future observation images at valuable robot views and translates the complex navigation planning process into a rather simple best-view image selection problem for VLM. To generate appropriate candidate robot views for imagination, we introduce the Where2Imagine module, which is distilled to align with human navigation habits. Finally, to reach the VLM preferred views, an off-the-shelf point-goal navigation policy is utilized. Empirical experiments on the challenging open-vocabulary object navigation benchmarks demonstrates the superiority of our proposed system.
Forward citations
Cited by 7 Pith papers
-
EAGOR: Embodied Reasoning in Omni-direction
EAGOR reformulates embodied 360-degree directional reasoning as recursive Bayesian estimation on a spherical manifold using spherical harmonics, achieving training-free, rotation-equivariant target tracking.
-
Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins
A VLM-driven model predictive controller that evaluates simulated future outcomes rendered from a physics-based digital twin.
-
AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
AREA3D fuses feed-forward 3D confidence and vision-language region reasoning to select informative viewpoints, improving sparse-view 3D reconstruction quality.
-
Imaginative World Modeling with Scene Graphs for Embodied Agent Navigation
SGImagineNav uses an imagined hierarchical scene graph, filled in by an LLM, that guides a robot to unseen objects and achieves 65.4% and 66.8% success on HM3D and HSSD.
-
RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models
RATE-Nav reduces redundant exploration in zero-shot object navigation by segmenting the map into regions and using VLM judgments to terminate unproductive region searches.
-
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
A dual-branch text-imagination system, with one LLM branch for state estimation and one for candidate-direction description, improves R2R navigation success over prior LLM-based VLN methods.
-
BeliefMapNav: 3D Voxel-Based Belief Map for Zero-Shot Object Navigation
A zero-shot navigation system that builds a 3D voxel belief map from LLM-generated landmarks and CLIP features, then plans frontier paths by expected search distance, achieving state-of-the-art success rate and succes...
Discussion (0). Continue with ORCID to comment.