Pith. sign in

REVIEW 7 cited by

ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.09874 v1 pith:FXZGQA4U submitted 2024-10-13 cs.RO

classification cs.RO
keywords navigationmodelsplanningviewsabilityimaginationimaginenavllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual navigation is an essential skill for home-assistance robots, providing the object-searching ability to accomplish long-horizon daily tasks. Many recent approaches use Large Language Models (LLMs) for commonsense inference to improve exploration efficiency. However, the planning process of LLMs is limited within texts and it is difficult to represent the spatial occupancy and geometry layout only by texts. Both are important for making rational navigation decisions. In this work, we seek to unleash the spatial perception and planning ability of Vision-Language Models (VLMs), and explore whether the VLM, with only on-board camera captured RGB/RGB-D stream inputs, can efficiently finish the visual navigation tasks in a mapless manner. We achieve this by developing the imagination-powered navigation framework ImagineNav, which imagines the future observation images at valuable robot views and translates the complex navigation planning process into a rather simple best-view image selection problem for VLM. To generate appropriate candidate robot views for imagination, we introduce the Where2Imagine module, which is distilled to align with human navigation habits. Finally, to reach the VLM preferred views, an off-the-shelf point-goal navigation policy is utilized. Empirical experiments on the challenging open-vocabulary object navigation benchmarks demonstrates the superiority of our proposed system.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EAGOR: Embodied Reasoning in Omni-direction

    cs.RO 2026-07 conditional novelty 7.0 of 10

    EAGOR reformulates embodied 360-degree directional reasoning as recursive Bayesian estimation on a spherical manifold using spherical harmonics, achieving training-free, rotation-equivariant target tracking.

  2. Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins

    cs.RO 2025-06 conditional novelty 7.0 of 10

    A VLM-driven model predictive controller that evaluates simulated future outcomes rendered from a physics-based digital twin.

  3. AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance

    cs.CV 2025-11 conditional novelty 6.0 of 10

    AREA3D fuses feed-forward 3D confidence and vision-language region reasoning to select informative viewpoints, improving sparse-view 3D reconstruction quality.

  4. Imaginative World Modeling with Scene Graphs for Embodied Agent Navigation

    cs.RO 2025-08 conditional novelty 6.0 of 10

    SGImagineNav uses an imagined hierarchical scene graph, filled in by an LLM, that guides a robot to unseen objects and achieves 65.4% and 66.8% success on HM3D and HSSD.

  5. RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    RATE-Nav reduces redundant exploration in zero-shot object navigation by segmenting the map into regions and using VLM judgments to terminate unproductive region searches.

  6. Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A dual-branch text-imagination system, with one LLM branch for state estimation and one for candidate-direction description, improves R2R navigation success over prior LLM-based VLN methods.

  7. BeliefMapNav: 3D Voxel-Based Belief Map for Zero-Shot Object Navigation

    cs.RO 2025-05 conditional novelty 5.0 of 10

    A zero-shot navigation system that builds a 3D voxel belief map from LLM-generated landmarks and CLIP features, then plans frontier paths by expected search distance, achieving state-of-the-art success rate and succes...

Pith tools