Pith. sign in

REVIEW 14 cited by

End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.05755 v1 pith:LHM7VL37 submitted 2024-11-08 cs.RO cs.CLcs.CV

End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering

classification cs.RO cs.CLcs.CV
keywords navigationend-to-endapproachdesignpolicyvlmnavactionsaddition
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We present VLMnav, an embodied framework to transform a Vision-Language Model (VLM) into an end-to-end navigation policy. In contrast to prior work, we do not rely on a separation between perception, planning, and control; instead, we use a VLM to directly select actions in one step. Surprisingly, we find that a VLM can be used as an end-to-end policy zero-shot, i.e., without any fine-tuning or exposure to navigation data. This makes our approach open-ended and generalizable to any downstream navigation task. We run an extensive study to evaluate the performance of our approach in comparison to baseline prompting methods. In addition, we perform a design analysis to understand the most impactful design decisions. Visual examples and code for our project can be found at https://jirl-upenn.github.io/VLMnav/

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LIME: Learning Intent-aware Camera Motion from Egocentric Video

    cs.RO 2026-07 unverdicted novelty 7.0

    LIME formulates language-conditioned camera motion as predicting SE(3) target poses from RGB and intent text, using mined multi-intent supervision from egocentric video and a flow-matching pose head.

  2. VTM-Nav: Harnessing Cross-Episode Experience for Object-Goal Navigation with Hierarchical Visual-Topological Memory

    cs.CV 2026-07 conditional novelty 6.0

    A hierarchical room-and-object memory that persists across independent ObjectNav episodes yields small success-rate gains, but most of the gain comes from within-episode memory rather than the cross-episode component.

  3. EvoMemNav: Efficient Self-Evolving Fine-Grained Memory for Zero-Shot Embodied Navigation

    cs.CV 2026-06 unverdicted novelty 6.0

    EvoMemNav builds a Visual-Semantic Memory Graph keeping raw views, applies a budgeted coarse-to-fine policy, and uses reflection-driven updates to improve zero-shot navigation on GOAT-Bench and HM3D.

  4. SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation

    cs.RO 2026-05 unverdicted novelty 6.0

    SEDualVLN proposes a spatially-enhanced dual-system VLN framework that pairs a fast VLM action generator with a slow MLLM waypoint planner and reports state-of-the-art results on VLN-CE benchmarks.

  5. Visually-grounded Humanoid Agents

    cs.CV 2026-04 unverdicted novelty 6.0

    A coupled world-agent framework uses 3D Gaussian reconstruction and first-person RGB-D perception with iterative planning to enable goal-directed, collision-avoiding humanoid behavior in novel reconstructed scenes.

  6. Structured Observation Language for Efficient and Generalizable Vision-Language Navigation

    cs.CV 2026-03 conditional novelty 6.0

    SOL-Nav encodes RGB-D observations as grid-organized text and uses a 0.6B text-embedding model with four classification heads to predict navigation action blocks, reporting SOTA/comparable R2R-CE/RxR-CE results with 1...

  7. ReMemNav: A Rethinking and Memory-Augmented Framework for Zero-Shot Object Navigation

    cs.RO 2026-03 conditional novelty 6.0

    ReMemNav improves zero-shot object navigation success and efficiency by integrating episodic memory and rethinking with VLMs, achieving SR/SPL gains of 1.7%/7.0% on HM3D v0.1, 18.2%/11.1% on HM3D v0.2, and 8.7%/7.9% on MP3D.

  8. DSCD-Nav: Dual-Stance Cooperative Debate for Object Navigation

    cs.RO 2026-01 conditional novelty 6.0

    A dual-stance debate between a goal-focused and a safety-focused VLM, plus arbitration and optional micro-probing, improves zero-shot object navigation success and path efficiency on HM3Dv1, HM3Dv2, MP3D, and GOAT.

  9. ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination

    cs.RO 2025-12 conditional novelty 6.0

    ImagineNav++ achieves SOTA mapless visual navigation by prompting VLMs to select imagined future views generated from a human-preference-distilled module and maintained via selective foveation memory.

  10. SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation

    cs.RO 2026-05 unverdicted novelty 5.0

    SEDualVLN introduces a spatially-enhanced dual-system VLN architecture that achieves state-of-the-art results on VLN-CE benchmarks through coordinated VLM action generation and MLLM waypoint planning.

  11. GUI Agents with Reinforcement Learning: Toward Digital Inhabitants

    cs.AI 2026-04 unverdicted novelty 5.0

    The paper delivers the first comprehensive overview of RL for GUI agents, organizing methods into offline, online, and hybrid strategies while analyzing trends in rewards, efficiency, and deliberation to outline a fut...

  12. OpenFrontier: General Navigation with Visual-Language Grounded Frontiers

    cs.RO 2026-03 unverdicted novelty 5.0

    OpenFrontier treats navigation as sparse visual-frontier subgoal selection guided by vision-language priors, claiming strong zero-shot and real-robot performance without task-specific training.

  13. OpenFrontier: General Navigation with Visual-Language Grounded Frontiers

    cs.RO 2026-03 unverdicted novelty 5.0

    OpenFrontier formulates robot navigation as sparse subgoal reaching via visual-language-grounded frontiers, achieving zero-shot performance without fine-tuning or dense semantic maps.

  14. Structured Observation Language for Efficient and Generalizable Vision-Language Navigation

    cs.CV 2026-03 reject novelty 4.0

    A 0.6B language model navigates by reading grid-structured text descriptions of depth, object class, and color instead of images, with reported R2R-CE/RxR-CE scores near the top of the leaderboard.