Pith. sign in

REVIEW 7 cited by

NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.16986 v3 pith:P4SLINML submitted 2023-05-26 cs.CV cs.AIcs.CLcs.RO

classification cs.CVcs.AIcs.CLcs.RO
keywords navigationllmsmodelsnavgptagentreasoninglanguageadapting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Trained with an unprecedented scale of data, large language models (LLMs) like ChatGPT and GPT-4 exhibit the emergence of significant reasoning abilities from model scaling. Such a trend underscored the potential of training LLMs with unlimited language data, advancing the development of a universal embodied agent. In this work, we introduce the NavGPT, a purely LLM-based instruction-following navigation agent, to reveal the reasoning capability of GPT models in complex embodied scenes by performing zero-shot sequential action prediction for vision-and-language navigation (VLN). At each step, NavGPT takes the textual descriptions of visual observations, navigation history, and future explorable directions as inputs to reason the agent's current status, and makes the decision to approach the target. Through comprehensive experiments, we demonstrate NavGPT can explicitly perform high-level planning for navigation, including decomposing instruction into sub-goal, integrating commonsense knowledge relevant to navigation task resolution, identifying landmarks from observed scenes, tracking navigation progress, and adapting to exceptions with plan adjustment. Furthermore, we show that LLMs is capable of generating high-quality navigational instructions from observations and actions along a path, as well as drawing accurate top-down metric trajectory given the agent's navigation history. Despite the performance of using NavGPT to zero-shot R2R tasks still falling short of trained models, we suggest adapting multi-modality inputs for LLMs to use as visual navigation agents and applying the explicit reasoning of LLMs to benefit learning-based models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ReferTrack: Referring Then Tracking for Embodied Visual Tracking

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A refer-then-track policy picks the target from indexed detections before planning waypoints, achieving state-of-the-art single-view results on EVT-Bench and approaching multi-camera performance.

  2. MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning

    cs.RO 2025-10 conditional novelty 6.0 of 10

    A four-camera VLA navigation model trained by distilling multiple RL experts achieves strong simulation performance and qualitative real-world transfer.

  3. DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation

    cs.RO 2025-09 conditional novelty 6.0 of 10

    DreamNav achieves new zero-shot SOTA on VLN-CE with an egocentric-only pipeline that generates candidate trajectories, imagines their futures, and selects the best by language alignment.

  4. StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A streaming navigation framework combining a sliding-window KV cache with depth-based token pruning achieves state-of-the-art results on VLN-CE benchmarks with bounded context and low latency.

  5. MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MVL-Loc fuses CLIP image and text features with per-scene pose heads and reports state-of-the-art multi-scene relocalization accuracy on 7Scenes and Cambridge Landmarks.

  6. SG-CoT: An Ambiguity-Aware Robotic Planning Framework using Scene Graph Representations

    cs.RO 2026-03 reject novelty 5.0 of 10

    SG-CoT grounds an LLM planner's chain-of-thought in a scene graph via iterative API queries, improving ambiguity detection and clarification in simulated manipulation, though its success metric credits any clarifying ...

  7. MSNav: Zero-Shot Vision-and-Language Navigation with Dynamic Memory and LLM Spatial Reasoning

    cs.CV 2025-08 conditional novelty 5.0 of 10

    MSNav integrates dynamic map pruning, fine-tuned spatial reasoning (Qwen-Sp), and GPT-4o planning to improve zero-shot vision-and-language navigation on R2R and REVERIE.

Pith tools