REVIEW 4 cited by
MC-GPT: Empowering Vision-and-Language Navigation with Memory Map and Reasoning Chains
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In the Vision-and-Language Navigation (VLN) task, the agent is required to navigate to a destination following a natural language instruction. While learning-based approaches have been a major solution to the task, they suffer from high training costs and lack of interpretability. Recently, Large Language Models (LLMs) have emerged as a promising tool for VLN due to their strong generalization capabilities. However, existing LLM-based methods face limitations in memory construction and diversity of navigation strategies. To address these challenges, we propose a suite of techniques. Firstly, we introduce a method to maintain a topological map that stores navigation history, retaining information about viewpoints, objects, and their spatial relationships. This map also serves as a global action space. Additionally, we present a Navigation Chain of Thoughts module, leveraging human navigation examples to enrich navigation strategy diversity. Finally, we establish a pipeline that integrates navigational memory and strategies with perception and action prediction modules. Experimental results on the REVERIE and R2R datasets show that our method effectively enhances the navigation ability of the LLM and improves the interpretability of navigation reasoning.
Forward citations
Cited by 4 Pith papers
-
AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness
A tool-calling harness replaces learned waypoints with pixel-level action, on-demand depth, and selective memory, achieving 55% SR and 48.41% SPL zero-shot on R2R-CE with GPT-5.5.
-
DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
DreamNav achieves new zero-shot SOTA on VLN-CE with an egocentric-only pipeline that generates candidate trajectories, imagines their futures, and selects the best by language alignment.
-
CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process Thinking
CogDDN uses a fast heuristic VLM paired with a slow analytic reflection process and a growing knowledge base to navigate to objects that implicitly satisfy a user's demand, with large reported gains on AI2Thor DDN benchmarks.
-
SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models
A self-evolving, training-free VLN agent with hierarchical memory, RAG plus chain-of-thought reasoning, and reflection reports state-of-the-art success rates on R2R and REVERIE.
Discussion (0). Sign in to comment.