REVIEW 21 cited by
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
read the original abstract
Vision-and-Language Navigation (VLN) has gained increasing attention over recent years and many approaches have emerged to advance their development. The remarkable achievements of foundation models have shaped the challenges and proposed methods for VLN research. In this survey, we provide a top-down review that adopts a principled framework for embodied planning and reasoning, and emphasizes the current methods and future opportunities leveraging foundation models to address VLN challenges. We hope our in-depth discussions could provide valuable resources and insights: on one hand, to milestone the progress and explore opportunities and potential roles for foundation models in this field, and on the other, to organize different challenges and solutions in VLN to foundation model researchers.
Forward citations
Cited by 21 Pith papers
-
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
EgoMemReason is a new benchmark showing that even the best multimodal models achieve only 39.6% accuracy on reasoning tasks that require integrating sparse evidence across days in egocentric video.
-
LIME: Learning Intent-aware Camera Motion from Egocentric Video
LIME formulates language-conditioned camera motion as predicting SE(3) target poses from RGB and intent text, using mined multi-intent supervision from egocentric video and a flow-matching pose head.
-
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
Frontier VLMs overconfidently answer spatial questions under occlusion (~30% accuracy) and perspective ambiguity (<10% accuracy) instead of abstaining, and often fail to select helpful additional views.
-
From Region Arrival to Instance-Level Grounding in Vision-and-Language Navigation
REALM, a visibility-aware plug-and-play last-meters module trained on the new REVERIE-AIM dataset, consistently raises instance proximity and grounding success on four VLN backbones.
-
3D-Aware VLMs with Implicit and Explicit Geometries
Fusing a VLM's normal 2D tokens with implicit geometry tokens from a video-geometry encoder plus tokens from its own reconstructed depth maps improves 3D detection, grounding, captioning, and spatial reasoning.
-
AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation
AdvNav disrupts multi-step vision-language navigation with gradient-free, behavior-guided visual noise, reaching 49.7–87.3% attack success on HAMT and MapGPT without model internals.
-
AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation
AdvNav is a gradient-free attack that overlays Perlin noise on a VLN agent's camera and uses behavior feedback plus genetic search, breaking 49.70-87.30% of successful R2R navigations.
-
SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation
SpaAct activates spatial awareness in VLMs using action retrospection, future frame prediction, and progressive curriculum learning to reach SOTA on VLN-CE benchmarks.
-
Instruction-as-State: Environment-Guided and State-Conditioned Semantic Understanding for Embodied Navigation
Instruction understanding is reframed as an evolving Instruction-as-State variable conditioned on perceptual state and realized via the S-EGIU coarse-to-fine framework, reporting a +2.68% SPL gain on REVERIE Test Unseen.
-
OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding
OpenGround grounds open-world 3D targets by planning a task chain and dynamically expanding the object lookup table through online 2D segmentation and 3D lifting, achieving SOTA zero-shot ScanRefer accuracy and 46.2% ...
-
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
GoViG decomposes goal-conditioned navigation instruction generation into visual state prediction and instruction synthesis using an autoregressive multimodal LLM with one-pass and interleaved reasoning, showing gains ...
-
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
Uni-NaVid unifies diverse embodied navigation tasks into one video-based vision-language-action model trained on 3.6 million samples from four sub-tasks, achieving state-of-the-art performance on benchmarks and real-w...
-
HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory
HoloAgent-0 is a unified embodied agent framework with Embodied AgentOS, 3D spatial memory, and embodied skills, deployed and evaluated on real robot hardware for navigation and manipulation tasks.
-
TARIC: Memory-Augmented Traversability-Aware Outdoor VLN under Interrupted Semantic Cues
TARIC maintains traversability-consistent guidance using 3D cue memory during semantic cue interruptions in outdoor VLN, improving success rates on long routes.
-
What Limits Vision-and-Language Navigation ?
StereoNav reaches new benchmark highs on R2R-CE and RxR-CE and improves real-robot reliability by supplying persistent target-location priors and stereo-derived geometry that stay stable under lighting changes and blur.
-
OpenFrontier: General Navigation with Visual-Language Grounded Frontiers
OpenFrontier formulates robot navigation as sparse subgoal reaching via visual-language-grounded frontiers, achieving zero-shot performance without fine-tuning or dense semantic maps.
-
OpenFrontier: General Navigation with Visual-Language Grounded Frontiers
OpenFrontier treats navigation as sparse visual-frontier subgoal selection guided by vision-language priors, claiming strong zero-shot and real-robot performance without task-specific training.
-
Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents
SkillNav decomposes VLN into skill-specific agents trained on synthetic data and routed by a VLM, achieving competitive benchmark results and SOTA generalization on GSA-R2R.
-
FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation
FlowDec is a novel image restoration framework using hybrid temporal conditioning and action-centroid filtering that claims to outperform prior decorruption methods on navigation accuracy and latency in VLN-CE.
-
A Deployable Embodied Vision-Language Navigation System with Hierarchical Cognition and Context-Aware Exploration
Introduces a hierarchical VLN architecture with asynchronous layers, incremental memory graph, and WTRP-based exploration that improves success and efficiency on resource-constrained robots.
-
A Deployable Embodied Vision-Language Navigation System with Hierarchical Cognition and Context-Aware Exploration
A modular VLN architecture builds a cognitive memory graph, decomposes it for VLM reasoning, and solves a weighted traveling repairman problem for context-aware exploration to achieve real-time performance and higher ...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.