Pith. sign in

REVIEW 8 cited by

Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.10756 v1 pith:MZZGOJFN submitted 2025-06-12 cs.RO cs.AI

Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding

classification cs.RO cs.AI
keywords vlflyenvironmentsgoalvision-languagelanguagenavigationinstructionsmodel
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Vision-and-language navigation (VLN) is a long-standing challenge in autonomous robotics, aiming to empower agents with the ability to follow human instructions while navigating complex environments. Two key bottlenecks remain in this field: generalization to out-of-distribution environments and reliance on fixed discrete action spaces. To address these challenges, we propose Vision-Language Fly (VLFly), a framework tailored for Unmanned Aerial Vehicles (UAVs) to execute language-guided flight. Without the requirement for localization or active ranging sensors, VLFly outputs continuous velocity commands purely from egocentric observations captured by an onboard monocular camera. The VLFly integrates three modules: an instruction encoder based on a large language model (LLM) that reformulates high-level language into structured prompts, a goal retriever powered by a vision-language model (VLM) that matches these prompts to goal images via vision-language similarity, and a waypoint planner that generates executable trajectories for real-time UAV control. VLFly is evaluated across diverse simulation environments without additional fine-tuning and consistently outperforms all baselines. Moreover, real-world VLN tasks in indoor and outdoor environments under direct and indirect instructions demonstrate that VLFly achieves robust open-vocabulary goal understanding and generalized navigation capabilities, even in the presence of abstract language input.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View

    cs.CV 2026-06 unverdicted novelty 6.0

    Introduces UAV-VLN-FOV task and 3DG-VLN framework for precise target-visible UAV navigation, reporting 13.82% success rate gain on a new 2,717-trajectory benchmark with code released.

  2. FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation

    cs.CV 2026-04 unverdicted novelty 6.0

    FineCog-Nav uses fine-grained cognitive modules driven by foundation models to outperform zero-shot baselines in UAV navigation and introduces the AerialVLN-Fine benchmark with refined instructions.

  3. GaussFly: Contrastive Reinforcement Learning for Visuomotor Policies in 3D Gaussian Fields

    cs.RO 2026-04 unverdicted novelty 6.0

    GaussFly decouples representation learning from policy optimization via 3D Gaussian Splatting reconstruction and contrastive features to achieve superior sample efficiency and zero-shot sim-to-real transfer for AAV vi...

  4. FSD-VLN: Fast-Slow Dual-System Modeling for Aerial Long-Horizon Vision-Language Navigation

    cs.RO 2026-07 conditional novelty 5.0

    An asynchronous fast-slow dual-system with DiT action modeling and time-weighted loss doubles unseen aerial VLN success rates and halves decision latency in simulation.

  5. No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation

    cs.CV 2026-07 conditional novelty 4.0

    Test-time scaling—parallel candidate generation, iterative self-correction, and multi-criteria selection—improves a frozen UAV navigation VLM's success rate by about 2 percentage points on the TravelUAV benchmark.

  6. Exp2VLA: Enabling Vision-Language-Action for Drone Navigation from Expert Demonstrations

    cs.RO 2026-07 conditional novelty 4.0

    Expert demonstrations distilled into LeRobot-format data let fine-tuned mid-scale VLAs navigate multi-object drone scenes from language, with partial zero-shot color-shape generalization in simulation and SITL.

  7. Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap

    cs.RO 2026-04 unverdicted novelty 4.0

    A survey of UAV vision-and-language navigation that establishes a methodological taxonomy, reviews resources and challenges, and proposes a forward-looking research roadmap.

  8. Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models

    cs.RO 2026-04 unverdicted novelty 4.0

    This survey organizes aerial vision-language navigation methods into five architectural categories, critically reviews evaluation infrastructure, and synthesizes seven open problems for LLM/VLM integration.