WorldVLN proposes the first autoregressive world action model for aerial vision-language navigation that predicts short-horizon latent world states, decodes them to waypoints in closed loop, and uses two-stage training with Action-aware GRPO to achieve over 12% success-rate gains on benchmarks plus零
Uav-flow colosseo: A real-world benchmark for flying-on-a-word uav imitation learning
9 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 9verdicts
UNVERDICTED 9roles
background 3polarities
background 3representative citing papers
ESARBench is the first unified benchmark for MLLM-driven UAV agents that must explore, locate clues, and decide on victim positions in photorealistic simulated SAR environments.
QuadAgent uses an asynchronous multi-agent architecture with an Impression Graph for scene memory and vision-based avoidance to enable training-free vision-language guided agile quadrotor flight, outperforming baselines in simulations and achieving real-world speeds up to 5 m/s.
Introduces UAV-VLN-FOV task and 3DG-VLN framework for precise target-visible UAV navigation, reporting 13.82% success rate gain on a new 2,717-trajectory benchmark with code released.
AIR-VLA+ introduces cascaded manipulation and movement decoders plus asymmetric MoE to decouple action scales in aerial manipulation, reporting 48.0 average score and 80.2% task completion gain over single-head baseline on AIR-VLA benchmark.
WorldFly integrates a world model into a VLA framework via dual-branch coupled flow matching to jointly generate future videos and actions, outperforming baselines on an urban canyon traversal benchmark especially in unseen environments.
Introduces CARLA-Air simulator for air-ground VLA evaluation and shows that current aerial VLA models track ground partners but fail to achieve stable cooperative behavior under text-based interfaces.
FineCog-Nav uses fine-grained cognitive modules driven by foundation models to outperform zero-shot baselines in UAV navigation and introduces the AerialVLN-Fine benchmark with refined instructions.
FlyMirage automates generation of large-scale, diverse, photorealistic aerial VLN datasets with dynamically feasible UAV trajectories by combining LLM scene design, 3D Gaussian Splatting world models, automated exploration, and trajectory planning.
citing papers explorer
-
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
WorldVLN proposes the first autoregressive world action model for aerial vision-language navigation that predicts short-horizon latent world states, decodes them to waypoints in closed loop, and uses two-stage training with Action-aware GRPO to achieve over 12% success-rate gains on benchmarks plus零
-
ESARBench: A Benchmark for Agentic UAV Embodied Search and Rescue
ESARBench is the first unified benchmark for MLLM-driven UAV agents that must explore, locate clues, and decide on victim positions in photorealistic simulated SAR environments.
-
QuadAgent: A Responsive Agent System for Vision-Language Guided Quadrotor Agile Flight
QuadAgent uses an asynchronous multi-agent architecture with an Impression Graph for scene memory and vision-based avoidance to enable training-free vision-language guided agile quadrotor flight, outperforming baselines in simulations and achieving real-world speeds up to 5 m/s.
-
See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View
Introduces UAV-VLN-FOV task and 3DG-VLN framework for precise target-visible UAV navigation, reporting 13.82% success rate gain on a new 2,717-trajectory benchmark with code released.
-
AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots
AIR-VLA+ introduces cascaded manipulation and movement decoders plus asymmetric MoE to decouple action scales in aerial manipulation, reporting 48.0 average score and 80.2% task completion gain over single-head baseline on AIR-VLA benchmark.
-
WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation
WorldFly integrates a world model into a VLA framework via dual-branch coupled flow matching to jointly generate future videos and actions, outperforming baselines on an urban canyon traversal benchmark especially in unseen environments.
-
Can Aerial VLA Models Cooperate? Evaluating Closed-Loop Air-Ground Coordination with CARLA-Air
Introduces CARLA-Air simulator for air-ground VLA evaluation and shows that current aerial VLA models track ground partners but fail to achieve stable cooperative behavior under text-based interfaces.
-
FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation
FineCog-Nav uses fine-grained cognitive modules driven by foundation models to outperform zero-shot baselines in UAV navigation and introduces the AerialVLN-Fine benchmark with refined instructions.
-
FlyMirage: A Fully Automated Generation Pipeline for Diverse and Scalable UAV Flight Data via Generative World Model
FlyMirage automates generation of large-scale, diverse, photorealistic aerial VLN datasets with dynamically feasible UAV trajectories by combining LLM scene design, 3D Gaussian Splatting world models, automated exploration, and trajectory planning.