Pith. sign in

REVIEW 12 cited by

UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.15725 v2 pith:7LB4YXTE submitted 2025-05-21 cs.RO cs.CV

UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning

classification cs.RO cs.CV
keywords controlfine-graineduavsuav-flowbenchmarkflightflowflying-on-a-word
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Unmanned Aerial Vehicles (UAVs) are evolving into language-interactive platforms, enabling more intuitive forms of human-drone interaction. While prior works have primarily focused on high-level planning and long-horizon navigation, we shift attention to language-guided fine-grained trajectory control, where UAVs execute short-range, reactive flight behaviors in response to language instructions. We formalize this problem as the Flying-on-a-Word (Flow) task and introduce UAV imitation learning as an effective approach. In this framework, UAVs learn fine-grained control policies by mimicking expert pilot trajectories paired with atomic language instructions. To support this paradigm, we present UAV-Flow, the first real-world benchmark for language-conditioned, fine-grained UAV control. It includes a task formulation, a large-scale dataset collected in diverse environments, a deployable control framework, and a simulation suite for systematic evaluation. Our design enables UAVs to closely imitate the precise, expert-level flight trajectories of human pilots and supports direct deployment without sim-to-real gap. We conduct extensive experiments on UAV-Flow, benchmarking VLN and VLA paradigms. Results show that VLA models are superior to VLN baselines and highlight the critical role of spatial grounding in the fine-grained Flow setting.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

    cs.RO 2026-07 conditional novelty 7.0

    ActiveFly-Bench defines Air-EQA, Observation Behavior Planning, and 7-DoF FLUC tasks on 10k real/sim trajectories so UAV agents must plan, fly, and answer questions they cannot solve from the start view.

  2. WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation

    cs.RO 2026-05 unverdicted novelty 7.0

    WorldVLN proposes the first autoregressive world action model for aerial vision-language navigation that predicts short-horizon latent world states, decodes them to waypoints in closed loop, and uses two-stage trainin...

  3. ESARBench: A Benchmark for Agentic UAV Embodied Search and Rescue

    cs.RO 2026-05 unverdicted novelty 7.0

    ESARBench is the first unified benchmark for MLLM-driven UAV agents that must explore, locate clues, and decide on victim positions in photorealistic simulated SAR environments.

  4. QuadAgent: A Responsive Agent System for Vision-Language Guided Quadrotor Agile Flight

    cs.RO 2026-04 unverdicted novelty 7.0

    QuadAgent uses an asynchronous multi-agent architecture with an Impression Graph for scene memory and vision-based avoidance to enable training-free vision-language guided agile quadrotor flight, outperforming baselin...

  5. See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View

    cs.CV 2026-06 unverdicted novelty 6.0

    Introduces UAV-VLN-FOV task and 3DG-VLN framework for precise target-visible UAV navigation, reporting 13.82% success rate gain on a new 2,717-trajectory benchmark with code released.

  6. AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots

    cs.RO 2026-06 unverdicted novelty 6.0

    AIR-VLA+ introduces cascaded manipulation and movement decoders plus asymmetric MoE to decouple action scales in aerial manipulation, reporting 48.0 average score and 80.2% task completion gain over single-head baseli...

  7. WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation

    cs.AI 2026-06 unverdicted novelty 6.0

    WorldFly integrates a world model into a VLA framework via dual-branch coupled flow matching to jointly generate future videos and actions, outperforming baselines on an urban canyon traversal benchmark especially in ...

  8. Can Aerial VLA Models Cooperate? Evaluating Closed-Loop Air-Ground Coordination with CARLA-Air

    cs.RO 2026-05 unverdicted novelty 6.0

    Introduces CARLA-Air simulator for air-ground VLA evaluation and shows that current aerial VLA models track ground partners but fail to achieve stable cooperative behavior under text-based interfaces.

  9. FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation

    cs.CV 2026-04 unverdicted novelty 6.0

    FineCog-Nav uses fine-grained cognitive modules driven by foundation models to outperform zero-shot baselines in UAV navigation and introduces the AerialVLN-Fine benchmark with refined instructions.

  10. FlyMirage: A Fully Automated Generation Pipeline for Diverse and Scalable UAV Flight Data via Generative World Model

    cs.RO 2026-05 unverdicted novelty 5.0

    FlyMirage automates generation of large-scale, diverse, photorealistic aerial VLN datasets with dynamically feasible UAV trajectories by combining LLM scene design, 3D Gaussian Splatting world models, automated explor...

  11. AION: Aerial Indoor Object-Goal Navigation Using Dual-Policy Reinforcement Learning

    cs.RO 2026-01 conditional novelty 5.0

    An end-to-end dual-policy RL method extends zero-shot object-goal navigation from ground robots to indoor drones, using depth-derived laser-scan and reachable-region features, and reports top AI2-THOR scores plus Isaa...

  12. Fly0: Persistent Metric Anchoring for Zero-Shot Aerial Vision-Language Navigation

    cs.RO 2026-02 conditional novelty 4.0

    A zero-shot aerial VLN system that has an MLLM output only 2D image coordinates, then uses depth unprojection and Ego-Planner to navigate, reporting >20 percentage-point SR gains and 31–37% NE reductions over baselines.