REVIEW 7 cited by
CityNav: A Large-Scale Dataset for Real-World Aerial Navigation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Vision-and-language navigation (VLN) aims to develop agents capable of navigating in realistic environments. While recent cross-modal training approaches have significantly improved navigation performance in both indoor and outdoor scenarios, aerial navigation over real-world cities remains underexplored primarily due to limited datasets and the difficulty of integrating visual and geographic information. To fill this gap, we introduce CityNav, the first large-scale real-world dataset for aerial VLN. Our dataset consists of 32,637 human demonstration trajectories, each paired with a natural language description, covering 4.65 km$^2$ across two real cities: Cambridge and Birmingham. In contrast to existing datasets composed of synthetic scenes such as AerialVLN, our dataset presents a unique challenge because agents must interpret spatial relationships between real-world landmarks and the navigation destination, making CityNav an essential benchmark for advancing aerial VLN. Furthermore, as an initial step toward addressing this challenge, we provide a methodology of creating geographic semantic maps that can be used as an auxiliary modality input during navigation. In our experiments, we compare performance of three representative aerial VLN agents (Seq2seq, CMA and AerialVLN models) and demonstrate that the semantic map representation significantly improves their navigation performance.
Forward citations
Cited by 7 Pith papers
-
UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning
The paper presents UAV-Flow, the first real-world benchmark for language-conditioned fine-grained UAV control, with 30K expert flight episodes and a simulation suite, and reports that VLA models outperform VLN baselines.
-
ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception
ActiveFly-Bench defines Air-EQA, Observation Behavior Planning, and 7-DoF FLUC tasks on 10k real/sim trajectories so UAV agents must plan, fly, and answer questions they cannot solve from the start view.
-
UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents
UAV-ON is a new benchmark of 14 Unreal Engine environments with 1270 annotated objects that tests whether aerial agents can navigate to goals described by semantic instance-level instructions.
-
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
An open-source ROS2/PX4 framework using locally hosted LLMs and VLMs enables natural language drone commands, with the best simulated mission success rate at 40%.
-
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
A dual-branch text-imagination system, with one LLM branch for state estimation and one for candidate-direction description, improves R2R navigation success over prior LLM-based VLN methods.
-
UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation
UAV-VLA generates drone flight plans from natural language using satellite imagery, GPT, and Molmo, and introduces a 30-image benchmark, but its evaluation against a single human operator is weak.
-
Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding
VLFly integrates LLaMA, CLIP, and a pretrained ViNT waypoint planner to guide a drone by matching language instructions to goal images, reporting 83% success on direct and 70% on indirect real-world instructions.
Discussion (0). Continue with ORCID to comment.