Pith. sign in

REVIEW 7 cited by

CityNav: A Large-Scale Dataset for Real-World Aerial Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14240 v3 pith:PQQ3PA2K submitted 2024-06-20 cs.CV cs.AI

classification cs.CVcs.AI
keywords navigationaerialdatasetreal-worldagentscitynavperformanceaerialvln
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Vision-and-language navigation (VLN) aims to develop agents capable of navigating in realistic environments. While recent cross-modal training approaches have significantly improved navigation performance in both indoor and outdoor scenarios, aerial navigation over real-world cities remains underexplored primarily due to limited datasets and the difficulty of integrating visual and geographic information. To fill this gap, we introduce CityNav, the first large-scale real-world dataset for aerial VLN. Our dataset consists of 32,637 human demonstration trajectories, each paired with a natural language description, covering 4.65 km$^2$ across two real cities: Cambridge and Birmingham. In contrast to existing datasets composed of synthetic scenes such as AerialVLN, our dataset presents a unique challenge because agents must interpret spatial relationships between real-world landmarks and the navigation destination, making CityNav an essential benchmark for advancing aerial VLN. Furthermore, as an initial step toward addressing this challenge, we provide a methodology of creating geographic semantic maps that can be used as an auxiliary modality input during navigation. In our experiments, we compare performance of three representative aerial VLN agents (Seq2seq, CMA and AerialVLN models) and demonstrate that the semantic map representation significantly improves their navigation performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning

    cs.RO 2025-05 conditional novelty 8.0 of 10

    The paper presents UAV-Flow, the first real-world benchmark for language-conditioned fine-grained UAV control, with 30K expert flight episodes and a simulation suite, and reports that VLA models outperform VLN baselines.

  2. ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

    cs.RO 2026-07 conditional novelty 7.0 of 10

    ActiveFly-Bench defines Air-EQA, Observation Behavior Planning, and 7-DoF FLUC tasks on 10k real/sim trajectories so UAV agents must plan, fly, and answer questions they cannot solve from the start view.

  3. UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents

    cs.RO 2025-08 unverdicted novelty 6.0 of 10

    UAV-ON is a new benchmark of 14 Unreal Engine environments with 1270 annotated objects that tests whether aerial agents can navigate to goals described by semantic instance-level instructions.

  4. Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent

    cs.RO 2025-06 conditional novelty 5.0 of 10

    An open-source ROS2/PX4 framework using locally hosted LLMs and VLMs enables natural language drone commands, with the best simulated mission success rate at 40%.

  5. Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A dual-branch text-imagination system, with one LLM branch for state estimation and one for candidate-direction description, improves R2R navigation success over prior LLM-based VLN methods.

  6. UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation

    cs.RO 2025-01 conditional novelty 5.0 of 10

    UAV-VLA generates drone flight plans from natural language using satellite imagery, GPT, and Molmo, and introduces a 30-image benchmark, but its evaluation against a single human operator is weak.

  7. Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding

    cs.RO 2025-06 conditional novelty 4.0 of 10

    VLFly integrates LLaMA, CLIP, and a pretrained ViNT waypoint planner to guide a drone by matching language instructions to goal images, reporting 83% success on direct and 70% on indirect real-world instructions.

Pith tools