Pith. sign in

REVIEW 11 cited by

CityNav: A Large-Scale Dataset for Real-World Aerial Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14240 v3 pith:PQQ3PA2K submitted 2024-06-20 cs.CV cs.AI

CityNav: A Large-Scale Dataset for Real-World Aerial Navigation

classification cs.CV cs.AI
keywords navigationaerialdatasetreal-worldagentscitynavperformanceaerialvln
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Vision-and-language navigation (VLN) aims to develop agents capable of navigating in realistic environments. While recent cross-modal training approaches have significantly improved navigation performance in both indoor and outdoor scenarios, aerial navigation over real-world cities remains underexplored primarily due to limited datasets and the difficulty of integrating visual and geographic information. To fill this gap, we introduce CityNav, the first large-scale real-world dataset for aerial VLN. Our dataset consists of 32,637 human demonstration trajectories, each paired with a natural language description, covering 4.65 km$^2$ across two real cities: Cambridge and Birmingham. In contrast to existing datasets composed of synthetic scenes such as AerialVLN, our dataset presents a unique challenge because agents must interpret spatial relationships between real-world landmarks and the navigation destination, making CityNav an essential benchmark for advancing aerial VLN. Furthermore, as an initial step toward addressing this challenge, we provide a methodology of creating geographic semantic maps that can be used as an auxiliary modality input during navigation. In our experiments, we compare performance of three representative aerial VLN agents (Seq2seq, CMA and AerialVLN models) and demonstrate that the semantic map representation significantly improves their navigation performance.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

    cs.RO 2026-07 conditional novelty 7.0

    ActiveFly-Bench defines Air-EQA, Observation Behavior Planning, and 7-DoF FLUC tasks on 10k real/sim trajectories so UAV agents must plan, fly, and answer questions they cannot solve from the start view.

  2. CosFly-Track: A Large-Scale Multi-Modal Dataset for UAV Visual Tracking via Multi-Constraint Trajectory Optimization

    cs.RO 2026-05 conditional novelty 7.0

    CosFlyTrack supplies 2.4 million timesteps of aligned RGB, depth, segmentation, pose, target state, and bilingual instructions from expert UAV trajectories, with experiments showing 53-69 point gains in SR@1m after fi...

  3. CosFly-Track: A Large-Scale Multi-Modal Dataset for UAV Visual Tracking via Multi-Constraint Trajectory Optimization

    cs.RO 2026-05 conditional novelty 7.0

    CosFlyTrack provides 12,000 expert UAV trajectories with aligned RGB, depth, segmentation, pose, target state, and bilingual instructions to train visual tracking agents, yielding 53-69 point gains in success rate aft...

  4. Beyond Isolation: A Unified Benchmark for General-Purpose Navigation

    cs.RO 2026-05 unverdicted novelty 7.0

    OmniNavBench is a unified benchmark for general-purpose navigation featuring composite multi-skill instructions, support for humanoid, quadrupedal and wheeled robots, and 1779 human teleoperated trajectories across 17...

  5. ESARBench: A Benchmark for Agentic UAV Embodied Search and Rescue

    cs.RO 2026-05 unverdicted novelty 7.0

    ESARBench is the first unified benchmark for MLLM-driven UAV agents that must explore, locate clues, and decide on victim positions in photorealistic simulated SAR environments.

  6. LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation

    cs.CV 2026-04 unverdicted novelty 7.0

    LookasideVLN improves aerial vision-and-language navigation by encoding directional cues from instructions into an egocentric graph and lightweight knowledge base, outperforming prior methods like CityNavAgent even wi...

  7. How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace

    cs.AI 2026-04 unverdicted novelty 7.0

    Large multimodal models display emerging but limited spatial action capabilities in goal-oriented urban 3D navigation, remaining far from human-level performance with errors diverging rapidly after critical decision points.

  8. See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View

    cs.CV 2026-06 unverdicted novelty 6.0

    Introduces UAV-VLN-FOV task and 3DG-VLN framework for precise target-visible UAV navigation, reporting 13.82% success rate gain on a new 2,717-trajectory benchmark with code released.

  9. FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation

    cs.CV 2026-04 unverdicted novelty 6.0

    FineCog-Nav uses fine-grained cognitive modules driven by foundation models to outperform zero-shot baselines in UAV navigation and introduces the AerialVLN-Fine benchmark with refined instructions.

  10. HTNav: A Hybrid Navigation Framework with Tiered Structure for Urban Aerial Vision-and-Language Navigation

    cs.RO 2026-04 unverdicted novelty 6.0

    HTNav combines imitation and reinforcement learning in a staged, tiered structure with map learning to reach state-of-the-art performance on the CityNav benchmark for urban aerial navigation.

  11. Diffusion-based 4D Trajectory Prediction and Distributed Control for UAV Swarms

    cs.RO 2026-06 unverdicted novelty 5.0

    A diffusion-based 4D trajectory predictor with dimension decoupling and residual refinement, integrated into DNMPC, reduces UAV swarm tracking error by 10-15% while running at 34 FPS on a new multi-scenario dataset.