Introduces a synchronized cross-view urban traffic dataset with drone ground truth for identity matching and monocular BEV localization tasks.
Vision Meets Drones: A Challenge
10 Pith papers cite this work. Polarity classification is still indexing.
abstract
In this paper we present a large-scale visual object detection and tracking benchmark, named VisDrone2018, aiming at advancing visual understanding tasks on the drone platform. The images and video sequences in the benchmark were captured over various urban/suburban areas of 14 different cities across China from north to south. Specifically, VisDrone2018 consists of 263 video clips and 10,209 images (no overlap with video clips) with rich annotations, including object bounding boxes, object categories, occlusion, truncation ratios, etc. With intensive amount of effort, our benchmark has more than 2.5 million annotated instances in 179,264 images/video frames. Being the largest such dataset ever published, the benchmark enables extensive evaluation and investigation of visual analysis algorithms on the drone platform. In particular, we design four popular tasks with the benchmark, including object detection in images, object detection in videos, single object tracking, and multi-object tracking. All these tasks are extremely challenging in the proposed dataset due to factors such as occlusion, large scale and pose variation, and fast motion. We hope the benchmark largely boost the research and development in visual analysis on drone platforms.
representative citing papers
DeTrack is a new benchmark for drone-embodied tracking in 3D environments and AaDWorlds is a dual world model that improves closed-loop performance by using altitude-aware predictions to balance visibility and safety.
MotionScape is a large-scale UAV video dataset with highly dynamic 6-DoF motions, geometric trajectories, and semantic annotations to train world models that better simulate complex 3D dynamics under large viewpoint changes.
The paper defines UAV reasoning segmentation, releases DRSeg—10,000 aerial images with reasoning QA and masks—and shows its PixDLM baseline beats prior reasoning-segmentation models under fine-tuning.
Proposes SASP for structure-aware data partitioning and CDRO for curriculum-based robust optimization to reduce leakage and stratification issues in spatiotemporally correlated datasets.
VFACamou uses UV-volume rendering and a diffusion-based texture generator with illumination consistency to produce environment-adaptive adversarial camouflage for physical evasion.
ERPPO adds a DSA-based ambiguity estimator to MAPPO and switches between L1 and L2 entropy regularization to improve exploration and stability in non-stationary multi-dimensional observations.
Introduces a multimodal UAV command dataset and shows image-augmented RNN language models outperform text-only versions despite imperfect training associations.
DroneIQA-VLE ensembles SigLIP2 vision encoders and a LoRA-adapted Qwen3.5-9B multimodal LLM to jointly predict global, target, and background quality scores for low-altitude UAV images, placing 2nd in the ICME 2026 Drone-IQA challenge.
Adding a P2 branch to YOLOX-Nano raises small-object AP by 31.10% on VisDrone; QIEA screens structures balancing accuracy, FLOPs, latency, memory and recall.
citing papers explorer
-
Cross-View Urban Traffic Dataset: Drone-Supervised Ground Truth for Monocular Bird's-Eye View Localization
Introduces a synchronized cross-view urban traffic dataset with drone ground truth for identity matching and monocular BEV localization tasks.
-
DeTrack: A Benchmark and Altitude-Aware Dual World Model for Drone-embodied Tracking
DeTrack is a new benchmark for drone-embodied tracking in 3D environments and AaDWorlds is a dual world model that improves closed-loop performance by using altitude-aware predictions to balance visibility and safety.
-
MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models
MotionScape is a large-scale UAV video dataset with highly dynamic 6-DoF motions, geometric trajectories, and semantic annotations to train world models that better simulate complex 3D dynamics under large viewpoint changes.
-
PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation
The paper defines UAV reasoning segmentation, releases DRSeg—10,000 aerial images with reasoning QA and masks—and shows its PixDLM baseline beats prior reasoning-segmentation models under fine-tuning.
-
Beyond the Performance Illusion: Structure-Aware Stratified Partitioning and Curriculum Distributionally Robust Optimization for Spatially Correlated Domains
Proposes SASP for structure-aware data partitioning and CDRO for curriculum-based robust optimization to reduce leakage and stratification issues in spatiotemporally correlated datasets.
-
VFACamou: View-Fused Adversarial Camouflage for Environment-Adaptive Physical Evasion
VFACamou uses UV-volume rendering and a diffusion-based texture generator with illumination consistency to produce environment-adaptive adversarial camouflage for physical evasion.
-
ERPPO: Entropy Regularization-based Proximal Policy Optimization
ERPPO adds a DSA-based ambiguity estimator to MAPPO and switches between L1 and L2 entropy regularization to improve exploration and stability in non-stationary multi-dimensional observations.
-
Kite: Automatic speech recognition for unmanned aerial vehicles
Introduces a multimodal UAV command dataset and shows image-augmented RNN language models outperform text-only versions despite imperfect training associations.
-
DroneIQA-VLE: Multi-Task Drone Image Quality Assessment via Vision-Language Ensemble
DroneIQA-VLE ensembles SigLIP2 vision encoders and a LoRA-adapted Qwen3.5-9B multimodal LLM to jointly predict global, target, and background quality scores for low-altitude UAV images, placing 2nd in the ICME 2026 Drone-IQA challenge.
-
Edge-Constrained UAV Small-Object Detection with P2 Enhancement and Quantum-Inspired Lightweight Structure Search
Adding a P2 branch to YOLOX-Nano raises small-object AP by 31.10% on VisDrone; QIEA screens structures balancing accuracy, FLOPs, latency, memory and recall.