Pith. sign in

Boost UAV-based Ojbect Detection via Scale-Invariant Feature Disentanglement and Adversarial Learning

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Detecting objects from Unmanned Aerial Vehicles (UAV) is often hindered by a large number of small objects, resulting in low detection accuracy. To address this issue, mainstream approaches typically utilize multi-stage inferences. Despite their remarkable detecting accuracies, real-time efficiency is sacrificed, making them less practical to handle real applications. To this end, we propose to improve the single-stage inference accuracy through learning scale-invariant features. Specifically, a Scale-Invariant Feature Disentangling module is designed to disentangle scale-related and scale-invariant features. Then an Adversarial Feature Learning scheme is employed to enhance disentanglement. Finally, scale-invariant features are leveraged for robust UAV-based object detection. Furthermore, we construct a multi-modal UAV object detection dataset, State-Air, which incorporates annotated UAV state parameters. We apply our approach to three lightweight detection frameworks on two benchmark datasets. Extensive experiments demonstrate that our approach can effectively improve model accuracy and achieve state-of-the-art (SoTA) performance on two datasets. Our code and dataset will be publicly available once the paper is accepted.

fields

cs.CV 1

years

2025 1

verdicts

REJECT 1

representative citing papers

RemoteSAM: Towards Segment Anything for Earth Observation

cs.CV · 2025-05-23 · reject · novelty 6.0

RemoteSAM unifies remote sensing classification, detection, segmentation, and grounding through a single referring expression segmentation model trained on 270K VLM-generated image-text-mask triplets.

citing papers explorer

Showing 1 of 1 citing paper.

  • RemoteSAM: Towards Segment Anything for Earth Observation cs.CV · 2025-05-23 · reject · none · ref 39 · internal anchor

    RemoteSAM unifies remote sensing classification, detection, segmentation, and grounding through a single referring expression segmentation model trained on 270K VLM-generated image-text-mask triplets.