Pith. sign in

REVIEW 3 major objections 6 minor 49 references

CollabOD raises high-IoU accuracy on tiny UAV objects by preserving structural cues early and aligning multi-path features before fusion, without raising inference cost.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-15 14:09 UTC pith:MIT7OI4A

load-bearing objection Solid UAV YOLO engineering with real multi-benchmark numbers, but VisDrone ablation tables are internally broken so module credit is not yet proven. the 3 major comments →

arxiv 2603.05905 v2 pith:MIT7OI4A submitted 2026-03-06 cs.CV

CollabOD: Collaborative Multi-Backbone with Cross-scale Vision for UAV Small Object Detection

classification cs.CV
keywords UAV small object detectionstructural detail preservationcross-path feature alignmentmulti-scale fusionlightweight detection headVisDroneUAVDTAI-TOD
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

High-altitude UAV images make small targets hard to localize: repeated downsampling erodes boundaries and textures, and multi-branch detectors often fuse misaligned feature streams. This paper claims that a lightweight detector can recover stable localization if it explicitly keeps structural detail at the input and backbone, then calibrates heterogeneous paths before multi-scale fusion, and finally uses a shared detail-aware head that reparameterizes away at inference. The resulting CollabOD stack—Dual-Path Fusion Stem, Dense Aggregation Block, Bilateral Reweighting Module, and Unified Detail-Aware Head—is built on a YOLO-style base and is tested on VisDrone, UAVDT, and AI-TOD. The reported result is higher strict-IoU accuracy with competitive or lower compute, including 52.4 AP50, 30.8 AP75, and 29.9 AP50:95 at 65.5 GFLOPs on VisDrone and 137 FPS on AI-TOD. A sympathetic reader cares because onboard UAV perception needs both precise boxes on sub-32-pixel objects and budgets that fit real flight hardware.

Core claim

CollabOD establishes that UAV small-object detection improves under stricter IoU thresholds when localization-related structural information is preserved early, heterogeneous multi-backbone streams are bilaterally reweighted and amplitude-calibrated before fusion, and the detection head shares detail enhancement without extra deployment cost. On VisDrone the model reports 52.4 AP50, 30.8 AP75, and 29.9 AP50:95 at 65.5 GFLOPs; on UAVDT it reports leading AP50 and AP50:95 among the compared methods; on AI-TOD it reports 45.4 AP50 and 20.0 AP50:95 at 137 FPS.

What carries the argument

CollabOD’s collaborative stack: Dual-Path Fusion Stem (structure vs detail streams before downsampling), Dense Aggregation Block (injects shallow structural cues into deeper features), Bilateral Reweighting Module (spatial masks plus learnable channel scaling to align two paths before fusion), and Unified Detail-Aware Head (shared detail enhancement with reparameterized, scale-shared prediction).

Load-bearing premise

The central claim rests on the premise that early structural preservation and pre-fusion path alignment—not other training or protocol differences—are what drive the reported high-IoU gains under the same evaluation setup as the main baselines.

What would settle it

A full re-run of VisDrone training and COCO-style evaluation under one fixed protocol where adding DPF-Stem, DABlock, BRM, and the UDA Head in turn fails to produce the claimed AP75 and AP50:95 lifts over YOLO11-M-P2 at comparable or lower GFLOPs.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Lightweight UAV detectors can raise AP75 without increasing inference FLOPs if structural cues are protected before deep downsampling.
  • Pre-fusion bilateral reweighting of multi-path features can reduce localization instability on dense small aerial objects.
  • A reparameterized shared detail head can improve boundary regression while remaining suitable for onboard budgets.
  • The same design is claimed to transfer across VisDrone, traffic-oriented UAVDT, and tiny-object AI-TOD benchmarks.
  • Stronger high-IoU boxes under fixed compute support tighter coupling to real-time aerial tracking and nest-based inspection pipelines.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If pre-fusion path calibration is the real lever, many post-fusion attention modules for tiny aerial objects may be solving the wrong stage of the pipeline.
  • The dual-stream stem idea is a natural candidate for other low-SNR remote-sensing tasks where edges vanish under hierarchical pooling.
  • A useful stress test the paper leaves open is whether the same modules keep their AP75 gains under stronger domain shift in weather, altitude, and motion blur without retuning.
  • On multi-UAV collaborative perception, aligning feature streams before fusion may matter as much between vehicles as between backbone paths inside one model.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. CollabOD is a YOLO11-M-P2-based detector for UAV small-object detection that aims to preserve localization-related structural cues and calibrate heterogeneous multi-path features before fusion. It introduces four modules: Dual-Path Fusion Stem (DPF-Stem) for early structure/detail separation, Dense Aggregation Block (DABlock) for hierarchical structural compensation, Bilateral Reweighting Module (BRM) for pre-fusion spatial/channel calibration, and Unified Detail-Aware Head (UDA Head) with shared detail enhancement and re-parameterization. The paper reports strong accuracy–efficiency results on VisDrone (52.4 AP50, 30.8 AP75, 29.9 AP50:95 at 65.5 GFLOPs), UAVDT (best AP50/AP50:95 among compared methods), and AI-TOD (best YOLO-series AP50/AP50:95 at 137 FPS), supported by stepwise ablations and qualitative visualizations.

Significance. If the reported gains are reliable, CollabOD is a practically useful contribution for UAV small-object detection: it targets a real deployment bottleneck (detail loss and cross-path misalignment under tight compute) and claims state-of-the-art or near-SOTA high-IoU accuracy with lower GFLOPs than several strong baselines, including transformer detectors. The modular design is clear, code is promised, and multi-benchmark evaluation (VisDrone, UAVDT, AI-TOD) is appropriate for the claim. The significance is currently limited by internal numerical inconsistencies that undermine the causal attribution of gains to the proposed modules; resolving those would make the paper a solid systems contribution rather than a provisional architecture report.

major comments (3)
  1. [§IV-B, Tables I–II] Table I vs Table II (VisDrone) are not reconcilable and currently break the causal claim that DPF-Stem/DABlock/BRM/UDA Head produce the reported gains. Table I lists YOLO11-M-P2 as 46.4 AP50 / 25.3 AP75 / 27.3 AP50:95 (91.3 GFLOPs) and CollabOD as 52.4 / 30.8 / 29.9 with 20.9M params and 65.5 GFLOPs. Table II’s baseline is 26.2 / 46.0 / 25.3 at the same 91.3 GFLOPs (AP50/AP75 appear swapped or mislabeled), and the full-model row plus surrounding prose report 50.7 AP50, 52.4 AP75, 30.8 AP50:95 with a 29.9-scale parameter figure. Intermediate AP75 values are also non-monotonic in a way that suggests column mislabeling. Please recompute and re-align all VisDrone numbers under one protocol, fix headers, and ensure the ablation baseline matches Table I’s YOLO11-M-P2.
  2. [Abstract; Tables I, II, IV] Parameter counts for CollabOD are inconsistent across the manuscript (Table I: 20.9M; Table II full model and AI-TOD Table IV: 29.9M). Because efficiency is a central claim (lowest GFLOPs among several competitors; favorable accuracy–efficiency trade-off), please report a single audited Params/GFLOPs/FPS measurement protocol for the final model and all ablation variants, and correct every table/abstract occurrence.
  3. [§IV-B.2, Table II] The ablation narrative depends on incremental module credit, but with the current VisDrone table conflict the module-level gains cannot be trusted even if the final detector number is real. After correcting metrics, please also clarify what the ablation baseline is (plain YOLO11-M-P2 vs a modified stem/backbone), why early DPF-Stem alone can drop AP75 so sharply if that drop is real, and whether intermediate rows use identical training schedules and evaluation settings as the main comparison.
minor comments (6)
  1. [§IV-B.2] In §IV-B.2 the prose says the baseline AP50 is 26.2, which matches neither Table I’s YOLO11-M-P2 nor a plausible COCO AP50 for that model; this should be corrected with the tables.
  2. [Table III] Table III caption cites UAVDT as [13], but [13] is AI-TOD; UAVDT is [12]. Fix reference numbering in captions.
  3. [Figs. 1–5] Several figure panels (Figs. 1–5) are rendered as placeholder boxes in the manuscript text; ensure final figures clearly show ERF growth, qualitative detections, and heatmaps with readable labels.
  4. [§III-A, Eq. (4)] Notation: Eq. (4) uses both concatenation-style aggregation and a residual switch δ without fully specifying how multi-level features are aligned before aggregation; a short implementation note would help reproducibility.
  5. [Abstract; Introduction] Minor language issues: “The code are available”, “UA V” spacing, and occasional duplicated phrasing between abstract and introduction.
  6. [§IV-D.2, Table V] AI-TOD ablation (Table V) shows temporary accuracy drops when adding DPF-Stem/DABlock before recovery with the full model; a brief discussion of this non-monotonic path would strengthen the ablation story once VisDrone is fixed.

Circularity Check

0 steps flagged

No circularity: empirical architecture paper whose reported APs are external measurements on public benchmarks, not algebraic restatements of fitted inputs or self-definitional loops.

full rationale

CollabOD is a standard empirical computer-vision architecture paper. Its load-bearing claims are measured detection metrics (AP50, AP75, AP50:95, GFLOPs, FPS) obtained by training and evaluating the proposed modules (DPF-Stem, DABlock, BRM, UDA Head) on fixed public datasets (VisDrone-2019-DET, UAVDT, AI-TOD) under the COCO protocol. The methodological equations (1)–(9) and Algorithm 1 merely define the network operators; they do not derive a numerical prediction from a quantity that was itself fitted to the same target. Related-work citations supply ordinary background and are not uniqueness theorems or load-bearing premises that force the experimental outcome. Table inconsistencies noted by the skeptic (baseline AP values, column swaps between Table I and Table II) constitute a reproducibility/correctness defect, not a circular derivation. Because no step reduces a claimed result to its own inputs by construction, the circularity score is 0 and the steps list is empty.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 4 invented entities

Load-bearing content is architectural design plus training recipe on public UAV benchmarks. Free parameters are optimizer/schedule/input choices and module widths. Axioms are standard detection evaluation and the YOLO11-M-P2 starting point. Invented entities are the four named modules; they are engineering constructs with only in-paper ablation evidence.

free parameters (5)
  • SGD learning rate = 0.01
    Initial lr 0.01 is a hand-chosen training hyperparameter that affects final AP.
  • SGD momentum = 0.937
    Momentum 0.937 is a fixed training choice from the YOLO recipe.
  • input resolution = 640x640
    640×640 is a design choice that trades small-object detail against compute.
  • batch size and epoch count = batch 8, 500 epochs
    Batch 8 and 500 epochs are training schedule knobs not derived from theory.
  • module channel/hidden widths (e.g., UDA Ch, DFL bins R)
    Architectural widths and DFL bin count are free design parameters controlling capacity and FLOPs.
axioms (4)
  • domain assumption COCO-style AP50/AP75/AP50:95 (and APS/APM) are valid proxies for localization quality on UAV small objects.
    All claims of improved localization stability are justified via these metrics (§IV-A3, Tables I–V).
  • domain assumption YOLO11-M-P2 is an appropriate and fairly trained baseline for attributing gains to the proposed modules.
    Framework is built on YOLO11-M-P2; ablations and comparisons treat it as the reference (§III, Table I–II).
  • ad hoc to paper Heterogeneous multi-path features are spatially/semantically misaligned enough that pre-fusion bilateral reweighting is necessary for stable small-object regression.
    Core motivation in §I and design of BRM in §III-B; not independently measured outside the proposed modules.
  • domain assumption Standard convolutional downsampling attenuates localization-critical high-frequency structure in UAV imagery.
    Stated as the problem premise motivating DPF-Stem and DABlock (§I, §III-A).
invented entities (4)
  • Dual-Path Fusion Stem (DPF-Stem) no independent evidence
    purpose: Split early features into structure and detail streams, fuse after pooling/conv to preserve boundaries before downsampling.
    New named module; evidence is only in-paper ablations and ERF-style narrative, not external independent validation.
  • Dense Aggregation Block (DABlock) no independent evidence
    purpose: Inject aligned multi-stage features into deeper layers to compensate hierarchical structural attenuation.
    Architectural block introduced in §III-A; support is ablation deltas and ERF visualization (Fig. 3).
  • Bilateral Reweighting Module (BRM) no independent evidence
    purpose: Generate spatial bilateral gates and learnable channel scales to align two-stream features before fusion.
    Key alignment mechanism (§III-B); no external measurement of misalignment reduction beyond end-task AP.
  • Unified Detail-Aware Head (UDA Head) no independent evidence
    purpose: Share detail enhancement across scales and reparameterize so inference cost stays low while improving box regression.
    Detection-head design in §III-C / Algorithm 1; gains claimed only via ablations.

pith-pipeline@v1.1.0-grok45 · 17823 in / 3817 out tokens · 48062 ms · 2026-07-15T14:09:39.524020+00:00 · methodology

0 comments
read the original abstract

Small object detection in unmanned aerial vehicle (UAV) imagery is challenging because high-altitude viewpoints produce severe scale variation, weak structural cues, and tight computational budgets. Existing lightweight detectors usually fuse multi-scale features after downsampling, where boundary and texture details have already been attenuated and heterogeneous feature streams may be spatially misaligned. To address these issues, we propose CollabOD, a collaborative detection framework that preserves structural details, aligns cross-path features before fusion, and keeps the detection head lightweight at inference time. CollabOD combines a Dual-Path Fusion Stem, a Dense Aggregation Block, a Bilateral Reweighting Module, and a Unified Detail-Aware Head to strengthen localization-oriented representation while limiting extra computation. On VisDrone, CollabOD obtains 52.4 AP50, 30.8 AP75, and 29.9 AP50:95 with 65.5 GFLOPs; on UAVDT it reaches 31.2 AP50 and 17.4 AP50:95; and on AI-TOD it reaches 45.4 AP50 and 20.0 AP50:95 at 137 FPS. The code is available at: https://github.com/Bai-Xuecheng/CollabOD.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

49 extracted references · 3 canonical work pages

  1. [1]

    Ad-det: Boosting object detection in uav images with focused small objects and balanced tail classes,

    Z. Li, S. Lian, D. Pan, Y . Wang, and W. Liu, “Ad-det: Boosting object detection in uav images with focused small objects and balanced tail classes,”Remote Sensing, vol. 17, no. 9, p. 1556, 2025

  2. [2]

    Small object detection: A comprehensive survey on challenges, techniques and real-world applications,

    M. Nikouei, B. Baroutian, S. Nabavi, F. Taraghi, A. Aghaei, A. Sajedi, and M. E. Moghaddam, “Small object detection: A comprehensive survey on challenges, techniques and real-world applications,”Intelligent Systems with Applications, vol. 27, p. 200561, 2025. [Online]. Available: https://www.sciencedirect.com/ science/article/pii/S2667305325000870

  3. [3]

    Msud-yolo: A novel multiscale small object detection model for uav aerial images,

    X. Zhao, H. Zhang, W. Zhang, J. Ma, C. Li, Y . Ding, and Z. Zhang, “Msud-yolo: A novel multiscale small object detection model for uav aerial images,”Drones, vol. 9, no. 6, 2025. [Online]. Available: https://www.mdpi.com/2504-446X/9/6/429

  4. [4]

    A uav aerial image small object detection algorithm based on fine-grained feature preservation and multi-scale feature pyramid balancing,

    J. Luo, K. Chang, J. Huanget al., “A uav aerial image small object detection algorithm based on fine-grained feature preservation and multi-scale feature pyramid balancing,”Complex & Intelligent Systems, vol. 12, p. 12, 2026. [Online]. Available: https://doi.org/10.1007/s40747-025-02126-x

  5. [5]

    Efsi-detr: Efficient frequency- semantic integration for real-time small object detection in uav imagery,

    Y . Xia, C. Liu, T. Xiang, and Z. Tu, “Efsi-detr: Efficient frequency- semantic integration for real-time small object detection in uav imagery,” 2026. [Online]. Available: https://arxiv.org/abs/2601.18597

  6. [6]

    Efficient feature fusion for uav object detection,

    X. Wang, Y . Peng, and C. Shen, “Efficient feature fusion for uav object detection,” 2025. [Online]. Available: https://arxiv.org/abs/2501.17983

  7. [7]

    MST-DETR: A multi-scale enhanced tiny object detection framework,

    L. Li, Z. Zhu, X. Zhaoet al., “MST-DETR: A multi-scale enhanced tiny object detection framework,”Signal, Image and Video Processing, vol. 20, p. 5, 2026. [Online]. Available: https://doi.org/10.1007/s11760-025-05074-8

  8. [8]

    Nas-fpn: Learning scalable feature pyramid architecture for object detection,

    G. Ghiasi, T.-Y . Lin, and Q. V . Le, “Nas-fpn: Learning scalable feature pyramid architecture for object detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 7036–7045

  9. [9]

    Augfpn: Improving multi-scale feature learning for object detection,

    C. Guo, B. Fan, Q. Zhang, S. Xiang, and C. Pan, “Augfpn: Improving multi-scale feature learning for object detection,” 2019. [Online]. Available: https://arxiv.org/abs/1912.05384

  10. [10]

    Ultralytics yolo11,

    G. Jocher and J. Qiu, “Ultralytics yolo11,” 2024. [Online]. Available: https://github.com/ultralytics/ultralytics

  11. [11]

    Detection and tracking meet drones challenge,

    P. Zhu, L. Wen, D. Du, X. Bian, H. Fan, Q. Hu, and H. Ling, “Detection and tracking meet drones challenge,”IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1–1, 2021

  12. [12]

    The unmanned aerial vehicle benchmark: Object detection and tracking,

    D. Du, Y . Qi, H. Yu, Y . Yang, K. Duan, G. Li, W. Zhang, Q. Huang, and Q. Tian, “The unmanned aerial vehicle benchmark: Object detection and tracking,” inProceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 370–386

  13. [13]

    Tiny object detection in aerial images,

    J. Wang, W. Yang, H. Guo, R. Zhang, and G.-S. Xia, “Tiny object detection in aerial images,” inProceedings of the 25th International Conference on Pattern Recognition (ICPR). IEEE, 2021, pp. 3791– 3798

  14. [14]

    Esod: Efficient small object detection on high-resolution images,

    K. Liu, Z. Fu, S. Jin, Z. Chen, F. Zhou, R. Jiang, Y . Chen, and J. Ye, “Esod: Efficient small object detection on high-resolution images,”IEEE Transactions on Image Processing, vol. 34, pp. 183–195, 2024

  15. [15]

    Sahi: A lightweight vision library for performing large scale object detection and instance segmentation,

    F. C. Akyon, C. Cengiz, S. O. Altinuc, D. Cavusoglu, K. Sahin, and O. Eryuksel, “Sahi: A lightweight vision library for performing large scale object detection and instance segmentation,” Nov. 2021. [Online]. Available: https://doi.org/10.5281/zenodo.5718950

  16. [16]

    Superyolo: Super resolution assisted object detection in multimodal remote sensing imagery,

    J. Zhang, J. Lei, W. Xie, Z. Fang, Y . Li, and Q. Du, “Superyolo: Super resolution assisted object detection in multimodal remote sensing imagery,”IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–15, 2023

  17. [17]

    Afpn: Asymptotic feature pyramid network for object detection,

    G. Yang, J. Lei, Z. Zhu, S. Cheng, Z. Feng, and R. Liang, “Afpn: Asymptotic feature pyramid network for object detection,”arXiv preprint arXiv:2306.15988, 2023

  18. [18]

    Legnet: A lightweight edge-gaussian network for low-quality remote sensing image object detection,

    W. Lu, S.-B. Chen, H.-D. Li, Q.-L. Shu, C. H. Ding, J. Tang, and B. Luo, “Legnet: A lightweight edge-gaussian network for low-quality remote sensing image object detection,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 2844–2853

  19. [19]

    Mgdfis: Multi-scale global-detail feature integration strategy for small object detection,

    Y . Wang, X. Bai, B. Hu, C. Xu, H. Chen, V . Chung, T. Li, and X. Chen, “Mgdfis: Multi-scale global-detail feature integration strategy for small object detection,”arXiv preprint arXiv:2506.12697, 2025

  20. [20]

    Yolo-ms: Rethinking multi-scale representation learning for real-time object detection,

    Y . Chen, X. Yuan, J. Wang, R. Wu, X. Li, Q. Hou, and M.-M. Cheng, “Yolo-ms: Rethinking multi-scale representation learning for real-time object detection,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 6, pp. 4240–4252, 2025

  21. [21]

    Feature pyramid networks for object detection,

    T.-Y . Lin, P. Doll´ar, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2017, pp. 2117–2125

  22. [22]

    Panet: Few-shot image semantic segmentation with prototype alignment,

    K. Wang, J. H. Liew, Y . Zou, D. Zhou, and J. Feng, “Panet: Few-shot image semantic segmentation with prototype alignment,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 9197–9206

  23. [23]

    Path aggregation network for instance segmentation,

    S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia, “Path aggregation network for instance segmentation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 8759–8768

  24. [24]

    Learning spatial fusion for single-shot object detection,

    S. Liu, D. Huang, and Y . Wang, “Learning spatial fusion for single-shot object detection,”arXiv preprint arXiv:1911.09516, 2019

  25. [25]

    Mhaf-yolo: Multi-branch heterogeneous auxiliary fusion yolo for accurate object detection,

    Z. Yang, Q. Guan, Z. Yu, X. Xu, H. Long, S. Lian, H. Hu, and Y . Tang, “Mhaf-yolo: Multi-branch heterogeneous auxiliary fusion yolo for accurate object detection,”arXiv preprint arXiv:2502.04656, 2025

  26. [26]

    Yolo-master: Moe- accelerated with specialized transformers for enhanced real-time detection,

    X. Lin, J. Peng, Z. Gan, J. Zhu, and J. Liu, “Yolo-master: Moe- accelerated with specialized transformers for enhanced real-time detection,”arXiv preprint arXiv:2512.23273, 2025

  27. [27]

    Asymmetric mamba–cnn collaborative architecture for large- size remote sensing image semantic segmentation,

    J. Zhang, M. Chen, Y . Zhao, L. Shan, C. Li, H. Hu, X. Ge, Q. Zhu, and B. Xu, “Asymmetric mamba–cnn collaborative architecture for large- size remote sensing image semantic segmentation,”IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1–19, 2025

  28. [28]

    Focal and global knowledge distillation for detectors,

    Z. Yang, Z. Li, X. Jiang, Y . Gong, Z. Yuan, D. Zhao, and C. Yuan, “Focal and global knowledge distillation for detectors,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 4643–4652

  29. [29]

    Task-specific context decoupling for object detection,

    J. Zhuang, Z. Qin, H. Yu, and X. Chen, “Task-specific context decoupling for object detection,”arXiv preprint arXiv:2303.01047, 2023

  30. [30]

    Generalized intersection over union: A metric and a loss for bounding box regression,

    H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, and S. Savarese, “Generalized intersection over union: A metric and a loss for bounding box regression,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 658–666

  31. [31]

    Focal and efficient iou loss for accurate bounding box regression,

    Y .-F. Zhang, W. Ren, Z. Zhang, Z. Jia, L. Wang, and T. Tan, “Focal and efficient iou loss for accurate bounding box regression,” Neurocomputing, vol. 506, pp. 146–157, 2022

  32. [32]

    Distance-iou loss: Faster and better learning for bounding box regression,

    Z. Zheng, P. Wang, W. Liu, J. Li, R. Ye, and D. Ren, “Distance-iou loss: Faster and better learning for bounding box regression,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 07, 2020, pp. 12 993–13 000

  33. [33]

    Pp-yoloe: An evolved version of yolo,

    S. Xu, X. Wang, W. Lv, Q. Chang, C. Cui, K. Deng, G. Wang, Q. Dang, S. Wei, Y . Du, and B. Lai, “Pp-yoloe: An evolved version of yolo,” arXiv preprint arXiv:2203.16250, 2022

  34. [34]

    Cross-layer feature pyramid transformer for small object detection in aerial images,

    Z. Du, Z. Hu, G. Zhao, Y . Jin, and H. Ma, “Cross-layer feature pyramid transformer for small object detection in aerial images,”IEEE Transactions on Geoscience and Remote Sensing, 2025

  35. [35]

    Querydet: Cascaded sparse query for accelerating high-resolution small object detection,

    C. Yang, Z. Huang, and N. Wang, “Querydet: Cascaded sparse query for accelerating high-resolution small object detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 13 668–13 677

  36. [36]

    Generalized uav object detection via frequency domain disentanglement,

    K. Wang, X. Fu, Y . Huang, C. Cao, G. Shi, and Z.-J. Zha, “Generalized uav object detection via frequency domain disentanglement,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 1064–1073

  37. [37]

    Uav-detr: efficient end-to-end object detection for unmanned aerial vehicle imagery,

    H. Zhang, H. Zhang, K. Liu, Z. Gan, and G.-N. Zhu, “Uav-detr: efficient end-to-end object detection for unmanned aerial vehicle imagery,” in 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2025, pp. 15 143–15 149

  38. [38]

    Brstd: Bio-inspired remote sensing tiny object detection,

    S. Huang, C. Lin, X. Jiang, and Z. Qu, “Brstd: Bio-inspired remote sensing tiny object detection,”IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–15, 2024

  39. [39]

    Uav- malo: Mamba-augmented yolo hybrid architecture for uav micro-object detection in autonomous robotics,

    L. Wei, S. Sun, J. Yao, Y . Mi, X. Sui, H. Chen, and S. Liu, “Uav- malo: Mamba-augmented yolo hybrid architecture for uav micro-object detection in autonomous robotics,” in2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2025, pp. 8187–8193

  40. [40]

    Yolo12: Attention-centric real-time object detectors,

    Y . Tian, Q. Ye, and D. Doermann, “Yolo12: Attention-centric real-time object detectors,”arXiv preprint arXiv:2502.12524, 2025

  41. [41]

    Ultralytics yolo26,

    G. Jocher and J. Qiu, “Ultralytics yolo26,” 2026. [Online]. Available: https://github.com/ultralytics/ultralytics

  42. [42]

    Microsoft coco: Common objects in context,

    T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll´ar, and C. L. Zitnick, “Microsoft coco: Common objects in context,” inProceedings of the European Conference on Computer Vision (ECCV). Springer, 2014, pp. 740–755

  43. [43]

    Clustered object detection in aerial images,

    F. Yang, H. Fan, P. Chu, E. Blasch, and H. Ling, “Clustered object detection in aerial images,” inThe IEEE International Conference on Computer Vision (ICCV), October 2019

  44. [44]

    A global-local self-adaptive network for drone-view object detection,

    S. Deng, S. Li, K. Xie, W. Song, X. Liao, A. Hao, and H. Qin, “A global-local self-adaptive network for drone-view object detection,” IEEE Transactions on Image Processing, vol. 30, pp. 1556–1569, 2021

  45. [45]

    How to fully exploit the abilities of aerial image detectors,

    J. Zhang, J. Huang, X. Chen, and D. Zhang, “How to fully exploit the abilities of aerial image detectors,” inProceedings of the IEEE/CVF International Conference on Computer Vision Workshops, 2019

  46. [46]

    Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection,

    X. Li, W. Wang, L. Wu, S. Chen, X. Hu, J. Li, J. Tang, and J. Yang, “Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection,” inAdvances in Neural Information Processing Systems (NeurIPS), 2020

  47. [47]

    Adaptive sparse convolutional networks with global context enhancement for faster object detection on drone images,

    B. Du, Y . Huang, J. Chen, and D. Huang, “Adaptive sparse convolutional networks with global context enhancement for faster object detection on drone images,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 13 435–13 444

  48. [48]

    Ultralytics yolov8,

    G. Jocher, A. Chaurasia, and J. Qiu, “Ultralytics yolov8,” 2023. [Online]. Available: https://github.com/ultralytics/ultralytics

  49. [49]

    Yolov10: Real-time end-to-end object detection,

    A. Wang, H. Chen, L. Liuet al., “Yolov10: Real-time end-to-end object detection,”arXiv preprint arXiv:2405.14458, 2024