Pith. sign in

REVIEW 28 cited by

YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.13616 v2 pith:AXDSLDV6 submitted 2024-02-21 cs.CV

YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information

classification cs.CV
keywords informationresultsgelangradientmodelsachievearchitecturedata
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Today's deep learning methods focus on how to design the most appropriate objective functions so that the prediction results of the model can be closest to the ground truth. Meanwhile, an appropriate architecture that can facilitate acquisition of enough information for prediction has to be designed. Existing methods ignore a fact that when input data undergoes layer-by-layer feature extraction and spatial transformation, large amount of information will be lost. This paper will delve into the important issues of data loss when data is transmitted through deep networks, namely information bottleneck and reversible functions. We proposed the concept of programmable gradient information (PGI) to cope with the various changes required by deep networks to achieve multiple objectives. PGI can provide complete input information for the target task to calculate objective function, so that reliable gradient information can be obtained to update network weights. In addition, a new lightweight network architecture -- Generalized Efficient Layer Aggregation Network (GELAN), based on gradient path planning is designed. GELAN's architecture confirms that PGI has gained superior results on lightweight models. We verified the proposed GELAN and PGI on MS COCO dataset based object detection. The results show that GELAN only uses conventional convolution operators to achieve better parameter utilization than the state-of-the-art methods developed based on depth-wise convolution. PGI can be used for variety of models from lightweight to large. It can be used to obtain complete information, so that train-from-scratch models can achieve better results than state-of-the-art models pre-trained using large datasets, the comparison results are shown in Figure 1. The source codes are at: https://github.com/WongKinYiu/yolov9.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 28 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Event-based Gaze Control System for Accurate Real-time Spin Estimation in Professional Ball Games

    cs.CV 2026-06 unverdicted novelty 7.0

    An event-camera system with active gaze control and contrast-maximization spin estimation achieves real-time performance in table tennis with 8.8% magnitude error, 6.4° axis error, 3 ms latency, and 750 Hz throughput.

  2. Event-based Gaze Control System for Accurate Real-time Spin Estimation in Professional Ball Games

    cs.CV 2026-06 unverdicted novelty 7.0

    Event-based gaze control system with s-CMax offline spin estimation and CNN online refinement achieves 8.8% magnitude error and 3 ms latency on professional table tennis matches.

  3. Gaze-DETR: Top-Down Guidance Through Priority Maps for Infrared Weak-Small UAV Detection with DETR

    cs.CV 2026-07 conditional novelty 6.0

    Supervising a pre-localization priority map, whether from boxes, real gaze, or transferred pseudo-gaze, improves infrared weak-small UAV detection over DINO-DETR.

  4. Dual-Correlation Hypergraph Network for Unaligned RGBT Video Object Detection and A Large-scale Benchmark

    cs.CV 2026-07 conditional novelty 6.0

    DHNet with patch alignment and dual hypergraph fusion reaches SOTA RGBT video object detection on VT-VOD50 and the new large-scale DVT-VOD1000 benchmark.

  5. InfraNet: Quality-Aware RGB Guidance for Efficient Infrared Object Detection

    cs.CV 2026-07 conditional novelty 6.0

    QualGate-regulated RGB guidance during training produces efficient IR-only and dual-modal detectors that match or beat equal-fusion baselines under low light and adverse weather.

  6. Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and Baseline

    cs.AI 2026-06 unverdicted novelty 6.0

    Presents MMIO benchmark and RTVP method achieving state-of-the-art 42.2% AP in zero-shot industrial defect detection.

  7. RefDiffNet: Learning to Expose Subtle PCB Defects Before Detection

    cs.CV 2026-05 unverdicted novelty 6.0

    RefDiffNet is a lightweight input enhancement block that uses reference image comparison to expose PCB defects, delivering up to 18% relative mAP50:95 gains across YOLO, RT-DETR, and Faster R-CNN detectors with 0.004-...

  8. Small Object Detection in Industrial Recycling: A New Dataset and YOLO Performance Evaluation

    cs.CV 2026-05 unverdicted novelty 6.0

    Releases a recycling-specific dataset of >10k images and evaluates YOLO variants on small dense overlapping objects with augmentation and anomaly detection.

  9. AnyDepth-DETR/-YOLO: Any-depth object detection with a single network

    cs.CV 2026-05 unverdicted novelty 6.0

    A single network achieves any-depth object detection by splitting stages into always-executed essential paths and skippable refinement paths, trained via self-distillation on the full and minimal extremes to maintain ...

  10. Street-Legal Physical-World Adversarial Rim for License Plates

    cs.CV 2026-04 conditional novelty 6.0

    SPAR is a street-legal physical rim that cuts modern ALPR accuracy by 60% and reaches 18% targeted impersonation while costing under $100 and requiring no plate modification.

  11. YOLOv12: Attention-Centric Real-Time Object Detectors

    cs.CV 2025-02 unverdicted novelty 6.0

    YOLOv12 is a new attention-based real-time object detector that reports higher accuracy than YOLOv10, YOLOv11, and RT-DETR variants at comparable or better speed and efficiency.

  12. Hippocampus-DETR: An Explicit Memory Object Detection Framework Based on Hippocampus Modeling

    cs.CV 2026-06 unverdicted novelty 5.0

    Hippocampus-DETR integrates a hippocampal memory network (HipNet) into DETR to simulate brain subregions for pattern separation, completion, and improved detection accuracy plus generalization.

  13. Fully Distributed Multi-View 3D Tracking in Real-Time

    cs.CV 2026-06 unverdicted novelty 5.0

    MV3DT is a fully distributed real-time multi-view 3D tracking framework achieving 94.3% IDF1 on WILDTRACK while scaling to 100 cameras at 30 FPS with under 10 ms latency and 2.2% communication overhead in a zero-shot regime.

  14. Collaborative Space Object Detection with Multi-Satellite Viewpoints in LEO Constellations

    cs.CV 2026-06 conditional novelty 5.0

    Multi-view fusion from three satellite viewpoints boosts mAP50 by up to 36.3% and mAP50-95 by 46.5% over single-view baselines in YOLO detectors for space object detection.

  15. Low-Cost Stereo Vision for Robust 3D Positioning of Thin Radiata Pine Branches in Autonomous Drone Pruning

    cs.CV 2026-05 unverdicted novelty 5.0

    A drone-mounted stereo camera pipeline with YOLO segmentation, deep stereo depth, centroid triangulation, and MAD outlier rejection achieves robust 3D positioning of thin pine branches at 1-2 m distances.

  16. A Real-time Scale-robust Network for Glottis Segmentation in Nasal Transnasal Intubation

    eess.IV 2026-04 unverdicted novelty 5.0

    A scale-robust lightweight CNN for glottis segmentation achieves 92.9% mDice at over 170 FPS with a 19 MB model size on three datasets.

  17. DocRevive: A Unified Pipeline for Document Text Restoration

    cs.CV 2026-04 unverdicted novelty 5.0

    DocRevive builds a unified pipeline using OCR, image analysis, language models, and diffusion to reconstruct degraded document text, backed by a 30k-image synthetic dataset and the UCSM metric.

  18. DocRevive: A Unified Pipeline for Document Text Restoration

    cs.CV 2026-04 unverdicted novelty 5.0

    A unified pipeline using OCR, inpainting, and diffusion models restores text in degraded documents on a new synthetic benchmark dataset, evaluated with the proposed UCSM metric.

  19. Deep Learning for Accurate Vision-based Catch Composition in Tropical Tuna Purse Seiners

    cs.CV 2025-11 conditional novelty 5.0

    YOLOv9+SAM2 segmentation with hierarchical classification estimates tuna catch composition from EM video with about 4.5% mean absolute error on controlled test operations.

  20. UAVDB: Point-Guided Masks for UAV Detection and Segmentation

    cs.CV 2024-09 unverdicted novelty 5.0

    Introduces UAVDB dataset for UAV detection/segmentation via PIC point-to-box conversion and SAM2 masks, with YOLO baselines showing PIC+SAM2 outperforms prior annotation methods on IoU.

  21. An Intelligent-Cloud Edge Multimodal Interaction System for Robots

    cs.RO 2026-07 conditional novelty 4.0

    A cloud-edge robot system combined a CBAM/DIoU-enhanced YOLO11n gesture detector with LLM/VLM agents, reporting 95–98.9% precision and 82–95% task success on small, validation-based evaluations.

  22. Hierarchically Decoupled Mixture-of-Experts for Robust Traffic Sign Recognition in Complex Driving Scenarios

    cs.CV 2026-06 unverdicted novelty 4.0

    A hierarchically decoupled heterogeneous MoE framework with YOLO experts and lightweight gating network reports 76.8% mAP50-95 on a composite traffic sign dataset, a 2.3% gain over baseline with 39.4% lower compute.

  23. PaveSync: A Unified and Comprehensive Dataset for Pavement Distress Analysis and Classification

    cs.CV 2025-12 conditional novelty 4.0

    PaveSync merges existing pavement imagery into a standardized 52,747-image, 13-class detection dataset and benchmarks seven object detectors on it.

  24. DFIR-DETR: Frequency-Domain Iterative Refinement and Dynamic Feature Aggregation for Small Object Detection

    cs.CV 2025-12 unverdicted novelty 4.0

    DFIR-DETR augments RT-DETR with frequency-domain iterative refinement and dynamic feature aggregation, reporting 92.9% mAP50 on NEU-DET and 51.6% on VisDrone at 11.7M parameters and 47.2 GFLOPs.

  25. AI-assisted radiographic analysis in detecting alveolar bone-loss severity and patterns

    cs.CV 2025-06 unverdicted novelty 4.0

    A deep learning pipeline with YOLOv8 and Keypoint R-CNN achieves ICC up to 0.80 for bone loss severity and 87% accuracy for horizontal vs. angular pattern classification on 1000 annotated IOPA radiographs.

  26. Positioning radiata pine branches requiring pruning by drone stereo vision

    cs.CV 2026-04 unverdicted novelty 3.0

    Drone stereo vision pipeline segments pine branches with YOLO variants and estimates depth with deep stereo networks, yielding more coherent maps than SGBM at 1-2 m distances.

  27. YOLOv8 to YOLO11: A Comprehensive Architecture In-depth Comparative Review

    cs.CV 2025-01 unverdicted novelty 2.0

    Comparative review of YOLOv8 to YOLO11 architectures based on papers, docs, and code inspection, noting incremental improvements and some unchanged blocks.

  28. YOLOv11: An Overview of the Key Architectural Enhancements

    cs.CV 2024-10 unverdicted novelty 1.0

    YOLOv11 adds blocks such as C3k2, SPPF, and C2PSA to improve feature extraction, mAP, and efficiency while supporting detection, segmentation, pose, and oriented detection across model sizes.