Pith. sign in

REVIEW 13 cited by

What is YOLOv5: A deep look into the internal features of the popular object detector

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.20892 v1 pith:C75B2AGW submitted 2024-07-30 cs.CV

What is YOLOv5: A deep look into the internal features of the popular object detector

classification cs.CV
keywords modelobjectyolov5detectionperformancepopularacrossadditionally
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

This study presents a comprehensive analysis of the YOLOv5 object detection model, examining its architecture, training methodologies, and performance. Key components, including the Cross Stage Partial backbone and Path Aggregation-Network, are explored in detail. The paper reviews the model's performance across various metrics and hardware platforms. Additionally, the study discusses the transition from Darknet to PyTorch and its impact on model development. Overall, this research provides insights into YOLOv5's capabilities and its position within the broader landscape of object detection and why it is a popular choice for constrained edge deployment scenarios.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Zero-Shot Quantization for Object Detectors using Off-the-Shelf Generative Models

    cs.LG 2026-06 unverdicted novelty 7.0

    GoodQ uses generative models with information-dense prompting, distribution-aware selection, and teacher-guided noise reduction to achieve SOTA low-bit (W4A4) and extreme-bit (W3A3) zero-shot quantization for object d...

  2. Tetris: Tile-level Sampling for Efficient and High-Fidelity Video Object Tracking

    cs.CV 2026-05 conditional novelty 7.0

    Tetris uses tile-level polyomino sampling and packing to materialize object tracks from stationary video with up to 68.8x throughput gain at under 5% HOTA loss.

  3. LEVIRDet: A Million-Scale 159-Category Dataset and Foundation Model for Universal Remote Sensing Object Detection

    cs.CV 2026-06 conditional novelty 6.5

    A million-box 159-class remote-sensing dataset plus a GSD- and hierarchy-aware detector yields ~5 mAP average gains over fully supervised baselines on nine external benchmarks with no target training.

  4. Zero-Shot Quantization for Object Detectors using Off-the-Shelf Generative Models

    cs.LG 2026-06 conditional novelty 6.0

    Diffusion-generated, distribution-matched synthetic images enable zero-shot quantized object detectors to outperform prior zero-shot methods and even real-data QAT at 4-bit and 3-bit precision.

  5. LEVIRDet: A Million-Scale 159-Category Dataset and Foundation Model for Universal Remote Sensing Object Detection

    cs.CV 2026-06 unverdicted novelty 6.0

    LEVIRDet-159 is a 159-category remote sensing detection dataset with 2.56M boxes exceeding prior scales; LEVIRDetNet achieves SOTA zero-shot performance on 9 external benchmarks with 5.02 mAP average improvement.

  6. Tetris: Tile-level Sampling for Efficient and High-Fidelity Video Object Tracking

    cs.CV 2026-05 unverdicted novelty 6.0

    Tetris decomposes stationary videos into tile polyominoes and applies classifier plus ILP pruning to cut detector calls, staying within 5% accuracy loss while delivering up to 17.4x throughput gains over priors.

  7. Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation

    cs.CV 2026-04 unverdicted novelty 6.0

    A pair-centric set-prediction model unifies present HOI detection and multi-horizon anticipation in video by modeling future interactions as residual transitions from current pair states, backed by a temporally correc...

  8. FSDC-DETR: A Frequency-Spatial Domain Collaborative DETR for Small Object Detection

    cs.CV 2026-07 accept novelty 5.5

    FSDC-DETR improves small-object AP by 6.8–6.9 points on VisDrone and AITODv2 by explicit frequency-spatial fusion and wavelet-style downsampling inside a DETR hybrid encoder.

  9. ISAC and Vision Fusion for Fine-Grained Low-Altitude Target Recognition

    eess.SP 2026-07 conditional novelty 5.0

    ISAC-guided PTZ imaging plus cGAN-denoised micro-Doppler fused with MobileViT reaches ~97.7% average accuracy on a synthetic 10-class UAV/bird dataset, beating single-modality baselines.

  10. FSDC-DETR: A Frequency-Spatial Domain Collaborative DETR for Small Object Detection

    cs.CV 2026-07 conditional novelty 5.0

    FSDC-DETR improves small object detection by explicitly modeling frequency-spatial representations through dual-branch adaptive fusion, shunt feature fusion, and wavelet-based dynamic downsampling, achieving state-of-...

  11. A Stereo Visual SLAM System Using Object-Level Motion Estimation and Geometric Filtering Based on Cross Disparity

    cs.RO 2026-07 unverdicted novelty 5.0

    OCD SLAM adds cross-disparity inconsistency checks and object-level motion classification to ORB-SLAM2, reporting better trajectory accuracy than prior dynamic SLAM methods on KITTI sequences.

  12. A Marine Debris Detection Framework for Ocean Robots via Self-Attention Enhancement and Feature Interaction Optimization

    cs.CV 2026-05 unverdicted novelty 4.0

    YOLO-MD improves underwater marine debris detection by adding a Dual-Branch Convolutional Enhanced Self-Attention module, a lightweight shift operation, and SFG-Loss for class imbalance, achieving 0.875 precision and ...

  13. YOLOv11: An Overview of the Key Architectural Enhancements

    cs.CV 2024-10 unverdicted novelty 1.0

    YOLOv11 adds blocks such as C3k2, SPPF, and C2PSA to improve feature extraction, mAP, and efficiency while supporting detection, segmentation, pose, and oriented detection across model sizes.