Pith. sign in

REVIEW 3 cited by

A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.00501 v7 pith:FRFVNPI4 submitted 2023-04-02 cs.CV

classification cs.CV
keywords yolocomprehensivedetectionobjectreal-timeyolo-nasyolov8analysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

YOLO has become a central real-time object detection system for robotics, driverless cars, and video monitoring applications. We present a comprehensive analysis of YOLO's evolution, examining the innovations and contributions in each iteration from the original YOLO up to YOLOv8, YOLO-NAS, and YOLO with Transformers. We start by describing the standard metrics and postprocessing; then, we discuss the major changes in network architecture and training tricks for each model. Finally, we summarize the essential lessons from YOLO's development and provide a perspective on its future, highlighting potential research directions to enhance real-time object detection systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VEIGAR: View-consistent Explicit Inpainting and Geometry Alignment for 3D object Removal

    cs.GR 2025-06 conditional novelty 5.0 of 10

    VEIGAR is a pipeline for 3D object removal in Gaussian Splatting that uses deep stereo depth projection and a scale-invariant depth loss to achieve faster training and comparable quality to prior state-of-the-art.

  2. Advancing from Automated to Autonomous Beamline by Leveraging Computer Vision

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A multi-camera computer vision system using segmentation, tracking, and pixel-distance checks detects beamline collisions in real time, reporting 99.8% accuracy on a small in-house dataset.

  3. SHeRL-FL: When Representation Learning Meets Split Learning in Hierarchical Federated Learning

    cs.LG 2025-08 unverdicted novelty 2.0 of 10

    The submitted body is an unrelated survey, not the SHeRL-FL method claimed in the metadata.

Pith tools