Pith. sign in

hub

Rt-detrv2: Improved base- line with bag-of-freebies for real-time detection transformer

16 Pith papers cite this work. Polarity classification is still indexing.

16 Pith papers citing it
abstract

In this report, we present RT-DETRv2, an improved Real-Time DEtection TRansformer (RT-DETR). RT-DETRv2 builds upon the previous state-of-the-art real-time detector, RT-DETR, and opens up a set of bag-of-freebies for flexibility and practicality, as well as optimizing the training strategy to achieve enhanced performance. To improve the flexibility, we suggest setting a distinct number of sampling points for features at different scales in the deformable attention to achieve selective multi-scale feature extraction by the decoder. To enhance practicality, we propose an optional discrete sampling operator to replace the grid_sample operator that is specific to RT-DETR compared to YOLOs. This removes the deployment constraints typically associated with DETRs. For the training strategy, we propose dynamic data augmentation and scale-adaptive hyperparameters customization to improve performance without loss of speed. Source code and pre-trained models will be available at https://github.com/lyuwenyu/RT-DETR.

hub tools

citation-role summary

background 1 baseline 1

citation-polarity summary

years

2026 15 2025 1

representative citing papers

Modular Diffusion Models for Structured Visual Recognition

cs.CV · 2026-06-21 · unverdicted · novelty 6.0

Modular Diffusion Models decompose diffusion into task-specific modules to model distributions over structured visual outputs for detection, segmentation, and scene graph generation.

YOLOv12: Attention-Centric Real-Time Object Detectors

cs.CV · 2025-02-18 · unverdicted · novelty 6.0

YOLOv12 is a new attention-based real-time object detector that reports higher accuracy than YOLOv10, YOLOv11, and RT-DETR variants at comparable or better speed and efficiency.

RT-SDGOD: Real-Time Single-Domain Generalized Object Detection

cs.CV · 2026-06-08 · unverdicted · novelty 4.0

RT-SDGDet applies one-to-many supervision, Discriminative Evidence Diversity Learning, and Dual-view Evidence Consistency Learning during training to reduce missed detections in real-time object detectors under unseen domain shifts.

Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models

cs.CV · 2026-06-02 · unverdicted · novelty 4.0

YOLO26 presents a unified real-time vision model family with dual-head end-to-end design, new training components, and task-specific heads that reports improved mAP-latency tradeoffs on COCO and LVIS benchmarks across detection, segmentation, pose, and oriented detection.

citing papers explorer

Showing 16 of 16 citing papers.