Pith. sign in

REVIEW 16 cited by

D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.13842 v1 pith:3LUDM6QH submitted 2024-10-17 cs.CV

D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement

classification cs.CV
keywords d-finelocalizationfine-grainedmodelsregressionaccuracyachievesd-fine-l
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce D-FINE, a powerful real-time object detector that achieves outstanding localization precision by redefining the bounding box regression task in DETR models. D-FINE comprises two key components: Fine-grained Distribution Refinement (FDR) and Global Optimal Localization Self-Distillation (GO-LSD). FDR transforms the regression process from predicting fixed coordinates to iteratively refining probability distributions, providing a fine-grained intermediate representation that significantly enhances localization accuracy. GO-LSD is a bidirectional optimization strategy that transfers localization knowledge from refined distributions to shallower layers through self-distillation, while also simplifying the residual prediction tasks for deeper layers. Additionally, D-FINE incorporates lightweight optimizations in computationally intensive modules and operations, achieving a better balance between speed and accuracy. Specifically, D-FINE-L / X achieves 54.0% / 55.8% AP on the COCO dataset at 124 / 78 FPS on an NVIDIA T4 GPU. When pretrained on Objects365, D-FINE-L / X attains 57.1% / 59.3% AP, surpassing all existing real-time detectors. Furthermore, our method significantly enhances the performance of a wide range of DETR models by up to 5.3% AP with negligible extra parameters and training costs. Our code and pretrained models: https://github.com/Peterande/D-FINE.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ContextShift: A Controlled Benchmark for Context Dependence in Object Detection

    cs.CV 2026-06 conditional novelty 7.0

    ContextShift benchmark on COCO reveals up to 227% more false negatives and 44% fewer predictions under controlled context changes, non-monotonic NPMI response, and gains from context-aware augmentation.

  2. Rethinking Event-Based Object Dtection through Representation-Level Temporal Aggregation and Model-Level Hypergraph Reasoning

    cs.CV 2026-05 unverdicted novelty 7.0

    Ev-DTAD improves event-based object detection accuracy and speed by using hierarchical temporal aggregation at the representation level and frequency-aware hypergraph fusion at the model level.

  3. Training-Free Semantic Multi-Object Tracking with Vision-Language Models

    cs.CV 2026-04 conditional novelty 7.0

    TF-SMOT composes pretrained vision-language models into a training-free pipeline that reaches state-of-the-art tracking and improved summary quality on the BenSMOT benchmark.

  4. WUTDet: A 100K-Scale Ship Detection Dataset and Benchmarks with Dense Small Objects

    cs.CV 2026-04 unverdicted novelty 7.0

    WUTDet is a 100K-image ship detection dataset with benchmarks indicating Transformer models outperform CNN and Mamba architectures in accuracy and small-object detection for complex maritime environments.

  5. SARES-DEIM: Sparse Mixture-of-Experts Meets DETR for Robust SAR Ship Detection

    cs.CV 2026-04 unverdicted novelty 7.0

    SARES-DEIM achieves 76.4% mAP50:95 and 93.8% mAP50 on HRSID by routing SAR features through sparse frequency and wavelet experts plus a high-resolution preservation neck, outperforming prior YOLO and SAR detectors.

  6. MORE: A Multilingual Document Parsing Benchmark and Evaluation

    cs.CV 2026-07 conditional novelty 6.0

    MORE provides a 149-language, structure-aware document parsing benchmark from real PDFs and reports baselines showing specialized OCR models still fail on tables and rare scripts.

  7. From Spatial to Spectral: An Efficient, Frequency-Guided Feature Representation Learner for Small Object Detection

    cs.CV 2026-06 unverdicted novelty 6.0

    Proposes DERNet with Decompose-Enhance-Reconstruct operator and three plug-and-play modules to shift small object detection from spatial to spectral feature processing, claiming better performance than YOLOv11 with 1/...

  8. Rethinking Event-Based Object Dtection through Representation-Level Temporal Aggregation and Model-Level Hypergraph Reasoning

    cs.CV 2026-05 unverdicted novelty 6.0

    Ev-DTAD combines hierarchical temporal aggregation into a pseudo-RGB representation with frequency-aware hypergraph fusion to improve accuracy and speed in event-based object detection on Gen1, Gen4, and eTraM datasets.

  9. Rethinking Event-Based Object Dtection through Representation-Level Temporal Aggregation and Model-Level Hypergraph Reasoning

    cs.CV 2026-05 unverdicted novelty 6.0

    Ev-DTAD introduces hierarchical temporal aggregation for event representation and frequency-aware hypergraph fusion for feature reasoning, delivering accuracy and speed gains on Gen1, 1Mpx/Gen4, and eTraM event detect...

  10. ZoomSpec: A Physics-Guided Coarse-to-Fine Framework for Wideband Spectrum Sensing

    cs.CV 2026-04 unverdicted novelty 6.0

    ZoomSpec achieves 78.1 mAP@0.5:0.95 on the SpaceNet dataset by combining log-space STFT, a coarse proposal net, adaptive heterodyne filtering, and dual-domain fine recognition to improve narrowband visibility in wideb...

  11. Sparse Hypergraph-Enhanced Frame-Event Object Detection with Fine-Grained MoE

    cs.CV 2026-04 unverdicted novelty 6.0

    Hyper-FEOD fuses RGB and event data via sparse hypergraph cross-modal fusion and region-specialized MoE experts to improve accuracy-efficiency in object detection.

  12. PS-MOT: Cultivating Instance Awareness from Point Seeds for Multi-Object Tracking

    cs.CV 2026-06 unverdicted novelty 5.0

    PS-Track sets a new state-of-the-art for point-supervised multi-object tracking by converting point seeds into temporally consistent pseudo-labels via Temporal-Feedback Prompting, Point-Excited Wavelet Attention, and ...

  13. M^2C-EvDet: Multi-Domain Multi-Order Cross-Modal Knowledge Distillation for Event-based Object Detection

    cs.CV 2026-06 unverdicted novelty 5.0

    M^2C-EvDet proposes Adaptive Frequency-Decoupled Feature Distillation (AF^2D^2) and Multi-Order Relational Distillation (MORD) modules to reduce the performance gap between event-based and frame-based object detection.

  14. When Detectors Forget Forensics: Blocking Semantic Shortcuts for Generalizable AI-Generated Image Detection

    cs.CV 2026-03 unverdicted novelty 5.0

    Forensic fine-tuning of vision foundation models leaves semantic structure intact ("semantic fallback"); suppressing CLIP-estimated semantic subspaces via SVD is claimed to yield more generalizable AI-image detectors.

  15. FMC-DETR: Frequency-Decoupled Multi-Domain Coordination for Aerial-View Object Detection

    cs.CV 2025-09 unverdicted novelty 5.0

    FMC-DETR proposes a frequency-decoupled fusion framework with WeKat backbone, MDFC coordination, and CPF fusion modules that claims state-of-the-art results on remote sensing object detection benchmarks.

  16. RT-SDGOD: Real-Time Single-Domain Generalized Object Detection

    cs.CV 2026-06 unverdicted novelty 4.0

    RT-SDGDet applies one-to-many supervision, Discriminative Evidence Diversity Learning, and Dual-view Evidence Consistency Learning during training to reduce missed detections in real-time object detectors under unseen...