Pith. sign in

REVIEW 2 cited by

Object Detection with Transformers: A Review

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.04670 v3 pith:D7PPQ7AK submitted 2023-06-07 cs.CV

classification cs.CV
keywords detectiontransformersdetrobjectperformanceimprovementsproposedresearchers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The astounding performance of transformers in natural language processing (NLP) has motivated researchers to explore their applications in computer vision tasks. DEtection TRansformer (DETR) introduces transformers to object detection tasks by reframing detection as a set prediction problem. Consequently, eliminating the need for proposal generation and post-processing steps. Initially, despite competitive performance, DETR suffered from slow training convergence and ineffective detection of smaller objects. However, numerous improvements are proposed to address these issues, leading to substantial improvements in DETR and enabling it to exhibit state-of-the-art performance. To our knowledge, this is the first paper to provide a comprehensive review of 21 recently proposed advancements in the original DETR model. We dive into both the foundational modules of DETR and its recent enhancements, such as modifications to the backbone structure, query design strategies, and refinements to attention mechanisms. Moreover, we conduct a comparative analysis across various detection transformers, evaluating their performance and network architectures. We hope that this study will ignite further interest among researchers in addressing the existing challenges and exploring the application of transformers in the object detection domain. Readers interested in the ongoing developments in detection transformers can refer to our website at: https://github.com/mindgarage-shan/trans_object_detection_survey

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 19 citations worldwide. Full citation record

  1. A multi-modal dataset for insect biodiversity with imagery and DNA at the trap and individual level

    cs.CV 2025-07 conditional novelty 7.0 of 10

    The authors present a multi-modal dataset of 45 Malaise-trap insect samples, pairing bulk images with individual-level DNA barcoding and per-specimen segmentation masks, plus benchmark instance segmentation results.

  2. Exploring Visual Embedding Spaces Induced by Vision Transformers for Online Auto Parts Marketplaces

    cs.CV 2025-02 conditional novelty 3.0 of 10

    Visual embeddings from a pretrained ViT produce poorly separated clusters of auto parts images (silhouette 0.015), far below the 0.38 reported for a multimodal model on similar data.

Pith tools