REVIEW 9 cited by
Real-time Transformer-based Open-Vocabulary Detection with Efficient Fusion Head
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
End-to-end transformer-based detectors (DETRs) have shown exceptional performance in both closed-set and open-vocabulary object detection (OVD) tasks through the integration of language modalities. However, their demanding computational requirements have hindered their practical application in real-time object detection (OD) scenarios. In this paper, we scrutinize the limitations of two leading models in the OVDEval benchmark, OmDet and Grounding-DINO, and introduce OmDet-Turbo. This novel transformer-based real-time OVD model features an innovative Efficient Fusion Head (EFH) module designed to alleviate the bottlenecks observed in OmDet and Grounding-DINO. Notably, OmDet-Turbo-Base achieves a 100.2 frames per second (FPS) with TensorRT and language cache techniques applied. Notably, in zero-shot scenarios on COCO and LVIS datasets, OmDet-Turbo achieves performance levels nearly on par with current state-of-the-art supervised models. Furthermore, it establishes new state-of-the-art benchmarks on ODinW and OVDEval, boasting an AP of 30.1 and an NMS-AP of 26.86, respectively. The practicality of OmDet-Turbo in industrial applications is underscored by its exceptional performance on benchmark datasets and superior inference speed, positioning it as a compelling choice for real-time object detection tasks. Code: \url{https://github.com/om-ai-lab/OmDet}
Forward citations
Cited by 9 Pith papers
-
Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
MoE fine-tuning with decomposed pre-trained FFN experts lets a real-time open-vocabulary detector beat a much larger-data baseline with similar active parameter count.
-
DDStereo: Efficient Dual Decoder Transformers for Stereo 3D Road Anomaly Detection
DDStereo uses two lightweight decoder branches sharing object queries plus a compact disparity extractor to deliver SOTA closed- and open-set accuracy with real-time inference on stereo 3D benchmarks.
-
Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
A multi-stage diffusion-based framework that generates labeled synthetic aerial images from weak image-level labels improves cross-domain vehicle detection AP50 over prior adaptation methods.
-
Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes?
Inpainting unusual objects into street scenes reveals that open-vocabulary detectors miss objects based on image location rather than object semantics.
-
Mcity Data Engine: Iterative Model Improvement Through Open-Vocabulary Data Selection
The Mcity Data Engine provides an open-source, end-to-end data engine whose ensemble of open-vocabulary detectors selected 100 frames that improved VRU detection mAP@0.5 by 17.45% in one iteration.
-
CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object Detection
CP-DETR combines progressive and gated multi-scale prompt-image fusion with text, visual, and optimized prompts to reach state-of-the-art universal detection with one model weight.
-
C-GAP: Class-Aware and Online Prompting Improves Vision-Language Models on Imbalanced Classes
Iterative LLM caption refinement guided by minority-class AP@0.5 lifts rare-object detection on frozen open-vocabulary detectors without labels or weight updates.
-
Towards Real-Time Open-Vocabulary Video Instance Segmentation
TROY-VIS makes open-vocabulary video instance segmentation run in real time (around 25 FPS) while keeping accuracy comparable to or better than slower state-of-the-art models.
-
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
DINO-X Pro sets new zero-shot detection records on COCO and LVIS, with large gains on rare classes, by scaling grounding pre-training and adding multiple prompt types and perception heads.
Discussion (0). Continue with ORCID to comment.