REVIEW 14 cited by
An Open and Comprehensive Pipeline for Unified Object Grounding and Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Grounding-DINO is a state-of-the-art open-set detection model that tackles multiple vision tasks including Open-Vocabulary Detection (OVD), Phrase Grounding (PG), and Referring Expression Comprehension (REC). Its effectiveness has led to its widespread adoption as a mainstream architecture for various downstream applications. However, despite its significance, the original Grounding-DINO model lacks comprehensive public technical details due to the unavailability of its training code. To bridge this gap, we present MM-Grounding-DINO, an open-source, comprehensive, and user-friendly baseline, which is built with the MMDetection toolbox. It adopts abundant vision datasets for pre-training and various detection and grounding datasets for fine-tuning. We give a comprehensive analysis of each reported result and detailed settings for reproduction. The extensive experiments on the benchmarks mentioned demonstrate that our MM-Grounding-DINO-Tiny outperforms the Grounding-DINO-Tiny baseline. We release all our models to the research community. Codes and trained models are released at https://github.com/open-mmlab/mmdetection/tree/main/configs/mm_grounding_dino.
Forward citations
Cited by 14 Pith papers
-
DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts
Global prompt integration, visual-textual relation distillation and selective fusion make visual prompts discriminative enough for DETR-ViP to beat prior visual-prompt detectors by several mAP points.
-
3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
3D-MOOD is the first end-to-end monocular 3D object detector for open-set classes and novel scenes, achieving SOTA on Omni3D and on new open-set benchmarks.
-
DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
DINO-R1 trains visual-prompt detectors with group-relative query rewards and KL regularization, improving zero-shot and fine-tuned detection over supervised fine-tuning.
-
Vision-Language Grounding as Bidirectional Concept Correspondence
Grounding can be treated as recovering all word-span to image-mask correspondences in one pass, and a bridge-token model, ConCor-1, does this better than existing grounding pipelines on the tested benchmarks.
-
RegionDet: A Benchmark for Region Detection Beyond Object Instances
RegionDet shows supervised detectors can partially learn to localize activity and context regions, while open-vocabulary detectors nearly fail, indicating current vision-language models are object-centric.
-
Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO
A recalibrated GRPO reinforcement learning method lets multimodal LLMs say 'None' for nonexistent referring expressions without sacrificing localization accuracy on objects that do exist.
-
UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery
UniRef-UAV defines multimodal universal referring for UAV scenes and shows a detection-style baseline with stronger no-target control than large MLLMs.
-
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
A two-stage bilingual CLIP-style model with region-text supervision and a new text-side contrastive loss outperforms prior open models on fine-grained vision-language tasks in English and Chinese.
-
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
Co-training an open-vocabulary detector with long, detailed image captions generated by a large vision-language model improves zero-shot detection, especially for rare classes.
-
Multi-Point Positional Insertion Tuning for Small Object Detection
Multi-point Positional Insertion tuning matches CoOp and VPT accuracy on small-object detection with 0.5M learnable parameters instead of 12M.
-
SemSegBench & DetecBench: Benchmarking Reliability and Generalization Beyond Classification
A large-scale benchmark of 76 segmentation and 61 detection models shows that robustness to attacks and corruptions does not reliably track clean accuracy, and that transformer backbones generalize better under shift.
-
UNCOVER: Unknown Class Object Detection for Autonomous Vehicles in Real-time
UNCOVER extends a real-time YOLO detector with an occupancy score, an extra OOD class, and a depth-based filter to catch objects outside the standard traffic classes.
-
DynamicEarth: How Far are We from Open-Vocabulary Change Detection?
The paper shows that composing mask proposal, feature comparison, and open-vocabulary classification models can detect arbitrary-category changes in satellite images without training.
-
RoboCup@Home 2024 OPL Winner NimbRo: Anthropomorphic Service Robots using Foundation Models for Perception and Planning
NimbRo's TIAGo++ robot won RoboCup@Home 2024 OPL using open-vocabulary segmentation and LLM-based task planning.
Discussion (0). Continue with ORCID to comment.