Pith. sign in

REVIEW 14 cited by

An Open and Comprehensive Pipeline for Unified Object Grounding and Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.02361 v2 pith:YEMGA3H7 submitted 2024-01-04 cs.CV

classification cs.CV
keywords comprehensivedetectiongroundingbaselinedatasetsgrounding-dinommdetectionmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Grounding-DINO is a state-of-the-art open-set detection model that tackles multiple vision tasks including Open-Vocabulary Detection (OVD), Phrase Grounding (PG), and Referring Expression Comprehension (REC). Its effectiveness has led to its widespread adoption as a mainstream architecture for various downstream applications. However, despite its significance, the original Grounding-DINO model lacks comprehensive public technical details due to the unavailability of its training code. To bridge this gap, we present MM-Grounding-DINO, an open-source, comprehensive, and user-friendly baseline, which is built with the MMDetection toolbox. It adopts abundant vision datasets for pre-training and various detection and grounding datasets for fine-tuning. We give a comprehensive analysis of each reported result and detailed settings for reproduction. The extensive experiments on the benchmarks mentioned demonstrate that our MM-Grounding-DINO-Tiny outperforms the Grounding-DINO-Tiny baseline. We release all our models to the research community. Codes and trained models are released at https://github.com/open-mmlab/mmdetection/tree/main/configs/mm_grounding_dino.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Global prompt integration, visual-textual relation distillation and selective fusion make visual prompts discriminative enough for DETR-ViP to beat prior visual-prompt detectors by several mAP points.

  2. 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection

    cs.CV 2025-07 conditional novelty 7.0 of 10

    3D-MOOD is the first end-to-end monocular 3D object detector for open-set classes and novel scenes, achieving SOTA on Omni3D and on new open-set benchmarks.

  3. DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models

    cs.CV 2025-05 conditional novelty 7.0 of 10

    DINO-R1 trains visual-prompt detectors with group-relative query rewards and KL regularization, improving zero-shot and fine-tuned detection over supervised fine-tuning.

  4. Vision-Language Grounding as Bidirectional Concept Correspondence

    cs.CV 2026-08 conditional novelty 6.5 of 10

    Grounding can be treated as recovering all word-span to image-mask correspondences in one pass, and a bridge-token model, ConCor-1, does this better than existing grounding pipelines on the tested benchmarks.

  5. RegionDet: A Benchmark for Region Detection Beyond Object Instances

    cs.CV 2026-08 conditional novelty 6.0 of 10

    RegionDet shows supervised detectors can partially learn to localize activity and context regions, while open-vocabulary detectors nearly fail, indicating current vision-language models are object-centric.

  6. Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A recalibrated GRPO reinforcement learning method lets multimodal LLMs say 'None' for nonexistent referring expressions without sacrificing localization accuracy on objects that do exist.

  7. UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery

    cs.CV 2026-07 conditional novelty 6.0 of 10

    UniRef-UAV defines multimodal universal referring for UAV scenes and shows a detection-style baseline with stronger no-target control than large MLLMs.

  8. FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A two-stage bilingual CLIP-style model with region-text supervision and a new text-side contrastive loss outperforms prior open models on fine-grained vision-language tasks in English and Chinese.

  9. LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Co-training an open-vocabulary detector with long, detailed image captions generated by a large vision-language model improves zero-shot detection, especially for rare classes.

  10. Multi-Point Positional Insertion Tuning for Small Object Detection

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Multi-point Positional Insertion tuning matches CoOp and VPT accuracy on small-object detection with 0.5M learnable parameters instead of 12M.

  11. SemSegBench & DetecBench: Benchmarking Reliability and Generalization Beyond Classification

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A large-scale benchmark of 76 segmentation and 61 detection models shows that robustness to attacks and corruptions does not reliably track clean accuracy, and that transformer backbones generalize better under shift.

  12. UNCOVER: Unknown Class Object Detection for Autonomous Vehicles in Real-time

    cs.CV 2024-12 conditional novelty 5.0 of 10

    UNCOVER extends a real-time YOLO detector with an occupancy score, an extra OOD class, and a depth-based filter to catch objects outside the standard traffic classes.

  13. DynamicEarth: How Far are We from Open-Vocabulary Change Detection?

    cs.CV 2025-01 reject novelty 4.0 of 10

    The paper shows that composing mask proposal, feature comparison, and open-vocabulary classification models can detect arbitrary-category changes in satellite images without training.

  14. RoboCup@Home 2024 OPL Winner NimbRo: Anthropomorphic Service Robots using Foundation Models for Perception and Planning

    cs.RO 2024-12 conditional novelty 4.0 of 10

    NimbRo's TIAGo++ robot won RoboCup@Home 2024 OPL using open-vocabulary segmentation and LLM-based task planning.

Pith tools