Pith. sign in

REVIEW 30 cited by

Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.10300 v2 pith:EVLFUC5U submitted 2024-05-16 cs.CV

Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

classification cs.CV
keywords groundingdinomodeledgedetectionobjectopen-setbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This paper introduces Grounding DINO 1.5, a suite of advanced open-set object detection models developed by IDEA Research, which aims to advance the "Edge" of open-set object detection. The suite encompasses two models: Grounding DINO 1.5 Pro, a high-performance model designed for stronger generalization capability across a wide range of scenarios, and Grounding DINO 1.5 Edge, an efficient model optimized for faster speed demanded in many applications requiring edge deployment. The Grounding DINO 1.5 Pro model advances its predecessor by scaling up the model architecture, integrating an enhanced vision backbone, and expanding the training dataset to over 20 million images with grounding annotations, thereby achieving a richer semantic understanding. The Grounding DINO 1.5 Edge model, while designed for efficiency with reduced feature scales, maintains robust detection capabilities by being trained on the same comprehensive dataset. Empirical results demonstrate the effectiveness of Grounding DINO 1.5, with the Grounding DINO 1.5 Pro model attaining a 54.3 AP on the COCO detection benchmark and a 55.7 AP on the LVIS-minival zero-shot transfer benchmark, setting new records for open-set object detection. Furthermore, the Grounding DINO 1.5 Edge model, when optimized with TensorRT, achieves a speed of 75.2 FPS while attaining a zero-shot performance of 36.2 AP on the LVIS-minival benchmark, making it more suitable for edge computing scenarios. Model examples and demos with API will be released at https://github.com/IDEA-Research/Grounding-DINO-1.5-API

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 30 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Comprehensive Ecosystem for Open-Domain Customized Video Generation

    cs.CV 2026-06 unverdicted novelty 7.0

    Introduces PexelsCustom-1M dataset, CustoMDiT parameter-efficient model, and OpenCustom benchmark for open-domain customized video generation.

  2. WHU-Infra3D: A Full-stack Multi-modal Dataset and Benchmark for 3D Roadside Infrastructure Inventory

    cs.CV 2026-06 unverdicted novelty 7.0

    WHU-Infra3D is a new large-scale multi-modal dataset and benchmark for 3D roadside infrastructure inventory, providing over 175k 2D boxes, thousands of 3D instances, and 181k annotations across five core tasks while e...

  3. FlowOVD: Learning Generative Latent Flows for Zero-shot Open-vocabulary Detection

    cs.CV 2026-05 unverdicted novelty 7.0

    FlowOVD applies rectified flow to generate continuous latent query dynamics for text-conditioned open-vocabulary detection, reporting 49.5 AP on COCO and 31.5 AP on LVIS.

  4. Vision Harnessing Agent for Open Ad-hoc Segmentation

    cs.CV 2026-05 unverdicted novelty 7.0

    VASA is a vision-guided agent for open ad-hoc segmentation that creates and validates masks through planning, tool use, and error recovery, outperforming baselines on the new PARS benchmark and RefCOCOm.

  5. DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model

    cs.CV 2026-02 unverdicted novelty 7.0

    DLEBench is the first benchmark for small-scale object editing in instruction-based image editing models, using 1889 samples, seven instruction types, and a dual-mode evaluation protocol to reveal performance gaps in ...

  6. SAM 3: Segment Anything with Concepts

    cs.CV 2025-11 unverdicted novelty 7.0

    SAM 3 introduces promptable concept segmentation that doubles accuracy of prior systems on images and videos while improving standard SAM segmentation performance.

  7. RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

    cs.RO 2026-07 conditional novelty 6.0

    RynnBrain 1.1 reports state-of-the-art embodied cognition and localization scores with a 122B-A10B model and improved real-robot VLA policies via joint multi-embodiment training.

  8. SceneBind: Binding What and Where Across Vision, Audio and Language

    cs.CV 2026-07 conditional novelty 6.0

    A scene is represented as a global semantic embedding plus object-centric semantic-spatial slots (azimuth, elevation, distance, confidence), which improves cross-modal retrieval and enables zero-shot audio-visual loca...

  9. SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding

    cs.CV 2026-05 unverdicted novelty 6.0

    SceneParser introduces hierarchical scene parsing as object-part-affordance chains, a VLM trained with pseudo labels and curriculum learning, and SceneParser-Bench with 1.74M affordance annotations, showing better str...

  10. MIRAGE: Benchmarking and Aligning Multi-Instance Image Editing

    cs.CV 2026-04 unverdicted novelty 6.0

    MIRAGE introduces a benchmark for multi-instance image editing and a training-free framework that uses vision-language parsing and parallel regional denoising to achieve precise edits without altering backgrounds.

  11. HERO: Learning Humanoid End-Effector Control for Visual Whole-Body Open-Vocabulary Object Grasping

    cs.RO 2026-02 conditional novelty 6.0

    HERO achieves 2.44 cm end-effector tracking error on a Unitree G1 humanoid and uses it, with open-vocabulary perception, to grasp novel objects at up to 90% success in diverse real scenes.

  12. Unify Robot Actions in Camera Frame

    cs.RO 2025-11 conditional novelty 6.0

    CalibAll estimates camera extrinsics on existing datasets to convert robot actions into a unified camera-frame representation, enabling stronger cross-embodiment pretraining.

  13. Inferring Dynamic Physical Properties from Video Foundation Models

    cs.CV 2025-10 unverdicted novelty 6.0

    Video foundation models infer dynamic physical properties such as elasticity, viscosity, and friction from videos at levels close to classical oracles while outperforming current MLLMs with suitable prompting.

  14. VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

    cs.CV 2025-04 unverdicted novelty 6.0

    VLM-R1 applies R1-style RL using rule-based rewards on visual tasks with clear ground truth to achieve competitive performance and superior generalization over SFT in vision-language models.

  15. TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

    cs.CV 2026-07 conditional novelty 5.5

    A 0.2B-parameter V+L→A policy matches larger LLM-centric VLAs on LIBERO at 31 ms latency and 0.9 GB VRAM by fusing DINOv3 and BERT features with bidirectional cross-attention and ACT-style action chunks.

  16. RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

    cs.RO 2026-07 conditional novelty 5.0

    RynnBrain 1.1 reports benchmark-leading embodied perception and cross-embodiment robot policies, adding 3D grounding and contact-point prediction.

  17. Embodied Operators and Benchmarking: Toward Reusable and Deployable Embodied Intelligence Systems

    cs.AI 2026-07 conditional novelty 5.0

    Embodied operators—deployable modules with task semantics and I/O contracts—should be the unit of optimization and multi-dimensional benchmarking for reusable robot intelligence systems.

  18. RelAfford6D: Relational 6D Affordance Graphs for Constraint-Driven Robotic Manipulation

    cs.RO 2026-06 unverdicted novelty 5.0

    RelAfford6D constructs relational 6D affordance graphs from instructions, uses vision foundation models for metric poses, and executes via closed-loop kinematic constraint tracking to achieve claimed superior zero-sho...

  19. VL-DINO: Leveraging CLIP Vision-Language Knowledge for Open-Vocabulary Object Detectio

    cs.CV 2026-06 unverdicted novelty 5.0

    VL-DINO improves open-vocabulary object detection by adding QPSC, VSE, and ORSA modules that inject CLIP knowledge into DINO, reaching 36.3 and 38.1 AP zero-shot on LVIS.

  20. TrackRef3D: Multi-View Consistent Track-then-Label for Open-World Referring Segmentation in 3D Gaussian Splatting

    cs.CV 2026-05 unverdicted novelty 5.0

    TrackRef3D proposes a fully automatic multi-view consistent track-then-label method for open-world referring segmentation in 3D Gaussian Splatting using TSCM, visibility-aware descriptions, and hybrid contrastive training.

  21. RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos

    cs.CV 2026-05 unverdicted novelty 5.0

    RHINO recovers 3D human, novel manipulated object, and static scene from monocular video by stabilizing SfM with foundation models, separating motions, and refining with compositional neural SDFs plus contact priors.

  22. DetRefiner: Model-Agnostic Detection Refinement with Feature Fusion Transformer

    cs.CV 2026-05 unverdicted novelty 5.0

    DetRefiner fuses global and local features with a Transformer to refine OVOD confidence scores, delivering up to +10.1 AP gains on novel categories across multiple datasets.

  23. Quantum orientation entanglement analysis of the interpolating helicity states between the instant form dynamics and the light-front dynamics

    hep-th 2026-03 unverdicted novelty 5.0

    Interpolating helicity states expanded in Jacob–Wick helicity via Wigner d-matrix probabilities reveal a critical angle that bifurcates instant-form and light-front spin dynamics in contact-interaction pair production.

  24. A Survey on Vision-Language-Action Models: An Action Tokenization Perspective

    cs.RO 2025-07 unverdicted novelty 5.0

    The survey frames VLA models as pipelines that generate progressively grounded action tokens and classifies those tokens into eight types to guide future development.

  25. Qwen2.5-VL Technical Report

    cs.CV 2025-02 unverdicted novelty 5.0

    Qwen2.5-VL reports a vision-language model family using native dynamic-resolution ViT and absolute time encoding that matches GPT-4o on document and diagram tasks while supporting hour-long videos with second-level lo...

  26. Benchmarking Vision Foundation Models for Input Monitoring in Autonomous Driving

    cs.CV 2025-01 unverdicted novelty 5.0

    Vision foundation model embeddings with density modeling outperform state-of-the-art methods for unsupervised semantic and covariate shift detection in autonomous driving inputs.

  27. Lightweight Neural Framework for Robust 3D Volume and Surface Estimation from Multi-View Images

    cs.CV 2026-06 unverdicted novelty 4.0

    A lightweight neural model fuses 3D point cloud reconstructions with view-aligned 2D features via a graph decoder to regress volume, surface area, and uncertainties from multi-view images without iterative optimization.

  28. Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models

    cs.CV 2026-06 unverdicted novelty 4.0

    YOLO26 presents a unified real-time vision model family with dual-head end-to-end design, new training components, and task-specific heads that reports improved mAP-latency tradeoffs on COCO and LVIS benchmarks across...

  29. One-Shot Crowd Counting With Density Guidance For Scene Adaptation

    cs.CV 2026-02 conditional novelty 4.0

    A one-shot crowd-counting method uses EM-clustered local and global density features from one labeled support image to adapt a fixed model to an unseen surveillance scene.

  30. LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation

    cs.CV 2026-04 unverdicted novelty 3.0

    This review organizes literature on large multimodal models and object-centric vision into four themes—understanding, referring segmentation, editing, and generation—while summarizing paradigms, strategies, and challe...