VocaDet detects arbitrary objects by retrieving multi-granularity visual tokens from a sample-built vector database of position-debiased DINOv3 features and topology, without detector training.
Ua-detrac: A new benchmark and protocol for multi-object detection and tracking.Computer Vision and Image Understanding, 193:102907, 2020
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Retrieval
VocaDet detects arbitrary objects by retrieving multi-granularity visual tokens from a sample-built vector database of position-debiased DINOv3 features and topology, without detector training.