REVIEW 2 cited by
LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained Descriptors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Inspired by the outstanding zero-shot capability of vision language models (VLMs) in image classification tasks, open-vocabulary object detection has attracted increasing interest by distilling the broad VLM knowledge into detector training. However, most existing open-vocabulary detectors learn by aligning region embeddings with categorical labels (e.g., bicycle) only, disregarding the capability of VLMs on aligning visual embeddings with fine-grained text description of object parts (e.g., pedals and bells). This paper presents DVDet, a Descriptor-Enhanced Open Vocabulary Detector that introduces conditional context prompts and hierarchical textual descriptors that enable precise region-text alignment as well as open-vocabulary detection training in general. Specifically, the conditional context prompt transforms regional embeddings into image-like representations that can be directly integrated into general open vocabulary detection training. In addition, we introduce large language models as an interactive and implicit knowledge repository which enables iterative mining and refining visually oriented textual descriptors for precise region-text alignment. Extensive experiments over multiple large-scale benchmarks show that DVDet outperforms the state-of-the-art consistently by large margins.
Forward citations
Cited by 2 Pith papers
-
Test-Time Optimization for Domain Adaptive Open Vocabulary Segmentation
A plug-and-play test-time optimization method improves zero-shot open-vocabulary segmentation on specialized-domain datasets by jointly tuning per-category text embeddings and aggregating visual features.
-
Open-Det: An Efficient Learning Framework for Open-Ended Detection
Open-Det trains an open-ended detector on Visual Genome and reports higher zero-shot LVIS accuracy than GenerateU using 1.5% of the data and 31 epochs instead of 149.
Discussion (0). Continue with ORCID to comment.