Pith. sign in

REVIEW 2 cited by

LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained Descriptors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.04630 v1 pith:VGWZLYCO submitted 2024-02-07 cs.CV

classification cs.CV
keywords detectiondescriptorsembeddingsobjectopenopen-vocabularytrainingvlms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Inspired by the outstanding zero-shot capability of vision language models (VLMs) in image classification tasks, open-vocabulary object detection has attracted increasing interest by distilling the broad VLM knowledge into detector training. However, most existing open-vocabulary detectors learn by aligning region embeddings with categorical labels (e.g., bicycle) only, disregarding the capability of VLMs on aligning visual embeddings with fine-grained text description of object parts (e.g., pedals and bells). This paper presents DVDet, a Descriptor-Enhanced Open Vocabulary Detector that introduces conditional context prompts and hierarchical textual descriptors that enable precise region-text alignment as well as open-vocabulary detection training in general. Specifically, the conditional context prompt transforms regional embeddings into image-like representations that can be directly integrated into general open vocabulary detection training. In addition, we introduce large language models as an interactive and implicit knowledge repository which enables iterative mining and refining visually oriented textual descriptors for precise region-text alignment. Extensive experiments over multiple large-scale benchmarks show that DVDet outperforms the state-of-the-art consistently by large margins.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Test-Time Optimization for Domain Adaptive Open Vocabulary Segmentation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A plug-and-play test-time optimization method improves zero-shot open-vocabulary segmentation on specialized-domain datasets by jointly tuning per-category text embeddings and aggregating visual features.

  2. Open-Det: An Efficient Learning Framework for Open-Ended Detection

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Open-Det trains an open-ended detector on Visual Genome and reports higher zero-shot LVIS accuracy than GenerateU using 1.5% of the data and 31 epochs instead of 149.

Pith tools