Pith. sign in

REVIEW 2 cited by

Prompt-Guided Transformers for End-to-End Open-Vocabulary Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.14386 v1 pith:TJFCOFSJ submitted 2023-03-25 cs.CV

classification cs.CV
keywords detectionopen-vocabularyclipend-to-endinferencenovelobjectprompt-ovd
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Prompt-OVD is an efficient and effective framework for open-vocabulary object detection that utilizes class embeddings from CLIP as prompts, guiding the Transformer decoder to detect objects in both base and novel classes. Additionally, our novel RoI-based masked attention and RoI pruning techniques help leverage the zero-shot classification ability of the Vision Transformer-based CLIP, resulting in improved detection performance at minimal computational cost. Our experiments on the OV-COCO and OVLVIS datasets demonstrate that Prompt-OVD achieves an impressive 21.2 times faster inference speed than the first end-to-end open-vocabulary detection method (OV-DETR), while also achieving higher APs than four two-stage-based methods operating within similar inference time ranges. Code will be made available soon.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Open-Vocabulary Gaze Object Prediction: Benchmark and Method

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A COCO+GazeFollow-derived benchmark (86 categories) plus a Grounding DINO + gaze-selection pipeline with selective tuning improves open-vocabulary gaze object prediction over existing closed-vocabulary methods.

  2. SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception

    cs.CV 2026-07 accept novelty 6.0 of 10

    SynCLIP aligns and refines spatial attention maps across synonyms via SSA/SAR modules and a new SEViC corpus, improving OVDP robustness and SOTA CLIP-based scores.

Pith tools