OW-CLIP reports that combining LLM-generated descriptions, user image sorting, and crop-based label smoothing lets a CLIP detector match most state-of-the-art open-world detection accuracy with under 4 percent of the usual training annotations.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
OW-CLIP: Data-Efficient Visual Supervision for Open-World Object Detection via Human-AI Collaboration
OW-CLIP reports that combining LLM-generated descriptions, user image sorting, and crop-based label smoothing lets a CLIP detector match most state-of-the-art open-world detection accuracy with under 4 percent of the usual training annotations.