Pith. sign in

REVIEW 1 cited by

Hyperbolic Learning with Synthetic Captions for Open-World Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.05016 v1 pith:TLMYCJDL submitted 2024-04-07 cs.CV

classification cs.CV
keywords detectioncaptionsnovelobjectopen-worldsyntheticcaptiondescriptions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Open-world detection poses significant challenges, as it requires the detection of any object using either object class labels or free-form texts. Existing related works often use large-scale manual annotated caption datasets for training, which are extremely expensive to collect. Instead, we propose to transfer knowledge from vision-language models (VLMs) to enrich the open-vocabulary descriptions automatically. Specifically, we bootstrap dense synthetic captions using pre-trained VLMs to provide rich descriptions on different regions in images, and incorporate these captions to train a novel detector that generalizes to novel concepts. To mitigate the noise caused by hallucination in synthetic captions, we also propose a novel hyperbolic vision-language learning approach to impose a hierarchy between visual and caption embeddings. We call our detector ``HyperLearner''. We conduct extensive experiments on a wide variety of open-world detection benchmarks (COCO, LVIS, Object Detection in the Wild, RefCOCO) and our results show that our model consistently outperforms existing state-of-the-art methods, such as GLIP, GLIPv2 and Grounding DINO, when using the same backbone.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

    cs.CV 2024-11 conditional novelty 6.0 of 10

    LangView uses per-view caption accuracy against view-agnostic narrations as pseudo-labels to train a view selector that outperforms heuristics and prior baselines on Ego-Exo4D and LEMMA.

Pith tools