ULAO, a CLIP-based framework with object-first sequential prediction and prediction-based hard negative contrastive training, reports state-of-the-art compositional zero-shot learning results on MIT-States, UT-Zappos, and C-GQA.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Compositional Zero-Shot Learning with Contextualized Cues and Adaptive Contrastive Training
ULAO, a CLIP-based framework with object-first sequential prediction and prediction-based hard negative contrastive training, reports state-of-the-art compositional zero-shot learning results on MIT-States, UT-Zappos, and C-GQA.