ULAO, a CLIP-based framework with object-first sequential prediction and prediction-based hard negative contrastive training, reports state-of-the-art compositional zero-shot learning results on MIT-States, UT-Zappos, and C-GQA.
Zero-Shot Compositional Concept Learning
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In this paper, we study the problem of recognizing compositional attribute-object concepts within the zero-shot learning (ZSL) framework. We propose an episode-based cross-attention (EpiCA) network which combines merits of cross-attention mechanism and episode-based training strategy to recognize novel compositional concepts. Firstly, EpiCA bases on cross-attention to correlate concept-visual information and utilizes the gated pooling layer to build contextualized representations for both images and concepts. The updated representations are used for a more in-depth multi-modal relevance calculation for concept recognition. Secondly, a two-phase episode training strategy, especially the transductive phase, is adopted to utilize unlabeled test examples to alleviate the low-resource learning problem. Experiments on two widely-used zero-shot compositional learning (ZSCL) benchmarks have demonstrated the effectiveness of the model compared with recent approaches on both conventional and generalized ZSCL settings.
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Compositional Zero-Shot Learning with Contextualized Cues and Adaptive Contrastive Training
ULAO, a CLIP-based framework with object-first sequential prediction and prediction-based hard negative contrastive training, reports state-of-the-art compositional zero-shot learning results on MIT-States, UT-Zappos, and C-GQA.