A single-stream vision-language transformer with top-K text selection and a sparse compositor achieves state-of-the-art open-world compositional zero-shot learning on MIT-States, C-GQA, and VAW-CZSL.
Multimodal pretraining unmasked: A meta-analysis and a unified framework of vision-and- language berts
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Unified Framework for Open-World Compositional Zero-shot Learning
A single-stream vision-language transformer with top-K text selection and a sparse compositor achieves state-of-the-art open-world compositional zero-shot learning on MIT-States, C-GQA, and VAW-CZSL.