A lightweight CLIP variant trained on one RTX3090 with an augmented 12M-image dataset and a teacher-distillation recipe reaches competitive zero-shot retrieval, though most of the gain comes from the pretrained teacher.
Unified contrastive learning in image-text-label space
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers
A lightweight CLIP variant trained on one RTX3090 with an augmented 12M-image dataset and a teacher-distillation recipe reaches competitive zero-shot retrieval, though most of the gain comes from the pretrained teacher.