LEViL performs annotation-free video pretraining via VLM-generated pseudo-label spaces and zero-shot distillation, then uses target-aware fine-tuning to outperform semi-supervised baselines on UCF101 and HMDB51 in limited-label settings.
Would mega- scale datasets further enhance spatiotemporal 3d cnns?
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
LEViL: Label-Efficient Video Learning via Zero-Shot Distillation over VLM-Generated Pseudo-Label Spaces
LEViL performs annotation-free video pretraining via VLM-generated pseudo-label spaces and zero-shot distillation, then uses target-aware fine-tuning to outperform semi-supervised baselines on UCF101 and HMDB51 in limited-label settings.