GVD guides a pre-trained video diffusion model with clustering-derived features to distill video datasets, outperforming prior methods on MiniUCF and HMDB51 while retaining over 70% of full-data accuracy using under 4% of frames.
Taming transformers for high-resolution image synthesis
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
GVD: Guiding Video Diffusion Model for Scalable Video Distillation
GVD guides a pre-trained video diffusion model with clustering-derived features to distill video datasets, outperforming prior methods on MiniUCF and HMDB51 while retaining over 70% of full-data accuracy using under 4% of frames.