KDC-MAE pretrains an audio-video transformer with two complementary masks plus KL self-distillation and reports small, partly inconsistent accuracy gains over CAV-MAE.
Vggsound: A large-scale audio-visual dataset
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
KDC-MAE: Knowledge Distilled Contrastive Mask Auto-Encoder
KDC-MAE pretrains an audio-video transformer with two complementary masks plus KL self-distillation and reports small, partly inconsistent accuracy gains over CAV-MAE.