CAV-MAE Sync improves audio-visual representation learning by aligning fine-grained audio segments with video frames and separating contrastive and reconstruction objectives, achieving state-of-the-art zero-shot retrieval.
Look, listen and learn
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.MM 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
CAV-MAE Sync improves audio-visual representation learning by aligning fine-grained audio segments with video frames and separating contrastive and reconstruction objectives, achieving state-of-the-art zero-shot retrieval.