Transferring I-JEPA's masked latent prediction to mel-spectrograms yields competitive audio representations on music and environmental sound tasks with a small fraction of the training data.
Efficient Self-super- vised Learning with Contextualized Target Representations for Vision, Speech and Language
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
Transferring I-JEPA's masked latent prediction to mel-spectrograms yields competitive audio representations on music and environmental sound tasks with a small fraction of the training data.