SLAP trains joint music-text embeddings with a BYOL-style, negative-free loss and outperforms CLAP on music-text retrieval, zero-shot classification, and several downstream MIR tasks while reducing the modality gap.
We train SLAP on an internal private dataset of 260,000 pairs of full-length production-quality music tracks and professionally annotated captions (PrivateCaps [16])
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding
SLAP trains joint music-text embeddings with a BYOL-style, negative-free loss and outperforms CLAP on music-text retrieval, zero-shot classification, and several downstream MIR tasks while reducing the modality gap.