A multi-loss self-supervised audio model that adds frame-level contrast and pitch-shift prediction to the COLA clip-level loss improves downstream audio classification, event detection, and pitch detection.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Self-supervised learning method using multiple sampling strategies for general-purpose audio representation
A multi-loss self-supervised audio model that adds frame-level contrast and pitch-shift prediction to the COLA clip-level loss improves downstream audio classification, event detection, and pitch detection.