A multi-loss self-supervised audio model that adds frame-level contrast and pitch-shift prediction to the COLA clip-level loss improves downstream audio classification, event detection, and pitch detection.
Our method improves the performance of all tasks compared to existing methods in the experiment of the subset of Audioset
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Self-supervised learning method using multiple sampling strategies for general-purpose audio representation
A multi-loss self-supervised audio model that adds frame-level contrast and pitch-shift prediction to the COLA clip-level loss improves downstream audio classification, event detection, and pitch detection.