REVIEW 2 cited by
Toward Fully Self-Supervised Multi-Pitch Estimation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multi-pitch estimation is a decades-long research problem involving the detection of pitch activity associated with concurrent musical events within multi-instrument mixtures. Supervised learning techniques have demonstrated solid performance on more narrow characterizations of the task, but suffer from limitations concerning the shortage of large-scale and diverse polyphonic music datasets with multi-pitch annotations. We present a suite of self-supervised learning objectives for multi-pitch estimation, which encourage the concentration of support around harmonics, invariance to timbral transformations, and equivariance to geometric transformations. These objectives are sufficient to train an entirely convolutional autoencoder to produce multi-pitch salience-grams directly, without any fine-tuning. Despite training exclusively on a collection of synthetic single-note audio samples, our fully self-supervised framework generalizes to polyphonic music mixtures, and achieves performance comparable to supervised models trained on conventional multi-pitch datasets.
Forward citations
Cited by 2 Pith papers
-
MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music
A self-supervised Swin/JEPA model on piano rolls reconstructs music at F1≈0.995, beats Haar scattering for emotion recognition, and steers flow-generated output via prompt embeddings.
-
Investigating an Overfitting and Degeneration Phenomenon in Self-Supervised Multi-Pitch Estimation
Adding self-supervised objectives to a supervised multi-pitch estimator improves closed-set performance but triggers degeneration to blank predictions on additional, unlabeled data.
Discussion (0). Continue with ORCID to comment.