Pith. sign in

REVIEW 2 cited by

Toward Fully Self-Supervised Multi-Pitch Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15569 v1 pith:23PLUXYY submitted 2024-02-23 eess.AS cs.LGcs.SD

classification eess.AScs.LGcs.SD
keywords multi-pitchestimationself-superviseddatasetsfullylearningmixturesmusic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-pitch estimation is a decades-long research problem involving the detection of pitch activity associated with concurrent musical events within multi-instrument mixtures. Supervised learning techniques have demonstrated solid performance on more narrow characterizations of the task, but suffer from limitations concerning the shortage of large-scale and diverse polyphonic music datasets with multi-pitch annotations. We present a suite of self-supervised learning objectives for multi-pitch estimation, which encourage the concentration of support around harmonics, invariance to timbral transformations, and equivariance to geometric transformations. These objectives are sufficient to train an entirely convolutional autoencoder to produce multi-pitch salience-grams directly, without any fine-tuning. Despite training exclusively on a collection of synthetic single-note audio samples, our fully self-supervised framework generalizes to polyphonic music mixtures, and achieves performance comparable to supervised models trained on conventional multi-pitch datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A self-supervised Swin/JEPA model on piano rolls reconstructs music at F1≈0.995, beats Haar scattering for emotion recognition, and steers flow-generated output via prompt embeddings.

  2. Investigating an Overfitting and Degeneration Phenomenon in Self-Supervised Multi-Pitch Estimation

    eess.AS 2025-06 conditional novelty 5.0 of 10

    Adding self-supervised objectives to a supervised multi-pitch estimator improves closed-set performance but triggers degeneration to blank predictions on additional, unlabeled data.

Pith tools