Pith. sign in

REVIEW 2 cited by

Toward Fully Self-Supervised Multi-Pitch Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15569 v1 pith:23PLUXYY submitted 2024-02-23 eess.AS cs.LGcs.SD

Toward Fully Self-Supervised Multi-Pitch Estimation

classification eess.AS cs.LGcs.SD
keywords multi-pitchestimationself-superviseddatasetsfullylearningmixturesmusic
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Multi-pitch estimation is a decades-long research problem involving the detection of pitch activity associated with concurrent musical events within multi-instrument mixtures. Supervised learning techniques have demonstrated solid performance on more narrow characterizations of the task, but suffer from limitations concerning the shortage of large-scale and diverse polyphonic music datasets with multi-pitch annotations. We present a suite of self-supervised learning objectives for multi-pitch estimation, which encourage the concentration of support around harmonics, invariance to timbral transformations, and equivariance to geometric transformations. These objectives are sufficient to train an entirely convolutional autoencoder to produce multi-pitch salience-grams directly, without any fine-tuning. Despite training exclusively on a collection of synthetic single-note audio samples, our fully self-supervised framework generalizes to polyphonic music mixtures, and achieves performance comparable to supervised models trained on conventional multi-pitch datasets.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music

    cs.SD 2026-07 conditional novelty 6.0

    A self-supervised Swin/JEPA model on piano rolls reconstructs music at F1≈0.995, beats Haar scattering for emotion recognition, and steers flow-generated output via prompt embeddings.

  2. Music102: An $D_{12}$-equivariant transformer for chord progression accompaniment

    cs.SD 2024-10 unverdicted novelty 5.0

    Music102 integrates D12-equivariance into a transformer for chord progression accompaniment and shows gains over Music101 on POP909.