Pith. sign in

REVIEW

Invariances and Data Augmentation for Supervised Music Transcription

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1711.04845 v1 pith:OZ5G3LFV submitted 2017-11-13 stat.ML cs.LGcs.SDeess.AS

classification stat.MLcs.LGcs.SDeess.AS
keywords datamodelsmusicfrequencymodelnetworkparameterstranscription
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper explores a variety of models for frame-based music transcription, with an emphasis on the methods needed to reach state-of-the-art on human recordings. The translation-invariant network discussed in this paper, which combines a traditional filterbank with a convolutional neural network, was the top-performing model in the 2017 MIREX Multiple Fundamental Frequency Estimation evaluation. This class of models shares parameters in the log-frequency domain, which exploits the frequency invariance of music to reduce the number of model parameters and avoid overfitting to the training data. All models in this paper were trained with supervision by labeled data from the MusicNet dataset, augmented by random label-preserving pitch-shift transformations.

Discussion (0). Continue with ORCID to comment.

Pith tools