Pith. sign in

REVIEW 4 cited by

End-to-end learning for music audio tagging at scale

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1711.02520 v4 pith:RRPILYBM submitted 2017-11-07 cs.SD eess.AS

classification cs.SDeess.AS
keywords datamodelsavailableend-to-endlearningmusicsongswhen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The lack of data tends to limit the outcomes of deep learning research, particularly when dealing with end-to-end learning stacks processing raw data such as waveforms. In this study, 1.2M tracks annotated with musical labels are available to train our end-to-end models. This large amount of data allows us to unrestrictedly explore two different design paradigms for music auto-tagging: assumption-free models - using waveforms as input with very small convolutional filters; and models that rely on domain knowledge - log-mel spectrograms with a convolutional neural network designed to learn timbral and temporal features. Our work focuses on studying how these two types of deep architectures perform when datasets of variable size are available for training: the MagnaTagATune (25k songs), the Million Song Dataset (240k songs), and a private dataset of 1.2M songs. Our experiments suggest that music domain assumptions are relevant when not enough training data are available, thus showing how waveform-based models outperform spectrogram-based ones in large-scale data scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Leave-One-EquiVariant: Alleviating invariance-related information loss in contrastive music representations

    cs.SD 2024-12 conditional novelty 6.0 of 10

    A contrastive music representation method, Leave-One-EquiVariant, keeps pitch and tempo information in separate embedding subspaces, improving key and tempo tasks without hurting tagging.

  2. Vision Language Models Are Few-Shot Audio Spectrogram Classifiers

    cs.SD 2024-11 conditional novelty 6.0 of 10

    GPT-4o classifies environmental sounds from spectrogram images in a few-shot setting with 59% cross-validated accuracy on ESC-10, beating a commercial audio language model and roughly matching human experts.

  3. Lukthung Classification Using Neural Networks on Lyrics and Audios

    cs.LG 2019-08 conditional novelty 5.0 of 10

    A combined bag-of-words lyrics network and spectrogram CNN classify Lukthung songs from other Thai genres with an F1 of 0.86 on a private 10,547-song dataset.

  4. The Good, The Efficient and the Inductive Biases: Exploring Efficiency in Deep Learning Through the Use of Inductive Biases

    cs.LG 2024-11 conditional novelty 3.0 of 10

    A dissertation synthesizing the author's papers on continuous kernel convolutions and symmetry-preserving architectures, claiming these inductive biases improve deep learning efficiency.

Pith tools