Pith. sign in

REVIEW 8 cited by

Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1704.01279 v1 pith:WKZR3WBN submitted 2017-04-05 cs.LG cs.AIcs.SD

classification cs.LGcs.AIcs.SD
keywords audioautoencoderdatasetshigh-qualitymodelmusicalnotesnsynth
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative models in vision have seen rapid progress due to algorithmic improvements and the availability of high-quality image datasets. In this paper, we offer contributions in both these areas to enable similar progress in audio modeling. First, we detail a powerful new WaveNet-style autoencoder model that conditions an autoregressive decoder on temporal codes learned from the raw audio waveform. Second, we introduce NSynth, a large-scale and high-quality dataset of musical notes that is an order of magnitude larger than comparable public datasets. Using NSynth, we demonstrate improved qualitative and quantitative performance of the WaveNet autoencoder over a well-tuned spectral autoencoder baseline. Finally, we show that the model learns a manifold of embeddings that allows for morphing between instruments, meaningfully interpolating in timbre to create new types of sounds that are realistic and expressive.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing

    cs.SD 2026-08 conditional novelty 6.0 of 10

    P-MUSE unifies MIDI-to-music generation and local editing under one flow-matching model that accepts paired, style, or mixed prompts, and introduces a Tail-Drop guidance schedule that improves quality.

  2. Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A zero-shot instrument cloning system feeds raw reference audio into a flow-matching DiT and uses asymmetric CFG to keep melody and timbre control separate.

  3. LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    LLaSO releases a 3.8B speech-language model, 25.5M training instances, and an evaluation benchmark, claiming a normalized score of 0.72.

  4. OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction

    cs.SD 2025-07 conditional novelty 5.0 of 10

    OMAR-RQ, an open 580M-parameter music audio model trained with multi-codebook, multi-feature masked token prediction on 330k hours, reports leading open-model results on several MIR benchmarks.

  5. Workflow-Based Evaluation of Music Generation Systems

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A single-producer workflow evaluation of eight music AI tools finds they work as idea and sound generators but not as complete composers, and proposes a reusable framework.

  6. Autoencoding sensory substitution

    q-bio.NC 2019-07 unverdicted novelty 4.0 of 10

    Deep recurrent autoencoders convert images to shortened audio signals that incorporate hearing models, enabling above-chance hand posture discrimination and object reaching after a few hours of training instead of months.

  7. Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

    cs.SD 2025-06 conditional novelty 3.0 of 10

    Transferring I-JEPA's masked latent prediction to mel-spectrograms yields competitive audio representations on music and environmental sound tasks with a small fraction of the training data.

  8. Classical Music Prediction and Composition by means of Variational Autoencoders

    cs.SD 2019-06 unverdicted novelty 3.0 of 10

    VAEs are trained on classical music to encode pieces into latent space and predict continuations, enabling composition of new music from existing pieces or random starts even with small training sets.

Pith tools