Pith. sign in

REVIEW 14 cited by

Music Source Separation in the Waveform Domain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.13254 v2 pith:44IPKYSU submitted 2019-11-27 cs.SD cs.LGeess.ASstat.ML

Music Source Separation in the Waveform Domain

classification cs.SD cs.LGeess.ASstat.ML
keywords sourceseparationdemucsmusicconv-tasnetwaveformarchitecturesaudio
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Source separation for music is the task of isolating contributions, or stems, from different instruments recorded individually and arranged together to form a song. Such components include voice, bass, drums and any other accompaniments.Contrarily to many audio synthesis tasks where the best performances are achieved by models that directly generate the waveform, the state-of-the-art in source separation for music is to compute masks on the magnitude spectrum. In this paper, we compare two waveform domain architectures. We first adapt Conv-Tasnet, initially developed for speech source separation,to the task of music source separation. While Conv-Tasnet beats many existing spectrogram-domain methods, it suffersfrom significant artifacts, as shown by human evaluations. We propose instead Demucs, a novel waveform-to-waveform model,with a U-Net structure and bidirectional LSTM.Experiments on the MusDB dataset show that, with proper data augmentation, Demucs beats allexisting state-of-the-art architectures, including Conv-Tasnet, with 6.3 SDR on average, (and up to 6.8 with 150 extra training songs, even surpassing the IRM oracle for the bass source).Using recent development in model quantization, Demucs can be compressed down to 120MBwithout any loss of accuracy.We also provide human evaluations, showing that Demucs benefit from a large advantagein terms of the naturalness of the audio. However, it suffers from some bleeding,especially between the vocals and other source.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics

    cs.SD 2026-04 unverdicted novelty 7.0

    ArtifactNet extracts codec residuals from spectrograms with a 4M-parameter network to detect AI music at F1=0.9829 and 1.49% FPR on unseen tracks from 22 generators, outperforming larger baselines.

  2. The Spheres Dataset: Multitrack Orchestral Recordings for Music Source Separation and Information Retrieval

    eess.AS 2025-11 accept novelty 7.0

    The Spheres dataset provides multitrack orchestral recordings with isolated instrument stems and acoustic characterizations to support supervised machine learning for music source separation in the classical domain.

  3. High Fidelity Neural Audio Compression

    eess.AS 2022-10 accept novelty 7.0

    EnCodec is an end-to-end trained streaming neural audio codec that uses a single multiscale spectrogram discriminator and a gradient-normalizing loss balancer to achieve higher fidelity than prior methods at the same ...

  4. MAGE: Modality-Agnostic Music Generation and Target-Source Extraction

    cs.SD 2026-04 unverdicted novelty 6.0

    MAGE unifies text, visual, and audio-conditioned music generation and editing in one flow-based latent model with dynamic modality masking and cross-gated control.

  5. Discrete Token Modeling for Multi-Stem Music Source Separation with Language Models

    eess.AS 2026-04 unverdicted novelty 6.0

    A Conformer-conditioned decoder-only language model generates discrete tokens via a neural audio codec to separate four music stems, reaching near state-of-the-art perceptual quality and top NISQA on vocals in MUSDB18...

  6. CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

    cs.SD 2025-09 reject novelty 6.0

    A FiLM-conditioned transformer masker on DAC codec latents performs text-guided sound separation with claimed efficiency, but the main comparison against AudioSep is confounded by asymmetric input processing.

  7. CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

    cs.SD 2025-09 unverdicted novelty 6.0

    CodecSep performs prompt-driven universal sound separation directly in neural audio codec latents by combining a frozen DAC backbone with a lightweight FiLM-conditioned Transformer masker driven by CLAP embeddings, yi...

  8. Improving Music Source Separation with Diffusion and Consistency Refinement

    cs.SD 2024-12 unverdicted novelty 6.0

    Diffusion-based refinement followed by consistency distillation improves music source separation quality and inference speed across U-Net and BS-RoFormer backbones on Slakh2100 and MUSDB18.

  9. MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

    cs.CV 2026-06 unverdicted novelty 5.0

    MaineCoon is presented as the first 22B-parameter real-time streaming audio-visual autoregressive model optimized for social-interactive applications, using novel training techniques and an agentic inference framework.

  10. MAGE: Modality-Agnostic Music Generation and Target-Source Extraction

    cs.SD 2026-04 unverdicted novelty 5.0

    A shared continuous-latent flow model generates music from text/vision or extracts a target source from a mixture via visual-audio alignment, gated modulation, and dynamic modality masking.

  11. Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM

    eess.AS 2025-09 conditional novelty 5.0

    iMTI-Net combines Whisper embeddings with uncertainty-proxy statistics and a CNN-sLSTM backbone to improve non-intrusive speech intelligibility prediction on TMHINT-QI(S).

  12. DTT-BSR+: A Generative-Regression Cascade for Music Source Restoration

    eess.AS 2026-06 unverdicted novelty 4.0

    DTT-BSR+ is a generative-then-regression cascade for music source restoration that reports MMSNR gains over single-stage DTT-BSR and X-LANCE on most stems while noting a distribution-vs-reconstruction trade-off via FAD.

  13. Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

    cs.SD 2026-07 conditional novelty 3.5

    MSST unifies training, validation, and inference for many music source-separation architectures and reports small quality gains from TTA, ensembling, and related engineering techniques.

  14. Improving Text-to-Music Generation with Human Preference Rewards

    cs.SD 2026-06 unverdicted novelty 2.0

    A text-to-music model is improved by conditioning on and selecting with a human preference reward, where expert iteration on top outputs contributes the largest measured gains on 100 Song Describer prompts.