Pith. sign in

REVIEW 14 cited by

Music Source Separation in the Waveform Domain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.13254 v2 pith:44IPKYSU submitted 2019-11-27 cs.SD cs.LGeess.ASstat.ML

classification cs.SDcs.LGeess.ASstat.ML
keywords sourceseparationdemucsmusicconv-tasnetwaveformarchitecturesaudio
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Source separation for music is the task of isolating contributions, or stems, from different instruments recorded individually and arranged together to form a song. Such components include voice, bass, drums and any other accompaniments.Contrarily to many audio synthesis tasks where the best performances are achieved by models that directly generate the waveform, the state-of-the-art in source separation for music is to compute masks on the magnitude spectrum. In this paper, we compare two waveform domain architectures. We first adapt Conv-Tasnet, initially developed for speech source separation,to the task of music source separation. While Conv-Tasnet beats many existing spectrogram-domain methods, it suffersfrom significant artifacts, as shown by human evaluations. We propose instead Demucs, a novel waveform-to-waveform model,with a U-Net structure and bidirectional LSTM.Experiments on the MusDB dataset show that, with proper data augmentation, Demucs beats allexisting state-of-the-art architectures, including Conv-Tasnet, with 6.3 SDR on average, (and up to 6.8 with 150 extra training songs, even surpassing the IRM oracle for the bass source).Using recent development in model quantization, Demucs can be compressed down to 120MBwithout any loss of accuracy.We also provide human evaluations, showing that Demucs benefit from a large advantagein terms of the naturalness of the audio. However, it suffers from some bleeding,especially between the vocals and other source.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MAGE: Modality-Agnostic Music Generation and Target-Source Extraction

    cs.SD 2026-04 unverdicted novelty 6.0 of 10

    MAGE unifies text, visual, and audio-conditioned music generation and editing in one flow-based latent model with dynamic modality masking and cross-gated control.

  2. Query-Based Asymmetric Modeling with Decoupled Input-Output Rates for Speech Restoration

    eess.AS 2025-09 conditional novelty 6.0 of 10

    TF-Restormer restores degraded speech at arbitrary input-output sampling rates in a single model, using a heavy encoder and a lightweight query-based decoder to generate missing high-frequency bands.

  3. CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

    cs.SD 2025-09 unverdicted novelty 6.0 of 10

    CodecSep performs prompt-driven universal sound separation directly in neural audio codec latents by combining a frozen DAC backbone with a lightweight FiLM-conditioned Transformer masker driven by CLAP embeddings, yi...

  4. Fx-Encoder++: Extracting Instrument-Wise Audio Effects Representations from Mixtures

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A contrastive-learning model, Fx-Encoder++, extracts per-instrument audio effects embeddings directly from music mixtures using audio or text queries, outperforming prior effects encoders at the mixture level.

  5. Training-Free Multi-Step Audio Source Separation

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Iteratively remixing and re-separating the input mixture, with the best blend chosen by a quality metric, improves pretrained one-step audio separation models without any retraining.

  6. Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Audio-SDS uses a pretrained text-to-audio diffusion model as a frozen critic to optimize parameters of FM synthesizers, impact simulators, and source separation latents, matching text prompts without task-specific training.

  7. How much to Dereverberate? Low-Latency Single-Channel Speech Enhancement in Distant Microphone Scenarios

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Distant-microphone, low-latency speech enhancement works better when synthetic room impulse responses link reverberation time to room volume and when the training target keeps early reflections with a 300 ms decay.

  8. A Mixture-Based Framework for Guiding Diffusion Models

    stat.ML 2025-02 conditional novelty 6.0 of 10

    MGDM approximates the intractable guided-diffusion posterior with a weighted mixture of likelihood approximations and samples the mixture using a Gibbs sampler with tunable repetitions.

  9. Hidden Echoes Survive Training in Audio To Audio Generative Instrument Models

    cs.SD 2024-12 conditional novelty 6.0 of 10

    Imperceptible echoes embedded in audio training data are reproduced by DDSP, RAVE, and Dance Diffusion models, enabling a simple watermarking method for audio-to-audio AI models.

  10. Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM

    eess.AS 2025-09 conditional novelty 5.0 of 10

    iMTI-Net combines Whisper embeddings with uncertainty-proxy statistics and a CNN-sLSTM backbone to improve non-intrusive speech intelligibility prediction on TMHINT-QI(S).

  11. Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation

    eess.AS 2025-01 conditional novelty 5.0 of 10

    VoiceFormer fuses text, video, and audio in a transformer to separate a target speaker, and stays robust when audio and video are misaligned by up to 200 ms.

  12. Evaluating the Impact of Discriminative and Generative E2E Speech Enhancement Models on Syllable Stress Preservation

    eess.AS 2024-12 conditional novelty 5.0 of 10

    A generative diffusion speech enhancement model preserves syllable stress better than discriminative enhancers for non-native English speech, and human perception matches automatic stress detection.

  13. High-Throughput Blind Co-Channel Interference Cancellation for Edge Devices Using Depthwise Separable Convolutions, Quantization, and Pruning

    eess.SP 2024-11 conditional novelty 4.0 of 10

    Depthwise separable U-Net variants for blind co-channel interference cancellation achieve higher MSE scores than WaveNet and ConvTasNet at a fraction of the computational cost.

  14. Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

    cs.SD 2026-07 conditional novelty 3.5 of 10

    MSST unifies training, validation, and inference for many music source-separation architectures and reports small quality gains from TTA, ensembling, and related engineering techniques.

Pith tools