Pith. sign in

REVIEW 5 cited by

KUIELab-MDX-Net: A Two-Stream Neural Network for Music Demixing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.12203 v1 pith:2CO4EFVX submitted 2021-11-24 eess.AS cs.SD

classification eess.AScs.SD
keywords musicbranchdemixingkuielab-mdx-netmanyperformancedeepleaderboard
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, many methods based on deep learning have been proposed for music source separation. Some state-of-the-art methods have shown that stacking many layers with many skip connections improve the SDR performance. Although such a deep and complex architecture shows outstanding performance, it usually requires numerous computing resources and time for training and evaluation. This paper proposes a two-stream neural network for music demixing, called KUIELab-MDX-Net, which shows a good balance of performance and required resources. The proposed model has a time-frequency branch and a time-domain branch, where each branch separates stems, respectively. It blends results from two streams to generate the final estimation. KUIELab-MDX-Net took second place on leaderboard A and third place on leaderboard B in the Music Demixing Challenge at ISMIR 2021. This paper also summarizes experimental results on another benchmark, MUSDB18. Our source code is available online.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves

    cs.SD 2026-07 conditional novelty 6.0 of 10

    SongSQA predicts both an overall singing-quality score and a temporal segment-level score curve for full-length songs, using teacher-generated pseudo labels and a learnable attention aggregator.

  2. Discrete Token Modeling for Multi-Stem Music Source Separation with Language Models

    eess.AS 2026-04 unverdicted novelty 6.0 of 10

    A Conformer-conditioned decoder-only language model generates discrete tokens via a neural audio codec to separate four music stems, reaching near state-of-the-art perceptual quality and top NISQA on vocals in MUSDB18...

  3. PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing

    eess.AS 2026-04 unverdicted novelty 6.0 of 10

    PS-TTS and PS-Comet TTS use isochrony via language model paraphrasing plus phonetic synchronization with DTW on vowel distances to achieve better lip-sync and semantic preservation in automated dubbing than standard T...

  4. Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling

    cs.SD 2026-01 conditional novelty 6.0 of 10

    A song-aesthetics model with multi-stem cross-attention and hierarchical interval regression beats two adapted MOS baselines on average, with some dimensions showing ties or losses.

  5. Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

    cs.SD 2026-07 conditional novelty 3.5 of 10

    MSST unifies training, validation, and inference for many music source-separation architectures and reports small quality gains from TTA, ensembling, and related engineering techniques.

Pith tools