REVIEW 5 cited by
KUIELab-MDX-Net: A Two-Stream Neural Network for Music Demixing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recently, many methods based on deep learning have been proposed for music source separation. Some state-of-the-art methods have shown that stacking many layers with many skip connections improve the SDR performance. Although such a deep and complex architecture shows outstanding performance, it usually requires numerous computing resources and time for training and evaluation. This paper proposes a two-stream neural network for music demixing, called KUIELab-MDX-Net, which shows a good balance of performance and required resources. The proposed model has a time-frequency branch and a time-domain branch, where each branch separates stems, respectively. It blends results from two streams to generate the final estimation. KUIELab-MDX-Net took second place on leaderboard A and third place on leaderboard B in the Music Demixing Challenge at ISMIR 2021. This paper also summarizes experimental results on another benchmark, MUSDB18. Our source code is available online.
Forward citations
Cited by 5 Pith papers
-
Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves
SongSQA predicts both an overall singing-quality score and a temporal segment-level score curve for full-length songs, using teacher-generated pseudo labels and a learnable attention aggregator.
-
Discrete Token Modeling for Multi-Stem Music Source Separation with Language Models
A Conformer-conditioned decoder-only language model generates discrete tokens via a neural audio codec to separate four music stems, reaching near state-of-the-art perceptual quality and top NISQA on vocals in MUSDB18...
-
PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing
PS-TTS and PS-Comet TTS use isochrony via language model paraphrasing plus phonetic synchronization with DTW on vowel distances to achieve better lip-sync and semantic preservation in automated dubbing than standard T...
-
Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling
A song-aesthetics model with multi-stem cross-attention and hierarchical interval regression beats two adapted MOS baselines on average, with some dimensions showing ties or losses.
-
Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models
MSST unifies training, validation, and inference for many music source-separation architectures and reports small quality gains from TTA, ensembling, and related engineering techniques.
Discussion (0). Sign in to comment.