Pith. sign in

REVIEW 2 cited by

MMDenseLSTM: An efficient combination of convolutional and recurrent neural networks for audio source separation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1805.02410 v2 pith:MJPXT5KH submitted 2018-05-07 cs.SD eess.AS

classification cs.SDeess.AS
keywords audiommdensenetnetworksproposedresultsseparationsourcearchitecture
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deep neural networks have become an indispensable technique for audio source separation (ASS). It was recently reported that a variant of CNN architecture called MMDenseNet was successfully employed to solve the ASS problem of estimating source amplitudes, and state-of-the-art results were obtained for DSD100 dataset. To further enhance MMDenseNet, here we propose a novel architecture that integrates long short-term memory (LSTM) in multiple scales with skip connections to efficiently model long-term structures within an audio context. The experimental results show that the proposed method outperforms MMDenseNet, LSTM and a blend of the two networks. The number of parameters and processing time of the proposed model are significantly less than those for simple blending. Furthermore, the proposed method yields better results than those obtained using ideal binary masks for a singing voice separation task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed

    cs.SD 2019-09 conditional novelty 6.0 of 10

    A waveform-domain convolutional-recurrent network with a remix-silence semi-supervised scheme reaches near state-of-the-art music source separation on MusDB without extra labeled data.

  2. Interleaved Multitask Learning for Audio Source Separation with Independent Databases

    cs.SD 2019-08 conditional novelty 6.0 of 10

    An interleaved multitask training procedure for a shared-encoder source separation network enables training on independent per-source databases and yields SIR improvements over simultaneous multitask training.

Pith tools