Pith. sign in

REVIEW 11 cited by

DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.00264 v4 pith:GSTQWTNY submitted 2020-08-01 eess.AS cs.SD

classification eess.AScs.SD
keywords networkcomplexconvolutiondccrndeeprecurrentcomplex-valuedspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speech enhancement has benefited from the success of deep learning in terms of intelligibility and perceptual quality. Conventional time-frequency (TF) domain methods focus on predicting TF-masks or speech spectrum, via a naive convolution neural network (CNN) or recurrent neural network (RNN). Some recent studies use complex-valued spectrogram as a training target but train in a real-valued network, predicting the magnitude and phase component or real and imaginary part, respectively. Particularly, convolution recurrent network (CRN) integrates a convolutional encoder-decoder (CED) structure and long short-term memory (LSTM), which has been proven to be helpful for complex targets. In order to train the complex target more effectively, in this paper, we design a new network structure simulating the complex-valued operation, called Deep Complex Convolution Recurrent Network (DCCRN), where both CNN and RNN structures can handle complex-valued operation. The proposed DCCRN models are very competitive over other previous networks, either on objective or subjective metric. With only 3.7M parameters, our DCCRN models submitted to the Interspeech 2020 Deep Noise Suppression (DNS) challenge ranked first for the real-time-track and second for the non-real-time track in terms of Mean Opinion Score (MOS).

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Ranking the Impact of Contextual Specialization in Neural Speech Enhancement

    eess.AS 2026-07 accept novelty 6.0 of 10

    Speaker-identity specialization produces the largest gains in neural speech enhancement, letting small models match or exceed generalists ten times their size.

  2. CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

    cs.SD 2025-09 unverdicted novelty 6.0 of 10

    CodecSep performs prompt-driven universal sound separation directly in neural audio codec latents by combining a frozen DAC backbone with a lightweight FiLM-conditioned Transformer masker driven by CLAP embeddings, yi...

  3. Affine Modulation-based Audiogram Fusion Network for Joint Noise Reduction and Hearing Loss Compensation

    eess.AS 2025-09 conditional novelty 6.0 of 10

    A hearing-aid network that injects the user's audiogram into a speech-enhancement model with affine modulation beats existing joint noise-reduction and compensation systems on objective quality metrics.

  4. UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling

    eess.AS 2025-08 conditional novelty 6.0 of 10

    UniFlow unifies four speech front-end tasks in one continuous-latent generative model with task-ID conditioning and reports competitive, but not uniformly superior, benchmark scores.

  5. User-guided Generative Source Separation

    cs.SD 2025-07 conditional novelty 6.0 of 10

    GuideSep separates arbitrary target instruments from a mixture using user-provided waveform mimicry and mel-spectrogram masks, and outperforms a same-architecture mask-prediction baseline in SDR and listening tests.

  6. FlowTSE: Target Speaker Extraction with Flow Matching

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Conditional flow matching on mel-spectrograms with a phase-conditioned vocoder matches or beats published TSE baselines on Libri2Mix.

  7. SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns

    eess.AS 2026-03 conditional novelty 5.5 of 10

    SEMamba++ combines Frequency GLP (FAN-based global-periodic + local conv) with multi-resolution parallel TFDP and learnable softplus mapping to outperform GSR baselines on VCTK, URGENT and AATC while remaining efficient.

  8. SaD: A Scenario-Aware Discriminator for Speech Enhancement

    cs.SD 2025-08 conditional novelty 5.0 of 10

    A scenario-aware discriminator that predicts a frequency division point and scores high/low bands separately improves GAN-based speech enhancement on several quality metrics, with some STOI declines.

  9. Infant Cry Detection In Noisy Environment Using Blueprint Separable Convolutions and Time-Frequency Recurrent Neural Network

    cs.SD 2025-08 conditional novelty 5.0 of 10

    A 1.54M-parameter CNN-RNN with blueprint separable convolutions, time-frequency BiLSTM/LSTM, and spatial/channel attention achieves state-of-the-art infant cry detection on a merged public dataset across SNRs down to -20 dB.

  10. FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching

    eess.AS 2025-05 reject novelty 5.0 of 10

    FlowSE applies rectified flow matching with a DiT backbone to speech enhancement, reporting better DNSMOS and WER results and a much lower real-time factor than diffusion baselines.

  11. Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation

    eess.AS 2025-05 conditional novelty 3.0 of 10

    A Transformer-Mamba model that adds a learned correction signal to degraded speech beats adapted active-noise-control baselines on denoising, dereverberation, and declipping in simulation.

Pith tools