Pith. sign in

REVIEW 8 cited by

Real Time Speech Enhancement in the Waveform Domain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.12847 v3 pith:JRNDELVT submitted 2020-06-23 eess.AS cs.LGcs.SDstat.ML

classification eess.AScs.LGcs.SDstat.ML
keywords modelwaveformcausaldirectlyenhancementperformanceproposedspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a causal speech enhancement model working on the raw waveform that runs in real-time on a laptop CPU. The proposed model is based on an encoder-decoder architecture with skip-connections. It is optimized on both time and frequency domains, using multiple loss functions. Empirical evidence shows that it is capable of removing various kinds of background noise including stationary and non-stationary noises, as well as room reverb. Additionally, we suggest a set of data augmentation techniques applied directly on the raw waveform which further improve model performance and its generalization abilities. We perform evaluations on several standard benchmarks, both using objective metrics and human judgements. The proposed model matches state-of-the-art performance of both causal and non causal methods while working directly on the raw waveform.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Echo-Aware Modulation for Compact-Latent Frequency-Time Modeling in Lightweight Acoustic Echo Cancellation

    eess.AS 2026-08 conditional novelty 6.0 of 10

    An echo-aware modulation module recovers frequency-time detail in Bark-domain lightweight acoustic echo cancellation, improving quality at modest extra cost.

  2. Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil

    eess.AS 2026-07 conditional novelty 6.0 of 10

    State-of-the-art audio deepfake detectors severely degrade on Brazilian Portuguese political speech, and the main source of performance gaps is the synthesis method, not demographic traits.

  3. Training-Free Intelligibility-Guided Observation Addition for Noisy ASR

    eess.AS 2026-02 conditional novelty 6.0 of 10

    Mixing noisy and enhanced speech with weights derived from the recognizer's confidence on each signal reduces ASR word error rate without any additional training.

  4. Affine Modulation-based Audiogram Fusion Network for Joint Noise Reduction and Hearing Loss Compensation

    eess.AS 2025-09 conditional novelty 6.0 of 10

    A hearing-aid network that injects the user's audiogram into a speech-enhancement model with affine modulation beats existing joint noise-reduction and compensation systems on objective quality metrics.

  5. Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition

    cs.LG 2025-05 conditional novelty 5.0 of 10

    On a single-speaker MRI speech corpus, adding vocal-tract video to audio does not improve phoneme recognition, but attention analysis shows articulatory cues can lead acoustic cues in time.

  6. FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching

    eess.AS 2025-05 reject novelty 5.0 of 10

    FlowSE applies rectified flow matching with a DiT backbone to speech enhancement, reporting better DNSMOS and WER results and a much lower real-time factor than diffusion baselines.

  7. DPDFNet: Boosting DeepFilterNet2 via Dual-Path RNN

    cs.SD 2025-12 conditional novelty 4.0 of 10

    DPDFNet inserts dual-path RNN blocks into DeepFilterNet2's encoder, adds an over-attenuation loss and long-context fine-tuning, and reports superior causal speech enhancement on a 12-language low-SNR test set.

  8. Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation

    eess.AS 2025-05 conditional novelty 3.0 of 10

    A Transformer-Mamba model that adds a learned correction signal to degraded speech beats adapted active-noise-control baselines on denoising, dereverberation, and declipping in simulation.

Pith tools