Pith. sign in

REVIEW 16 cited by

Real Time Speech Enhancement in the Waveform Domain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.12847 v3 pith:JRNDELVT submitted 2020-06-23 eess.AS cs.LGcs.SDstat.ML

classification eess.AScs.LGcs.SDstat.ML
keywords modelwaveformcausaldirectlyenhancementperformanceproposedspeech
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present a causal speech enhancement model working on the raw waveform that runs in real-time on a laptop CPU. The proposed model is based on an encoder-decoder architecture with skip-connections. It is optimized on both time and frequency domains, using multiple loss functions. Empirical evidence shows that it is capable of removing various kinds of background noise including stationary and non-stationary noises, as well as room reverb. Additionally, we suggest a set of data augmentation techniques applied directly on the raw waveform which further improve model performance and its generalization abilities. We perform evaluations on several standard benchmarks, both using objective metrics and human judgements. The proposed model matches state-of-the-art performance of both causal and non causal methods while working directly on the raw waveform.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Let SSMs be ConvNets: State-space Modeling with Optimal Tensor Contractions

    cs.LG 2025-01 conditional novelty 7.0 of 10

    Treating state-space layers as tensor networks with CNN-style connectivity and optimized contraction orders yields hybrid SSM networks that outperform homogeneous SSMs on raw audio tasks and enable competitive streami...

  2. HybridSB-MoE: Dual-Domain Schr\"odinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A dual-domain Schrödinger Bridge mixture-of-experts system with asymmetric uncertainty fusion reports a PESQ of 3.88 on VoiceBank+DEMAND, though the supporting theory is left loose and unvalidated.

  3. Echo-Aware Modulation for Compact-Latent Frequency-Time Modeling in Lightweight Acoustic Echo Cancellation

    eess.AS 2026-08 conditional novelty 6.0 of 10

    An echo-aware modulation module recovers frequency-time detail in Bark-domain lightweight acoustic echo cancellation, improving quality at modest extra cost.

  4. Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil

    eess.AS 2026-07 conditional novelty 6.0 of 10

    State-of-the-art audio deepfake detectors severely degrade on Brazilian Portuguese political speech, and the main source of performance gaps is the synthesis method, not demographic traits.

  5. Training-Free Intelligibility-Guided Observation Addition for Noisy ASR

    eess.AS 2026-02 conditional novelty 6.0 of 10

    Mixing noisy and enhanced speech with weights derived from the recognizer's confidence on each signal reduces ASR word error rate without any additional training.

  6. Affine Modulation-based Audiogram Fusion Network for Joint Noise Reduction and Hearing Loss Compensation

    eess.AS 2025-09 conditional novelty 6.0 of 10

    A hearing-aid network that injects the user's audiogram into a speech-enhancement model with affine modulation beats existing joint noise-reduction and compensation systems on objective quality metrics.

  7. How much to Dereverberate? Low-Latency Single-Channel Speech Enhancement in Distant Microphone Scenarios

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Distant-microphone, low-latency speech enhancement works better when synthetic room impulse responses link reverberation time to room volume and when the training target keeps early reflections with a 300 ms decay.

  8. Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment

    cs.SD 2025-01 conditional novelty 6.0 of 10

    MUTUD trains audiovisual speech models with both modalities but lets them run with audio only, recovering much of the multimodal benefit at a fraction of the compute.

  9. From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology

    eess.AS 2024-12 conditional novelty 6.0 of 10

    On VoiceBank-DEMAND, replacing dense layers or ReLU activations with GR-KAN layers in MP-SENet and Demucs improved PESQ by up to 0.1 and cut parameters by up to 4x in one comparison.

  10. Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition

    cs.LG 2025-05 conditional novelty 5.0 of 10

    On a single-speaker MRI speech corpus, adding vocal-tract video to audio does not improve phoneme recognition, but attention analysis shows articulatory cues can lead acoustic cues in time.

  11. FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching

    eess.AS 2025-05 reject novelty 5.0 of 10

    FlowSE applies rectified flow matching with a DiT backbone to speech enhancement, reporting better DNSMOS and WER results and a much lower real-time factor than diffusion baselines.

  12. Scalable Speech Enhancement with Dynamic Channel Pruning

    eess.AS 2024-12 conditional novelty 5.0 of 10

    A custom convolutional speech enhancement network with a learned gating module skips individual channels at runtime, saving up to 29.6% of MACs on VoiceBank+DEMAND with a negligible PESQ drop.

  13. Evaluating the Impact of Discriminative and Generative E2E Speech Enhancement Models on Syllable Stress Preservation

    eess.AS 2024-12 conditional novelty 5.0 of 10

    A generative diffusion speech enhancement model preserves syllable stress better than discriminative enhancers for non-native English speech, and human perception matches automatic stress detection.

  14. DPDFNet: Boosting DeepFilterNet2 via Dual-Path RNN

    cs.SD 2025-12 conditional novelty 4.0 of 10

    DPDFNet inserts dual-path RNN blocks into DeepFilterNet2's encoder, adds an over-attenuation loss and long-context fine-tuning, and reports superior causal speech enhancement on a 12-language low-SNR test set.

  15. Noisereduce: Domain General Noise Reduction for Time Series Signals

    eess.SP 2024-12 conditional novelty 4.0 of 10

    Noisereduce, a no-training spectral gating method, outperforms classical noise reduction baselines across speech, bioacoustics, neurophysiology, and seismology, and is a fast, domain-general baseline.

  16. Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation

    eess.AS 2025-05 conditional novelty 3.0 of 10

    A Transformer-Mamba model that adds a learned correction signal to degraded speech beats adapted active-noise-control baselines on denoising, dereverberation, and declipping in simulation.

Pith tools