REVIEW 8 cited by
Real Time Speech Enhancement in the Waveform Domain
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present a causal speech enhancement model working on the raw waveform that runs in real-time on a laptop CPU. The proposed model is based on an encoder-decoder architecture with skip-connections. It is optimized on both time and frequency domains, using multiple loss functions. Empirical evidence shows that it is capable of removing various kinds of background noise including stationary and non-stationary noises, as well as room reverb. Additionally, we suggest a set of data augmentation techniques applied directly on the raw waveform which further improve model performance and its generalization abilities. We perform evaluations on several standard benchmarks, both using objective metrics and human judgements. The proposed model matches state-of-the-art performance of both causal and non causal methods while working directly on the raw waveform.
Forward citations
Cited by 8 Pith papers
-
Echo-Aware Modulation for Compact-Latent Frequency-Time Modeling in Lightweight Acoustic Echo Cancellation
An echo-aware modulation module recovers frequency-time detail in Bark-domain lightweight acoustic echo cancellation, improving quality at modest extra cost.
-
Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil
State-of-the-art audio deepfake detectors severely degrade on Brazilian Portuguese political speech, and the main source of performance gaps is the synthesis method, not demographic traits.
-
Training-Free Intelligibility-Guided Observation Addition for Noisy ASR
Mixing noisy and enhanced speech with weights derived from the recognizer's confidence on each signal reduces ASR word error rate without any additional training.
-
Affine Modulation-based Audiogram Fusion Network for Joint Noise Reduction and Hearing Loss Compensation
A hearing-aid network that injects the user's audiogram into a speech-enhancement model with affine modulation beats existing joint noise-reduction and compensation systems on objective quality metrics.
-
Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition
On a single-speaker MRI speech corpus, adding vocal-tract video to audio does not improve phoneme recognition, but attention analysis shows articulatory cues can lead acoustic cues in time.
-
FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
FlowSE applies rectified flow matching with a DiT backbone to speech enhancement, reporting better DNSMOS and WER results and a much lower real-time factor than diffusion baselines.
-
DPDFNet: Boosting DeepFilterNet2 via Dual-Path RNN
DPDFNet inserts dual-path RNN blocks into DeepFilterNet2's encoder, adds an over-attenuation loss and long-context fine-tuning, and reports superior causal speech enhancement on a 12-language low-SNR test set.
-
Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation
A Transformer-Mamba model that adds a learned correction signal to degraded speech beats adapted active-noise-control baselines on denoising, dereverberation, and declipping in simulation.
Discussion (0). Sign in to comment.