Pith. sign in

REVIEW 3 cited by

Speech Denoising with Deep Feature Losses

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.10522 v2 pith:FOERN3JY submitted 2018-06-27 eess.AS cs.SD

classification eess.AScs.SD
keywords speechdeepnetworkapproachdenoisingfeatureaudiobackground
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present an end-to-end deep learning approach to denoising speech signals by processing the raw waveform directly. Given input audio containing speech corrupted by an additive background signal, the system aims to produce a processed signal that contains only the speech content. Recent approaches have shown promising results using various deep network architectures. In this paper, we propose to train a fully-convolutional context aggregation network using a deep feature loss. That loss is based on comparing the internal feature activations in a different network, trained for acoustic environment detection and domestic audio tagging. Our approach outperforms the state-of-the-art in objective speech quality metrics and in large-scale perceptual experiments with human listeners. It also outperforms an identical network trained using traditional regression losses. The advantage of the new approach is particularly pronounced for the hardest data with the most intrusive background noise, for which denoising is most needed and most challenging.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs

    cs.SD 2026-08 conditional novelty 6.0 of 10

    An inaudible 5-20 Hz intermittent waveform degrades performance across six audio-language models, and a spectral-clustering requery guard partially recovers it.

  2. Deep Speech Synthesis from Multimodal Articulatory Representations

    eess.AS 2024-12 conditional novelty 5.0 of 10

    Multimodal pre-training across EMA, MRI, and EMG articulatory streams cuts MRI-to-speech word error rate from 69.5% to 33.4% in a single-speaker low-resource setting, with similar gains for EMG-to-speech.

  3. Coarse-to-fine Optimization for Speech Enhancement

    cs.SD 2019-08 conditional novelty 5.0 of 10

    A coarse-to-fine cosine similarity loss with a dynamic perceptual loss improves objective speech enhancement scores on the Valentini benchmark.

Pith tools