Pith. sign in

REVIEW 6 cited by

MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.03538 v2 pith:VJ5U4WAG submitted 2021-04-08 cs.SD cs.AIeess.AS

MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement

classification cs.SD cs.AIeess.AS
keywords metricganspeechmetricstrainingenhancementevaluationhumanobjective
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The discrepancy between the cost function used for training a speech enhancement model and human auditory perception usually makes the quality of enhanced speech unsatisfactory. Objective evaluation metrics which consider human perception can hence serve as a bridge to reduce the gap. Our previously proposed MetricGAN was designed to optimize objective metrics by connecting the metric with a discriminator. Because only the scores of the target evaluation functions are needed during training, the metrics can even be non-differentiable. In this study, we propose a MetricGAN+ in which three training techniques incorporating domain-knowledge of speech processing are proposed. With these techniques, experimental results on the VoiceBank-DEMAND dataset show that MetricGAN+ can increase PESQ score by 0.3 compared to the previous MetricGAN and achieve state-of-the-art results (PESQ score = 3.15).

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil

    eess.AS 2026-07 conditional novelty 6.0

    State-of-the-art audio deepfake detectors severely degrade on Brazilian Portuguese political speech, and the main source of performance gaps is the synthesis method, not demographic traits.

  2. Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis

    eess.AS 2026-07 conditional novelty 6.0

    Magnitude strength, not estimated phase, drives SE-induced ASR degradation, and the optimal strength is recognizer-dependent (strong for wav2vec 2.0, mild for Whisper).

  3. LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement

    cs.SD 2026-03 conditional novelty 6.0

    Using LLM-generated text descriptions of enhanced speech converted to sentiment scores as PPO rewards improves PESQ, STOI, and neural quality scores over supervised and DNSMOS-reward baselines on AVSEC-4.

  4. SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns

    eess.AS 2026-03 conditional novelty 5.5

    SEMamba++ combines Frequency GLP (FAN-based global-periodic + local conv) with multi-resolution parallel TFDP and learnable softplus mapping to outperform GSR baselines on VCTK, URGENT and AATC while remaining efficient.

  5. SaD: A Scenario-Aware Discriminator for Speech Enhancement

    cs.SD 2025-08 conditional novelty 5.0

    A scenario-aware discriminator that predicts a frequency division point and scores high/low bands separately improves GAN-based speech enhancement on several quality metrics, with some STOI declines.

  6. Audio Editing in the Era of Foundation Models: A Survey

    eess.AS 2026-06 unverdicted novelty 3.0

    A survey that presents a unified taxonomy of audio editing tasks, summarizes training-based and training-free foundation model approaches, reviews datasets and evaluation protocols, and identifies future challenges.