Pith. sign in

REVIEW 4 cited by

Universal Score-based Speech Enhancement with High Content Preservation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.12194 v1 pith:KFZ6V5HW submitted 2024-06-18 eess.AS cs.SD

classification eess.AScs.SD
keywords speechenhancementuniversaluniverseadversarialcontentdiffusionhigh
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose UNIVERSE++, a universal speech enhancement method based on score-based diffusion and adversarial training. Specifically, we improve the existing UNIVERSE model that decouples clean speech feature extraction and diffusion. Our contributions are three-fold. First, we make several modifications to the network architecture, improving training stability and final performance. Second, we introduce an adversarial loss to promote learning high quality speech features. Third, we propose a low-rank adaptation scheme with a phoneme fidelity loss to improve content preservation in the enhanced speech. In the experiments, we train a universal enhancement model on a large scale dataset of speech degraded by noise, reverberation, and various distortions. The results on multiple public benchmark datasets demonstrate that UNIVERSE++ compares favorably to both discriminative and generative baselines for a wide range of qualitative and intelligibility metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. User-guided Generative Source Separation

    cs.SD 2025-07 conditional novelty 6.0 of 10

    GuideSep separates arbitrary target instruments from a mixture using user-provided waveform mimicry and mel-spectrogram masks, and outperforms a same-architecture mask-prediction baseline in SDR and listening tests.

  2. GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling

    eess.AS 2025-02 conditional novelty 6.0 of 10

    GenSE enhances speech by first denoising semantic tokens with a language model and then generating acoustic tokens from a single-quantizer codec, reporting higher DNSMOS, speaker similarity, and lower WER than prior systems.

  3. SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns

    eess.AS 2026-03 conditional novelty 5.5 of 10

    SEMamba++ combines Frequency GLP (FAN-based global-periodic + local conv) with multi-resolution parallel TFDP and learnable softplus mapping to outperform GSR baselines on VCTK, URGENT and AATC while remaining efficient.

  4. DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration

    eess.AS 2025-05 conditional novelty 4.0 of 10

    A 3.58M-parameter, two-stage causal speech enhancer adds a GAN second stage to DeepFilterNet2 and improves NISQA-MOS on the URGENT test set.

Pith tools