REVIEW 4 cited by
Universal Score-based Speech Enhancement with High Content Preservation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose UNIVERSE++, a universal speech enhancement method based on score-based diffusion and adversarial training. Specifically, we improve the existing UNIVERSE model that decouples clean speech feature extraction and diffusion. Our contributions are three-fold. First, we make several modifications to the network architecture, improving training stability and final performance. Second, we introduce an adversarial loss to promote learning high quality speech features. Third, we propose a low-rank adaptation scheme with a phoneme fidelity loss to improve content preservation in the enhanced speech. In the experiments, we train a universal enhancement model on a large scale dataset of speech degraded by noise, reverberation, and various distortions. The results on multiple public benchmark datasets demonstrate that UNIVERSE++ compares favorably to both discriminative and generative baselines for a wide range of qualitative and intelligibility metrics.
Forward citations
Cited by 4 Pith papers
-
User-guided Generative Source Separation
GuideSep separates arbitrary target instruments from a mixture using user-provided waveform mimicry and mel-spectrogram masks, and outperforms a same-architecture mask-prediction baseline in SDR and listening tests.
-
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
GenSE enhances speech by first denoising semantic tokens with a language model and then generating acoustic tokens from a single-quantizer codec, reporting higher DNSMOS, speaker similarity, and lower WER than prior systems.
-
SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
SEMamba++ combines Frequency GLP (FAN-based global-periodic + local conv) with multi-resolution parallel TFDP and learnable softplus mapping to outperform GSR baselines on VCTK, URGENT and AATC while remaining efficient.
-
DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration
A 3.58M-parameter, two-stage causal speech enhancer adds a GAN second stage to DeepFilterNet2 and improves NISQA-MOS on the URGENT test set.
Discussion (0). Continue with ORCID to comment.