Introduces a universal real-time speech enhancement model with controllable algorithmic latency via parallel convolutional layers and computational latency via early exits, using a two-stage training approach.
V oicefixer: Toward general speech restoration with neural vocoder.arXiv:2109.13731
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.SD 3verdicts
UNVERDICTED 3representative citing papers
Replacing early-reflected speech with time-shifted anechoic clean speech as the training target, combined with a two-stage distortion-perception framework, yields state-of-the-art universal speech enhancement.
SonicMaster is a text-conditioned flow-matching generative model for unified music restoration and mastering, trained on a dataset of simulated degradations across equalization, dynamics, reverb, amplitude, and stereo.
citing papers explorer
-
One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications
Introduces a universal real-time speech enhancement model with controllable algorithmic latency via parallel convolutional layers and computational latency via early exits, using a two-stage training approach.
-
Rethinking Training Targets, Architectures and Data Quality for Universal Speech Enhancement
Replacing early-reflected speech with time-shifted anechoic clean speech as the training target, combined with a two-stage distortion-perception framework, yields state-of-the-art universal speech enhancement.
-
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
SonicMaster is a text-conditioned flow-matching generative model for unified music restoration and mastering, trained on a dataset of simulated degradations across equalization, dynamics, reverb, amplitude, and stereo.