Pith. sign in

REVIEW 5 cited by

VoiceFixer: A Unified Framework for High-Fidelity Speech Restoration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.05841 v2 pith:GAUTWXAL submitted 2022-04-12 eess.AS cs.SDeess.SP

classification eess.AScs.SDeess.SP
keywords speechvoicefixerdegradeddistortionsrestorationhigh-fidelityframeworkmultiple
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Speech restoration aims to remove distortions in speech signals. Prior methods mainly focus on a single type of distortion, such as speech denoising or dereverberation. However, speech signals can be degraded by several different distortions simultaneously in the real world. It is thus important to extend speech restoration models to deal with multiple distortions. In this paper, we introduce VoiceFixer, a unified framework for high-fidelity speech restoration. VoiceFixer restores speech from multiple distortions (e.g., noise, reverberation, and clipping) and can expand degraded speech (e.g., noisy speech) with a low bandwidth to 44.1 kHz full-bandwidth high-fidelity speech. We design VoiceFixer based on (1) an analysis stage that predicts intermediate-level features from the degraded speech, and (2) a synthesis stage that generates waveform using a neural vocoder. Both objective and subjective evaluations show that VoiceFixer is effective on severely degraded speech, such as real-world historical speech recordings. Samples of VoiceFixer are available at https://haoheliu.github.io/voicefixer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How much to Dereverberate? Low-Latency Single-Channel Speech Enhancement in Distant Microphone Scenarios

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Distant-microphone, low-latency speech enhancement works better when synthetic room impulse responses link reverberation time to room volume and when the training target keeps early reflections with a 300 ms decay.

  2. Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A diffusion voice conversion model, conditioned on clean speaker embeddings and HuBERT content features, is applied after a generative speech restorer to achieve state-of-the-art-comparable speech quality.

  3. Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

    cs.SD 2025-02 conditional novelty 5.0 of 10

    A masked generative model pre-trained on unlabeled speech then fine-tuned per task matches or beats task-specific systems across TTS, voice conversion, speaker extraction, enhancement, and lip-to-speech.

  4. Overview of the Amphion Toolkit (v0.2)

    cs.SD 2025-01 conditional novelty 4.0 of 10

    Amphion v0.2 is an open-source toolkit for audio, music, and speech generation, adding a 101K-hour multilingual dataset, processing pipelines, and pretrained models.

  5. Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation

    eess.AS 2025-05 conditional novelty 3.0 of 10

    A Transformer-Mamba model that adds a learned correction signal to degraded speech beats adapted active-noise-control baselines on denoising, dereverberation, and declipping in simulation.

Pith tools