Pith. sign in

REVIEW 7 cited by

VoiceFixer: Toward General Speech Restoration with Neural Vocoder

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.13731 v3 pith:634T7KWV submitted 2021-09-28 cs.SD eess.AS

classification cs.SDeess.AS
keywords speechvoicefixerrestorationmodelstageanalysisdistortionsgeneral
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speech restoration aims to remove distortions in speech signals. Prior methods mainly focus on single-task speech restoration (SSR), such as speech denoising or speech declipping. However, SSR systems only focus on one task and do not address the general speech restoration problem. In addition, previous SSR systems show limited performance in some speech restoration tasks such as speech super-resolution. To overcome those limitations, we propose a general speech restoration (GSR) task that attempts to remove multiple distortions simultaneously. Furthermore, we propose VoiceFixer, a generative framework to address the GSR task. VoiceFixer consists of an analysis stage and a synthesis stage to mimic the speech analysis and comprehension of the human auditory system. We employ a ResUNet to model the analysis stage and a neural vocoder to model the synthesis stage. We evaluate VoiceFixer with additive noise, room reverberation, low-resolution, and clipping distortions. Our baseline GSR model achieves a 0.499 higher mean opinion score (MOS) than the speech enhancement SSR model. VoiceFixer further surpasses the GSR baseline model on the MOS score by 0.256. Moreover, we observe that VoiceFixer generalizes well to severely degraded real speech recordings, indicating its potential in restoring old movies and historical speeches. The source code is available at https://github.com/haoheliu/voicefixer_main.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films

    cs.CV 2026-08 conditional novelty 7.0 of 10

    A single 22B-parameter model jointly restores degraded video and audio of historical films, beating prior separate restorers on visual, audio, and sync metrics.

  2. On the Geometry of Music Bandwidth Extension in Latent Spaces of Audio Codecs

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A single mean shift in the latent space of several neural codecs achieves competitive music bandwidth extension on some metrics, implying a largely linear structure.

  3. Semantic Sampling via Learnable Observation Front Ends

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Learnable acoustic filterbanks plus constrained mixing and temporal readout produce more informative low-rate observations for speech reconstruction than fixed waveform sampling at the same budget.

  4. Audio-Visual World Models: Learning Physically Grounded Multisensory Dynamics

    cs.MM 2025-11 conditional novelty 6.0 of 10

    Future video frames, binaural audio, and task rewards can be generated jointly under action control by a diffusion transformer trained on a new 30-hour simulated audio-visual benchmark, with small downstream navigation gains.

  5. UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement

    cs.SD 2025-10 conditional novelty 6.0 of 10

    A 63M-parameter decoder-only LM, UniSE, unifies speech restoration, target speaker extraction, and speech separation by generating BiCodec discrete tokens under task-specific prompts.

  6. Music Source Restoration

    cs.SD 2025-05 conditional novelty 6.0 of 10

    The paper defines Music Source Restoration, releases a 578-song annotated RawStems dataset of supposedly unprocessed hierarchical stems, and shows a U-Former baseline that improves mel-SSIM but often leaves SI-SDR nea...

  7. Inference-time Scaling for Diffusion-based Audio Super-resolution

    cs.SD 2025-08 conditional novelty 4.0 of 10

    Generating 120 candidate super-resolved audios and choosing the best by task-specific verifiers improves speech, music, and sound effects over single-sample diffusion output, at 120x compute.

Pith tools