REVIEW 7 cited by
VoiceFixer: Toward General Speech Restoration with Neural Vocoder
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Speech restoration aims to remove distortions in speech signals. Prior methods mainly focus on single-task speech restoration (SSR), such as speech denoising or speech declipping. However, SSR systems only focus on one task and do not address the general speech restoration problem. In addition, previous SSR systems show limited performance in some speech restoration tasks such as speech super-resolution. To overcome those limitations, we propose a general speech restoration (GSR) task that attempts to remove multiple distortions simultaneously. Furthermore, we propose VoiceFixer, a generative framework to address the GSR task. VoiceFixer consists of an analysis stage and a synthesis stage to mimic the speech analysis and comprehension of the human auditory system. We employ a ResUNet to model the analysis stage and a neural vocoder to model the synthesis stage. We evaluate VoiceFixer with additive noise, room reverberation, low-resolution, and clipping distortions. Our baseline GSR model achieves a 0.499 higher mean opinion score (MOS) than the speech enhancement SSR model. VoiceFixer further surpasses the GSR baseline model on the MOS score by 0.256. Moreover, we observe that VoiceFixer generalizes well to severely degraded real speech recordings, indicating its potential in restoring old movies and historical speeches. The source code is available at https://github.com/haoheliu/voicefixer_main.
Forward citations
Cited by 7 Pith papers
-
OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films
A single 22B-parameter model jointly restores degraded video and audio of historical films, beating prior separate restorers on visual, audio, and sync metrics.
-
On the Geometry of Music Bandwidth Extension in Latent Spaces of Audio Codecs
A single mean shift in the latent space of several neural codecs achieves competitive music bandwidth extension on some metrics, implying a largely linear structure.
-
Semantic Sampling via Learnable Observation Front Ends
Learnable acoustic filterbanks plus constrained mixing and temporal readout produce more informative low-rate observations for speech reconstruction than fixed waveform sampling at the same budget.
-
Audio-Visual World Models: Learning Physically Grounded Multisensory Dynamics
Future video frames, binaural audio, and task rewards can be generated jointly under action control by a diffusion transformer trained on a new 30-hour simulated audio-visual benchmark, with small downstream navigation gains.
-
UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement
A 63M-parameter decoder-only LM, UniSE, unifies speech restoration, target speaker extraction, and speech separation by generating BiCodec discrete tokens under task-specific prompts.
-
Music Source Restoration
The paper defines Music Source Restoration, releases a 578-song annotated RawStems dataset of supposedly unprocessed hierarchical stems, and shows a U-Former baseline that improves mel-SSIM but often leaves SI-SDR nea...
-
Inference-time Scaling for Diffusion-based Audio Super-resolution
Generating 120 candidate super-resolved audios and choosing the best by task-specific verifiers improves speech, music, and sound effects over single-sample diffusion output, at 120x compute.
Discussion (0). Sign in to comment.