REVIEW 2 cited by
SpeechPainter: Text-conditioned Speech Inpainting
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose SpeechPainter, a model for filling in gaps of up to one second in speech samples by leveraging an auxiliary textual input. We demonstrate that the model performs speech inpainting with the appropriate content, while maintaining speaker identity, prosody and recording environment conditions, and generalizing to unseen speakers. Our approach significantly outperforms baselines constructed using adaptive TTS, as judged by human raters in side-by-side preference and MOS tests.
Forward citations
Cited by 2 Pith papers
-
Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing
A multi-codebook discrete diffusion model with coarse-to-fine RVQ generation and span-localized guidance improves speech inpainting and editing on RealEdit.
-
A2SB: Audio-to-Audio Schrodinger Bridges
A2SB applies Schrödinger bridges to music restoration, achieving state-of-the-art bandwidth extension and inpainting at 44.1kHz in a single vocoder-free model.
Discussion (0). Continue with ORCID to comment.