REVIEW 2 cited by
FluentSpeech: Stutter-Oriented Automatic Speech Editing with Context-Aware Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Stutter removal is an essential scenario in the field of speech editing. However, when the speech recording contains stutters, the existing text-based speech editing approaches still suffer from: 1) the over-smoothing problem in the edited speech; 2) lack of robustness due to the noise introduced by stutter; 3) to remove the stutters, users are required to determine the edited region manually. To tackle the challenges in stutter removal, we propose FluentSpeech, a stutter-oriented automatic speech editing model. Specifically, 1) we propose a context-aware diffusion model that iteratively refines the modified mel-spectrogram with the guidance of context features; 2) we introduce a stutter predictor module to inject the stutter information into the hidden sequence; 3) we also propose a stutter-oriented automatic speech editing (SASE) dataset that contains spontaneous speech recordings with time-aligned stutter labels to train the automatic stutter localization model. Experimental results on VCTK and LibriTTS datasets demonstrate that our model achieves state-of-the-art performance on speech editing. Further experiments on our SASE dataset show that FluentSpeech can effectively improve the fluency of stuttering speech in terms of objective and subjective metrics. Code and audio samples can be found at https://github.com/Zain-Jiang/Speech-Editing-Toolkit.
Forward citations
Cited by 2 Pith papers
-
ClaritySpeech: Dementia Obfuscation in Speech
An ASR, text-obfuscation, and zero-shot TTS pipeline lowers automatic dementia detection in speech by 10 to 16 percent F1 while improving intelligibility, with only moderate speaker similarity.
-
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
R-VC performs zero-shot voice conversion in two sampling steps while transferring the target speaker's rhythm, matching or exceeding prior systems in naturalness and intelligibility.
Discussion (0). Continue with ORCID to comment.