REVIEW 3 cited by
Learning Explicit Prosody Models and Deep Speaker Embeddings for Atypical Voice Conversion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Though significant progress has been made for the voice conversion (VC) of typical speech, VC for atypical speech, e.g., dysarthric and second-language (L2) speech, remains a challenge, since it involves correcting for atypical prosody while maintaining speaker identity. To address this issue, we propose a VC system with explicit prosodic modelling and deep speaker embedding (DSE) learning. First, a speech-encoder strives to extract robust phoneme embeddings from atypical speech. Second, a prosody corrector takes in phoneme embeddings to infer typical phoneme duration and pitch values. Third, a conversion model takes phoneme embeddings and typical prosody features as inputs to generate the converted speech, conditioned on the target DSE that is learned via speaker encoder or speaker adaptation. Extensive experiments demonstrate that speaker adaptation can achieve higher speaker similarity, and the speaker encoder based conversion model can greatly reduce dysarthric and non-native pronunciation patterns with improved speech intelligibility. A comparison of speech recognition results between the original dysarthric speech and converted speech show that absolute reduction of 47.6% character error rate (CER) and 29.3% word error rate (WER) can be achieved.
Forward citations
Cited by 3 Pith papers
-
ClaritySpeech: Dementia Obfuscation in Speech
An ASR, text-obfuscation, and zero-shot TTS pipeline lowers automatic dementia detection in speech by 10 to 16 percent F1 while improving intelligibility, with only moderate speaker similarity.
-
Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters
Fast-VGAN is a lightweight GAN-based voice converter that explicitly controls F0, phoneme timing, and intensity, achieving near-perfect intelligibility and competitive speaker similarity on a small test set.
-
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
A latent diffusion model with SSL-based content restoration and in-context speaker prompts improves dysarthric speech intelligibility and speaker similarity on UASpeech.
Discussion (0). Sign in to comment.