A conditional flow matching model using WavLM-derived discrete units converts dysarthric speech to a synthesized clean voice with 31.3% WER and 3.9 MOS, outperforming a mel-spectrogram model (84.1% WER).
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching
A conditional flow matching model using WavLM-derived discrete units converts dysarthric speech to a synthesized clean voice with 31.3% WER and 3.9 MOS, outperforming a mel-spectrogram model (84.1% WER).