Multimodal pre-training across EMA, MRI, and EMG articulatory streams cuts MRI-to-speech word error rate from 69.5% to 33.4% in a single-speaker low-resource setting, with similar gains for EMG-to-speech.
Since these mod- els are biologically grounded, they can be applied to decod- ing speech from biosignals for health technology applications [6, 7, 8, 9, 10, 11, 12, 13]
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Deep Speech Synthesis from Multimodal Articulatory Representations
Multimodal pre-training across EMA, MRI, and EMG articulatory streams cuts MRI-to-speech word error rate from 69.5% to 33.4% in a single-speaker low-resource setting, with similar gains for EMG-to-speech.