ASR transcriptions and word-boundary timestamps as classification features raise balanced accuracy for dysarthria severity to 83.72 percent on a Korean dataset, beating waveform and deep-model baselines.
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Due to the subjective nature of current clinical evaluation, the need for automatic severity evaluation in dysarthric speech has emerged. DNN models outperform ML models but lack user-friendly explainability. ML models offer explainable results at a feature level, but their performance is comparatively lower. Current ML models extract various features from raw waveforms to predict severity. However, existing methods do not encompass all dysarthric features used in clinical evaluation. To address this gap, we propose a feature extraction method that minimizes information loss. We introduce an ASR transcription as a novel feature extraction source. We finetune the ASR model for dysarthric speech, then use this model to transcribe dysarthric speech and extract word segment boundary information. It enables capturing finer pronunciation and broader prosodic features. These features demonstrated an improved severity prediction performance to existing features: balanced accuracy of 83.72%.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
ASR transcriptions and word-boundary timestamps as classification features raise balanced accuracy for dysarthria severity to 83.72 percent on a Korean dataset, beating waveform and deep-model baselines.