Pith. sign in

DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Dysarthric speech recognition often suffers from performance degradation due to the intrinsic diversity of dysarthric severity and extrinsic disparity from normal speech. To bridge these gaps, we propose a Dynamic Phoneme-level Contrastive Learning (DyPCL) method, which leads to obtaining invariant representations across diverse speakers. We decompose the speech utterance into phoneme segments for phoneme-level contrastive learning, leveraging dynamic connectionist temporal classification alignment. Unlike prior studies focusing on utterance-level embeddings, our granular learning allows discrimination of subtle parts of speech. In addition, we introduce dynamic curriculum learning, which progressively transitions from easy negative samples to difficult-to-distinguishable negative samples based on phonetic similarity of phoneme. Our approach to training by difficulty levels alleviates the inherent variability of speakers, better identifying challenging speeches. Evaluated on the UASpeech dataset, DyPCL outperforms baseline models, achieving an average 22.10\% relative reduction in word error rate (WER) across the overall dysarthria group.

citation-role summary

background 1

citation-polarity summary

fields

eess.AS 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Towards Temporally Explainable Dysarthric Speech Clarity Assessment

eess.AS · 2025-05-31 · conditional · novelty 6.0

A therapist-annotated dysarthric speech dataset and a three-stage ASR framework show that Whisper-large localizes mispronunciations precisely, with substitution errors detected best and 70.1% of ASR error descriptions matching therapist labels.

citing papers explorer

Showing 1 of 1 citing paper.

  • Towards Temporally Explainable Dysarthric Speech Clarity Assessment eess.AS · 2025-05-31 · conditional · none · ref 9 · internal anchor

    A therapist-annotated dysarthric speech dataset and a three-stage ASR framework show that Whisper-large localizes mispronunciations precisely, with substitution errors detected best and 70.1% of ASR error descriptions matching therapist labels.