A comparative study of dysfluency detection models finds UDM best balances accuracy and clinical interpretability, while SSDM is not reproducible from its published description.
LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Phonetic speech transcription is crucial for fine-grained linguistic analysis and downstream speech applications. While Connectionist Temporal Classification (CTC) is a widely used approach for such tasks due to its efficiency, it often falls short in recognition performance, especially under unclear and nonfluent speech. In this work, we propose LCS-CTC, a two-stage framework for phoneme-level speech recognition that combines a similarity-aware local alignment algorithm with a constrained CTC training objective. By predicting fine-grained frame-phoneme cost matrices and applying a modified Longest Common Subsequence (LCS) algorithm, our method identifies high-confidence alignment zones which are used to constrain the CTC decoding path space, thereby reducing overfitting and improving generalization ability, which enables both robust recognition and text-free forced alignment. Experiments on both LibriSpeech and PPA demonstrate that LCS-CTC consistently outperforms vanilla CTC baselines, suggesting its potential to unify phoneme modeling across fluent and non-fluent speech.
citation-role summary
citation-polarity summary
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
A Comparative Study of Controllability, Explainability, and Performance in Dysfluency Detection Models
A comparative study of dysfluency detection models finds UDM best balances accuracy and clinical interpretability, while SSDM is not reproducible from its published description.