REVIEW 4 major objections 5 minor 44 references
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper claims that speech-recognition-based features, built from transcriptions and word boundaries, classify dysarthria severity with higher balanced accuracy than waveform-based features or tested deep models on a Korean post-stroke…
desk verdict A useful, well-designed feature idea for dysarthria severity classification, but the headline 83.72% balanced accuracy is unverifiable from the paper as written because the train/test split may not be speaker-disjoint. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a two-stage pipeline whose carrier is a fine-tuned speech recognizer. Whisper is fine-tuned on dysarthric Korean speech with a special pause token added, and the resulting model outputs both a transcription and word segment boundaries obtained by aligning the audio to the inferred transcription. From these outputs the paper computes 'SR-features' in two families: Pronunciation Correctness, using syntactic edit metrics, semantic BERT score, and disfluency measures; and Structural Prosody, using pause-sequence edit distances, pause duration statistics, word-segment duration statistics, and rhythm measures such as speed change rate and syllables per second. These features carry the argument because they encode clinical description criteria that pure waveform features cannot express, such as whether a pause occurs at a linguistically expected location.
What would settle it
Run the same feature extraction and AutoML pipeline on a speaker-disjoint split of the same dataset, so that no patient contributes utterances to both training and test. If balanced accuracy drops to the waveform baseline or below, the method's advantage is an artifact of speaker leakage; if it stays near 83.72%, the ASR-feature claim survives.
Extended reading notes
Core claim
The paper's central discovery is that the mismatch between what a speaker intended to read and what a dysarthria-adapted recognizer transcribes is itself a diagnostic signal, and so is the timing of the word segments the recognizer produces. Using Whisper fine-tuned on dysarthric speech, the authors obtain transcriptions and word-level timestamps, then define two feature families: Pronunciation Correctness, covering ASR-style edit metrics, BERT score, repetition and filler-word similarity; and Structural Prosody, covering pause locations and durations, articulation durations, and rhythm. In the main experiment these features give balanced accuracy of 83.72% (accuracy 72.85%) with an AutoML classifier, versus 68.96% balanced accuracy for waveform features and 62.47% for the best DNN baseline. The confusion matrix shows the SR-feature model finds 18 of 18 severity-2 utterances, while Whisper+Linear finds 0 of 18 and waveform features find 8 of 18. The paper also reports that the method sacrifices some accuracy on the majority class (severity 1), which it attributes to remaining ASR error.
Load-bearing premise
The load-bearing premise is that the utterance-level 8:1:1 data split keeps training and test independent; if utterances from the same patient appear in both, the reported balanced accuracy could reflect speaker identity rather than severity.
Editorial extensions
If this is right
- If the central claim is right, explainable machine-learning classifiers can match or beat deep models on dysarthria severity grading, at least on balanced accuracy, while returning feature-level explanations clinicians can inspect.
- The most severe class—the one missed by the tested deep models—becomes the best-recalled class (18/18 in the test set), which matters for screening and therapy planning.
- The top-ranked features all come from Structural Prosody, suggesting that timing and rhythm carry the most severity signal among the proposed features.
- Because the features rely on ASR output rather than manual transcription, the approach could reduce the annotation burden of prosodic analysis if it transfers to other languages and datasets.
- Combining waveform and SR-features did not improve balanced accuracy, so the advantage is not simply a matter of feature count.
Reading between the lines
- The reported 8:1:1 split is at the utterance level, not the speaker level; if the same patient contributes utterances to both training and test, part of the balanced-accuracy gain could be speaker-specific rather than severity-specific. A speaker-disjoint split is the natural test of this.
- The feature families are language-agnostic in principle, but they depend on having a dysarthria-adapted recognizer for the target language; in languages without such a model, the method's advantage may shrink.
- The same transcription-alignment pipeline could be extended to predict continuous severity scores or to flag which words a patient mispronounced as direct therapy feedback, neither of which the paper evaluates.
- Because the authors note severity-1 accuracy falls, the method currently trades majority-class accuracy for minority-class recall; clinical use may require a decision threshold tuned to the cost of missing severe cases.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an automatic dysarthria severity classification method based on features extracted from a fine-tuned Whisper ASR model (DysarthricWhisper). The features fall into two families: Pronunciation Correctness (ASR error metrics, BERTScore, disfluency measures) and Structural Prosody (pause location/duration, articulation duration, rhythm). These are fed to an AutoGluon tabular classifier and evaluated on a Korean dysarthric speech dataset from AI-Hub, with a reported balanced accuracy of 83.72% for the proposed SR-features, compared with 68.96% for waveform-based features and 53.69–62.47% for the tested DNN baselines. The paper includes an ablation of the two feature families and an integration experiment to argue that the gain is not simply due to a larger feature count. Code is made publicly available.
Significance. If the evaluation is valid, the contribution is practically useful: the feature taxonomy maps naturally onto clinically used perceptual categories, the pipeline retains feature-level explainability, and the reported balanced accuracy is competitive with DNNs while improving recall for the severe class. The paper's strengths are its clearly specified feature definitions (Table 2), the inclusion of ablation and integration experiments that rule out feature-count as the sole explanation, and the stated intention to release code. The central claim is not circular: severity labels are not used to train the ASR feature extractor, and the pause-token method is external prior work. However, the empirical headline rests on a very small severe-class test set and an underspecified data-split procedure, so the key quantitative claim is not yet verifiable from the manuscript.
major comments (4)
- [Section 4.1, Table 5] The evaluation does not establish that the train and test sets are speaker-disjoint. Section 4.1 states only that the corpus is split 'in the ratio of 8:1:1, considering the ratio of each severity level,' but does not state whether the unit of random assignment is an utterance or a speaker. The test-set sizes in Table 5 (54 severity-0, 300 severity-1, 18 severity-2; 372 total) are also inconsistent with a strict 10% holdout of the 2,567 utterances in Table 1, so the actual split procedure is unclear. If speakers contribute multiple utterances and the same speaker appears in both training and test, the classifier and the fine-tuned Whisper model can exploit speaker-specific voice characteristics, and the reported 83.72% balanced accuracy would not reflect generalization to unseen patients. Please report the number of speakers, the assignment unit, and results under a speaker-exclusive split, or demonstrate that each speaker contributes exactly one utterance.
- [Equations (1)–(2), Section 4.2] The balanced-accuracy definition is given only for a binary classifier, but the task has three severity classes. The reported 83.72% appears to match macro-average recall across the three classes, but this is not stated, and Equations (1)–(2) do not define a multiclass balanced accuracy. Please correct the metric definition, state explicitly how the multiclass balanced accuracy was computed, and report per-class recall and precision alongside the aggregate number.
- [Section 4.2, Table 5] The main improvement is driven by perfect recall on the 18 severity-2 test utterances. The 95% confidence interval for 18/18 successes is approximately [81.5%, 100%], and the balanced-accuracy gap between SR-features (83.72%) and the waveform baseline (68.96%) is reported without any variance estimate or significance test. Please provide bootstrap confidence intervals or a permutation test, ideally clustered by speaker, to show that the difference is not attributable to the small size of the severe class.
- [Section 3.2] The reference duration sequence used for WS DTW and Abnormal Speed is described as 'the average duration sequence of 360 samples read by healthy individuals,' but the paper does not specify the provenance of these 360 samples. If they are drawn from the severity-0 utterances of the same corpus and if those speakers appear in the test split, the reference statistics are partly constructed from test data, which can inflate the reported performance for severity 0 and potentially for other classes. Please state where the reference samples come from, and if they are from the corpus, ensure they are disjoint from the test set or re-estimate the reference on training data only.
minor comments (5)
- [Section 4.4, Table 8] There are typos in the labels: 'Integerated-without-FS', 'Integarted-before-FS', and 'Integrated-wihtout-FS' should be 'Integrated-without-FS' and 'Integrated-before-FS'.
- [Table 5(a)] The entry '00.00%' in the last row of the severity-2 column should be '0.00%'.
- [Table 2, Section 3.1] The filler-word tokens used in Filler Words Similarity ([2], [Wm], [W], [kW]) are not defined; please describe the token set and how fillers were identified in the transcription.
- [Section 3.2, Table 2] The Top-30% short/long word-segment threshold is a free parameter, and no sensitivity analysis is reported; a short experiment varying this threshold would clarify how dependent the results are on this choice.
- [Section 4.1] The training details for the DNN baselines are incomplete (e.g., number of epochs, learning-rate schedule, whether Whisper/wav2vec2.0 were fine-tuned or used as frozen feature extractors); adding these details would improve reproducibility. The Discussion also acknowledges the severity-1 accuracy drop, but because severity 1 is the majority class, the paper should discuss whether the balanced-accuracy improvement corresponds to a clinically preferable operating point.
Circularity Check
No significant circularity: the proposed ASR-derived features feed a standard supervised classifier, and the headline balanced accuracy is an empirical result, not a rearrangement of fitted inputs.
full rationale
The paper's derivation chain is not circular. Severity labels are used only at the final classification stage; the ASR feature extractor is fine-tuned on the same corpus with the same split, but the test partition is held out from both ASR training and classifier training, as stated in Section 4.1: 'the samples in test sets have not been seen in training for both DysarthricWhisper and ML-classifier.' The SR-features are ASR evaluation metrics (WER, CER, BERTScore), pause-sequence comparisons, and duration/rhythm statistics computed from ASR word boundaries and a pause token; none of these are defined in terms of the severity labels. The reported balanced accuracy of 83.72% is an empirical value matching macro-averaged recall from the confusion matrix in Table 5(c), not a quantity forced by construction. The only overlapping-author citation, [25], supplies a pause-detection token method from prior ICASSP work; it is a reusable technical component, not an unverified premise that the conclusion reduces to. The acknowledged limitation that severity-1 accuracy degrades, and the reviewer's concern about an unreported speaker-independent split, are validity and generalization risks, not circularity under the definitions used here. Accordingly, no circular step is identified.
Assumptions & free parameters
free parameters (1)
- Top-30% short/long word-segment duration threshold =
30%
assumptions (5)
- domain assumption DysarthricWhisper transcribes dysarthric speech and predicts word boundaries accurately enough to reflect clinically meaningful pronunciation and prosody.
- domain assumption The hand-designed features based on Duffy's five subsystems capture the clinically relevant severity attributes.
- domain assumption Expert NIHSS severity labels on the Korean dataset are reliable enough to serve as ground truth.
- domain assumption No speaker appears in both the training and test splits, so the test estimate is unbiased.
- domain assumption Standard ASR metrics (WER, CER, BERTscore, DTW) are valid proxies for pronunciation and prosody deficits in dysarthria.
Cite this review
Pith. "Pith review of Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech." pith.science (2026). https://pith.science/paper/MB45AZSV
@misc{pith2026241203784,
author = {Pith},
title = {Pith review of: Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech},
year = {2026},
howpublished = {\url{https://pith.science/paper/MB45AZSV}},
note = {Machine review of arXiv:2412.03784}
}
read the original abstract
Due to the subjective nature of current clinical evaluation, the need for automatic severity evaluation in dysarthric speech has emerged. DNN models outperform ML models but lack user-friendly explainability. ML models offer explainable results at a feature level, but their performance is comparatively lower. Current ML models extract various features from raw waveforms to predict severity. However, existing methods do not encompass all dysarthric features used in clinical evaluation. To address this gap, we propose a feature extraction method that minimizes information loss. We introduce an ASR transcription as a novel feature extraction source. We finetune the ASR model for dysarthric speech, then use this model to transcribe dysarthric speech and extract word segment boundary information. It enables capturing finer pronunciation and broader prosodic features. These features demonstrated an improved severity prediction performance to existing features: balanced accuracy of 83.72%.
Figures
Reference graph
Works this paper leans on
-
[1]
INTRODUCTION Dysarthria, a speech-motor disorder leading to significant speech intelligibility loss [1], can result from factors like stroke, tumors, Parkinson’s Disease, and cerebral palsy. It impacts physical and psychological well-being, reducing the overall quality of life [2]. The primary cause of dysarthria is stroke, affecting around 50% of stroke ...
work page Pith review arXiv 2024
-
[2]
DATASET Table 1: Dataset distribution according to severity scale. Severity 0 1 2 Total Number of Utterances 431 1950 186 2567 We employ the proposed method to the Korean Dysarthric Speech dataset 2. It consists of recorded speech for sev- eral tasks used in clinical settings to evaluate the severity of dysarthria. Patients read the standard Korean paragr...
work page 1950
-
[3]
Ta- ble 2 describes features we extracted from ASR transcription
SPEECH RECOGNITION-BASED FEATURE EXTRACTION We categorize speech recognition-based features (SR-features) into Pronunciation Correctness and Structural Prosody. Ta- ble 2 describes features we extracted from ASR transcription. We use the OpenAI’s Whisper [24]. To accurately diag- nose severity, we need an ASR model that can transcribe the pronunciation er...
-
[4]
EXPERIMENTS 4.1. Experimental Setup We used the same corpus to train the ASR system and au- tomatic severity classification model. We split the Korean dysarthric speech dataset (Section 2) into the train, valida- tion, and test sets in the ratio of 8:1:1, considering the ratio of each severity level. The divided sets were maintained dur- ing ASR system tr...
-
[5]
DISCUSSION Our approach can offer feature-level analysis of the predicted severity, providing insight into the patient’s severity and speech therapy planning. For example, as in Figure 1, our method could reveal the impact on severity through feature importance analysis aligned with clinical indicators. More- over, our method can provide an easily compreh...
-
[6]
We employed an ML model, which has feature-level interpretability and improved its performance
CONCLUSION We propose the speech recognition-based feature extraction method for dysarthric speech’s automatic severity classifica- tion. We employed an ML model, which has feature-level interpretability and improved its performance. To this end, we quantified the clinically utilized dysarthric features. We fine-tuned the ASR model with dysarthric speech ...
-
[7]
ACKNOWLEDGMENT This work was supported by Institute of Information & com- munications Technology Planning & Evaluation (IITP) grant funded by the Korea government(MSIT) (No.2022-0-00621, Development of artificial intelligence technology that pro- vides dialog-based multi-modal explainability)
work page 2022
-
[8]
Acceptability and intelligibility of moderately dysarthric speech by four types of listeners,
Paul A Dagenais and Amy F Wilson, “Acceptability and intelligibility of moderately dysarthric speech by four types of listeners,” inInvestigations in clinical phonetics and linguistics, pp. 379–388. Psychology Press, 2012
work page 2012
Show all 44 references
-
[9]
Improving the intelligibility of dysarthric speech,
Alexander B. Kain, John-Paul Hosom, Xiaochuan Niu, Jan P.H. van Santen, Melanie Fried-Oken, and Janice Staehely, “Improving the intelligibility of dysarthric speech,” Speech Communication , vol. 49, no. 9, pp. 743–759, 2007
2007
-
[10]
Patients’ experiences of disruptions associated with post-stroke dysarthria,
Sylvia Dickson, Rosaline S. Barbour, Marian Brady, Alexander M. Clark, and Gillian Paton, “Patients’ experiences of disruptions associated with post-stroke dysarthria,” International Journal of Language & Com- munication Disorders, vol. 43, no. 2, pp. 135–153, 2008
2008
-
[11]
A feasibility randomized controlled trial of readyspeech for people with dysarthria after stroke,
Claire Mitchell, Audrey Bowen, Sarah Tyson, and Paul Conroy, “A feasibility randomized controlled trial of readyspeech for people with dysarthria after stroke,” Clinical Rehabilitation, vol. 32, no. 8, pp. 1037–1046, 2018
2018
-
[12]
Socioeconomic status and stroke inci- dence, prevalence, mortality, and worldwide burden: an ecological analysis from the global burden of disease study 2017,
Abolfazl Avan, Hadi Digaleh, Mario Di Napoli, Save- rio Stranges, Reza Behrouz, Golnaz Shojaeianbabaei, Amin Amiri, Reza Tabrizi, Naghmeh Mokhber, J David Spence, et al., “Socioeconomic status and stroke inci- dence, prevalence, mortality, and worldwide burden: an ecological a...
2017
-
[13]
Where the ear fits: A perceptual evaluation of motor speech disorders,
Robert Wertz and John Rosenbek, “Where the ear fits: A perceptual evaluation of motor speech disorders,” Semin. Speech Lang. , vol. 13, no. 01, pp. 39–54, Feb. 1992
1992
-
[14]
Implementing speech supplementation strategies: effects on intelligibility and speech rate of individuals with chronic severe dysarthria,
Katherine C Hustad, Tabitha Jones, and Suzanne Dai- ley, “Implementing speech supplementation strategies: effects on intelligibility and speech rate of individuals with chronic severe dysarthria,” J. Speech Lang. Hear. Res., vol. 46, no. 2, pp. 462–474, Apr. 2003
2003
-
[15]
Listener agreement for auditory-perceptual ratings of dysarthria,
Kate Bunton, Raymond D. Kent, Joseph R. Duffy, John C. Rosenbek, and Jane F. Kent, “Listener agreement for auditory-perceptual ratings of dysarthria,” Journal of Speech, Language, and Hearing Research , vol. 50, no. 6, pp. 1481–1495, 2007
2007
-
[16]
Classification of dysarthric speech according to the severity of impairment: an analysis of acoustic fea- tures,
Bassam Ali Al-Qatab and Mumtaz Begum Mustafa, “Classification of dysarthric speech according to the severity of impairment: an analysis of acoustic fea- tures,” IEEE Access, vol. 9, pp. 18183–18194, 2021
2021
-
[17]
Perceptual classification of motor speech disorders: The role of severity, speech task, and listener’s expertise,
Michaela Pernon, Fr ´ed´eric Assal, Ina Kodrasi, and Ma- rina Laganaro, “Perceptual classification of motor speech disorders: The role of severity, speech task, and listener’s expertise,” Journal of Speech, Language, and Hearing Research, vol. 65, no. 8, pp. 2727–2747, 2022
2022
-
[18]
Reliability of auditory- perceptual scaling of dysarthria,
Joslin Zeplin and Ray D Kent, “Reliability of auditory- perceptual scaling of dysarthria,” Disorders of motor speech: Assessment, treatment, and clinical characteri- zation, pp. 145–154, 1996
1996
-
[19]
Zhengjun Yue, Erfan Loweimi, and Zoran Cvetkovic, Dysarthric Speech Recognition, Detection and Clas- sification using Raw Phase and Magnitude Spectra , ISCA-INST SPEECH COMMUNICATION ASSOC, May 2023
2023
-
[20]
Automatic severity classification of dysarthric speech by using self-supervised model with multi-task learning,
Eun Jung Yeo, Kwanghee Choi, Sunhee Kim, and Minhwa Chung, “Automatic severity classification of dysarthric speech by using self-supervised model with multi-task learning,” in ICASSP 2023 - 2023 IEEE In- ternational Conference on Acoustics, Speech and Signal Processing (ICASSP...
2023
-
[21]
Adversarial- Free Speaker Identity-Invariant Representation Learn- ing for Automatic Dysarthric Speech Classification,
Parvaneh Janbakhshi and Ina Kodrasi, “Adversarial- Free Speaker Identity-Invariant Representation Learn- ing for Automatic Dysarthric Speech Classification,” in Proc. Interspeech 2022, 2022, pp. 2138–2142
2022
-
[22]
Inspired by these attributes, we designed two main categories of fea- tures: Pronunciation Correctness, and Structural Prosody
categorized dysarthric speech features into five subsys- tems (respiration, phonation, articulation, resonance, and prosody), outlining specific attributes for each. Inspired by these attributes, we designed two main categories of fea- tures: Pronunciation Correctness, and Str...
-
[23]
Grad-cam: Visual explanations from deep net- works via gradient-based localization,
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra, “Grad-cam: Visual explanations from deep net- works via gradient-based localization,” in 2017 IEEE International Conference on Computer Vision (ICCV) , 2017, pp. 618–626
2017
-
[24]
Dysarthric Speech Classification Using Glottal Features Computed from Non-words, Words and Sentences,
Narendra N P and Paavo Alku, “Dysarthric Speech Classification Using Glottal Features Computed from Non-words, Words and Sentences,” inProc. Interspeech 2018, 2018, pp. 3403–3407
2018
-
[25]
Phonological features in discrimina- tive classification of dysarthric speech,
Frank Rudzicz, “Phonological features in discrimina- tive classification of dysarthric speech,” in 2009 IEEE International Conference on Acoustics, Speech and Sig- nal Processing, 2009, pp. 4605–4608
2009
-
[26]
Neurospeech: An open-source software for parkinson’s speech analysis,
Juan Rafael Orozco-Arroyave, Juan Camilo V ´asquez- Correa, Jes ´us Francisco Vargas-Bonilla, R. Arora, N. Dehak, P.S. Nidadavolu, H. Christensen, F. Rudz- icz, M. Yancheva, H. Chinaei, A. Vann, N. V ogler, T. Bocklet, M. Cernak, J. Hannink, and Elmar N ¨oth, “Neurospeech: An ...
2018
-
[27]
On using the ua-speech and torgo databases to validate automatic dysarthric speech classification approaches,
Guilherme Schu, Parvaneh Janbakhshi, and Ina Kodrasi, “On using the ua-speech and torgo databases to validate automatic dysarthric speech classification approaches,” in ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023, pp. 1–5
2023
-
[28]
Speech Intelligibility of Dysarthric Speech: Human Scores and Acoustic-Phonetic Features,
Wei Xue, Roeland van Hout, Fleur Boogmans, Mario Ganzeboom, Catia Cucchiarini, and Helmer Strik, “Speech Intelligibility of Dysarthric Speech: Human Scores and Acoustic-Phonetic Features,” in Proc. In- terspeech 2021, 2021, pp. 2911–2915
2021
-
[29]
Praat, a system for doing phonetics by computer,
Paul Boersma, “Praat, a system for doing phonetics by computer,” Glot. Int., vol. 5, no. 9, pp. 341–345, 2001
2001
-
[30]
Duffy, Motor speech disorders: Substrates, differential diagnosis, and management, Elsevier, 2020
Joseph R. Duffy, Motor speech disorders: Substrates, differential diagnosis, and management, Elsevier, 2020
2020
-
[31]
Dysarthria evaluation,
HyangHee Kim, “Dysarthria evaluation,” Communica- tion Sciences & Disorders, pp. 23–28, 2005
2005
-
[32]
Robust speech recognition via large-scale weak supervision,
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brock- man, Christine McLeavey, and Ilya Sutskever, “Robust speech recognition via large-scale weak supervision,” arXiv preprint arXiv:2212.04356, 2022
2022 arXiv
-
[33]
Inappropriate pause detection in dysarthric speech using large-scale speech recognition,
Jeehyun Lee, Yerin Choi, Tae-Jin Song, and Myoung- Wan Koo, “Inappropriate pause detection in dysarthric speech using large-scale speech recognition,” inICASSP 2024 - 2024 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP), 2024, pp. 12486–12490
2024
-
[34]
Bertscore: Evaluating text generation with BERT,
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi, “Bertscore: Evaluating text generation with BERT,” in8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. 2020, OpenReview.net
2020
-
[35]
BERT: pre-training of deep bidi- rectional transformers for language understanding,
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, “BERT: pre-training of deep bidi- rectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North Amer- ican Chapter of the Association for Computational Lin- guistics: Hum...
2019
-
[36]
As- sessing ASR model quality on disordered speech using bertscore,
Jimmy Tobin, Qisheng Li, Subhashini Venugopalan, Katie Seaver, Richard Cave, and Katrin Tomanek, “As- sessing ASR model quality on disordered speech using bertscore,” CoRR, vol. abs/2209.10591, 2022
2022 arXiv
-
[37]
KLUE: korean language understanding evaluation,
Sungjoon Park, Jihyung Moon, Sungdong Kim, Won-Ik Cho, Jiyoon Han, Jangwon Park, Chisung Song, Jun- seong Kim, Youngsook Song, Tae Hwan Oh, Joohong Lee, Juhyun Oh, Sungwon Lyu, Younghoon Jeong, Inkwon Lee, Sangwoo Seo, Dongjun Lee, Hyunwoo Kim, Myeonghwa Lee, Seongbo Jang, Seu...
2021
-
[38]
Ef- fect of speech rate for sentences on speech intelligibil- ity,
Aihong Du, Chundan Lin, and Jingjing Wang, “Ef- fect of speech rate for sentences on speech intelligibil- ity,” 2014 IEEE International Conference on Commu- nication Problem-Solving, ICCP 2014, pp. 233–236, 03 2015
2014
-
[39]
Autogluon-tabular: Robust and accurate automl for structured data,
Nick Erickson, Jonas Mueller, Alexander Shirkov, Hang Zhang, Pedro Larroy, Mu Li, and Alexander Smola, “Autogluon-tabular: Robust and accurate automl for structured data,” arXiv preprint arXiv:2003.06505 , 2020
2003 arXiv
-
[40]
Decoupled weight decay regularization,
Ilya Loshchilov and Frank Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations, 2017
2017
-
[41]
Cross-lingual dysarthria severity classi- fication for english, korean, and tamil,
Eun Jung Yeo, Kwanghee Choi, Sunhee Kim, and Min- hwa Chung, “Cross-lingual dysarthria severity classi- fication for english, korean, and tamil,” in 2022 Asia- Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) . IEEE, 2022, pp. 566–574
2022
-
[42]
In- troducing Parselmouth: A Python interface to Praat,
Yannick Jadoul, Bill Thompson, and Bart de Boer, “In- troducing Parselmouth: A Python interface to Praat,” Journal of Phonetics, vol. 71, pp. 1–15, 2018
2018
-
[43]
wav2vec 2.0: A framework for self-supervised learning of speech representations,
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Ad- vances in neural information processing systems , vol. 33, pp. 12449–12460, 2020
2020
-
[44]
Whisper Features for Dysarthric Severity-Level Classification,
Siddharth Rathod, Monil Charola, Akshat V ora, Yash Jogi, and Hemant A. Patil, “Whisper Features for Dysarthric Severity-Level Classification,” in Proc. IN- TERSPEECH 2023, 2023, pp. 1523–1527
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.