REVIEW 4 major objections 5 minor 1 cited by
Improving ASR Fairness for Cleft Lip and Palate Speech: A Study on Severity-Aware Data Mixing
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Severity-aware mixing of cleft lip and palate (CLP) speech with normal speech during ASR training improves accuracy and fairness, cutting pooled word error rate from 22.64% to 18.76% on Kannada and from 28.45% to 18.89% on English child…
desk verdict A first fairness evaluation for CLP speech with a sensible augmentation idea, but the evidence is confounded by speaker overlap and uncontrolled training set size. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the severity ordering of CLP speech, measured by dynamic time warping distance between voiced spectrogram frames of normal and CLP utterances, together with the two-group fairness score. The dynamic time warping distance supplies a graded distortion axis from mild to moderate to severe, motivating which utterances to add during training. The fairness score, defined as $FS = -\alpha \cdot \text{Average Error Rate} - \beta \cdot \text{Error Disparity}$, collapses the normal-versus-CLP word error rates into a single number, with values closer to zero meaning fairer. The mixing strategy trains a model on normal speech plus a severity subset of CLP speech and evaluates on the full test set, allowing the choice of severity composition to be tuned for a given model and language.
What would settle it
Re-run the best severity-mixed training condition with a speaker-disjoint split, where all recordings of each speaker are assigned wholly to either training or testing, on both corpora, and compare the pooled word error rate and fairness score against the paper's utterance-level split. If the improvement disappears or falls below a meaningful margin, the reported fairness gain is speaker memorization rather than generalization to unseen CLP speakers.
Extended reading notes
Core claim
The central claim is that ASR fairness for cleft lip and palate speech can be improved by severity-aware data mixing: rather than training only on normal speech or only on CLP speech, the authors train on normal speech augmented with CLP utterances in order of increasing severity. They introduce a fairness score that trades off the average error rate against the error disparity between normal and CLP groups, with values closer to zero indicating a fairer system. Spectrogram, formant-contour, and dynamic time warping diagnostics show that spectral distortion grows progressively from mild to moderate to severe CLP speech, which justifies mixing by severity groups. In their experiments, adding mild, moderate, and severe CLP utterances to the normal training data improves pooled word error rate and moves the fairness score closer to zero on both corpora; the best configuration for GMM-HMM on the Kannada corpus includes severe speech, while for Whisper on the English corpus the best configuration stops at moderate. The reported fairness-score improvements are 17.89% on the Kannada corpus and 47.50% on the English corpus.
Load-bearing premise
The evaluation uses a random 80/20 split of individual recordings, so the same speaker can appear in both the training and testing sets; if that overlap drives the results, the reported gains may come from remembering a speaker's voice rather than from learning to recognize cleft-palate speech in general.
Editorial extensions
If this is right
- On both Kannada child speech and English child speech, replacing a normal-only training set with a severity-mixed set reduces the gap between normal and CLP word error rates.
- The best mixing recipe is not universal: for GMM-HMM on the Kannada corpus it includes severe CLP speech, while for Whisper on the English corpus adding severe utterances hurts, so the recipe must be chosen per model and language.
- The fairness score supports deployment choices: setting $\alpha=0.9$, $\beta=0.1$ prioritizes overall accuracy, while $\alpha=0.1$, $\beta=0.9$ prioritizes reducing disparity between normal and CLP groups.
- Training with CLP data barely degrades performance on normal speech, so the fairness improvement is not bought by sacrificing typical speech recognition.
- Severity-aware mixing improves fairness for a conventional GMM-HMM model, a self-supervised foundation model, and a supervised foundation model, suggesting the mechanism is not tied to one architecture.
Reading between the lines
- If the effect is driven by spectral overlap rather than speaker identity, the same severity-graded mixing protocol may transfer to other graded speech disorders, such as dysarthria or stuttering, where distortion increases with severity.
- Because the reported experiments split recordings randomly rather than by speaker, a speaker-disjoint evaluation is the direct test of whether the gain is acoustic generalization or speaker memorization; that check is not reported.
- A natural extension is to condition the model on the severity label itself, for instance through an auxiliary loss or an embedding, which could combine the benefit of mixing with explicit severity awareness.
- The two-group fairness score could be extended to multi-severity fairness by summing pairwise disparities or using the maximum group error, rather than collapsing all CLP speakers into one group.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses automatic speech recognition (ASR) fairness for cleft lip and palate (CLP) speech. It introduces a fairness score (FS) that balances average word error rate and error disparity between normal and CLP groups, evaluates the fairness of a public Google ASR API, and proposes severity-aware mixing of CLP and normal speech during ASR training. Experiments are conducted with GMM-HMM, Whisper, and XLS-R on two datasets: AIISH (Kannada) and NMCPC (English). The authors report that severity-aware augmentation improves FS and reduces WER; for example, the full-text abstract reports WER decreasing from 22.64% to 18.76% (GMM-HMM, AIISH) and from 28.45% to 18.89% (Whisper, NMCPC).
Significance. If the reported effect is robust, the paper would be a useful empirical contribution to ASR fairness for disordered speech: it provides a simple fairness metric, a concrete data-mixing recipe, and evaluations across a public API, a traditional HMM system, and two foundation models in two languages. The paper also documents a striking performance gap in a public ASR service for CLP speech. However, the central claim depends on unaddressed confounds and internal inconsistencies. The utterance-level split, lack of size-matched controls, and post hoc selection of the best augmentation condition make the current evidence suggestive rather than demonstrative; the reported headline numbers are also internally inconsistent.
major comments (4)
- [Section 3 (Database setup)] The 80/20 split is utterance-level, not speaker-disjoint. With only 60 speakers in AIISH and 65 in NMCPC, the same speaker very likely contributes utterances to both training and evaluation. CLP speech carries speaker-specific hypernasality and articulation patterns, so the reported WER gains may reflect speaker-specific memorization rather than generalization to unseen CLP speakers. A speaker-disjoint split, and ideally speaker-independent evaluation, is needed to support the claim that severity-aware mixing improves fairness for new CLP speakers.
- [Tables 7 and 8] The augmentation conditions are not matched in training-set size or total CLP content. 'Mild+Normal', 'Mild+Moderate+Normal', and 'Mild+Moderate+Severe+Normal' add monotonically more CLP utterances, and there is no control adding an equivalent amount of normal-only data or an equivalent amount of severity-blind CLP data. Moreover, the 'CLP' training condition alone improves FS substantially (AIISH GMM-HMM FS rises from -31.57 with Normal training to -26.67 with CLP training), so the reported gains are not specifically attributable to severity-aware selection.
- [Abstract and Tables 5/7] The headline numbers are internally inconsistent. The metadata abstract states WER decreases from 37.58% to 25.47% (GMM-HMM, AIISH) and from 35.74% to 21.72% (Whisper, NMCPC), while the full-text abstract states 22.64% to 18.76% and 28.45% to 18.89%, respectively. In addition, Table 5 lists the GMM-HMM AIISH CLP-trained CLP WER as 36.44, whereas Table 7 lists the same quantity as 11.00; the Whisper values also differ (55.66 vs. 44.77). These discrepancies undermine the reliability of the reported improvements and must be reconciled.
- [Section 6.2, Tables 7-9] The evaluation lacks error bars, confidence intervals, and significance tests. All WER and FS comparisons are point estimates from a single random split, and the best augmentation recipe is chosen post hoc per dataset (Mild+Moderate+Severe+Normal for AIISH, Mild+Moderate+Normal for NMCPC). Without multiple seeds, bootstrap confidence intervals, or a pre-specified selection rule, the paper cannot rule out chance variation or selection effects.
minor comments (5)
- [Section 6.2.1] The phrase 'crisis cross' should be 'criss-cross'.
- [Table 3] The header 'datsets' should be 'datasets'.
- [Reference list] Reference [49] contains an encoding artifact ('children²s'); references [55] and [56] are identical and one should be removed.
- [Table 7] The Whisper AIISH CLP-trained row reports '7' for the normal test column without a decimal place; clarify whether this is 7.00 or another value.
- [Section 4.2] The FS definition states 'α,β> = 0' but the experiments later use α+β=1; clarify the constraint on the weights.
Circularity Check
No significant circularity: the central results are empirical evaluations on held-out test partitions, the fairness score is a definition rather than a derived prediction, and self-citations are contextual prior work, not load-bearing premises.
full rationale
The paper's derivation chain is empirical rather than circular. The central claims—that public ASR systems are less fair on CLP speech, and that severity-aware mixing of CLP and normal speech improves WER and fairness scores—are supported by direct measurements on held-out evaluation sets reported in Tables 3 through 9. The fairness score FS = −α·AverageErrorRate − β·ErrorDisparity is explicitly introduced as a metric definition; it does not by construction force any particular WER outcome, and the improvement claims are evaluated independently of how FS is weighted. The DTW proximity analysis is an independent acoustic-motivation step: it supports the hypothesis that mild and moderate CLP speech retains spectral alignment with normal speech, but the DTW distances are not fitted parameters used to compute the reported WERs. Self-citations to the authors' earlier CLP classification and enhancement work (e.g., [10], [15], [37]) are contextual literature references and are not invoked as uniqueness theorems or as premises that force the empirical results. Possible weaknesses such as the utterance-level 80/20 split potentially allowing speaker overlap, the lack of size-matched controls, and internal inconsistencies between Table 5 and Table 7 are experimental-validity and reporting concerns, not circularity under the specified patterns. There is no step where a prediction reduces by construction to its input or where a fitted parameter is renamed as a prediction.
Assumptions & free parameters
free parameters (4)
- alpha (FS weight on average error) =
0.5 (also 0.1, 0.9)
- beta (FS weight on error disparity) =
0.5 (also 0.9, 0.1)
- VAD energy threshold =
6% of average energy
- GMM-HMM senones and Gaussians =
50 senones, 500 Gaussians
assumptions (5)
- domain assumption The 80/20 random utterance split yields independent training and evaluation sets.
- domain assumption CLP severity labels (mild/moderate/severe) in AIISH and NMCPC are accurate and consistent.
- domain assumption WER is an appropriate and sufficient proxy for ASR fairness between normal and CLP groups.
- standard math Dynamic time warping distance between voiced spectrogram frames is a valid measure of spectral distortion.
- ad hoc to paper The linear combination of average error and error disparity captures fairness.
Cite this review
Pith. "Pith review of Improving ASR Fairness for Cleft Lip and Palate Speech: A Study on Severity-Aware Data Mixing." pith.science (2026). https://pith.science/paper/QDNMIG3M
@misc{pith2026250503697,
author = {Pith},
title = {Pith review of: Improving ASR Fairness for Cleft Lip and Palate Speech: A Study on Severity-Aware Data Mixing},
year = {2026},
howpublished = {\url{https://pith.science/paper/QDNMIG3M}},
note = {Machine review of arXiv:2505.03697}
}
read the original abstract
Speech produced by individuals with cleft lip and palate (CLP) is often hypernasal (and sometimes breathy) due to structural anomalies, yielding shifts in formant structure that degrade automatic speech recognition (ASR) performance and fairness. Building on evidence that mainstream ASR systems underperform on atypical and disordered speech, we posit that widely used services (e.g., Google Speech-to-Text) exhibit reduced fairness for CLP speech, and we evaluate this claim empirically. To quantify fairness consistently, we introduce a simple fairness score (FS) that trades off overall error and between-group disparity. Despite formant disruptions, mild and moderate CLP speech retains partial spectro-temporal alignment with typical speech, motivating the use of mixing strategies to improve fairness. We systematically investigated severity-aware mixing of CLP and normal speech at different severity levels and assessed its effect on fairness. Three ASR models GMM-HMM, Whisper, and XLSR were evaluated on the AIISH (Kannada language) and NMCPC (English language) datasets. A mixing strategy that leverages severity-aware mixing of CLP and normal speech improves fairness on both English (NMCPC) and Kannada (AIISH) corpora. Notably, the word error rate (WER) decreased from 37.58% to 25.47% (GMM-HMM, AIISH) and from 35.74% to 21.72% (Whisper, NMCPC).
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Normal-Anchored First-Order Model-Agnostic Meta-Learning based Whisper Fine-Tuning for Enhancing Fairness of Cleft Lip and Palate Speech Recognition
Normal-anchored FOMAML fine-tuning of Whisper lowers word error rates on cleft lip and palate speech in two datasets, but the key comparison to conventional fine-tuning uses different training data.
Reference graph
Works this paper leans on
-
[1]
D. J. Zajac, L. Vallino, Evaluation and Management of Cleft Lip and Palate: A Developmental Perspective, 2017. doi:10.1109/ICME.2016. 7552917
-
[2]
A. Lohmander, M. Olsson, Methodology for perceptual assessment of speech in patients with cleft palate: A critical review of the literature, The Cleft Palate-Craniofacial Journal 41 (2004) 64 – 70
work page 2004
-
[3]
J. Stengelhofen, Cleft palate: The nature and remediation of communication problems, Churchill Livingstone (1993)
work page 1993
- [4]
-
[5]
K. M. V . Lierde, S. Claeys, M. D. Bodt, P. V . Cauwenberge, V ocal quality characteristics in children with cleft palate: a multiparameter approach., Journal of voice : o fficial journal of the V oice Foundation 18 3 (2004) 354–62
work page 2004
-
[6]
D. J. Zajac, C. Plante, A. Lloyd, K. L. Haley, Reliability and validity of a computer-mediated, single-word intelligibility test: Preliminary findings for children with repaired cleft lip and palate, The Cleft Palate-Craniofacial Journal 48 (2011) 538 – 549
work page 2011
-
[7]
T. L. Whitehill, C. H. F. Chau, Single-word intelligibility in speakers with repaired cleft palate, Clinical Linguistics & Phonetics 18 (2004) 341 – 355
work page 2004
-
[8]
A. W. Kummer, Cleft Palate and Craniofacial Anomalies: E ffects on Speech and Resonance, 2007
work page 2007
Show all 58 references
-
[9]
Baumann, D
I. Baumann, D. Wagner, F. Braun, S. P. Bayerl, E. Noth, K. Riedhammer, T. Bocklet, Influence of utterance and speaker characteristics on the classification of children with cleft lip and palate, INTERSPEECH 2023 (2022)
2022
-
[10]
Bhattacharjee, H
S. Bhattacharjee, H. S. Shekhawat, S. R. M. Prasanna, Classification of cleft lip and palate speech using fine-tuned transformer pretrained models, in: B. J. Choi, D. Singh, U. S. Tiwary, W.-Y . Chung (Eds.), Intelligent Human Computer Interaction, Springer Nature Switzerland,...
2024
-
[11]
Kalita, G
S. Kalita, G. K S, P. Mariswamy, S. Prasanna, S. Dandapat, Objective assessment of cleft lip and palate speech intelligibility using articulation and hypernasality measures, The Journal of the Acoustical Society of America 146 (2019) 1164–1175. doi: 10.1121/1.5121310
2019 doi
-
[12]
Kalita, S
S. Kalita, S. R. M. Prasanna, S. Dandapat, Importance of glottis landmarks for the assessment of cleft lip and palate speech intelligibility., The Journal of the Acoustical Society of America 144 5 (2018) 2656
2018
-
[13]
Kalita, S
S. Kalita, S. R. M. Prasanna, S. Dandapat, Intelligibility assessment of cleft lip and palate speech using gaussian posteriograms based on joint spectro-temporal features., The Journal of the Acoustical Society of America 144 4 (2018) 2413
2018
-
[14]
C. M. Vikram, S. Macha, S. Kalita, S. R. M. Prasanna, Acoustic analysis of misarticulated trills in cleft lip and palate children., The Journal of the Acoustical Society of America 143 6 (2018) EL474
2018
-
[15]
Bhattacharjee, R
S. Bhattacharjee, R. Sinha, Sensitivity analysis of maskcyclegan based voice conversion for enhancing cleft lip and palate speech recognition, 2022, pp. 1–5. doi:10.1109/SPCOM55316.2022.9840769
2022
-
[16]
X. Wang, S. Yang, M. Tang, H. Yin, H. Huang, L. He, Hypernasalitynet: Deep recurrent neural network for automatic hypernasality detection, International Journal of Medical Informatics 129 (2019) 1–12. doi:https://doi.org/10.1016/j.ijmedinf.2019.05.023
2019 doi
-
[17]
J. H. Ha, H. Lee, S. M. Kwon, H. Joo, G. Lin, D. Y . Kim, S. Kim, J. Y . Hwang, J. H. Chung, H. J. Kong, Deep learning–based diagnostic system for velopharyngeal insu fficiency based on videofluoroscopy in patients with repaired cleft palates, Journal of Craniofacial Surgery 1...
2023
-
[18]
A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y . Saraf, J. M. Pino, A. Baevski, A. Conneau, M. Auli, Xls-r: Self-supervised cross-lingual speech representation learning at scale, ArXiv abs /2111.09296 (2021)
2021 arXiv
-
[19]
Radford, J
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, I. Sutskever, Robust speech recognition via large-scale weak supervision, in: Proceedings of the 40th International Conference on Machine Learning, ICML’23, JMLR.org, 2023
2023
-
[20]
A. K. Dubey, S. R. M. Prasanna, S. Dandapat, Zero time windowing based severity analysis of hypernasal speech, 2016 IEEE Region 10 Conference (TENCON) (2016) 970–974
2016
-
[21]
A. K. Dubey, S. R. M. Prasanna, S. Dandapat, Hypernasality detection using zero time windowing, 2018 International Conference on Signal Processing and Communications (SPCOM) (2018) 105–109
2018
-
[22]
Nikitha, S
K. Nikitha, S. Kalita, M. VikramC., M. Pushpavathi, S. R. M. Prasanna, Hypernasality severity analysis in cleft lip and palate speech using vowel space area, in: Interspeech, 2017
2017
-
[23]
VikramC., A
M. VikramC., A. Tripathi, S. Kalita, S. R. M. Prasanna, Estimation of hypernasality scores from cleft lip and palate speech, in: Interspeech, 2018
2018
-
[24]
V . C. Mathad, N. J. Scherer, K. Chapman, J. M. Liss, V . Berisha, An attention model for hypernasality prediction in children with cleft palate, ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2021) 7248–7252
2021
-
[25]
A. K. Dubey, S. R. M. Prasanna, S. Dandapat, Pitch-adaptive front-end feature for hypernasality detection, in: Interspeech, 2018
2018
-
[26]
A. K. Dubey, S. R. M. Prasanna, S. Dandapat, Hypernasality severity detection using constant q cepstral coe fficients, in: Interspeech, 2019. 16
2019
-
[27]
A. K. Dubey, S. R. M. Prasanna, S. Dandapat, Sinusoidal model-based hypernasality detection in cleft palate speech using cvcv sequence, Speech Commun. 124 (2020) 1–12
2020
-
[28]
A. K. Dubey, S. R. M. Prasanna, S. Dandapat, Detection and assessment of hypernasality in repaired cleft palate speech using vocal tract and residual features., The Journal of the Acoustical Society of America 146 6 (2019) 4211
2019
-
[29]
Kalita, S
S. Kalita, S. R. M. Prasanna, S. Dandapat, Self-similarity matrix based intelligibility assessment of cleft lip and palate speech, in: Interspeech, 2018
2018
-
[30]
Kalita, K
S. Kalita, K. S. Girish, P. M., S. R. M. Prasanna, S. Dandapat, Objective assessment of cleft lip and palate speech intelligibility using articulation and hypernasality measures., The Journal of the Acoustical Society of America 146 2 (2019) 1164
2019
-
[31]
Kalita, P
S. Kalita, P. N. Sudro, S. R. M. Prasanna, S. Dandapat, Nasal air emission in sibilant fricatives of cleft lip and palate speech, in: Interspeech, 2019
2019
-
[32]
C. M. Vikram, N. Adiga, S. R. M. Prasanna, Detection of nasalized voiced stops in cleft palate speech using epoch-synchronous features, IEEE/ACM Transactions on Audio, Speech, and Language Processing 27 (2019) 1189–1200
2019
-
[33]
P. N. Sudro, S. Kalita, S. R. M. Prasanna, Processing transition regions of glottal stop substituted /s/ for intelligibility enhancement of cleft palate speech, in: Interspeech, 2018
2018
-
[34]
P. N. Sudro, S. M. Prasanna, Enhancement of cleft palate speech using temporal and spectral processing, Speech Communication 123 (2020) 70–82
2020
-
[35]
P. N. Sudro, S. M. Prasanna, Modification of misarticulated fricative /s/in cleft lip and palate speech, Biomedical Signal Processing and Control 67 (2021) 102088
2021
-
[36]
P. N. Sudro, C. M. Vikram, S. R. M. Prasanna, Event-based transformation of misarticulated stops in cleft lip and palate speech, Circuits, Systems, and Signal Processing 40 (2021) 4064 – 4088
2021
-
[37]
P. N. Sudro, R. K. Das, R. Sinha, S. R. M. Prasanna, Enhancing the intelligibility of cleft lip and palate speech using cycle-consistent adversarial networks, 2021 IEEE Spoken Language Technology Workshop (SLT) (2021) 720–727
2021
-
[38]
P. N. Sudro, R. K. Das, R. Sinha, S. R. Mahadeva Prasanna, Significance of data augmentation for improving cleft lip and palate speech recognition, in: 2021 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2021, pp. 484–490
2021
-
[39]
Baumann, D
I. Baumann, D. Wagner, F. Braun, S. P. Bayerl, E. N ¨oth, K. Riedhammer, T. Bocklet, Influence of utterance and speaker characteristics on the classification of children with cleft lip and palate, in: Interspeech, 2022
2022
-
[40]
K. Song, T. Wan, B. Wang, H. Jiang, L. K. Qiu, J. Xu, L. ping Jiang, Q. Lou, Y . Yang, D. Li, X. Wang, L. Qiu, Improving hypernasality estimation with automatic speech recognition in cleft palate speech, in: Interspeech, 2022
2022
-
[41]
VikramC., S
M. VikramC., S. R. M. Prasanna, A. K. Abraham, M. Pushpavathi, S. GirishK., Detection of glottal activity errors in production of stop consonants in children with cleft lip and palate, in: Interspeech, 2018
2018
-
[42]
V . C. Mathad, S. R. M. Prasanna, V owel onset point based screening of misarticulated stops in cleft lip and palate speech, IEEE /ACM Transactions on Audio, Speech, and Language Processing 28 (2020) 450–460
2020
-
[43]
V . C. Mathad, N. J. Scherer, K. Chapman, J. M. Liss, V . Berisha, A deep learning algorithm for objective assessment of hypernasality in children with cleft palate, IEEE Transactions on Biomedical Engineering 68 (2020) 2986–2996
2020
-
[44]
Fujiwara, R
K. Fujiwara, R. Takashima, C. Sugiyama, N. Tanaka, K. Nohara, K. Nozaki, T. Takiguchi, Data augmentation based on frequency warping for recognition of cleft palate speech, in: 2021 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ...
2021
-
[45]
M. H. Javid, K. Gurugubelli, A. K. Vuppala, Single frequency filter bank based long-term average spectra for hypernasality detection and assessment in cleft lip and palate speech, in: ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (...
2020
-
[46]
Bocklet, A
T. Bocklet, A. K. Maier, K. Riedhammer, U. Eysholdt, E. N ¨oth, Erlangen-clp: A large annotated corpus of speech from children with cleft lip and palate, in: International Conference on Language Resources and Evaluation, 2014
2014
-
[47]
Eshky, M
A. Eshky, M. S. Ribeiro, J. Cleland, K. Richmond, Z. Roxburgh, J. M. Scobbie, A. A. Wrench, Ultrasuite: A repository of ultrasound and acoustic data from child speech therapy sessions, ArXiv abs/1907.00835 (2018)
2018 arXiv
-
[48]
Nikitha, S
K. Nikitha, S. Kalita, C. Vikram, M. Pushpavathi, S. M. Prasanna, Hypernasality severity analysis in cleft lip and palate speech using vowel space area., in: Interspeech, 2017, pp. 1829–1833
2017
-
[49]
Russell, S
M. Russell, S. D’Arcy, Challenges for computer recognition of children²s speech, in: Proc. Speech and Language Technology in Education (SLaTE 2007), 2007, pp. 108–111. doi:10.21437/SLaTE.2007-26
2007 doi
-
[50]
J. J. Howard, E. J. Laird, Y . B. Sirotin, R. E. Rubin, J. L. Tipton, A. R. Vemury, Evaluating proposed fairness models for face recognition algorithms, in: ICPR Workshops, 2022
2022
-
[51]
Liang, J
A. Liang, J. Lu, X. Mu, Algorithm design: Fairness and accuracy *, 2022
2022
-
[52]
Sakoe, Dynamic programming algorithm optimization for spoken word recognition, IEEE Transactions on Acoustics, Speech, and Signal Processing 26 (1978) 159–165
H. Sakoe, Dynamic programming algorithm optimization for spoken word recognition, IEEE Transactions on Acoustics, Speech, and Signal Processing 26 (1978) 159–165
1978
-
[53]
D. A. Reynolds, T. F. Quatieri, R. B. Dunn, Speaker verification using adapted gaussian mixture models, Digital Signal Processing 10 (2000) 19–41
2000
-
[54]
M. J. F. Gales, S. J. Young, The application of hidden markov models in speech recognition, Found. Trends Signal Process. 1 (2007) 195–304
2007
-
[56]
L. R. Rabiner, A tutorial on hidden markov models and selected applications in speech recognition, Proc. IEEE 77 (1989) 257–286
1989
-
[57]
Salton, M
G. Salton, M. J. McGill, Introduction to Modern Information Retrieval, McGraw-Hill, 1983
1983
-
[58]
Conneau, A
A. Conneau, A. Baevski, R. Collobert, A. rahman Mohamed, M. Auli, Unsupervised cross-lingual representation learning for speech recog- nition, ArXiv abs/2006.13979 (2020)
2020 arXiv
-
[59]
Baevski, H
A. Baevski, H. Zhou, A. rahman Mohamed, M. Auli, wav2vec 2.0: A framework for self-supervised learning of speech representations, ArXiv abs/2006.11477 (2020). 17
2020 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.