Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Improving ASR Fairness for Cleft Lip and Palate Speech: A Study on Severity-Aware Data Mixing

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Severity-aware mixing of cleft lip and palate (CLP) speech with normal speech during ASR training improves accuracy and fairness, cutting pooled word error rate from 22.64% to 18.76% on Kannada and from 28.45% to 18.89% on English child…

desk verdict A first fairness evaluation for CLP speech with a sensible augmentation idea, but the evidence is confounded by speaker overlap and uncontrolled training set size. read the letter →

arxiv 2505.03697 v2 pith:QDNMIG3M submitted 2025-05-06 eess.AS

classification eess.AS
keywords automaticspeechrecognitioncleftlipandpalateASRfairnessseverity-awaredatamixinghypernasalityWhisperXLS-Rscore
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether automatic speech recognition (ASR) is fair to people with cleft lip and palate, whose speech is hypernasal and breathy because of structural differences in the mouth and nose. It first shows that a commercial speech-to-text API errs far more on CLP speech than on typical speech. It then argues that mild and moderate CLP speech still preserves enough spectral structure to be useful in training, and proposes a simple intervention: mix severity-graded CLP utterances with normal speech when training ASR models. On Kannada and English child speech, this lowers word error rate and improves a two-group fairness score across GMM-HMM, Whisper, and XLS-R models. The headline numbers are pooled word error rate falling from 22.64% to 18.76% for GMM-HMM on the Kannada corpus and from 28.45% to 18.89% for Whisper on the English corpus.

What carries the argument

The load-bearing object is the severity ordering of CLP speech, measured by dynamic time warping distance between voiced spectrogram frames of normal and CLP utterances, together with the two-group fairness score. The dynamic time warping distance supplies a graded distortion axis from mild to moderate to severe, motivating which utterances to add during training. The fairness score, defined as $FS = -\alpha \cdot \text{Average Error Rate} - \beta \cdot \text{Error Disparity}$, collapses the normal-versus-CLP word error rates into a single number, with values closer to zero meaning fairer. The mixing strategy trains a model on normal speech plus a severity subset of CLP speech and evaluates on the full test set, allowing the choice of severity composition to be tuned for a given model and language.

What would settle it

Re-run the best severity-mixed training condition with a speaker-disjoint split, where all recordings of each speaker are assigned wholly to either training or testing, on both corpora, and compare the pooled word error rate and fairness score against the paper's utterance-level split. If the improvement disappears or falls below a meaningful margin, the reported fairness gain is speaker memorization rather than generalization to unseen CLP speakers.

Watch

Extended reading notes

Core claim

The central claim is that ASR fairness for cleft lip and palate speech can be improved by severity-aware data mixing: rather than training only on normal speech or only on CLP speech, the authors train on normal speech augmented with CLP utterances in order of increasing severity. They introduce a fairness score that trades off the average error rate against the error disparity between normal and CLP groups, with values closer to zero indicating a fairer system. Spectrogram, formant-contour, and dynamic time warping diagnostics show that spectral distortion grows progressively from mild to moderate to severe CLP speech, which justifies mixing by severity groups. In their experiments, adding mild, moderate, and severe CLP utterances to the normal training data improves pooled word error rate and moves the fairness score closer to zero on both corpora; the best configuration for GMM-HMM on the Kannada corpus includes severe speech, while for Whisper on the English corpus the best configuration stops at moderate. The reported fairness-score improvements are 17.89% on the Kannada corpus and 47.50% on the English corpus.

Load-bearing premise

The evaluation uses a random 80/20 split of individual recordings, so the same speaker can appear in both the training and testing sets; if that overlap drives the results, the reported gains may come from remembering a speaker's voice rather than from learning to recognize cleft-palate speech in general.

Editorial extensions

If this is right

  • On both Kannada child speech and English child speech, replacing a normal-only training set with a severity-mixed set reduces the gap between normal and CLP word error rates.
  • The best mixing recipe is not universal: for GMM-HMM on the Kannada corpus it includes severe CLP speech, while for Whisper on the English corpus adding severe utterances hurts, so the recipe must be chosen per model and language.
  • The fairness score supports deployment choices: setting $\alpha=0.9$, $\beta=0.1$ prioritizes overall accuracy, while $\alpha=0.1$, $\beta=0.9$ prioritizes reducing disparity between normal and CLP groups.
  • Training with CLP data barely degrades performance on normal speech, so the fairness improvement is not bought by sacrificing typical speech recognition.
  • Severity-aware mixing improves fairness for a conventional GMM-HMM model, a self-supervised foundation model, and a supervised foundation model, suggesting the mechanism is not tied to one architecture.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the effect is driven by spectral overlap rather than speaker identity, the same severity-graded mixing protocol may transfer to other graded speech disorders, such as dysarthria or stuttering, where distortion increases with severity.
  • Because the reported experiments split recordings randomly rather than by speaker, a speaker-disjoint evaluation is the direct test of whether the gain is acoustic generalization or speaker memorization; that check is not reported.
  • A natural extension is to condition the model on the severity label itself, for instance through an auxiliary loss or an embedding, which could combine the benefit of mixing with explicit severity awareness.
  • The two-group fairness score could be extended to multi-severity fairness by summing pairwise disparities or using the maximum group error, rather than collapsing all CLP speakers into one group.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper addresses automatic speech recognition (ASR) fairness for cleft lip and palate (CLP) speech. It introduces a fairness score (FS) that balances average word error rate and error disparity between normal and CLP groups, evaluates the fairness of a public Google ASR API, and proposes severity-aware mixing of CLP and normal speech during ASR training. Experiments are conducted with GMM-HMM, Whisper, and XLS-R on two datasets: AIISH (Kannada) and NMCPC (English). The authors report that severity-aware augmentation improves FS and reduces WER; for example, the full-text abstract reports WER decreasing from 22.64% to 18.76% (GMM-HMM, AIISH) and from 28.45% to 18.89% (Whisper, NMCPC).

Significance. If the reported effect is robust, the paper would be a useful empirical contribution to ASR fairness for disordered speech: it provides a simple fairness metric, a concrete data-mixing recipe, and evaluations across a public API, a traditional HMM system, and two foundation models in two languages. The paper also documents a striking performance gap in a public ASR service for CLP speech. However, the central claim depends on unaddressed confounds and internal inconsistencies. The utterance-level split, lack of size-matched controls, and post hoc selection of the best augmentation condition make the current evidence suggestive rather than demonstrative; the reported headline numbers are also internally inconsistent.

major comments (4)
  1. [Section 3 (Database setup)] The 80/20 split is utterance-level, not speaker-disjoint. With only 60 speakers in AIISH and 65 in NMCPC, the same speaker very likely contributes utterances to both training and evaluation. CLP speech carries speaker-specific hypernasality and articulation patterns, so the reported WER gains may reflect speaker-specific memorization rather than generalization to unseen CLP speakers. A speaker-disjoint split, and ideally speaker-independent evaluation, is needed to support the claim that severity-aware mixing improves fairness for new CLP speakers.
  2. [Tables 7 and 8] The augmentation conditions are not matched in training-set size or total CLP content. 'Mild+Normal', 'Mild+Moderate+Normal', and 'Mild+Moderate+Severe+Normal' add monotonically more CLP utterances, and there is no control adding an equivalent amount of normal-only data or an equivalent amount of severity-blind CLP data. Moreover, the 'CLP' training condition alone improves FS substantially (AIISH GMM-HMM FS rises from -31.57 with Normal training to -26.67 with CLP training), so the reported gains are not specifically attributable to severity-aware selection.
  3. [Abstract and Tables 5/7] The headline numbers are internally inconsistent. The metadata abstract states WER decreases from 37.58% to 25.47% (GMM-HMM, AIISH) and from 35.74% to 21.72% (Whisper, NMCPC), while the full-text abstract states 22.64% to 18.76% and 28.45% to 18.89%, respectively. In addition, Table 5 lists the GMM-HMM AIISH CLP-trained CLP WER as 36.44, whereas Table 7 lists the same quantity as 11.00; the Whisper values also differ (55.66 vs. 44.77). These discrepancies undermine the reliability of the reported improvements and must be reconciled.
  4. [Section 6.2, Tables 7-9] The evaluation lacks error bars, confidence intervals, and significance tests. All WER and FS comparisons are point estimates from a single random split, and the best augmentation recipe is chosen post hoc per dataset (Mild+Moderate+Severe+Normal for AIISH, Mild+Moderate+Normal for NMCPC). Without multiple seeds, bootstrap confidence intervals, or a pre-specified selection rule, the paper cannot rule out chance variation or selection effects.
minor comments (5)
  1. [Section 6.2.1] The phrase 'crisis cross' should be 'criss-cross'.
  2. [Table 3] The header 'datsets' should be 'datasets'.
  3. [Reference list] Reference [49] contains an encoding artifact ('children²s'); references [55] and [56] are identical and one should be removed.
  4. [Table 7] The Whisper AIISH CLP-trained row reports '7' for the normal test column without a decimal place; clarify whether this is 7.00 or another value.
  5. [Section 4.2] The FS definition states 'α,β> = 0' but the experiments later use α+β=1; clarify the constraint on the weights.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central results are empirical evaluations on held-out test partitions, the fairness score is a definition rather than a derived prediction, and self-citations are contextual prior work, not load-bearing premises.

full rationale

The paper's derivation chain is empirical rather than circular. The central claims—that public ASR systems are less fair on CLP speech, and that severity-aware mixing of CLP and normal speech improves WER and fairness scores—are supported by direct measurements on held-out evaluation sets reported in Tables 3 through 9. The fairness score FS = −α·AverageErrorRate − β·ErrorDisparity is explicitly introduced as a metric definition; it does not by construction force any particular WER outcome, and the improvement claims are evaluated independently of how FS is weighted. The DTW proximity analysis is an independent acoustic-motivation step: it supports the hypothesis that mild and moderate CLP speech retains spectral alignment with normal speech, but the DTW distances are not fitted parameters used to compute the reported WERs. Self-citations to the authors' earlier CLP classification and enhancement work (e.g., [10], [15], [37]) are contextual literature references and are not invoked as uniqueness theorems or as premises that force the empirical results. Possible weaknesses such as the utterance-level 80/20 split potentially allowing speaker overlap, the lack of size-matched controls, and internal inconsistencies between Table 5 and Table 7 are experimental-validity and reporting concerns, not circularity under the specified patterns. There is no step where a prediction reduces by construction to its input or where a fitted parameter is renamed as a prediction.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The paper's conclusions rest on hand-picked FS weights (alpha, beta), a thresholded VAD for the motivating DTW analysis, unvalidated severity labels, and an unclear data split. The FS is a new metric with no normative grounding, so the fairness improvement percentages are metric-dependent. No new physical or computational entities are introduced.

free parameters (4)
  • alpha (FS weight on average error) = 0.5 (also 0.1, 0.9)
    Weights in the proposed fairness score FS = -alpha*average_error - beta*disparity; chosen by hand in Section 4.2 and Section 6.1.2. Reported FS improvements (17.89%, 47.50%) depend on this choice.
  • beta (FS weight on error disparity) = 0.5 (also 0.9, 0.1)
    Companion weight in FS; no principled method is given for selecting alpha and beta.
  • VAD energy threshold = 6% of average energy
    Section 5 uses energy-based VAD with threshold 6% of average energy to select voiced frames for the DTW proximity analysis; arbitrary and not justified by data.
  • GMM-HMM senones and Gaussians = 50 senones, 500 Gaussians
    Section 6.1 sets these Kaldi hyperparameters without sensitivity analysis; WER could change with other settings.
assumptions (5)
  • domain assumption The 80/20 random utterance split yields independent training and evaluation sets.
    Section 3 describes only a random split by utterances; if the same speaker appears in both partitions, the evaluation is not speaker-independent and the central WER improvements are inflated.
  • domain assumption CLP severity labels (mild/moderate/severe) in AIISH and NMCPC are accurate and consistent.
    All augmentation analysis relies on these labels; no validation against clinical ratings is reported for either dataset.
  • domain assumption WER is an appropriate and sufficient proxy for ASR fairness between normal and CLP groups.
    The fairness score is built directly from WER values; other dimensions such as intelligibility or user experience are not modeled.
  • standard math Dynamic time warping distance between voiced spectrogram frames is a valid measure of spectral distortion.
    Used in Section 5 to justify the severity ordering; the DTW recursion is standard but the specific feature (voiced spectrogram frames, energy threshold) is a modeling choice.
  • ad hoc to paper The linear combination of average error and error disparity captures fairness.
    FS = -alpha*average_error - beta*|disparity| is introduced without justification from social-choice or fairness literature; the reported fairness improvements are relative movements of this hand-defined scalar.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Improving ASR Fairness for Cleft Lip and Palate Speech: A Study on Severity-Aware Data Mixing." pith.science (2026). https://pith.science/paper/QDNMIG3M

@misc{pith2026250503697,
  author       = {Pith},
  title        = {Pith review of: Improving ASR Fairness for Cleft Lip and Palate Speech: A Study on Severity-Aware Data Mixing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QDNMIG3M}},
  note         = {Machine review of arXiv:2505.03697}
}
read the original abstract

Speech produced by individuals with cleft lip and palate (CLP) is often hypernasal (and sometimes breathy) due to structural anomalies, yielding shifts in formant structure that degrade automatic speech recognition (ASR) performance and fairness. Building on evidence that mainstream ASR systems underperform on atypical and disordered speech, we posit that widely used services (e.g., Google Speech-to-Text) exhibit reduced fairness for CLP speech, and we evaluate this claim empirically. To quantify fairness consistently, we introduce a simple fairness score (FS) that trades off overall error and between-group disparity. Despite formant disruptions, mild and moderate CLP speech retains partial spectro-temporal alignment with typical speech, motivating the use of mixing strategies to improve fairness. We systematically investigated severity-aware mixing of CLP and normal speech at different severity levels and assessed its effect on fairness. Three ASR models GMM-HMM, Whisper, and XLSR were evaluated on the AIISH (Kannada language) and NMCPC (English language) datasets. A mixing strategy that leverages severity-aware mixing of CLP and normal speech improves fairness on both English (NMCPC) and Kannada (AIISH) corpora. Notably, the word error rate (WER) decreased from 37.58% to 25.47% (GMM-HMM, AIISH) and from 35.74% to 21.72% (Whisper, NMCPC).

Figures

Figures reproduced from arXiv: 2505.03697 by the authors.

Figure 1
Figure 1. Speech signal produced by Normal, mild, moderate and severe subjects, and their corresponding spectrogram and formant contour, while [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. (a-h) Speech signal produced by Normal, mild, moderate, and severe subjects, and their corresponding spectrogram, and (i) showing the [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Fairness in ASR computation pipeline . 5.1. Description of the used ASR models The block diagram in [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Bar graph showing FS for AIISH dataset. No represents Normal, Mi represents Mild, Mo represents Moderate and Se represents Severe. [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: Bar graph showing FS for NMCPC dataset. No represents Normal, Mi represents Mild, Mo represents Moderate and Se represents [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Normal-Anchored First-Order Model-Agnostic Meta-Learning based Whisper Fine-Tuning for Enhancing Fairness of Cleft Lip and Palate Speech Recognition

    eess.SP 2026-07 conditional novelty 5.0 of 10

    Normal-anchored FOMAML fine-tuning of Whisper lowers word error rates on cleft lip and palate speech in two datasets, but the key comparison to conventional fine-tuning uses different training data.

Reference graph

Works this paper leans on

58 extracted references · 51 canonical work pages · cited by 1 Pith paper

  1. [1]

    D. J. Zajac, L. Vallino, Evaluation and Management of Cleft Lip and Palate: A Developmental Perspective, 2017. doi:10.1109/ICME.2016. 7552917

  2. [2]

    Lohmander, M

    A. Lohmander, M. Olsson, Methodology for perceptual assessment of speech in patients with cleft palate: A critical review of the literature, The Cleft Palate-Craniofacial Journal 41 (2004) 64 – 70

  3. [3]

    Stengelhofen, Cleft palate: The nature and remediation of communication problems, Churchill Livingstone (1993)

    J. Stengelhofen, Cleft palate: The nature and remediation of communication problems, Churchill Livingstone (1993)

  4. [4]

    D. Wang, B. Zhang, Q. Zhang, Y . Wu, Global, regional and national burden of orofacial clefts from 1990 to 2019: an analysis of the global burden of disease study 2019, Annals of Medicine 55 (2023). doi:10.1080/07853890.2023.2215540

  5. [5]

    K. M. V . Lierde, S. Claeys, M. D. Bodt, P. V . Cauwenberge, V ocal quality characteristics in children with cleft palate: a multiparameter approach., Journal of voice : o fficial journal of the V oice Foundation 18 3 (2004) 354–62

  6. [6]

    D. J. Zajac, C. Plante, A. Lloyd, K. L. Haley, Reliability and validity of a computer-mediated, single-word intelligibility test: Preliminary findings for children with repaired cleft lip and palate, The Cleft Palate-Craniofacial Journal 48 (2011) 538 – 549

  7. [7]

    T. L. Whitehill, C. H. F. Chau, Single-word intelligibility in speakers with repaired cleft palate, Clinical Linguistics & Phonetics 18 (2004) 341 – 355

  8. [8]

    A. W. Kummer, Cleft Palate and Craniofacial Anomalies: E ffects on Speech and Resonance, 2007

Show all 58 references
  1. [9]

    Baumann, D

    I. Baumann, D. Wagner, F. Braun, S. P. Bayerl, E. Noth, K. Riedhammer, T. Bocklet, Influence of utterance and speaker characteristics on the classification of children with cleft lip and palate, INTERSPEECH 2023 (2022)

  2. [10]

    Bhattacharjee, H

    S. Bhattacharjee, H. S. Shekhawat, S. R. M. Prasanna, Classification of cleft lip and palate speech using fine-tuned transformer pretrained models, in: B. J. Choi, D. Singh, U. S. Tiwary, W.-Y . Chung (Eds.), Intelligent Human Computer Interaction, Springer Nature Switzerland,...

  3. [11]

    Kalita, G

    S. Kalita, G. K S, P. Mariswamy, S. Prasanna, S. Dandapat, Objective assessment of cleft lip and palate speech intelligibility using articulation and hypernasality measures, The Journal of the Acoustical Society of America 146 (2019) 1164–1175. doi: 10.1121/1.5121310

  4. [12]

    Kalita, S

    S. Kalita, S. R. M. Prasanna, S. Dandapat, Importance of glottis landmarks for the assessment of cleft lip and palate speech intelligibility., The Journal of the Acoustical Society of America 144 5 (2018) 2656

  5. [13]

    Kalita, S

    S. Kalita, S. R. M. Prasanna, S. Dandapat, Intelligibility assessment of cleft lip and palate speech using gaussian posteriograms based on joint spectro-temporal features., The Journal of the Acoustical Society of America 144 4 (2018) 2413

  6. [14]

    C. M. Vikram, S. Macha, S. Kalita, S. R. M. Prasanna, Acoustic analysis of misarticulated trills in cleft lip and palate children., The Journal of the Acoustical Society of America 143 6 (2018) EL474

  7. [15]

    Bhattacharjee, R

    S. Bhattacharjee, R. Sinha, Sensitivity analysis of maskcyclegan based voice conversion for enhancing cleft lip and palate speech recognition, 2022, pp. 1–5. doi:10.1109/SPCOM55316.2022.9840769

  8. [16]

    X. Wang, S. Yang, M. Tang, H. Yin, H. Huang, L. He, Hypernasalitynet: Deep recurrent neural network for automatic hypernasality detection, International Journal of Medical Informatics 129 (2019) 1–12. doi:https://doi.org/10.1016/j.ijmedinf.2019.05.023

  9. [17]

    J. H. Ha, H. Lee, S. M. Kwon, H. Joo, G. Lin, D. Y . Kim, S. Kim, J. Y . Hwang, J. H. Chung, H. J. Kong, Deep learning–based diagnostic system for velopharyngeal insu fficiency based on videofluoroscopy in patients with repaired cleft palates, Journal of Craniofacial Surgery 1...

  10. [18]

    A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y . Saraf, J. M. Pino, A. Baevski, A. Conneau, M. Auli, Xls-r: Self-supervised cross-lingual speech representation learning at scale, ArXiv abs /2111.09296 (2021)

  11. [19]

    Radford, J

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, I. Sutskever, Robust speech recognition via large-scale weak supervision, in: Proceedings of the 40th International Conference on Machine Learning, ICML’23, JMLR.org, 2023

  12. [20]

    A. K. Dubey, S. R. M. Prasanna, S. Dandapat, Zero time windowing based severity analysis of hypernasal speech, 2016 IEEE Region 10 Conference (TENCON) (2016) 970–974

  13. [21]

    A. K. Dubey, S. R. M. Prasanna, S. Dandapat, Hypernasality detection using zero time windowing, 2018 International Conference on Signal Processing and Communications (SPCOM) (2018) 105–109

  14. [22]

    Nikitha, S

    K. Nikitha, S. Kalita, M. VikramC., M. Pushpavathi, S. R. M. Prasanna, Hypernasality severity analysis in cleft lip and palate speech using vowel space area, in: Interspeech, 2017

  15. [23]

    VikramC., A

    M. VikramC., A. Tripathi, S. Kalita, S. R. M. Prasanna, Estimation of hypernasality scores from cleft lip and palate speech, in: Interspeech, 2018

  16. [24]

    V . C. Mathad, N. J. Scherer, K. Chapman, J. M. Liss, V . Berisha, An attention model for hypernasality prediction in children with cleft palate, ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2021) 7248–7252

  17. [25]

    A. K. Dubey, S. R. M. Prasanna, S. Dandapat, Pitch-adaptive front-end feature for hypernasality detection, in: Interspeech, 2018

  18. [26]

    A. K. Dubey, S. R. M. Prasanna, S. Dandapat, Hypernasality severity detection using constant q cepstral coe fficients, in: Interspeech, 2019. 16

  19. [27]

    A. K. Dubey, S. R. M. Prasanna, S. Dandapat, Sinusoidal model-based hypernasality detection in cleft palate speech using cvcv sequence, Speech Commun. 124 (2020) 1–12

  20. [28]

    A. K. Dubey, S. R. M. Prasanna, S. Dandapat, Detection and assessment of hypernasality in repaired cleft palate speech using vocal tract and residual features., The Journal of the Acoustical Society of America 146 6 (2019) 4211

  21. [29]

    Kalita, S

    S. Kalita, S. R. M. Prasanna, S. Dandapat, Self-similarity matrix based intelligibility assessment of cleft lip and palate speech, in: Interspeech, 2018

  22. [30]

    Kalita, K

    S. Kalita, K. S. Girish, P. M., S. R. M. Prasanna, S. Dandapat, Objective assessment of cleft lip and palate speech intelligibility using articulation and hypernasality measures., The Journal of the Acoustical Society of America 146 2 (2019) 1164

  23. [31]

    Kalita, P

    S. Kalita, P. N. Sudro, S. R. M. Prasanna, S. Dandapat, Nasal air emission in sibilant fricatives of cleft lip and palate speech, in: Interspeech, 2019

  24. [32]

    C. M. Vikram, N. Adiga, S. R. M. Prasanna, Detection of nasalized voiced stops in cleft palate speech using epoch-synchronous features, IEEE/ACM Transactions on Audio, Speech, and Language Processing 27 (2019) 1189–1200

  25. [33]

    P. N. Sudro, S. Kalita, S. R. M. Prasanna, Processing transition regions of glottal stop substituted /s/ for intelligibility enhancement of cleft palate speech, in: Interspeech, 2018

  26. [34]

    P. N. Sudro, S. M. Prasanna, Enhancement of cleft palate speech using temporal and spectral processing, Speech Communication 123 (2020) 70–82

  27. [35]

    P. N. Sudro, S. M. Prasanna, Modification of misarticulated fricative /s/in cleft lip and palate speech, Biomedical Signal Processing and Control 67 (2021) 102088

  28. [36]

    P. N. Sudro, C. M. Vikram, S. R. M. Prasanna, Event-based transformation of misarticulated stops in cleft lip and palate speech, Circuits, Systems, and Signal Processing 40 (2021) 4064 – 4088

  29. [37]

    P. N. Sudro, R. K. Das, R. Sinha, S. R. M. Prasanna, Enhancing the intelligibility of cleft lip and palate speech using cycle-consistent adversarial networks, 2021 IEEE Spoken Language Technology Workshop (SLT) (2021) 720–727

  30. [38]

    P. N. Sudro, R. K. Das, R. Sinha, S. R. Mahadeva Prasanna, Significance of data augmentation for improving cleft lip and palate speech recognition, in: 2021 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2021, pp. 484–490

  31. [39]

    Baumann, D

    I. Baumann, D. Wagner, F. Braun, S. P. Bayerl, E. N ¨oth, K. Riedhammer, T. Bocklet, Influence of utterance and speaker characteristics on the classification of children with cleft lip and palate, in: Interspeech, 2022

  32. [40]

    K. Song, T. Wan, B. Wang, H. Jiang, L. K. Qiu, J. Xu, L. ping Jiang, Q. Lou, Y . Yang, D. Li, X. Wang, L. Qiu, Improving hypernasality estimation with automatic speech recognition in cleft palate speech, in: Interspeech, 2022

  33. [41]

    VikramC., S

    M. VikramC., S. R. M. Prasanna, A. K. Abraham, M. Pushpavathi, S. GirishK., Detection of glottal activity errors in production of stop consonants in children with cleft lip and palate, in: Interspeech, 2018

  34. [42]

    V . C. Mathad, S. R. M. Prasanna, V owel onset point based screening of misarticulated stops in cleft lip and palate speech, IEEE /ACM Transactions on Audio, Speech, and Language Processing 28 (2020) 450–460

  35. [43]

    V . C. Mathad, N. J. Scherer, K. Chapman, J. M. Liss, V . Berisha, A deep learning algorithm for objective assessment of hypernasality in children with cleft palate, IEEE Transactions on Biomedical Engineering 68 (2020) 2986–2996

  36. [44]

    Fujiwara, R

    K. Fujiwara, R. Takashima, C. Sugiyama, N. Tanaka, K. Nohara, K. Nozaki, T. Takiguchi, Data augmentation based on frequency warping for recognition of cleft palate speech, in: 2021 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ...

  37. [45]

    M. H. Javid, K. Gurugubelli, A. K. Vuppala, Single frequency filter bank based long-term average spectra for hypernasality detection and assessment in cleft lip and palate speech, in: ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (...

  38. [46]

    Bocklet, A

    T. Bocklet, A. K. Maier, K. Riedhammer, U. Eysholdt, E. N ¨oth, Erlangen-clp: A large annotated corpus of speech from children with cleft lip and palate, in: International Conference on Language Resources and Evaluation, 2014

  39. [47]

    Eshky, M

    A. Eshky, M. S. Ribeiro, J. Cleland, K. Richmond, Z. Roxburgh, J. M. Scobbie, A. A. Wrench, Ultrasuite: A repository of ultrasound and acoustic data from child speech therapy sessions, ArXiv abs/1907.00835 (2018)

  40. [48]

    Nikitha, S

    K. Nikitha, S. Kalita, C. Vikram, M. Pushpavathi, S. M. Prasanna, Hypernasality severity analysis in cleft lip and palate speech using vowel space area., in: Interspeech, 2017, pp. 1829–1833

  41. [49]

    Russell, S

    M. Russell, S. D’Arcy, Challenges for computer recognition of children²s speech, in: Proc. Speech and Language Technology in Education (SLaTE 2007), 2007, pp. 108–111. doi:10.21437/SLaTE.2007-26

  42. [50]

    J. J. Howard, E. J. Laird, Y . B. Sirotin, R. E. Rubin, J. L. Tipton, A. R. Vemury, Evaluating proposed fairness models for face recognition algorithms, in: ICPR Workshops, 2022

  43. [51]

    Liang, J

    A. Liang, J. Lu, X. Mu, Algorithm design: Fairness and accuracy *, 2022

  44. [52]

    Sakoe, Dynamic programming algorithm optimization for spoken word recognition, IEEE Transactions on Acoustics, Speech, and Signal Processing 26 (1978) 159–165

    H. Sakoe, Dynamic programming algorithm optimization for spoken word recognition, IEEE Transactions on Acoustics, Speech, and Signal Processing 26 (1978) 159–165

  45. [53]

    D. A. Reynolds, T. F. Quatieri, R. B. Dunn, Speaker verification using adapted gaussian mixture models, Digital Signal Processing 10 (2000) 19–41

  46. [54]

    M. J. F. Gales, S. J. Young, The application of hidden markov models in speech recognition, Found. Trends Signal Process. 1 (2007) 195–304

  47. [56]

    L. R. Rabiner, A tutorial on hidden markov models and selected applications in speech recognition, Proc. IEEE 77 (1989) 257–286

  48. [57]

    Salton, M

    G. Salton, M. J. McGill, Introduction to Modern Information Retrieval, McGraw-Hill, 1983

  49. [58]

    Conneau, A

    A. Conneau, A. Baevski, R. Collobert, A. rahman Mohamed, M. Auli, Unsupervised cross-lingual representation learning for speech recog- nition, ArXiv abs/2006.13979 (2020)

  50. [59]

    Baevski, H

    A. Baevski, H. Zhou, A. rahman Mohamed, M. Auli, wav2vec 2.0: A framework for self-supervised learning of speech representations, ArXiv abs/2006.11477 (2020). 17

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.