Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

Overview of Automatic Speech Analysis and Technologies for Neurodegenerative Disorders: Diagnosis and Assistive Applications

T0 review · 2 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This review claims to be the first comprehensive survey of automatic speech technologies for neurodegenerative disorders, spanning detection, recognition, intelligibility enhancement, and assessment, and it provides a catalog of the…

desk verdict A useful survey map of pathological speech processing, but the dataset catalog (Table I) has arithmetic errors that must be fixed before it serves as a reliable reference. read the letter →

arxiv 2501.03536 v2 pith:SRDWB665 submitted 2025-01-07 eess.AS

classification eess.AS
keywords pathologicalspeechneurodegenerativedisordersdetectionautomaticrecognitionintelligibilityenhancementseverityassessmentdataaugmentationdatasets
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper aims to be the first comprehensive survey of automatic speech technologies for neurodegenerative disorders, covering both the clinical goal of diagnosis and monitoring and the technological goal of making speech-based systems work for impaired speakers. It organizes the field into five problem areas—pathological speech detection, automatic speech recognition, intelligibility enhancement, intelligibility and severity assessment, and data augmentation—and reviews the speech representations used in each. It also compiles a table of more than a dozen pathological speech datasets, marking each as public, accessible, or non-accessible. The paper then identifies cross-cutting challenges such as robustness, privacy, interpretability, and speech mode, and proposes future directions including multimodal analysis and large language models. A sympathetic reader would take the paper's contribution to be a structured map of the field plus a dataset catalog that researchers can use to orient themselves.

What carries the argument

The organizing device is a taxonomy of the field into five research problems—detection, recognition, enhancement, assessment, and data augmentation—together with a catalog of pathological speech datasets that maps each dataset's size, impairment type, language, modality, and accessibility. This structure does the argument's work: by aligning each subfield with the same set of datasets and speech representations (handcrafted features, time-frequency representations, raw waveforms, self-supervised embeddings), the survey makes the field's fragmentation visible and lets the authors claim comprehensiveness. The definition of pathological speech as speech that deviates from neurotypical patterns due to underlying impairments, with deviations in voice, articulation, prosody, and language, sets the boundaries of the review.

What would settle it

Compare every row of Table I against the primary dataset publications: for example, the text says the AMSDC corpus has 99 patients but the table lists 62 control and 37 pathological speakers, and the mPower table entries do not match the stated 6,805 total participants; a systematic mismatch would show the catalog is not a reliable reference.

Watch

Extended reading notes

Core claim

The paper's central claim is that no existing review covers pathological speech from both clinical and technological perspectives across detection, recognition, enhancement, and assessment, and that this survey fills that gap. On the paper's own terms, the discovery is the scope itself: a unified treatment of how automatic systems can decide whether speech is pathological, recognize what a pathological speaker is saying, make pathological speech more intelligible, estimate intelligibility and severity, and generate or perturb data to train such systems. The paper also documents the field's reliance on a small set of mostly small, often private datasets, and its shift from handcrafted acoustic features toward self-supervised embeddings as the current state of the art. It concludes that robustness, privacy, interpretability, and generalization across speech modes and languages are the open problems, and points to multimodal and large-language-model approaches as the likely next steps.

Load-bearing premise

The survey's usefulness rests on the accuracy of its dataset table and on its summaries of the cited studies being faithful to what those studies actually report.

Editorial extensions

If this is right

  • Researchers entering the field can use the survey as a single entry point: the dataset list marks which corpora are available, under what conditions, and with which impairments.
  • The paper's review implies that self-supervised embeddings such as wav2vec2, HuBERT, and WavLM are currently the strongest input representations for detection and recognition, so new work should build on those rather than handcrafted features.
  • The catalog's accessibility column makes it possible to see why replication is hard: many standard datasets are private, and even public ones such as TORGO carry recording artifacts.
  • The challenges section argues that future progress will come from robustness to environmental and adversarial distortions, privacy-preserving training, interpretable models, and use of spontaneous speech.
  • Following the paper's logic, multimodal and large-language-model systems are the designated next research frontier for both clinical assessment and assistive devices.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 'first comprehensive survey' claim is definition-dependent: prior reviews cover single disorders or single tasks, so whether this is genuinely first hinges on accepting the clinical-plus-technological scope as the relevant frame.
  • The survey's emphasis on self-supervised embeddings implies a concrete testable prediction: detection and recognition systems built on SSL features should dominate leaderboards on any new pathological speech benchmark, while handcrafted-feature systems should lag on larger datasets.
  • The authors' call for LLM-based and multimodal approaches could be sharpened by a shared benchmark that reports intelligibility gains and word error rates on standardized pathological speech corpora, allowing direct comparison across future systems.
  • The dataset catalog could naturally evolve into a living community resource, with corrections and additions tracked over time, which would make the survey's practical utility outlast its publication date.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This manuscript is an overview article on automatic speech analysis and processing technologies for speech affected by neurodegenerative disorders. It reviews pathological speech datasets (Section II), speech representations (Section III), automatic detection (Section IV), ASR for pathological speech (Section V), intelligibility enhancement (Section VI), intelligibility and severity assessment (Section VII), data augmentation (Section VIII), and challenges and future directions (Section IX). The two stated contributions are a comprehensive survey spanning detection, recognition, enhancement, and assessment, and a structured catalog of accessible and non-accessible pathological speech datasets.

Significance. If corrected, the paper would be a useful entry point for researchers and clinicians working on pathological speech. Its organization by task rather than by disorder is helpful, and the inclusion of dataset accessibility labels, robustness concerns, privacy, explainability, and multimodal/LLM directions is timely. The manuscript does not claim new algorithms or experimental results; its value lies in synthesis and bibliographic coverage, so the accuracy of its dataset table and the fidelity of its summaries of cited works are central to its reliability.

major comments (2)
  1. [Section II, Table I] The dataset catalog, one of the paper's two stated contributions, contains multiple verified numerical errors and must be corrected against the primary sources. (i) AMSDC: the text says 99 patients (62 male, 37 female) and no controls, but the table lists Control=62, Pathological=37, placing sex counts in the wrong columns. (ii) Italian Parkinson's Database: the text says 28 patients and 37 controls, but the table lists Control=28, Pathological=37, a second column swap. (iii) COPAS: the text says 197 pathological and 122 control speakers, but the table lists Control=197, Pathological=122. (iv) mPower: the text states 6,805 participants, 1,087 PD and 5,581 controls, while 1,087+5,581=6,668; the table repeats the 1,087/5,581 breakdown, so either the total or the group sizes are wrong. (v) Saarbrücken: the text states 2,255 German speakers, but 1,356 patients + 869 controls = 2,225, not 2,255; the table matches the components, so the total in the text is the error. (vi) NeuroVoz: the text states 54 patients and 58 controls, but the subgroup counts (33+20 patients; 28+26+1 controls) sum to 53 and 55, respectively. These arithmetic and column-semantics errors undermine the reliability of the dataset catalog and must be fixed before publication.
  2. [Section I, Contribution 1] The claim to present 'the first comprehensive survey' is load-bearing for the paper's novelty. The related works [41-51] include broad reviews of pathological speech processing and of PD/dementia detection, and the manuscript currently distinguishes them only by saying they are narrower or outdated. To make the novelty claim verifiable, the authors should either provide a systematic comparison of tasks, disorders, and time coverage against those reviews, or revise the claim to 'a comprehensive survey' with the intended scope stated explicitly.
minor comments (5)
  1. [Section II, Table I] The table contains typographical errors that should be corrected: 'Dysarthira' appears twice, 'CV A' and 'EW A-DB' should be 'CVA' and 'EWA-DB', and 'Saarbrucken' should be 'Saarbrücken'.
  2. [Section II, CUDYS] The text describes the impairment as 'cerebellar degeneration' while Table I says 'spino-cerebellar ataxia'; use consistent terminology or clarify that these refer to the same diagnostic category.
  3. [Section I, Contribution 1] The phrase 'from a clinical and technological perspectives' is ungrammatical; use 'from clinical and technological perspectives' or 'from a clinical and a technological perspective'.
  4. [References, [30]] The bibliographic entry lists IEEE Transactions on Speech and Audio Processing with volume 255, which is not a plausible volume number; verify and correct the reference.
  5. [Section II] The paper does not describe how the surveyed literature and dataset information were selected or verified; adding a brief statement on search databases, years, and inclusion criteria would make the comprehensiveness claim more transparent.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular reasoning found: this is a literature survey with no derived predictions, and self-citations are descriptive rather than load-bearing.

full rationale

This manuscript is an overview/review paper. It does not fit parameters, derive quantities, or make predictions from first principles; its content is a structured summary of existing datasets, features, methods, challenges, and future directions. The one substantive checkable contribution, the dataset catalog in Table I, contains arithmetic and column-semantics errors (e.g., AMSDC Control/Pathological entries do not match the text, COPAS control and pathological counts are swapped, the mPower row does not sum to the stated total, and the Saarbruecken total is inconsistent with the listed subgroups), but these are factual accuracy defects, not circular reasoning: nothing in the table is constructed from the paper's own outputs. Although several cited works are by the same authors (e.g., [70], [71], [76], [92], [109], [159], [204]-[206], [231], [232], [234], [236]), these citations are used descriptively to summarize prior research and to motivate open challenges; they do not function as the justification of any derived result. The 'first comprehensive survey' claim is a scope statement relative to prior reviews [41]-[51], and assessing it is a matter of coverage and interpretation, not circularity. No equation, fitted parameter, or self-citation chain reduces the paper's conclusions to its inputs, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

A review paper introduces no free parameters or invented entities. Its load-bearing assumptions are the accuracy of its literature summaries and the validity of its 'first comprehensive survey' claim.

assumptions (2)
  • domain assumption The descriptions of the cited algorithms and datasets accurately represent the original publications.
    The review's utility rests on the fidelity of its summaries; this is not proven within the paper.
  • domain assumption The absence of a prior comprehensive survey covering detection, recognition, enhancement, and assessment simultaneously.
    The paper's novelty claim depends on this, but it only loosely addresses overlapping surveys [41-51].

how reviews work

0 comments
Cite this review

Pith. "Pith review of Overview of Automatic Speech Analysis and Technologies for Neurodegenerative Disorders: Diagnosis and Assistive Applications." pith.science (2026). https://pith.science/paper/SRDWB665

@misc{pith2026250103536,
  author       = {Pith},
  title        = {Pith review of: Overview of Automatic Speech Analysis and Technologies for Neurodegenerative Disorders: Diagnosis and Assistive Applications},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SRDWB665}},
  note         = {Machine review of arXiv:2501.03536}
}
read the original abstract

Advancements in spoken language technologies for neurodegenerative speech disorders are crucial for meeting both clinical and technological needs. This overview paper is vital for advancing the field, as it presents a comprehensive review of state-of-the-art methods in pathological speech detection, automatic speech recognition, pathological speech intelligibility enhancement, intelligibility and severity assessment, and data augmentation approaches for pathological speech. It also highlights key challenges, such as ensuring robustness, privacy, and interpretability. The paper concludes by exploring promising future directions, including the adoption of multimodal approaches and the integration of large language models to further advance speech technologies for neurodegenerative speech disorders.

Figures

Figures reproduced from arXiv: 2501.03536 by the authors.

Figure 1
Figure 1. Traditional auditory-perceptual assessment in clinical practice (bounded by the dashed box A) and automatic pathological speech analysis system (bounded by the solid box B). The clinician listens to the (potential) patient and assesses by ear the various characteristics of the speech. The automatic model is trained to detect and analyze speech impairments. Clinicians may use the insights provided by the automatic mo… view at source ↗
Figure 2
Figure 2. Overview of key components discussed in this manuscript, including datasets, features, and research directions for pathological speech. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating the Usefulness of Non-Diagnostic Speech Data for Developing Parkinson's Disease Classifiers

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Non-diagnostic dialogue speech, after concatenation and demographic balancing, can perform nearly as well as diagnostic-task speech for Parkinson's disease classification, though validation leakage clouds the result.

Reference graph

Works this paper leans on

245 extracted references · 80 canonical work pages · cited by 1 Pith paper

  1. [1]

    New developments in understanding the complexity of human speech production,

    K. Simonyan, H. Ackermann, E. F. Chang, and J. D. Greenlee, “New developments in understanding the complexity of human speech production,” Journal of Neuroscience , vol. 36, no. 45, Nov. 2016

  2. [2]

    Clinical evaluation of Parkinson’s-related dysphonia,

    G. K. Sewall, J. Jiang, and C. N. Ford, “Clinical evaluation of Parkinson’s-related dysphonia,” Laryngoscope, vol. 116, no. 10, pp. 1740–1744, Oct. 2024

  3. [3]

    Dif- ferential diagnosis of hoarseness,

    N. Isshiki, H. Okamura, M. Tanabe, and M. Morimoto, “Dif- ferential diagnosis of hoarseness,” Folia Phoniatrica, vol. 21, no. 1, pp. 9–19, 1969

  4. [4]

    J. R. Duffy, Motor Speech Disorders: Substrates, Differential Diagnosis, and Management , 4th ed. Elsevier Health Sci- ences, 2019

  5. [5]

    Apraxia of speech: An overview,

    J. Ogar, H. Slama, N. Dronkers, S. Amici, and M. L. Gorno-Tempini, “Apraxia of speech: An overview,”Neurocase, vol. 11, no. 6, pp. 427–459, Dec. 2005

  6. [6]

    An automatic measure for speech intelligibility in dysarthrias—validation across multiple languages and neurological disorders,

    J. Tr ¨oger, F. D ¨orr, L. Schwed, N. Linz, A. K ¨onig, T. Thies, J. R. Orozco-Arroyave, and J. Rusz, “An automatic measure for speech intelligibility in dysarthrias—validation across multiple languages and neurological disorders,” Frontiers in Digital Health, vol. 6, p. 1440986, 2024

  7. [7]

    Aphasia in dementia of the Alzheimer type,

    J. L. Cummings, F. Benson, M. A. Hill, and S. Read, “Aphasia in dementia of the Alzheimer type,” Neurology, vol. 35, no. 3, pp. 394–401, Mar. 1985

  8. [8]

    J. S. Damico, N. M ¨uller, and M. J. Ball, The Handbook of Language and Speech Disorders. Wiley Online Library, 2010

Show all 245 references
  1. [9]

    Huang, A

    X. Huang, A. Acero, H.-W. Hon, and R. Reddy, Spoken Language Processing: A Guide to Theory, Algorithm, and System Development, 1st ed. USA: Prentice Hall PTR, 2001

  2. [10]

    The apraxia of speech rating scale : A tool for diagnosis and description of apraxia of speech,

    E. A. Strand, J. R. Duffy, H. M. Clark, and K. Josephs, “The apraxia of speech rating scale : A tool for diagnosis and description of apraxia of speech,” Journal of Communication Disorders, vol. 51, pp. 43–50, Sep. 2014

  3. [11]

    Frequency of speech disruptions in Parkinson’s disease and developmental stuttering: A comparison among speech tasks,

    F. S. Juste, F. C. Sassi, J. B. Costa, and C. R. F. de Andrade, “Frequency of speech disruptions in Parkinson’s disease and developmental stuttering: A comparison among speech tasks,” PLoS One, vol. 13, no. 6, June 2018

  4. [12]

    R. T. Sataloff, Clinical Assessment of Voice, Second Edition . Plural Publishing, Sep 2017

  5. [13]

    A forced Gaussians based methodology for the differential evaluation of Parkinson’s disease by means of speech processing,

    L. Moro-Velazquez, J. A. Gomez-Garcia, J. I. Godino- Llorente, J. Villalba, J. Rusz, S. Shattuck-Hufnagel, and N. Dehak, “A forced Gaussians based methodology for the differential evaluation of Parkinson’s disease by means of speech processing,” Biomedical Signal Processing an...

  6. [14]

    W. H. Organization, https://www.who.int/news-room/ fact-sheets/detail/parkinson-disease#: ∼:text=Global% 20estimates%20in%202019%20showed,of%20over%20100% 25%20since%202000., 2023, [Online; accessed 23.12.2024]

  7. [15]

    Global, re- gional, and national burden of Parkinson’s disease, 1990–2016: a systematic analysis for the global burden of disease study 2016,

    GBD 2016 Parkinson’s Disease Collaborators, “Global, re- gional, and national burden of Parkinson’s disease, 1990–2016: a systematic analysis for the global burden of disease study 2016,” The Lancet Neurology , vol. 17, no. 11, Nov. 2016

  8. [16]

    World Alzheimer report 2015,

    A. Wimo, G.-C. Ali, M. Guerchet, M. Prince, M. Prina, and Y .- T. Wu, “World Alzheimer report 2015,” Alzheimer’s Disease International, Tech. Rep., 2015

  9. [17]

    Projected increase in Amyotrophic Lateral Sclerosis from 2015 to 2040,

    K. C. Arthur, A. Calvo, T. R. Price, J. Geiger, A. Chip, and B. J. Traynor, “Projected increase in Amyotrophic Lateral Sclerosis from 2015 to 2040,” Nature Communications, vol. 7, no. 12408, Aug. 2016

  10. [18]

    Imprecise vowel articulation as a potential early marker of parkinson’s disease: Effect of speaking task,

    J. Rusz, R. Cmejla, T. Tykalova, H. Ruzickova, J. Klem- pir, V . Majerova, J. Picmausova, J. Roth, and E. Ruzicka, “Imprecise vowel articulation as a potential early marker of parkinson’s disease: Effect of speaking task,” The Journal of the Acoustical Society of America , vol...

  11. [19]

    P. Rong, Y . Yunusova, J. Wang, and J. Green, “Predicting early bulbar decline in Amyotrophic Lateral Sclerosis: A speech IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, VOL. XX, NO. XX, DECEMBER 2024 13 subsystem approach,” Behavioural Neurology, vol. 2015, pp. 1–11, Jul. 2015

  12. [20]

    Investigating voice as a biomarker: Deep phenotyping meth- ods for early detection of Parkinson’s disease,

    J. M. Tracy, Y . ¨Ozkanca, D. C. Atkins, and R. H. Ghomi, “Investigating voice as a biomarker: Deep phenotyping meth- ods for early detection of Parkinson’s disease,” Journal of Biomedical Informatics, vol. 104, p. 103362, April 2020

  13. [21]

    Progressive Apraxia of Speech as a sign of motor neuron disease,

    J. R. Duffy, R. K. Peach, and E. A. Strand, “Progressive Apraxia of Speech as a sign of motor neuron disease,” Ameri- can Journal of Speech-Language Pathology, vol. 16, no. 3, pp. 198–208, Aug. 2007

  14. [22]

    J. R. Duffy, Parkinson’s disease and movement disorders: di- agnosis and treatment guidelines for the practicing physician . Humana Press, 2000, ch. Motor speech disorders: Clues to neurologic diagnosis, pp. 35–53

  15. [23]

    Consensus auditory-perceptual evaluation of voice: Development of a standardized clinical protocol,

    G. B. Kempster, B. R. Gerratt, K. V . Abbott, J. Barkmeier- Kraemer, and R. E. Hillman, “Consensus auditory-perceptual evaluation of voice: Development of a standardized clinical protocol,” American Journal of Speech Language Pathology , vol. 18, no. 2, pp. 124–132, Oct. 2009

  16. [24]

    Diagnosis of voice disorders,

    K. Omori, “Diagnosis of voice disorders,” Japan Medical Association Journal, vol. 54, no. 4, pp. 248–253, July 2011

  17. [25]

    Automatic assess- ment of voice quality according to the GRBAS scale,

    N. S ´aenz-Lech´on, J. I. Godino-Llorente, V . Osma-Ruiz, M. Blanco-Velasco, and F. Cruz-Rold ´an, “Automatic assess- ment of voice quality according to the GRBAS scale,” in Proc. Annual International Conference of the IEEE Engineering in Medicine and Biology Society , vol. 20...

  18. [26]

    Factor structure of the unified Parkinson’s disease rating scale: Motor examination section,

    G. T. Stebbins and C. G. Goetz, “Factor structure of the unified Parkinson’s disease rating scale: Motor examination section,” Movement Disorders: Journal of the Movement Disorder So- ciety, vol. 13, no. 4, pp. 633–636, July 1998

  19. [27]

    Gauging the auditory dimensions of dysarthric impairment: Reliability and construct validity of the Bogenhausen dysarthria scales (BoDyS),

    W. Ziegler, A. Staiger, T. Sch ¨olderle, and M. V ogel, “Gauging the auditory dimensions of dysarthric impairment: Reliability and construct validity of the Bogenhausen dysarthria scales (BoDyS),” Journal of Speech, Language, and Hearing Re- search, vol. 60, no. 6, pp. 1516–15...

  20. [28]

    K. M. Yorkston and D. R. Beukelman, Assessment of Intelli- gibility of Dysarthric Speech . C.C. Publications, 1981

  21. [29]

    Lis- tener agreement for auditory-perceptual ratings of dysarthria,

    K. Bunton, R. Kent, J. R. Duffy, J. Rosenbek, and J. Kent, “Lis- tener agreement for auditory-perceptual ratings of dysarthria,” Journal of Speech, Language, and Hearing Research , vol. 50, no. 6, pp. 1481–1495, Jan. 2008

  22. [30]

    Accuracy and inter-observer varia- tion in the classification of dysarthria from speech recordings,

    S. Fonville, H. B. van der Worp, P. Maat, M. Aldenhoven, A. Algra, and J. van Gijn, “Accuracy and inter-observer varia- tion in the classification of dysarthria from speech recordings,” IEEE Transactions on Speech and Audio Processing, vol. 255, no. 10, pp. 1545–1548, Oct. 2008

  23. [31]

    On the impact of dysarthric speech on contemporary ASR cloud platforms,

    L. De Russis and F. Corno, “On the impact of dysarthric speech on contemporary ASR cloud platforms,” Journal of Reliable Intelligent Environments, vol. 5, pp. 163–172, 2019

  24. [32]

    Recent progress in the CUHK dysarthric speech recognition system,

    S. Liu, M. Geng, S. Hu, X. Xie, M. Cui, J. Yu, X. Liu, and H. Meng, “Recent progress in the CUHK dysarthric speech recognition system,” IEEE/ACM Transactions on Au- dio, Speech, and Language Processing , vol. 29, pp. 2267– 2281, 2021

  25. [33]

    Formant re-synthesis of dysarthric speech,

    A. Kain, X. Niu, J.-P. Hosom, Q. Miao, and J. P. H. van Santen, “Formant re-synthesis of dysarthric speech,” in Proc. 5th ISCA Workshop on Speech Synthesis (SSW 5) , Pittsburgh, PA, USA, July 2004, pp. 25–30

  26. [34]

    Synthesizing dysarthric speech using multi-speaker TTS for dysarthric speech recognition,

    M. Soleymanpour, M. T. Johnson, R. Soleymanpour, and J. Berry, “Synthesizing dysarthric speech using multi-speaker TTS for dysarthric speech recognition,” in Proc. of Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, May 2022, pp. 7382–7386

  27. [35]

    Using HMM-based speech synthesis to reconstruct the voice of individuals with degenerative speech disorders,

    C. Veaux, J. Yamagishi, and S. King, “Using HMM-based speech synthesis to reconstruct the voice of individuals with degenerative speech disorders,” in Proc. of Annual Conference of the International Speech Communication , Portland, OR, USA, Sept. 2012, pp. 967–970

  28. [36]

    Speech-input speech-output communication for dysarthric speakers using HMM-based speech recognition and adaptive synthesis system,

    M. Dhanalakshmi, T. A. Mariya Celin, T. Nagarajan, and P. Vijayalakshmi, “Speech-input speech-output communication for dysarthric speakers using HMM-based speech recognition and adaptive synthesis system,” Circuits, Systems, and Signal Processing, vol. 37, no. 2, pp. 674–703, ...

  29. [37]

    Improving dysarthric speech intelligibility using cycle-consistent adversarial training,

    S. H. Yang and M. Chung, “Improving dysarthric speech intelligibility using cycle-consistent adversarial training,” in Proc. of International Conference on Bio-inspired Systems and Signal Processing, Valletta, Malta, Feb. 2020

  30. [38]

    An objective evaluation frame- work for pathological speech synthesis,

    B. M. Halpern, J. Fritsch, E. Hermann, R. van Son, O. Scharen- borg, and M. Magimai-Doss, “An objective evaluation frame- work for pathological speech synthesis,” in Proc. of Speech Communication; 14th ITG Conference , Kiel, Germany, Sep. 2021, pp. 1–5

  31. [39]

    Spectro- temporal representation of speech for intelligibility assessment of dysarthria,

    H. M. Chandrashekar, V . Karjigi, and N. Sreedevi, “Spectro- temporal representation of speech for intelligibility assessment of dysarthria,” IEEE Journal of Selected Topics in Signal Processing, vol. 14, no. 2, pp. 390–399, Feb. 2020

  32. [40]

    Improved speaker independent dysarthria intelligibility classification us- ing deepspeech posteriors,

    A. Tripathi, S. Bhosale, and S. K. Kopparapu, “Improved speaker independent dysarthria intelligibility classification us- ing deepspeech posteriors,” in Proc. of International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP) , Virtual Barcelona, Spain, May 2020...

  33. [41]

    On the design of automatic voice condition analysis systems. part ii: Review of speaker recognition techniques and study on the effects of different variability factors,

    J. A. G ´omez-Garc´ıa, L. Moro-Vel ´azquez, and J. I. Godino- Llorente, “On the design of automatic voice condition analysis systems. part ii: Review of speaker recognition techniques and study on the effects of different variability factors,”Biomedical Signal Processing and C...

  34. [42]

    Advances in Parkinson’s disease detection and assessment using voice and speech: A review of the articulatory and phonatory aspects,

    L. Moro-Velazquez, J. A. Gomez-Garcia, J. D. Arias-Londo ˜no, N. Dehak, and J. I. Godino-Llorente, “Advances in Parkinson’s disease detection and assessment using voice and speech: A review of the articulatory and phonatory aspects,” Biomedical Signal Processing and Control , ...

  35. [43]

    Speech based detection of Alzheimer’s disease: A survey of AI techniques, datasets and challenges,

    K. Ding, M. Chetty, A. Noori Hoshyar, T. Bhattacharya, and B. Klein, “Speech based detection of Alzheimer’s disease: A survey of AI techniques, datasets and challenges,” Artificial Intelligence Review, vol. 57, no. 12, p. 325, Oct. 2024

  36. [44]

    Quantifying articulatory impairments in neurode- generative motor diseases: A scoping review and meta-analysis of interpretable acoustic features,

    H. P. Rowe, S. Shellikeri, Y . Yunusova, K. V . Chenausky, and J. R. Green, “Quantifying articulatory impairments in neurode- generative motor diseases: A scoping review and meta-analysis of interpretable acoustic features,” International Journal of Speech-Language Pathology, ...

  37. [45]

    Innovative speech-based deep learning approaches for Parkinson’s disease classification: A systematic review,

    L. van Gelderen and C. Tejedor-Garc ´ıa, “Innovative speech-based deep learning approaches for Parkinson’s disease classification: A systematic review,” arXiv preprint arXiv:2407.17844, 2024

  38. [46]

    An intro- duction to machine learning for speech-language pathologists: Concepts, terminology, and emerging applications,

    C. Cordella, M. J. Marte, H. Liu, and S. Kiran, “An intro- duction to machine learning for speech-language pathologists: Concepts, terminology, and emerging applications,” Perspec- tives of the ASHA Special Interest Groups , pp. 1–19, 2024

  39. [47]

    Exploring the role of machine learning in diagnosing and treating speech disorders: A systematic literature review,

    Z. Brahmi, M. Mahyoob, M. Al-Sarem, J. Algaraady, K. Bous- selmi, and A. Alblwi, “Exploring the role of machine learning in diagnosing and treating speech disorders: A systematic literature review,” Psychology Research and Behavior Man- agement, pp. 2205–2232, May 2024

  40. [48]

    V oice disorder recognition using machine learning: A scoping review protocol,

    R. Gupta, D. R. Gunjawate, D. D. Nguyen, C. Jin, and C. Madill, “V oice disorder recognition using machine learning: A scoping review protocol,” BMJ Open , vol. 14, no. 2, Feb. 2024

  41. [49]

    V oice disorder detection using machine learning algorithms: An application in speech and language pathology,

    M. U. Rehman, A. Shafique, S. S. Jamal, Y . Gheraibia, A. B. Usman et al., “V oice disorder detection using machine learning algorithms: An application in speech and language pathology,” Engineering Applications of Artificial Intelligence, vol. 133, p. 108047, July 2024

  42. [50]

    Machine learning-and statistical-based voice analysis of Parkinson’s disease patients: A survey,

    F. Amato, G. Saggio, V . Cesarini, G. Olmo, and G. Costan- tini, “Machine learning-and statistical-based voice analysis of Parkinson’s disease patients: A survey,” Expert Systems with IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, VOL. XX, NO. XX, DECEMBER 2024 14 Appli...

  43. [51]

    Pathological speech processing: State-of-the- art, current challenges, and future directions,

    R. Gupta, T. Chaspari, J. Kim, N. Kumar, D. Bone, and S. Narayanan, “Pathological speech processing: State-of-the- art, current challenges, and future directions,” in Proc. of International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP) , Shanghai, China, Mar...

  44. [52]

    The TORGO database of acoustic and articulatory speech from speakers with dysarthria,

    F. Rudzicz, A. K. Namasivayam, and T. Wolff, “The TORGO database of acoustic and articulatory speech from speakers with dysarthria,” Language Resources and Evaluation, vol. 46, no. 4, pp. 523–541, Dec. 2012

  45. [53]

    New Spanish speech corpus database for the analysis of people suffering from Parkinson’s disease,

    J. R. Orozco-Arroyave, J. D. Arias-Londo ˜no, J. F. Vargas- Bonilla, M. C. Gonz ´alez-R´ativa, and E. N ¨oth, “New Spanish speech corpus database for the analysis of people suffering from Parkinson’s disease,” in Proc. of International Con- ference on Language Resources and Ev...

  46. [54]

    MoSpeeDi-ChaSpeePro dataset,

    “MoSpeeDi-ChaSpeePro dataset,” https://www.unige.ch/fapse/ mospeedi/mospeedi-dataset, database available within the MoSpeeDi-ChaSpeePro project consortium

  47. [55]

    The Nemours database of dysarthric speech,

    X. Menendez-Pidal, J. Polikoff, S. Peters, J. Leonzio, and H. Bunnell, “The Nemours database of dysarthric speech,” in Proc. of International Conference on Spoken Language Processing. ICSLP ’96 , vol. 3, Philadelphia, USA, oct. 1996, pp. 1962–1965 vol.3

  48. [56]

    Development of a Cantonese dysarthric speech corpus,

    K. H. Wong, Y . T. Yeung, E. H. Y . Chan, P. C. M. Wong, G.-A. Levow, and H. Meng, “Development of a Cantonese dysarthric speech corpus,” inProc. of Annual Conference of the International Speech Communication Association , Sept. 2015, pp. 329–333

  49. [57]

    The Atlanta Motor Speech Disorders Corpus: Motivation, Devel- opment, and Utility,

    J. Laures-Gore, S. Russell, R. Patel, and M. Frankel, “The Atlanta Motor Speech Disorders Corpus: Motivation, Devel- opment, and Utility,” Folia Phoniatrica et Logopaedica: Inter- national Association of Logopedics and Phoniatrics (IALP) , vol. 68, no. 2, pp. 99–105, 2016

  50. [58]

    A Dutch dysarthric speech database for individual- ized speech therapy research,

    E. Yilmaz, M. Ganzeboom, L. Beijer, C. Cucchiarini, and H. Strik, “A Dutch dysarthric speech database for individual- ized speech therapy research,” in Proc. of International Con- ference on Language Resources and Evaluation (LREC’16) , Portoroˇz, Slovenia, May 2016, pp. 792–795

  51. [59]

    EasyCall corpus: A dysarthric speech dataset,

    R. Turrisi, A. Braccia, M. Emanuele, S. Giulietti, M. Pugliatti, M. Sensi, L. Fadiga, and L. Badino, “EasyCall corpus: A dysarthric speech dataset,” in Proc. of Annual Conference of the International Speech Communication , Brno, Czech Republic, Sept. 2021, pp. 41–45

  52. [60]

    Dutch corpus of pathological and normal speech (COPAS),

    G. Van Nuffelen, M. De Bodt, C. Middag, and J.-P. Martens, “Dutch corpus of pathological and normal speech (COPAS),” Antwerp University Hospital and Ghent University, Tech. Rep., 2009

  53. [61]

    Unveiling early signs of Parkinson’s dis- ease via a longitudinal analysis of celebrity speech recordings,

    A. Favaro, A. Butala, T. Thebaud, J. Villalba, N. Dehak, and L. Moro-Vel´azquez, “Unveiling early signs of Parkinson’s dis- ease via a longitudinal analysis of celebrity speech recordings,” NPJ Parkinson’s Disease, vol. 10, no. 1, p. 207, 2024

  54. [62]

    Assessment of Speech Intelligibility in Parkinson’s Disease Using a Speech-To-Text System,

    G. Dimauro, V . Di Nicola, V . Bevilacqua, D. Caivano, and F. Girardi, “Assessment of Speech Intelligibility in Parkinson’s Disease Using a Speech-To-Text System,”IEEE Access, vol. 5, pp. 22 199–22 208, 2017

  55. [63]

    NeuroV oz: A Castillian Spanish corpus of Parkinsonian speech,

    J. Mendes-Laureano, J. A. G ´omez-Garc´ıa, A. Guerrero-L ´opez, E. Luque-Buzo, J. D. Arias-Londo ˜no, F. J. Grandas-P´erez, and J. I. Godino-Llorente, “NeuroV oz: A Castillian Spanish corpus of Parkinsonian speech,” Scientific Data , vol. 11, no. 1, p. 1367, Dec. 2024

  56. [64]

    Saarbr ¨ucken V oice Database,

    M. Putzer and W. Barry, “Saarbr ¨ucken V oice Database,” Institute of Phonetics, University of Saarland, 2021, accessed July 2025. [Online]. Available: https://stimmdb. coli.uni-saarland.de/

  57. [65]

    V oice pathology detection on the Saarbr¨ucken voice database with calibration and fusion of scores using multifocal toolkit,

    D. Mart ´ınez, E. Lleida, A. Ortega, A. Miguel, and J. Villalba, “V oice pathology detection on the Saarbr¨ucken voice database with calibration and fusion of scores using multifocal toolkit,” in Proc. of Advances in Speech and Language Technologies for Iberian Languages: Iber...

  58. [66]

    Collection and analysis of a Parkinson speech dataset with multiple types of sound recordings,

    B. E. Sakar, M. E. Isenkul, C. O. Sakar, A. Sertbas, F. Gur- gen, S. Delil, H. Apaydin, and O. Kursun, “Collection and analysis of a Parkinson speech dataset with multiple types of sound recordings,” IEEE Journal of Biomedical and Health Informatics, vol. 17, no. 4, pp. 828–83...

  59. [67]

    Acoustic tracking of pitch, modal, and subharmonic vibrations of vocal folds in Parkinson’s disease and Parkinsonism,

    J. Hlavni ˇcka, R. ˇCmejla, J. Klemp ´ıˇr, E. R ˚uˇziˇcka, and J. Rusz, “Acoustic tracking of pitch, modal, and subharmonic vibrations of vocal folds in Parkinson’s disease and Parkinsonism,” IEEE Access, vol. 7, pp. 150 339–150 354, Oct. 2019

  60. [68]

    Slovak database of speech affected by neurodegenerative diseases,

    M. Rusko, R. Sabo, M. Trnka, A. Zimmermann, R. Malaschitz, E. Ru ˇzick`y, P. Brandoburov´a, V . Kevick´a, and M. ˇSkorv´anek, “Slovak database of speech affected by neurodegenerative diseases,” Scientific Data, vol. 11, no. 1, pp. 1–16, Dec. 2024

  61. [69]

    The mPower study, Parkinson disease mobile data collected using researchkit,

    B. M. Bot, C. Suver, E. C. Neto, M. Kellen, A. Klein, C. Bare, M. Doerr, A. Pratap, J. Wilbanks, E. Dorsey et al. , “The mPower study, Parkinson disease mobile data collected using researchkit,” Scientific data, vol. 3, no. 1, pp. 1–9, March 2016

  62. [70]

    On using the UA- Speech and Torgo databases to validate automatic dysarthric speech classification approaches,

    G. Schu, P. Janbakhshi, and I. Kodrasi, “On using the UA- Speech and Torgo databases to validate automatic dysarthric speech classification approaches,” in Proc. of International Conference on Acoustics, Speech, and Signal Processing , Rhodes Island, Greece, June 2023, pp. 1–5

  63. [71]

    Suppressing noise disparity in train- ing data for automatic pathological speech detection,

    M. Amiri and I. Kodrasi, “Suppressing noise disparity in train- ing data for automatic pathological speech detection,” in Proc. of International Workshop on Acoustic Signal Enhancement , Aalborg, Denmark, Sept. 2024, pp. 110–114

  64. [72]

    V owel acoustics in dysarthria: Speech disorder diagnosis and classification,

    K. L. Lansford and J. M. Liss, “V owel acoustics in dysarthria: Speech disorder diagnosis and classification,” Journal of Speech, Language, and Hearing Research , vol. 57, pp. 57–67, Feb. 2014

  65. [73]

    Speech’s syllabic rhythm and articulatory features produced under different auditory feedback conditions identify Parkinsonism,

    ´A. Pi ˜na M ´endez, A. Taitz, O. Palacios Rodr ´ıguez, I. Rodr ´ıguez Leyva, and M. F. Assaneo, “Speech’s syllabic rhythm and articulatory features produced under different auditory feedback conditions identify Parkinsonism,” Scientific Reports, vol. 14, no. 1, p. 15787, July 2024

  66. [74]

    Acoustic characteristics of normal and patholog- ical voices,

    S. B. Davis, “Acoustic characteristics of normal and patholog- ical voices,” in Speech and Language . Elsevier, June 1979, vol. 1, pp. 271–335

  67. [75]

    V oiced/unvoiced transitions in speech as a potential bio- marker to detect Parkinson’s disease,

    J. R. Orozco-Arroyave, F. H ¨onig, J. D. Arias-Londo ˜no, J. F. Vargas-Bonilla, S. Skodda, J. Rusz, and E. N ¨oth, “V oiced/unvoiced transitions in speech as a potential bio- marker to detect Parkinson’s disease,” in Proc. of Annual Conference of the International Speech Commu...

  68. [76]

    Spectro-temporal sparsity charac- terization for dysarthric speech detection,

    I. Kodrasi and H. Bourlard, “Spectro-temporal sparsity charac- terization for dysarthric speech detection,” IEEE/ACM Trans- actions on Audio, Speech, and Language Processing , vol. 28, pp. 1210–1222, April 2020

  69. [77]

    Experimental investigation on STFT phase representations for deep learning-based dysarthric speech detection,

    P. Janbakhshi and I. Kodrasi, “Experimental investigation on STFT phase representations for deep learning-based dysarthric speech detection,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, May 2022, pp. 6477–6481

  70. [78]

    Raw speech waveform based classification of patients with ALS, Parkinson’s disease and healthy controls using CNN-BLSTM,

    J. Mallela, A. Illa, Y . Belur, N. Atchayaram, R. Yadav, P. Reddy, D. Gope, and P. K. Ghosh, “Raw speech waveform based classification of patients with ALS, Parkinson’s disease and healthy controls using CNN-BLSTM,” in Proc. of An- nual Conference of the International Speech C...

  71. [79]

    Wav2vec-based detection and severity level clas- sification of dysarthria from speech,

    F. Javanmardi, S. Tirronen, M. Kodali, S. R. Kadiri, and P. Alku, “Wav2vec-based detection and severity level clas- sification of dysarthria from speech,” in Proc. of Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, June 20...

  72. [80]

    Investigating self- supervised pretraining frameworks for pathological speech IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, VOL. XX, NO. XX, DECEMBER 2024 15 recognition,

    L. P. Violeta, W. C. Huang, and T. Toda, “Investigating self- supervised pretraining frameworks for pathological speech IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, VOL. XX, NO. XX, DECEMBER 2024 15 recognition,” in Proc. of Annual Conference of the Interna- tional Sp...

  73. [81]

    OpenSmile: The Munich versatile and fast open-source audio feature extractor,

    F. Eyben, M. W ¨ollmer, and B. Schuller, “OpenSmile: The Munich versatile and fast open-source audio feature extractor,” in Proc. of ACM International Conference on Multimedia , Feirenze, Italy, Oct. 2010, pp. 1459–1462

  74. [82]

    Dysarthric speech classification using glottal features computed from non-words, words and sentences,

    N. N. Prabhakera and P. Alku, “Dysarthric speech classification using glottal features computed from non-words, words and sentences,” in Proc. of Annual Conference of the International Speech Communication , Hyderabad, India, Sept. 2018, pp. 3403–3407

  75. [83]

    Learning to detect dysarthria from raw speech,

    J. Millet and N. Zeghidour, “Learning to detect dysarthria from raw speech,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Brighton, UK, May 2019, pp. 5831–5835

  76. [84]

    Dysarthric speech classification from coded telephone speech using glottal features,

    N. P. Narendra and P. Alku, “Dysarthric speech classification from coded telephone speech using glottal features,” Speech Communication, vol. 110, pp. 47–55, Jul. 2019

  77. [85]

    Recognising emotions in dysarthric speech using typical speech data,

    L. Alhinti, S. Cunningham, and H. Christensen, “Recognising emotions in dysarthric speech using typical speech data,” in Proc. of Annual Conference of the International Speech Communication, Shanghai, China, Oct. 2020

  78. [86]

    Auto- matic discrimination of apraxia of speech and dysarthria using a minimalistic set of handcrafted features,

    M. L. Ina Kodrasi, Michaela Pernon and H. Bourlard, “Auto- matic discrimination of apraxia of speech and dysarthria using a minimalistic set of handcrafted features,” in Proc. of An- nual Conference of the International Speech Communication , Shanghai, China, Oct. 2020, pp. 4991–4995

  79. [87]

    Automatic assessment of intelligi- bility in speakers with dysarthria from coded telephone speech using glottal features,

    N. P. Narendra and P. Alku, “Automatic assessment of intelligi- bility in speakers with dysarthria from coded telephone speech using glottal features,” Computer Speech & Language, vol. 65, p. 101117, Jan. 2021

  80. [88]

    Au- tomatic and perceptual discrimination between dysarthria, apraxia of speech, and neurotypical speech,

    I. Kodrasi, M. Pernon, M. Laganaro, and H. Bourlard, “Au- tomatic and perceptual discrimination between dysarthria, apraxia of speech, and neurotypical speech,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, Canada, June 2021...

  81. [89]

    Automatic speaker independent dysarthric speech intelligibility assess- ment system,

    A. Tripathi, S. Bhosale, and S. K. Kopparapu, “Automatic speaker independent dysarthric speech intelligibility assess- ment system,” Computer Speech & Language , vol. 69, p. 101213, Sep. 2021

  82. [90]

    Automated dysarthria severity classification: A study on acoustic features and deep learning techniques,

    A. A. Joshy and R. Rajan, “Automated dysarthria severity classification: A study on acoustic features and deep learning techniques,” IEEE Transactions on Neural Systems and Reha- bilitation Engineering, vol. 30, pp. 1147–1157, May 2022

  83. [91]

    Pre-trained models for detection and severity level classification of dysarthria from speech,

    F. Javanmardi, S. R. Kadiri, and P. Alku, “Pre-trained models for detection and severity level classification of dysarthria from speech,” Speech Communication , vol. 158, p. 103047, Mar. 2024

  84. [92]

    Impact of speech mode in automatic pathological speech detection,

    S. Shakeel, A. and K. Ina, “Impact of speech mode in automatic pathological speech detection,” in Proc. of European Signal Processing Conference, Lyon, France, Aug. 2024

  85. [93]

    Paralinguistic and spectral feature extrac- tion for speech emotion classification using machine learning techniques,

    T. Liu and X. Yuan, “Paralinguistic and spectral feature extrac- tion for speech emotion classification using machine learning techniques,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2023, no. 1, p. 23, May 2023

  86. [94]

    Use of global and acoustic features associated with contextual factors to adapt language models for spontaneous speech recognition

    S. Toyama, D. Saito, and N. Minematsu, “Use of global and acoustic features associated with contextual factors to adapt language models for spontaneous speech recognition.” in Proc. of Annual Conference of the International Speech Communication, Stockholm, Sweden, Aug. 2017, p...

  87. [95]

    Artificial neural net- works as speech recognisers for dysarthric speech: Identifying the best-performing set of MFCC parameters and studying a speaker-independent approach,

    S. R. Shahamiri and S. S. Binti Salim, “Artificial neural net- works as speech recognisers for dysarthric speech: Identifying the best-performing set of MFCC parameters and studying a speaker-independent approach,” Advanced Engineering Infor- matics, vol. 28, no. 1, pp. 102–11...

  88. [96]

    Recognition of dysarthric speech using voice parameters for speaker adap- tation and multi-taper spectral estimation,

    C. Bhat, B. Vachhani, and S. Kopparapu, “Recognition of dysarthric speech using voice parameters for speaker adap- tation and multi-taper spectral estimation,” in Proc. of Annual Conference of the International Speech Communication , San Francisco, USA, Sept. 2016, pp. 228–232

  89. [97]

    An automatic diagnosis and assessment of dysarthric speech using speech disorder specific prosodic features,

    G. Vyas, M. K. Dutta, J. Prinosil, and P. Har ´ar, “An automatic diagnosis and assessment of dysarthric speech using speech disorder specific prosodic features,” in Proc. of International Conference on Telecommunications and Signal Processing (TSP), Vienna, Austria, June 2016,...

  90. [98]

    Automatic recognition system for dysarthric speech based on MFCC’s, PNCC’s, jitter and shimmer coefficients,

    B.-F. Zaidi, M. Boudraa, S.-A. Selouani, D. Addou, and M. S. Yakoub, “Automatic recognition system for dysarthric speech based on MFCC’s, PNCC’s, jitter and shimmer coefficients,” in Proc. of Advances in Computer Vision, Advances in Intelligent Systems and Computing , vol. 944...

  91. [99]

    Classification of dysarthric speech according to the severity of impairment: An analysis of acoustic features,

    B. A. Al-Qatab and M. B. Mustafa, “Classification of dysarthric speech according to the severity of impairment: An analysis of acoustic features,” IEEE Access, vol. 9, pp. 18 183– 18 194, Feb. 2021

  92. [100]

    Deep neural network architectures for dysarthric speech analysis and recognition,

    B. F. Zaidi, S. A. Selouani, M. Boudraa, and M. Sidi Yak- oub, “Deep neural network architectures for dysarthric speech analysis and recognition,” Neural Computing and Applications, vol. 33, no. 15, pp. 9089–9108, Jan. 2021

  93. [101]

    Dysarthric speech recognition using multi-taper Mel frequency cepstrum coefficients,

    P. Sahane, S. Pangaonkar, and S. Khandekar, “Dysarthric speech recognition using multi-taper Mel frequency cepstrum coefficients,” in Proc. of International Conference on Comput- ing, Communication and Green Engineering (CCGE) , Pune, India, Sept. 2021, pp. 1–4

  94. [102]

    Analysis of short-time magnitude spectra for improving intelligibility assessment of dysarthric speech,

    L. P. Sahu and G. Pradhan, “Analysis of short-time magnitude spectra for improving intelligibility assessment of dysarthric speech,” Circuits, Systems, and Signal Processing , vol. 41, no. 10, pp. 5676–5698, Oct. 2022

  95. [103]

    Enhancing dysarthria detection: Harnessing ensemble models and MFCC,

    J. Jothieswari, T. Manicka Sundara Valli, and S. Suguna, “Enhancing dysarthria detection: Harnessing ensemble models and MFCC,” in Proc. of Smart Trends in Computing and Communications, Pune, India, 2024, pp. 135–147

  96. [104]

    Automatic detection of voice impairments by means of short-term cepstral param- eters and neural network based detectors,

    J. Godino-Llorente and P. Gomez-Vilda, “Automatic detection of voice impairments by means of short-term cepstral param- eters and neural network based detectors,” IEEE Transactions on Biomedical Engineering , vol. 51, no. 2, pp. 380–384, Feb 2004

  97. [105]

    Analysis of speaker recognition methodologies and the in- fluence of kinetic changes to automatically detect Parkinson’s disease,

    L. Moro-Vel ´azquez, J. A. G ´omez-Garc´ıa, J. I. Godino- Llorente, J. Villalba, J. R. Orozco-Arroyave, and N. Dehak, “Analysis of speaker recognition methodologies and the in- fluence of kinetic changes to automatically detect Parkinson’s disease,” Applied Soft Computing , vo...

  98. [106]

    Using X- Vectors to automatically detect Parkinson’s disease from speech,

    L. Moro-Vel ´azquez, J. Villalba, and N. Dehak, “Using X- Vectors to automatically detect Parkinson’s disease from speech,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Virtual Barcelona, Spain, May 2020, pp. 1155–1159

  99. [107]

    On combining information from modulation spectra and Mel-frequency cepstral coefficients for automatic detection of pathological voices,

    J. D. Arias-Londo ˜no, J. I. Godino-Llorente, M. Markaki, and Y . Stylianou, “On combining information from modulation spectra and Mel-frequency cepstral coefficients for automatic detection of pathological voices,” Logopedics Phoniatrics Vo- cology, vol. 36, no. 2, pp. 60–69,...

  100. [108]

    Analysis and detection of patho- logical voice using glottal source features,

    S. R. Kadiri and P. Alku, “Analysis and detection of patho- logical voice using glottal source features,” IEEE Journal of Selected Topics in Signal Processing , vol. 14, no. 2, pp. 367– 379, Dec. 2019

  101. [109]

    Statistical modeling of speech spectral coefficients in patients with Parkinson’s disease,

    I. Kodrasi and H. Bourlard, “Statistical modeling of speech spectral coefficients in patients with Parkinson’s disease,” in Proc. of Speech Communication; 13th ITG-Symposium , Oldenburg, Germany, Oct. 2018, pp. 1–5

  102. [110]

    Automatic pathological speech assessment,

    P. Janbakhshi, “Automatic pathological speech assessment,” Ph.D. dissertation, EPFL, June 2022

  103. [111]

    Dysarthric speech recognition using time-delay neural net- work based denoising autoencoder,

    C. Bhat, B. Das, B. Vachhani, and S. K. Kopparapu, “Dysarthric speech recognition using time-delay neural net- work based denoising autoencoder,” inProc. of Annual Confer- ence of the International Speech Communication , Hyderabad, IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PRO...

  104. [112]

    Whisper features for dysarthric severity level classification,

    S. Rathod, M. Charola, A. V ora, Y . Jogi, and H. A. Patil, “Whisper features for dysarthric severity level classification,” in Proc. of Annual Conference of the International Speech Communication, Dublin, Ireland, Aug. 2023, pp. 1523–1527

  105. [113]

    Syn- thesis of new words for improved dysarthric speech recognition on an expanded vocabulary,

    J. Harvill, D. Issa, M. Hasegawa-Johnson, and C. Yoo, “Syn- thesis of new words for improved dysarthric speech recognition on an expanded vocabulary,” in Proc. of International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP) , Tornonto, Canada, June 2021, pp. ...

  106. [114]

    Improving dysarthric speech segmentation with emulated and synthetic augmentation,

    S. A. Naeini, L. Simmatis, D. Jafari, Y . Yunusova, and B. Taati, “Improving dysarthric speech segmentation with emulated and synthetic augmentation,” IEEE Journal of Translational Engineering in Health and Medicine , vol. 12, pp. 382–389, March 2024

  107. [115]

    Sensitive quantification of cerebellar speech abnor- malities using deep learning models,

    K. Vattis, B. Oubre, A. C. Luddy, J. S. Ouillon, N. M. Eklund, C. D. Stephen, J. D. Schmahmann, A. S. Nunes, and A. S. Gupta, “Sensitive quantification of cerebellar speech abnor- malities using deep learning models,” IEEE Access , vol. 12, pp. 62 328–62 340, April 2024

  108. [116]

    Acoustic modelling from raw source and filter components for dysarthric speech recognition,

    Z. Yue, E. Loweimi, H. Christensen, J. Barker, and Z. Cvetkovic, “Acoustic modelling from raw source and filter components for dysarthric speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 2968–2980, Sept. 2022

  109. [117]

    Per- sonalizing TTS voices for progressive dysarthria,

    Y . Zhao, M. Song, Y . Yue, and M. Kuruvilla-Dugdale, “Per- sonalizing TTS voices for progressive dysarthria,” in Proc. of IEEE EMBS International Conference on Biomedical and Health Informatics (BHI) , Virtual Athens, Greece, July 2021, pp. 1–4

  110. [118]

    A sequential contrastive learning framework for robust dysarthric speech recognition,

    L. Wu, D. Zong, S. Sun, and J. Zhao, “A sequential contrastive learning framework for robust dysarthric speech recognition,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, Canada, June 2021, pp. 7303–7307

  111. [119]

    Raw source and filter modelling for dysarthric speech recognition,

    Z. Yue, E. Loweimi, and Z. Cvetkovic, “Raw source and filter modelling for dysarthric speech recognition,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, May 2022, pp. 7377–7381

  112. [120]

    Speaker adaptation using spectro-temporal deep features for dysarthric and elderly speech recognition,

    M. Geng, X. Xie, Z. Ye, T. Wang, G. Li, S. Hu, X. Liu, and H. Meng, “Speaker adaptation using spectro-temporal deep features for dysarthric and elderly speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, pp. 2597–2611, July 2022

  113. [121]

    Hybrid CNN-LSTM network to detect dysarthria using Mel-frequency cepstral coefficients,

    K. V ora, D. Padalia, D. Mehta, and D. Sharma, “Hybrid CNN-LSTM network to detect dysarthria using Mel-frequency cepstral coefficients,” in Proc. of International Conference on Advances in Science and Technology (ICAST), Mumbai, India, Dec. 2022, pp. 615–621

  114. [122]

    Convolutional neural network to model articulation impairments in patients with Parkinson’s disease

    J. C. V ´asquez-Correa et al. , “Convolutional neural network to model articulation impairments in patients with Parkinson’s disease.” in Proc. of Annual Conference of the International Speech Communication , Stockholm, Sweden, Aug. 2017, pp. 314–318

  115. [123]

    End-to-end Parkinson’s disease detection using a deep convolutional recurrent network,

    C. D. Rios-Urrego, S. A. Moreno-Acevedo, E. N ¨oth, and J. R. Orozco-Arroyave, “End-to-end Parkinson’s disease detection using a deep convolutional recurrent network,” in Proc. of International Conference on Text, Speech, and Dialogue, Brno, Czech Republic, Sept. 2022, pp. 326–338

  116. [124]

    The detection of Parkinson’s disease from speech using voice source informa- tion,

    N. Narendra, B. Schuller, and P. Alku, “The detection of Parkinson’s disease from speech using voice source informa- tion,” IEEE/ACM Transactions on Audio, Speech, and Lan- guage Processing, vol. 29, pp. 1925–1936, May 2021

  117. [125]

    Wav2vec 2.0: A framework for self-supervised learning of speech representations,

    A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “Wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. of Annual Conference on Neural Information Processing Systems, vol. 33, Virtual Online, Dec. 2020, pp. 12 449–12 460

  118. [126]

    HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

    W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhut- dinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Trans. Audio, Speech and Lang. Proc. , vol. 29, p. 3451–3460, Oct. 2021

  119. [127]

    WavLM: Large-scale self-supervised pre-training for full stack speech processing,

    S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selecte...

  120. [128]

    Perceiver-prompt: Flexible speaker adaptation in Whisper for Chinese disordered speech recognition,

    Y . Jiang, T. Wang, X. Xie, J. Liu, W. Sun, N. Yan, H. Chen, L. Wang, X. Liu, and F. Tian, “Perceiver-prompt: Flexible speaker adaptation in Whisper for Chinese disordered speech recognition,” in Proc. of Annual Conference of the Interna- tional Speech Communication , Kos, Gre...

  121. [129]

    Investigation of self-supervised pre-trained models for classification of voice quality from speech and neck surface accelerometer signals,

    S. R. Kadiri, F. Javanmardi, and P. Alku, “Investigation of self-supervised pre-trained models for classification of voice quality from speech and neck surface accelerometer signals,” Computer Speech & Language , vol. 83, p. 101550, Jan. 2024

  122. [130]

    Hierarchical multi- class classification of voice disorders using self-supervised models and glottal features,

    S. Tirronen, S. R. Kadiri, and P. Alku, “Hierarchical multi- class classification of voice disorders using self-supervised models and glottal features,” IEEE Open Journal of Signal Processing, vol. 4, pp. 80–88, Feb. 2023

  123. [131]

    Deep transfer learning for automatic speech recognition: Towards better generalization,

    H. Kheddar, Y . Himeur, S. Al-Maadeed, A. Amira, and F. Bensaali, “Deep transfer learning for automatic speech recognition: Towards better generalization,” Knowledge-Based Systems, vol. 277, p. 110851, Oct. 2023

  124. [132]

    Novel methods for detection and analysis of atypical aspects in speech,

    J. D. Fritsch, “Novel methods for detection and analysis of atypical aspects in speech,” Ph.D. dissertation, EPFL, Lau- sanne, May 2023

  125. [133]

    Automatic voice disorder detection using self- supervised representations,

    D. Ribas, M. A. Pastor, A. Miguel, D. Mart ´ınez, A. Ortega, and E. Lleida, “Automatic voice disorder detection using self- supervised representations,” IEEE Access, vol. 11, pp. 14 915– 14 927, Feb. 2023

  126. [134]

    On-the-fly feature based rapid speaker adaptation for dysarthric and elderly speech recognition,

    M. Geng, X. Xie, R. Su, J. Yu, Z. Jin, T. Wang, S. Hu, Z. Ye, H. Meng, and X. Liu, “On-the-fly feature based rapid speaker adaptation for dysarthric and elderly speech recognition,” in Proc. of Annual Conference of the International Speech Communication, Dublin, Ireland, Aug. 2022

  127. [135]

    Inclusive ASR for disfluent speech: Cascaded large-scale self-supervised learning with targeted fine-tuning and data augmentation,

    D. Mujtaba, N. R. Mahapatra, M. Arney, J. S. Yaruss, C. Her- ring, and J. Bin, “Inclusive ASR for disfluent speech: Cascaded large-scale self-supervised learning with targeted fine-tuning and data augmentation,” in Proc. of Annual Conference of the International Speech Communi...

  128. [136]

    Domain and language adaptation using heterogeneous datasets for Wav2vec2.0-based speech recognition of low-resource lan- guage,

    K. Soky, S. Li, C. Chu, and T. Kawahara, “Domain and language adaptation using heterogeneous datasets for Wav2vec2.0-based speech recognition of low-resource lan- guage,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Rhodes Island, ...

  129. [137]

    Towards objective and interpretable speech disorder assessment: A comparative analysis of CNN and transformer-based models,

    M. Maisonneuve, C. Fredouille, M. Lalain, A. Ghio, and V . Woisard, “Towards objective and interpretable speech disorder assessment: A comparative analysis of CNN and transformer-based models,” in Proc. of Annual Conference of the International Speech Communication , Kos, Gree...

  130. [138]

    Impact of including pathological speech in pre-training on pathology detection,

    T. Weise, A. Maier, K. C. Demir, P. A. P ´erez-Toro, T. Arias- Vergara, B. Heismann, E. N ¨oth, M. Schuster, and S. H. Yang, “Impact of including pathological speech in pre-training on pathology detection,” in Proc. of Text, Speech, and Dialogue , Pilsen, Czech Republic, Sept....

  131. [139]

    Automatic assessment of dysarthria using audio-visual vowel graph attention network,

    X. Liu, X. Du, J. Liu, R. Su, M. L. Ng, Y . Zhang, Y . Yang, S. Zhao, L. Wang, and N. Yan, “Automatic assessment of dysarthria using audio-visual vowel graph attention network,” arXiv preprint arXiv:2405.03254 , 2024

  132. [140]

    Kheirkhahzadeh, “Speech classification using acoustic em- bedding and large language models applied on Alzheimer’s IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, VOL

    M. Kheirkhahzadeh, “Speech classification using acoustic em- bedding and large language models applied on Alzheimer’s IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, VOL. XX, NO. XX, DECEMBER 2024 17 disease prediction task,” Ph.D. dissertation, KTH, Sweden, Aug. 2023

  133. [141]

    Speaker adaptation for Wav2vec2 based dysarthric ASR,

    M. K. Baskar, T. Herzig, D. Nguyen, M. Diez, T. Polzehl, L. Burget, and J. H. ˇCernock´y, “Speaker adaptation for Wav2vec2 based dysarthric ASR,” in Proc. of Annual Con- ference of the International Speech Communication , Incheon, South Korea, Sept. 2022

  134. [142]

    Improving dysarthric speech recognition by enrich- ing training datasets,

    S. Cullen, “Improving dysarthric speech recognition by enrich- ing training datasets,” Ph.D. dissertation, TU Dublin, Ireland, March 2022

  135. [143]

    Exploring pathological speech quality assessment with ASR-powered Wav2Vec2 in data-scarce context,

    T. Nguyen, C. Fredouille, A. Ghio, M. Balaguer, and V . Wois- ard, “Exploring pathological speech quality assessment with ASR-powered Wav2Vec2 in data-scarce context,” in Proc. of Joint International Conference on Computational Linguistics, Language Resources and Evaluation, T...

  136. [144]

    CFDRN: A cognition- inspired feature decomposition and recombination network for dysarthric speech recognition,

    Y . Lin, L. Wang, Y . Yang, and J. Dang, “CFDRN: A cognition- inspired feature decomposition and recombination network for dysarthric speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 3824– 3836, Sept. 2023

  137. [145]

    Benefits of pre-trained mono- and cross-lingual speech representations for spoken language understanding of Dutch dysarthric speech,

    P. Wang and H. Van hamme, “Benefits of pre-trained mono- and cross-lingual speech representations for spoken language understanding of Dutch dysarthric speech,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2023, no. 1, p. 15, Apr. 2023

  138. [146]

    Exploring self-supervised pre-trained ASR models for dysarthric and elderly speech recognition,

    S. Hu, X. Xie, Z. Jin, M. Geng, Y . Wang, M. Cui, J. Deng, X. Liu, and H. Meng, “Exploring self-supervised pre-trained ASR models for dysarthric and elderly speech recognition,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Is...

  139. [147]

    Enhancing pre-trained ASR system fine-tuning for dysarthric speech recognition using adversarial data aug- mentation,

    H. Wang, Z. Jin, M. Geng, S. Hu, G. Li, T. Wang, H. Xu, and X. Liu, “Enhancing pre-trained ASR system fine-tuning for dysarthric speech recognition using adversarial data aug- mentation,” in Proc. of International Conference on Acoustics, Speech, and Signal Processing (ICASSP)...

  140. [148]

    Cross-lingual self- supervised speech representations for improved dysarthric speech recognition,

    A. Hernandez, P. A. P ´erez-Toro, E. N ¨oth, J. R. Orozco- Arroyave, A. Maier, and S. H. Yang, “Cross-lingual self- supervised speech representations for improved dysarthric speech recognition,” in Proc. of Annual Conference of the International Speech Communication , Incheon,...

  141. [149]

    Using voice conversion and time-stretching to enhance the quality of dysarthric speech for automatic speech recognition,

    M. Spijkerman, “Using voice conversion and time-stretching to enhance the quality of dysarthric speech for automatic speech recognition,” Ph.D. dissertation, University of Gronin- gen, Netherlands, July 2022

  142. [150]

    TRILLsson: Distilled Universal Paralinguistic Speech Representations,

    J. Shor and S. Venugopalan, “TRILLsson: Distilled Universal Paralinguistic Speech Representations,” in Proc. of Annual Conference of the International Speech Communication , In- cheon, Korea, Sep. 2022, pp. 356–360

  143. [151]

    Interpretable speech fea- tures vs. DNN embeddings: What to use in the automatic assessment of Parkinson’s disease in multi-lingual scenarios,

    A. Favaro, Y .-T. Tsai, A. Butala, T. Thebaud, J. Villalba, N. Dehak, and L. Moro-Vel ´azquez, “Interpretable speech fea- tures vs. DNN embeddings: What to use in the automatic assessment of Parkinson’s disease in multi-lingual scenarios,” Computers in Biology and Medicine, vo...

  144. [152]

    Novel speech signal processing algorithms for high-accuracy classification of Parkinson’s disease,

    A. Tsanas, M. A. Little, P. E. McSharry, J. Spielman, and L. O. Ramig, “Novel speech signal processing algorithms for high-accuracy classification of Parkinson’s disease,” IEEE Transactions on Biomedical Engineering , vol. 59, no. 5, pp. 1264–1271, May 2012

  145. [153]

    Detecting Parkinson’s disease from sustained phonation and speech signals,

    E. Vaiciukynas, A. Verikas, A. Gelzinis, and M. Bacauskiene, “Detecting Parkinson’s disease from sustained phonation and speech signals,” PLoS One , vol. 12, no. 10, pp. 1–16, Oct. 2017

  146. [154]

    Automatic system to detect the type of voice pathology,

    S. Jothilakshmi, “Automatic system to detect the type of voice pathology,” Applied Soft Computing , vol. 21, pp. 244–249, Aug. 2014

  147. [155]

    Spectral and cepstral analy- ses for Parkinson’s disease detection in Spanish vowels and words,

    J. R. Orozco-Arroyave, F. H ¨onig, J. D. Arias-Londo ˜no, J. F. Vargas-Bonilla, and E. N ¨oth, “Spectral and cepstral analy- ses for Parkinson’s disease detection in Spanish vowels and words,” Expert Systems , vol. 32, no. 6, pp. 688–697, Dec. 2015

  148. [156]

    Automatic evaluation of articulatory disorders in Parkinson’s disease,

    M. Novotn ´y, J. Rusz, R. ˇCmejla, and E. R ˚uˇziˇcka, “Automatic evaluation of articulatory disorders in Parkinson’s disease,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 22, no. 9, pp. 1366–1378, Sept. 2014

  149. [157]

    Towards automatic detection of Amyotrophic Lateral Sclerosis from speech acoustic and articulatory samples,

    J. Wang, P. V . Kothalkar, B. Cao, and D. Heitzman, “Towards automatic detection of Amyotrophic Lateral Sclerosis from speech acoustic and articulatory samples,” in Proc. of Annual Conference of the International Speech Communication , San Francisco, USA, Sept. 2016, pp. 1195–1199

  150. [158]

    Detection of Amyotrophic Lateral Sclerosis (ALS) via acoustic analysis,

    R. Norel, M. Pietrowicz, C. Agurto, S. Rishoni, and G. Cec- chi, “Detection of Amyotrophic Lateral Sclerosis (ALS) via acoustic analysis,” in Proc. of Annual Conference of the International Speech Communication,, Hyderabad, India, Sept. 2018, pp. 377–381

  151. [159]

    Subspace-based learning for automatic dysarthric speech detection,

    P. Janbakhshi, I. Kodrasi, and H. Bourlard, “Subspace-based learning for automatic dysarthric speech detection,” IEEE Signal Processing Letters , vol. 28, pp. 96–100, Dec. 2021

  152. [160]

    Cross-database models for the classification of dysarthria presence,

    S. Gillespie, Y .-Y . Logan, E. Moore, J. Laures-Gore, S. Russell, and R. Patel, “Cross-database models for the classification of dysarthria presence,” in Proc. of Annual Conference of the International Speech Communication, Stockholm, Sweden, Aug. 2017, pp. 3127–3131

  153. [161]

    Dysarthric-speech detection using transfer learning with convolutional neural networks,

    S. R. Mani Sekhar, G. Kashyap, A. Bhansali, A. A. Andrew, and K. Singh, “Dysarthric-speech detection using transfer learning with convolutional neural networks,” Information & Communications Technology Express, vol. 8, no. 1, pp. 61–64, March 2022

  154. [162]

    A multitask learning approach to assess the dysarthria severity in patients with Parkinson’s disease,

    J. C. V ´asquez Correa, T. Arias, J. R. Orozco-Arroyave, and E. N ¨oth, “A multitask learning approach to assess the dysarthria severity in patients with Parkinson’s disease,” in Proc. of Annual Conference of the International Speech Com- munication, Hyderabad, India, Sept. 20...

  155. [163]

    Diagnosing dysarthria with long short- term memory networks

    A. Mayle et al. , “Diagnosing dysarthria with long short- term memory networks.” in Proc. of Annual Conference of the International Speech Communication , Graz, Austria, Sept. 2019, pp. 4514–4518

  156. [164]

    Automatic assessment of sentence- level dysarthria intelligibility using BLSTM,

    C. Bhat and H. Strik, “Automatic assessment of sentence- level dysarthria intelligibility using BLSTM,” IEEE Journal of Selected Topics in Signal Processing , vol. 14, no. 2, pp. 322–330, Feb. 2020

  157. [165]

    Exploring the impact of fine-tuning the wav2vec2 model in database-independent detection of dysarthric speech,

    F. Javanmardi, S. R. Kadiri, and P. Alku, “Exploring the impact of fine-tuning the wav2vec2 model in database-independent detection of dysarthric speech,” IEEE Journal of Biomedical and Health Informatics , vol. 28, no. 8, pp. 4951–4962, April 2024

  158. [166]

    Two-step acoustic model adaptation for dysarthric speech recognition,

    R. Takashima, T. Takiguchi, and Y . Ariki, “Two-step acoustic model adaptation for dysarthric speech recognition,” in Proc. of International Conference on Acoustics, Speech and Signal Processing, Virtual Barcelona, Spain, May 2020, pp. 6104– 6108

  159. [167]

    Dysarthric speech recognition based on deep metric learning,

    Y . Takashima, R. Takashima, T. Takiguchi, and Y . Ariki, “Dysarthric speech recognition based on deep metric learning,” in Proc. of Annual Conference of the International Speech Communication, Shanghai, China, Oct. 2020, pp. 4796–4800

  160. [168]

    Automatic speech recognition of disordered speech: Personalized models outper- forming human listeners on short phrases,

    J. R. Green, R. L. MacDonald et al. , “Automatic speech recognition of disordered speech: Personalized models outper- forming human listeners on short phrases,” in Proc. of Annual Conference of the International Speech Communication , Brno, Czech Republic, Sept. 2021, pp. 4778–4782

  161. [169]

    Dysarthric speech recog- nition with lattice-free MMI,

    E. Hermann and M. Magimai.-Doss, “Dysarthric speech recog- nition with lattice-free MMI,” in Proc. of International Con- ference on Acoustics, Speech and Signal Processing (ICASSP), Virtual Barcelona, Spain, May 2020, pp. 6109–6113. IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PR...

  162. [170]

    Source domain data selection for improved transfer learning targeting dysarthric speech recognition,

    F. Xiong, J. Barker, Z. Yue, and H. Christensen, “Source domain data selection for improved transfer learning targeting dysarthric speech recognition,” in Proc. of International Con- ference on Acoustics, Speech and Signal Processing (ICASSP), Virtual Barcelona, Spain, May 202...

  163. [171]

    Autoencoder bottleneck features with multi-task optimisation for improved continuous dysarthric speech recognition,

    Z. Yue, H. Christensen, and J. Barker, “Autoencoder bottleneck features with multi-task optimisation for improved continuous dysarthric speech recognition,” in Proc. of Annual Conference of the International Speech Communication , Shanghai, China, Oct. 2020, pp. 4581–4585

  164. [172]

    Speech vision: An end-to-end deep learning- based dysarthric automatic speech recognition system,

    S. R. Shahamiri, “Speech vision: An end-to-end deep learning- based dysarthric automatic speech recognition system,” IEEE Transactions on Neural Systems and Rehabilitation Engineer- ing, vol. 29, pp. 852–861, May 2021

  165. [173]

    LoRA: Low-rank adaptation of large language models,

    E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in Proc. of International Conference on Learning Representations, Virtual, Apr. 2022

  166. [174]

    LoRA-whisper: Parameter-efficient and extensible multilin- gual ASR,

    Z. Song, J. Zhuo, Y . Yang, Z. Ma, S. Zhang, and X. Chen, “LoRA-whisper: Parameter-efficient and extensible multilin- gual ASR,” in Proc. of Annual Conference of the International Speech Communication , Kos, Greece, Sept. 2024, pp. 3934– 3938

  167. [175]

    Adjusting dysarthric speech signals to be more intelligible,

    F. Rudzicz, “Adjusting dysarthric speech signals to be more intelligible,” Computer Speech & Language , vol. 27, no. 6, pp. 1163–1177, Sep. 2013

  168. [176]

    Intelligibility of modifications to dysarthric speech,

    J.-P. Hosom, A. Kain, T. Mishra, J. van Santen, M. Fried-Oken, and J. Staehely, “Intelligibility of modifications to dysarthric speech,” in Proc. of International Conference on Acoustics, Speech, and Signal Processing, (ICASSP) , vol. 1, Hong Kong, China, Apr. 2003, pp. I–I

  169. [177]

    A Kepstrum based approach for enhancement of dysarthric speech,

    V . Lalitha, P. Prema, and L. Mathew, “A Kepstrum based approach for enhancement of dysarthric speech,” in Proc. of International Congress on Image and Signal Processing , vol. 7, Yantai, China, Oct. 2010, pp. 3474–3478

  170. [178]

    Intelligibility mod- ification of dysarthric speech using HMM-based adaptive synthesis system,

    M. Dhanalakshmi and P. Vijayalakshmi, “Intelligibility mod- ification of dysarthric speech using HMM-based adaptive synthesis system,” in Proc. of International Conference on Biomedical Engineering (ICoBE) , Penang, Malaysia, Mar. 2015, pp. 1–5

  171. [179]

    Dysarthric speech enhancement based on convolution neural network,

    S. Wang, Y . Tsao, W. Zheng, H. Yeh, P. Li, S. Fang, and Y . Lai, “Dysarthric speech enhancement based on convolution neural network,” in Proc. of Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC) , Glasgow, Scotland, UK, July 2022, pp. 60–64

  172. [180]

    Improving the intelligibility of dysarthric speech towards enhancing the effectiveness of speech therapy,

    S. A. Kumar and C. S. Kumar, “Improving the intelligibility of dysarthric speech towards enhancing the effectiveness of speech therapy,” in Proc. of International Conference on Advances in Computing, Communications and Informatics (ICACCI), Jaipur, India, Sep. 2016, pp. 1000–1005

  173. [181]

    Towards improving the intelligibility of dysarthric speech,

    A. Roy, L. Thakur, G. Vyas, and G. Raj, “Towards improving the intelligibility of dysarthric speech,” in Proc. of Interna- tional Conference on Soft Computing and Signal Processing , vol. 898, Hyderabad, India, Feb. 2019, pp. 547–559

  174. [182]

    Intelligibility improvement of dysarthric speech using MMSE DiscoGAN,

    M. Purohit, M. Patel, H. Malaviya et al. , “Intelligibility improvement of dysarthric speech using MMSE DiscoGAN,” in Proc. of International Conference on Signal Processing and Communications (SPCOM) , Bangalore, India, July 2020, pp. 1–5

  175. [183]

    The effectiveness of time stretching for enhancing dysarthric speech for improved dysarthric speech recognition,

    L. Prananta, B. M. Halpern, S. Feng, and O. Scharenborg, “The effectiveness of time stretching for enhancing dysarthric speech for improved dysarthric speech recognition,” in Proc. of Annual Conference of the International Speech Communi- cation, Incheon, Korea, Sept. 2022, pp. 36–40

  176. [184]

    High-intelligibility speech synthesis for dysarthric speakers with LPCNet-based TTS and CycleV AE-based VC,

    K. Matsubara, T. Okamoto, R. Takashima, T. Takiguchi, T. Toda, Y . Shiga, and H. Kawai, “High-intelligibility speech synthesis for dysarthric speakers with LPCNet-based TTS and CycleV AE-based VC,” inProc. of International Conference on Acoustics, Speech and Signal Processing ...

  177. [185]

    End-to-end voice conversion via cross- modal knowledge distillation for dysarthric speech recon- struction,

    D. Wang, J. Yu et al., “End-to-end voice conversion via cross- modal knowledge distillation for dysarthric speech recon- struction,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Virtual, Barcelona, Spain, May. 2020, pp. 7744–7748

  178. [186]

    UNIT- DSR: Dysarthric speech reconstruction system using speech unit normalization,

    Y . Wang, X. Wu, D. Wang, L. Meng, and H. Meng, “UNIT- DSR: Dysarthric speech reconstruction system using speech unit normalization,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Seoul, Korea, Apr. 2024, pp. 12 306–12 310

  179. [187]

    Automated dysarthria severity classification for improved objective intelligibility assessment of spastic dysarthric speech,

    M. S. Paja and T. H. Falk, “Automated dysarthria severity classification for improved objective intelligibility assessment of spastic dysarthric speech,” in Proc. of Annual Conference of the International Speech Communication,, Portland, OR, USA, Sept. 2012, pp. 62–65

  180. [188]

    Dysarthria in- telligibility assessment in a factor analysis total variability space,

    D. Mart ´ınez, P. Green, and H. Christensen, “Dysarthria in- telligibility assessment in a factor analysis total variability space,” in Proc. of Annual Conference of the International Speech Communication Association, Lyon, France, Aug. 2013, pp. 2133–2137

  181. [189]

    Speech intelligibility estimation using multi-resolution spectral features for speakers undergoing cancer treatment,

    J. C. Kim, H. Rao, and M. A. Clements, “Speech intelligibility estimation using multi-resolution spectral features for speakers undergoing cancer treatment,” The Journal of the Acoustical Society of America , vol. 136, no. 4, pp. 315–321, Oct. 2014

  182. [190]

    Spectral fea- tures for automatic blind intelligibility estimation of spastic dysarthric speech,

    R. Hummel, W.-Y . Chan, and T. H. Falk, “Spectral fea- tures for automatic blind intelligibility estimation of spastic dysarthric speech,” in Proc. of Annual Conference of the International Speech Communication Association , Florence, Italy, Aug. 2011, pp. 3017–3020

  183. [191]

    Characterization of atypical vocal source excitation, temporal dynamics and prosody for objective measurement of dysarthric word intelli- gibility,

    T. H. Falk, W.-Y . Chan, and F. Shein, “Characterization of atypical vocal source excitation, temporal dynamics and prosody for objective measurement of dysarthric word intelli- gibility,” Speech Communication, vol. 54, no. 5, pp. 622–631, June 2012

  184. [192]

    Robust automatic eval- uation of intelligibility in voice rehabilitation using prosodic analysis,

    T. Haderlein, A. Sch ¨utzenberger et al., “Robust automatic eval- uation of intelligibility in voice rehabilitation using prosodic analysis,” in Proc. of International Conference on Text, Speech, and Dialogue,, Prague, Czech Republic, Aug. 2017, pp. 11–19

  185. [193]

    Predicting intelligibility gains in dysarthria through automated speech feature analysis,

    A. R. Fletcher, A. A. Wisler, M. J. McAuliffe, K. L. Lansford, and J. M. Liss, “Predicting intelligibility gains in dysarthria through automated speech feature analysis,” Journal of Speech, Language, and Hearing Research , vol. 60, no. 11, pp. 3058– 3068, Nov. 2017

  186. [194]

    Automatic recognition and evaluation of tracheoe- sophageal speech,

    T. Haderlein, S. Steidl, E. N ¨oth, F. Rosanowski, and M. Schus- ter, “Automatic recognition and evaluation of tracheoe- sophageal speech,” in Proc. of International Conference on Text, Speech and Dialogue, Brno, Czech Republic, Sept. 2004, pp. 331–338

  187. [195]

    Objective intelligibility assessment of pathological speakers,

    C. Middag, G. Van Nuffelen, J. P. Martens, and M. De Bodt, “Objective intelligibility assessment of pathological speakers,” in Proc. of Annual Conference of the International Speech Communication Association , Brisbane, Australia, Sept. 2008, pp. 1745–1748

  188. [196]

    Automated intelligibility assessment of pathological speech using phonological features,

    C. Middag, J.-P. Martens, G. V . Nuffelen, and M. De Bodt, “Automated intelligibility assessment of pathological speech using phonological features,” EURASIP Journal on Advances in Signal Processing , vol. 2009, no. 1, pp. 1–9, May 2009

  189. [197]

    Towards an ASR-free objective analysis of pathological speech,

    C. Middag, Y . Saeys, and J.-P. Martens, “Towards an ASR-free objective analysis of pathological speech,” in Proc. of Annual Conference of the International Speech Communication Asso- ciation, Makuhari, Chiba, Japan, Sept. 2010, pp. 294–297

  190. [198]

    PEAKS - A system for the automatic evaluation of voice and speech disorders,

    A. Maier, T. Haderlein, U. Eysholdt, F. Rosanowski, A. Bat- liner, M. Schuster, and E. N ¨oth, “PEAKS - A system for the automatic evaluation of voice and speech disorders,” Speech Communication, vol. 51, no. 5, pp. 425–437, May 2009

  191. [199]

    Speech technology-based assessment of phoneme intelligi- bility in dysarthria,

    G. V . Nuffelen, C. Middag, M. De Bodt, and J.-P. Martens, “Speech technology-based assessment of phoneme intelligi- bility in dysarthria,” International Journal of Language & IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, VOL. XX, NO. XX, DECEMBER 2024 19 Communication...

  192. [200]

    Automatic intelligibility assessment of speakers after laryn- geal cancer by means of acoustic modeling,

    T. Bocklet, K. Riedhammer, U. Eysholdt, and T. Haderlein, “Automatic intelligibility assessment of speakers after laryn- geal cancer by means of acoustic modeling,” Journal of Voice, vol. 26, no. 3, pp. 390–397, May 2012

  193. [201]

    Intelligibility assessment and speech recog- nizer word accuracy rate prediction for dysarthric speakers in a factor analysis subspace,

    D. Mart ´ınez, E. Lleida, P. Green, H. Christensen, A. Ortega, and A. Miguel, “Intelligibility assessment and speech recog- nizer word accuracy rate prediction for dysarthric speakers in a factor analysis subspace,” ACM Transactions on Accessible Computing, vol. 6, no. 3, pp. ...

  194. [202]

    Automatic prediction of speech evaluation metrics for dysarthric speech,

    L. Imed, B. K. Waad, F. Corinne, and M. Christine, “Automatic prediction of speech evaluation metrics for dysarthric speech,” in Proc. of Annual Conference of the International Speech Communication Association, Stockholm, Sweden, Aug. 2017, pp. 1834–1838

  195. [203]

    Intel- ligibility assessment of cleft lip and palate speech using Gaus- sian posteriograms based on joint spectro-temporal features,

    S. Kalita, S. R. Mahadeva Prasanna, and S. Dandapat, “Intel- ligibility assessment of cleft lip and palate speech using Gaus- sian posteriograms based on joint spectro-temporal features,” Journal of the Acoustical Society of America , vol. 144, no. 4, pp. 2413–2423, Oct. 2018

  196. [204]

    Pathological speech intelligibility assessment based on the short-time ob- jective intelligibility measure,

    P. Janbakhshi, I. Kodrasi, and H. Bourlard, “Pathological speech intelligibility assessment based on the short-time ob- jective intelligibility measure,” in Proc. of International Con- ference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK, May 2019, pp. 6405–6409

  197. [205]

    Synthetic speech references for automatic patholog- ical speech intelligibility assessment,

    ——, “Synthetic speech references for automatic patholog- ical speech intelligibility assessment,” in Proc. of Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP), Virtual Barcelona, Spain, 2020, pp. 6099–6103

  198. [206]

    Automatic pathological speech intelligibility assess- ment exploiting subspace-based analyses,

    ——, “Automatic pathological speech intelligibility assess- ment exploiting subspace-based analyses,” IEEE/ACM Trans- actions on Audio, Speech, and Language Processing , vol. 28, pp. 1717–1728, 2020

  199. [207]

    Automatic severity clas- sification of Korean dysarthric speech using phoneme-level pronunciation features,

    E. J. Yeo, S. Kim, and M. Chung, “Automatic severity clas- sification of Korean dysarthric speech using phoneme-level pronunciation features,” in Proc. of Annual Conference of the International Speech Communication, , Brno, Czech Republic, Sept. 2021, pp. 4838–4842

  200. [208]

    Fully automated speaker identification and intelligibility as- sessment in dysarthria disease using auditory knowledge,

    K. L. Kadi, S. A. Selouani, B. Boudraa, and M. Boudraa, “Fully automated speaker identification and intelligibility as- sessment in dysarthria disease using auditory knowledge,” Biocybernetics and Biomedical Engineering , vol. 36, no. 1, pp. 233–247, Jan. 2016

  201. [209]

    Precision medicine in ALS: Identification of new acoustic markers for dysarthria severity assessment,

    R. Dubbioso, M. Spisto, L. Verde, V . V . Iuzzolino, G. Sener- chia, G. De Pietro, I. De Falco, and G. Sannino, “Precision medicine in ALS: Identification of new acoustic markers for dysarthria severity assessment,” Biomedical Signal Processing and Control, vol. 89, p. 105706,...

  202. [210]

    Increasing the precision of dysarthric speech intelligibility and severity level estimate,

    M. Soleymanpour, M. T. Johnson, and J. Berry, “Increasing the precision of dysarthric speech intelligibility and severity level estimate,” Lecture Notes in Computer Science , vol. 12997, pp. 670–679, 2021

  203. [211]

    Automated dysarthria severity classification using deep learning frameworks,

    A. A. Joshy and R. Rajan, “Automated dysarthria severity classification using deep learning frameworks,” in Proc. of European Signal Processing Conference (EUSIPCO), Amster- dam, Netherlands, Aug. 2021, pp. 116–120

  204. [212]

    Residual neural network precisely quantifies dysarthria severity-level based on short-duration speech segments,

    S. Gupta, A. T. Patil, M. Purohit, M. Parmar, M. Patel, H. A. Patil, and R. C. Guido, “Residual neural network precisely quantifies dysarthria severity-level based on short-duration speech segments,” Neural Networks , vol. 139, pp. 105–117, Jul. 2021

  205. [213]

    Dysarthria severity assessment using squeeze-and-excitation networks,

    A. A. Joshy and R. Rajan, “Dysarthria severity assessment using squeeze-and-excitation networks,” Biomedical Signal Processing and Control, vol. 82, p. 104606, Apr. 2023

  206. [214]

    Variable STFT layered CNN model for automated dysarthria detection and severity assessment using raw speech,

    K. Radha, M. Bansal, and V . R. Dulipalla, “Variable STFT layered CNN model for automated dysarthria detection and severity assessment using raw speech,” Circuits, Systems, and Signal Processing, vol. 43, no. 5, pp. 3261–3278, May 2024

  207. [215]

    Automatic dysarthria detection and severity level assessment using CWT-layered CNN model,

    S. Sajiha, K. Radha, D. Venkata Rao, N. Sneha, S. Gunnam, and D. P. Bavirisetti, “Automatic dysarthria detection and severity level assessment using CWT-layered CNN model,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2024, no. 1, p. 33, June 2024

  208. [216]

    Perceptually enhanced single frequency filtering for dysarthric speech detection and intelligibility assessment,

    K. Gurugubelli and A. K. Vuppala, “Perceptually enhanced single frequency filtering for dysarthric speech detection and intelligibility assessment,” in Proc. of International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP) , Brighton, UK, May 2019, pp. 6410–6414

  209. [217]

    Towards an automatic evaluation of the dysarthria level of patients with parkinson’s disease,

    J. C. V ´asquez-Correa, J. R. Orozco-Arroyave, T. Bocklet, and E. N ¨oth, “Towards an automatic evaluation of the dysarthria level of patients with parkinson’s disease,” Journal of Com- munication Disorders, vol. 76, pp. 21–36, Nov. 2018

  210. [218]

    Dysarthria detection and severity assessment using rhythm-based met- rics,

    A. Hernandez, E. J. Yeo, S. Kim, and M. Chung, “Dysarthria detection and severity assessment using rhythm-based met- rics,” in Proc. of Annual Conference of the International Speech Communication , Shanghai, China, Oct. 2020, pp. 2897–2901

  211. [219]

    Prosody-based mea- sures for automatic severity assessment of dysarthric speech,

    A. Hernandez, S. Kim, and M. Chung, “Prosody-based mea- sures for automatic severity assessment of dysarthric speech,” Applied Sciences, vol. 10, no. 19, p. 6999, Jan. 2020

  212. [220]

    Automatic assessment of dysarthric severity level using audio-video cross- modal approach in deep learning,

    H. Tong, H. Sharifzadeh, and I. McLoughlin, “Automatic assessment of dysarthric severity level using audio-video cross- modal approach in deep learning,” in Proc. of Annual Confer- ence of the International Speech Communication , Shanghai, China, Oct. 2020, pp. 4786–4790

  213. [221]

    Dysarthria severity classification using multi-head attention and multi-task learning,

    A. A. Joshy and R. Rajan, “Dysarthria severity classification using multi-head attention and multi-task learning,” Speech Communication, vol. 147, pp. 1–11, Feb. 2023

  214. [222]

    End-to-end dysarthric speech recognition using multiple databases,

    Y . Takashima, T. Takiguchi, and Y . Ariki, “End-to-end dysarthric speech recognition using multiple databases,” in Proc. of International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Brighton, UK, May 2019, pp. 6395–6399

  215. [223]

    Data augmenta- tion using healthy speech for dysarthric speech recognition,

    B. Vachhani, C. Bhat, and S. K. Kopparapu, “Data augmenta- tion using healthy speech for dysarthric speech recognition,” in Proc. of Annual Conference of the International Speech Communication, Hyderabad, India, Sept. 2018, pp. 471–475

  216. [224]

    Investigation of data augmentation techniques for disordered speech recognition,

    M. Geng, X. Xie, S. Liu, J. Yu, S. Hu, X. Liu, and H. Meng, “Investigation of data augmentation techniques for disordered speech recognition,” in Proc. of Annual Conference of the International Speech Communication , Shanghai, China, Oct 2020, pp. 696–700

  217. [225]

    Phonetic analysis of dysarthric speech tempo and applications to robust per- sonalised dysarthric speech recognition,

    F. Xiong, J. Barker, and H. Christensen, “Phonetic analysis of dysarthric speech tempo and applications to robust per- sonalised dysarthric speech recognition,” in Proc. of Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK, May. 2019,...

  218. [226]

    Training data augmentation for dysarthric automatic speech recognition by text-to-dysarthric-speech synthesis,

    W.-Z. Leung, M. Cross, A. Ragni, and S. Goetze, “Training data augmentation for dysarthric automatic speech recognition by text-to-dysarthric-speech synthesis,” in Proc. of Annual Conference of the International Speech Communication , Kos, Greece, June 2024, pp. 2494–2498

  219. [227]

    Few-shot dysarthric speech recognition with text-to-speech data augmentation,

    E. Hermann and M. Magimai. Doss, “Few-shot dysarthric speech recognition with text-to-speech data augmentation,” in Proc. of Annual Conference of the International Speech Communication, Dublin, Ireland, Aug. 2023, pp. 156–160

  220. [228]

    Adversarial data augmentation for disordered speech recog- nition,

    Z. Jin, M. Geng, X. Xie, J. Yu, S. Liu, X. Liu, and H. Meng, “Adversarial data augmentation for disordered speech recog- nition,” in Proc. of Annual Conference of the International Speech Communication, , Brno, Czech Republic, Sept. 2021, pp. 4803–4807

  221. [229]

    Personalized adversarial data augmentation for dysarthric and elderly speech recognition,

    Z. Jin, M. Geng, J. Deng, T. Wang, S. Hu, G. Li, and X. Liu, “Personalized adversarial data augmentation for dysarthric and elderly speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 32, pp. 413– 429, Oct. 2024

  222. [230]

    V oxTester, software for digital evaluation of speech changes in Parkinson disease,

    G. Dimauro, D. Caivano, V . Bevilacqua, F. Girardi, and IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, VOL. XX, NO. XX, DECEMBER 2024 20 V . Napoletano, “V oxTester, software for digital evaluation of speech changes in Parkinson disease,” in Proc. of Interna- tional Sym...

  223. [231]

    Graph neural networks for Parkinsons disease detection,

    S. A. Sheikh, Y . Kaloga, and I. Kodrasi, “Graph neural networks for Parkinsons disease detection,” inProc. of Interna- tional Conference on Acoustics, Speech and Signal Processing, Hyderabad, India, April 2024, pp. 1–5

  224. [232]

    Multiview canon- ical correlation analysis for automatic pathological speech detection,

    Y . Kaloga, S. A. Sheikh, and I. Kodrasi, “Multiview canon- ical correlation analysis for automatic pathological speech detection,” in Proc. of International Conference on Acoustics, Speech and Signal Processing , Hyderabad, India, April 2024, pp. 1–5

  225. [233]

    Noise robust dysarthric speech classification using domain adaptation,

    A. Wisler, V . Berisha, A. Spanias, and J. Liss, “Noise robust dysarthric speech classification using domain adaptation,” in Proc. of Digital Media Industry & Academic Forum (DMIAF), Santorini, Greece, July 2016, pp. 135–138

  226. [234]

    Test-time adaptation for automatic pathological speech detection in noisy environments,

    M. Amiri and I. Kodrasi, “Test-time adaptation for automatic pathological speech detection in noisy environments,” in Proc. of Europenan Signal Porcessing Conference , Lyon, France, Aug. 2024, pp. 86–90

  227. [235]

    Towards a corpus (and language)-independent screening of Parkinson’s disease from voice and speech through domain adaptation,

    E. J. Ibarra, J. D. Arias-Londo ˜no, M. Za˜nartu, and J. I. Godino- Llorente, “Towards a corpus (and language)-independent screening of Parkinson’s disease from voice and speech through domain adaptation,” Bioengineering, vol. 10, no. 11, p. 1316, 2023

  228. [236]

    Adversarial robustness analysis in automatic pathological speech detection approaches,

    M. Amiri and I. Kodrasi, “Adversarial robustness analysis in automatic pathological speech detection approaches,” in Proc. of Annual Conference of the International Speech Communi- cation, Rhodes Islands, Greece, Sept. 2024, pp. 1415–1419

  229. [237]

    Learn- ing audio-visual speech representation by masked multimodal cluster prediction,

    B. Shi, W.-N. Hsu, K. Lakhotia, and A. Mohamed, “Learn- ing audio-visual speech representation by masked multimodal cluster prediction,” in Proc. of International Conference on Learning Representations, Virtual, April 2022

  230. [238]

    Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead,

    C. Rudin, “Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead,” Nature Machine Intelligence, vol. 1, no. 5, pp. 206–215, May 2019

  231. [239]

    Interpretable phonological fea- tures for clinical applications,

    Y . Jiao, V . Berisha, and J. Liss, “Interpretable phonological fea- tures for clinical applications,” in Proc. of International Con- ference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA, Mar. 2017, pp. 5045–5049

  232. [240]

    Acoustic-based articulatory phenotypes of Amyotrophic Lat- eral Sclerosis and Parkinson’s disease: Towards an inter- pretable, hypothesis-driven framework of motor control,

    H. P. Rowe, S. E. Gutz, M. F. Maffei, and J. R. Green, “Acoustic-based articulatory phenotypes of Amyotrophic Lat- eral Sclerosis and Parkinson’s disease: Towards an inter- pretable, hypothesis-driven framework of motor control,” in Proc. of Annual Conference of the Internatio...

  233. [241]

    Interpretable dysarthric speaker adaptation based on optimal-transport,

    R. Turrisi and L. Badino, “Interpretable dysarthric speaker adaptation based on optimal-transport,” in Proc. of Annual Conference of the International Speech Communication , In- cheon, Korea, Sep. 2022, pp. 26–30

  234. [242]

    Unveiling interpretability in self- supervised speech representations for Parkinson’s diagnosis,

    D. Gimeno-G ´omez, C. Botelho, A. Pompili, A. Abad, and C.-D. Mart ´ınez-Hinarejos, “Unveiling interpretability in self- supervised speech representations for Parkinson’s diagnosis,” IEEE Journal of Selected Topics in Signal Processing , vol. 1, pp. 1–14, Feb. 2025

  235. [243]

    Dysarthria detection based on a deep learning model with a clinically-interpretable layer,

    L. Xu, J. Liss, and V . Berisha, “Dysarthria detection based on a deep learning model with a clinically-interpretable layer,” JASA Express Letters, vol. 3, no. 1, p. 015201, Jan. 2023

  236. [244]

    Large language models: A comprehensive survey of its applications, challenges, limitations, and future prospects,

    M. U. Hadi, Q. A. Tashi, A. Shah, R. Qureshi, A. Muneer, M. Irfan, A. Zafar, M. B. Shaikh, N. Akhtar, J. Wu, S. Mirjalili, and M. Shah, “Large language models: A comprehensive survey of its applications, challenges, limitations, and future prospects,” Authorea Preprints, Aug. 2024

  237. [245]

    Large language models for dysfluency detec- tion in stuttered speech,

    D. Wagner, S. P. Bayerl, I. Baumann, K. Riedhammer, E. N¨oth, and T. Bocklet, “Large language models for dysfluency detec- tion in stuttered speech,” in Proc. of Annual Conference of the International Speech Communication , Kos, Greece, Sept. 2024, pp. 5118–5122. Shakeel A. Sh...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.