REVIEW 1 major objections 42 references
Continuous-Speech Parkinson's Disease Detection Using Acoustic and Inharmonicity Features
T0 review · 1 major / 0 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read Continuous speech models detect Parkinson's disease more effectively than sustained vowel models across two datasets.
desk verdict The paper reports that continuous-speech models with acoustic and inharmonicity features beat sustained-vowel baselines on two datasets, but the abstract supplies no performance numbers, dataset sizes, or test details. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The continuous-speech PD classifier that fuses conventional acoustic representations with an inharmonicity-based feature framework, evaluated under speaker-level partitioning to avoid leakage.
What would settle it
A new dataset experiment in which the continuous-speech model shows no accuracy gain over the sustained-vowel model when the same speaker-level and leakage controls are applied.
Extended reading notes
Core claim
The proposed continuous-speech framework for Parkinson's disease identification outperforms the best sustained-vowel model on two distinct datasets; the added inharmonicity features improve results on one dataset but produce no significant change on the other.
Load-bearing premise
That the speaker-level evaluation and leakage-prevention steps produce a fair, unbiased comparison between continuous-speech and sustained-vowel models.
Editorial extensions
If this is right
- Continuous speech enables practical background monitoring of voice changes without requiring controlled phonations.
- Inharmonicity features supply complementary information that can raise performance on at least some datasets when added to acoustic features.
- Speaker-level evaluation and leakage prevention are required to produce trustworthy comparisons between speech types.
- Vowel content can be extracted from continuous recordings in a manner that supports downstream classification.
Reading between the lines
- If continuous-speech superiority holds across more datasets, screening protocols could move from clinic-controlled vowels to natural conversation recordings.
- The dataset-dependent effect of inharmonicity points to the need for studies that isolate which voice traits make the feature helpful.
- Integration with mobile devices could allow passive, real-world tracking of vocal changes before clinical symptoms appear.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces a Parkinson's disease (PD) detection framework for continuous speech that combines traditional acoustic features with a novel inharmonicity representation. Using two distinct datasets, it compares the proposed continuous-speech pipeline against the best-performing sustained-vowel baseline and reports superior performance for the continuous-speech approach. The work also examines speaker-level cross-validation, data-leakage safeguards, and automatic extraction of vowel segments from running speech.
Significance. If the reported performance advantage is shown to be robust under a fully leakage-free, speaker-disjoint protocol with matched hyper-parameter search, the result would be significant: it would support the shift from controlled sustained-vowel recordings to ecologically valid continuous-speech monitoring for PD. The cautious finding that inharmonicity features improve one dataset but are neutral on the other is also useful, as it highlights the need for further validation before claiming general complementarity.
major comments (1)
- [Abstract and §3] Abstract and §3 (evaluation protocol): the central claim that the continuous-speech model 'clearly illustrates the preferential performance' over the best sustained-vowel model rests on an empirical comparison whose validity cannot be assessed from the supplied information. No details are given on (a) whether identical speakers appear in the train/test folds for both tasks, (b) the precise automatic method used to isolate vowels from continuous speech and whether its errors systematically favor the continuous-speech condition, or (c) whether the 'best' sustained-vowel model was selected under the identical hyper-parameter search used for the continuous-speech model. Any of these three conditions being incompletely satisfied would render the reported performance gap non-interpretable.
Simulated Author's Rebuttal
We thank the referee for the constructive and detailed feedback. The major comment concerns the transparency of the evaluation protocol; we address each sub-point below and will incorporate clarifications to ensure the comparison is fully interpretable.
read point-by-point responses
-
Referee: [Abstract and §3] Abstract and §3 (evaluation protocol): the central claim that the continuous-speech model 'clearly illustrates the preferential performance' over the best sustained-vowel model rests on an empirical comparison whose validity cannot be assessed from the supplied information. No details are given on (a) whether identical speakers appear in the train/test folds for both tasks, (b) the precise automatic method used to isolate vowels from continuous speech and whether its errors systematically favor the continuous-speech condition, or (c) whether the 'best' sustained-vowel model was selected under the identical hyper-parameter search used for the continuous-speech model. Any of these three conditions being incompletely satisfied would render the reported performance gap non-interpretable.
Authors: We appreciate the referee highlighting these critical aspects of reproducibility. (a) The manuscript already specifies speaker-disjoint, speaker-level cross-validation for both tasks and states that the same speaker partitioning is used across conditions to eliminate leakage; we will add an explicit sentence confirming identical folds were applied to the sustained-vowel and continuous-speech pipelines. (b) Section 3 describes the automatic vowel-segment extraction procedure from running speech. While extraction errors are possible, the sustained-vowel baseline also relies on vowel material, and we will add a short error-analysis paragraph quantifying extraction accuracy on held-out data to address potential systematic bias. (c) Hyper-parameter selection for both models was performed with the identical grid search and the same speaker-disjoint CV folds; we will make this explicit in the revised §3. These additions will be included in the revision. revision: yes
Circularity Check
Empirical ML comparison on two datasets; no derivation chain present
full rationale
The paper reports an empirical classification study that extracts acoustic and inharmonicity features from continuous speech and sustained vowels, then compares model performance via speaker-level cross-validation on two independent datasets. The central claim is a performance advantage for the continuous-speech pipeline; this rests on measured accuracy/F1 numbers rather than any algebraic derivation, parameter fit renamed as prediction, or self-referential definition. No equations are offered that reduce to their own inputs, and the abstract explicitly frames the work as data-driven comparison with leakage controls. Consequently the result is self-contained against external benchmarks and carries no circularity.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Continuous-Speech Parkinson's Disease Detection Using Acoustic and Inharmonicity Features." pith.science (2026). https://pith.science/paper/BSXAHI7Y
@misc{pith2026260619125,
author = {Pith},
title = {Pith review of: Continuous-Speech Parkinson's Disease Detection Using Acoustic and Inharmonicity Features},
year = {2026},
howpublished = {\url{https://pith.science/paper/BSXAHI7Y}},
note = {Machine review of arXiv:2606.19125}
}
read the original abstract
Notable efforts have been made to identify Parkinson's disease (PD) from vocal data, primarily using sustained vowel phonations. In this work, we extend on these efforts introducing a PD identification approach for continuous speech, enabling a practical background monitoring of voice data to detect vocal changes indicative of PD. Using two distinct data sets, we compare the best sustained vowel model with that of the proposed continuous speech model, clearly illustrating the preferential performance of the latter. We examine approaches for speaker level evaluation and data leakage preventions, as well as how vowel information may be reliable extracted from continuous speech. The proposed method framework exploits both traditional acoustic representations and a promising novel inharmonicity based framework, showing how the latter provides complementary information improving the performance for one of the data sets; however, for the other data set, this information did not significantly improve (nor reduce) the performance, suggesting that further studies are required before being able to draw firm conclusions in its use. Overall, the work clearly illustrates the benefit of forming PD classification using continuous speech compared to using sustained vowel sounds.
Figures
Reference graph
Works this paper leans on
-
[1]
Detecting Parkinson’s Disease Using V oice Recordings From Mobile Devices
Momeni N, Whitling S, Jakobsson A. Detecting Parkinson’s Disease Using V oice Recordings From Mobile Devices. In: Proceedings of the 32nd European Signal Processing Conference; 2024. p. 1516-20
2024
-
[2]
Parkinson’s Disease [web page]
World Health Organization. Parkinson’s Disease [web page]
-
[3]
Accessed: 5 February 2024
Available from: https://www.who.int/news-room/fact-sheets/detail/ parkinson-disease. Accessed: 5 February 2024
2024
-
[4]
Speech disorders in Parkinson’s disease: early diagnostics and effects of medication and brain stimulation
Brabenec L, Mekyska J, Galaz Z, Rektorova I. Speech disorders in Parkinson’s disease: early diagnostics and effects of medication and brain stimulation. Journal of Neural Transmission. 2017;124(3):303-34
2017
-
[5]
Hypokinetic Dysarthria in Parkin- son’s Disease: A Narrative Review
Atalar MS, Oguz O, Genc G. Hypokinetic Dysarthria in Parkin- son’s Disease: A Narrative Review. Sisli Etfal Hastanesi Tip Bulteni. 2023;57(2):163-70
2023
-
[6]
Speech and language biomarkers for Parkinson’s disease prediction, early diagnosis and pro- gression
Cao F, V ogel AP, Gharahkhani P, Renteria ME. Speech and language biomarkers for Parkinson’s disease prediction, early diagnosis and pro- gression. npj Parkinson’s Disease. 2025;11:57
2025
-
[7]
Parkinson’s disease
Kalia LV , Lang AE. Parkinson’s disease. The Lancet. 2015;386(9996):896-912
2015
-
[8]
Early treatment of Parkinson’s disease: Opportunities for managed care
Murman DL. Early treatment of Parkinson’s disease: Opportunities for managed care. The American Journal of Managed Care. 2012;18(7 Suppl):S183-8. LIet al.:CONTINUOUS-SPEECH PARKINSON’S DISEASE DETECTION USING ACOUSTIC AND INHARMONICITY FEATURES 11 TABLE A-I RETAINEDPERSON-LEVELINHARMONICITYDESCRIPTORSAFTERSPARSE FEATURESELECTION Data set Retained descript...
2012
Show all 42 references
-
[9]
Machine learning for the diagnosis of Parkinson’s disease: A review of literature
Mei J, Desrosiers C, Frasnelli J. Machine learning for the diagnosis of Parkinson’s disease: A review of literature. Frontiers in Aging Neuroscience. 2021;13:633752
2021
-
[10]
Advances in Parkinson’s disease detection and as- sessment using voice and speech: A review of the articulatory and phona- tory aspects
Moro-Vel ´azquez L, G ´omez-Garc´ıa JA, Arias-Londo ˜no JD, Dehak N, Godino-Llorente JI. Advances in Parkinson’s disease detection and as- sessment using voice and speech: A review of the articulatory and phona- tory aspects. Biomedical Signal Processing and Control. 2021;66:102418
2021
-
[11]
Early detection of Parkinson’s disease through speech features and machine learning: A review
Gullapalli AS, Mittal VK. Early detection of Parkinson’s disease through speech features and machine learning: A review. In: ICT with Intelligent Applications. vol. 248 of Smart Innovation, Systems and Technologies. Singapore: Springer; 2022
2022
-
[12]
Interpretable Parkinson’s Disease Detection Using Group-Wise Scaling
Momeni N, Whitling S, Jakobsson A. Interpretable Parkinson’s Disease Detection Using Group-Wise Scaling. IEEE Access. 2025;13:29147-61
2025
-
[13]
Use of machine learning and voice for multi- class classification of Parkinson’s disease, chronic obstructive pulmonary disease, and healthy controls
Idrisoglu A, Behrens A. Use of machine learning and voice for multi- class classification of Parkinson’s disease, chronic obstructive pulmonary disease, and healthy controls. Scientific Reports. 2026;16:15485
2026
-
[14]
Automated analysis of connected speech reveals early biomarkers of Parkinson’s disease in patients with rapid eye movement sleep behaviour disorder
Hlavni ˇcka J, ˇCmejla R, Tykalov ´a T, ˇSonka K, R ˚uˇziˇcka E, Rusz J. Automated analysis of connected speech reveals early biomarkers of Parkinson’s disease in patients with rapid eye movement sleep behaviour disorder. Scientific Reports. 2017;7:12
2017
-
[15]
Detecting Parkinson’s disease from sustained phonation and speech signals
Vaiciukynas E, Verikas A, Gelzinis A, Bacauskiene M. Detecting Parkinson’s disease from sustained phonation and speech signals. PLOS ONE. 2017;12(10):e0185613
2017
-
[16]
Things to Consider When Automatically Detecting Parkinson’s Disease Using the Phonation of Sustained V owels: Analysis of Methodological Issues
Ozbolt AS, Moro-Vel ´azquez L, Lina I, Butala AA, Dehak N. Things to Consider When Automatically Detecting Parkinson’s Disease Using the Phonation of Sustained V owels: Analysis of Methodological Issues. Applied Sciences. 2022;12(3):991
2022
-
[17]
Automatic Detection of Parkin- son’s Disease from Continuous Speech Recorded in Non-Controlled Noise Conditions
V ´asquez-Correa JC, Arias-Vergara T, Orozco-Arroyave JR, Vargas- Bonilla JF, Arias-Londo ˜no JD, N ¨oth E. Automatic Detection of Parkin- son’s Disease from Continuous Speech Recorded in Non-Controlled Noise Conditions. In: Proceedings of Interspeech 2015; 2015. p. 105-9
2015
-
[18]
Automatic Detection of Parkinson’s Disease in Running Speech Spoken in Three Different Languages
Orozco-Arroyave JR, H ¨onig F, Arias-Londo ˜no JD, Vargas-Bonilla JF, Daqrouq K, Skodda S, et al. Automatic Detection of Parkinson’s Disease in Running Speech Spoken in Three Different Languages. The Journal of the Acoustical Society of America. 2016;139(1):481-500
2016
-
[19]
Parkinson’s Disease Classification Framework Using V ocal Dynamics in Connected Speech
Appakaya SB, Pratihar R, Sankar R. Parkinson’s Disease Classification Framework Using V ocal Dynamics in Connected Speech. Algorithms. 2023;16(11):509
2023
-
[20]
CNN-Based Identification of Parkinson’s Disease from Continuous Speech in Noisy Environments
Farag ´o P, S ¸tef˘anig˘a SA, Cordos ¸ CG, Mih˘ail˘a LI, Hintea S, Pes ¸tean AS, et al. CNN-Based Identification of Parkinson’s Disease from Continuous Speech in Noisy Environments. Bioengineering. 2023;10(5):531
2023
-
[21]
Evaluating the Effectiveness of Pre-Trained Audio Embeddings for Classification of Parkinson’s Disease Speech Data
Postma E, Tejedor-Garcia C. Evaluating the Effectiveness of Pre-Trained Audio Embeddings for Classification of Parkinson’s Disease Speech Data. In: Proceedings of Interspeech; 2025. p. 4603-7
2025
-
[22]
Has machine learning over- promised in healthcare?: A critical analysis and a proposal for improved evaluation, with evidence from Parkinson’s disease
Ge W, Lueck C, Suominen H, Apthorp D. Has machine learning over- promised in healthcare?: A critical analysis and a proposal for improved evaluation, with evidence from Parkinson’s disease. Artificial Intelligence in Medicine. 2023;139:102524
2023
-
[23]
NeuroV oz: a Castillian Spanish corpus of parkinsonian speech
Mendes-Laureano J, G ´omez-Garc´ıa JA, Guerrero-L´opez A, Luque-Buzo E, Arias-Londo ˜no JD, Grandas-P ´erez FJ, et al. NeuroV oz: a Castillian Spanish corpus of parkinsonian speech. Scientific Data. 2024;11:1367
2024
-
[24]
openSMILE: The Munich Versatile and Fast Open-Source Audio Feature Extractor
Eyben F, W ¨ollmer M, Schuller B. openSMILE: The Munich Versatile and Fast Open-Source Audio Feature Extractor. In: Proceedings of the 18th ACM International Conference on Multimedia; 2010. p. 1459-62
2010
-
[25]
The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for V oice Research and Affective Computing
Eyben F, Scherer KR, Schuller BW, Sundberg J, Andr ´e E, Busso C, et al. The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for V oice Research and Affective Computing. IEEE Transactions on Affective Computing. 2016;7(2):190-202
2016
-
[26]
Statistics; 2024
Parkinson’s Foundation. Statistics; 2024. Accessed: 5 February 2024. Available from: https://www.parkinson.org/understanding-parkinsons/ statistics
2024
-
[27]
How does Parkinson’s disease affect quality of life? A comparison with quality of life in the general population
Schrag A, Jahanshahi M, Quinn N. How does Parkinson’s disease affect quality of life? A comparison with quality of life in the general population. Movement Disorders. 2000 Nov;15(6):1112-8
2000
-
[28]
Defining Fundamental Frequency for Al- most Harmonic Signals
Elvander F, Jakobsson A. Defining Fundamental Frequency for Al- most Harmonic Signals. IEEE Transactions on Signal Processing. 2020;68:6453-66
2020
-
[29]
Estimating Inharmonic Signals with Optimal Transport Priors
Elvander F. Estimating Inharmonic Signals with Optimal Transport Priors. In: 2023 IEEE International Conference on Acoustics, Speech and Signal Processing; 2023. p. 1-5
2023
-
[30]
V osk Speech Recognition Toolkit [web page]; 2024
Alpha Cephei. V osk Speech Recognition Toolkit [web page]; 2024. Available from: https://alphacephei.com/vosk/. Accessed: 24 May 2026
2024
-
[31]
Introduction to Speech Processing: 2nd Edition
B ¨ackstr¨om T, R ¨as¨anen O, Zewoudie A, P ´erez Zarazaga P, Koivusalo L, Das S, et al. Introduction to Speech Processing: 2nd Edition. Zenodo; 2022
2022
-
[32]
Harmonic to Noise Ratio Measurement – Selection of Window and Length
Fernandes J, Teixeira F, Guedes V , Junior A, Teixeira JP. Harmonic to Noise Ratio Measurement – Selection of Window and Length. Procedia Computer Science. 2018;138:280-5
2018
-
[33]
Multi-Pitch Estimation
Christensen MG, Jakobsson A. Multi-Pitch Estimation. vol. 5 of Synthe- sis Lectures on Speech and Audio Processing. San Rafael, CA: Morgan & Claypool Publishers; 2009
2009
-
[34]
Multiple Fundamental Frequency Estimation Based on Harmonicity and Spectral Smoothness
Klapuri AP. Multiple Fundamental Frequency Estimation Based on Harmonicity and Spectral Smoothness. IEEE Transactions on Speech and Audio Processing. 2003 Nov;11(6):804-16
2003
-
[35]
Optimizing Parkinson’s Disease Prediction: A Comparative Analysis of Data Aggregation Methods Using Multiple V oice Recordings via an Automated Artificial Intelligence Pipeline
Yang Z, Zhou H, Srivastav S, Shaffer JG, Abraham KE, Naandam SM, et al. Optimizing Parkinson’s Disease Prediction: A Comparative Analysis of Data Aggregation Methods Using Multiple V oice Recordings via an Automated Artificial Intelligence Pipeline. Data. 2025;10(1):4
2025
-
[36]
Things to Consider When Automatically Detecting Parkinson’s Disease Using the Phonation of Sustained V owels: Analysis of Methodological Issues
Ozbolt AS, Moro-Velazquez L, Lina I, Butala AA, Dehak N. Things to Consider When Automatically Detecting Parkinson’s Disease Using the Phonation of Sustained V owels: Analysis of Methodological Issues. Applied Sciences. 2022;12(3):991
2022
-
[37]
Reliable AI via Age-Balanced Validation: Fair Model Selection for Parkinson’s Detection from V oice
Momeni N, Whitling S, Jakobsson A. Reliable AI via Age-Balanced Validation: Fair Model Selection for Parkinson’s Detection from V oice. In: ICASSP 2026 – 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE; 2026. p. 3031-5
2026
-
[38]
XGBoost: A Scalable Tree Boosting System
Chen T, Guestrin C. XGBoost: A Scalable Tree Boosting System. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; 2016. p. 785-94
2016
-
[39]
Regularization and Variable Selection via the Elastic Net
Zou H, Hastie T. Regularization and Variable Selection via the Elastic Net. Journal of the Royal Statistical Society: Series B. 2005;67(2):301- 20
2005
-
[40]
Regularization Paths for Generalized Linear Models via Coordinate Descent
Friedman J, Hastie T, Tibshirani R. Regularization Paths for Generalized Linear Models via Coordinate Descent. Journal of Statistical Software. 2010;33(1):1-22
2010
-
[41]
Small Area Shrinkage Estimation
Datta G, Ghosh M. Small Area Shrinkage Estimation. Statistical Science. 2012;27(1):95-114
2012
-
[42]
On Combining Classi- fiers
Kittler J, Hatef M, Duin RPW, Matas J. On Combining Classi- fiers. IEEE Transactions on Pattern Analysis and Machine Intelligence. 1998;20(3):226-39
1998
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.