REVIEW 4 major objections 5 minor 70 references
LSEAD: A Privacy-Preserving LLM-Based Speech Analysis Framework for Early Alzheimer's Disease Screening
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read LSEAD claims that a fully local, text-only pipeline—Whisper transcription, Zephyr-7B-β mean-pooled embeddings, PCA, and logistic regression—reaches 90.0% accuracy on the combined ADReSS20 and ADReSSo2021 Alzheimer's speech test set…
desk verdict Useful incremental engineering, but the 90% headline is an unblinded test-set selection bound; the cross-dataset study is the more solid part. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the combination of attention-mask-aware mean pooling of the penultimate-layer hidden states of Zephyr-7B-β, PCA dimensionality reduction, and logistic regression. Mean pooling turns variable-length transcripts into a fixed vector while ignoring padding tokens; PCA removes redundant dimensions so a linear classifier can find separating directions; logistic regression makes the final AD-versus-CN decision. The paper argues that this specific chain, rather than the largest model, is what yields the reported accuracy, since removing PCA degrades all classifiers and both Llama 2-7B and Qwen3-30B underperform Zephyr-7B-β in the same pipeline.
What would settle it
Run an acoustic-only classifier using prosody, pause structure, and voice quality on the same combined ADReSS20 and ADReSSo2021 test set; if acoustics alone matches or exceeds the reported 90.0% accuracy, the paper's text-only premise is false. Alternatively, report a 95% confidence interval for the 119-sample test accuracy; a confidence interval spanning the 84.9% comparator would mean the headline margin is not statistically established.
Extended reading notes
Core claim
The central claim is that diagnostic information in Alzheimer's speech is recoverable from the lexical-semantic content of automatically transcribed text, and that a modest open-source LLM embedding plus PCA plus a linear classifier extracts it better than larger or cloud-based models. On the combined benchmark, logistic regression achieves 90.0% accuracy (precision 91.2%, recall 88.1%, F1 89.7%), surpassing the 84.9% of the best compared method, with cross-dataset transfer between ADReSS20 and ADReSSo2021 staying near 87% in both directions, and Whisper outperforming Wav2Vec 2.0 transcripts. The paper further claims that the embedding, not model scale, drives performance: Zephyr-7B-β beats Qwen3-30B (87.4%) and Llama 2-7B (83.2%) within the same pipeline. It also claims early-stage sensitivity, with correctly classified AD subjects averaging MMSE 19.4±7.3 and a substantial share of true positives falling in the mild impairment range of 19–23.
Load-bearing premise
The framework's accuracy depends on the assumption that the Alzheimer's signal in the Cookie Theft recordings survives automatic transcription and is carried by the words and their meaning, rather than by how they are spoken.
Editorial extensions
If this is right
- A clinic could run the entire screening loop on its own servers and still get benchmark-competitive accuracy, removing a major privacy obstacle to using LLMs on patient speech.
- Automatic Whisper transcription is sufficient; manual transcripts and specialized recording hardware are not required, which widens the population that could be screened.
- Because a 7B-parameter model with PCA and logistic regression beats a 30B model in the same pipeline, model scale is not the main driver of accuracy, and small instruction-tuned models may be enough for this task.
- Cross-dataset transfer staying near 87% in both directions suggests the learned linguistic markers generalize across recording protocols, a precondition for real-world deployment.
- The MMSE-stratified analysis indicates mild cases with scores 19–23 are captured, supporting use of the framework as an early-screening trigger rather than only a confirmatory test.
- Going beyond the paper, an immediate testable extension is to report bootstrap confidence intervals around the 90.0% figure; with only 119 test samples the 5.1-point gap over the closest comparator could be within sampling noise.
- Going beyond the paper, running the same pipeline on manual ADReSS20 transcripts would isolate how much information is lost by automatic transcription, and comparing the Zephyr pipeline with an acoustic-only classifier would test whether the text-only premise omits a large part of the signal.
- Going beyond the paper, the concentration of false negatives near the MMSE 24–30 boundary suggests that adding prosodic or acoustic features to the text embeddings is the most direct route to improving detection of preclinical and very early Alzheimer's disease.
Reading between the lines
- Going beyond the paper, an immediate testable extension is to report bootstrap confidence intervals around the 90.0% figure; with only 119 test samples the 5.1-point gap over the closest comparator could be within sampling noise.
- Going beyond the paper, running the same pipeline on manual ADReSS20 transcripts would isolate how much information is lost by automatic transcription, and comparing the Zephyr pipeline with an acoustic-only classifier would test whether the text-only premise omits a large part of the signal.
- Going beyond the paper, the concentration of false negatives near the MMSE 24–30 boundary suggests that adding prosodic or acoustic features to the text embeddings is the most direct route to improving detection of preclinical and very early Alzheimer's disease.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes LSEAD, a privacy-preserving speech-based framework for early Alzheimer's disease (AD) screening. Raw audio is transcribed with an ASR system (Whisper or Wav2Vec 2.0), transcript-level embeddings are extracted from the penultimate layer of locally deployed open-source LLMs (Zephyr-7B-beta, with Llama 2 and Qwen3-30B as baselines), PCA reduces dimensionality, and a downstream classifier (logistic regression, SVC, XGBoost, or a neural network) performs binary AD/CN classification. The method is evaluated on ADReSS20 and ADReSSo2021 under a combined setting and cross-dataset transfer settings. The central claim is that this pipeline achieves 90.0% accuracy on the combined test set, outperforming published methods by up to about 5 percentage points, while keeping all data processing on-premises.
Significance. If the 90.0% accuracy claim were supported by an unbiased evaluation, this would be a practically useful contribution: it combines open-source, locally deployable components, addresses an important privacy constraint in clinical deployment, and shows systematic robustness across two ASR systems and two datasets. The cross-dataset experiments and the comparison of multiple LLM backbones are also useful empirical evidence. However, the headline performance is currently not an unbiased estimate because the best classifier and the PCA variance threshold appear to be selected using the test set; the comparison with prior work also mixes incompatible evaluation protocols. The paper has strengths in scope and reproducibility (a code link is provided) but needs substantial additional experimental rigor before its central claims can be accepted.
major comments (4)
- [§4.2 (Table 2) and §3.2.4 (Eq. 13)] The headline 90.0% accuracy is the best result among four classifiers evaluated on the same 119-sample combined test set, and the PCA variance threshold V_a is selected by a grid search over {0.9, 0.95, 0.97, 0.99, 0.999} without reporting the selected value or reserving a validation split for this choice. Because each test sample is approximately 0.84 percentage points, the reported 5.1% margin over Mortensen et al. corresponds to about six samples. To support the claim, the authors should use nested cross-validation or a pre-specified protocol in which model selection (classifier and V_a) is performed only on training folds, and should report the selected V_a and K as well as confidence intervals or a significance test for the 90.0% accuracy.
- [§4.4 (Table 3)] The comparative claim 'consistently outperforms existing approaches' is not established because the comparison mixes a combined-dataset result (LSEAD, 90.0%) with published per-dataset results from different protocols. The differences may reflect evaluation setup rather than algorithmic superiority. The authors should either re-implement the baselines under the same combined training/test protocol, or restrict the comparison to published results obtained under the same protocol, and state explicitly which protocol each row follows.
- [§4.2 and §3.2.4] The paper asserts that removing PCA 'consistently degrades performance across all models' and that PCA is 'a crucial step in the proposed framework,' but no results for the no-PCA condition are reported in any table or figure. Given that the PCA variance threshold is a free parameter and that the performance of the downstream classifier depends on K, this assertion is not currently supported by the evidence. The authors should include a table or plot comparing classifiers with and without PCA, using a fixed model-selection protocol.
- [§4.4 (Table 3) and §2 (references [36], [61], [44], [65])] There are citation and identification errors that affect the comparison. In Table 3, both Mortensen et al. and Bang et al. are cited as [36], while the text cites Mortensen et al. as [61]; the Zephyr-7B-beta model is introduced with citation [44] (the GPT-3 paper) rather than the Zephyr paper [65]; and 'Kheirkhahzadeh et al. [62]' refers to a single-author work. These errors need to be corrected, especially because the comparison with locally deployable LLMs (Mortensen et al./ADetectoLocum) is central to the contribution.
minor comments (5)
- [Throughout] The notation is used inconsistently: 'V AD' appears with a space in Eq. (3), 'ADReSSo21' and 'ADReSSo2021' are both used, and 'Wav2Vec 2.0' is written as 'Wav2Vec2.0' in Table 6. The text should be unified.
- [Table 2] The 5-fold cross-validation rows report only point estimates; standard deviations or confidence intervals over folds and over repeated runs would help assess the stability of the classifier ranking, especially given the small dataset.
- [§3.2.4] The PCA variance threshold is described as determined by grid search 'based on the suggested PCA variance [0.9,0.95,0.97,0.99,0.999] [51]'; the reference is a lecture note and the selection criterion is not defined. The authors should specify the objective used to choose among these values (e.g., validation accuracy) and report the chosen value.
- [§4.3 (Figure 3)] The means and standard deviations for MMSE scores of true positives and false negatives are reported without sample sizes or confidence intervals; reporting counts per MMSE band would make the early-detection analysis more informative.
- [§5 (Conclusions)] The conclusions state that the framework 'consistently achieves the best performance' and 'strongly generalizes' across datasets, but the cross-dataset results in Tables 7 and 8 also appear to involve selecting the best classifier on the target test set; this should be acknowledged or the protocol should be clarified.
Circularity Check
The 90.0% headline is not an independently predicted accuracy: the classifier, ASR model, and LLM backbone are chosen after inspecting the same 119-sample test results, so the reported number is the best of the grid by construction.
-
fitted input called prediction
[Section 4.2, Table 2 (and Section 4.5, Table 6)]
"The independent test results further reinforce these observations. Among all evaluated classifiers, LR achieves the best overall performance, reaching an accuracy of 90.0% and an F1 score of 89.7%, demonstrating strong generalization to unseen data."
Table 2 reports test-set accuracies for NNs, SVC, XGBoost, and LR on the same 119-sample combined test set, and Section 4.2 selects LR because it has the highest test accuracy. Section 4.5 then compares Wav2Vec 2.0 and Whisper on the same test set and again selects the combination (Whisper + LR) with the best test accuracy. The reported 90.0% is therefore, by construction, the maximum of the evaluated configurations on the test set, not the accuracy of a pre-specified pipeline evaluated once. Calling this an 'independent test result demonstrating strong generalization to unseen data' is circular in the statistical sense: the same test labels were used to choose the model and then to compute the reported performance. The 5.1% margin over Mortensen et al.
full rationale
The pipeline itself is not definitionally circular: Zephyr-7B-β mean-pooled embeddings, PCA, and logistic regression are genuinely different processing stages, and the LR coefficients are trained on the training split rather than copied from the test labels. There are also no load-bearing self-citations: the cited prior work used for comparison (Mortensen et al., Bang et al., Agbavor et al., Luz et al.) is external, and no uniqueness theorem or ansatz is imported from the authors' own prior papers. The circularity-adjacent defect is concentrated in the evaluation protocol: four classifiers, two ASR systems, three LLM backbones, and a PCA variance grid are all compared on the same 119-sample test set, and the best combination is then reported as the method's 'independent test' accuracy. Because model selection and performance estimation share the same labels, the headline 90.0% and the resulting 5.1% SOTA margin are selected quantities rather than unbiased predictions. This is a test-set overfitting issue and a statistical-validity problem, not a definitional equivalence of the sort 'Eq. X = Eq. Y'; it nevertheless fits the fitted-input-called-prediction pattern because the supposedly predicted accuracy is the objective used to choose the final configuration. Other concerns raised by the text (mixing combined-dataset results with published per-dataset numbers in Table 3, not reporting the selected PCA variance or the data used to select it, and the text-only design's exclusion of acoustic features) affect correctness and external validity but are not circularity.
Assumptions & free parameters
free parameters (3)
- PCA variance threshold V_a =
not reported (grid searched over 0.9, 0.95, 0.97, 0.99, 0.999)
- Number of retained principal components K =
not reported
- Downstream classifier family =
logistic regression (best of LR, SVC, XGBoost, NNs)
assumptions (4)
- standard math Eigen-decomposition of the covariance matrix yields a variance-ordered subspace and the PCA variance threshold selects informative components.
- domain assumption Transcribed lexical and semantic content from Cookie Theft picture descriptions carries the AD-related signal, so acoustic and prosodic features can be discarded.
- domain assumption ADReSS20 and ADReSSo2021 are sufficiently comparable, despite different segmentation and transcript availability, for pooling and cross-dataset transfer to be meaningful.
- domain assumption Zephyr-7B-beta produces stable and discriminative mean-pooled penultimate-layer embeddings despite being a beta-stage model with an early 2023 training cutoff.
Cite this review
Pith. "Pith review of LSEAD: A Privacy-Preserving LLM-Based Speech Analysis Framework for Early Alzheimer's Disease Screening." pith.science (2026). https://pith.science/paper/IJRXKUX7
@misc{pith2026260807378,
author = {Pith},
title = {Pith review of: LSEAD: A Privacy-Preserving LLM-Based Speech Analysis Framework for Early Alzheimer's Disease Screening},
year = {2026},
howpublished = {\url{https://pith.science/paper/IJRXKUX7}},
note = {Machine review of arXiv:2608.07378}
}
read the original abstract
Early diagnosis of Alzheimer's disease (AD) is critical for enabling timely interventions that may slow disease progression and improve patient outcomes. There is a growing need for AD detection methods that are non-invasive and cost-effective, especially in real-world clinical settings with diverse patient populations and recording conditions. Speech-based screening addresses these needs by using natural speech collected without specialized equipment. Recent advances in large language models (LLMs) have improved speech analysis by providing rich linguistic representations and strong generalization. In this study, we propose LSEAD, a speech-based AD detection framework using pretrained open-source LLMs. Speech recordings are automatically transcribed, and text embeddings are extracted using locally deployed LLMs. Principal component analysis (PCA) is applied to reduce dimensionality before classification. Because the framework relies only on speech transcripts and locally deployed models, it supports privacy-preserving AD risk assessment without external data exchange. We evaluate LSEAD on the ADReSS20 and ADReSSo2021 benchmark datasets. Experimental results show that LLM-based embeddings generalize well across datasets and improve AD classification accuracy by up to 5 percent over existing methods, especially for early-stage detection. These results demonstrate that LSEAD provides a practical, secure, and scalable approach for early AD screening.
Figures
Reference graph
Works this paper leans on
-
[36]
J.-U. Bang, S.-H. Han, B.-O. Kang, Alzheimer’s disease recognition from spontaneous speech using large lan- guage models, ETRI Journal 46 (1) (2024) 96–105.arXiv:https://onlinelibrary.wiley.com/doi/pdf/ 10.4218/etrij.2023-0356,doi:https://doi.org/10.4218/etrij.2023-0356. URLhttps://onlinelibrary.wiley.com/doi/abs/10.4218/etrij.2023-0356
-
[61]
G. A. Mortensen, R. Zhu, Early alzheimer’s detection through voice analysis: Harnessing locally deployable llms via adetectolocum, a privacy-preserving diagnostic system, AMIA Joint Summits on Translational Science Proceedings 2025 (2025) 365–374
work page 2025
-
[44]
Brown, B
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al., Language models are few-shot learners, Advances in neural information processing systems 33 (2020) 1877–1901
2020
-
[65]
L. Tunstall, E. Beeching, N. Lambert, N. Rajani, K. Rasul, Y . Belkada, S. Huang, L. von Werra, C. Fourrier, N. Habib, N. Sarrazin, O. Sanseviero, A. M. Rush, T. Wolf, Zephyr: Direct distillation of lm alignment (2023). arXiv:2310.16944. URLhttps://arxiv.org/abs/2310.16944
arXiv 2023
-
[62]
M. Kheirkhahzadeh, Speech classification using acoustic embedding and large language models applied on alzheimer’s disease prediction task (2023)
work page 2023
- [1]
-
[2]
C. Reitz, C. Brayne, R. Mayeux, Epidemiology of alzheimer disease, Nature Reviews Neurology 7 (3) (2011) 137–152, epub 2011 Feb 8.doi:10.1038/nrneurol.2011.2
-
[3]
World Health Organization, Global action plan on the public health re- sponse to dementia 2017–2025,https://www.who.int/publications/i/item/ global-action-plan-on-the-public-health-response-to-dementia-2017---2025, accessed: 2025-04-15 (2017)
work page 2017
Show all 70 references
-
[4]
Crous-Bou, C
M. Crous-Bou, C. Minguillón, N. Gramunt, J. L. Molinuevo, Alzheimer’s disease prevention: from risk factors to early intervention, Alzheimer’s Research & Therapy 9 (1) (2017) 71.doi:10.1186/s13195-017-0297-z. URLhttps://doi.org/10.1186/s13195-017-0297-z
2017 doi
-
[5]
Pressman, G
P. Pressman, G. Rabinovici, Alzheimer’s Disease, Elsevier Inc., 2014, pp. 122–127, publisher Copyright:© 2014 Elsevier Inc. All rights reserved.doi:10.1016/B978-0-12-385157-4.00475-9
2014 doi
-
[6]
A. Kaur, M. Mittal, J. S. Bhatti, S. Thareja, S. Singh, A systematic literature review on the significance of deep learning and machine learning in predicting alzheimer’s disease, Artificial Intelligence in Medicine 154 (2024) 102928.doi:https://doi.org/10.1016/j.artmed.2024.1...
2024
-
[7]
Khojaste-Sarakhsi, S
M. Khojaste-Sarakhsi, S. S. Haghighi, S. F. Ghomi, E. Marchiori, Deep learning for alzheimer’s disease diag- nosis: A survey, Artificial Intelligence in Medicine 130 (2022) 102332.doi:https://doi.org/10.1016/j. artmed.2022.102332. URLhttps://www.sciencedirect.com/science/artic...
2022
-
[8]
Lazli, F
L. Lazli, F. Cheriet, M. Boukadoum, Multiclass prediction of alzheimer’s disease using balanced multimodal data and deep ensemble learning, Biomedical Signal Processing and Control 114 (2026) 109026.doi:https: //doi.org/10.1016/j.bspc.2025.109026. URLhttps://www.sciencedirect....
2026
-
[9]
Sudharsan, G
M. Sudharsan, G. Thailambal, An recognition of alzheimer disease using brain mri images with dpnmm through adaptive model, in: 2022 International Conference on Edge Computing and Applications (ICECAA), IEEE, 2022, pp. 952–959
2022
-
[10]
M. S. Safi, S. M. M. Safi, Early detection of alzheimer’s disease from eeg signals using hjorth parameters, Biomedical Signal Processing and Control 65 (2021) 102338
2021
-
[11]
T ˘au¸ tan, B
A.-M. T ˘au¸ tan, B. Ionescu, E. Santarnecchi, Artificial intelligence in neurodegenerative diseases: A review of available tools with a focus on machine learning techniques, Artificial Intelligence in Medicine 117 (2021) 102081.doi:https://doi.org/10.1016/j.artmed.2021.102081...
2021
-
[12]
S. I. Tokushige, H. Matsumoto, S. I. Matsuda, S. Inomata-Terada, N. Kotsuki, M. Hamada, S. Tsuji, Y . Ugawa, Y . Terao, Early detection of cognitive decline in alzheimer’s disease using eye tracking, Frontiers in Aging Neuroscience 15 (2023) 1123456.doi:10.3389/fnagi.2023.1123456
2023
-
[13]
P. S. Pressman, K. H. Chen, J. Casey, S. Sillau, H. J. Chial, C. M. Filley, B. L. Miller, R. W. Leven- son, Incongruences between facial expression and self-reported emotional reactivity in frontotemporal de- mentia and related disorders, The Journal of Neuropsychiatry and Cli...
2023 doi
-
[14]
Zheng, M
C. Zheng, M. Bouazizi, T. Ohtsuki, M. Kitazawa, T. Horigome, T. Kishimoto, Detecting dementia from face- related features with automated computational methods, Bioengineering 10 (7) (2023) 862.doi:10.3390/ bioengineering10070862
2023
-
[15]
König, N
A. König, N. Linz, J. Tröger, Novel digital speech biomarker for early detection of alzheimer’s disease, Alzheimer’s & Dementia 20 (Suppl 3) (2025) e083421.doi:10.1002/alz.083421
2025 doi
-
[16]
J. T. Becker, F. Boller, O. L. Lopez, J. Saxton, K. L. McGonigle, The natural history of alzheimer’s disease: Description of study cohort and accuracy of diagnosis, Archives of Neurology 51 (6) (1994) 585–594.doi: 10.1001/archneur.1994.00540180063015
1994
-
[17]
Filiou, N
R.-P. Filiou, N. Bier, A. Slegers, B. Houzé, P. Belchior, S. M. Brambati, Connected speech assessment in the early detection of alzheimer’s disease and mild cognitive impairment: a scoping review, Aphasiology 34 (6) (2020) 723–755.arXiv:https://doi.org/10.1080/02687038.2019.16...
2020
-
[18]
S. Luz, F. Haider, S. de la Fuente, D. Fromm, B. MacWhinney, Alzheimer’s dementia recognition through spontaneous speech: The adress challenge (2020).arXiv:2004.06833. URLhttps://arxiv.org/abs/2004.06833
2020 arXiv
-
[19]
Alsuhaibani, A
M. Alsuhaibani, A. Pourramezan Fard, J. Sun, F. Far Poor, P. S. Pressman, M. H. Mahoor, A review of machine learning approaches for non-invasive cognitive impairment detection, IEEE access 13 (2025) 56355–56384
2025
-
[20]
Agbavor, H
F. Agbavor, H. Liang, Predicting dementia from spontaneous speech using large language models, PLOS digital health 1 (12) (2022) e0000168. 18
2022
-
[21]
Panesar, M
K. Panesar, M. B. Pérez Cabello de Alba, Natural language processing-driven framework for the early detection of language and cognitive decline, Language and Health 1 (2) (2023) 20–35.doi:https://doi.org/10. 1016/j.laheal.2023.09.002. URLhttps://www.sciencedirect.com/science/a...
2023
-
[22]
Maity, M
S. Maity, M. J. Saikia, Large language models in healthcare and medical applications: A review, Bioengineering 12 (6) (2025).doi:10.3390/bioengineering12060631. URLhttps://www.mdpi.com/2306-5354/12/6/631
2025 doi
-
[23]
Shadle, Phonetics, acoustic, in: K
C. Shadle, Phonetics, acoustic, in: K. Brown (Ed.), Encyclopedia of Language & Linguistics (Second Edi- tion), second edition Edition, Elsevier, Oxford, 2006, pp. 442–460.doi:https://doi.org/10.1016/ B0-08-044854-2/00001-8. URLhttps://www.sciencedirect.com/science/article/pii/...
2006
-
[24]
Mahon, M
E. Mahon, M. E. Lachman, V oice biomarkers as indicators of cognitive changes in middle and later adulthood, Neurobiology of Aging 119 (2022) 22–35.doi:https://doi.org/10.1016/j.neurobiolaging.2022. 06.010. URLhttps://www.sciencedirect.com/science/article/pii/S0197458022001415
2022 doi
-
[25]
Zolnour, H
A. Zolnour, H. Azadmaleki, Y . Haghbin, F. Taherinezhad, M. J. M. Nezhad, S. Rashidi, M. Khani, A. Taleban, S. M. Sani, M. Dadkhah, J. M. Noble, S. Bakken, Y . Yaghoobzadeh, A. H. Vahabie, M. Rouhizadeh, M. Zolnoori, Llmcare: Early detection of cognitive impairment via transfo...
2025
-
[26]
Shankar, Z
R. Shankar, Z. Goh, F. Devi, et al., A systematic review of explainable artificial intelligence methods for speech- based cognitive decline detection, npj Digital Medicine 8 (2025) 724.doi:10.1038/s41746-025-02105-z
2025 doi
-
[27]
P. M. Naim, S. Sadeh-Sharvit, S. Jefroykin, E. Silber, D. P. Morrison, A. Goldstein, Preprocessing large-scale conversational datasets: A framework and its application to behavioral health transcripts, JMIR Form Res 9 (2025) e78082.doi:10.2196/78082. URLhttps://formative.jmir....
2025 doi
-
[28]
S. O. Russell, I. Gessinger, A. Krason, G. Vigliocco, N. Harte, What automatic speech recognition can and cannot do for conversational speech transcription, Research Methods in Applied Linguistics 3 (3) (2024) 100163. doi:https://doi.org/10.1016/j.rmal.2024.100163. URLhttps://...
2024
-
[29]
Zhang, Q
M. Zhang, Q. Cui, W. Li, W. Yu, L. Chen, W. Li, C. Zhu, Y . Lü, Augmented dialectal speech recognition for ai-based neuropsychological scale assessment in alzheimer’s disease, Biomedical Signal Processing and Control 99 (2025) 106821.doi:https://doi.org/10.1016/j.bspc.2024.106...
2025
-
[30]
Agbavor, H
F. Agbavor, H. Liang, Artificial intelligence-enabled end-to-end detection and assessment of alzheimer’s disease using voice, Brain Sciences 13 (1) (2023) 28.doi:10.3390/brainsci13010028
2023 doi
-
[31]
D. S. Asudani, N. K. Nagwani, P. Singh, Impact of word embedding models on text analytics in deep learning environment: a review, Artificial Intelligence Review 56 (2023) 10345–10425.doi:10.1007/ s10462-023-10419-1
2023
-
[32]
Devlin, M.-W
J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional transformers for language understanding, in: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, v...
2019
-
[33]
Radford, J
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, Language models are unsupervised multitask learners, OpenAI Blog 1 (8) (2019) 9. 19
2019
-
[34]
Searle, Z
T. Searle, Z. Ibrahim, R. Dobson, Comparing natural language processing techniques for alzheimer’s dementia prediction in spontaneous speech, in: Proceedings of Interspeech, 2020, pp. 2192–2196
2020
-
[35]
Roshanzamir, H
A. Roshanzamir, H. Aghajan, M. Soleymani Baghshah, Transformer-based deep neural network language mod- els for alzheimer’s disease risk assessment from targeted speech, BMC Medical Informatics and Decision Mak- ing 21 (1) (2021) 92
2021
-
[37]
Y . Guo, C. Li, C. Roan, S. Pakhomov, T. Cohen, Crossing the ‘cookie theft’ corpus chasm: Applying what bert learns from outside data to the adress challenge dementia detection task, Frontiers in Computer Science 3 (2021) 642517
2021
-
[38]
J. Yuan, X. Cai, Y . Bian, Z. Ye, K. Church, Pauses for detection of alzheimer’s disease, Frontiers in Computer Science V olume 2 - 2020 (2021).doi:10.3389/fcomp.2020.624488. URLhttps://www.frontiersin.org/journals/computer-science/articles/10.3389/fcomp. 2020.624488
2021
-
[39]
W. N. Price, I. G. Cohen, Privacy in the age of medical big data, Nature Medicine 25 (1) (2019) 37–43.doi: 10.1038/s41591-018-0272-7
2019 doi
-
[40]
Moore, S
W. Moore, S. Frye, Review of hipaa, part 1: History, protected health information, and privacy and security rules, Journal of Nuclear Medicine Technology 47 (4) (2019) 269–272.doi:10.2967/jnmt.119.227819
2019 doi
-
[41]
Yadav, S
N. Yadav, S. Pandey, A. Gupta, P. Dudani, S. Gupta, K. Rangarajan, Data privacy in healthcare: In the era of artificial intelligence, Indian Dermatology Online Journal 14 (6) (2023) 788–792.doi:10.4103/idoj.idoj_ 543_23
2023 doi
-
[42]
Li, Security implications of ai chatbots in health care, Journal of Medical Internet Research 25 (1) (2023) e47551.doi:10.2196/47551
J. Li, Security implications of ai chatbots in health care, Journal of Medical Internet Research 25 (1) (2023) e47551.doi:10.2196/47551
2023 doi
-
[45]
S. Luz, F. Haider, S. de la Fuente, D. Fromm, B. MacWhinney, Detecting cognitive decline using speech only: The adresso challenge, arXiv preprint arXiv:2104.09356 (2021). URLhttps://arxiv.org/abs/2104.09356
2021 arXiv
-
[46]
Goodglass, E
H. Goodglass, E. Kaplan, B. Barresi, Boston Diagnostic Aphasia Examination, 3rd Edition, Lippincott Williams & Wilkins, Philadelphia, 2001
2001
-
[47]
Teipel, D
S. Teipel, D. Gustafson, R. Ossenkoppele, O. Hansson, C. Babiloni, M. Wagner, S. G. Riedel-Heller, I. Kilimann, Y . Tang, Alzheimer disease: Standard of diagnosis, treatment, care, and prevention, Journal of Nuclear Medicine 63 (7) (2022) 981–985.arXiv:https://jnm.snmjournals....
2022
-
[48]
Y . Pan, B. Mirheidari, J. M. Harris, J. C. Thompson, M. Jones, J. S. Snowden, D. Blackburn, H. Christensen, Using the outputs of different automatic speech recognition paradigms for acoustic- and bert-based alzheimer’s dementia detection through spontaneous speech, in: Inters...
2021
-
[49]
Tsukagoshi, R
H. Tsukagoshi, R. Sasano, Redundancy, isotropy, and intrinsic dimensionality of prompt-based text embeddings, in: W. Che, J. Nabende, E. Shutova, M. T. Pilehvar (Eds.), Findings of the Association for Computational Linguistics: ACL 2025, Association for Computational Linguisti...
2025 doi
-
[50]
Liashchynskyi, P
P. Liashchynskyi, P. Liashchynskyi, Grid search, random search, genetic algorithm: A big comparison for nas (2019).arXiv:1912.06059. URLhttps://arxiv.org/abs/1912.06059
2019 arXiv
-
[51]
Chen, Principal component analysis (pca), Lecture notes, San José State University (2020)
G. Chen, Principal component analysis (pca), Lecture notes, San José State University (2020). URLhttps://www.sjsu.edu/faculty/guangliang.chen/Math253S20/lec8pca.pdf
2020
-
[52]
A. M. Kashyap, D. Rao, M. R. Boland, L. Shen, C. Callison-Burch, Predicting explainable dementia types with llm-aided feature engineering, Bioinformatics 41 (4) (2025) btaf156.doi:10.1093/bioinformatics/ btaf156
2025 doi
-
[53]
B. A. Llaca-Sánchez, L. R. García-Noguez, M. A. Aceves-Fernández, A. Takacs, S. Tovar-Arriaga, Exploring llm embedding potential for dementia detection using audio transcripts, Eng 6 (7) (2025).doi:10.3390/ eng6070163. URLhttps://www.mdpi.com/2673-4117/6/7/163
2025
-
[54]
T. Mo, J. C. K. Lam, V . O. K. Li, L. Y . L. Cheung, Leveraging large language models for identifying interpretable linguistic markers and enhancing alzheimer’s disease diagnostics, medRxiv (2024).arXiv: https://www.medrxiv.org/content/early/2024/08/23/2024.08.22.24312463.full...
2024
-
[55]
R. Xiao, X. Cui, H. Qiao, X. Zheng, Y . Zhang, C. Zhang, X. Liu, Early diagnosis model of alzheimer’s disease based on sparse logistic regression with the generalized elastic net, Biomedical Signal Processing and Control 66 (2021) 102362.doi:https://doi.org/10.1016/j.bspc.2020...
2021
-
[56]
J. V . Shanmugam, B. Duraisamy, B. C. Simon, P. Bhaskaran, Alzheimer’s disease classification using pre-trained deep networks, Biomedical Signal Processing and Control 71 (2022) 103217.doi:https://doi.org/10. 1016/j.bspc.2021.103217. URLhttps://www.sciencedirect.com/science/ar...
2022
-
[57]
Botros, F
J. Botros, F. Mourad-Chehade, D. Laplanche, Explainable multimodal data fusion framework for heart failure detection: Integrating cnn and xgboost, Biomedical Signal Processing and Control 100 (2025) 106997.doi: https://doi.org/10.1016/j.bspc.2024.106997. URLhttps://www.science...
2025
-
[58]
Goodfellow, Y
I. Goodfellow, Y . Bengio, A. Courville, Y . Bengio, Deep learning, V ol. 1, MIT press Cambridge, 2016
2016
-
[59]
Hastie, R
T. Hastie, R. Tibshirani, J. Friedman, The Elements of Statistical Learning: Data Mining, Inference, and Predic- tion, 2nd Edition, Springer, New York, NY , 2009
2009
-
[60]
Kurlowicz, M
L. Kurlowicz, M. Wallace, The mini-mental state examination (mmse), Journal of Gerontological Nursing 25 (5) (1999) 8–9.doi:10.3928/0098-9134-19990501-08. URLhttps://doi.org/10.3928/0098-9134-19990501-08 21
1999 doi
-
[63]
Touvron, L
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, et al., Llama 2: Open foun- dation and fine-tuned chat models,https://ai.meta.com/research/publications/ llama-2-open-foundation-and-fine-tuned-chat-models/, accessed: 2024-03-18 (2023)
2023
-
[64]
A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. L...
2025 arXiv
-
[66]
Fawcett, An introduction to roc analysis, Pattern Recognition Letters 27 (8) (2006) 861–874.doi:10.1016/ j.patrec.2005.10.010
T. Fawcett, An introduction to roc analysis, Pattern Recognition Letters 27 (8) (2006) 861–874.doi:10.1016/ j.patrec.2005.10.010
2006
-
[67]
Walther, M
F. Walther, M. Eberlein-Gonska, R. T. Hoffmann, J. Schmitt, S. F. U. Blum, Measuring appropriate- ness of diagnostic imaging: A scoping review, Insights into Imaging 14 (1) (2023) 62.doi:10.1186/ s13244-023-01409-6. URLhttps://doi.org/10.1186/s13244-023-01409-6
2023 doi
-
[68]
Wolfsgruber, J
S. Wolfsgruber, J. L. Molinuevo, M. Wagner, et al., Prevalence of abnormal alzheimer’s disease biomarkers in patients with subjective cognitive decline: Cross-sectional comparison of three european memory clinic samples, Alzheimer’s Research & Therapy 11 (2019) 8.doi:10.1186/s...
2019 doi
-
[69]
Baevski, Y
A. Baevski, Y . Zhou, A. Mohamed, M. Auli, wav2vec 2.0: A framework for self-supervised learning of speech representations, Advances in Neural Information Processing Systems 33 (2020) 12449–12460
2020
-
[70]
Radford, J
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, I. Sutskever, Robust speech recognition via large- scale weak supervision, arXiv preprint arXiv:2212.04356 (2023)
2023 arXiv
-
[71]
V . D. Badal, J. M. Reinen, E. W. Twamley, E. E. Lee, R. P. Fellows, E. Bilal, C. A. Depp, Investigating acoustic and psycholinguistic predictors of cognitive impairment in older adults: Modeling study, JMIR Aging 7 (2024) e54655.doi:10.2196/54655. URLhttps://aging.jmir.org/20...
2024 doi
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.