Pith. sign in

REVIEW 4 major objections 4 minor 40 references

Beyond Manual Transcripts: The Potential of Automated Speech Recognition Errors in Improving Alzheimer's Disease Detection

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Automatically transcribed speech, errors included, can outperform human transcripts for detecting Alzheimer's disease.

desk verdict Broad ASR survey that extends a known result, but the headline comparisons sit inside seed noise on a 48-sample test set. read the letter →

arxiv 2505.19448 v1 pith:L3AGJQK5 submitted 2025-05-26 eess.AS

classification eess.AS
keywords Alzheimer'sdiseasedetectionautomaticspeechrecognitionASRerrorscross-attentioninterpretabilitysynthesisADReSSdementiascreening
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that errors introduced by automatic speech recognition are not merely noise to be removed before screening for Alzheimer's disease; in some cases they are diagnostic cues in their own right. On the ADReSS benchmark, certain ASR transcripts and speech re-synthesized from those transcripts classify Alzheimer's versus healthy controls more accurately than manual transcripts and their synthesized counterparts. The authors attribute the gain to asymmetric biases: recognizer mistakes mirror AD speakers' indistinct pronunciation and disfluencies, amplifying group differences in lexical and timing features. The paper also contributes a cross-attention model that performs the classification while exposing which features the errors strengthen. If the claim holds, fully automated, low-cost dementia screening could rely directly on recognizer output rather than expensive human transcription.

What carries the argument

The machinery is a cross-attention interpretability model that fuses two input types: knowledge-based features (35 textual or 60 acoustic descriptors previously linked to AD) and embeddings from a pre-trained language or speech model. The key-projection matrix is fixed to identity, so the knowledge-based feature vector acts directly as the keys while BERT or Wav2Vec2 embeddings supply queries and values; the resulting attention matrix drives the classification and, averaged over correctly predicted samples, reveals which features receive heightened attention. This design converts the model's decision into a readable map of which cues—such as disfluency ratios or pause durations—are emphasized when ASR output is used, which is how the paper identifies the error-derived signals.

What would settle it

Re-run the synthesized-speech comparison using the exact same prompt segment from the original recording for both the ASR transcript and the manual transcript of each subject; if the ASR-synthesized advantage disappears, the reported gain is an artifact of unequal prompts rather than evidence about ASR errors.

Watch

Extended reading notes

Core claim

The central discovery is that noisy machine transcription can beat the gold-standard human transcript for detecting Alzheimer's disease. After fine-tuning eighteen variants of four recognizer families and producing thirty-six transcripts with a wide range of word error rates, the paper finds that several ASR transcripts exceed manual transcripts in mean classification accuracy, and the same pattern appears when the transcripts are converted to speech by a neural text-to-speech model. The proposed cross-attention model reproduces the advantage and traces it to specific knowledge-based features—syllable and lexicon counts, average sentence length, filler-pause and repetition ratios, readability indices, and prosodic timing variables—that ASR errors amplify in the Alzheimer's group compared with healthy controls. These results support the paper's hypothesis that ASR errors introduce asymmetric, group-dependent biases that function as usable diagnostic signals.

Load-bearing premise

The synthesized-speech result assumes the random voice prompt chosen for each transcript does not differ between ASR-derived and manually derived conditions in ways that affect detection; if it does, the apparent accuracy gain may come from acoustics rather than from transcript errors.

Editorial extensions

If this is right

  • End-to-end Alzheimer's screening can run from raw audio to diagnosis without human transcription, since recognizer output is sufficient and sometimes more accurate.
  • Word error rate is not a reliable guide to diagnostic value: the best-performing detector in the tables is not the lowest-error recognizer.
  • ASR errors that reflect articulation and disfluency can be treated as deliberate biomarkers instead of artifacts to be discarded.
  • Because synthesized speech scores below original audio, TTS-based analysis is best used as an experimental tool for isolating transcript content, not as a clinical replacement for real recordings.
  • The same cross-attention classifier can be applied to other transcript- or speech-based cognitive assessments, carrying interpretability with it.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An operational screener could select the recognizer that maximizes downstream detection accuracy on a held-out set, even if that recognizer is not the most accurate transcriber; the paper's numbers show accuracy and word error rate do not move together.
  • The cross-attention diagnostic could be transferred to other conditions marked by speech disfluency, such as Parkinson's disease or primary progressive aphasia, to test whether recognition bias is a general clinical signal.
  • A natural follow-up is to correlate error subtypes (insertions, deletions, substitutions) with cognitive severity scores, since the feature analysis suggests errors concentrate in disfluent, low-diversity stretches of speech.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper studies whether automatic speech recognition (ASR) transcripts, and speech synthesized from them, can outperform manual transcripts for Alzheimer's disease (AD) detection on the ADReSS dataset. The authors fine-tune 18 ASR model variants on DementiaBank data, produce 36 ASR transcripts per ADReSS subject, and compare AD detection accuracy using BERT/Wav2Vec2 embeddings and knowledge-based features in self-attention and cross-attention models. They report that certain ASR transcripts and ASR-synthesized speech yield higher accuracy than manual counterparts, and they propose a cross-attention interpretability model to identify which knowledge-based features are amplified by ASR errors. The paper also claims that ASR errors introduce asymmetric biases between AD and healthy control groups that help detection.

Significance. If the central empirical claim is robust, the finding is practically significant: ASR transcripts are cheaper to obtain than manual transcripts, and the possibility that ASR errors carry informative cues about AD could improve scalable screening. The study is also unusually broad, covering 36 ASR variants and two input modalities (text and synthesized speech), and the proposed cross-attention model offers a concrete interpretability mechanism. However, the significance hinges entirely on whether the reported accuracy differences are statistically meaningful and on whether the synthesized-speech comparison controls for acoustic confounds. The paper does not currently provide the needed evidence for these load-bearing points.

major comments (4)
  1. [Section 4.3, Table 2] The central claim that certain ASR transcripts and ASR-synthesized speech outperform manual counterparts is not supported by significance testing. The table reports mean accuracy over 10 seeds on a 48-sample test set, but no standard deviations, confidence intervals, or paired significance tests are given. With n=48, the standard error of an accuracy estimate is roughly 7 percentage points, and the largest headline margin (85.42% vs. 81.67% for fine-tuned whisper large v2 in the embedding self-attention row) is 3.75 points, i.e., about 1.8 subjects. Since 36 ASR variants and three detection methods are compared, some apparent winners are expected by chance. Please report per-seed variance, perform paired tests (e.g., McNemar or bootstrap) for the manual-vs-ASR comparisons, and correct for multiple comparisons, or the abstract's assertion is not established.
  2. [Section 3.2, Section 4.3] The synthesized-speech comparison is potentially confounded by the choice of TTS prompt. The paper states that for each transcript the authors 'randomly selected a segment from the subject's original speech as a prompt' for CosyVoice2. If the prompt segment is not matched across the manual and ASR conditions, the synthesized audio may differ in speaker identity, recording conditions, or segment-specific acoustic content that is unrelated to whether the transcript came from ASR or manual transcription. The measured accuracy difference between ASR-synthesized and manual-synthesized speech cannot then be attributed purely to transcript content. Please specify whether the same prompt segment was used for all transcript conditions of a given subject, and if not, re-run the comparison with matched prompts or otherwise control for this variable.
  3. [Section 4.4] The interpretability analysis selects the best-performing random seed on the test set and then analyzes attention scores from that same test set. This is circular: the model is chosen for its test-set accuracy, so the subsequent attention analysis is conditioned on the test set and may reflect chance patterns rather than stable ASR-related cues. A proper protocol would select hyperparameters and seeds on a validation split, or report the distribution of attention patterns across seeds. Additionally, the claim that 'additional statistical experiments support the hypothesis' is not backed by any reported results; the experiments are mentioned but not described or shown, so the asymmetric-bias explanation is currently unsupported.
  4. [Section 4.3, observation (a)] The paper's repeated use of 'certain ASR transcripts' is insufficiently precise. Across 36 ASR variants and three detection methods, the identity of the 'certain' winners changes by row, and no criterion is given for what counts as a systematic advantage. A reader cannot tell whether the reported superiority is a consistent property of specific ASR error patterns or a set of isolated high-scoring configurations. Please define the comparison criterion a priori, and report the number of variants/methods that exceed manual performance and whether this number is itself significant relative to the 36-by-3 comparison space.
minor comments (4)
  1. [Section 1] There is a typo in the introduction: 'withih' should be 'within', and 'manual transcripts is labor-intensive' should be 'manual transcripts are labor-intensive'.
  2. [Section 4.5] The paper says the attention heatmap for original speech is 'omitted due to space limitations', but the corresponding analysis is a substantive part of the claimed contribution. Please include the figure or a quantitative summary, otherwise Section 4.5 cannot be verified.
  3. [Section 3.4.2, Equation (1)] The use of an identity projection Wk so that knowledge-based features directly serve as K is an interesting design choice, but the paper should clarify whether this restricts the learnable interaction between embeddings and features, and whether alternative projections were considered in preliminary experiments.
  4. [Table 2] The table caption defines the three slash-separated values, but the layout makes it difficult to compare rows quickly. Consider restructuring the table so that manual-transcript and manual-synthesized baselines are repeated or clearly marked in each block, and add a column for the WER of each ASR model to help the reader connect transcription quality with detection accuracy.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central ASR-vs-manual comparison is an external empirical benchmark; the only self-citation is non-load-bearing.

full rationale

The paper's central claim is empirical: it compares AD detection accuracy on the held-out ADReSS test set using manual transcripts, 36 ASR transcripts, and their TTS-synthesized counterparts. The ASR models are fine-tuned on DementiaBank (WLS, Lu, Kempler) and the detection models are trained on the ADReSS training set, so the headline numbers are genuine out-of-sample results rather than fitted quantities renamed as predictions. The only self-citation, [21], appears in a literature-gap statement ('a comprehensive evaluation of various ASR models for AD detection is limited [20, 21]') and is not load-bearing for any result. The cross-attention interpretability model does set K equal to the knowledge-based features through an identity projection (Section 3.4.2), so the attention scores are computed directly from those features; however, the paper uses these scores for post-hoc comparison between ASR and manual conditions, and the main 'ASR errors help' claim is supported by the accuracy differences in Table 2, not by the attention ranking. This design choice limits the strength of the interpretability analysis but does not make any prediction equivalent to its input by construction. Concerns about missing significance tests and unreported statistical experiments are correctness risks, not circularity. No uniqueness theorem, no ansatz smuggled in via citation, and no renaming of a known result were found.

Assumptions & free parameters 4 free parameters · 5 assumptions · 1 invented entities

The central empirical comparison rests on the ADReSS benchmark, on the assumption that ASR-vs-manual accuracy differences are not confounded with acoustic or speaker factors, and on the interpretability assumption that attention scores denote feature importance. The main free choices are the embedding layer indices, the random seed selection for interpretation, and the TTS prompt segments.

free parameters (4)
  • BERT layer index = 11
    Chosen based on preliminary experiments (Section 3.3.2), not derived; different layer choices can change embeddings and affect results.
  • Wav2Vec2 layer index = 8
    Selected empirically (Section 3.3.2); the paper relies on this layer for acoustic embeddings.
  • Best-performing random seed selection = Seed with highest test accuracy per condition
    Section 4.4 selects the best-performing seed on the ADReSS test set for interpretability analysis, a data-dependent choice that can inflate reported effects.
  • TTS prompt segment = Random segment from subject's original speech
    Section 3.2; the random prompt choice can vary between ASR and manual conditions, confounding the synthesized-speech comparison.
assumptions (5)
  • domain assumption ADReSS dataset is a valid benchmark for AD detection and the subjects are representative.
    The paper treats ADReSS as the gold standard; the test set has only 48 samples, making group comparisons noisy.
  • domain assumption ASR-vs-manual transcript accuracy differences are not confounded with subject identity or recording quality.
    The comparison assumes transcript differences are due to ASR errors, not acoustic or speaker factors that also influence the classifier.
  • domain assumption Attention scores in the cross-attention model are meaningful estimates of feature importance.
    Sections 4.4 and 4.5 interpret attention scores as revealing 'valuable cues'; this is a strong assumption about model interpretability.
  • domain assumption CosyVoice2 TTS preserves the lexical content and relevant para-linguistic cues equally for ASR and manual transcripts.
    The synthesized-speech experiment relies on this, but the paper reports that synthesized speech performs worse than original speech, indicating the TTS loses some pathological features.
  • domain assumption Pre-trained BERT and Wav2Vec2 embeddings capture AD-relevant information in the chosen layers.
    Standard practice, but the layer selection is empirical and the embeddings are not verified against external benchmarks here.
invented entities (1)
  • Asymmetric bias hypothesis (ASR errors introduce systematic differences between AD and HC groups)
    purpose: To explain why ASR transcripts outperform manual transcripts in AD detection.
    The paper asserts this hypothesis in Section 4.3(a) and Section 5, but does not test it directly; the claimed supporting 'additional statistical experiments' are not reported (Section 4.4).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Beyond Manual Transcripts: The Potential of Automated Speech Recognition Errors in Improving Alzheimer's Disease Detection." pith.science (2026). https://pith.science/paper/L3AGJQK5

@misc{pith2026250519448,
  author       = {Pith},
  title        = {Pith review of: Beyond Manual Transcripts: The Potential of Automated Speech Recognition Errors in Improving Alzheimer's Disease Detection},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/L3AGJQK5}},
  note         = {Machine review of arXiv:2505.19448}
}
read the original abstract

Recent breakthroughs in Automatic Speech Recognition (ASR) have enabled fully automated Alzheimer's Disease (AD) detection using ASR transcripts. Nonetheless, the impact of ASR errors on AD detection remains poorly understood. This paper fills the gap. We conduct a comprehensive study on AD detection using transcripts from various ASR models and their synthesized speech on the ADReSS dataset. Experimental results reveal that certain ASR transcripts (ASR-synthesized speech) outperform manual transcripts (manual-synthesized speech) in detection accuracy, suggesting that ASR errors may provide valuable cues for improving AD detection. Additionally, we propose a cross-attention-based interpretability model that not only identifies these cues but also achieves superior or comparable performance to the baseline. Furthermore, we utilize this model to unveil AD-related patterns within pre-trained embeddings. Our study offers novel insights into the potential of ASR models for AD detection.

Figures

Figures reproduced from arXiv: 2505.19448 by the authors.

Figure 1
Figure 1. (a) refers to the overall workflow, (b) refers to the AD detection models, and (c) refers to the interpretability analysis. and Kempler (4 samples). The ADReSS challenge dataset [8] from Interspeech 2020 was used for binary AD detection, with 108 training samples (54 AD, 54 HC) and 48 test samples (24 AD, 24 HC). Each sample consists of an audio record￾ing and a corresponding manual transcript, which include sub￾jec… view at source ↗
Figure 2
Figure 2. Attention analysis of transcripts. best detection accuracy of 87.5%), then input the correspond￾ing ADReSS test set into this model. For each correctly pre￾dicted sample, we obtained an attention score matrix (i.e., ma￾trix A in Section 3.4.2), then averaged all matrices, as shown in Figure 2a. The x-axis (0-34) corresponded to each of the knowledge-based text features introduced sequentially in Sec￾tion 3.3.1. This… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

40 extracted references · 20 canonical work pages

  1. [1]

    Introduction Alzheimer’s Disease (AD), the leading cause of dementia, is a progressive neurodegenerative disorder causing irreversible brain damage and cognitive decline in memory, language, atten- tion, and executive function [1]. Early detection is crucial, and compared to traditional clinical methods, speech-and-language- based automatic AD diagnosis h...

  2. [2]

    Beyond Manual Transcripts: The Potential of Automated Speech Recognition Errors in Improving Alzheimer's Disease Detection

    Dataset We selected and organized three datasets from DB for fine- tuning the ASR models: WLS (187 samples), Lu (54 samples), arXiv:2505.19448v1 [eess.AS] 26 May 2025 Fine-tune 18 fine-tuned ASR models ADReSS 36 ASR transcripts Manual transcripts 18 original ASR models Overall workflow DementiaBank WLS Lu Kempler (a) CosyVoice2 Synthesized speech Origi...

  3. [3]

    Fine-tuning ASR models We selected 18 variants from Wav2Vec2 [14], HuBERT [15], WavLM [16], and Whisper [17] for fine-tuning, as these models provide varying WER

    Methods 3.1. Fine-tuning ASR models We selected 18 variants from Wav2Vec2 [14], HuBERT [15], WavLM [16], and Whisper [17] for fine-tuning, as these models provide varying WER. The selected variants are: wav2vec2-{base-100h, base-960h, large-960h, large-960h- lv60, large-960h-lv60-self, large-xlsr-53-english, xls-r-1b- english}, hubert-{large-ls960-ft, xla...

  4. [4]

    Experimental setup For fine-tuning the ASR models, we employed the AdamW op- timizer with a learning rate of 1 × 10−5 and weight decay of 5 × 10−3

    Experiments and results 4.1. Experimental setup For fine-tuning the ASR models, we employed the AdamW op- timizer with a learning rate of 1 × 10−5 and weight decay of 5 × 10−3. The models were trained for 20 epochs with a batch size of 8, and performance was evaluated using WER. For AD detection, we used AdamW with a learning rate of 4 × 10−4 and weight d...

  5. [5]

    Conclusions In this paper, we conducted a comprehensive study on AD de- tection using ASR transcripts with various WER and their cor- responding synthesized speech. Our results indicated that ASR errors may offer valuable cues for improving AD detection, which we attributed to the asymmetric biases introduced by these errors between the AD and HC groups. ...

  6. [6]

    S2023Z20004), by the National Social Science Foundation of China (Grant No

    Acknowledgements This work was partially supported by the Anhui Province Major Science and Technology Research Project (Grant No. S2023Z20004), by the National Social Science Foundation of China (Grant No. 23AYY012), and by the Supercomputing Center of the University of Science and Technology of China

  7. [7]

    Advances in the early detection of alzheimer’s disease,

    P. J. Nestor, P. Scheltens, and J. R. Hodges, “Advances in the early detection of alzheimer’s disease,” Nature medicine, vol. 10, no. Suppl 7, pp. S34–S41, 2004

  8. [8]

    De- tection of cognitive impairment and alzheimer’s disease using a speech- and language-based protocol,

    T. Talkar, S. Charles, C. Krantsevich, and K. Kawabata, “De- tection of cognitive impairment and alzheimer’s disease using a speech- and language-based protocol,” inProc. Interspeech, 2024, pp. 3025–3029

Show all 40 references
  1. [9]

    Clever hans ef- fect found in automatic detection of alzheimer’s disease through speech,

    Y .-L. Liu, R. Feng, J.-H. Yuan, and Z.-H. Ling, “Clever hans ef- fect found in automatic detection of alzheimer’s disease through speech,” in Proc. Interspeech, 2024, pp. 2435–2439

  2. [10]

    Macro-descriptors for alzheimer’s disease detection using large language models,

    C. Botelho, J. Mendonc ¸a, A. Pompili, T. Schultz, A. Abad, and I. Trancoso, “Macro-descriptors for alzheimer’s disease detection using large language models,” in Proc. Interspeech, 2024, pp. 1975–1979

  3. [11]

    Using the outputs of different auto- matic speech recognition paradigms for acoustic-and BERT-based Alzheimer’s dementia detection through spontaneous speech

    Y . Pan, B. Mirheidari, J. M. Harris, J. C. Thompson, M. Jones, J. S. Snowden et al. , “Using the outputs of different auto- matic speech recognition paradigms for acoustic-and BERT-based Alzheimer’s dementia detection through spontaneous speech.” in Proc. Interspeech, 2021, p...

  4. [12]

    The USTC system for ADReSS-M challenge,

    K. Mei, X. Ding, Y . Liu, Z. Guo, F. Xu, X. Li, T. Naren, J. Yuan, and Z. Ling, “The USTC system for ADReSS-M challenge,” in Proc. ICASSP, 2023, pp. 1–2

  5. [13]

    Leveraging prompt learning and pause encoding for alzheimer’s disease de- tection,

    Y .-L. Liu, R. Feng, J.-H. Yuan, and Z.-H. Ling, “Leveraging prompt learning and pause encoding for alzheimer’s disease de- tection,” in Proc. ISCSLP, 2024, pp. 486–490

  6. [14]

    Alzheimer’s dementia recognition through spontaneous speech: The ADReSS challenge,

    S. Luz, F. Haider, S. de la Fuente Garcia, D. Fromm, and B. MacWhinney, “Alzheimer’s dementia recognition through spontaneous speech: The ADReSS challenge,” in Proc. Inter- speech, 2020, pp. 2172–2176

  7. [15]

    Detecting cognitive decline using speech only: The ADReSSo challenge,

    S. Luz, F. Haider, S. de la Fuente, D. Fromm, and B. MacWhin- ney, “Detecting cognitive decline using speech only: The ADReSSo challenge,” arXiv preprint arXiv:2104.09356, 2021

  8. [16]

    Multilingual Alzheimer’s dementia recogni- tion through spontaneous speech: a signal processing grand chal- lenge,

    S. Luz, F. Haider, D. Fromm, I. Lazarou, I. Kompatsiaris, and B. MacWhinney, “Multilingual Alzheimer’s dementia recogni- tion through spontaneous speech: a signal processing grand chal- lenge,” in Proc. ICASSP, 2023, pp. 1–2

  9. [17]

    Connected speech-based cognitive assessment in chinese and english,

    S. D. L. F. Garcia, F. Haider, D. Fromm, B. MacWhin- ney, A. Lanzi, Y .-N. Chang et al. , “Connected speech-based cognitive assessment in chinese and english,” arXiv preprint arXiv:2406.10272, 2024

  10. [18]

    Automated screening for Alzheimer’s dementia through spontaneous speech

    M. S. S. Syed, Z. S. Syed, M. Lech, and E. Pirogova, “Automated screening for Alzheimer’s dementia through spontaneous speech.” in Proc. Interspeech, 2020, pp. 2222–2226

  11. [19]

    A comparative study of acoustic and linguistic fea- tures classification for Alzheimer’s disease detection,

    J. Li, J. Yu, Z. Ye, S. Wong, M. Mak, B. Mak, X. Liu, and H. Meng, “A comparative study of acoustic and linguistic fea- tures classification for Alzheimer’s disease detection,” in Proc. ICASSP, 2021, pp. 6423–6427

  12. [20]

    wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,

    A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,” in Proc. NeurIPS, 2020, pp. 12 449–12 460

  13. [21]

    HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

    W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhut- dinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM transactions on audio, speech, and language process- ing, vol. 29, pp. 3451–3460, 2021

  14. [22]

    WavLM: Large-scale self-supervised pre-training for full stack speech processing,

    S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen et al. , “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, 2022

  15. [23]

    Robust speech recognition via large-scale weak su- pervision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak su- pervision,” in Proc. ICML, 2023, pp. 28 492–28 518

  16. [24]

    Speech recognition in Alzheimer’s disease and in its assessment

    L. Zhou, K. C. Fraser, F. Rudzicz et al., “Speech recognition in Alzheimer’s disease and in its assessment.” in Proc. Interspeech, 2016, pp. 1948–1952

  17. [25]

    Impact of ASR on Alzheimer’s disease detection: All errors are equal, but deletions are more equal than others,

    A. Balagopalan, K. Shkaruta, and J. Novikova, “Impact of ASR on Alzheimer’s disease detection: All errors are equal, but deletions are more equal than others,” arXiv preprint arXiv:1904.01684 , 2019

  18. [26]

    Useful blunders: Can automated speech recognition errors improve downstream demen- tia classification?

    C. Li, W. Xu, T. Cohen, and S. Pakhomov, “Useful blunders: Can automated speech recognition errors improve downstream demen- tia classification?” Journal of Biomedical Informatics, vol. 150, p. 104598, 2024

  19. [27]

    Can automated speech recognition errors pro- vide valuable clues for alzheimer’s disease detection?

    Y .-L. Liu, R. Feng, Y .-X. Lu, J.-X. Chen, Y . Ai, J.-H. Yuan, and Z.-H. Ling, “Can automated speech recognition errors pro- vide valuable clues for alzheimer’s disease detection?” in Proc. ICASSP, 2025, pp. 1–5

  20. [28]

    BERT: Pre- training of deep bidirectional transformers for language under- standing,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre- training of deep bidirectional transformers for language under- standing,” arXiv preprint arXiv:1810.04805, 2018

  21. [29]

    Exploring linguistic feature and model combination for speech recognition based automatic ad detection,

    Y . Wang, T. Wang, Z. Ye, L. Meng, S. Hu, X. Wu, X. Liu, and H. Meng, “Exploring linguistic feature and model combination for speech recognition based automatic ad detection,” inProc. In- terspeech, 2022, pp. 3328–3332

  22. [30]

    Disflu- encies and fine-tuning pre-trained language models for detection of Alzheimer’s disease

    J. Yuan, Y . Bian, X. Cai, J. Huang, Z. Ye, and K. Church, “Disflu- encies and fine-tuning pre-trained language models for detection of Alzheimer’s disease.” in Proc. Interspeech, 2020, pp. 2162– 2166

  23. [31]

    Cosyvoice 2: Scalable stream- ing speech synthesis with large language models,

    Z. Du, Y . Wang, Q. Chen, X. Shi, X. Lv, T. Zhao, Z. Gao, Y . Yang, C. Gao, H. Wanget al., “Cosyvoice 2: Scalable stream- ing speech synthesis with large language models,” arXiv preprint arXiv:2412.10117, 2024

  24. [32]

    Leveraging pretrained representations with task-related key- words for alzheimer’s disease detection,

    J. Li, K. Song, J. Li, B. Zheng, D. Li, X. Wu, X. Liu, and H. Meng, “Leveraging pretrained representations with task-related key- words for alzheimer’s disease detection,” in Proc. ICASSP, 2023, pp. 1–5

  25. [33]

    A two- step attention-based feature combination cross-attention system for speech-based dementia detection,

    Y . Pan, B. Mirheidari, D. Blackburn, and H. Christensen, “A two- step attention-based feature combination cross-attention system for speech-based dementia detection,” IEEE Transactions on Au- dio, Speech and Language Processing, 2025

  26. [34]

    An explain- able ai approach to speech-based alzheimer’s dementia screen- ing,

    F. Iqbal, Z. S. Syed, M. S. S. Syed, and A. S. Syed, “An explain- able ai approach to speech-based alzheimer’s dementia screen- ing,” in Proc. SMM, 2024, pp. 11–15

  27. [35]

    Unveiling interpretability in self-supervised speech representations for parkinson’s diagnosis,

    D. Gimeno-G ´omez, C. Botelho, A. Pompili, A. Abad, and C.-D. Mart´ınez-Hinarejos, “Unveiling interpretability in self-supervised speech representations for parkinson’s diagnosis,” arXiv preprint arXiv:2412.02006, 2024

  28. [36]

    Influence of the interviewer on the automatic assessment of alzheimer’s disease in the context of the adresso challenge,

    P. P ´erez-Toro, S. Bayerl, T. Arias-Vergara, J. V ´asquez-Correa, P. Klumpp, M. Schuster, E. N ¨oth, J. Orozco-Arroyave, and K. Riedhammer, “Influence of the interviewer on the automatic assessment of alzheimer’s disease in the context of the adresso challenge,” in Proc. Inte...

  29. [37]

    Mrc psycholinguistic database: Machine-usable dic- tionary, version 2.00,

    M. Wilson, “Mrc psycholinguistic database: Machine-usable dic- tionary, version 2.00,”Behavior research methods, instruments, & computers, vol. 20, no. 1, pp. 6–10, 1988

  30. [38]

    What does bert learn about the structure of language?

    G. Jawahar, B. Sagot, and D. Seddah, “What does bert learn about the structure of language?” in Proc. ACL, 2019

  31. [39]

    Layer-wise analysis of a self-supervised speech representation model,

    A. Pasad, J.-C. Chou, and K. Livescu, “Layer-wise analysis of a self-supervised speech representation model,” in Proc. ASRU, 2021, pp. 914–921

  32. [40]

    Attentive pooling networks,

    C. d. Santos, M. Tan, B. Xiang, and B. Zhou, “Attentive pooling networks,” arXiv preprint arXiv:1602.03609, 2016

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.