{"id":"2529e7b1-ef1f-4a3f-85c5-872536469937","arxiv_id":"2501.03536","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A comprehensive overview of machine learning methods for detecting, recognizing, assessing, and enhancing pathological speech caused by neurodegenerative disorders, along with a dataset catalog.","lead":"This paper is a review of automated speech technologies for people with neurodegenerative diseases such as Parkinson's or ALS. It organizes research on detecting, recognizing, and enhancing atypical speech, and lists the datasets used.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table I contains verified internal inconsistencies (AMSDC columns swapped, COPAS control/pathological reversed, mPower speaker counts not summing to stated total, Saarbrücken total mismatch), so the dataset reference contribution is not yet reliable; these errors must be corrected before the…","rationale":"I focused on the dataset table rather than the \"first comprehensive survey\" claim because the numerical inconsistencies are concrete, internally verifiable, and directly undercut the paper's usefulness as a reference, which is the strongest interpretation of its central contribution. The \"first\" claim is contestable but hinges on literature coverage and timing; even if another review existed, the survey's value would be its accurate snapshot, so the table errors are more load-bearing. The reader identified AMSDC, COPAS, and mPower; my independent check confirms those and adds Saarbrücken. I recommend keeping the verdict CONDITIONAL: the narrative is useful, the organization is clear, and the errors are localized to Table I and its accompanying text, but a reference-quality contribution cannot ship with these numbers. If the authors correct the table, the paper would be acceptable; no evidence requires rejection.","tokens_in":34966,"tokens_out":5281,"duration_ms":48641,"concrete_test":"Reconstruct Table I row by row from the primary dataset papers [52]-[69], recording the exact speaker counts and group definitions. Verify at minimum: AMSDC (99 patients, 62M/37F, 0 controls), COPAS (122 controls/197 patients), mPower (1,087 PD + 5,581 controls = 6,668 vs. text's 6,805 total), and Saarbrücken (1,356 + 869 = 2,225 vs. text's 2,255 total). If any of these checks fails, the table and the corresponding text must be corrected and the paper reissued; if all pass, the reader's condition is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section II presents Table I as the \"summary of these datasets and their characteristics.\" Four rows show objective arithmetic or column-semantics failures. (1) AMSDC: the text (Sec. II) says the corpus has 99 patients (62 male, 37 female) and no control speakers, but Table I lists Control=62, Pathological=37, i.e., sex counts are misplaced into the wrong columns. (2) COPAS: the text states \"197 pathological speakers and 122 control speakers,\" but Table I lists Control=197, Pathological=122, swapping the two quantities. (3) mPower: the text reports 6,805 participants, with 1,087 PD and 5,581 controls; Table I's row sums to 6,668, an unexplained 137-person discrepancy (or an error in the stated total). (4) Saarbrücken Voice Database: the text says \"2,255 German speakers\" and lists 1,356 patients and 869 controls; 1,356+869=2,225, not 2,255. These are not interpretive disagreements; they are factual and arithmetic errors in the paper's explicitly claimed \"thorough list\" of datasets. Because the dataset catalog is one of the paper's two stated contributions, the central claim that this is a reliable comprehensive survey is currently weakened.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is an overview article on automatic speech analysis and processing technologies for speech affected by neurodegenerative disorders. It reviews pathological speech datasets (Section II), speech representations (Section III), automatic detection (Section IV), ASR for pathological speech (Section V), intelligibility enhancement (Section VI), intelligibility and severity assessment (Section VII), data augmentation (Section VIII), and challenges and future directions (Section IX). The two stated contributions are a comprehensive survey spanning detection, recognition, enhancement, and assessment, and a structured catalog of accessible and non-accessible pathological speech datasets.","tokens_in":35359,"tokens_out":10387,"duration_ms":96389,"significance":"If corrected, the paper would be a useful entry point for researchers and clinicians working on pathological speech. Its organization by task rather than by disorder is helpful, and the inclusion of dataset accessibility labels, robustness concerns, privacy, explainability, and multimodal/LLM directions is timely. The manuscript does not claim new algorithms or experimental results; its value lies in synthesis and bibliographic coverage, so the accuracy of its dataset table and the fidelity of its summaries of cited works are central to its reliability.","major_comments":[{"comment":"The dataset catalog, one of the paper's two stated contributions, contains multiple verified numerical errors and must be corrected against the primary sources. (i) AMSDC: the text says 99 patients (62 male, 37 female) and no controls, but the table lists Control=62, Pathological=37, placing sex counts in the wrong columns. (ii) Italian Parkinson's Database: the text says 28 patients and 37 controls, but the table lists Control=28, Pathological=37, a second column swap. (iii) COPAS: the text says 197 pathological and 122 control speakers, but the table lists Control=197, Pathological=122. (iv) mPower: the text states 6,805 participants, 1,087 PD and 5,581 controls, while 1,087+5,581=6,668; the table repeats the 1,087/5,581 breakdown, so either the total or the group sizes are wrong. (v) Saarbrücken: the text states 2,255 German speakers, but 1,356 patients + 869 controls = 2,225, not 2,255; the table matches the components, so the total in the text is the error. (vi) NeuroVoz: the text states 54 patients and 58 controls, but the subgroup counts (33+20 patients; 28+26+1 controls) sum to 53 and 55, respectively. These arithmetic and column-semantics errors undermine the reliability of the dataset catalog and must be fixed before publication.","section":"Section II, Table I"},{"comment":"The claim to present 'the first comprehensive survey' is load-bearing for the paper's novelty. The related works [41-51] include broad reviews of pathological speech processing and of PD/dementia detection, and the manuscript currently distinguishes them only by saying they are narrower or outdated. To make the novelty claim verifiable, the authors should either provide a systematic comparison of tasks, disorders, and time coverage against those reviews, or revise the claim to 'a comprehensive survey' with the intended scope stated explicitly.","section":"Section I, Contribution 1"}],"minor_comments":[{"comment":"The table contains typographical errors that should be corrected: 'Dysarthira' appears twice, 'CV A' and 'EW A-DB' should be 'CVA' and 'EWA-DB', and 'Saarbrucken' should be 'Saarbrücken'.","section":"Section II, Table I"},{"comment":"The text describes the impairment as 'cerebellar degeneration' while Table I says 'spino-cerebellar ataxia'; use consistent terminology or clarify that these refer to the same diagnostic category.","section":"Section II, CUDYS"},{"comment":"The phrase 'from a clinical and technological perspectives' is ungrammatical; use 'from clinical and technological perspectives' or 'from a clinical and a technological perspective'.","section":"Section I, Contribution 1"},{"comment":"The bibliographic entry lists IEEE Transactions on Speech and Audio Processing with volume 255, which is not a plausible volume number; verify and correct the reference.","section":"References, [30]"},{"comment":"The paper does not describe how the surveyed literature and dataset information were selected or verified; adding a brief statement on search databases, years, and inclusion criteria would make the comprehensiveness claim more transparent.","section":"Section II"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of an overview journal. The main obstacle to acceptance is data accuracy: Table I and the Section II bullet counts need an independent verification pass. The 'first comprehensive survey' claim should be moderated or substantiated with a systematic scope comparison. I found no evidence of methodological circularity, since the paper makes no derivational claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good to see this survey. The genuine value here is organizational: it brings detection, ASR, intelligibility enhancement, severity/intelligibility assessment, and data augmentation under one roof, with a structured dataset table and feature taxonomy. For a graduate student entering pathological speech processing, or a clinician trying to understand what the field offers, this saves real time. The narrative in the main sections is accurate in broad strokes and the reference list is wide; the per-area summaries of classical ML, deep learning, SSL, and augmentation all match the state of the art as I know it.\n\nThe soft spot is exactly where the reader's report points: the dataset table (Table I) has four verifiable arithmetic or column-semantics failures — AMSDC's sex counts are placed in Control/Pathological columns, COPAS's control and pathological counts are swapped, mPower's row sums to 6,668 against a stated 6,805, and the Saarbrücken row sums to 2,225 against a stated 2,255. These are not interpretive disagreements; they are errors in a table the paper explicitly promotes as a key deliverable. A catalog with wrong and inconsistent numbers compromises the paper's usefulness as a reference, even though the surrounding narrative is sound.\n\nAbout the 'first comprehensive survey' claim: it is contestable, since earlier reviews cover substantial subsets, but the scope across detection + recognition + enhancement + assessment is broader than most, so I would not treat that as a fatal flaw. The paper would benefit from toning down the claim, but the contribution stands on its own as a synthesis.\n\nMy verdict: this deserves peer review, conditional on fixing Table I. The authors should re-check every number in the table against the original dataset papers, and add a note about the discrepancy. Once that is done, I would cite this as a comprehensive overview. If the table is not corrected, the paper is still readable but loses its main reference value.\n\nFor your reading group: maybe worth a slot once the corrected version is out, to get everyone pointing at the same map.","headline":"A useful survey map of pathological speech processing, but the dataset catalog (Table I) has arithmetic errors that must be fixed before it serves as a reliable reference.","tokens_in":35676,"tokens_out":1601,"would_cite":true,"duration_ms":14545,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review claims to be the first comprehensive survey of automatic speech technologies for neurodegenerative disorders, spanning detection, recognition, intelligibility enhancement, and assessment, and it provides a catalog of the…","keywords":["pathological speech","neurodegenerative disorders","speech detection","automatic speech recognition","intelligibility enhancement","severity assessment","data augmentation","speech datasets"],"falsifier":"Compare every row of Table I against the primary dataset publications: for example, the text says the AMSDC corpus has 99 patients but the table lists 62 control and 37 pathological speakers, and the mPower table entries do not match the stated 6,805 total participants; a systematic mismatch would show the catalog is not a reliable reference.","tokens_in":34820,"feed_emoji":"🗣️","tokens_out":5821,"duration_ms":49252,"temperature":0.7,"pith_summary":"This paper aims to be the first comprehensive survey of automatic speech technologies for neurodegenerative disorders, covering both the clinical goal of diagnosis and monitoring and the technological goal of making speech-based systems work for impaired speakers. It organizes the field into five problem areas—pathological speech detection, automatic speech recognition, intelligibility enhancement, intelligibility and severity assessment, and data augmentation—and reviews the speech representations used in each. It also compiles a table of more than a dozen pathological speech datasets, marking each as public, accessible, or non-accessible. The paper then identifies cross-cutting challenges such as robustness, privacy, interpretability, and speech mode, and proposes future directions including multimodal analysis and large language models. A sympathetic reader would take the paper's contribution to be a structured map of the field plus a dataset catalog that researchers can use to orient themselves.","feed_headline":"First full survey maps speech AI for neurodegenerative disorders","feed_subtitle":"Detection, recognition, enhancement and assessment of disordered speech, plus a dataset catalog and future directions.","key_machinery":"The organizing device is a taxonomy of the field into five research problems—detection, recognition, enhancement, assessment, and data augmentation—together with a catalog of pathological speech datasets that maps each dataset's size, impairment type, language, modality, and accessibility. This structure does the argument's work: by aligning each subfield with the same set of datasets and speech representations (handcrafted features, time-frequency representations, raw waveforms, self-supervised embeddings), the survey makes the field's fragmentation visible and lets the authors claim comprehensiveness. The definition of pathological speech as speech that deviates from neurotypical patterns due to underlying impairments, with deviations in voice, articulation, prosody, and language, sets the boundaries of the review.","core_discovery":"The paper's central claim is that no existing review covers pathological speech from both clinical and technological perspectives across detection, recognition, enhancement, and assessment, and that this survey fills that gap. On the paper's own terms, the discovery is the scope itself: a unified treatment of how automatic systems can decide whether speech is pathological, recognize what a pathological speaker is saying, make pathological speech more intelligible, estimate intelligibility and severity, and generate or perturb data to train such systems. The paper also documents the field's reliance on a small set of mostly small, often private datasets, and its shift from handcrafted acoustic features toward self-supervised embeddings as the current state of the art. It concludes that robustness, privacy, interpretability, and generalization across speech modes and languages are the open problems, and points to multimodal and large-language-model approaches as the likely next steps.","pith_inferences":["The 'first comprehensive survey' claim is definition-dependent: prior reviews cover single disorders or single tasks, so whether this is genuinely first hinges on accepting the clinical-plus-technological scope as the relevant frame.","The survey's emphasis on self-supervised embeddings implies a concrete testable prediction: detection and recognition systems built on SSL features should dominate leaderboards on any new pathological speech benchmark, while handcrafted-feature systems should lag on larger datasets.","The authors' call for LLM-based and multimodal approaches could be sharpened by a shared benchmark that reports intelligibility gains and word error rates on standardized pathological speech corpora, allowing direct comparison across future systems.","The dataset catalog could naturally evolve into a living community resource, with corrections and additions tracked over time, which would make the survey's practical utility outlast its publication date."],"forward_implications":["Researchers entering the field can use the survey as a single entry point: the dataset list marks which corpora are available, under what conditions, and with which impairments.","The paper's review implies that self-supervised embeddings such as wav2vec2, HuBERT, and WavLM are currently the strongest input representations for detection and recognition, so new work should build on those rather than handcrafted features.","The catalog's accessibility column makes it possible to see why replication is hard: many standard datasets are private, and even public ones such as TORGO carry recording artifacts.","The challenges section argues that future progress will come from robustness to environmental and adversarial distortions, privacy-preserving training, interpretable models, and use of spontaneous speech.","Following the paper's logic, multimodal and large-language-model systems are the designated next research frontier for both clinical assessment and assistive devices."],"supporting_citations":[{"why":"Prior review of speaker-recognition techniques for voice condition analysis; defines the narrower scope this survey widens.","marker":"[41]"},{"why":"Review of articulatory and phonatory aspects of Parkinson's disease detection; limits the field to one disorder.","marker":"[42]"},{"why":"Survey of AI techniques, datasets, and challenges for speech-based Alzheimer's detection; single-disorder scope.","marker":"[43]"},{"why":"Scoping review of interpretable acoustic features in neurodegenerative motor diseases; covers only features, not technologies.","marker":"[44]"},{"why":"Systematic review of deep learning approaches for Parkinson's disease classification; detection task only.","marker":"[45]"},{"why":"Survey of machine-learning and statistical voice analysis for Parkinson's disease; single disorder and analysis-focused.","marker":"[50]"},{"why":"Earlier overview of pathological speech processing challenges that the paper positions itself as updating and broadening.","marker":"[51]"},{"why":"Overview of a dysarthric speech recognition system; supplies the recognition subfield's baseline coverage and adaptation techniques.","marker":"[32]"}],"fun_headline_variants":["Speech AI review covers detection to assistive applications","First survey maps speech AI for neurodegenerative disorders","Unified view of speech tech for diagnosis and assistive care","Speech analysis AI: from disorder detection to support tools"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's usefulness rests on the accuracy of its dataset table and on its summaries of the cited studies being faithful to what those studies actually report.","fun_headline_variants_meta":{"raw":{"variants":["Speech AI review covers detection to assistive applications","First survey maps speech AI for neurodegenerative disorders","Unified view of speech tech for diagnosis and assistive care","Speech analysis AI: from disorder detection to support tools"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000851,"raw_usage":{"total_tokens":3630,"prompt_tokens":806,"completion_tokens":2824,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":422,"completion_tokens_details":{"reasoning_tokens":2762}},"tokens_in":422,"tokens_out":2824,"duration_ms":21246,"temperature":1.0,"reasoning_tokens":2762,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:51:12.344744+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare every row of Table I against the primary dataset publications: for example, the text says the AMSDC corpus has 99 patients but the table lists 62 control and 37 pathological speakers, and the mPower table entries do not match the stated 6,805 total participants; a systematic mismatch would show the catalog is not a reliable reference.","supporting_citations":[{"cited_title":"Innovative Speech-Based Deep Learning Approaches for Parkinson's Disease Classification: A Systematic Review","cited_arxiv_id":"2407.17844","evidence_quote":"Systematic review of deep learning approaches for Parkinson's disease classification; detection task only."}],"review_version":1}