{"id":"4221976f-81d1-43e8-80c0-580dd001322f","arxiv_id":"2501.14994","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Whisper fine-tuned on the SAP-1005 Parkinson's speech dataset reaches 10.71% word error on held-out speakers and 39.56% word error on the cross-etiology TORGO dataset.","lead":"The authors fine-tuned OpenAI's Whisper speech recognition model on a new dataset of speech from people with Parkinson's disease, then tested it on people with different speech disorders. It reports low word error rates on held-out speakers and moderate but speaker-dependent results when transferred to a second dataset.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Speaker-independent claim is not fully established because the dev-shared validation set and dev-unshared test set may share the same PD speakers; checkpoint selection on the test voices can inflate the reported CER/WER.","rationale":"The reader's weakest assumption is that the official SAP-1005 split guarantees no speaker appears in both train and dev-unshared, and no dev-unshared prompt appears in training. That is a necessary condition, and I agree it is unverified. However, the more specific and equally load-bearing condition is that the validation set used for checkpoint selection (dev-shared) does not share speakers with the test set (dev-unshared). The paper's own description of SAP-1005 defines shared/unshared by prompt overlap, not by speaker identity, so the same 21 dev speakers may appear in both portions. If so, choosing the best checkpoint on dev-shared is a model-selection step performed on target-speaker voices, which can bias the reported speaker-independent numbers upward even when train and dev-unshared are perfectly speaker-disjoint. I do not see this as a reason to reject the paper outright: the measured numbers are internally plausible, the TORGO cross-etiology result is clearly separated from training, and the concern is empirically checkable. But until the metadata check and the speaker-disjoint validation rerun are done, the central claim should remain conditional. My agreement with the reader is therefore partial: the train/dev-unshared check is important, but the validation/test speaker-overlap path is the sharper version of the same underlying worry.","tokens_in":8835,"tokens_out":11491,"duration_ms":111786,"concrete_test":"Obtain the SAP-1005 partition metadata and compute three set intersections: speaker IDs between train and dev-unshared, prompt strings between train transcripts and dev-unshared, and speaker IDs between dev-shared and dev-unshared. If the third intersection is nonempty, re-run fine-tuning with identical hyperparameters but with validation drawn from a speaker-disjoint subset of dev-unshared (or from held-out train speakers), select the checkpoint on that validation set, and recompute dev-unshared CER/WER. If the re-run values are within about 1 absolute WER point of 10.71, the concern is minor; if they move by several points or alter severity/category rankings, the speaker-independent claim is inflated and the verdict should tighten.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing condition for the headline 6.99% CER / 10.71% WER is that dev-unshared is isolated from all training-time model-selection decisions. The paper uses dev-shared as the validation set (§IV-B). The SAP-1005 partition described in §III-A defines 'shared'/'unshared' by text-prompt overlap with the train split, not by speaker identity. Since dev is a 21-speaker set split into shared and unshared portions, the same speakers can appear in both dev-shared (validation) and dev-unshared (test). If they do, the early-stopping/checkpoint-selection step is performed on the very speakers later scored, which is a form of target-speaker leakage. A checkpoint that happens to generalize to those 21 voices is preferred, so the reported speaker-independent CER/WER may overstate what a model selected on speaker-disjoint validation would achieve. The reader's train/dev-unshared disjointness assumption is necessary but not sufficient: even with train/dev-unshared speaker disjointness, validation/test overlap creates the same inflation. Because the paper provides no metadata check or code, this is the least secure point in the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a speaker-independent dysarthric speech recognition system based on fine-tuning the multilingual Whisper medium model on the SAP-1005 dataset of Parkinson's disease speech. The authors use the dev unshared subset as a test set and the dev shared subset as validation, reporting a CER of 6.99% and WER of 10.71% on dev unshared, with breakdowns by severity and sentence category. They also evaluate the SAP-fine-tuned model on continuous speech from the TORGO dataset (cerebral palsy and ALS) without further adaptation, reporting an average CER of 25.08% and WER of 39.56%, and interpret this as evidence of cross-etiology generalization.","tokens_in":9080,"tokens_out":4765,"duration_ms":42599,"significance":"If the speaker-independence claim holds, the Whisper-based system provides a strong and useful baseline for the recently released SAP-1005 dataset and a benchmark for cross-etiology transfer from Parkinson's disease to CP/ALS dysarthria. The stratified results by severity and sentence category are informative, and the cross-etiology TORGO results are a valuable data point for the community. However, the validity of the headline numbers depends on the exact speaker partition between training, validation, and test, which is not established in the manuscript; the paper also relies on an admitted comparison to a different evaluation split.","major_comments":[{"comment":"The speaker-independent claim is not fully established because the validation and test sets may share speakers. Section III-A defines 'shared' and 'unshared' partitions by text-prompt overlap with the training set, not by speaker identity, and Section IV-B states that the dev shared set is used as the validation set for checkpoint selection while the dev unshared set is used as the test set. Since the 21-speaker dev set is split into shared and unshared portions, the same speakers can appear in both partitions. If so, selecting the best checkpoint on dev shared tunes the model to those speakers, and the reported 6.99% CER / 10.71% WER on dev unshared would overstate speaker-independent performance. The authors should verify speaker-ID disjointness between dev shared and dev unshared, and ideally provide per-speaker results or a validation split that is speaker-disjoint from the test speakers.","section":"Section IV-B and Section III-A"},{"comment":"The headline comparison with Zheng et al.'s 26.92% WER is not apples-to-apples, as the manuscript itself admits in the text: 'the comparison is limited by the fact that we evaluated different portions of the dataset.' Reporting a 60.24% relative improvement over a method evaluated on a different split overstates the contribution. The authors should either reproduce the baseline on the exact same split (e.g., dev shared or dev unshared with comparable settings) or clearly label the comparison as indicative rather than as a relative improvement.","section":"Section V-A and Table I"},{"comment":"The reported aggregate error rates lack confidence intervals, per-speaker breakdowns, and significance tests. In Table I, the overall averages (6.99% CER, 10.71% WER) are computed across severity groups that themselves vary widely (e.g., High severity CER 20.71% vs. Very Low 4.73%), so the average is sensitive to the composition of the dev unshared set. In Table II, per-speaker WER on TORGO ranges from 2.12% (M03) to 76.4% (M04), so the overall average of 39.56% is highly variable across only eight speakers. At minimum, per-speaker results and bootstrap confidence intervals are needed to assess the stability of the speaker-independent and cross-etiology claims.","section":"Tables I and II"},{"comment":"The chunking procedure for utterances longer than 30 seconds is described as splitting into 30-second chunks with a 5-second overlap, with each chunk decoded individually and the transcripts concatenated. The paper does not specify how the overlapped regions are resolved (e.g., deduplication of repeated words at boundaries) and does not provide a comparison against a no-chunking baseline or against different chunking overlaps. Since the authors attribute residual errors in spontaneous speech partly to hallucination in long utterances, an ablation would help establish the contribution of the chunking step.","section":"Section IV-B"}],"minor_comments":[{"comment":"The sentence 'these prompts are recorded by entirely different speakers, making SAP-1005 a speaker-independent dataset' conflates prompt sharing with speaker independence; the speaker-independent property requires explicit speaker disjointness between the train set and the dev/test sets, which is not described.","section":"Section III-A"},{"comment":"The legend order 'HMLVL' on the y-axis is not self-explanatory; the authors should spell out the severity levels or provide a note that H, M, L, VL denote High, Median (or Moderate), Low, and Very Low.","section":"Figure 2"},{"comment":"The text states that 'error rates exceeding 100% highlight these challenges' in relation to Figure 1; the figure's y-axis extends beyond 100%, but the caption should explicitly mention that CER and WER can exceed 100% due to insertions.","section":"Section V-A"},{"comment":"Reference [28] is listed as 'accepted' without a venue or page numbers, and reference [25] lacks university and year details; these should be completed for reproducibility.","section":"References"},{"comment":"The TORGO section reports averages over severity groups but does not state how many utterances per speaker were used in the evaluation; this information is needed to interpret the per-speaker error rates and the overall average.","section":"Section V-B"}],"recommendation":"major_revision","confidential_remarks":"The paper is a straightforward application of Whisper fine-tuning to a new public dataset; novelty is incremental, but the SAP-1005 evaluation and the cross-etiology TORGO results are timely and useful. The main technical risk is the potential speaker overlap between the validation (dev shared) and test (dev unshared) sets, which, if present, would invalidate the headline speaker-independent claim. The authors can address this by verifying speaker disjointness and, if needed, re-running the analysis with a proper speaker-disjoint validation protocol. I recommend major revision rather than rejection because the issue is fixable within the scope of the manuscript and the underlying data are valuable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful thing here is the baseline: fine-tuned Whisper medium multilingual on SAP-1005 gives 6.99% CER / 10.71% WER on dev-unshared, and transfers to TORGO at 25.08% / 39.56% CER/WER without adaptation. That is new for the SAP-1005 community, and the paper writes it up honestly.\n\nThe main soft spot is the split. SAP-1005's dev shared/unshared division is by text-prompt overlap with train, not by speaker. The 21 dev speakers appear in both shared and unshared portions. The authors use dev-shared as the validation set and dev-unshared as the test set. That means the checkpoint-selection step sees the same speakers that are later scored, a target-speaker leak. The reported speaker-independent numbers are therefore optimistic. This is not a trivial nit; the central claim is 'speaker-independent,' and the leak affects exactly that claim.\n\nCredit where due: the paper handles Whisper's 30-second context by chunking long utterances with overlap, breaks down results by severity and sentence category (useful for the community), and explicitly acknowledges that the headline comparison to Zheng et al. uses different evaluation splits. The cross-etiology experiment on TORGO is suggestive, though it rests on eight dysarthric speakers and severity subgroups of two or three, so the per-group averages are noisy.\n\nOther issues are more minor: no confidence intervals, no code or checkpoints, no per-speaker SAP scores. The self-citation to the prior Whisper evaluation is reasonable and does not bother me.\n\nWho should read it: anyone building on SAP-1005 or looking for a strong generic ASR baseline in dysarthria research. It deserves a serious referee because the numbers are useful and the split problem is fixable—the authors should verify speaker IDs, rerun with a speaker-disjoint validation set, or at least report both analyses. I would send it out.\n\nCandidly, the paper's value is as a benchmark data point, not a method contribution.","headline":"Whisper baseline on SAP-1005 is useful, but validation and test share the same speakers, undercutting the speaker-independence claim.","tokens_in":9633,"tokens_out":2797,"would_cite":true,"duration_ms":26073,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a single Whisper model fine-tuned on Parkinson's dysarthric speech reaches near-typical error rates on unseen speakers and transfers without retraining to cerebral palsy and ALS speech.","keywords":["dysarthric speech recognition","speaker-independent ASR","Whisper fine-tuning","Speech Accessibility Project","SAP-1005","TORGO dataset","cross-etiology transfer","Parkinson's disease"],"falsifier":"Re-identify every speaker ID and text prompt in the SAP-1005 dev-unshared set and check whether any appears in the training partition; then re-run the same Whisper fine-tune on a verified disjoint split and compare the CER and WER to 6.99% and 10.71%.","tokens_in":8638,"feed_emoji":"🗣️","tokens_out":5622,"duration_ms":44580,"temperature":0.7,"pith_summary":"The paper tries to show that a single speaker-independent automatic speech recognition system can handle dysarthric speech well enough to be useful, across both unseen speakers and different underlying neurological conditions. It fine-tunes a medium-sized multilingual Whisper model on the SAP-1005 corpus of Parkinson's disease speech and reports 6.99% character error rate and 10.71% word error rate on the dev-unshared set, a 60.24% relative WER improvement over the earlier wav2vec 2.0 baseline. The same model, without any retraining, reaches 25.08% CER and 39.56% WER on the TORGO corpus of cerebral palsy and ALS speech. If the data split is truly speaker- and prompt-disjoint, this establishes a strong baseline for speaker-independent dysarthric ASR and suggests that acoustic regularities are shared across etiologies.","feed_headline":"One Whisper model hits 6.99% CER on Parkinson's speech","feed_subtitle":"Fine-tuned on SAP-1005, it also transcribes cerebral palsy and ALS speech without retraining.","key_machinery":"The workhorse is the Whisper encoder-decoder transformer, in its medium multilingual variant, fine-tuned on SAP-1005 and decoded with beam search (num_beams=10, no_repeat_ngram_size=3, length_penalty=1.0). Because Whisper's receptive field is 30 seconds, utterances longer than 30 seconds are split into 30-second chunks with a 5-second overlap and decoded separately, then the transcripts are concatenated; this chunking is what keeps long spontaneous prompts from being lost to repetition hallucinations. The fine-tuned model is then applied unchanged to TORGO continuous speech.","core_discovery":"The central claim is that one fine-tuned Whisper model generalizes across dysarthric speakers in a way that conventional speaker-dependent systems do not. Trained only on PD speech from SAP-1005, the model reaches 6.99% CER and 10.71% WER on the dev-unshared subset, and its cross-etiology evaluation on TORGO yields 25.08% CER and 39.56% WER. The authors interpret the TORGO result as evidence that the model captures shared dysarthric speech characteristics rather than PD-specific artifacts. They also find that error rates rise with severity, from 4.73% CER for very low severity to 20.71% CER for high severity, and that spontaneous speech is much harder than digital-assistant commands, with WER of 19.29% versus 7.92%.","pith_inferences":["If the speaker and prompt isolation in SAP-1005 holds, then Whisper's pretraining already supplies enough phonetic generality that a small amount of fine-tuning on one dysarthria population transfers to others; a testable extension is to add a small amount of CP/ALS data and measure whether cross-etiology WER drops proportionally.","The reported comparison to the earlier wav2vec 2.0 baseline is not apples-to-apples because the earlier system was evaluated on a different portion of the corpus; a matched evaluation on the same dev-unshared partition would settle which backbone is actually better.","Because the official SAP-1005 test set is reserved for a competition, the dev-unshared numbers may be optimistic relative to a true held-out test, and readers should treat them as indicative until official test-set results appear.","Some high-severity spontaneous utterances show error rates exceeding 100%, meaning the model sometimes produces transcripts longer than the reference; reporting CER alone may hide such hallucination behavior."],"forward_implications":["A single fine-tuned Whisper model can serve as a strong speaker-independent baseline for dysarthric ASR without per-speaker adaptation.","Learned representations transfer across dysarthria types: a PD-trained model partially handles CP and ALS speech, suggesting shared acoustic regularities across etiologies.","Performance degrades sharply with severity and with spontaneous speech, identifying the high-severity and spontaneous-speech regimes as the main remaining challenges.","Long spontaneous utterances remain a failure mode; chunking mitigates but does not eliminate Whisper's 30-second receptive-field hallucinations."],"supporting_citations":[{"why":"Supplies the Whisper encoder-decoder model that is fine-tuned and defines its 30-second receptive field.","marker":"[13]"},{"why":"Defines the SAP-1005 split, severity labels, and the 26.92% WER wav2vec 2.0 baseline this paper improves on.","marker":"[32]"},{"why":"Earlier comparison study that selected the medium multilingual Whisper variant for dysarthric speech.","marker":"[28]"},{"why":"Provides the TORGO dataset used for the cross-etiology evaluation.","marker":"[29]"},{"why":"Documents the availability of the SAP-1005 corpus used for training.","marker":"[31]"}],"fun_headline_variants":["Whisper model generalizes across dysarthria etiologies: 25.08% CER on TORGO","Speaker-independent dysarthria ASR: Whisper fine-tune achieves 6.99% CER on PD","From PD to CP and ALS: Whisper model proves robust for dysarthric speech","Whisper fine-tuned on PD speech also handles CP and ALS: 25.08% CER","Cross-etiology dysarthria recognition: one Whisper model, unseen speakers, 10.71% WER"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the SAP-1005 train and dev-unshared sets share no speakers and no text prompts, so that the reported error rates measure generalization to truly unseen speakers; if the official split leaks speakers or prompts, the headline numbers are optimistic.","fun_headline_variants_meta":{"raw":{"variants":["Whisper model generalizes across dysarthria etiologies: 25.08% CER on TORGO","Speaker-independent dysarthria ASR: Whisper fine-tune achieves 6.99% CER on PD","From PD to CP and ALS: Whisper model proves robust for dysarthric speech","Whisper fine-tuned on PD speech also handles CP and ALS: 25.08% CER","Cross-etiology dysarthria recognition: one Whisper model, unseen speakers, 10.71% WER"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001205,"raw_usage":{"total_tokens":4965,"prompt_tokens":948,"completion_tokens":4017,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":3884}},"tokens_in":564,"tokens_out":4017,"duration_ms":25109,"temperature":1.0,"reasoning_tokens":3884,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:43:04.339805+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-identify every speaker ID and text prompt in the SAP-1005 dev-unshared set and check whether any appears in the training partition; then re-run the same Whisper fine-tune on a verified disjoint split and compare the CER and WER to 6.99% and 10.71%.","supporting_citations":[{"cited_title":"Robust speech recognition via large-scale weak supervi- sion,","cited_arxiv_id":null,"evidence_quote":"Supplies the Whisper encoder-decoder model that is fine-tuned and defines its 30-second receptive field."},{"cited_title":"Fine-Tuning Auto- matic Speech Recognition for People with Parkinson’s: An Effective Strategy for Enhancing Speech Technology Accessibility,","cited_arxiv_id":null,"evidence_quote":"Defines the SAP-1005 split, severity labels, and the 26.92% WER wav2vec 2.0 baseline this paper improves on."},{"cited_title":"A Comprehensive Performance Evaluation of Whisper Models in Dysarthric Speech Recognition,","cited_arxiv_id":null,"evidence_quote":"Earlier comparison study that selected the medium multilingual Whisper variant for dysarthric speech."},{"cited_title":"The TORGO database of acoustic and articulatory speech from speakers with dysarthria,","cited_arxiv_id":null,"evidence_quote":"Provides the TORGO dataset used for the cross-etiology evaluation."},{"cited_title":"UIUC Leads Effort to Make Speech Recognition Technology More Inclusive,","cited_arxiv_id":null,"evidence_quote":"Documents the availability of the SAP-1005 corpus used for training."}],"review_version":1}