{"id":"78eb3711-0d05-4e26-90b7-85e4c2eb03c3","arxiv_id":"2606.31722","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Fine-tuning Whisper on over 100 hours of personalized dysarthric speech data reduces WER to 9.7% for a single speaker.","lead":"This paper adapts the Whisper foundation model to one dysarthric speaker using 92 hours of read speech plus 8.8 hours of app corrections, reaching 9.7% word error rate. A smart generalist might read it to see how large pre-trained models can be made usable for people with speech impairments through personalization.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Single-speaker read-speech results plus corrections do not establish suitability for practical deployment","rationale":"The reader's weakest assumption directly identifies the same load-bearing gap for the deployment claim. Because the paper is framed as a single-speaker case study, the absence of spontaneous-speech testing is the precise point where the strongest claim becomes unsupported rather than an internal inconsistency or statistical flaw.","tokens_in":1652,"tokens_out":281,"duration_ms":14511,"concrete_test":"Run the final fine-tuned model on a new held-out set of spontaneous conversational utterances from the same speaker (minimum 30 min); if WER exceeds 20% or rises >5 points above the 9.7% read-speech figure, the practical-deployment claim is unsupported by the current evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the reported WER reductions (15.8% at 1.4 h, 10.7% at 22.5 h, 9.7% with all data) translate to usable performance in everyday communication. The study uses only one speaker, 92 h read speech plus 8.8 h app corrections, and reports no evaluation on spontaneous conversational dysarthric speech. Without that test, the deployment suitability assertion rests on an unverified extrapolation from the controlled collection protocol.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper presents a case study adapting the Whisper foundation ASR model to one dysarthric speaker via fine-tuning on 92 hours of read speech (collected with TEQST) plus 8.8 hours of mobile-app corrections. It reports WER dropping to 15.8% with 1.4 h adaptation data, 10.7% with 22.5 h, and 9.7% with all data; LoRA adaptation and Qwen3-ASR performed worse. The abstract concludes that personalized fine-tuning makes foundation models substantially more effective for dysarthric speech and suitable for practical deployment.","tokens_in":1761,"tokens_out":488,"duration_ms":21173,"significance":"If the empirical results hold under broader testing, the work supplies concrete evidence that large amounts of speaker-specific data, including real user corrections, can produce usable WERs for dysarthric speech in a read-speech setting. This could guide data-collection practices for accessibility applications. The single-speaker, read-speech design, however, restricts claims about generalizability or everyday use.","major_comments":[{"comment":"Abstract: the baseline WER of the unadapted Whisper model on the target speaker is never stated. Without it, the magnitude of the reported reductions (15.8 %, 10.7 %, 9.7 %) cannot be assessed, which is load-bearing for the central claim that fine-tuning makes the model “substantially more effective.”","section":"Abstract"},{"comment":"Abstract and Results: all reported numbers come from read speech of a single speaker; no evaluation on spontaneous conversational dysarthric speech is described. This directly undermines the assertion of suitability for “practical deployment,” as everyday communication is not limited to read material.","section":"Abstract"},{"comment":"Results: no error bars, statistical tests, or repeated runs are mentioned, and the only comparisons are to LoRA and Qwen3-ASR. The absence of these controls leaves the reliability of the WER improvements open to question.","section":"Results"}],"minor_comments":[{"comment":"Abstract: the acronym TEQST is introduced without expansion.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We address each major comment below and will make targeted revisions to improve the manuscript.","responses":[{"response":"We agree that the baseline is required to evaluate the improvements. We will revise the abstract to explicitly state the unadapted Whisper WER on the target speaker (reported in the results section of the full manuscript).","revision_made":"yes","referee_comment":"[Abstract] Abstract: the baseline WER of the unadapted Whisper model on the target speaker is never stated. Without it, the magnitude of the reported reductions (15.8 %, 10.7 %, 9.7 %) cannot be assessed, which is load-bearing for the central claim that fine-tuning makes the model “substantially more effective.”"},{"response":"The paper is framed as a single-speaker case study on read speech (as stated in the title and introduction). We will revise the abstract and conclusion to qualify the 'practical deployment' language, making clear that results apply to read speech and that spontaneous speech evaluation is left for future work.","revision_made":"partial","referee_comment":"[Abstract] Abstract and Results: all reported numbers come from read speech of a single speaker; no evaluation on spontaneous conversational dysarthric speech is described. This directly undermines the assertion of suitability for “practical deployment,” as everyday communication is not limited to read material."},{"response":"Given the single-speaker design and the substantial effort required to collect over 100 hours of data, multiple independent runs were not performed. We will add a limitations paragraph noting this constraint and highlighting the consistent WER trend across data scales. The LoRA and Qwen3-ASR comparisons address relevant adaptation and model alternatives.","revision_made":"partial","referee_comment":"[Results] Results: no error bars, statistical tests, or repeated runs are mentioned, and the only comparisons are to LoRA and Qwen3-ASR. The absence of these controls leaves the reliability of the WER improvements open to question."}],"tokens_in":1379,"tokens_out":458,"duration_ms":24689,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this paper reports new WER numbers from fine-tuning Whisper on a large personalized dataset for one dysarthric speaker, including corrections collected through a deployed app.\n\nThey gathered 92 hours of read speech plus 8.8 hours of app corrections, then showed WER falling to 15.8% with 1.4 hours of adaptation data, 10.7% with 22.5 hours, and 9.7% with the full set. LoRA and Qwen3-ASR performed worse. This gives a clear data point on how adaptation data volume affects performance in this domain and demonstrates a workable collection pipeline.\n\nThe soft spots are straightforward. Everything is one speaker and mostly read speech. The abstract supplies no baseline WER for the unadapted model on this data, no error bars, no statistical tests, and no results on spontaneous conversational speech. The assertion that the system is suitable for practical deployment therefore rests on an untested jump from controlled read material to everyday use.\n\nThis paper is for people working on ASR adaptation for impaired speech or accessibility applications. Readers who want empirical numbers on data scale for dysarthric fine-tuning will find it useful. It deserves a serious referee because the results are new empirical measurements on a substantial dataset, even though the deployment language needs to be walked back.\n\nI would send it to peer review and ask the authors to add the missing baseline, report variance, and test on spontaneous speech.","headline":"Fine-tuning Whisper on 100+ hours from one dysarthric speaker reaches 9.7% WER, but the single-speaker read-speech setup does not support the practical deployment claim.","tokens_in":2234,"tokens_out":388,"would_cite":false,"duration_ms":18947,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Personalized fine-tuning of Whisper on one dysarthric speaker's data reaches 9.7 percent word error rate.","keywords":["dysarthric speech","ASR adaptation","Whisper model","personalized fine-tuning","word error rate","read speech","mobile app corrections","foundation models"],"falsifier":"A new test of the adapted model on spontaneous conversational speech or on recordings from additional dysarthric speakers that shows error rates remaining high would falsify the suitability-for-deployment claim.","tokens_in":2557,"feed_emoji":"🎙️","tokens_out":656,"duration_ms":21060,"temperature":0.7,"pith_summary":"The paper demonstrates that a general foundation ASR model can be adapted to handle dysarthric speech through speaker-specific data collection and training. Standard systems fail on such speech, restricting daily communication for affected individuals. Using 92 hours of read speech plus 8.8 hours of corrections gathered via a mobile app, fine-tuning Whisper steadily lowers error rates as data volume increases. The approach outperforms both LoRA adaptation and an alternative base model in this single-speaker setting.","feed_headline":"Whisper fine-tuned on dysarthric speech reaches 9.7% WER","feed_subtitle":"One speaker's 100 hours of read data plus app feedback makes a foundation model usable for daily communication.","key_machinery":"Personalized fine-tuning of the Whisper foundation ASR model on speaker-specific read speech collected via TEQST and user corrections from a deployed mobile application.","core_discovery":"Starting from the Whisper foundation model, fine-tuning on 1.4 hours of adaptation data yields 15.8 percent word error rate, on 22.5 hours yields 10.7 percent, and on the full set of 92 hours read speech plus 8.8 hours of user corrections yields 9.7 percent; this shows personalized fine-tuning makes foundation ASR models substantially more effective for dysarthric speech and suitable for practical deployment.","pith_inferences":["The single-speaker protocol could be repeated with additional dysarthric individuals to test whether similar data volumes produce comparable gains.","Performance on read speech may not directly predict results in unscripted conversation, suggesting a need for spontaneous-speech test sets.","The same data-collection and fine-tuning steps could be applied to other forms of atypical speech to check for similar error-rate reductions."],"forward_implications":["Word error rate decreases as the amount of speaker-specific adaptation data grows from 1.4 hours to the full collection.","Incorporating 8.8 hours of corrections collected through the mobile app produces the lowest error rate achieved.","Applying LoRA adaptation instead of full fine-tuning results in higher word error rates.","Using Qwen3-ASR as the starting foundation model also produces worse performance than starting from Whisper."],"fun_headline_variants":["Whisper fine-tuned to 9.7% WER for dysarthric speech","9.7% WER with Whisper adapted to dysarthric speaker","Dysarthric ASR improved to 9.7% WER via Whisper fine-tuning","Fine-tuning Whisper on dysarthric data achieves 9.7% WER"],"cache_read_input_tokens":64,"weakest_assumption_plain":"That results from one speaker using read speech and app corrections indicate the method suits practical deployment for dysarthric users in general.","fun_headline_variants_meta":{"raw":{"variants":["Whisper fine-tuned to 9.7% WER for dysarthric speech","9.7% WER with Whisper adapted to dysarthric speaker","Dysarthric ASR improved to 9.7% WER via Whisper fine-tuning","Fine-tuning Whisper on dysarthric data achieves 9.7% WER"]},"model":"grok-4.3","cost_usd":0.009152,"raw_usage":{"total_tokens":4087,"prompt_tokens":637,"num_sources_used":0,"completion_tokens":86,"cost_in_usd_ticks":91524500,"prompt_tokens_details":{"text_tokens":637,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3364,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":637,"tokens_out":86,"duration_ms":28419,"temperature":1.0,"reasoning_tokens":3364,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-01T05:49:44.493609+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A new test of the adapted model on spontaneous conversational speech or on recordings from additional dysarthric speakers that shows error rates remaining high would falsify the suitability-for-deployment claim.","supporting_citations":[],"review_version":1}