Pith. sign in

REVIEW 4 cited by

Enhancing ASR for Stuttered Speech with Limited Data Using Detect and Pass

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.05396 v1 pith:4DM5BEVA submitted 2022-02-08 eess.AS cs.CLcs.LGcs.SD

classification eess.AScs.CLcs.LGcs.SD
keywords speechstutterdatadetectlimitedpeoplesystemsaccessible
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

It is estimated that around 70 million people worldwide are affected by a speech disorder called stuttering. With recent advances in Automatic Speech Recognition (ASR), voice assistants are increasingly useful in our everyday lives. Many technologies in education, retail, telecommunication and healthcare can now be operated through voice. Unfortunately, these benefits are not accessible for People Who Stutter (PWS). We propose a simple but effective method called 'Detect and Pass' to make modern ASR systems accessible for People Who Stutter in a limited data setting. The algorithm uses a context aware classifier trained on a limited amount of data, to detect acoustic frames that contain stutter. To improve robustness on stuttered speech, this extra information is passed on to the ASR model to be utilized during inference. Our experiments show a reduction of 12.18% to 71.24% in Word Error Rate (WER) across various state of the art ASR systems. Upon varying the threshold of the associated posterior probability of stutter for each stacked frame used in determining low frame rate (LFR) acoustic features, we were able to determine an optimal setting that reduced the WER by 23.93% to 71.67% across different ASR systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection

    eess.AS 2025-05 reject novelty 6.0 of 10

    The paper introduces LLM-Dys, a 12,790-hour synthetic dysfluent speech corpus generated by LLM plus TTS, and claims state-of-the-art dysfluency detection with a Whisper-based transcriber.

  2. Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches

    cs.SD 2025-06 conditional novelty 5.0 of 10

    Personalized LoRA fine-tuning of Whisper on both read and spontaneous stuttered speech from one speaker cuts word error rate roughly in half compared to a generalized model.

  3. Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection

    cs.SD 2025-05 conditional novelty 5.0 of 10

    An LLM-driven multi-task system reports a 5.45% CER and 73.63% average SED F1 on the AS-70 Mandarin stuttering benchmark, though key baselines and uncertainty are missing.

  4. Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A training-free WFST decoder that uses the reference text to constrain phoneme decoding reports large gains in dysfluent speech transcription and detection, though the baselines lack the same text information.

Pith tools