Pith. sign in

REVIEW 1 cited by

Analysis and Tuning of a Voice Assistant System for Dysfluent Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.11759 v1 pith:WF5FPIJF submitted 2021-06-18 eess.AS cs.AIcs.CLcs.CVcs.LGcs.SD

classification eess.AScs.AIcs.CLcs.CVcs.LGcs.SD
keywords speechrecognitionindividualssystemdisorderstuningvoiceanalysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Dysfluencies and variations in speech pronunciation can severely degrade speech recognition performance, and for many individuals with moderate-to-severe speech disorders, voice operated systems do not work. Current speech recognition systems are trained primarily with data from fluent speakers and as a consequence do not generalize well to speech with dysfluencies such as sound or word repetitions, sound prolongations, or audible blocks. The focus of this work is on quantitative analysis of a consumer speech recognition system on individuals who stutter and production-oriented approaches for improving performance for common voice assistant tasks (i.e., "what is the weather?"). At baseline, this system introduces a significant number of insertion and substitution errors resulting in intended speech Word Error Rates (isWER) that are 13.64\% worse (absolute) for individuals with fluency disorders. We show that by simply tuning the decoding parameters in an existing hybrid speech recognition system one can improve isWER by 24\% (relative) for individuals with fluency disorders. Tuning these parameters translates to 3.6\% better domain recognition and 1.7\% better intent recognition relative to the default setup for the 18 study participants across all stuttering severities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection

    cs.SD 2025-05 conditional novelty 5.0 of 10

    An LLM-driven multi-task system reports a 5.45% CER and 73.63% average SED F1 on the AS-70 Mandarin stuttering benchmark, though key baselines and uncertainty are missing.

Pith tools