REVIEW 2 cited by
Investigation of Data Augmentation Techniques for Disordered Speech Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Disordered speech recognition is a highly challenging task. The underlying neuro-motor conditions of people with speech disorders, often compounded with co-occurring physical disabilities, lead to the difficulty in collecting large quantities of speech required for system development. This paper investigates a set of data augmentation techniques for disordered speech recognition, including vocal tract length perturbation (VTLP), tempo perturbation and speed perturbation. Both normal and disordered speech were exploited in the augmentation process. Variability among impaired speakers in both the original and augmented data was modeled using learning hidden unit contributions (LHUC) based speaker adaptive training. The final speaker adapted system constructed using the UASpeech corpus and the best augmentation approach based on speed perturbation produced up to 2.92% absolute (9.3% relative) word error rate (WER) reduction over the baseline system without data augmentation, and gave an overall WER of 26.37% on the test set containing 16 dysarthric speakers.
Forward citations
Cited by 2 Pith papers
-
Adaptive Data Augmentation with NaturalSpeech3 for Far-field Speaker Verification
A voice-conversion augmentation that preserves far-field acoustics while transplanting near-field speaker identity improves FFSVC2020 verification in the training phase, but the headline test-time results use the test...
-
Not All Errors Are Equal: Investigation of Speech Recognition Errors in Alzheimer's Disease Detection
The analysis of ASR errors in BERT-based Alzheimer's detection shows that stopwords dominate the errors but barely affect classification, while task-related keywords are transcribed accurately and are the decisive factor.
Discussion (0). Continue with ORCID to comment.