Pith. sign in

REVIEW 1 cited by

Text-To-Speech Data Augmentation for Low Resource Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.00291 v1 pith:DS4NTRON submitted 2022-04-01 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords syntheticdataaugmentationmodelspeechmethodquechuatext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Nowadays, the main problem of deep learning techniques used in the development of automatic speech recognition (ASR) models is the lack of transcribed data. The goal of this research is to propose a new data augmentation method to improve ASR models for agglutinative and low-resource languages. This novel data augmentation method generates both synthetic text and synthetic audio. Some experiments were conducted using the corpus of the Quechua language, which is an agglutinative and low-resource language. In this study, a sequence-to-sequence (seq2seq) model was applied to generate synthetic text, in addition to generating synthetic speech using a text-to-speech (TTS) model for Quechua. The results show that the new data augmentation method works well to improve the ASR model for Quechua. In this research, an 8.73% improvement in the word-error-rate (WER) of the ASR model is obtained using a combination of synthetic text and synthetic speech.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Diarization-guided full SFT, synthetic-speech LoRA, and GRPO RL adapt Qwen3-ASR-1.7B to 23.70 average tcpMER on the MLC-SLM 2026 Task 1 development set.

Pith tools