Pith. sign in

REVIEW 3 cited by

Towards Improved Speech Recognition through Optimized Synthetic Data Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.21631 v1 pith:JUTVWQOG submitted 2025-08-29 eess.AS

Towards Improved Speech Recognition through Optimized Synthetic Data Generation

classification eess.AS
keywords dataspeechsyntheticgenerationrecognitionaudiomodelmodels
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Supervised training of speech recognition models requires access to transcribed audio data, which often is not possible due to confidentiality issues. Our approach to this problem is to generate synthetic audio from a text-only corpus using a state-of-the-art text-to-speech model with voice cloning capabilities. Our goal is to achieve automatic speech recognition (ASR) performance comparable to models trained on real data. We explore ways to optimize synthetic data generation through finetuning, filtering and evaluation, and its use for training an end-to-end encoder-decoder ASR model. Experiments were conducted using two datasets of spontaneous, conversational speech in Qu\'ebec French. We show that improving data generation leads to large improvements in the final ASR system trained on synthetic data.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. When Synthetic Speech Is All You Have: Better Call GRPO

    cs.CL 2026-07 conditional novelty 6.0

    On synthetic banking speech alone, GRPO cuts ASR WER 40% relative to SFT (36.71%→22.09%) by improving stopping calibration and attention anchoring to audio.

  2. How to Leverage Synthetic Speech for LLM-Based ASR Systems?

    cs.CL 2026-06 conditional novelty 6.0

    Layer-wise pooling plus RIR-augmented synthetic speech matches a 100%-real ASR baseline with only 25% real data (13.6 h) and beats it at higher real fractions.

  3. How to Leverage Synthetic Speech for LLM-Based ASR Systems?

    cs.CL 2026-06 unverdicted novelty 5.0

    Layer selection plus RIR augmentation on synthetic speech matches full real-data ASR performance using 25% real speech in SLAM-ASR.