Pith. sign in

REVIEW 3 cited by

Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.13198 v1 pith:PCGVU6TK submitted 2024-10-17 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords correctionerrorsentitieserrorgenerativescenariosapproachdarag
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative Error Correction (GEC) has emerged as a powerful post-processing method to enhance the performance of Automatic Speech Recognition (ASR) systems. However, we show that GEC models struggle to generalize beyond the specific types of errors encountered during training, limiting their ability to correct new, unseen errors at test time, particularly in out-of-domain (OOD) scenarios. This phenomenon amplifies with named entities (NEs), where, in addition to insufficient contextual information or knowledge about the NEs, novel NEs keep emerging. To address these issues, we propose DARAG (Data- and Retrieval-Augmented Generative Error Correction), a novel approach designed to improve GEC for ASR in in-domain (ID) and OOD scenarios. We augment the GEC training dataset with synthetic data generated by prompting LLMs and text-to-speech models, thereby simulating additional errors from which the model can learn. For OOD scenarios, we simulate test-time errors from new domains similarly and in an unsupervised fashion. Additionally, to better handle named entities, we introduce retrieval-augmented correction by augmenting the input with entities retrieved from a database. Our approach is simple, scalable, and both domain- and language-agnostic. We experiment on multiple datasets and settings, showing that DARAG outperforms all our baselines, achieving 8\% -- 30\% relative WER improvements in ID and 10\% -- 33\% improvements in OOD settings.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction

    cs.CL 2025-06 conditional novelty 6.0 of 10

    DeRAGEC explicitly denoises retrieved named-entity candidates with phonetic scores, definitions, and synthetic rationales, improving ASR error-correction WER and NE hit ratio without additional training.

  2. CHSER: A Dataset and Case Study on Generative Speech Error Correction for Child ASR

    eess.AS 2025-05 conditional novelty 6.0 of 10

    CHSER is a new 200K-pair dataset and case study showing that fine-tuned language models can correct child ASR errors, reducing WER by up to 28.5% relative.

  3. PHRASED: Phrase Dictionary Biasing for Speech Translation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Phrase dictionary biasing, which matches source phrases in intermediate ASR text and then boosts or prompts the matching target phrases, improves phrase recall in streaming and LLM-based speech translation.

Pith tools