Pith. sign in

REVIEW 7 cited by

Can Generative Large Language Models Perform ASR Error Correction?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.04172 v2 pith:6KSY5GLI submitted 2023-07-09 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords correctionerrorgenerativelanguagemodelssystemapproachfashion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

ASR error correction is an interesting option for post processing speech recognition system outputs. These error correction models are usually trained in a supervised fashion using the decoding results of a target ASR system. This approach can be computationally intensive and the model is tuned to a specific ASR system. Recently generative large language models (LLMs) have been applied to a wide range of natural language processing tasks, as they can operate in a zero-shot or few shot fashion. In this paper we investigate using ChatGPT, a generative LLM, for ASR error correction. Based on the ASR N-best output, we propose both unconstrained and constrained, where a member of the N-best list is selected, approaches. Additionally, zero and 1-shot settings are evaluated. Experiments show that this generative LLM approach can yield performance gains for two different state-of-the-art ASR architectures, transducer and attention-encoder-decoder based, and multiple test sets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Can Large Language Models Reliably Correct Errors in Low-Resource ASR? A Contamination-Aware Case Study on West Frisian

    cs.CL 2026-05 conditional novelty 7.0 of 10

    LLM generative error correction improves low-resource Frisian ASR performance, with comparable gains on a contamination-controlled offline dataset confirming true correction ability.

  2. Phonemes vs. Projectors: An Investigation of Speech-Language Interfaces for LLM-based ASR

    eess.AS 2026-04 unverdicted novelty 7.0 of 10

    Phoneme-based interfaces match or surpass projector-based ones for LLM ASR, especially in low-resource languages, and a BPE-phoneme hybrid offers additional improvements.

  3. Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection

    eess.AS 2026-01 conditional novelty 7.0 of 10

    A learnable prompt projector added to LLM-based ASR reduces prompt sensitivity, lowers performance variability, and beats the best fixed prompts on four datasets.

  4. Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy

    eess.AS 2026-06 unverdicted novelty 6.0 of 10

    READ is a reference-free ASR hypothesis scorer that measures acoustic discrepancy via conditional likelihood from a pretrained auto-regressive TTS model and yields up to 20% relative error rate reduction when used for...

  5. Error-Aware TF-IDF Retrieval-Augmented Generation for ASR Error Correction

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    Introduces error-aware TF-IDF RAG that raises error-aware hit rate from 53.7% to 90.9% and lowers WER from 23.06% to 18.83% on Persian FLEURS data.

  6. An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications

    eess.AS 2025-07 reject novelty 5.0 of 10

    The paper introduces AER, an LLM-judged question-answering metric for evaluating ASR output in LLM applications, and shows it does not correlate strongly with WER.

  7. Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition

    cs.SD 2025-09 reject novelty 4.0 of 10

    An LLM-based ASR error correction framework with noise-adaptive encoding and dynamic multi-modal fusion reports WER gains, but its fusion weights require ground-truth text at inference.

Pith tools