Pith. sign in

REVIEW 1 cited by

Automatic Speech Recognition System-Independent Word Error Rate Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.16743 v2 pith:ZHOYCKY4 submitted 2024-04-25 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords dataestimationerrorestimatorsperformancespeechapplicationsautomatic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Word error rate (WER) is a metric used to evaluate the quality of transcriptions produced by Automatic Speech Recognition (ASR) systems. In many applications, it is of interest to estimate WER given a pair of a speech utterance and a transcript. Previous work on WER estimation focused on building models that are trained with a specific ASR system in mind (referred to as ASR system-dependent). These are also domain-dependent and inflexible in real-world applications. In this paper, a hypothesis generation method for ASR System-Independent WER estimation (SIWE) is proposed. In contrast to prior work, the WER estimators are trained using data that simulates ASR system output. Hypotheses are generated using phonetically similar or linguistically more likely alternative words. In WER estimation experiments, the proposed method reaches a similar performance to ASR system-dependent WER estimators on in-domain data and achieves state-of-the-art performance on out-of-domain data. On the out-of-domain data, the SIWE model outperformed the baseline estimators in root mean square error and Pearson correlation coefficient by relative 17.58% and 18.21%, respectively, on Switchboard and CALLHOME. The performance was further improved when the WER of the training set was close to the WER of the evaluation dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Weak Supervision Techniques towards Enhanced ASR Models in Industry-level CRM Systems

    cs.SD 2025-07 conditional novelty 4.0 of 10

    Using LLM- and TTS-generated synthetic speech to fine-tune Whisper models cuts character error rates by up to 63% in a luxury retail CRM transcription task, though the evaluation has several methodological weaknesses.

Pith tools