Pith. sign in

REVIEW 3 cited by

An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.16456 v1 pith:KJ3PBVED submitted 2025-07-22 eess.AS cs.SD

classification eess.AScs.SD
keywords applicationslargeperformanceautomaticerrorslanguagellmsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatic Speech Recognition (ASR) plays a crucial role in human-machine interaction and serves as an interface for a wide range of applications. Traditionally, ASR performance has been evaluated using Word Error Rate (WER), a metric that quantifies the number of insertions, deletions, and substitutions in the generated transcriptions. However, with the increasing adoption of large and powerful Large Language Models (LLMs) as the core processing component in various applications, the significance of different types of ASR errors in downstream tasks warrants further exploration. In this work, we analyze the capabilities of LLMs to correct errors introduced by ASRs and propose a new measure to evaluate ASR performance for LLM-powered applications.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach

    cs.AI 2026-07 conditional novelty 6.0 of 10

    An ASR–Refiner system emits and revises clean transcripts online over a bounded sliding context, outperforming offline spoken-to-written baselines on a new bilingual rubric benchmark.

  2. Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    Agentic ASR adds closed-loop semantic correction to ASR and introduces S²ER, an LLM judge for meaning-level errors, showing larger gains on semantic than token metrics across multilingual benchmarks.

  3. From Text Metrics to Model Internals: A Study of Whisper ASR Hallucination Detection

    cs.SD 2026-06 unverdicted novelty 5.0 of 10

    Internal decoder probing of Whisper yields strongest hallucination detection without references, with late fusion of text and internal features performing best overall.

Pith tools