Pith. sign in

REVIEW 6 cited by

Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12414 v2 pith:L5OAMSL2 submitted 2025-02-18 cs.CL

Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models

classification cs.CL
keywords hallucinationmodelsmodelspeechdistributionerrorlikeperformance
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Speech foundation models trained at a massive scale, both in terms of model and data size, result in robust systems capable of performing multiple speech tasks, including automatic speech recognition (ASR). These models transcend language and domain barriers, yet effectively measuring their performance remains a challenge. Traditional metrics like word error rate (WER) and character error rate (CER) are commonly used to evaluate ASR performance but often fail to reflect transcription quality in critical contexts, particularly when detecting fabricated outputs. This phenomenon, known as hallucination, is especially concerning in high-stakes domains such as healthcare, legal, and aviation, where errors can have severe consequences. In our work, we address this gap by investigating hallucination in ASR models. We examine how factors such as distribution shifts, model size, and model architecture influence the hallucination error rate (HER), a metric we introduce to quantify hallucinations. Our analysis of over 20 ASR models reveals \numinsights~key insights: (1) High WERs can mask low hallucination rates, while low WERs may conceal dangerous hallucinations. (2) Synthetic noise, both adversarial and common perturbations like white noise, pitch shift, and time stretching, increase HER. (3) Distribution shift correlates strongly with HER ($\alpha = 0.91$). Our findings highlight the importance of incorporating HER alongside traditional metrics like WER to better assess ASR model performance, particularly in high-stakes domains.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales

    cs.LG 2026-03 conditional novelty 7.5

    The Spectral Sensitivity Theorem identifies a phase transition in Whisper models where scaling causes self-attention to collapse into rank-1 attractors, decoupling output from acoustic evidence.

  2. REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing

    cs.CL 2026-07 conditional novelty 7.0

    A two-stage replay-based post-training method corrects ASR timestamp drift across non-speech gaps while preserving recognition far better than ordinary timestamp fine-tuning.

  3. HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems

    cs.SD 2026-06 unverdicted novelty 7.0

    HALAS is a human-annotated dataset of ASR hallucinations on unprocessed real audio that shows simple metrics outperform current detection methods at 81% ROC-AUC versus 53.1% F1.

  4. REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing

    cs.CL 2026-07 conditional novelty 6.0

    REDDIT corrects non-speech-induced timestamp drift in autoregressive ASR by editing timestamp targets under cached replay context while anchoring non-timestamp behavior to the frozen base distribution.

  5. From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales

    cs.LG 2026-03 unverdicted novelty 5.5

    Whisper models show a scale-dependent spectral phase transition from mid-size cross-attention rank collapse under stress to large-model self-attention compression that decouples from acoustic evidence.

  6. From Text Metrics to Model Internals: A Study of Whisper ASR Hallucination Detection

    cs.SD 2026-06 unverdicted novelty 5.0

    Internal decoder probing of Whisper yields strongest hallucination detection without references, with late fusion of text and internal features performing best overall.