Pith. sign in

REVIEW 10 cited by

Hallucinations in Neural Automatic Speech Recognition: Identifying Errors and Hallucinatory Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.01572 v1 pith:XXPRU2SO submitted 2024-01-03 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords hallucinationshallucinatoryautomaticerrormodelmodelsnoiserecognition
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hallucinations are a type of output error produced by deep neural networks. While this has been studied in natural language processing, they have not been researched previously in automatic speech recognition. Here, we define hallucinations in ASR as transcriptions generated by a model that are semantically unrelated to the source utterance, yet still fluent and coherent. The similarity of hallucinations to probable natural language outputs of the model creates a danger of deception and impacts the credibility of the system. We show that commonly used metrics, such as word error rates, cannot differentiate between hallucinatory and non-hallucinatory models. To address this, we propose a perturbation-based method for assessing the susceptibility of an automatic speech recognition (ASR) model to hallucination at test time, which does not require access to the training dataset. We demonstrate that this method helps to distinguish between hallucinatory and non-hallucinatory models that have similar baseline word error rates. We further explore the relationship between the types of ASR errors and the types of dataset noise to determine what types of noise are most likely to create hallucinatory outputs. We devise a framework for identifying hallucinations by analysing their semantic connection with the ground truth and their fluency. Finally, we discover how to induce hallucinations with a random noise injection to the utterance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models

    cs.SD 2026-04 unverdicted novelty 8.0 of 10

    HalluAudio is the first large-scale benchmark spanning speech, environmental sound, and music that uses human-verified QA pairs, adversarial prompts, and mixed-audio tests to measure hallucinations in large audio-lang...

  2. HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems

    cs.SD 2026-06 unverdicted novelty 7.0 of 10

    HALAS is a human-annotated dataset of ASR hallucinations on unprocessed real audio that shows simple metrics outperform current detection methods at 81% ROC-AUC versus 53.1% F1.

  3. Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition

    cs.CL 2026-04 unverdicted novelty 7.0 of 10

    LLM decoders in speech recognition show no racial bias amplification and fewer repetition hallucinations under degradation than Whisper, with audio encoder design mattering more than model scale for fairness and robustness.

  4. LLMs and Speech: Integration vs. Combination

    eess.AS 2026-03 conditional novelty 6.0 of 10

    With matched data and sizes, CTC+LLM shallow fusion beats tight speech-LLM integration on in-domain ASR, while prefix LLMs win average WER on out-of-domain HuggingFace sets.

  5. TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling

    cs.SD 2026-03 unverdicted novelty 6.0 of 10

    TW-Sound580K dataset plus Tai-LALM model with dynamic Dual-ASR arbitration lifts localized Taiwanese audio-language accuracy to 49.1% on the TAU benchmark.

  6. From Text Metrics to Model Internals: A Study of Whisper ASR Hallucination Detection

    cs.SD 2026-06 unverdicted novelty 5.0 of 10

    Internal decoder probing of Whisper yields strongest hallucination detection without references, with late fusion of text and internal features performing best overall.

  7. Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    Four attention metrics enable logistic regression classifiers that detect hallucinations in SpeechLLMs with up to +0.23 PR-AUC gains over baselines on ASR and translation tasks.

  8. Group Relative Policy Optimization for Speech Recognition

    eess.AS 2025-09 conditional novelty 5.0 of 10

    Applying GRPO with rule-based rewards to LLM-based ASR improves WER by up to 18.4% relative and reduces hallucination errors on unseen acoustic conditions.

  9. LLMs and Speech: Integration vs. Combination

    eess.AS 2026-03 unverdicted novelty 4.0 of 10

    Tight integration of acoustic models with LLMs for ASR is ablated against shallow fusion across label units, fine-tuning strategies, LLM sizes, and joint CTC decoding to mitigate hallucinations.

  10. Too Good to Be True: A Study on Modern Automatic Speech Recognition for the Evaluation of Speech Enhancement

    eess.AS 2026-05 unverdicted novelty 3.0 of 10

    Modern ASR models with noisy training and language models correlate better with human WER for speech enhancement evaluation than simpler models, yet their robustness makes them less suitable for purely acoustic assessments.

Pith tools