Pith. sign in

REVIEW 3 cited by

Can Contextual Biasing Remain Effective with Whisper and GPT-2?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.01942 v1 pith:IDGGIM66 submitted 2023-06-02 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords biasingwhispercontextualgpt-2datatrainingwordseffective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

End-to-end automatic speech recognition (ASR) and large language models, such as Whisper and GPT-2, have recently been scaled to use vast amounts of training data. Despite the large amount of training data, infrequent content words that occur in a particular task may still exhibit poor ASR performance, with contextual biasing a possible remedy. This paper investigates the effectiveness of neural contextual biasing for Whisper combined with GPT-2. Specifically, this paper proposes integrating an adapted tree-constrained pointer generator (TCPGen) component for Whisper and a dedicated training scheme to dynamically adjust the final output without modifying any Whisper model parameters. Experiments across three datasets show a considerable reduction in errors on biasing words with a biasing list of 1000 words. Contextual biasing was more effective when applied to domain-specific data and can boost the performance of Whisper and GPT-2 without losing their generality.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Synthetic Data Training for Contextual Biasing Models with a Keyword-Aware Cost Function

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A keyword-aware loss with masked cross-entropy and binary gating terms reduces overfitting in synthetic-data training of TCPGen, improving Whisper WER on NSC Part 2 from 14.16% (AGEM baseline) to 11.81%.

  2. Improving Contextual ASR via Multi-grained Fusion with Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A multi-grained fusion method that jointly uses token-level and phrase-level scores from ASR and LLM improves keyword recognition in contextual ASR.

  3. Efficient Trie-based Biasing using K-step Prediction for Rare Word Recognition

    cs.CL 2025-09 conditional novelty 5.0 of 10

    A future-token prediction branch in Whisper gates trie-based biasing rewards, letting greedy decoding recognize rare words without a beam-search reward revocation step.

Pith tools