Pith. sign in

REVIEW 2 cited by

Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.05554 v1 pith:XZO3VWJA submitted 2024-08-10 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords performancespeechdatawhisperimprovekazakhlanguagesrecognition
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Whisper and other large-scale automatic speech recognition models have made significant progress in performance. However, their performance on many low-resource languages, such as Kazakh, is not satisfactory. It is worth researching how to utilize low-cost data to improve the performance of Whisper on under-represented languages. In this study, we utilized easily accessible unpaired speech and text data and combined the language model GPT with Whisper on Kazakh. We implemented end of transcript (EOT) judgment modification and hallucination penalty to improve the performance of speech recognition. Further, we employed the decoding average token log probability as a criterion to select samples from unlabeled speech data and used pseudo-labeled data to fine-tune the model to further improve its performance. Ultimately, we achieved more than 10\% absolute WER reduction in multiple experiments, and the whole process has the potential to be generalized to other under-represented languages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling

    eess.AS 2024-12 conditional novelty 6.0 of 10

    A weighted sum of Whisper's language embeddings, optionally refined by a small MLP, improves ASR on unseen languages in zero-shot and fine-tuning settings.

  2. Enhancing Whisper's Accuracy and Speed for Indian Languages through Prompt-Tuning and Tokenization

    cs.CL 2024-12 conditional novelty 4.0 of 10

    Language-family prompt tuning and a BPE-token-extended tokenizer improve Whisper's WER and inference speed on eight Indian languages, with 250 added tokens per language as the best configuration.

Pith tools