Pith. sign in

REVIEW

Empowering Low-Resource Language ASR via Large-Scale Pseudo Labeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.14026 v1 pith:7F2JIYC7 submitted 2024-08-26 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords datalanguageslow-resourcemultiplebenchmarkexistingframeworkindicyt
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this study, we tackle the challenge of limited labeled data for low-resource languages in ASR, focusing on Hindi. Specifically, we explore pseudo-labeling, by proposing a generic framework combining multiple ideas from existing works. Our framework integrates multiple base models for transcription and evaluators for assessing audio-transcript pairs, resulting in robust pseudo-labeling for low resource languages. We validate our approach with a new benchmark, IndicYT, comprising diverse YouTube audio files from multiple content categories. Our findings show that augmenting pseudo labeled data from YouTube with existing training data leads to significant performance improvements on IndicYT, without affecting performance on out-of-domain benchmarks, demonstrating the efficacy of pseudo-labeled data in enhancing ASR capabilities for low-resource languages. The benchmark, code and models developed as a part of this work will be made publicly available.

Discussion (0). Sign in to comment.

Pith tools