Pith. sign in

REVIEW 2 cited by

Semi-supervised Learning for Code-Switching ASR with Large Language Model Filter

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.04219 v2 pith:N6OQWXT7 submitted 2024-07-05 eess.AS

classification eess.AS
keywords datalearningmonolingualsemi-supervisedwhenachievecode-switchingcs-asr
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Code-switching (CS) phenomenon occurs when words or phrases from different languages are alternated in a single sentence. Due to data scarcity, building an effective CS Automatic Speech Recognition (ASR) system remains challenging. In this paper, we propose to enhance CS-ASR systems by utilizing rich unsupervised monolingual speech data within a semi-supervised learning framework, particularly when access to CS data is limited. To achieve this, we establish a general paradigm for applying noisy student training (NST) to the CS-ASR task. Specifically, we introduce the LLM-Filter, which leverages well-designed prompt templates to activate the correction capability of large language models (LLMs) for monolingual data selection and pseudo-labels refinement during NST. Our experiments on the supervised ASRU-CS and unsupervised AISHELL-2 and LibriSpeech datasets show that our method not only achieves significant improvements over supervised and semi-supervised learning baselines for the CS task, but also attains better performance compared with the fully-supervised oracle upper-bound on the CS English part. Additionally, we further investigate the influence of accent on AESRC dataset and demonstrate that our method can get achieve additional benefits when the monolingual data contains relevant linguistic characteristic.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fotheidil: an Automatic Transcription System for the Irish Language

    cs.CL 2024-12 conditional novelty 5.0 of 10

    An Irish-language transcription web service whose ASR is improved by semi-supervised learning on 3,230 hours of radio speech, and whose punctuation/capitalisation restoration uses a sequence-to-sequence transformer.

  2. Optimizing Estonian TV Subtitles with Semi-supervised Learning and LLMs

    cs.CL 2025-01 conditional novelty 4.0 of 10

    Fine-tuning Whisper on Estonian subtitles with iterative pseudo-labeling and test-time LLM editing improves subtitle quality, while LLM editing during training yields no gain.

Pith tools