Pith. sign in

REVIEW 5 cited by

Iterative Pseudo-Labeling for Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.09267 v2 pith:6O2UVR42 submitted 2020-05-19 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords datapseudo-labelingmodeliterativelanguagelibrispeechlow-resourcerecognition
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pseudo-labeling has recently shown promise in end-to-end automatic speech recognition (ASR). We study Iterative Pseudo-Labeling (IPL), a semi-supervised algorithm which efficiently performs multiple iterations of pseudo-labeling on unlabeled data as the acoustic model evolves. In particular, IPL fine-tunes an existing model at each iteration using both labeled data and a subset of unlabeled data. We study the main components of IPL: decoding with a language model and data augmentation. We then demonstrate the effectiveness of IPL by achieving state-of-the-art word-error rate on the Librispeech test sets in both standard and low-resource setting. We also study the effect of language models trained on different corpora to show IPL can effectively utilize additional text. Finally, we release a new large in-domain text corpus which does not overlap with the Librispeech training transcriptions to foster research in low-resource, semi-supervised ASR

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Emerging Properties in Self-Supervised Vision Transformers

    cs.CV 2021-04 conditional novelty 8.0 of 10

    Self-supervised ViTs show emergent semantic segmentation and 78.3% k-NN accuracy on ImageNet; DINO reaches 80.1% linear evaluation with ViT-Base.

  2. Vision Transformers Need Registers

    cs.CV 2023-09 unverdicted novelty 6.0 of 10

    Adding register tokens to Vision Transformers eliminates high-norm background artifacts and raises state-of-the-art performance on dense visual prediction tasks.

  3. Align-Consistency: Improving Non-autoregressive and Semi-supervised ASR with Consistency Regularization

    eess.AS 2026-02 conditional novelty 5.0 of 10

    Adding consistency regularization to every refinement step of Align-Refine reduces LibriSpeech word error rate and improves semi-supervised self-training with pseudo-labels.

  4. Pitch Accent Detection improves Pretrained Automatic Speech Recognition

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Jointly training pitch accent detection with ASR on wav2vec2 reduces LibriSpeech WER from 6.0 to 4.3 in a one-hour fine-tuning setting.

  5. Progressive Refinement: An Iterative Pseudo-Labeling Approach for Mandarin-English Code-Switching ASR

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Iterative pseudo-labeling on 22.4k hours of unlabeled code-switching audio reduces Mix Error Rate on SEAME devman to 12.88% and devsge to 18.89%.

Pith tools