Pith. sign in

REVIEW 3 cited by

Effectiveness of self-supervised pre-training for speech recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.03912 v3 pith:5A6QK6ZC submitted 2019-11-10 cs.CL cs.LG

classification cs.CLcs.LG
keywords dataaudiobertlabeledlearningmodelrepresentationsspeech
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We compare self-supervised representation learning algorithms which either explicitly quantize the audio data or learn representations without quantization. We find the former to be more accurate since it builds a good vocabulary of the data through vq-wav2vec [1] to enable learning of effective representations in subsequent BERT training. Different to previous work, we directly fine-tune the pre-trained BERT models on transcribed speech using a Connectionist Temporal Classification (CTC) loss instead of feeding the representations into a task-specific model. We also propose a BERT-style model learning directly from the continuous audio data and compare pre-training on raw audio to spectral features. Fine-tuning a BERT model on 10 hour of labeled Librispeech data with a vq-wav2vec vocabulary is almost as good as the best known reported system trained on 100 hours of labeled data on testclean, while achieving a 25% WER reduction on test-other. When using only 10 minutes of labeled data, WER is 25.2 on test-other and 16.3 on test-clean. This demonstrates that self-supervision can enable speech recognition systems trained on a near-zero amount of transcribed data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language

    cs.CL 2025-02 conditional novelty 6.0 of 10

    The paper releases Sagalee, a 100-hour, 283-speaker Oromo ASR dataset, and reports baseline WERs of 15.32% (Conformer AED), 18.74% (Conformer CTC), and 10.82% (Whisper Large-v3 fine-tuned).

  2. Pitch Accent Detection improves Pretrained Automatic Speech Recognition

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Jointly training pitch accent detection with ASR on wav2vec2 reduces LibriSpeech WER from 6.0 to 4.3 in a one-hour fine-tuning setting.

  3. Is Self-Pretraining really useful to improve diagnosis in medical Time Series?

    cs.LG 2026-08 reject novelty 4.0 of 10

    Self-pretraining with masked reconstruction on the target medical time-series dataset improves transformer classification accuracy over training from scratch in most tested configurations, but gains vary by masking st...

Pith tools