Pith. sign in

REVIEW

Supervision-Guided Codebooks for Masked Prediction in Speech Pre-training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.10125 v1 pith:AKECC7S7 submitted 2022-06-21 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords speechpre-trainingclusteringcodebookhybridmaskedmodelsnamed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, masked prediction pre-training has seen remarkable progress in self-supervised learning (SSL) for speech recognition. It usually requires a codebook obtained in an unsupervised way, making it less accurate and difficult to interpret. We propose two supervision-guided codebook generation approaches to improve automatic speech recognition (ASR) performance and also the pre-training efficiency, either through decoding with a hybrid ASR system to generate phoneme-level alignments (named PBERT), or performing clustering on the supervised speech features extracted from an end-to-end CTC model (named CTC clustering). Both the hybrid and CTC models are trained on the same small amount of labeled speech as used in fine-tuning. Experiments demonstrate significant superiority of our methods to various SSL and self-training baselines, with up to 17.0% relative WER reduction. Our pre-trained models also show good transferability in a non-ASR speech task.

Discussion (0). Continue with ORCID to comment.

Pith tools