Pith. sign in

REVIEW 2 cited by

DeCoAR 2.0: Deep Contextualized Acoustic Representations with Vector Quantization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.06659 v1 pith:5WAOIBOV submitted 2020-12-11 eess.AS cs.CLcs.LG

classification eess.AScs.CLcs.LG
keywords speechdecoarrepresentationdataquantizationrepresentationsvectorlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent success in speech representation learning enables a new way to leverage unlabeled data to train speech recognition model. In speech representation learning, a large amount of unlabeled data is used in a self-supervised manner to learn a feature representation. Then a smaller amount of labeled data is used to train a downstream ASR system using the new feature representations. Based on our previous work DeCoAR and inspirations from other speech representation learning, we propose DeCoAR 2.0, a Deep Contextualized Acoustic Representation with vector quantization. We introduce several modifications over the DeCoAR: first, we use Transformers in encoding module instead of LSTMs; second, we introduce a vector quantization layer between encoder and reconstruction modules; third, we propose an objective that combines the reconstructive loss with vector quantization diversity loss to train speech representations. Our experiments show consistent improvements over other speech representations in different data-sparse scenarios. Without fine-tuning, a light-weight ASR model trained on 10 hours of LibriSpeech labeled data with DeCoAR 2.0 features outperforms the model trained on the full 960-hour dataset with filterbank features.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pitch Accent Detection improves Pretrained Automatic Speech Recognition

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Jointly training pitch accent detection with ASR on wav2vec2 reduces LibriSpeech WER from 6.0 to 4.3 in a one-hour fine-tuning setting.

  2. Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Unsupervised k-means clusters of SSL features and i-vectors identify speaker-relevant feed-forward neurons; protecting them during pruning preserves speaker identification.

Pith tools