REVIEW 7 cited by
Opening the Black Box of wav2vec Feature Encoder
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Self-supervised models, namely, wav2vec and its variants, have shown promising results in various downstream tasks in the speech domain. However, their inner workings are poorly understood, calling for in-depth analyses on what the model learns. In this paper, we concentrate on the convolutional feature encoder where its latent space is often speculated to represent discrete acoustic units. To analyze the embedding space in a reductive manner, we feed the synthesized audio signals, which is the summation of simple sine waves. Through extensive experiments, we conclude that various information is embedded inside the feature encoder representations: (1) fundamental frequency, (2) formants, and (3) amplitude, packed with (4) sufficient temporal detail. Further, the information incorporated inside the latent representations is analogous to spectrograms but with a fundamental difference: latent representations construct a metric space so that closer representations imply acoustic similarity.
Forward citations
Cited by 7 Pith papers
-
Phone Segmentation and Recognition through Phonological Activation Mapping
SPAM projects S3M frames onto phonological vectors and uses gradient-free heads to jointly segment and recognize phones from under a minute of labels, generalizing to unseen phones and languages.
-
Discrete Speech Unit Extraction via Independent Component Analysis
ICA and whitening preprocessing before k-means improves DSU-based ASR for XLS-R-300M, and ICA axes show interpretable phonetic contrasts.
-
Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition
Transliterated zero-shot domain adaptation reduces ASR word error rate by 9.2% relative to wav2vec 2.0 by pre-training on transliterated pseudo-labels from a related source language.
-
Automatic classification of stop realisation with wav2vec2.0
wav2vec2.0 models classify stop burst presence with 88 to 94 percent accuracy, and a Japanese-pretrained model reaches near-peak accuracy on English with only 500 annotated tokens.
-
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
Modeling each phoneme as a 32-component Gaussian mixture of self-supervised speech features improves atypical pronunciation scoring on four of five datasets, with S3Ms showing stronger allophonic structure than MFCCs ...
-
Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation
A transformer ASR model with one temporally smoothed attention head per layer yields explicit speaker embeddings that improve diarization without hurting recognition.
-
Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context
Combining language-universal and language-specific acoustic features in XGBoost improves multilingual dysarthria severity classification by 7.33% relative over a language-universal-only baseline.
Discussion (0). Continue with ORCID to comment.