Pith. sign in

REVIEW 7 cited by

Opening the Black Box of wav2vec Feature Encoder

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.15386 v1 pith:5LIPVQLZ submitted 2022-10-27 cs.SD cs.CLcs.LGeess.AS

classification cs.SDcs.CLcs.LGeess.AS
keywords representationsencoderfeaturelatentspaceacousticfundamentalinformation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Self-supervised models, namely, wav2vec and its variants, have shown promising results in various downstream tasks in the speech domain. However, their inner workings are poorly understood, calling for in-depth analyses on what the model learns. In this paper, we concentrate on the convolutional feature encoder where its latent space is often speculated to represent discrete acoustic units. To analyze the embedding space in a reductive manner, we feed the synthesized audio signals, which is the summation of simple sine waves. Through extensive experiments, we conclude that various information is embedded inside the feature encoder representations: (1) fundamental frequency, (2) formants, and (3) amplitude, packed with (4) sufficient temporal detail. Further, the information incorporated inside the latent representations is analogous to spectrograms but with a fundamental difference: latent representations construct a metric space so that closer representations imply acoustic similarity.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Phone Segmentation and Recognition through Phonological Activation Mapping

    eess.AS 2026-07 accept novelty 6.0 of 10

    SPAM projects S3M frames onto phonological vectors and uses gradient-free heads to jointly segment and recognize phones from under a minute of labels, generalizing to unseen phones and languages.

  2. Discrete Speech Unit Extraction via Independent Component Analysis

    eess.AS 2025-01 conditional novelty 6.0 of 10

    ICA and whitening preprocessing before k-means improves DSU-based ASR for XLS-R-300M, and ICA axes show interpretable phonetic contrasts.

  3. Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition

    eess.AS 2024-12 conditional novelty 6.0 of 10

    Transliterated zero-shot domain adaptation reduces ASR word error rate by 9.2% relative to wav2vec 2.0 by pre-training on transliterated pseudo-labels from a related source language.

  4. Automatic classification of stop realisation with wav2vec2.0

    cs.CL 2025-05 conditional novelty 5.0 of 10

    wav2vec2.0 models classify stop burst presence with 88 to 94 percent accuracy, and a Japanese-pretrained model reaches near-peak accuracy on English with only 500 annotated tokens.

  5. Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Modeling each phoneme as a 32-component Gaussian mixture of self-supervised speech features improves atypical pronunciation scoring on four of five datasets, with S3Ms showing stronger allophonic structure than MFCCs ...

  6. Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation

    eess.AS 2024-11 conditional novelty 5.0 of 10

    A transformer ASR model with one temporally smoothed attention head per layer yields explicit speaker embeddings that improve diarization without hurting recognition.

  7. Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context

    cs.SD 2024-12 conditional novelty 4.0 of 10

    Combining language-universal and language-specific acoustic features in XGBoost improves multilingual dysarthria severity classification by 7.33% relative over a language-universal-only baseline.

Pith tools