Pith. sign in

Opening the Black Box of wav2vec Feature Encoder

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Self-supervised models, namely, wav2vec and its variants, have shown promising results in various downstream tasks in the speech domain. However, their inner workings are poorly understood, calling for in-depth analyses on what the model learns. In this paper, we concentrate on the convolutional feature encoder where its latent space is often speculated to represent discrete acoustic units. To analyze the embedding space in a reductive manner, we feed the synthesized audio signals, which is the summation of simple sine waves. Through extensive experiments, we conclude that various information is embedded inside the feature encoder representations: (1) fundamental frequency, (2) formants, and (3) amplitude, packed with (4) sufficient temporal detail. Further, the information incorporated inside the latent representations is analogous to spectrograms but with a fundamental difference: latent representations construct a metric space so that closer representations imply acoustic similarity.

citation-role summary

background 1

citation-polarity summary

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

support 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • Automatic classification of stop realisation with wav2vec2.0 cs.CL · 2025-05-29 · conditional · none · ref 26 · internal anchor

    wav2vec2.0 models classify stop burst presence with 88 to 94 percent accuracy, and a Japanese-pretrained model reaches near-peak accuracy on English with only 500 annotated tokens.