Pith. sign in

REVIEW 3 cited by

Opening the Black Box of wav2vec Feature Encoder

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.15386 v1 pith:5LIPVQLZ submitted 2022-10-27 cs.SD cs.CLcs.LGeess.AS

classification cs.SDcs.CLcs.LGeess.AS
keywords representationsencoderfeaturelatentspaceacousticfundamentalinformation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised models, namely, wav2vec and its variants, have shown promising results in various downstream tasks in the speech domain. However, their inner workings are poorly understood, calling for in-depth analyses on what the model learns. In this paper, we concentrate on the convolutional feature encoder where its latent space is often speculated to represent discrete acoustic units. To analyze the embedding space in a reductive manner, we feed the synthesized audio signals, which is the summation of simple sine waves. Through extensive experiments, we conclude that various information is embedded inside the feature encoder representations: (1) fundamental frequency, (2) formants, and (3) amplitude, packed with (4) sufficient temporal detail. Further, the information incorporated inside the latent representations is analogous to spectrograms but with a fundamental difference: latent representations construct a metric space so that closer representations imply acoustic similarity.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Phone Segmentation and Recognition through Phonological Activation Mapping

    eess.AS 2026-07 accept novelty 6.0 of 10

    SPAM projects S3M frames onto phonological vectors and uses gradient-free heads to jointly segment and recognize phones from under a minute of labels, generalizing to unseen phones and languages.

  2. Automatic classification of stop realisation with wav2vec2.0

    cs.CL 2025-05 conditional novelty 5.0 of 10

    wav2vec2.0 models classify stop burst presence with 88 to 94 percent accuracy, and a Japanese-pretrained model reaches near-peak accuracy on English with only 500 annotated tokens.

  3. Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Modeling each phoneme as a 32-component Gaussian mixture of self-supervised speech features improves atypical pronunciation scoring on four of five datasets, with S3Ms showing stronger allophonic structure than MFCCs ...

Pith tools