Pith. sign in

REVIEW 2 cited by

Layer-wise Analysis of a Self-supervised Speech Representation Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.04734 v3 pith:LCQHOB22 submitted 2021-07-10 cs.CL cs.LGeess.AS

classification cs.CLcs.LGeess.AS
keywords informationmodelbeenrepresentationspeechanalysisdownstreamfine-tuning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently proposed self-supervised learning approaches have been successful for pre-training speech representation models. The utility of these learned representations has been observed empirically, but not much has been studied about the type or extent of information encoded in the pre-trained representations themselves. Developing such insights can help understand the capabilities and limits of these models and enable the research community to more efficiently develop their usage for downstream applications. In this work, we begin to fill this gap by examining one recent and successful pre-trained model (wav2vec 2.0), via its intermediate representation vectors, using a suite of analysis tools. We use the metrics of canonical correlation, mutual information, and performance on simple downstream tasks with non-parametric probes, in order to (i) query for acoustic and linguistic information content, (ii) characterize the evolution of information across model layers, and (iii) understand how fine-tuning the model for automatic speech recognition (ASR) affects these observations. Our findings motivate modifying the fine-tuning protocol for ASR, which produces improved word error rates in a low-resource setting.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BrainWavLM: Fine-tuning Speech Representations with Brain Responses to Language

    cs.CL 2025-02 conditional novelty 5.0 of 10

    LoRA fine-tuning of WavLM on fMRI responses to stories improves held-out brain encoding performance by about 12.5 percent over the linearized pretrained model, and improves semantic probe accuracy without labeled annotations.

  2. Lillama: Large Language Models Compression via Low-Rank Feature Distillation

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Lillama compresses LLMs by SVD-initialized low-rank layers trained with a local Teacher plus Student activation distillation loss, achieving 20-40% parameter reduction with only 13 million calibration tokens.

Pith tools