Pith. sign in

REVIEW 1 cited by

Investigating Pre-trained Audio Encoders in the Low-Resource Condition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.17733 v1 pith:FDWEHI42 submitted 2023-05-28 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords encoderslow-resourcespeechtasksacrosscapabilitiesconvergencegeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-trained speech encoders have been central to pushing state-of-the-art results across various speech understanding and generation tasks. Nonetheless, the capabilities of these encoders in low-resource settings are yet to be thoroughly explored. To address this, we conduct a comprehensive set of experiments using a representative set of 3 state-of-the-art encoders (Wav2vec2, WavLM, Whisper) in the low-resource setting across 7 speech understanding and generation tasks. We provide various quantitative and qualitative analyses on task performance, convergence speed, and representational properties of the encoders. We observe a connection between the pre-training protocols of these encoders and the way in which they capture information in their internal layers. In particular, we observe the Whisper encoder exhibits the greatest low-resource capabilities on content-driven tasks in terms of performance and convergence speed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Continual Speech Learning with Fused Speech Features

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Gated fusion of frozen Whisper layers improves continual learning on six speech tasks, with the double-stage variant best overall.

Pith tools