Pith. sign in

REVIEW 2 cited by

Knowledge Distillation from Multiple Foundation Models for End-to-End Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.10917 v1 pith:TG6FBKSH submitted 2023-03-20 eess.AS cs.SD

classification eess.AScs.SD
keywords modelsstudentknowledgemodelmultipleencoderfoundationframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although large foundation models pre-trained by self-supervised learning have achieved state-of-the-art performance in many tasks including automatic speech recognition (ASR), knowledge distillation (KD) is often required in practice to transfer the knowledge learned by large teacher models into much smaller student models with affordable computation and memory costs. This paper proposes a novel two-stage KD framework to distil the knowledge from multiple speech foundation models as teachers into a single student neural transducer model for ASR. In the first stage, the student model encoder is pre-trained using the embeddings extracted from multiple teacher models. In the second stage, the student encoder is fine-tuned with the audio-text pairs based on the ASR task. Experiments on the LibriSpeech 100-hour subset show that the proposed KD framework improves the performance of both streaming and non-streaming student models when using only one teacher. The performance of the student model can be further enhanced when multiple teachers are used jointly, achieving word error rate reductions (WERRs) of 17.5% and 10.6%. Our proposed framework can be combined with other existing KD methods to achieve further improvements. Further WERRs were obtained by incorporating extra unlabelled data during encoder pre-training, leading to a total relative WERR of 55.0% on the non-streaming student model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language

    cs.CL 2025-02 conditional novelty 6.0 of 10

    The paper releases Sagalee, a 100-hour, 283-speaker Oromo ASR dataset, and reports baseline WERs of 15.32% (Conformer AED), 18.74% (Conformer CTC), and 10.82% (Whisper Large-v3 fine-tuned).

  2. Multi-Distillation from Speech and Music Representation Models

    eess.AS 2025-06 conditional novelty 4.0 of 10

    A 23M-parameter student distilled from HuBERT/WavLM and MERT gets close to teacher-level average accuracy on speech and music benchmarks and outperforms its teachers in few-shot classification.

Pith tools