Pith. sign in

REVIEW

DistilXLSR: A Light Weight Cross-Lingual Speech Representation Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.01303 v1 pith:AE6AWZGE submitted 2023-06-02 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords modelscross-lingualspeechrepresentationlanguagesmethodteacherdistilxlsr
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multilingual self-supervised speech representation models have greatly enhanced the speech recognition performance for low-resource languages, and the compression of these huge models has also become a crucial prerequisite for their industrial application. In this paper, we propose DistilXLSR, a distilled cross-lingual speech representation model. By randomly shuffling the phonemes of existing speech, we reduce the linguistic information and distill cross-lingual models using only English data. We also design a layer-jumping initialization method to fully leverage the teacher's pre-trained weights. Experiments on 2 kinds of teacher models and 15 low-resource languages show that our method can reduce the parameters by 50% while maintaining cross-lingual representation ability. Our method is proven to be generalizable to various languages/teacher models and has the potential to improve the cross-lingual performance of the English pre-trained models.

Discussion (0). Sign in to comment.

Pith tools