Pith. sign in

REVIEW 1 cited by

Distilling a speech and music encoder with task arithmetic

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.13270 v1 pith:ETSPL4NI submitted 2025-05-19 cs.SD eess.AS

classification cs.SDeess.AS
keywords musicspeechdistillationmodelmodelsunifiedaudiogeneral
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the progress in self-supervised learning (SSL) for speech and music, existing models treat these domains separately, limiting their capacity for unified audio understanding. A unified model is desirable for applications that require general representations, e.g. audio large language models. Nonetheless, directly training a general model for speech and music is computationally expensive. Knowledge Distillation of teacher ensembles may be a natural solution, but we posit that decoupling the distillation of the speech and music SSL models allows for more flexibility. Thus, we propose to learn distilled task vectors and then linearly interpolate them to form a unified speech+music model. This strategy enables flexible domain emphasis through adjustable weights and is also simpler to train. Experiments on speech and music benchmarks demonstrate that our method yields superior overall performance compared to ensemble distillation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Distillation from Speech and Music Representation Models

    eess.AS 2025-06 conditional novelty 4.0 of 10

    A 23M-parameter student distilled from HuBERT/WavLM and MERT gets close to teacher-level average accuracy on speech and music benchmarks and outperforms its teachers in few-shot classification.

Pith tools