UME upcycles pretrained dense ASR checkpoints into larger MoE models via weight copying, layer freezing, and expert balancing, yielding up to 11.9% relative CER reduction and 86.7% training time savings versus training from scratch.
Wavlm: Large- scale self-supervised pre-training for full stack speech processing,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition
UME upcycles pretrained dense ASR checkpoints into larger MoE models via weight copying, layer freezing, and expert balancing, yielding up to 11.9% relative CER reduction and 86.7% training time savings versus training from scratch.