UME upcycles pretrained dense ASR checkpoints into larger MoE models via weight copying, layer freezing, and expert balancing, yielding up to 11.9% relative CER reduction and 86.7% training time savings versus training from scratch.
Bigssl: Exploring the frontier of large-scale semi-supervised learning for automatic speech recognition,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition
UME upcycles pretrained dense ASR checkpoints into larger MoE models via weight copying, layer freezing, and expert balancing, yielding up to 11.9% relative CER reduction and 86.7% training time savings versus training from scratch.