A supervised mixture of experts, routed by fixed bandwidth and task labels, improves multi-task ASR/ST over hard parameter sharing while keeping active parameters constant.
By us- ing guiding tokens instead of dynamic gating functions, S-MoE ensures efficient training and inference while improving per- formance across various tasks
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Beyond Hard Sharing: Efficient Multi-Task Speech-to-Text Modeling with Supervised Mixture of Experts
A supervised mixture of experts, routed by fixed bandwidth and task labels, improves multi-task ASR/ST over hard parameter sharing while keeping active parameters constant.