Guiding Soft MoE dispatch weights with foreground segmentation masks plus a zero-initialized LayerScale improves ImageNet-1K top-1 by 0.6% and ImageNet-100 by 1.4% over a reproduced baseline.
Tutel: Adaptive mixture-of-experts at scale.Proceedings of Machine Learning and Systems, 5:269–287, 2023
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Guiding the Experts: Semantic Priors for Efficient and Focused MoE Routing
Guiding Soft MoE dispatch weights with foreground segmentation masks plus a zero-initialized LayerScale improves ImageNet-1K top-1 by 0.6% and ImageNet-100 by 1.4% over a reproduced baseline.