Cluster-aware Upcycling breaks MoE expert symmetry by initializing experts from semantic activation clusters using SVD subspaces and cluster centroids for the router, plus self-distillation, yielding better zero- and few-shot performance on CLIP ViT models than standard upcycling.
Scaling vision with sparse mix- ture of experts.NeurIPS, 2021
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
Cluster-aware Upcycling breaks MoE expert symmetry by initializing experts from semantic activation clusters using SVD subspaces and cluster centroids for the router, plus self-distillation, yielding better zero- and few-shot performance on CLIP ViT models than standard upcycling.