HierMoE reduces MoE training time by removing duplicate token copies at each GPU-hierarchy level and swapping experts for load balance, measured at 1.18-1.27x end-to-end speedup on 32 GPUs.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
HierMoE reduces MoE training time by removing duplicate token copies at each GPU-hierarchy level and swapping experts for load balance, measured at 1.18-1.27x end-to-end speedup on 32 GPUs.