D2MoE dynamically selects bit-width per expert per token, stores weights in a nested matryoshka-like layout, and schedules expert loading to improve on-device MoE LLM throughput by up to 1.39x with up to 53% memory reduction.
Gonzalez, Matei Zaharia, and Ion Stoica
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
D2MoE dynamically selects bit-width per expert per token, stores weights in a nested matryoshka-like layout, and schedules expert loading to improve on-device MoE LLM throughput by up to 1.39x with up to 53% memory reduction.