A shared per-token controller jointly routes attention resolution, FFN experts, and KV bit-width and is claimed to Pareto-dominate independently tuned MoD+MoE+KV-quant at matched cost while protecting rare-token accuracy.
DeepSpeed-MoE: Advancing mixture-of- experts inference and training to power next-generation AI scale
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
MoP combines specialized parallelisms and a new optimizer step for memory-efficient MoE training, delivering 4.7x-8.2x higher per-GPU throughput than FSDP2 while supporting 1M-token contexts on 12 8x H200 nodes.
citing papers explorer
-
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation
A shared per-token controller jointly routes attention resolution, FFN experts, and KV bit-width and is claimed to Pareto-dominate independently tuned MoD+MoE+KV-quant at matched cost while protecting rare-token accuracy.
-
Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models
MoP combines specialized parallelisms and a new optimizer step for memory-efficient MoE training, delivering 4.7x-8.2x higher per-GPU throughput than FSDP2 while supporting 1M-token contexts on 12 8x H200 nodes.