A cross-layer micro-batch scheduler reduces GPU idling in distributed MoE inference by jointly executing workloads from different layers and deferring others to later steps.
Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles , pages=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference
A cross-layer micro-batch scheduler reduces GPU idling in distributed MoE inference by jointly executing workloads from different layers and deferring others to later steps.