REVIEW 3 major objections 6 minor 60 references
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper argues that by predicting one iteration ahead, MoE load balancing can be reduced to lightweight expert replication and hidden communication, achieving up to 2.66x end-to-end speedup.
desk verdict A solid MoE load-balancing system with a genuinely new placement-plus-overlap combination, whose abstract overstates and whose locality assumption needs stress-testing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the lightweight expert placement: each expert is mapped independently to one or more devices, and only its parameters (forward Trans) and gradients (backward Agg) move among those devices, rather than whole model states. The planner's performance model estimates a MoE layer's time as $T' = 4T_{A2A}(R)+3T_{FEC}(H)+T_{Trans}(s,n)+T_{Agg}(s,n)$, with the last two terms replaced by their optimally overlapped versions after scheduling, and a locality-based greedy algorithm iteratively moves the most-loaded experts' parameters to the devices holding the most of their inputs until the load spread satisfies $\max(H)-\min(H)<\alpha I/E$. The scheduler's block-wise strategy partitions Trans and Agg into sub-operators and launches them in parallel with the forward expert and non-expert computations and the backward non-expert and expert computations of the neighboring block, while Plan is precomputed during the previous iteration's all-to-all communication.
What would settle it
Feed a MoE model inputs whose routing is deliberately perturbed every few iterations so that adjacent-iteration expert-load vectors have high divergence, and measure Pro-Prophet's per-iteration time; if it degrades toward the no-overlap baseline and planning or prefetch corrections dominate, the locality assumption is falsified. Quantitatively, compute the Jensen-Shannon divergence between consecutive iterations' expert-load distributions; Pro-Prophet's speedup should decrease monotonically as this divergence increases.
Extended reading notes
Core claim
Pro-Prophet's central claim is that the dynamic device load imbalance that plagues expert-parallel training of Mixture-of-Experts models can be corrected cheaply instead of expensively. The paper observes that the per-expert input distribution of a MoE layer changes little between adjacent training iterations, and builds two components on this locality. The planner searches over lightweight expert placements, placements where each expert is replicated to exactly the subset of devices that currently receives its tokens and only that expert's parameters and gradients are communicated, using a performance model and a greedy search to produce a communication-efficient placement. The scheduler then moves the three data-dependent load-balancing primitives, planning, parameter transfer, and gradient aggregation, off the critical path by computing the plan one iteration ahead, prefetching expert parameters, and splitting transfers into sub-operators overlapped with neighboring blocks' forward and backward computations. The paper reports end-to-end speedups of 1.18-2.66x over the main baseline and 1.01-1.50x over the dynamic-shadowing baseline across four clusters and five MoE-GPT variants, with up to 11.01x improvement in its load-balance metric.
Load-bearing premise
The scheduling rests on locality, the premise that each expert's share of inputs changes slowly between adjacent training iterations, and the paper tests this only in the first 100 iterations, where the distribution is still stabilizing; a sharp shift in routing would make the precomputed plan and prefetched parameters stale.
Editorial extensions
If this is right
- If Pro-Prophet's central claim holds, system-level load balancing for MoE training no longer requires moving whole model states: replicating only the parameters and gradients of selected experts to a subset of devices is enough to even out device load.
- The locality of expert input distributions implies the load-balancing plan can be computed one iteration ahead and expert parameters prefetched, so balancing overhead is largely removed from the critical path.
- Splitting communication primitives into sub-operators and overlapping them with forward and backward computation of neighboring MoE blocks can hide most remaining balancing communication, improving device utilization without changing model convergence behavior.
- Across four clusters of 8-32 GPUs and five MoE-GPT model sizes, the method reports end-to-end speedups of 1.18-2.66x over the primary baseline, up to 1.50x over the dynamic-shadowing baseline, and up to 11.01x improvement in its load-balance metric.
Reading between the lines
- Beyond the paper: the same one-iteration-ahead prediction could be applied to expert-parallel inference serving, where request routing is also temporally correlated, so prefetching expert weights to the devices that will receive the next batch could hide much of the serving latency.
- Beyond the paper: the scheduler uses a static offline split of Trans and Agg based on fixed non-MoE computation durations; making the split adaptive to measured per-iteration load divergence is a natural testable extension that would also stress the locality assumption.
- Beyond the paper: the performance model prices communication with a fixed average bandwidth, so on clusters where achievable bandwidth depends on which device pairs communicate, the placement search would likely benefit from topology-aware communication costs.
- Beyond the paper: the planner's search space grows with the number of experts per layer, so the method's advantage should be largest in the regime MoE models are headed toward, many more experts than devices.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Pro-Prophet, a system-level load-balancing method for training large-scale Mixture-of-Experts (MoE) models. It consists of a planner, which searches over 'lightweight expert placements' using a cost model and a greedy algorithm, and a scheduler, which exploits the observed locality of input distributions across adjacent iterations to precompute placement plans and to overlap parameter/gradient communication with expert and non-expert computation. The authors evaluate Pro-Prophet on four clusters of up to 32 GPUs with five MoE-GPT model variants and report end-to-end speedups of up to 2.66x over Deepspeed-MoE and up to 1.50x over FasterMoE, together with a load-balancing improvement of up to 11.01x over FasterMoE.
Significance. If the reported speedups are robust, Pro-Prophet would be a useful systems contribution: it addresses a real bottleneck in MoE training and combines two mechanisms (placement search and communication-computation overlap) that are usually treated separately. The paper gives a concrete performance model, a greedy planner, a block-wise scheduling strategy, and an ablation study separating the contribution of the planner and the scheduler. The evaluation covers multiple clusters, models, and top-k settings, which is a reasonable breadth for a systems paper. However, the central claim depends on an inter-iteration locality assumption whose empirical support is currently narrow, and the measurement methodology lacks variance information and a released artifact. These issues are load-bearing because the scheduler's precomputation and early parameter transfer are only beneficial if the next iteration's routing distribution is predictable.
major comments (3)
- [Section V-A and Section VI (Default settings)] The load-bearing assumption of the scheduler is the locality property: the input distribution of iteration j+1 is estimated from iteration j, so Plan_{i}^{j+1} is launched during A2A_i^j and Trans is started early (Algorithm 2, line 5). However, the only evidence for this property is Fig. 4, which shows a qualitative plot for iterations 2500-2600 of the second MoE layer, while the default evaluation is explicitly restricted to the first 100 iterations. The paper does not quantify how similar adjacent-iteration distributions are in the first 100 iterations, nor does it test a scenario with an abrupt distribution shift (e.g., changed data order, checkpoint restart, or altered batch composition). Because there is no described fallback if the predicted Plan or early Trans is stale, the reported end-to-end speedups may be confined to a regime where the locality assumption happens to hold. Please add locality measurements within the evaluation window itself and a stress test with a distribution shift, or describe a recovery mechanism.
- [Section VI (Figs. 10-12, Tables IV-V)] All speedup results are reported as single numbers without error bars, standard deviations, or the number of repeated runs. The headline ranges (e.g., 1.47-2.66x vs. Deepspeed-MoE, 1.01-1.48x vs. FasterMoE) are therefore not yet distinguished from run-to-run noise. In addition, Fig. 12 shows per-iteration times only for iterations 62-70, even though the evaluation is said to cover the first 100 iterations; a complete time curve or an average over the full window would be more convincing. I ask for at least three repeated runs per configuration and reporting of the mean and spread, together with the average throughput over the full 100-iteration window.
- [Section IV-B and Section V-C] The performance model uses constants B (average communication bandwidth) and t (computation throughput) in Eqs. (1)-(5), and the scheduler's split of Trans/Agg in Algorithm 2 relies on pre-training estimates of FNEC and BNEC duration. The paper does not state how these constants are obtained, whether they are re-measured for each cluster, or how sensitive the planner's chosen placement is to their values. Since the planner's greedy search uses this model to select the placement, and the scheduler's split depends on static estimates, the portability of the method to a new cluster is not yet demonstrated. Please report the calibration procedure and, if possible, a sensitivity analysis over B and t.
minor comments (6)
- [Abstract and Introduction] The abstract and introduction state that Pro-Prophet achieves 'up to 2.66x speedup compared to Deepspeed-MoE and FasterMoE,' but the measured speedups against FasterMoE are at most about 1.50x. Please rephrase to avoid implying that 2.66x was obtained against both baselines.
- [Section VI (Default settings)] The sentence 'We evaluate Pro-Prophet within the first 100 iterations as the input distribution tends to stabilize with the training process' is ambiguous: it could mean the first 100 iterations are chosen because the distribution is stabilizing, or that the distribution only stabilizes later. Please clarify the intended meaning, since it bears directly on the locality assumption.
- [Section VI-B (Single-iteration speedups)] The text introducing Fig. 12 says 'We also evaluate the single-layer performance of Pro-Prophet' but the experiment measures per-iteration time; this should read 'single-iteration performance.'
- [Eqs. (4)-(5)] The denominator in T_Trans(s,n) and T_Agg(s,n) is D*B, but only D-n devices participate in the communication. If the intended model is aggregate bandwidth across D devices, please justify this; otherwise the denominator should presumably involve D-n.
- [Algorithm 1 and Section IV-C] The pseudocode uses the variable i both for the index of the heaviest device (line 6) and for 'expert-i' in the comment on line 11; since a lightweight placement maps experts to subsets of devices, please disambiguate the notation so the mapping between devices and experts is clear.
- [General] No code or artifact release is mentioned. For a systems paper whose claims rest on careful implementation details (e.g., the split of Trans into sub-primitives), releasing the implementation or at least a detailed reproducibility appendix would substantially strengthen the paper.
Circularity Check
No significant circularity: the end-to-end speedup claims are measured against external baselines, and the performance model and locality assumption are internal heuristics with independent validation.
full rationale
The paper's central claim—end-to-end training speedups of 1.36–2.66x over Deepspeed-MoE and 1.01–1.50x over FasterMoE—is established by direct measurement on four clusters, not by derivation from the paper's own model. The performance model (Eqs. 1–6 and 8) is used internally to guide the greedy placement search; its accuracy is checked against measured operation times in Fig. 13, and the final speedups are external end-to-end comparisons, so no fitted constant is renamed as a prediction. The locality assumption (Section II-B and V-A) is an empirical premise that the input distribution of adjacent iterations is similar; it is load-bearing for the scheduler, but it is not defined in terms of the claimed speedup, and a failure of locality would be a validation gap rather than a circular derivation. Self-citations (e.g., Merak [37], AutoPipe [47], and the communication-scheduling work [58]) appear only in related-work surveys and do not supply the load-bearing result; no uniqueness theorem or prior ansatz is imported to force the design. Thus the derivation chain is self-contained with respect to its external benchmarks, and no circular step can be exhibited from the paper's equations or citations.
Assumptions & free parameters
free parameters (4)
- alpha (balance threshold) =
not stated
- n (number of devices an expert avoids) =
not stated
- Planner launch frequency =
not stated
- Trans/Agg split ratios in block-wise scheduling =
not stated
assumptions (4)
- domain assumption Input routing distributions of adjacent training iterations are similar (locality).
- domain assumption Execution times combine linearly according to Eqs. 1-5 with fixed bandwidth and compute throughput.
- domain assumption Plan for iteration j+1 can be computed from iteration j's distribution, and Agg can be shifted within the iteration without breaking correctness.
- domain assumption A device can store parameters of several experts at once in lightweight placements.
Cite this review
Pith. "Pith review of Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models." pith.science (2026). https://pith.science/paper/DCBNNXXP
@misc{pith2026241110003,
author = {Pith},
title = {Pith review of: Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/DCBNNXXP}},
note = {Machine review of arXiv:2411.10003}
}
read the original abstract
The size of deep learning models has been increasing to enhance model quality. The linear increase in training computation budget with model size means that training an extremely large-scale model is exceedingly time-consuming. Recently, the Mixture of Expert (MoE) has drawn significant attention as it can scale models to extra-large sizes with a stable computation budget. However, inefficient distributed training of large-scale MoE models hinders their broader application. Specifically, a considerable dynamic load imbalance occurs among devices during training, significantly reducing throughput. Several load-balancing works have been proposed to address the challenge. System-level solutions draw more attention for their hardware affinity and non-disruption of model convergence compared to algorithm-level ones. However, they are troubled by high communication costs and poor communication-computation overlapping. To address these challenges, we propose a systematic load-balancing method, Pro-Prophet, which consists of a planner and a scheduler for efficient parallel training of large-scale MoE models. To adapt to the dynamic load imbalance, we profile training statistics and use them to design Pro-Prophet. For lower communication volume, Pro-Prophet planner determines a series of lightweight load-balancing strategies and efficiently searches for a communication-efficient one for training based on the statistics. For sufficient overlapping of communication and computation, Pro-Prophet scheduler schedules the data-dependent operations based on the statistics and operation features, further improving the training throughput. Experimental results indicate that Pro-Prophet achieves up to 2.66x speedup compared to Deepspeed-MoE and FasterMoE. Additionally, Pro-Prophet achieves a load-balancing enhancement of up to 11.01x when compared to FasterMoE.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Gshard: Scaling giant models with conditional computation and automatic sharding,
D. Lepikhin, H. Lee, Y . Xu, D. Chen, O. Firat, Y . Huang, M. Krikun, N. Shazeer, and Z. Chen, “Gshard: Scaling giant models with conditional computation and automatic sharding,” in International Conference on Learning Representations, 2021
2021
-
[2]
Glam: Efficient scaling of language models with mixture-of-experts,
N. Du, Y . Huang, A. M. Dai, S. Tong, D. Lepikhin, Y . Xu, M. Krikun, Y . Zhou, A. W. Yu, O. Firat et al. , “Glam: Efficient scaling of language models with mixture-of-experts,” in International Conference on Machine Learning . PMLR, 2022, pp. 5547–5569
2022
-
[3]
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model,
A. Liu, B. Feng, B. Wang, B. Wang, B. Liu, C. Zhao, C. Dengr, C. Ruan, D. Dai, D. Guo et al., “Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model,” arXiv preprint arXiv:2405.04434 , 2024
arXiv 2024
-
[4]
Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,
S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y . Aminabadi, A. A. Awan, J. Rasley, and Y . He, “Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,” in International Conference on Machine Learning . PMLR, 2022, pp. 18 332–18 346
2022
-
[5]
Tutel: Adaptive mixture-of-experts at scale,
C. Hwang, W. Cui, Y . Xiong, Z. Yang, Z. Liu, H. Hu, Z. Wang, R. Salas, J. Jose, P. Ram et al. , “Tutel: Adaptive mixture-of-experts at scale,” Proceedings of Machine Learning and Systems , vol. 5, pp. 269–287, 2023
2023
-
[6]
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,”The Journal of Machine Learning Research , vol. 23, no. 1, pp. 5232–5270, 2022
2022
-
[7]
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in International Conference on Learning Representations, 2017
2017
-
[8]
Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models,
J. He, J. Zhai, T. Antunes, H. Wang, F. Luo, S. Shi, and Q. Li, “Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models,” in Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming, 2022, pp. 120–134. 13
work page 2022
Show all 60 references
-
[9]
Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,
X. Nie, X. Miao, Z. Wang, Z. Yang, J. Xue, L. Ma, G. Cao, and B. Cui, “Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,” Proceedings of the ACM on Management of Data, vol. 1, no. 1, pp. 1–19, 2023
2023
-
[10]
Zero: Memory optimizations toward training trillion parameter models,
S. Rajbhandari, J. Rasley, O. Ruwase, and Y . He, “Zero: Memory optimizations toward training trillion parameter models,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2020, pp. 1–16
2020
-
[11]
Scaling laws for neural language models,
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models,” arXiv preprint arXiv:2001.08361 , 2020
2001 arXiv
-
[12]
Bert: Pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805, 2018
2018 arXiv
-
[13]
Exploring the limits of transfer learning with a unified text-to-text transformer,
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research, vol. 21, no. 1, pp. 5485–5551, 2020
2020
-
[14]
Xlnet: Generalized autoregressive pretraining for language understanding,
Z. Yang, Z. Dai, Y . Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V . Le, “Xlnet: Generalized autoregressive pretraining for language understanding,” Advances in neural information processing systems , vol. 32, 2019
2019
-
[15]
Roberta: A robustly optimized bert pretraining approach,
Y . Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V . Stoyanov, “Roberta: A robustly optimized bert pretraining approach,” arXiv preprint arXiv:1907.11692 , 2019
1907 arXiv
-
[16]
Language models are unsupervised multitask learners,
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al., “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
-
[17]
Language mod- els are few-shot learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., “Language mod- els are few-shot learners,” Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020
1901
-
[18]
Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model,
S. Smith, M. Patwary, B. Norick, P. LeGresley, S. Rajbhandari, J. Casper, Z. Liu, S. Prabhumoye, G. Zerveas, V . Korthikanti et al. , “Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model,” arXiv preprint arXiv:2201.11990 , 2022
2022 arXiv
-
[19]
Mixture of a million experts,
X. O. He, “Mixture of a million experts,” arXiv preprint arXiv:2407.04153, 2024
2024 arXiv
-
[20]
Taming sparsely activated transformer with stochastic experts,
S. Zuo, X. Liu, J. Jiao, Y . J. Kim, H. Hassan, R. Zhang, T. Zhao, and J. Gao, “Taming sparsely activated transformer with stochastic experts,” arXiv preprint arXiv:2110.04260 , 2021
2021 arXiv
-
[21]
Scaling laws for fine-grained mixture of experts,
J. Ludziejewski, J. Krajewski, K. Adamczewski, M. Pi ´oro, M. Krutul, S. Antoniak, K. Ciebiera, K. Kr ´ol, T. Odrzyg ´o´zd´z, P. Sankowski et al., “Scaling laws for fine-grained mixture of experts,” in Forty-first Inter- national Conference on Machine Learning
-
[22]
Sparse upcycling: Training mixture-of-experts from dense checkpoints,
A. Komatsuzaki, J. Puigcerver, J. Lee-Thorp, C. R. Ruiz, B. Mustafa, J. Ainslie, Y . Tay, M. Dehghani, and N. Houlsby, “Sparse upcycling: Training mixture-of-experts from dense checkpoints,” arXiv preprint arXiv:2212.05055, 2022
2022 arXiv
-
[23]
Go wider instead of deeper,
F. Xue, Z. Shi, F. Wei, Y . Lou, Y . Liu, and Y . You, “Go wider instead of deeper,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 8, 2022, pp. 8779–8787
2022
-
[24]
One student knows all experts know: From sparse to dense,
F. Xue, X. He, X. Ren, Y . Lou, and Y . You, “One student knows all experts know: From sparse to dense,” arXiv preprint arXiv:2201.10890, 2022
2022 arXiv
-
[25]
St-moe: Designing stable and transferable sparse expert models,
B. Zoph, I. Bello, S. Kumar, N. Du, Y . Huang, J. Dean, N. Shazeer, and W. Fedus, “St-moe: Designing stable and transferable sparse expert models,” arXiv preprint arXiv:2202.08906 , 2022
2022 arXiv
-
[26]
Gpt-4 technical report,
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023
2023 arXiv
-
[27]
Stabilization of planar collective motion: All-to-all communication,
R. Sepulchre, D. A. Paley, and N. E. Leonard, “Stabilization of planar collective motion: All-to-all communication,” IEEE Transactions on automatic control, vol. 52, no. 5, pp. 811–824, 2007
2007
-
[28]
Optimization of all-to-all communication on the blue gene/l supercomputer,
S. Kumar, Y . Sabharwal, R. Garg, and P. Heidelberger, “Optimization of all-to-all communication on the blue gene/l supercomputer,” in 2008 37th International Conference on Parallel Processing. IEEE, 2008, pp. 320–329
2008
-
[29]
The hierarchical factor algorithm for all-to-all communication,
P. Sanders and J. L. Tr ¨aff, “The hierarchical factor algorithm for all-to-all communication,” in Euro-Par 2002 Parallel Processing: 8th International Euro-Par Conference Paderborn, Germany, August 27–30, 2002 Proceedings 8 . Springer, 2002, pp. 799–803
2002
-
[30]
Megatron-lm: Training multi-billion parameter language models using model parallelism,
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catan- zaro, “Megatron-lm: Training multi-billion parameter language models using model parallelism,” arXiv preprint arXiv:1909.08053 , 2019
1909 arXiv
-
[31]
Hetumoe: An effi- cient trillion-scale mixture-of-expert distributed training system,
X. Nie, P. Zhao, X. Miao, T. Zhao, and B. Cui, “Hetumoe: An effi- cient trillion-scale mixture-of-expert distributed training system,” arXiv preprint arXiv:2203.14685, 2022
2022 arXiv
-
[32]
Fastmoe: A fast mixture-of-expert training system,
J. He, J. Qiu, A. Zeng, Z. Yang, J. Zhai, and J. Tang, “Fastmoe: A fast mixture-of-expert training system,” arXiv preprint arXiv:2103.13262 , 2021
2021 arXiv
-
[33]
Moesys: A distributed and efficient mixture-of-experts training and inference system for internet services,
D. Yu, L. Shen, H. Hao, W. Gong, H. Wu, J. Bian, L. Dai, and H. Xiong, “Moesys: A distributed and efficient mixture-of-experts training and inference system for internet services,” IEEE Transactions on Services Computing, 2024
2024
-
[34]
Accelerating distributed moe training and inference with lina,
J. Li, Y . Jiang, Y . Zhu, C. Wang, and H. Xu, “Accelerating distributed moe training and inference with lina,” in2023 USENIX Annual Technical Conference (USENIX ATC 23) , 2023, pp. 945–959
2023
-
[35]
Janus: A unified distributed training framework for sparse mixture-of-experts models,
J. Liu, J. H. Wang, and Y . Jiang, “Janus: A unified distributed training framework for sparse mixture-of-experts models,” in Proceedings of the ACM SIGCOMM 2023 Conference, ACM SIGCOMM 2023, New York, NY, USA, 10-14 September 2023 . ACM, 2023, pp. 486–498
2023
-
[36]
Amp: Automatically finding model parallel strategies with heterogeneity awareness,
D. Li, H. Wang, E. Xing, and H. Zhang, “Amp: Automatically finding model parallel strategies with heterogeneity awareness,” Advances in Neural Information Processing Systems , vol. 35, pp. 6630–6639, 2022
2022
-
[37]
Merak: An efficient distributed dnn training framework with automated 3d parallelism for giant foundation models,
Z. Lai, S. Li, X. Tang, K. Ge, W. Liu, Y . Duan, L. Qiao, and D. Li, “Merak: An efficient distributed dnn training framework with automated 3d parallelism for giant foundation models,” IEEE Transactions on Parallel and Distributed Systems , vol. 34, no. 5, pp. 1466–1478, 2023
2023
-
[38]
Hippie: A data-paralleled pipeline approach to improve memory-efficiency and scalability for large dnn training,
X. Ye, Z. Lai, S. Li, L. Cai, D. Sun, L. Qiao, and D. Li, “Hippie: A data-paralleled pipeline approach to improve memory-efficiency and scalability for large dnn training,” in 50th International Conference on Parallel Processing, 2021, pp. 1–10
2021
-
[39]
Colossal-ai: A unified deep learning system for large-scale parallel training,
S. Li, H. Liu, Z. Bian, J. Fang, H. Huang, Y . Liu, B. Wang, and Y . You, “Colossal-ai: A unified deep learning system for large-scale parallel training,” in Proceedings of the 52nd International Conference on Parallel Processing, 2023, pp. 766–775
2023
-
[40]
Piper: Mul- tidimensional planner for dnn parallelization,
J. M. Tarnawski, D. Narayanan, and A. Phanishayee, “Piper: Mul- tidimensional planner for dnn parallelization,” Advances in Neural Information Processing Systems , vol. 34, pp. 24 829–24 840, 2021
2021
-
[41]
Parallel intelligent computing: development and challenges,
K. Lu, Z. Lai, S. Li, W. Liu, K. Ge, X. Lu, and D. Li, “Parallel intelligent computing: development and challenges,” SCIENTIA SINICA Informationis, vol. 53, no. 8, pp. 1441–1468, 2023
2023
-
[42]
Zero- infinity: Breaking the gpu memory wall for extreme scale deep learning,
S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith, and Y . He, “Zero- infinity: Breaking the gpu memory wall for extreme scale deep learning,” in Proceedings of the international conference for high performance computing, networking, storage and analysis , 2021, pp. 1–14
2021
-
[43]
Zero-offload: Democratizing billion-scale model training,
J. Ren, S. Rajbhandari, R. Y . Aminabadi, O. Ruwase, S. Yang, M. Zhang, D. Li, and Y . He, “Zero-offload: Democratizing billion-scale model training,” in 2021 USENIX Annual Technical Conference (USENIX ATC 21), 2021, pp. 551–564
2021
-
[44]
Pytorch fsdp: experiences on scaling fully sharded data parallel,
Y . Zhao, A. Gu, R. Varma, L. Luo, C.-C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer et al., “Pytorch fsdp: experiences on scaling fully sharded data parallel,” arXiv preprint arXiv:2304.11277 , 2023
2023 arXiv
-
[45]
Mics: near-linear scaling for training gigantic model on public cloud,
Z. Zhang, S. Zheng, Y . Wang, J. Chiu, G. Karypis, T. Chilimbi, M. Li, and X. Jin, “Mics: near-linear scaling for training gigantic model on public cloud,” arXiv preprint arXiv:2205.00119 , 2022
2022 arXiv
-
[46]
Maximizing parallelism in distributed training for huge neural networks,
Z. Bian, Q. Xu, B. Wang, and Y . You, “Maximizing parallelism in distributed training for huge neural networks,” arXiv preprint arXiv:2105.14450, 2021
2021 arXiv
-
[47]
Autopipe: A fast pipeline parallelism approach with balanced partitioning and micro- batch slicing,
W. Liu, Z. Lai, S. Li, Y . Duan, K. Ge, and D. Li, “Autopipe: A fast pipeline parallelism approach with balanced partitioning and micro- batch slicing,” in 2022 IEEE International Conference on Cluster Computing (CLUSTER). IEEE, 2022, pp. 301–312
2022
-
[48]
Hph: Hybrid parallelism on heterogeneous clusters for accelerating large-scale dnns training,
Y . Duan, Z. Lai, S. Li, W. Liu, K. Ge, P. Liang, and D. Li, “Hph: Hybrid parallelism on heterogeneous clusters for accelerating large-scale dnns training,” in 2022 IEEE International Conference on Cluster Computing (CLUSTER). IEEE, 2022, pp. 313–323
2022
-
[49]
Sequence paral- lelism: Long sequence training from system perspective,
S. Li, F. Xue, C. Baranwal, Y . Li, and Y . You, “Sequence paral- lelism: Long sequence training from system perspective,” arXiv preprint arXiv:2105.13120, 2021
2021 arXiv
-
[50]
Reducing activation recomputation in large transformer models,
V . A. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro, “Reducing activation recomputation in large transformer models,” Proceedings of Machine Learning and Systems, vol. 5, pp. 341–353, 2023
2023
-
[51]
Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models,
S. A. Jacobs, M. Tanaka, C. Zhang, M. Zhang, S. L. Song, S. Rajbhan- dari, and Y . He, “Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models,” arXiv preprint arXiv:2309.14509, 2023. 14
2023 arXiv
-
[52]
Ring attention with blockwise transformers for near-infinite context,
H. Liu, M. Zaharia, and P. Abbeel, “Ring attention with blockwise transformers for near-infinite context,” arXiv preprint arXiv:2310.01889, 2023
2023 arXiv
-
[53]
Bagualu: targeting brain scale pretrained models with over 37 million cores,
Z. Ma, J. He, J. Qiu, H. Cao, Y . Wang, Z. Sun, L. Zheng, H. Wang, S. Tang, T. Zheng et al. , “Bagualu: targeting brain scale pretrained models with over 37 million cores,” in Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Program- ming, 2...
2022
-
[54]
Parm: Efficient training of large sparsely-activated models with dedicated schedules,
X. Pan, W. Lin, S. Shi, X. Chu, W. Sun, and B. Li, “Parm: Efficient training of large sparsely-activated models with dedicated schedules,” in IEEE INFOCOM 2024-IEEE Conference on Computer Communications. IEEE, 2024, pp. 1880–1889
2024
-
[55]
A hybrid tensor-expert-data parallelism approach to optimize mixture- of-experts training,
S. Singh, O. Ruwase, A. A. Awan, S. Rajbhandari, Y . He, and A. Bhatele, “A hybrid tensor-expert-data parallelism approach to optimize mixture- of-experts training,” in Proceedings of the 37th International Conference on Supercomputing, 2023, pp. 203–214
2023
-
[56]
Mg-wfbp: Efficient data communication for distributed synchronous sgd algorithms,
S. Shi, X. Chu, and B. Li, “Mg-wfbp: Efficient data communication for distributed synchronous sgd algorithms,” in IEEE INFOCOM 2019- IEEE Conference on Computer Communications. IEEE, 2019, pp. 172– 180
2019
-
[57]
Pipetransformer: Automated elastic pipelining for distributed training of transformers,
C. He, S. Li, M. Soltanolkotabi, and S. Avestimehr, “Pipetransformer: Automated elastic pipelining for distributed training of transformers,” arXiv preprint arXiv:2102.03161 , 2021
2021 arXiv
-
[58]
A multidimensional communication scheduling method for hybrid parallel dnn training,
S. Li, K. Lu, Z. Lai, W. Liu, K. Ge, and D. Li, “A multidimensional communication scheduling method for hybrid parallel dnn training,” IEEE Transactions on Parallel and Distributed Systems , 2024
2024
-
[59]
Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,
S. Shi, X. Pan, Q. Wang, C. Liu, X. Ren, Z. Hu, Y . Yang, B. Li, and X. Chu, “Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,” in Proceedings of the Nineteenth European Conference on Computer Systems , 2024, pp. 236–249
2024
-
[60]
Pipemoe: Accelerating mixture- of-experts through adaptive pipelining,
S. Shi, X. Pan, X. Chu, and B. Li, “Pipemoe: Accelerating mixture- of-experts through adaptive pipelining,” in IEEE INFOCOM 2023-IEEE Conference on Computer Communications . IEEE, 2023, pp. 1–10
2023
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.