REVIEW 2 major objections 5 minor 1 cited by
TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training
T0 review · 2 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read TAOT claims a topology-aware optimal-transport planner for dynamic expert replicas cuts MoE training time by 1.43x and achieves the lowest expert-communication cost across all tested configurations.
desk verdict Solid MoE load-balancing paper with a real measured speedup and a communication-cost claim calibrated to its own lambda=3 assumption. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is a balanced entropy-regularized optimal transport problem at rank level: supply is the per-rank overload, demand is the per-rank spare capacity, and the cost matrix W has entry 1 for intra-node moves and lambda for inter-node moves. Its regularized solution has Gibbs-kernel form T = diag(u) exp(-W/eps) diag(v), with eps=lambda, so the intra-node kernel entry is exp((lambda-1)/lambda) times the inter-node one: a soft topology preference rather than a hard graph constraint. That flow hint feeds a column-first integer matching (balancing gain first, topology preference second, OT hint as tiebreaker) and then a Lagrangian auction assigns token counts to spare slots; system-level ove
What would settle it
Measure the actual intra-node versus inter-node weight-transfer cost on the target cluster and rerun Table 1 with that ratio; if TAOT no longer has the lowest weighted communication cost in all ten rows, or if the end-to-end speedup on a cluster with a different interconnect gap drops below 1.43x, the quantitative claim is hardware-specific rather than general.
Extended reading notes
Core claim
The central claim is that replica placement for hot experts should be solved as a joint optimization over residual imbalance and the topology-dependent cost of moving expert weights, rather than for balance alone. TAOT treats the excess load of hot ranks as supply and the spare capacity of cold ranks as demand, with a cost matrix that charges 1 for same-node expert-weight transfer and lambda=3 for cross-node transfer. Entropy-regularized Sinkhorn-Knopp iterations produce a soft flow hint, phase-2 column-first matching turns the hint into binary replica decisions that exhaust intra-node capacity before reaching across nodes, and a phase-3 Lagrangian auction assigns exact token counts to spare
Load-bearing premise
The load-bearing premise is that moving expert weights across nodes costs exactly three times moving them within a node, and that the same 1:3 ratio is the right yardstick for comparing communication costs; if the hardware ratio differs, the planner's preferences and the headline reduction are measured in that assumed scale.
Editorial extensions
If this is right
- End-to-end training time drops from 155.4 ms to 108.8 ms per forward-plus-backward step, a 1.43x speedup, while loss trajectories stay within roughly ±3 per mille over 100 steps.
- Balance quality remains competitive with existing methods and becomes best or tied for best at EP=32, while weighted expert-communication cost is lowest in all ten tested imbalance/EP configurations.
- The gain grows with expert-parallel scale (up to 1.79x at EP=16) and with initial imbalance (up to 1.75x at 90%), because more spare ranks enlarge the low-cost offload search space.
- Online planning overhead stays below 1% of forward time, making micro-batch-level replanning practical; about two spare slots per rank balances performance against memory in the tested setup.
Reading between the lines
- The same supply/demand/cost decomposition could be applied to other dynamic weight-migration decisions, such as optimizer-state migration or inference-time expert offloading, wherever distribution of a movable payload trades balance against interconnect cost.
- The framework is agnostic to the 1:3 ratio; measuring the true intra-node-to-inter-node bandwidth gap per cluster and rerunning placement would shift preferences toward even more intra-node bias if the gap is larger, or allow more cross-node offload if it is smaller.
- Because routing semantics are untouched, TAOT could be stacked with router-level balance losses or expert-choice routing, with statistical balancing and instantaneous peak shaving addressing different timescales.
- A useful external check is to report the ten-configuration comparison in raw transferred bytes as well as in the weighted metric, so the claimed 74% reduction can be evaluated independently of the assumed weighting.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper addresses load imbalance in expert-parallel MoE training by dynamically replicating hot experts onto spare slots of lightly loaded ranks. TAOT decomposes replica placement into three phases: a rank-level balanced entropy-regularized optimal transport problem with a topology-dependent cost matrix (intra-node cost 1, inter-node cost λ), solved by Sinkhorn-Knopp iterations to produce a soft flow hint; a column-priority integer matching phase that assigns expert replicas to cold ranks using balance gain, topology preference, and the OT hint; and a Lagrange-auction token-allocation phase. The system also overlaps guest-weight transfer with home-expert GEMM. Experiments on 32 A800 GPUs with Qwen3-30B-A3B report a 1.43x end-to-end speedup over Megatron-LM, balance quality comparable to or better than ECHO, LLEP, and LPLB on synthetic imbalance traces, and the lowest weighted expert-communication cost in all ten tested configurations, with up to a 74% reduction.
Significance. If the claims hold, the paper makes a useful systems contribution: it is, to my knowledge, the first replica-placement scheme for MoE training that explicitly includes both peak-shaving gain and topology-dependent weight-movement cost in the objective, and it ships a complete, GPU-friendly algorithm (Algorithms 1–3 plus appendix derivations) with measured planning overhead below 1% of forward time. The 100-step loss-consistency check (mean absolute relative error 0.297 per mille) is a good guard against routing-semantics change, and the ablations show each phase contributes. However, the headline differentiator—lowest weighted expert-communication cost—is evaluated under the same 1:3 intra/inter weighting that the planner itself assumes, and no hardware microbenchmark is provided to justify that ratio. The quantitative communication-cost advantage is therefore conditional on an unmeasured parameter, and the end-to-end speedup is a single-configuration, single-run measurement against Megatron-LM only.
major comments (2)
- [Experimental Setup; A.1 Eq. (10); Table 1] The central claim 'lowest weighted expert-communication cost across all ten configurations, up to 74% reduction' is evaluated with a 1:3 intra/inter ratio that is identical to the planner's own cost matrix (W=1 intra, λ=3 inter in A.1 Eq. 10; Experimental Setup reports the weighted cost 'with a 1:3 intra/inter ratio'). No microbenchmark is given for the A800 NVLink-to-InfiniBand bandwidth ratio. This is not merely a cosmetic issue: for the EP=16, 70% row, the reported TAOT weighted cost 27 corresponds to 9 intra + 6 inter transfers (9+6·3), while LPLB's cost 32 is 32 intra + 0 inter (32+0·3). The comparison flips for any λ>3.83, and at λ=10 TAOT costs 69 against LPLB's 32. Thus the 'lowest cost' and 'up to 74% reduction' statements are artifacts of an uncalibrated coefficient in at least one configuration. Since the same λ steers the Sinkhorn kernel (Eq. 15) and the topology bonuses in P
- [Experimental Setup; B.2; Figure 2] The end-to-end speedup evidence is thinner than the headline suggests. The 1.43x/42.82% figure is one measurement on one configuration (TP4/PP2/EP16, 32 A800s, Qwen3-30B-A3B on Pile-test), reported as a mean of 10 consecutive steps after 20 warm-up steps. No repeated runs, seeds, or standard deviations are given, and the end-to-end comparison is against Megatron-LM only; ECHO, LLEP, and LPLB—the methods compared in Table 1—are not run end-to-end. The paper therefore does not currently establish that TAOT improves end-to-end throughput over the dynamic-replica baselines it claims to outperform, only that it beats a static baseline in one configuration. At minimum, report variance across multiple trials and add an end-to-end comparison against at least ECHO or LPLB in the same harness.
minor comments (5)
- [Table 1] The table formatting is corrupted or ambiguous: several TAOT rows read as '178 3' or '3817 7' with no clear separation between weighted cost, intra-node, and inter-node counters. The reader cannot reliably parse the numbers. Please render as a proper table with distinct columns.
- [Experimental Setup] The statement that 'balancing and communication behavior is corpus-independent' is asserted, not demonstrated. Router load statistics can be corpus-sensitive; if this is meant only as a claim about the algorithm being corpus-agnostic, please rephrase or provide evidence.
- [Experiments, after Table 1] The 1:7 expert-communication/computation ratio and the resulting '1pp imbalance ~ 7% communication cost' critical value are introduced without a measurement or derivation. This is a load-bearing justification for reporting balance and cost separately; please provide the measurement or mark it as an assumption.
- [A.3, Eq. (21)] The notation (Ter)_norm is used in the Phase 2 score before it is defined in the appendix; define it in the main text or move the definition up.
- [Abstract and Conclusion] The abstract says '1.43x end-to-end MoE training speedup' without specifying the baseline; the body correctly says 'over Megatron-LM.' Please make the baseline explicit in the abstract. Also, 'up to 74% reduction' refers to LPLB at EP=32, which is not clear at first read.
Circularity Check
The headline communication-cost advantage is evaluated with the same 1:3 intra/inter weighting that the planner itself uses, so Table 1's 'lowest cost' is partly self-defined; the end-to-end speedup is independent.
-
self definitional
[Experimental Setup (Metrics); A.1 Eq. (10) and Eq. (2)/(9)]
"We also report the number of intra-/inter-node expert transfers and the weighted expert-communication cost (with a 1:3 intra/inter ratio). ... W_{r_e,r} = 1 if r_e and r are on the same node; λ if r_e and r are on different nodes. ... we set λ = 3 in our experiments."
The evaluation metric for the paper's central differentiator — 'lowest weighted expert-communication cost across all configurations, up to 74% reduction' — is the same cost model that the TAOT planner optimizes. Eq. (9) minimizes Σ z_er·W_{r_e,r} with λ=3, and the Experimental Setup measures weighted cost with the identical 1:3 intra/inter ratio. Beating ECHO/LLEP/LPLB on this metric is therefore partly a comparison of each baseline against TAOT's own objective; the baselines optimize different objectives and are not expected to minimize this particular weighting. The ranking is also sensitive to λ: at EP=16, 70%, TAOT's 9 intra + 6 inter transfers give 27 units at λ=3, below LPLB's fixed 32 intra, but at λ≥4 TAOT becomes 9+24=33 and LPLB is cheaper. Thus the 'lowest cost' claim is conditi
full rationale
TAOT's derivation is mostly self-contained: the three-phase OT planner, integer matching, and auction assignment are described with explicit equations and algorithms, and the end-to-end speedup is measured as real F+B time against Megatron-LM, which is an external benchmark. The loss-consistency check provides independent evidence that routing semantics are preserved. The only significant circular element is the communication-cost evaluation. The paper's headline 'lowest weighted expert-communication cost' is computed using the same 1:3 intra/inter ratio that appears as λ=3 in the planner's cost matrix W (Eq. 10) and objective (Eq. 9). Consequently, Table 1's cost comparison is essentially TAOT being judged on its own objective function, not on an independently calibrated hardware cost. Moreover, the paper provides no microbenchmark justifying λ=3 on A800 clusters, and the ranking flips at λ≥4 in at least one configuration, so the 'up to 74% reduction' claim is partly an artifact of the chosen coefficient. This does not invalidate the measured speedup or the algorithm's internal consistency, but it does mean the communication-cost superiority claim is partially circular and should be treated as conditional on the unverified cost model. No other circularity was found: claims about balance quality use the standard imbalance metric, and no load-bearing self-citations or uniqueness-import arguments appear.
Assumptions & free parameters
free parameters (5)
- lambda (inter-node cost factor) =
3
- epsilon (entropy regularization) =
lambda (3)
- mu (trade-off coefficient in Eq. 2/9) =
not assigned
- w_inter (inter-node topology-preference value, Eq. 22) =
not assigned
- 0.1 (OT-hint coefficient in Phase 2 score) =
0.1
assumptions (5)
- domain assumption Per-rank time is proportional to routed token count; iteration time equals the max-load rank (Eqs. 5-7)
- domain assumption Guest replicas can be fetched and computed on cold ranks, with guest gradients returned and accumulated at the home rank, without changing training semantics
- ad hoc to paper Expert communication and computation take about a 1:7 ratio, making 1pp imbalance roughly equivalent to 7% communication cost
- ad hoc to paper Routing traffic is corpus-independent, so Pile-test is representative
- standard math Total supply equals total demand in exact arithmetic
Cite this review
Pith. "Pith review of TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training." pith.science (2026). https://pith.science/paper/UHXMYLI5
@misc{pith2026260803676,
author = {Pith},
title = {Pith review of: TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training},
year = {2026},
howpublished = {\url{https://pith.science/paper/UHXMYLI5}},
note = {Machine review of arXiv:2608.03676}
}
read the original abstract
Mixture-of-Experts (MoE) has become a key architecture for scaling large language models (LLMs), yet its dynamic routing causes severe load imbalance in expert-parallel training. Existing dynamic-replica methods copy hot experts onto idle ranks to share computation, but they optimize load balance alone and ignore the cost of moving expert weights across a multi-node topology, so the resulting cross-node communication can outweigh the balancing gain and inflate training cost. We present TAOT, a topology-aware optimal transport method for dynamic expert-replica placement. TAOT models the overload on hot ranks and the spare capacity on lightly loaded ranks as a balanced entropy-regularized optimal transport problem with a communication-cost matrix, solves it with Sinkhorn-Knopp iterations to produce rank-level flow hints, and combines integer replica matching with token assignment into an executable schedule. At the system level, it overlaps guest-weight transfer with home-expert computation to hide the communication overhead. Experiments show TAOT achieves a 1.43x end-to-end MoE training speedup, reaches balance quality competitive with or better than existing state-of-the-art methods, and attains the lowest weighted expert-communication cost across all configurations, with up to a 74% reduction.
Figures
Forward citations
Cited by 1 Pith paper
-
TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes
Expert cost in MoE serving follows a two-regime max-affine law, and a makespan-aware dispatcher over this model tracks the best fixed policy and wins where regimes mix.
Reference graph
Works this paper leans on
-
[2]
Cai, W.; Jiang, J.; Wang, F.; et al. 2025. A Survey on Mixture of Experts in Large Language Models. IEEE Transactions on Knowledge and Data Engineering
work page 2025
-
[3]
DeepSeek-AI . 2025. LPLB : Linear-Programming-Based Load Balancer for Mixture-of-Experts Models. https://github.com/deepseek-ai/LPLB
work page 2025
-
[4]
Du, N.; et al. 2022. GLaM : Efficient Scaling of Language Models with Mixture-of-Experts. In International Conference on Machine Learning (ICML), 5547--5569. PMLR
work page 2022
-
[5]
Fedus, W.; Zoph, B.; and Shazeer, N. 2022. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. Journal of Machine Learning Research, 23(120): 1--39
work page 2022
-
[7]
He, J.; Zhai, J.; Antunes, T.; Wang, H.; Luo, F.; Shi, S.; and Li, Q. 2022. FasterMoE : Modeling and Optimizing Training of Large-Scale Dynamic Pre-Trained Models. In Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP), 120--134
work page 2022
-
[8]
Hwang, C.; Cui, W.; Xiong, Y.; Yang, Z.; Liu, Z.; Hu, H.; Wang, Z.; Salas, R.; Jose, J.; Ram, P.; et al. 2023. Tutel: Adaptive Mixture-of-Experts at Scale. Proceedings of Machine Learning and Systems, 5: 269--287
work page 2023
-
[11]
Lepikhin, D.; Lee, H.; Xu, Y.; Chen, D.; Firat, O.; Huang, Y.; Krikun, M.; Shazeer, N.; and Chen, Z. 2021. GShard : Scaling Giant Models with Conditional Computation and Automatic Sharding. In International Conference on Learning Representations (ICLR)
work page 2021
-
[12]
Lewis, M.; Bhosale, S.; Dettmers, T.; Goyal, N.; and Zettlemoyer, L. 2021. BASE Layers: Simplifying Training of Large, Sparse Models. In International Conference on Machine Learning (ICML), 6265--6274. PMLR
work page 2021
Show all 50 references
-
[14]
Liu, X.; Wang, Y.; Fu, F.; Xiao, X.; Li, H.; Li, J.; and Cui, B. 2026. LAER-MoE : Load-Adaptive Expert Re-Layout for Efficient Mixture-of-Experts Training. In International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), volume 2, ...
2026
-
[15]
Nguyen, X.-P.; Pandit, S.; Xu, A.; Xiong, C.; and Joty, S. 2026. Least-Loaded Expert Parallelism: Load Balancing an Imbalanced Mixture-of-Experts. In International Conference on Machine Learning (ICML)
2026
-
[16]
NVIDIA Megatron-LM Team . 2025. MoE ECHO : Unlocking Sync-Free, Full CUDA-Graph Support for Dropless MoE via Elastic Cloning. Megatron-LM Pull Request \#2368. https://github.com/NVIDIA/Megatron-LM/pull/2368
2025
-
[18]
Y.; Awan, A
Rajbhandari, S.; Li, C.; Yao, Z.; Zhang, M.; Aminabadi, R. Y.; Awan, A. A.; Rasley, J.; and He, Y. 2022. DeepSpeed-MoE : Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale. In International Conference on Machine Learning (ICML), 18332--18346. PMLR
2022
-
[19]
Roller, S.; Sukhbaatar, S.; Szlam, A.; and Weston, J. 2021. Hash Layers for Large Sparse Models. In Advances in Neural Information Processing Systems (NeurIPS), volume 34, 17555--17566
2021
-
[20]
Shazeer, N.; Mirhoseini, A.; Maziarz, K.; Davis, A.; Le, Q.; Hinton, G.; and Dean, J. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In International Conference on Learning Representations (ICLR)
2017
-
[22]
Skiadopoulos, A.; Zhao, M.; Gandhi, S.; Norrie, T.; Mukherjee, S.; and Kozyrakis, C. 2026. SYMI : Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling. In USENIX Symposium on Networked Systems Design and Implementation (NSDI), 75--92
2026
-
[23]
N.; Kaiser, L.; and Polosukhin, I
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention Is All You Need. In Advances in Neural Information Processing Systems (NeurIPS), volume 30
2017
-
[24]
Wen, J.-L.; Li, X.-J.; Yao, J.-P.; Sun, H.-F.; and An, X.-R. 2026. Consensus-Expert DynamicMoE : ARIMA -Based Capacity Prediction with Adaptive Load Balancing for Sparse Models. International Journal of Computational Intelligence Systems, 19(1): 157
2026
-
[26]
Zeng, Y.; Huang, C.; Mei, Y.; Zhang, L.; Su, T.; Ye, W.; Shi, W.; and Wang, S. 2025. EfficientMoE : Optimizing Mixture-of-Experts Model Training with Adaptive Load Balance. IEEE Transactions on Parallel and Distributed Systems, 36(4): 677--688
2025
-
[27]
Zhai, M.; He, J.; Ma, Z.; Zong, Z.; Zhang, R.; and Zhai, J. 2023. SmartMoE : Efficiently Training Sparsely-Activated Models through Combining Offline and Online Parallelization. In 2023 USENIX Annual Technical Conference (USENIX ATC 23), 961--975
2023
-
[28]
Zhang, B.; Chen, X.; Zhang, S.; Zhang, S.; Zhou, X.; and Sun, M. 2026. FLEX-MoE : Federated Mixture-of-Experts with Load-Balanced Expert Assignment. In Proceedings of the AAAI Conference on Artificial Intelligence
2026
-
[30]
M.; Chen, Z.; Le, Q
Zhou, Y.; Lei, T.; Liu, H.; Du, N.; Huang, Y.; Zhao, V.; Dai, A. M.; Chen, Z.; Le, Q. V.; and Laudon, J. 2022. Mixture-of-Experts with Expert Choice Routing. In Advances in Neural Information Processing Systems (NeurIPS), 7103--7114
2022
-
[32]
and Kaiser, Lukasz and Polosukhin, Illia , title =
Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N. and Kaiser, Lukasz and Polosukhin, Illia , title =. Advances in Neural Information Processing Systems (NeurIPS) , volume =
-
[33]
International Conference on Learning Representations (ICLR) , year =
Shazeer, Noam and Mirhoseini, Azalia and Maziarz, Krzysztof and Davis, Andy and Le, Quoc and Hinton, Geoffrey and Dean, Jeff , title =. International Conference on Learning Representations (ICLR) , year =
-
[34]
International Conference on Learning Representations (ICLR) , year =
Lepikhin, Dmitry and Lee, HyoukJoong and Xu, Yuanzhong and Chen, Dehao and Firat, Orhan and Huang, Yanping and Krikun, Maxim and Shazeer, Noam and Chen, Zhifeng , title =. International Conference on Learning Representations (ICLR) , year =
-
[35]
Journal of Machine Learning Research , volume =
Fedus, William and Zoph, Barret and Shazeer, Noam , title =. Journal of Machine Learning Research , volume =
-
[36]
IEEE Transactions on Knowledge and Data Engineering , year =
Cai, Weilin and Jiang, Juyong and Wang, Fan and others , title =. IEEE Transactions on Knowledge and Data Engineering , year =
-
[37]
International Conference on Machine Learning (ICML) , pages =
Du, Nan and others , title =. International Conference on Machine Learning (ICML) , pages =
-
[38]
and others , title =
Jiang, Albert Q. and others , title =. arXiv preprint arXiv:2401.04088 , year =
-
[39]
arXiv preprint arXiv:2412.19437 , year =
Liu, Aixin and others , title =. arXiv preprint arXiv:2412.19437 , year =
-
[40]
arXiv preprint arXiv:2505.09388 , year =
Yang, An and others , title =. arXiv preprint arXiv:2505.09388 , year =
-
[41]
arXiv preprint arXiv:2507.20534 , year =
-
[42]
arXiv preprint arXiv:2601.11659 , year =
Adcock, Andrew and Srivastava, Aakanksha and Dubey, Abhimanyu and others , title =. arXiv preprint arXiv:2601.11659 , year =
-
[43]
arXiv preprint arXiv:1909.08053 , year =
Shoeybi, Mohammad and Patwary, Mostofa and Puri, Raul and LeGresley, Patrick and Casper, Jared and Catanzaro, Bryan , title =. arXiv preprint arXiv:1909.08053 , year =
1909 arXiv
-
[44]
International Conference on Machine Learning (ICML) , pages =
Rajbhandari, Samyam and Li, Conglong and Yao, Zhewei and Zhang, Minjia and Aminabadi, Reza Yazdani and Awan, Ammar Ahmad and Rasley, Jeff and He, Yuxiong , title =. International Conference on Machine Learning (ICML) , pages =
-
[45]
Proceedings of Machine Learning and Systems , volume =
Hwang, Changho and Cui, Wei and Xiong, Yifan and Yang, Ziyue and Liu, Ze and Hu, Han and Wang, Zilong and Salas, Rafael and Jose, Jithin and Ram, Prabhat and others , title =. Proceedings of Machine Learning and Systems , volume =
-
[46]
Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP) , pages =
He, Jiaao and Zhai, Jidong and Antunes, Tiago and Wang, Haojie and Luo, Fuwen and Shi, Shangfeng and Li, Qin , title =. Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP) , pages =
-
[47]
International Conference on Machine Learning (ICML) , year =
Nguyen, Xuan-Phi and Pandit, Shrey and Xu, Austin and Xiong, Caiming and Joty, Shafiq , title =. International Conference on Machine Learning (ICML) , year =
-
[48]
USENIX Symposium on Networked Systems Design and Implementation (NSDI) , pages =
Skiadopoulos, Athinagoras and Zhao, Mark and Gandhi, Swapnil and Norrie, Thomas and Mukherjee, Shreya and Kozyrakis, Christos , title =. USENIX Symposium on Networked Systems Design and Implementation (NSDI) , pages =
-
[49]
arXiv preprint arXiv:2202.08906 , year =
Zoph, Barret and Bello, Irwan and Kumar, Sameer and Du, Nan and Huang, Yanping and Dean, Jeff and Shazeer, Noam and Fedus, William , title =. arXiv preprint arXiv:2202.08906 , year =
-
[50]
and Chen, Zhifeng and Le, Quoc V
Zhou, Yanqi and Lei, Tao and Liu, Hanxiao and Du, Nan and Huang, Yanping and Zhao, Vincent and Dai, Andrew M. and Chen, Zhifeng and Le, Quoc V. and Laudon, James , title =. Advances in Neural Information Processing Systems (NeurIPS) , pages =
-
[51]
2023 USENIX Annual Technical Conference (USENIX ATC 23) , pages =
Zhai, Mingshu and He, Jiaao and Ma, Zixuan and Zong, Zan and Zhang, Runqing and Zhai, Jidong , title =. 2023 USENIX Annual Technical Conference (USENIX ATC 23) , pages =
2023
-
[52]
IEEE Transactions on Parallel and Distributed Systems , volume =
Zeng, Yan and Huang, Chengchuang and Mei, Yipeng and Zhang, Lifu and Su, Ting and Ye, Wei and Shi, Wenqi and Wang, Sheng , title =. IEEE Transactions on Parallel and Distributed Systems , volume =
-
[53]
arXiv preprint arXiv:2511.16947 , year =
Zhao, Chen and Wu, Wei and Song, Lei and Xu, Yang and Yuan, Yang , title =. arXiv preprint arXiv:2511.16947 , year =
-
[54]
arXiv preprint arXiv:2604.19654 , year =
Qi, Shaohuai and Liu, Hao and Zhao, Shuai , title =. arXiv preprint arXiv:2604.19654 , year =
-
[55]
International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) , volume =
Liu, Xin and Wang, Yujie and Fu, Fangcheng and Xiao, Xupeng and Li, Huanran and Li, Jinbao and Cui, Bin , title =. International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) , volume =
-
[56]
International Conference on Machine Learning (ICML) , pages =
Lewis, Mike and Bhosale, Shruti and Dettmers, Tim and Goyal, Naman and Zettlemoyer, Luke , title =. International Conference on Machine Learning (ICML) , pages =
-
[57]
Advances in Neural Information Processing Systems (NeurIPS) , volume =
Roller, Stephen and Sukhbaatar, Sainbayar and Szlam, Arthur and Weston, Jason , title =. Advances in Neural Information Processing Systems (NeurIPS) , volume =
-
[58]
International Journal of Computational Intelligence Systems , volume =
Wen, Jia-Long and Li, Xiao-Jun and Yao, Jian-Ping and Sun, Hai-Feng and An, Xin-Ran , title =. International Journal of Computational Intelligence Systems , volume =
-
[59]
Proceedings of the AAAI Conference on Artificial Intelligence , year =
Zhang, Bo and Chen, Xu and Zhang, Shuai and Zhang, Sheng and Zhou, Xu and Sun, Maosong , title =. Proceedings of the AAAI Conference on Artificial Intelligence , year =
-
[60]
arXiv preprint arXiv:2101.00027 , year =
Gao, Leo and Biderman, Stella and Black, Sid and Golding, Laurence and Hoppe, Travis and Foster, Charles and Phang, Jason and He, Horace and Thite, Anish and Nabeshima, Noa and Presser, Shawn and Leahy, Connor , title =. arXiv preprint arXiv:2101.00027 , year =
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.