Pith. sign in

REVIEW 4 major objections 6 minor 75 references

On Topology's Role in ML Training Performance

T0 review · 4 major / 6 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read The Clos topology outperforms the torus for most ML training collectives; AllGather is the exception.

desk verdict Useful analytic extension of collective bounds to failures and placement; the 'Clos mostly wins' claim is honest under the model's assumptions, but those assumptions are untested and the case-study table has numbers that don't obviously follow from their own formulas. read the letter →

arxiv 2608.01707 v1 pith:4JDC76ZF submitted 2026-08-03 cs.NI

classification cs.NI
keywords collectivecommunicationAllReduceGatherAlltoClostopologytorusMLtrainingnetworkslinkfailures
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Distributed ML training runs on networks that are almost always either a fat-tree Clos or a torus, and this paper asks which one actually delivers faster training traffic. It derives closed-form completion-time formulas for the three collective operations that dominate training communication—AllReduce, AllGather, and AlltoAll—under a simple latency-plus-bandwidth model, then extends the comparison to link failures, job placement, cost, multicast, and hybrid scale-up/scale-out designs. The central finding is that no single topology wins everywhere: the torus is better for AllGather and can match the Clos at small node counts, but the Clos wins AllReduce and AlltoAll as scale grows and is markedly less sensitive to failures and placement. The paper's bottom line is that unless AllGather is the bottleneck, the Clos is the better choice. This matters because designers currently pick between the two largely on tradition, and the paper offers analytic grounds for that choice.

What carries the argument

The load-bearing object is the collective completion time (CCT) under the alpha-beta model, in which a message costs per-link latency alpha plus transmission time beta times message size, plus reduction time gamma where reductions happen. The paper derives minimum CCT formulas for AllReduce, AllGather, and AlltoAll on each topology; the formulas expose which term dominates. The key structural difference is per-node access bandwidth: a torus node has 2k_t incident links while a Clos node has one, which makes AllGather cheaper on the torus; meanwhile the Clos's path diversity enables log-depth AllReduce and full-bisection AlltoAll. A secondary device is the bandwidth ratio r_b, the relative pe

What would settle it

Run the same collective schedules on a Clos and a same-size torus with equal per-link bandwidth, at n=512, using a real transport with congestion control, and measure collective completion time for AllReduce, AllGather, and AlltoAll. If AllReduce or AlltoAll completes faster on the torus at scale, the paper's main ordering is wrong; if AllGather completes faster on the Clos at equal bandwidth, its AllGather claim is wrong.

Watch

Extended reading notes

Core claim

The paper's central claim is that for the collective communication patterns that dominate ML training, the Clos and torus trade wins in a predictable, scale-dependent way, and the Clos is the safer default. Using optimal routing and an alpha-beta model (per-link latency plus transmission time, no queueing), it derives minimum collective completion times on each topology. The torus wins AllGather because each node has 2k_t access links rather than one, so its access bandwidth is higher; it also does well at small n because its diameter is low. But AllReduce on a Clos completes in log n rounds via recursive doubling, and AlltoAll benefits from the Clos's full bisection bandwidth, so both colle

Load-bearing premise

The load-bearing premise is that flows can be placed to avoid congestion whenever capacity exists and that queueing delay can be ignored; under real congestion control and bursty traffic, absolute completion times rise and the ordering between topologies could change.

Editorial extensions

If this is right

  • At scale, AllReduce and AlltoAll finish faster on a Clos than on an equal-bandwidth torus, and the gap widens with node count.
  • AllGather is the torus's one reliable win at equal bandwidth, and it comes from per-node access bandwidth, not diameter.
  • With realistic link-bandwidth ratios (torus links around half the Clos's), AllGather stays torus-favored but AllReduce and AlltoAll favor the Clos at scale.
  • A single link failure costs the Clos little for AllGather and AlltoAll but hits the torus harder; only for AllReduce is the torus more failure-tolerant.
  • Hybrid scale-up/scale-out designs (torus pods behind a Clos) can beat both single topologies for AllReduce and AllGather, though AlltoAll gains little.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the model assumes congestion-free routing, real-world queueing is likely to raise absolute completion times on both topologies; whether it preserves the ordering is an open question, and the Clos's path diversity may make its relative advantage larger under congestion.
  • The r_b analysis suggests a practical rule of thumb the paper does not spell out: a torus needs several times the Clos's per-link bandwidth to make AllReduce or AlltoAll competitive, so a designer can decide on link technology and scale from the crossover lines rather than running a full simulation.
  • The placement results imply that torus-based job schedulers should co-optimize parallelism dimensions with physical placement, while Clos-based schedulers can ignore placement; a random torus placement can cost up to a factor k_t in latency and bandwidth.
  • Extending the single-failure analysis to correlated or multiple simultaneous failures would likely widen the Clos's resilience advantage, since the Clos has many edge-disjoint paths per flow while the torus's detours share scarce links.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper develops an analytical alpha-beta (Hockney) model for the completion time of three collective operations—AllReduce, AllGather, and AlltoAll—on Clos and torus topologies, and uses the resulting formulas to compare the topologies across scale, message size, link failures, placement flexibility, cost, hybrid designs, multicast support, and two modern LLM training case studies. The main qualitative conclusions are that the torus generally wins for AllGather because of higher per-node access bandwidth, while the Clos wins for AllReduce and AlltoAll at larger scales, and that the Clos is less sensitive to placement and failures. The paper is explicit that the model assumes store-and-forward switches, optimal routing, and no queuing, and it marks several formulas as extensions of prior work rather than full derivations.

Significance. If the derivations are correct, the paper provides a useful reference framework for comparing two dominant ML-training interconnects, and it extends the comparison to failure, placement, cost, and hybrid settings that are often treated only informally. The paper is honest about its modeling assumptions and includes case studies grounded in real model parameters. However, the quantitative completion times are not validated against packet-level simulation or measurements, and several load-bearing formulas are cited rather than derived, so the strength of the practical recommendation exceeds what the current evidence strictly supports.

major comments (4)
  1. [§2, Performance and footnote 6; Table 2] The congestion-free alpha-beta model is load-bearing for the quantitative CCT values and for the final recommendation. The paper acknowledges that queuing is ignored, but the body uses absolute times (Figures 1, 4, 6, 11 and Table 8) to conclude that 'unless AllGather is your bottleneck, the Clos topology is likely the better choice.' Real congestion control, ECMP hash collisions, and bursty collective traffic can add superlinear delay and may reorder the topologies at the n=64–256 scales highlighted in the paper. Please either validate the ordering with packet-level simulation for representative collectives, or explicitly reframe the claims as upper bounds under optimal congestion-free routing and temper the concluding recommendation accordingly.
  2. [§3, Table 2] Several central entries are marked with a dagger and are not derived in the paper: the torus AllGather formula, the torus AlltoAll formula, and the Clos AllGather formula. Appendix A only covers failure derivations, not the baseline collectives. These formulas are reused throughout Sections 4–9, so a reader cannot verify the claimed optimality or the constants (e.g., the 1/(2k_t) and 1/8 factors). Please provide complete derivations in an appendix, or exact pointers to theorems/lemmas in the cited works [12, 52, 66, 70].
  3. [§4, Table 3] The torus entries marked with a double dagger are lower bounds rather than known achievable schedules, but the text describes them as the 'fundamental, best-case impact' and Figure 6 plots them as 'increase in CCT' alongside exact Clos values. This conflates a lower bound with an exact additional time. Since the torus AlltoAll and AllGather comparisons in the failure section depend on these entries, the lower-bound status should be stated prominently in the body and figure, and the conclusions should be phrased as conservative comparisons.
  4. [§5, Table 4] The locality-optimized AllGather bandwidth term is derived as a lower bound (total messages divided by four incoming links), but the text presents the resulting expression as a completion time 'upper bound.' The distinction matters because the locality-optimized placement is then recommended on the basis of these formulas. Please clarify whether the AllGather cost is an achievable schedule or a bound, and state the same for the AlltoAll cost in that table.
minor comments (6)
  1. [Abstract] Typo: 'the elucidate' should be 'that elucidate'.
  2. [§2, Topologies] Typo: 'we a traditional 3-tier k_c-ary fat tree' should be 'we use a traditional 3-tier k_c-ary fat tree'.
  3. [Table 2] Notation is ambiguous: 'log2(n^{1/k})' is missing the subscript t on k, and 'β_t m n−1/2k_t' should be written as β_t m (n−1)/(2k_t). Similar parenthesization issues appear in the AlltoAll row.
  4. [§4, Network-Level Mitigation] Typo: 'we donotfix k_c' should be 'we do not fix k_c'.
  5. [Appendix B.2.1] The sentence 'the diameter of each group may be up to 2√N=2√N' is a self-evident typo; presumably the second expression should be 2 N^{1/4} or similar.
  6. [§9, Takeaway and Table 8] The takeaway 'A torus-Clos hybrid outperforms in all cases' is contradicted by Table 8: DeepSeek v3 AllGather is 114 ms on the hybrid versus 36 ms on both Torus and Clos. The sentence immediately after acknowledges the exception, but the takeaway should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: Table 2 formulas extend externally cited collective-communication results and are not fitted to the conclusions.

full rationale

The paper's central contribution is an analytical comparison of collective completion time (CCT) formulas for Clos and torus under the alpha-beta model. The formulas in Table 2 are either taken from external prior work (e.g., [12], [66], [70], [4]) or are flagged as extensions (†), and the paper states in Section 3 that it provides descriptions of algorithms and bounds rather than full derivations. Nothing is fitted to a target outcome: the predicted ordering (torus wins AllGather; Clos wins AllReduce/AlltoAll at scale) follows algebraically from the stated per-link bandwidth, latency, and access-link counts, not from a parameter fitted to that ordering. The failure (Section 4), placement (Section 5), hybrid (Section 7), and multicast (Section 8) results are derived by composing or extending the Table 2 expressions; lower-bound entries are explicitly marked (‡) and derived in Appendix A rather than assumed. The case studies (Section 9) plug externally sourced model parameters into the same formulas. The self-citations ([13], [63], both involving author Minlan Yu) are present but not load-bearing: [63] is one of two sources for a default message size in Table 1, and [13] is a related-work pointer for placement; neither supplies a uniqueness theorem or the scaling conclusions. The paper's explicit limitations—ignoring queuing delay and assuming congestion-free flow placement (footnote 6), and omitting full derivations in Section 3—are modeling caveats, not circular reductions. Therefore no step in the derivation chain is equivalent to its own inputs by construction.

Assumptions & free parameters 7 free parameters · 6 assumptions · 0 invented entities

All parameters are hand-chosen operating points from cited hardware references, not fitted to the model. The main modeling axioms are the queueless alpha-beta model, optimal congestion-free routing, store-and-forward switching, and optimality of known collective schedules. No new physical entities are introduced.

free parameters (7)
  • per-link latency alpha = 1 us
    Hand-picked default from Table 1, based on typical link technology; not fitted.
  • link bandwidth beta = 400 Gbps
    Default from Table 1, based on modern switches [8,20]; not fitted.
  • reduction time per byte gamma = 1 ps
    Default from Table 1 [45]; not fitted.
  • message size m = 100 MB
    Default from Table 1 [38,63]; not fitted.
  • torus dimensionality k_t = 3
    Default from Table 1 [33]; not fitted.
  • Clos switch radix k_c = 128
    Default from Table 1 [8,33]; not fitted.
  • TPU/GPU bandwidth ratio r_b = 0.5
    Derived from Figure 3 recent per-link bandwidths [24,26,27,46,49,51]; used as an operating point in Section 3, not fitted to the model.
assumptions (6)
  • domain assumption Alpha-beta (Hockney) model with no queuing delay
    Section 2 Performance: 'This simplified model ignores queuing delay, allowing us to evaluate the impact of topology without having to model fine-grained transport layer choices.' This is the core modeling assumption.
  • domain assumption Optimal routing and congestion-free flow placement
    Section 3 footnote 6: 'We assume that flows are placed to avoid congestion as long as there is sufficient capacity.' The bounds are best-case.
  • domain assumption Store-and-forward switches
    Section 3: 'These results were derived assuming store-and-forward switches.'
  • domain assumption Known collective algorithms are optimal
    Section 3 builds on prior work [12,52,66,70] for the best-performing known algorithms; some entries marked with a dagger are extensions but not fully derived.
  • domain assumption Single non-partitioning link failure, no downlink failures
    Section 4: 'we model link failures as single, bidirectional, failures for the entire duration of a given collective' and exclude downlink failures that would partition the network.
  • domain assumption Best-possible locality for Clos subsets and n=N for failures
    Sections 2 and 4: 'When taking a subset, we assume the best-possible locality' and 'moving forward, we assume n=N for both topologies'.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On Topology's Role in ML Training Performance." pith.science (2026). https://pith.science/paper/4JDC76ZF

@misc{pith2026260801707,
  author       = {Pith},
  title        = {Pith review of: On Topology's Role in ML Training Performance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4JDC76ZF}},
  note         = {Machine review of arXiv:2608.01707}
}
read the original abstract

Modern machine learning training workloads run on large-scale networks of compute accelerators. The networks commonly deployed in these systems are typically variations of two basic topologies: the fat-tree Clos and the torus. In this paper, we derive analytical results the elucidate how the choice of topology shapes achievable performance for the small set of collective communication operations that underlies modern machine learning workloads. We also consider how these results change when we include additional factors such as network failures and job placement strategies. Overall, we find that one topology does not dominate in all cases, but that the Clos achieves better collective completion time in most cases and provides benefits in resilience and flexibility.

Figures

Figures reproduced from arXiv: 2608.01707 by the authors.

Figure 1
Figure 1. Completion times (using default values in [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 4
Figure 4. Completion times for each collective on a 128 node network as the message size is varied. 10 100 1000 0 200 400 600 Number of Nodes (n) CCT (us) [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 7
Figure 7. Torus link failure alternate routes example for a single-hop path and a multi-hop path. Alternate paths of equal or best-possible length are shown (dashed). would cause a network partition and a collective could not complete.10 Importantly, here we do not fix 𝑘𝑐 for the Clos and instead scale it to the minimum size necessary to fit 𝑁 (i.e., 𝑛=𝑁). This ensures there is not a large portion of the network (and bandwidt… view at source ↗
Figures from the paper (5 more)
Figure 7
Figure 7. Figure 7 [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 9
Figure 9. Figure 9: Placement of two parallelisms on a 2-D, 9×9 torus. One parallelism is nodes of the same color, the other is nodes that are boxed together. groups which tile the torus with √ 𝑁 smaller 2-D grids15, as in [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 10
Figure 10. Figure 10: Switch cost comparisons of the two topologies. [PITH_FULL_IMAGE:figures/full_fig_p009_10.png]
Figure 11
Figure 11. Figure 11: Comparison of Clos + Torus hybrid topology with [PITH_FULL_IMAGE:figures/full_fig_p010_11.png]
Figure 12
Figure 12. Figure 12: Locality-Opt torus placement for the Megatron-LM [PITH_FULL_IMAGE:figures/full_fig_p012_12.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

75 extracted references · 58 canonical work pages

  1. [1]

    Mohammad Al-Fares, Alexander Loukissas, and Amin Vahdat. 2008. A scalable, commodity data center network architecture. InProceedings of the ACM SIGCOMM 2008 Conference on Data Communication (SIGCOMM ’08). Association for Computing Machinery, New York, NY , USA, 63–74. https://doi.org/10.1145/1402958.1402967

  2. [2]

    Carl Albing, Norm Troullier, Stephen Whalen, Ryan Olson, Joe Glenski, Howard Pritchard, and Hugo Mills. 2011. Scalable node allocation for improved performance in regular and anisotropic 3D torus supercomputers. InProceedings of the 18th European MPI Users’ Group Conference on Recent Advances in the Message Passing Interface (EuroMPI’11). Springer-Verlag,...

  3. [3]

    Jacob Austin, Sholto Douglas, Roy Frostig, Anselm Levskaya, Charlie Chen, Sharad Vikram, Federico Lebron, Peter Choy, Vinay Ramasesh, Albert Webson, and Reiner Pope. 2025. How to Scale Your Model. Online. (2025). Retrieved from https://jax-ml.github.io/scaling-book/

  4. [4]

    Prithwish Basu, Liangyu Zhao, Jason Fantl, Siddharth Pal, Arvind Krish- namurthy, and Joud Khoury. 2024. Efficient all-to-all Collective Commu- nication Schedules for Direct-connect Topologies. InProceedings of the 33rd International Symposium on High-Performance Parallel and Dis- tributed Computing (HPDC ’24). Association for Computing Machinery, New Yor...

  5. [5]

    Borkar, R

    S. Borkar, R. Cohn, G. Cox, S. Gleason, T. Gross, H.T. Kung, M. Lam, B. Moore, C. Peterson, J. Pieper, L. Rankin, P.S. Tseng, J. Sutton, J. Urbanski, and J. Webb. 1988. iWarp: an integrated solution to high-speed parallel computing. InSupercomputing ’88:Proceedings of the 1988 ACM/IEEE Conference on Supercomputing, Vol. I (SC ’88). 330–339. https://doi.or...

  6. [6]

    Nuclear multipole responses from chiral effective field theory interaction

    G. Bosilca, A. Bouteiller, F. Cappello, S. Djilali, G. Fedak, C. Germain, T. Herault, P. Lemarinier, O. Lodygensky, F. Magniette, V . Neri, and A. Selikhov. 2002. MPICH-V: Toward a Scalable Fault Tolerant MPI for V olatile Nodes. InSC ’02: Proceedings of the 2002 ACM/IEEE Conference on Supercomputing. 29–29. https://doi.org/10.1109/SC.2002.10048

  7. [7]

    Pat Bosshart, Glen Gibb, Hun-Seok Kim, George Varghese, Nick McKeown, Martin Izzard, Fernando Mujica, and Mark Horowitz

  8. [8]

    Broadcom. 2026. Tomahawk 6 / BCM78910 Series. (2026). https://www.broadcom.com/products/ethernet-connectivity/switch ing/strataxgs/bcm78910-series

Show all 75 references
  1. [9]

    Jehoshua Bruck, Ching-Tien Ho, Shlomo Kipnis, and Derrick Weath- ersby. 1994. Efficient algorithms for all-to-all communications in multi-port message-passing systems. InProceedings of the Sixth Annual ACM Symposium on Parallel Algorithms and Architectures (SPAA ’94). Associat...

  2. [10]

    Zixian Cai, Zhengyang Liu, Saeed Maleki, Madanlal Musuvathi, Todd Mytkowicz, Jacob Nelson, and Olli Saarikivi. 2021. Synthesizing optimal collective algorithms. InProceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP ’21). Asso...

  3. [11]

    Jiamin Cao, Shangfeng Shi, Jiaqi Gao, Weisen Liu, Yifan Yang, Yichi Xu, Zhilong Zheng, Yu Guan, Kun Qian, Ying Liu, Mingwei Xu, Tianshu Wang, Ning Wang, Jianbo Dong, Binzhang Fu, Dennis Cai, and Ennan Zhai. 2025. SyCCL: Exploiting Symmetry for Efficient Collective Communicatio...

  4. [12]

    Ernie Chan, Marcel Heimlich, Avi Purkayastha, and Robert van de Geijn

  5. [13]

    Shawn Shuoshuo Chen, Daiyaan Arfeen, Minlan Yu, Peter Steenkiste, and Srinivasan Seshan. 2025. Toward Co-adapting Machine Learning Job Shape and Cluster Topology. (2025). arXiv:cs.DC/2510.03891 https://arxiv.org/abs/2510.03891

  6. [14]

    Charles Clos. 1953. A study of non-blocking switching net- works.The Bell System Technical Journal32, 2 (1953), 406–424. https://doi.org/10.1002/j.1538-7305.1953.tb01433.x

  7. [15]

    W.J. Dally. 1990. Performance analysis of k-ary n-cube intercon- nection networks.IEEE Trans. Comput.39, 6 (1990), 775–785. https://doi.org/10.1109/12.53599

  8. [16]

    Dally and C.L

    W.J. Dally and C.L. Seitz. 1986. The Torus Routing Chip. InDistributed Computing, V ol. 1. 187–196

  9. [17]

    Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

    DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...

  10. [18]

    Fagg and Jack J

    Graham E. Fagg and Jack J. Dongarra. 2000. FT-MPI: Fault Tolerant MPI, Supporting Dynamic Applications in a Dynamic World. InRecent Advances in Parallel Virtual Machine and Message Passing Interface, Jack Dongarra, Peter Kacsuk, and Norbert Podhorszki (Eds.). Springer Berlin H...

  11. [19]

    Nathan Farrington, George Porter, Sivasankar Radhakrishnan, Hamid Hajabdolali Bazzaz, Vikram Subramanya, Yeshaiahu Fain- man, George Papen, and Amin Vahdat. 2010. Helios: a hybrid electrical/optical switch architecture for modular data centers. In Proceedings of the ACM SIGCOM...

  12. [21]

    Gherghescu, Vlad-Andrei B˘adoiu, Alexandru Agache, Mihai-Valentin Dumitru, Iuliu Vasilescu, Radu Mantu, and Costin Raiciu

    Alexandru M. Gherghescu, Vlad-Andrei B˘adoiu, Alexandru Agache, Mihai-Valentin Dumitru, Iuliu Vasilescu, Radu Mantu, and Costin Raiciu. 2024. I’ve Got 99 Problems But FLOPS Ain’t One. In Proceedings of the 23rd ACM Workshop on Hot Topics in Networks (HotNets ’24). Association ...

  13. [22]

    Phillipa Gill, Navendu Jain, and Nachiappan Nagappan. 2011. Under- standing network failures in data centers: measurement, analysis, and implications. InProceedings of the ACM SIGCOMM 2011 Conference (SIGCOMM ’11). Association for Computing Machinery, New York, NY , USA, 350–3...

  14. [23]

    Ivan Goldwasser, Harry Petty, Pradyumna Desale, and Kirthi Devleker. 2024. NVIDIA GB200 NVL72 Delivers Trillion- Parameter LLM Training and Real-Time Inference. (18 Apr 2024). https://developer.nvidia.com/blog/nvidia-gb200-nvl72-delivers-tri llion-parameter-llm-training-and-re...

  15. [24]

    Google. 2026. TPU7x. (2026). https://docs.cloud.google.com/tpu/ docs/tpu7x

  16. [25]

    Google. 2026. TPUv4. (2026). https://docs.cloud.google.com/tpu/ docs/v4

  17. [26]

    Google. 2026. TPUv5p. (2026). https://docs.cloud.google.com/tpu/ docs/v5p

  18. [27]

    Google. 2026. TPUv6e. (2026). https://docs.cloud.google.com/tpu/ docs/v6e

  19. [28]

    Hamilton, Navendu Jain, Srikanth Kandula, Changhoon Kim, Parantap Lahiri, David A

    Albert Greenberg, James R. Hamilton, Navendu Jain, Srikanth Kandula, Changhoon Kim, Parantap Lahiri, David A. Maltz, Parveen Patel, and Sudipta Sengupta. 2009. VL2: a scalable and flexible data center network. InProceedings of the ACM SIGCOMM 2009 Conference on Data Com- munic...

  20. [29]

    Le, Yonghui Wu, and Zhifeng Chen

    Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V . Le, Yonghui Wu, and Zhifeng Chen. 2019.GPipe: efficient training of giant neural networks using pipeline parallelism. Curran Associates Inc., Red Hook, NY , USA

  21. [30]

    Sam Ade Jacobs, Ammar Ahmad Awan, Martin Cai, Elton Zheng, Zhen Zheng, Rangan Majumder, and Junhua Wang. [n. d.]. Deep- Speed: Extreme-scale Model Training for Everyone. ([n. d.]). https://www.microsoft.com/en-us/research/blog/deepspeed-extreme -scale-model-training-for-everyo...

  22. [31]

    Nikhil Jain and Yogish Sabharwal. 2010. Optimal bucket algorithms for large MPI collectives on torus interconnects. InProceedings of the 24th ACM International Conference on Supercomputing (ICS ’10). Association for Computing Machinery, New York, NY , USA, 27–36. https://doi.o...

  23. [32]

    Nisha Mariam Johnson and Andi Gavrilescu. 2023. How to scale AI training to up to tens of thousands of Cloud TPU chips with Multislice. (31 Aug 2023). https://cloud.google.com/blog/products/compute/u sing-cloud-tpu-multislice-to-scale-ai-workloads

  24. [33]

    Norm Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, Clifford Young, Xiang Zhou, Zongwei Zhou, and David A Patter- son. 2023. TPU v4: An Optically Reconfigurable Supercomputer for Ma- chine L...

  25. [34]

    Kessler and J.L

    R.E. Kessler and J.L. Schwarzmeier. 1993. Cray T3D: a new dimension for Cray Research. InCompcon Spring ’93’. 176–182. https://doi.org/10.1109/CMPCON.1993.289660

  26. [35]

    John Kim. 2009. Low-cost router microarchitecture for on-chip networks. InProceedings of the 42nd Annual IEEE/ACM Inter- national Symposium on Microarchitecture (MICRO 42). Associ- ation for Computing Machinery, New York, NY , USA, 255–266. https://doi.org/10.1145/1669112.1669145

  27. [36]

    Apostolos Kokolis, Michael Kuchnik, John Hoffman, Adithya Kumar, Parth Malani, Faye Ma, Zachary DeVito, Shubho Sengupta, Kalyan Saladi, and Carole-Jean Wu. 2025. Revisiting Reliability in Large-Scale Machine Learning Research Clusters. (2025). arXiv:cs.DC/2410.21680 https://ar...

  28. [37]

    Shenggui Li and Siqi Mai. [n. d.]. Colossal-AI Paradigms of Parallelism. ([n. d.]). https://colossalai.org/docs/concepts/paradigms_of_parallelism/ Accessed: 2026-01-28

  29. [38]

    Wenxue Li, Xiangzhou Liu, Yuxuan Li, Yilun Jin, Han Tian, Zhizhen Zhong, Guyue Liu, Ying Zhang, and Kai Chen. 2024. Understanding Communication Characteristics of Distributed Training. InProceed- ings of the 8th Asia-Pacific Workshop on Networking (APNet ’24). Association for ...

  30. [40]

    Karthik Mandakolathur and Sylvain Jeaugey. 2022. Doubling all2all Performance with NVIDIA Collective Communication Library 2.12. (Feb. 2022). https://developer.nvidia.com/blog/doubling-all2all-per formance-with-nvidia-collective-communication-library-2-12/

  31. [41]

    Mellette, Alex C

    William M. Mellette, Alex C. Snoeren, and George Porter. 2016. P- FatTree: A multi-channel datacenter network topology. InProceedings of the 15th ACM Workshop on Hot Topics in Networks (HotNets ’16). Association for Computing Machinery, New York, NY , USA, 78–84. https://doi.o...

  32. [42]

    Justin Meza, Tianyin Xu, Kaushik Veeraraghavan, and Onur Mutlu

  33. [43]

    Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGres- ley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. 2021. Efficient large-scale language model training on GPU cl...

  34. [44]

    Nickolls

    J.R. Nickolls. 1990. The design of the MasPar MP-1: a cost effective massively parallel computer. InCompcon Spring ’90. Thirty-Fifth IEEE Computer Society International Conference on Intellectual Leverage. 25–28. https://doi.org/10.1109/CMPCON.1990.63649

  35. [45]

    NVIDIA. [n. d.]. NVIDIA A100 Tensor Core GPU. ([n. d.]). https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/ a100/pdf/nvidia-a100-datasheet-us-nvidia-1758950-r4-web.pdf

  36. [46]

    NVIDIA. 2022. NVIDIA DGX A100. (2022). https://images.nvidia. com/aem-dam/Solutions/Data-Center/nvidia-dgx-a100-datasheet.pdf

  37. [47]

    NVIDIA. 2025. Megatron-LM. https://github.com/NVIDIA/Megatron- LM. (2025)

  38. [48]

    NVIDIA. 2026. NVLink Switch Chip NVIDIA NVLink and NVLink Switch. (2026). https://www.nvidia.com/en-us/data-center/nvlink/

  39. [49]

    NVIDIA. 2026. Introduction to NVIDIA DGX H100/H200 Systems. (2026). https://docs.nvidia.com/dgx/dgxh100-user-guide/introduct ion-to-dgxh100.html

  40. [50]

    NVIDIA. 2026. NVIDIA Collective Communications Library (NCCL). (2026). https://developer.nvidia.com/nccl

  41. [51]

    NVIDIA. 2026. NVIDIA DGX Rubin NVL8. (2026). https://www.nvidia.com/en-us/data-center/dgx-rubin-nvl8/

  42. [52]

    Pjesivac-Grbovic, T

    J. Pjesivac-Grbovic, T. Angskun, G. Bosilca, G.E. Fagg, E. Gabriel, and J.J. Dongarra. 2005. Performance analysis of MPI collective operations. In19th IEEE International Parallel and Distributed Processing Symposium. 8 pp.–. https://doi.org/10.1109/IPDPS.2005.335

  43. [53]

    Lucian Popa, Sylvia Ratnasamy, Gianluca Iannaccone, Arvind Krishna- murthy, and Ion Stoica. 2010. A cost comparison of datacenter network architectures. InProceedings of the 6th International COnference (Co-NEXT ’10). Association for Computing Machinery, New York, NY , USA, Ar...

  44. [54]

    Leon Poutievski, Omid Mashayekhi, Joon Ong, Arjun Singh, Mukarram Tariq, Rui Wang, Jianan Zhang, Virginia Beauregard, Patrick Conner, Steve Gribble, Rishi Kapoor, Stephen Kratzer, Nanfang Li, Hong Liu, Karthik Nagaraj, Jason Ornstein, Samir Sawhney, Ryohei Urata, Lorenzo Vicis...

  45. [55]

    Kun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao, Yichi Xu, Yu Guan, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao, Chao Wang, Peng Wang, Pengcheng Zhang, Xianlong Zeng, Eddie Ruan, Zhiping Yao, Ennan Zhai, and Dennis Cai. 2024. Alibaba HPN: A Data Center Network for Large Languag...

  46. [56]

    Le Qin, Junwei Cui, Weilin Cai, Meng Niu, Yan Yang, and Jiayi Huang. 2025. Optimizing All-to-All Collective Communication with Fault Tolerance on Torus Networks. InProceedings of the 58th IEEE/ACM International Symposium on Microarchitecture (MICRO ’25). Association for Comput...

  47. [57]

    Martin Ruefenacht, Mark Bull, and Stephen Booth. 2016. Gen- eralisation of Recursive Doubling for AllReduce. InProceedings of the 23rd European MPI Users’ Group Meeting (EuroMPI ’16). Association for Computing Machinery, New York, NY , USA, 23–31. https://doi.org/10.1145/29668...

  48. [58]

    Paul Sack and William Gropp. 2012. Faster topology-aware collective algorithms through non-minimal communication.SIGPLAN Not.47, 8 (Feb. 2012), 45–54. https://doi.org/10.1145/2370036.2145823

  49. [59]

    Sriram Sankaran, Jeffrey M Squyres, Brian Barrett, Vishal Sahay, Andrew Lumsdaine, Jason Duell, Paul Hargrove, and Eric Roman

  50. [60]

    Aashaka Shah, Vijay Chidambaram, Meghan Cowan, Saeed Maleki, Madan Musuvathi, Todd Mytkowicz, Jacob Nelson, Olli Saarikivi, and Rachee Singh. 2023. TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches. In20th USENIX Symposium on Networked Systems Design a...

  51. [61]

    Shipman, Richard L

    Galen M. Shipman, Richard L. Graham, and George Bosilca. 2007. Network Fault Tolerance in Open MPI. InEuro-Par 2007 Parallel Processing, Anne-Marie Kermarrec, Luc Bougé, and Thierry Priol (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 868–878

  52. [62]

    Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. arXiv preprint arXiv:1909.08053(2019)

  53. [63]

    Min Si, Pavan Balaji, Yongzhou Chen, Ching-Hsiang Chu, Adi Gangidi, Saif Hasan, Subodh Iyengar, Dan Johnson, Bingzhe Liu, Regina Ren, Deep Shah, Ashmitha Jeevaraj Shetty, Greg Steinbrecher, Yulun Wang, Bruce Wu, Xinfeng Xie, Jingyi Yang, Mingran Yang, Kenny Yu, Minlan Yu, Cen ...

  54. [64]

    Arjun Singh, Joon Ong, Amit Agarwal, Glen Anderson, Ashby Armistead, Roy Bannon, Seb Boving, Gaurav Desai, Bob Felderman, Paulie Germano, Anand Kanagala, Jeff Provost, Jason Simmons, Eiichi Tanda, Jim Wanderer, Urs Hölzle, Stephen Stuart, and Amin Vahdat

  55. [65]

    Georg Stellner. 1996. CoCheck: Checkpointing and process migration for MPI. InProceedings of International Conference on Parallel Processing. IEEE, 526–531

  56. [66]

    Rajeev Thakur and William D. Gropp. 2003. Improving the Performance of Collective Operations in MPICH. InRecent Advances in Parallel Virtual Machine and Message Passing Interface, Jack Dongarra, Domenico Laforenza, and Salvatore Orlando (Eds.). Springer Berlin Heidelberg, Berl...

  57. [67]

    Weiyang Wang, Manya Ghobadi, Kayvon Shakeri, Ying Zhang, and Naader Hasani. 2024. Rail-only: A Low-Cost High- Performance Network for Training LLMs with Trillion Parameters . In2024 IEEE Symposium on High-Performance Interconnects (HOTI). IEEE Computer Society, Los Alamitos, C...

  58. [68]

    Weiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi, Zhihao Jia, Dheevatsa Mudigere, Ying Zhang, and Anthony Ke- witsch. 2023. TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training Jobs. In20th USENIX Symposium on Networked System...

  59. [69]

    Udayanga Wickramasinghe and Andrew Lumsdaine. 2016. A Survey of Methods for Collective Communication Optimization and Tuning.CoRRabs/1611.06334 (2016). arXiv:1611.06334 http://arxiv.org/abs/1611.06334

  60. [70]

    Eitan Zahavi. 2011. Fat-Trees Routing and Node Ordering Providing Contention Free Traffic for MPI Global Collectives. In2011 IEEE Inter- national Symposium on Parallel and Distributed Processing Workshops and Phd Forum. 761–770. https://doi.org/10.1109/IPDPS.2011.219

  61. [71]

    Liangyu Zhao, Saeed Maleki, Ziyue Yang, Hossein Pourreza, and Arvind Krishnamurthy. 2025. ForestColl: Throughput-Optimal Collective Communications on Heterogeneous Network Fabrics. (2025). arXiv:cs.NI/2402.06787 https://arxiv.org/abs/2402.06787

  62. [72]

    schedule

    Yazhou Zu, Alireza Ghaffarkhah, Hoang-Vu Dang, Brian Towles, Steven Hand, Safeen Huda, Adekunle Bello, Alexander Kolbasov, Arash Rezaei, Dayou Du, Steve Lacy, Hang Wang, Aaron Wisner, Chris Lewis, and Henri Bahini. 2024. Resiliency at Scale: Man- aging Google’s TPUv4 Machine L...

  63. [2005]

    The LAM/MPI checkpoint/restart framework: System-initiated checkpointing.The International Journal of High Performance Computing Applications19, 4 (2005), 479–493

  64. [2007]

    Comput.: Pract

    Collective communication: theory, practice, and experience: Research Articles.Concurr. Comput.: Pract. Exper.19, 13 (Sept. 2007), 1749–1783

  65. [2013]

    Forwarding metamorphosis: fast programmable match-action processing in hardware for SDN.SIGCOMM Comput. Commun. Rev. 43, 4 (Aug. 2013), 99–110. https://doi.org/10.1145/2534169.2486011

  66. [2015]

    InProceedings of the 2015 ACM Conference on Special Interest Group on Data Communication (SIGCOMM ’15)

    Jupiter Rising: A Decade of Clos Topologies and Centralized Control in Google’s Datacenter Network. InProceedings of the 2015 ACM Conference on Special Interest Group on Data Communication (SIGCOMM ’15). Association for Computing Machinery, New York, NY , USA, 183–197. https:/...

  67. [2018]

    In Proceedings of the Internet Measurement Conference 2018 (IMC ’18)

    A Large Scale Study of Data Center Network Reliability. In Proceedings of the Internet Measurement Conference 2018 (IMC ’18). Association for Computing Machinery, New York, NY , USA, 393–407. https://doi.org/10.1145/3278532.3278566

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.