Pith. sign in

REVIEW 4 major objections 7 minor 83 references

Photonic Rails in ML Datacenters

T0 review · 4 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper claims that replacing electrical rail switches with optical circuit switches, reconfigured between parallelism phases inside a training job, can cut networking cost by 70% and power by 96% while adding only 3% to iteration time.

desk verdict Fresh idea with a real feasibility trace; the headline numbers overreach, but the design is worth a serious look. read the letter →

arxiv 2507.08119 v1 pith:VK5BK3YH submitted 2025-07-10 cs.NI

classification cs.NI
keywords rail-optimizedfabricopticalcircuitswitchMLtrainingnetworkparallelism-awarereconfigurationhybridparallelismdatacentercollectivecommunicationOpuscontrolplane
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the rail-optimized fabrics used to train large language models are overbuilt: they rely on high-radix electrical packet switches to give every GPU rank full all-to-all connectivity, which costs enormous power and money. The authors propose to replace those electrical rails with photonic optical circuit switches that form one-to-one optical paths, and to reconfigure those paths during the job, between the communication phases of different parallelism axes (tensor, data, pipeline). They build a control plane called Opus that intercepts collective communication calls, predicts the next parallelism phase, and provisions the optical switches during the idle windows that naturally occur between phases. Early experiments with a Llama3-8B workload show that more than 75% of these windows are over one millisecond, and simulation indicates that reconfiguration can be hidden well enough to keep iteration time within a few percent of an electrical rail baseline while cutting infrastructure cost by about 70% and power by about 96%.

What carries the argument

Parallelism-driven rail reconfiguration: the observation that communication from different parallelism dimensions (e.g., pipeline Send/Recv versus data-parallel AllGather/ReduceScatter) occurs in sequential phases separated by measurable idle windows, defined by the paper as the gap between the end of one communication group's traffic and the start of the next. Opus uses this window to hide circuit reconfiguration, and its provisioning mechanism speculatively issues reconfiguration requests immediately after the preceding collective finishes, based on a profiled schedule, so that the optical circuits are ready before the next collective begins. The window count per training iteration is estimated by a formula (Eq. 1) that counts phase transitions across PP, FSDP/DP, CP/EP, and microbatches.

What would settle it

Run a production-scale LLM training job (e.g., Llama3-405B with TP, PP, and FSDP) on a real OCS-based rail fabric with Opus-style reconfiguration, and measure the per-iteration wall-clock time and the distribution of idle windows between parallelism phases. If the median window between a data-parallel and a pipeline phase shrinks below the OCS reconfiguration time (e.g., below 25 ms for MEMS switches) or becomes unpredictable, the reconfiguration cannot be hidden and iteration time will grow substantially beyond the claimed 3%.

Watch

Extended reading notes

Core claim

The central claim is that the rail abstraction—the illusion of a dedicated, congestion-free all-to-all network connecting GPUs of the same rank across scale-up domains—can be preserved without electrical packet switching. Because ML collectives are issued in a strict, predictable order dictated by the model's execution graph, the network can be time-multiplexed: an optical circuit switch is reconfigured only when the parallelism axis changes, and the reconfiguration delay is hidden inside the idle window between the end of one collective and the start of the next. The paper introduces Opus, a control layer that sits between the application and the collective communication library, profiles traffic demands, and issues reconfiguration requests to a controller that programs the optical switches. This turns network connectivity into an allocatable resource that co-evolves with the job's parallelism phases, rather than a static topology configured before the job starts.

Load-bearing premise

The whole scheme depends on real training jobs having idle gaps between parallelism phases that are long enough and reliably placed to hide the time it takes to reconfigure an optical switch, and that these gaps stay large as jobs scale to thousands of GPUs.

Editorial extensions

If this is right

  • If the claim holds, ML datacenters could replace multi-tier electrical Clos networks with flat, single-hop optical rails, eliminating switch ASICs and OEO conversions from the datapath.
  • The cost and power savings (70% and 96%) would let the same budget support significantly larger training clusters, or cut operational costs for existing ones.
  • The paper's control-plane design suggests a new software abstraction where applications expose circuit connectivity as a schedulable resource, enabling co-scheduling of compute, communication, and reconfiguration.
  • Because Opus only reconfigures on parallelism shifts, the number of reconfigurations per iteration is small (estimated at 127 for a Llama3.1-405B iteration), making millisecond-scale OCS technology sufficient.
  • The design is compatible with existing rail topologies and GPU-NIC configurations, so adoption would not require changing cabling or the GPU-to-rail mapping.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The window-based feasibility argument likely extends to other hybrid-parallel schemes, but its validity depends on the assumption that the measured millisecond-scale idle windows persist for larger models, deeper pipelines, and higher-bandwidth NICs; a natural test is to instrument a production-scale job (e.g., Llama3-405B) and measure window distributions.
  • The paper's emphasis on per-communication-group reconfiguration granularity suggests that the approach could be extended to expert-parallel AllToAll traffic by configuring circuits to prioritize bottleneck flows, though the paper itself notes that AllToAll does not map cleanly onto a ring.
  • A co-designed scheduling interface that lets the training framework delay or reorder collectives could reduce the need for spacing back-to-back traffic, potentially eliminating the GPU bubbles the paper acknowledges.
  • The cost/power numbers assume equal bandwidth between optical and electrical rails; if optical links deliver higher bandwidth, as the paper hints, the iteration overhead could become negative, meaning the optical rail could be faster than the electrical baseline.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. This paper proposes replacing the electrical packet switches in rail-optimized ML training fabrics with optical circuit switches, and reconfiguring the OCSes between parallelism phases within a training iteration to emulate the all-to-all rail abstraction. The authors present a control plane called Opus, measure communication idle windows in a small Llama3-8B TorchTitan workload, extrapolate window counts to larger jobs, report a trace-driven simulation of iteration-time overhead, and use a vendor-price model to claim 70.5% cost and 95.84% power savings.

Significance. If the results hold, the paper opens a credible new direction: reconfiguring the network at the granularity of ML collectives rather than per job, potentially eliminating switch ASICs from the data path. The strengths are the clean problem framing, the concrete Opus design, the use of a real (if small) workload trace, and the falsifiable cost/power model. However, the headline performance-comparability claim rests on one favorable workload and an unvalidated extrapolation, so the current evidence supports a position paper more than a definitive systems claim.

major comments (4)
  1. [§3.1 and Fig. 4] The paper's central performance claim (3% iteration-time increase) is not adequately supported by the reported window statistics. The text states only that more than 75% of windows exceed 1 ms, while the OCS technologies in Table 3 (3D MEMS 15 ms, Polatis 25 ms, liquid crystal 100 ms) have reconfiguration times of tens of milliseconds. Please report the fraction of windows exceeding each of these thresholds, and describe how the §4.2 simulator handles a window that is shorter than the reconfiguration time. Without this, the reader cannot tell whether the 3.5% overhead at 100 ms in Fig. 8 is robust or is an artifact of a favorable trace with unusually long windows.
  2. [§3.1, Eq. (1)] The extrapolation to large jobs is a count heuristic, not a size predictor. Equation (1) yields 127 windows per 20 s iteration for Llama3.1-405B, but the window-size distribution is measured only for a 4-node A100 Llama3-8B run with TP=4, FSDP=2, PP=2. Window sizes in larger jobs, different pipeline schedules, or with more aggressive overlap could be much smaller, and the paper provides no model for how window durations scale. Please validate the window-size distribution on at least one larger configuration, or provide a principled model of window durations (not just counts) before claiming that reconfiguration can be hidden.
  3. [§5, Opportunities] The paper concedes that "back-to-back or partially-overlapped traffic from two parallelism axes need to be spaced, resulting in bubbles and GPU idling." This admission is in tension with the headline 3% overhead, because the §4.2 simulation does not appear to model the insertion of such spacing/bubbles. The simulation should be extended to include the scheduling constraints described in §5 and report iteration time for workloads whose parallelism phases overlap. As written, the performance parity claim is demonstrated for one favorable workload only.
  4. [§4.2 and Fig. 7] The 70.5% cost and 95.84% power savings are computed with a vendor-list-price model that explicitly excludes fiber cable cost and power (Fig. 7 caption). For an optical replacement of electrical rails, the fiber plant and transceiver count are first-order cost and power components, and list prices can differ substantially from negotiated datacenter prices. Please quantify the excluded components and provide a sensitivity analysis over transceiver/OCS price assumptions, and state whether the savings hold at the cluster sizes shown (the plot begins at 1024 GPUs).
minor comments (7)
  1. [Abstract and §4.2] The Abstract states a 3% iteration-time increase, while §4.2/Fig. 8 reports 3.5% at 100 ms with provisioning; please reconcile these numbers.
  2. [Eq. (1)] The symbols n_layer, n_microbatch, and PP are used in Eq. (1) but not defined before first use; please define them in the text or in a notation table.
  3. [Fig. 3] The arrows that denote windows in Fig. 3 are not explained in the caption; please state explicitly what the arrows mark and how their lengths relate to T_window.
  4. [§4.2 and Fig. 7] The Fig. 7 caption states that fiber cable cost and power are excluded, but the main-text cost analysis in §4.2 does not; move this caveat into the main text and discuss its implications for the savings figures.
  5. [§4.2 and Fig. 8] The simulator behind Fig. 8 is described in a single sentence; please provide a precise model (e.g., how circuit setup is serialized across rails, how FC-FS is enforced, whether reconfiguration of different rails proceeds in parallel) so the result is reproducible.
  6. [§5] The text reads "Course-grained reconfiguration" and should read "Coarse-grained reconfiguration."
  7. [General] Please state whether the workload trace and simulation code will be released; this would significantly improve reproducibility of the central claims.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central claims are supported by measured traces, trace-driven simulation, and vendor-based cost modeling, not by definitional identity or fitted inputs.

full rationale

The paper's quantitative claims do not reduce to their own inputs. The cost and power savings (70% and 96%) come from a component-level comparison against fat-tree and rail-optimized baselines using vendor datasheets and the methodology of prior work [71,72], which is external, not self-referential. The iteration-time result comes from a trace-driven simulation of a measured Llama3-8B workload in which reconfiguration latency is added and the resulting normalized iteration time is reported in Fig. 8; it is an experimental result, not a quantity forced by construction. The window definition in Section 3.1 is descriptive: windows are measured idle gaps, and the paper reports their empirical CDF (more than 75% over 1ms) rather than assuming that every reconfiguration is hidden. The 405B extrapolation via Eq. 1 is a schedule-based count heuristic, not a fitted parameter designed to produce a particular overhead, and the paper explicitly admits in Section 5 that back-to-back or partially overlapped traffic must be spaced, creating bubbles and GPU idling. The self-citations ([28], [29], [60]) are used for context and prior limitations, not as load-bearing justification for Opus's savings, and no uniqueness theorem or ansatz is imported from them. Any concern about whether idle windows remain large enough at scale is a validation or robustness question, not a circularity in the derivation chain.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central claims rest on four domain assumptions: schedulable idle windows, equal bandwidth between optical and electrical rails, stable and provisionable communication schedules, and the representativeness of the vendor-based cost model. Equation 1, used to extrapolate window counts to large jobs, is an unvalidated heuristic. No numeric parameters are fitted to data, and no new physical entities are proposed; Opus is a software control-plane design, not an invented physical entity.

assumptions (5)
  • domain assumption ML communication phases from different parallelism dimensions are sequentially ordered and separated by idle windows large enough to hide OCS reconfiguration.
    Measured in §3.1 on one 4-node Llama3-8B workload; the entire Opus design depends on these windows existing at scale. The paper itself notes in §5 that some back-to-back traffic must be spaced, creating bubbles.
  • domain assumption Optical rails provide bandwidth equal to electrical rails for the collectives being compared.
    Stated in §4.2: 'our simulation assumes equal bandwidth between optical and electrical rails.' If optical links deliver less effective bandwidth for ring collectives, the iteration-time overhead would be higher.
  • domain assumption Profiled communication schedules from the first training iteration remain stable enough for speculative provisioning across later iterations.
    Opus provisioning in §4.1 fetches cached circuit configurations based on the profiled schedule; this assumes schedule stability and that controller/OCS programming completes within the window.
  • domain assumption The vendor-based cost and power model from [71,72] and current list prices is representative of real deployments.
    Used to produce the 70.5% cost and 95.84% power figures in §4.2; the model excludes fiber cable cost and power and is not validated against a deployed system.
  • ad hoc to paper Equation 1 correctly counts reconfiguration windows for arbitrary parallelization strategies and hardware scales.
    Introduced in §3.1 and used to claim 127 windows per Llama3.1-405B iteration; it is a hand-derived combinatorial count with no derivation or validation against large-scale traces.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Photonic Rails in ML Datacenters." pith.science (2026). https://pith.science/paper/VK5BK3YH

@misc{pith2026250708119,
  author       = {Pith},
  title        = {Pith review of: Photonic Rails in ML Datacenters},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VK5BK3YH}},
  note         = {Machine review of arXiv:2507.08119}
}
read the original abstract

Rail-optimized network fabrics have become the de facto datacenter scale-out fabric for large-scale ML training. However, the use of high-radix electrical switches to provide all-to-all connectivity in rails imposes massive power, cost, and complexity overheads. We propose a rethinking of the rail abstraction by retaining its communication semantics, but realizing it using optical circuit switches. The key challenge is that optical switches support only one-to-one connectivity at a time, limiting the fan-out of traffic in ML workloads using hybrid parallelisms. We introduce parallelism-driven rail reconfiguration as a solution that leverages the sequential ordering between traffic from different parallelisms. We design a control plane, Opus, to enable time-multiplexed emulation of electrical rail switches using optical switches. More broadly, our work discusses a new research agenda: datacenter fabrics that co-evolve with the model parallelism dimensions within each job, as opposed to the prevailing mindset of reconfiguring networks before a job begins.

Figures

Figures reproduced from arXiv: 2507.08119 by the authors.

Figure 1
Figure 1. Rail-optimized fabrics.We propose to replace packet [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Traffic in a training iteration with 3D parallelism. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Communication pattern for PP and FSDP in one [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: (a) CDF of window size from 10 iterations. (b) Rail 0 [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Reconfiguration during the warm-up stage of rank [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 8
Figure 8. Figure 8: Iteration time for varying network reconfiguration [PITH_FULL_IMAGE:figures/full_fig_p006_8.png]
Figure 7
Figure 7. Figure 7: GPU-backend network cost and power comparison [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

83 extracted references · 60 canonical work pages

  1. [1]

    Saksham Agarwal, Qizhe Cai, Rachit Agarwal, David Shmoys, and Amin Vahdat. 2024. Harmony: A congestion-free datacenter archi- tecture. In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24). 329–343

  2. [2]

    Daniel Amir, Nitika Saran, Tegan Wilson, Robert Kleinberg, Vishal Shrivastav, and Hakim Weatherspoon. 2024. Shale: A practical, scalable oblivious reconfigurable network. InProceedings of the ACM SIGCOMM 2024 Conference. 449–464

  3. [3]

    Hitesh Ballani, Paolo Costa, Raphael Behrendt, Daniel Cletheroe, Ist- van Haller, Krzysztof Jozwik, Fotini Karinou, Sophie Lange, Kai Shi, Benn Thomsen, et al . 2020. Sirius: A flat datacenter network with nanosecond optical switching. In Proceedings of the Annual confer- ence of the ACM Special Interest Group on Data Communication on the applications, te...

  4. [4]

    Maciej Besta and Torsten Hoefler. 2014. Slim Fly: A Cost Effective Low- Diameter Network Topology. InSC ’14: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 348–359. https://doi.org/10.1109/SC.2014.34

  5. [5]

    Broadcom Inc. 2025. BCM78909 51.2 -Tb/s Multilayer Co-Packaged Optics Switch. Online; accessed July 5, 2025. https://www.broadcom.com/products/fiber-optic-modules- components/co-packaged-optics/switches/bcm78909 A high -radix, high-bandwidth CPO switch supporting up to 64 ×800GbE or 128×400GbE

  6. [6]

    Broadcom Inc. 2025. Co -Packaged Optics (CPO). https://www. broadcom.com/info/optics/cpo. Accessed: 2025-07-03

  7. [7]

    Optical Systems Division

    Broadcom Inc. Optical Systems Division. 2021.SiPh Chiplets In Package (SCIP). Technical Report. Broadcom Inc., Irvine, CA, USA. https: //docs.broadcom.com/doc/siph-chiplets-in-package-scip OSD CPO SCIP_20211106 V5

  8. [8]

    CALIENT Technologies, Inc. 2022. Calient’s Optical Circuit Switch (S-Series) Datasheet. https://www.calient.net/wp-content/uploads/ 2022/06/Datasheet_Calients-Optical-Circuit-Switches.pdf. Accessed: 2025-07-03

Show all 83 references
  1. [9]

    Li Chen, Kai Chen, Zhonghua Zhu, Minlan Yu, George Porter, Chun- ming Qiao, and Shan Zhong. 2017. Enabling {Wide-Spread} Com- munications on Optical Fabric with{MegaSwitch}. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). 577–593

  2. [10]

    Weiwei Chu, Xinfeng Xie, Jiecao Yu, Jie Wang, Amar Phanishayee, Chunqiang Tang, Yuchen Hao, Jianyu Huang, Mustafa Ozdal, Jun Wang, et al. 2025. Scaling Llama 3 Training with Efficient Parallelism Strategies. In Proceedings of the 52nd Annual International Symposium on Computer...

  3. [11]

    Coherent Corp. 2025. Optical Circuit Switch (OCS). https://www. coherent.com/networking/optical-circuit-switch. Accessed: 2025- 07-10; Based on press release published March 25,2024; Coherent’s liquid-crystal-based OCS architecture supports up to 300×300 ports and is optimized...

  4. [12]

    EpiPhotonics Corp. 2025. Products. http://epiphotonics.com/products. html. Accessed: 2025-07-03

  5. [13]

    Nathan Farrington, George Porter, Sivasankar Radhakrishnan, Hamid Hajabdolali Bazzaz, Vikram Subramanya, Yeshaiahu Fainman, George Papen, and Amin Vahdat. 2010. Helios: a hybrid electri- cal/optical switch architecture for modular data centers. InProceedings of the ACM SIGCOMM...

  6. [14]

    William Fedus, Barret Zoph, and Noam Shazeer. 2022. Switch trans- formers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research 23, 120 (2022), 1–39

  7. [15]

    FS.COM. n.d.. Cisco Compatible 400GBASE-XDR4 QSFP-DD PAM4 1310nm 2km Module. https://www.fs.com/products/110530.html? attribute=94270&id=4477813. Accessed: 2025-07-02

  8. [16]

    FS.COM. n.d.. N9510 -64D 64-Port Ethernet L3 Data Center Switch (Broadcom Tomahawk-4, 64×400GbE). https://www.fs.com/products/ 149853.html. Accessed: 2025-07-02

  9. [17]

    Adithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu, Guilherme Goes, Hany Morsy, Rohit Puri, Mohammad Riftadi, Ashmitha Jeevaraj Shetty, Jingyi Yang, et al. 2024. Rdma over ether- net for distributed training at meta scale. In Proceedings of the ACM SIGCOMM 2024 Confer...

  10. [18]

    Alexandru M Gherghescu, Vlad-Andrei Bădoiu, Alexandru Agache, Mihai-Valentin Dumitru, Iuliu Vasilescu, Radu Mantu, and Costin Raiciu. 2024. I’ve Got 99 Problems But FLOPS Ain’t One. InProceedings of the 23rd ACM Workshop on Hot Topics in Networks . 195–204

  11. [19]

    Monia Ghobadi, Ratul Mahajan, Amar Phanishayee, Nikhil Deva- nur, Janardhan Kulkarni, Gireeja Ranade, Pierre-Alexandre Blanche, Houman Rastegarfar, Madeleine Glick, and Daniel Kilper. 2016. Pro- jecToR: Agile Reconfigurable Data Center Interconnect. InProceedings of the 2016 A...

  12. [20]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)

  13. [21]

    Hamilton, Navendu Jain, Srikanth Kandula, Changhoon Kim, Parantap Lahiri, David A

    Albert Greenberg, James R. Hamilton, Navendu Jain, Srikanth Kandula, Changhoon Kim, Parantap Lahiri, David A. Maltz, Parveen Patel, and Sudipta Sengupta. 2009. VL2: A Scalable and Flexible Data Center Network. In Proceedings of the ACM SIGCOMM 2009 Conference on Data Communica...

  14. [22]

    Das, Jon P

    Navid Hamedazimi, Zafar Qazi, Himanshu Gupta, Vyas Sekar, Samir R. Das, Jon P. Longtin, Himanshu Shah, and Ashish Tanwer. 2014. FireFly: 7 Eric Ding, Chuhan Ouyang, and Rachee Singh A Reconfigurable Wireless Data Center Fabric Using Free-Space Op- tics. In Proceedings of the 2...

  15. [23]

    Vipul Harsh, Sangeetha Abdu Jyothi, and P Brighten Godfrey. 2020. Spineless data centers. In Proceedings of the 19th ACM Workshop on Hot Topics in Networks. 67–73

  16. [24]

    Hewlett Packard Enterprise. 2021. HPE Cray EX Supercomputer Overview. https://www.hpe.com/psnow/doc/a50002546enw. Ac- cessed: 2025-07-09

  17. [25]

    Norm Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, et al. 2023. Tpu v4: An optically reconfigurable supercom- puter for machine learning with hardware support for embeddings. In Proceedings...

  18. [26]

    Mehrdad Khani, Manya Ghobadi, Mohammad Alizadeh, Ziyi Zhu, Madeleine Glick, Keren Bergman, Amin Vahdat, Benjamin Klenk, and Eiman Ebrahimi. [n. d.]. SiP-ML: High-Bandwidth Optical Network Interconnects for Machine Learning Training. In Proceedings of the 2021 ACM SIGCOMM 2021 ...

  19. [27]

    Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catan- zaro. 2023. Reducing activation recomputation in large transformer models. Proceedings of Machine Learning and Systems 5 (2023), 341– 353

  20. [28]

    Abhishek Vijaya Kumar, Arjun Devraj, Darius Bunandar, and Rachee Singh. 2024. A case for server-scale photonic connectivity. In Proceed- ings of the 23rd ACM Workshop on Hot Topics in Networks (Irvine, CA, USA) (HotNets ’24). Association for Computing Machinery, New York, NY, ...

  21. [29]

    Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj, Darius Bunandar, and Rachee Singh. 2025. LUMION: Fast Fault Recovery for ML Jobs Using Programmable Optical Fabrics. arXiv:2505.23105 [cs.LG] https: //arxiv.org/abs/2505.23105

  22. [30]

    Cong Liang, Xiangli Song, Jing Cheng, Mowei Wang, Yashe Liu, Zhen- hua Liu, Shizhen Zhao, and Yong Cui. 2024. NegotiaToR: Towards A Simple Yet Effective On-demand Reconfigurable Datacenter Network. In Proceedings of the ACM SIGCOMM 2024 Conference . 415–432

  23. [31]

    Wanchao Liang, Tianyu Liu, Less Wright, Will Constable, Andrew Gu, Chien-Chin Huang, Iris Zhang, Wei Feng, Howard Huang, Junjie Wang, et al. 2024. TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training. arXiv preprint arXiv:2410.06511 (2024)

  24. [32]

    Xudong Liao, Yijun Sun, Han Tian, Xinchen Wan, Yilun Jin, Zilong Wang, Zhenghang Ren, Xinyang Huang, Wenxue Li, Kin Fai Tse, et al

  25. [33]

    Lightmatter, Inc. 2025. Passage Technology. https://lightmatter.co/ products/passage/. Accessed: 2025-07-03

  26. [34]

    Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al . 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437 (2024)

  27. [35]

    Hong Liu, Ryohei Urata, Kevin Yasumura, Xiang Zhou, Roy Bannon, Jill Berger, Pedram Dashti, Norm Jouppi, Cedric Lam, Sheng Li, et al. 2023. Lightwave fabrics: at-scale optical circuit switching for datacenter and machine learning systems. In Proceedings of the ACM SIGCOMM 2023...

  28. [36]

    Hao Liu, Matei Zaharia, and Pieter Abbeel. 2023. Ring attention with blockwise transformers for near-infinite context. arXiv preprint arXiv:2310.01889 (2023)

  29. [37]

    Lumentum Holdings Inc. 2025. Lumentum Optical Circuit Switch to Improve Next -Generation AI Data Center Scalability. https: //www.lumentum.com/en/media-room/news-releases/lumentum- optical-circuit-switch-improve-next-generation-ai-data-center. Accessed June 20, 2025

  30. [38]

    William M Mellette, Rob McGuinness, Arjun Roy, Alex Forencich, George Papen, Alex C Snoeren, and George Porter. 2017. Rotornet: A scalable, low-complexity, optical datacenter network. In Proceedings of the Conference of the ACM Special Interest Group on Data Communi- cation. 267–280

  31. [39]

    Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. 2017. Mixed precision training. arXiv preprint arXiv:1710.03740 (2017)

  32. [40]

    National Energy Research Scientific Computing Center (NERSC). 2025. Perlmutter Architecture — NERSC Documentation. https://docs.nersc. gov/systems/perlmutter/architecture/. Accessed: 2025-07-04

  33. [41]

    NVIDIA. 2025. Llama-3.1-405B DGXC Benchmarking Recipe. https://catalog.ngc.nvidia.com/orgs/nvidia/teams/dgxc- benchmarking/resources/llama31-405b-dgxc-benchmarking-a. Version 24.11.1, modified January 29, 2025

  34. [42]

    NVIDIA Corporation. 2020. NVIDIA Collective Communication Library (NCCL): Creating a Communicator . NVIDIA. https://docs.nvidia.com/ deeplearning/nccl/user-guide/docs/usage/communicators.html Ac- cessed July 6, 2025

  35. [43]

    NVIDIA Corporation. 2022. Doubling all -to-all Performance with NCCL 2.12: Introducing PXN (PCI X NVLink). NVIDIA Devel- oper Blog. https://developer.nvidia.com/blog/doubling-all2all- performance-with-nvidia-collective-communication-library-2-12/ Describes PXN, which enables G...

  36. [44]

    NVIDIA Corporation. 2024. ConnectX -7 400G Adapters Datasheet. https://resources.nvidia.com/en-us-accelerated-networking- resource-library/connectx-7-datasheet. Accessed: 2025-07-02

  37. [45]

    NVIDIA Corporation. 2025. Co -Packaged Silicon Photonics Network- ing Switches. Online; accessed July 5,2025. https://www.nvidia.com/ en-us/networking/products/silicon-photonics/ Describes NVIDIA’s co-packaged optics (CPO) switches with integrated silicon photonics

  38. [46]

    NVIDIA Corporation. 2025. NVIDIA Announces Spectrum-X Photon- ics, Co-Packaged Optics Networking Switches to Scale AI Factories to Millions of GPUs . Press Release. NVIDIA Corporation, Santa Clara, CA, USA. https://nvidianews.nvidia.com/news/nvidia-spectrum-x- co-packaged-opti...

  39. [47]

    NVIDIA Corporation. 2025. NVIDIA Collective Communications Li- brary (NCCL). NVIDIA Developer. https://developer.nvidia.com/nccl Version 2.x; MPI-compatible multi-GPU / multi-node collective com- munication library

  40. [48]

    NVIDIA Corporation. 2025. NVIDIA DGX H200 Datasheet . Datasheet. NVIDIA Corporation, Santa Clara, CA. https://resources.nvidia.com/ en-us-dgx-systems/dgx-h200-datasheet Includes specifications of the DGX H200 system, featuring 8× H200 GPUs, dual Xeon Platinum 8480C CPUs, 2 TB ...

  41. [49]

    NVIDIA Corporation. 2025. NVIDIA DGX SuperPOD . NVIDIA. https://www.nvidia.com/en-us/data-center/dgx-superpod/ Full-stack data center platform scaling to tens of thousands of GPUs; includes compute, networking, storage, and software

  42. [50]

    NVIDIA Corporation. 2025. NVIDIA HGX Platform . NVIDIA. https: //www.nvidia.com/en-us/data-center/hgx/ Reference architecture 8 Photonic Rails in ML Datacenters combining GPUs, NVLink/NVSwitch, networking, and AI/HPC soft- ware stack

  43. [51]

    NVIDIA Corporation. 2025. Rail Optimized Topology Vali- dation. NVIDIA Networking, Santa Clara, CA. https: //docs.nvidia.com/networking/display/ibdiagnetusermanualv221/ Rail+Optimized+Topology+Validation Part of the ibdiagnet InfiniBand Fabric Diagnostic Tool User Manual; desc...

  44. [52]

    Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. 2013. On the difficulty of training recurrent neural networks. In International con- ference on machine learning . Pmlr, 1310–1318

  45. [53]

    Polatis (a HUBER+SUHNER company). n.d.. Series 7000 — 384 ×384-port Software -Defined Optical Circuit Switch. https://www.polatis.com/series-7000-384x384-port-software- controlled-optical-circuit-switch-sdn-enabled.asp. Accessed: 2025-07-01

  46. [54]

    Leon Poutievski, Omid Mashayekhi, Joon Ong, Arjun Singh, Mukar- ram Tariq, Rui Wang, Jianan Zhang, Virginia Beauregard, Patrick Con- ner, Steve Gribble, Rishi Kapoor, Stephen Kratzer, Nanfang Li, Hong Liu, Karthik Nagaraj, Jason Ornstein, Samir Sawhney, Ryohei Urata, Lorenzo V...

  47. [55]

    PyTorch Team. 2025. Automatic Mixed Precision package (torch.amp). https://pytorch.org/docs/stable/amp.html. Accessed: 2025-07-09

  48. [56]

    Penghui Qi, Xinyi Wan, Guangxing Huang, and Min Lin. 2023. Zero bubble pipeline parallelism. arXiv preprint arXiv:2401.10241 (2023)

  49. [57]

    Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He

  50. [58]

    Peter Sanders, Jochen Speck, and Jesper Larsson Träff. 2009. Two-tree algorithms for full bandwidth broadcast, reduction and scan. Parallel Comput. 35, 12 (2009), 581–594

  51. [59]

    Ken-ichi Sato. 2023. Optical switching will innovate intra data center networks [Invited Tutorial]. Journal of Optical Communications and Networking 16, 1 (2023), A1–A23

  52. [60]

    Aashaka Shah, Vijay Chidambaram, Meghan Cowan, Saeed Maleki, Madan Musuvathi, Todd Mytkowicz, Jacob Nelson, Olli Saarikivi, and Rachee Singh. 2023. TACCL: Guiding Collective Algorithm Synthe- sis using Communication Sketches. In 20th USENIX Symposium on Networked Systems Desig...

  53. [61]

    Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053 (2019)

  54. [62]

    Vishal Shrivastav, Asaf Valadarsky, Hitesh Ballani, Paolo Costa, Ki Suh Lee, Han Wang, Rachit Agarwal, and Hakim Weatherspoon. 2019. Shoal: A Network Architecture for Disaggregated Racks. In 16th USENIX Symposium on Networked Systems Design and Implementa- tion (NSDI 19) . USE...

  55. [63]

    Arjun Singh, Joon Ong, Amit Agarwal, Glen Anderson, Ashby Armis- tead, Roy Bannon, Seb Boving, Gaurav Desai, Bob Felderman, Paulie Germano, Anand Kanagala, Jeff Provost, Jason Simmons, Eiichi Tanda, Jim Wanderer, Urs Hölzle, Stephen Stuart, and Amin Vahdat. 2015. Jupiter Risin...

  56. [64]

    Ankit Singla, P Brighten Godfrey, and Alexandra Kolla. 2014. High throughput data center topology design. In 11th USENIX Symposium on Networked Systems Design and Implementation (NSDI 14) . 29–41

  57. [65]

    Brighten Godfrey

    Ankit Singla, Chi-Yao Hong, Lucian Popa, and P. Brighten Godfrey

  58. [66]

    Sithara P Sreenilayam, Dermot Brabazon, and Yuri P Panarin. 2019. Fast ferroelectric liquid crystal based optical switch: simulation and experiments. Crystals 9, 8 (2019), 388

  59. [67]

    Nouamane Tazi, Ferdinand Mom, Haojun Zhao, Phuc Nguyen, Mo- hamed Mekkouri, Leandro Werra, and Thomas Wolf. 2025. The Ultra- Scale Playbook: Training LLMs on GPU Clusters. https://huggingface. co/spaces/nanotron/ultrascale-playbook. Accessed: 2025-05-16

  60. [68]

    Telescent Inc. n.d.. Products | Telescent. https://www.telescent.com/ products. Accessed: 2025-07-01

  61. [69]

    Rajeev Thakur and William D Gropp. 2003. Improving the perfor- mance of collective operations in MPICH. In European Parallel Virtual Machine/Message Passing Interface Users’ Group Meeting . Springer, 257–267

  62. [70]

    Andersen, Michael Kaminsky, Konstantina Papagiannaki, T.S

    Guohui Wang, David G. Andersen, Michael Kaminsky, Konstantina Papagiannaki, T.S. Eugene Ng, Michael Kozuch, and Michael Ryan

  63. [71]

    Weiyang Wang, Manya Ghobadi, Kayvon Shakeri, Ying Zhang, and Naader Hasani. 2024. Rail-only: A low-cost high-performance network for training LLMs with trillion parameters. In 2024 IEEE Symposium on High-Performance Interconnects (HOTI) . IEEE, 1–10

  64. [72]

    2023.{TopoOpt}: Co-optimizing network topology and parallelization strategy for distributed training jobs

    Weiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi, Zhihao Jia, Dheevatsa Mudigere, Ying Zhang, and Anthony Kewitsch. 2023.{TopoOpt}: Co-optimizing network topology and parallelization strategy for distributed training jobs. In 20th USENIX Symposium on Networked System...

  65. [73]

    Zhenguo Wu, Liang Yuan Dai, Ziyi Zhu, Asher Novick, Madeleine Glick, and Keren Bergman. 2023. SiP Architecture For Accelerat- ing Collective Communication in Distributed Deep Learning. In 2023 Optical Fiber Communications Conference and Exhibition (OFC) . 1–3. https://doi.org/...

  66. [74]

    Eric P Xing, Qirong Ho, Wei Dai, Jin-Kyu Kim, Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, and Yaoliang Yu

  67. [75]

    Sharada Yeluri. 2023. Optimizing Power Consumption in High -End Routers

  68. [76]

    Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, et al

  69. [77]

    Yazhou Zu, Alireza Ghaffarkhah, Hoang-Vu Dang, Brian Towles, Steven Hand, Safeen Huda, Adekunle Bello, Alexander Kolbasov, Arash Rezaei, Dayou Du, Steve Lacy, Hang Wang, Aaron Wisner, Chris Lewis, and Henri Bahini. 2024. Resiliency at Scale: Managing Google’s TPUv4 Machine Lea...

  70. [2010]

    In Proceedings of the ACM SIGCOMM 2010 Conference(New Delhi, India)(SIGCOMM ’10)

    C-Through: Part-Time Optics in Data Centers. In Proceedings of the ACM SIGCOMM 2010 Conference(New Delhi, India)(SIGCOMM ’10). Association for Computing Machinery, New York, NY, USA, 327–338. https://doi.org/10.1145/1851182.1851222

  71. [2012]

    In 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12)

    Jellyfish: Networking Data Centers Randomly. In 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12). USENIX Association, San Jose, CA, 225–238. https://www.usenix. org/conference/nsdi12/technical-sessions/presentation/singla

  72. [2015]

    In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining

    Petuum: A new platform for distributed machine learning on big data. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . 1335–1344

  73. [2020]

    In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining

    Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining. 3505–3506

  74. [2023]

    arXiv preprint arXiv:2304.11277 (2023)

    Pytorch fsdp: experiences on scaling fully sharded data parallel. arXiv preprint arXiv:2304.11277 (2023)

  75. [2025]

    arXiv preprint arXiv:2501.03905 (2025)

    mFabric: An Efficient and Scalable Fabric for Mixture-of-Experts Training. arXiv preprint arXiv:2501.03905 (2025)

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.