REVIEW 3 major objections 5 minor 101 references
By time-multiplexing optical circuits across parallelism phases, photonic rails can cut ML network power over 23× and cost 4× while adding under 7% to training time.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 23:45 UTC pith:TE3CXUCK
load-bearing objection A genuinely novel phase-multiplexed rail idea, but the headline overhead numbers rest on an unproven non-overlap assumption and a NIC firmware fix that the hardware doesn't deliver yet. the 3 major comments →
Opus: Photonic Rail-Optimized Fabric in ML Datacenters
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Opus's central claim is that the all-to-all connectivity invariant of rail-optimized fabrics need not be physically provisioned at all times: it can be presented as an illusion by an application-aware control plane that reconfigures a single optical circuit switch between the communication phases of different parallelism dimensions. Because collectives from different parallelisms (e.g., data-parallel ReduceScatter and pipeline-parallel Send/Recv) are separated by data dependencies in the model's compute graph, there are idle windows—often milliseconds long—between phases. If the optical switch can be reprogrammed within such a window, a GPU's few physical NIC ports can be time-multiplexed ac
What carries the argument
The central mechanism is parallelism-driven rail reconfiguration: within a single training iteration, the optical circuit switch is reprogrammed at parallelism phase boundaries to present a circuit topology tailored to the upcoming collective, using the same physical ports for every phase. The load-bearing object is the communication window—the idle interval between the end of one phase's collectives and the start of the next phase's—which is formalized as the minimum over next-phase collectives of the slowest-rank start time minus the maximum end time of the previous phase. Opus exploits this window in two ways: on-demand reconfiguration at phase transitions, and speculative provisioning wh
Load-bearing premise
The load-bearing premise is that the communication phases of different parallelism dimensions never overlap in time, so a single optical switch can be reprogrammed between them without stalling traffic—if DP and PP (or other) collectives overlap, the time-multiplexing mechanism collapses.
What would settle it
Record collective traces of a hybrid-parallel LLM training step and check whether any idle window between a pipeline-parallel Send/Recv and the following data-parallel ReduceScatter is ever shorter than the OCS reconfiguration time; a single such overlap, or a window below the switching latency, would invalidate the reported overhead.
If this is right
- The number of parallelisms a job can use is no longer bounded by the number of NIC ports per GPU; Opus's topology encoding supports up to 10 parallelism dimensions with only 2-degree ring connectivity.
- At production-relevant OCS reconfiguration latencies (up to 100 ms), the training overhead stays below 7%—around 5% on current GPU clusters with provisioning, and lower at 10 ms.
- Network power and cost scale with cluster size; at 2,048 GPUs the photonic rail shows 15–24× lower power and 3–4× lower cost than electrical rail fabrics, with the absolute savings growing as clusters grow.
- The datapath becomes GPU→NIC→optical fiber→NIC→GPU, removing OEO conversions and switch ASIC processing, so bandwidth scaling is no longer limited by ASIC speed.
- Opus works with existing training frameworks through a single backend flag, requiring no changes to model code or parallelism constructs, and its locking protocol ensures circuits are never torn down with traffic in flight.
Where Pith is reading between the lines
- An extension the paper leaves implicit: the same phase-window mechanism could serve inference pipelines and RL post-training, whose prefill/decode and rollout/optimize phases also have structured idle windows.
- The headline 23× power savings is measured against a fully electrical rail baseline; compared to hybrid fabrics that already use co-packaged optics, the relative gain would be smaller, so the figure is best read as the opportunity against today's standard deployment.
- A testable consequence of the central assumption: the paper's window measurements come from three LLM configurations; workloads that aggressively overlap communication with compute (e.g., zero-bubble pipeline schedules) may shrink inter-phase windows below the OCS reconfiguration time, and measuring window distributions across a broader workload space would bound Opus's applicability.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Opus proposes replacing electrical rail switches in ML datacenter fabrics with optical circuit switches (OCSes), while retaining the rail abstraction through 'parallelism-driven rail reconfiguration.' The key idea is to time-multiplex a single set of physical ports across circuit configurations that are optimized for each parallelism phase (DP, PP, FSDP, etc.), reconfiguring the OCS during the idle windows between communication phases of different parallelism dimensions. The paper describes a control plane (shim, controller, network orchestrator), implements it as a PyTorch backend, and evaluates it on a small physical OCS testbed, on the Perlmutter supercomputer via emulation (up to 64 GPUs), and in simulation at up to 2,048 GPUs. The headline results are >23× network power reduction and >4× cost savings, with less than 6.7% training iteration-time overhead at OCS reconfiguration latencies up to 100 ms.
Significance. If the central claims hold, this would be a significant contribution: it is one of the first systems to make a concrete case for replacing electrical packet switches with OCSes inside a widely deployed rail-optimized topology, without changing the number of NICs per GPU or the job's parallelism strategy. The paper includes a real hardware testbed, a working control-plane implementation, open-source code, and a three-scale evaluation, which are strengths for reproducibility. The power and cost numbers, if substantiated with a transparent methodology, would be important for datacenter designers. However, the headline performance claims rest on two assumptions that are not fully validated: that communication phases of different parallelism dimensions are strictly non-overlapping, and that NIC firmware can be made to support fast link-up after circuit reconfiguration. The physical testbed only demonstrates ~3 s end-to-end reconfiguration, and the simulator enforces the non-overlap assumption rather than testing it.
major comments (3)
- [§3.2, Eq. (1)–(2), Fig. 3, and §5.3 (AstraSim backend)] The central time-multiplexing mechanism assumes that communication phases of different parallelism dimensions never overlap. The window definition in Eq. (1) presupposes a gap between the end of all comm_i in P1 and the start of all comm_j in P2. The empirical support is limited to three TorchTitan workloads with PP=2/FSDP=2 and PP=3/FSDP=2; zero-bubble pipeline schedules, MoE AllToAll/AllGather, and aggressive compute-communication overlap are not covered. Critically, the simulation backend in §5.3 'rejects reconfiguration requests while collectives are in flight,' so the simulator enforces the non-overlap assumption rather than testing it. If a collective from the next phase arrives before reconfiguration completes, Opus's lock stalls that collective for the full reconfiguration latency. Please add evidence for schedules with overlap (e.g., zero-bubble, MoE, interleaved microbatches),
- [§5.1, Fig. 9(c)–(d)] The hardware testbed does not demonstrate the production-relevant reconfiguration latencies (≤100 ms) used in the paper's headline. The measured end-to-end reconfiguration is dominated by the NIC firmware: the Polatis switch returns optical power within ~200 ms, but the Mellanox firmware takes ~3 s (or ~6 s with auto-negotiation) to report link-up. The paper attributes this to firmware assumptions and says fast link-up is available with firmware support, but no such firmware is demonstrated or simulated at the hardware layer. Consequently, the physical system validates the control plane only at ~3 s reconfiguration time, while the 100 ms results come from emulation/simulation. Please temper the claim that the physical testbed validates production-relevant performance, or add a concrete path (e.g., modified/emulated firmware behavior) to bring the NIC link-up time into the OCS switching r
- [§5.3, Figure 14, 'Cost and power'] The cost and power savings—4.27× cost and 23.86× power for H200, 3.17× and 15.44× for GB200—are central to the paper's contribution, but the methodology is not described. The text only cites [16–18,44,52,63] and states that fiber cables are excluded. There is no bill of materials, no unit power/cost table, no switch/transceiver counts, and no sensitivity analysis. As written, the savings factors are not reproducible. Please provide a component-level cost/power model, including OCS, transceivers, NICs, and switch ASICs, and show how the savings vary with the assumed OCS port count, link rate, and pricing source.
minor comments (5)
- [Abstract and §1] The abstract states 'less than 6% training overhead,' while §1 and §5.3 report 'less than 6.7%' and specific values such as 5.31% and 11.22%. Please reconcile the abstract with the empirical numbers.
- [Eq. (5) / Fig. 5] The formula for the number of windows is presented as an equation with symbols (n_layer, n_microbatch, PP) but the terms in the right-hand side are not individually derived or defined in the text. Please define all terms and give a brief derivation of each additive component.
- [§5.1 / Fig. 9] The RDMA RETRY_CNT=7 setting is mentioned in the text but not discussed as a potential limitation. Since the paper claims no transport modifications, please clarify whether this setting affects fault tolerance or timeout behavior during reconfiguration.
- [Figure 3 and 4] The axis labels and legends in Figures 3 and 4 are hard to read in the provided text; e.g., repeated '0481204812' tick labels and 'Rail 0 window break-down' should be 'breakdown.' Please improve figure clarity.
- [§5.3 / Table 2] The simulation baseline 'EPS' is described as having all links active that Opus could form, but the paper does not specify the exact EPS topology (e.g., rail-optimized vs. fat-tree). Since the power/cost comparison uses 'EPS Rail-Opt' in Figure 14, please clarify whether the performance baseline is the same rail-optimized EPS or a generic electrical fabric.
Circularity Check
No significant circularity: the claimed power/cost and overhead results are derived from component models, measurements, and simulations whose inputs are not fitted to the claimed outputs; the self-citations are not load-bearing.
full rationale
I walked the paper's derivation chain and found no step in which a claimed prediction reduces by construction to its inputs. The headline quantitative results—over 23× power reduction, ~4× cost savings, and <6.7% iteration-time overhead—are produced by two independent mechanisms. The cost/power comparison is based on component counts and unit costs (e.g., "Cost and power exclude fiber cables. [16–18, 44, 52, 63]"), not on fitted values derived from the claimed savings. The iteration-time overhead is measured on a physical testbed, emulated on Perlmutter with injected reconfiguration delays, and simulated in AstraSim with Chakra traces; in all cases the reconfiguration latency is swept as an independent input (0–1000 ms) rather than tuned to reproduce the <6.7% number. The window measurements in §3.2 motivate the choice of reconfiguration latency target, but the simulation derives communication and compute times from model/workload configurations, so the overhead result is not an artifact of the window definition. The paper's central mechanism does rely on the empirical assumption that parallelism phases are non-overlapping, and the evidence for this is limited to three measured workloads; however, that is a generalizability/correctness concern, not circularity, because the assumption is stated as an observation and is not used to define the measured overhead. The citations to the authors' own prior work ([32] and [33]) appear in related-work and survey contexts (e.g., "A wide range of general datacenter fabric designs... [29, 30, 32, 33, 83, 85, 91]"), and are not invoked as proof of Opus's mechanism, uniqueness, or optimality. Accordingly, no load-bearing self-citation chain exists. The paper is a systems design with components, measurements, and simulations whose inputs are independent of the final claims, so the honest finding is no significant circularity.
Axiom & Free-Parameter Ledger
free parameters (3)
- profiling_steps =
5
- OCS reconfiguration latency in simulations =
swept: 0–1000 ms
- RDMA RETRY_CNT =
7
axioms (4)
- domain assumption Communication phases of different parallelisms do not overlap in time (there is a non-empty window between the end of one parallelism's last collective and the start of the next's first collective).
- domain assumption OCS reconfiguration latency can be hidden within phase-transition windows (i.e., the windows are larger than the reconfiguration delay).
- domain assumption The parallelism phase structure is stable across training iterations (the phase table learned in the first 5 steps remains valid).
- domain assumption OCS radix is sufficient to connect all GPUs of a rail (e.g., 384-512 ports for large scale-up domains).
read the original abstract
Rail-optimized network fabrics have become the de facto datacenter scale-out fabric for large-scale ML training. However, the use of high-radix electrical switches to provide all-to-all connectivity in rails imposes substantial power and cost. We propose a rethinking of the rail abstraction by retaining its communication semantics, but realizing it using optical circuit switches. The key challenge is that optical switches support one-to-one connectivity at a time, limiting the fan-out of traffic in ML workloads using hybrid parallelisms. We overcome this through \emph{parallelism-driven rail reconfiguration}, which exploits the non-overlapping communication phases of different parallelism dimensions. This time-multiplexes a single set of physical ports across circuit configurations tailored to each phase within a training iteration. We design and implement Opus, a control plane that orchestrates this in-job reconfiguration of photonic rails at parallelism phase boundaries, and evaluate it on a physical OCS testbed, the Perlmutter supercomputer, and in simulation at up to 2,048 GPUs. Our results show that photonic rails can achieve over $23\times$ network power reduction and $4\times$ cost savings while incurring only modest training overhead at production-relevant OCS reconfiguration latencies.
Figures
Reference graph
Works this paper leans on
-
[1]
NVIDIA GB200 NVL72
2025. NVIDIA GB200 NVL72. https://www.nvidia.com/en-us/ data-center/gb200-nvl72/. (2025). https://www.nvidia.com/en-us/ data-center/gb200-nvl72/ Accessed: 2026-02-07
2025
-
[2]
Saksham Agarwal, Qizhe Cai, Rachit Agarwal, David Shmoys, and Amin Vahdat. 2024. Harmony: A congestion-free datacenter archi- tecture. In21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24). 329–343
2024
-
[3]
Daniel Amir, Nitika Saran, Tegan Wilson, Robert Kleinberg, Vishal Shrivastav, and Hakim Weatherspoon. 2024. Shale: A practical, scalable oblivious reconfigurable network. InProceedings of the ACM SIGCOMM 2024 Conference. 449–464
2024
-
[4]
Hitesh Ballani, Paolo Costa, Raphael Behrendt, Daniel Cletheroe, Istvan Haller, Krzysztof Jozwik, Fotini Karinou, Sophie Lange, Kai Shi, Benn Thomsen, et al. 2020. Sirius: A flat datacenter network with nanosecond optical switching. InProceedings of the Annual conference of the ACM Special Interest Group on Data Communication on the applications, technolo...
2020
-
[5]
Kaoutar Benyahya, Ariel Gomez Diaz, Junyi Liu, Vassily Lyutsarev, Marianna Pantouvaki, Kai Shi, Shawn Yohanes Siew, Hitesh Ballani, Thomas Burridge, Daniel Cletheroe, et al. 2025. Mosaic: Breaking the Optics versus Copper Trade-off with a Wide-and-Slow Architecture and MicroLEDs. InProceedings of the ACM SIGCOMM 2025 Conference. 234–247
2025
-
[6]
Pankaj Berde, Matteo Gerola, Jonathan Hart, Yuta Higuchi, Masayoshi Kobayashi, Toshio Koide, Bob Lantz, Brian O’Connor, Pavlin Ra- doslavov, William Snow, et al . 2014. ONOS: towards an open, dis- tributed SDN OS. InProceedings of the third workshop on Hot topics in software defined networking. 1–6
2014
-
[7]
Maciej Besta and Torsten Hoefler. 2014. Slim Fly: A Cost Effective Low- Diameter Network Topology. InSC ’14: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 348–359. https://doi.org/10.1109/SC.2014.34
-
[8]
Broadcom Inc. 2025. BCM78909 51.2 -Tb/s Multilayer Co-Packaged Optics Switch. Online; accessed July 5, 2025. (2025). https: //www.broadcom.com/products/fiber-optic-modules-components/ co-packaged-optics/switches/bcm78909 A high-radix, high-bandwidth CPO switch supporting up to 64×800GbE or 128×400GbE
2025
-
[9]
Broadcom Inc. 2025. Co -Packaged Optics (CPO). https://www. broadcom.com/info/optics/cpo. (2025). Accessed: 2025-07-03
2025
-
[10]
Optical Systems Division
Broadcom Inc. Optical Systems Division. 2021.SiPh Chiplets In Package (SCIP). Technical Report. Broadcom Inc., Irvine, CA, USA. https: //docs.broadcom.com/doc/siph-chiplets-in-package-scip OSD CPO SCIP_20211106 V5
2021
-
[11]
Li Chen, Kai Chen, Zhonghua Zhu, Minlan Yu, George Porter, Chun- ming Qiao, and Shan Zhong. 2017. Enabling {Wide-Spread} Com- munications on Optical Fabric with{MegaSwitch}. In14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). 577–593
2017
-
[12]
Weiwei Chu, Xinfeng Xie, Jiecao Yu, Jie Wang, Amar Phanishayee, Chunqiang Tang, Yuchen Hao, Jianyu Huang, Mustafa Ozdal, Jun Wang, et al. 2025. Scaling Llama 3 Training with Efficient Parallelism Strategies. InProceedings of the 52nd Annual International Symposium on Computer Architecture. 1703–1716
2025
-
[13]
Coherent Corp. 2025. Optical Circuit Switch (OCS). https://www. coherent.com/networking/optical-circuit-switch. (2025). Accessed: 2025-07-10; Based on press release published March 25,2024; Coher- ent’s liquid-crystal-based OCS architecture supports up to 300×300 ports and is optimized for AI/ML data center fabrics
2025
-
[14]
Nathan Farrington, George Porter, Sivasankar Radhakrishnan, Hamid Hajabdolali Bazzaz, Vikram Subramanya, Yeshaiahu Fainman, George Papen, and Amin Vahdat. 2010. Helios: a hybrid electri- cal/optical switch architecture for modular data centers. InProceedings of the ACM SIGCOMM 2010 Conference. 339–350
2010
-
[15]
William Fedus, Barret Zoph, and Noam Shazeer. 2022. Switch trans- formers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research23, 120 (2022), 1–39
2022
-
[16]
FS.COM. n.d.. Cisco Compatible 400GBASE -XDR4 QSFP-DD PAM4 1310nm 2km Module. https://www.fs.com/products/110530.html? attribute=94270&id=4477813. (n.d.). Accessed: 2025-07-02
2025
-
[17]
FS.COM. n.d.. N9510 -64D 64-Port Ethernet L3 Data Center Switch (Broadcom Tomahawk-4, 64×400GbE). https://www.fs.com/products/ 149853.html. (n.d.). Accessed: 2025-07-02
2025
-
[18]
FS.com Inc. 2025. NVIDIA/Mellanox MMA4Z00-NS Optical Transceiver Module. https://www.fs.com/products/229253.html. (2025). Product page, Accessed: 2026-02-06
2025
-
[19]
Swapnil Gandhi, Mark Zhao, Athinagoras Skiadopoulos, and Christos Kozyrakis. 2024. Recycle: Resilient training of large dnns using pipeline adaptation. InProceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles. 211–228
2024
-
[20]
Adithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu, Guilherme Goes, Hany Morsy, Rohit Puri, Mohammad Riftadi, Ashmitha Jeevaraj Shetty, Jingyi Yang, et al . 2024. Rdma over eth- ernet for distributed training at meta scale. InProceedings of the ACM SIGCOMM 2024 Conference. 57–70
2024
-
[21]
Alexandru M Gherghescu, Vlad-Andrei Bădoiu, Alexandru Agache, Mihai-Valentin Dumitru, Iuliu Vasilescu, Radu Mantu, and Costin Raiciu. 2024. I’ve Got 99 Problems But FLOPS Ain’t One. InProceedings of the 23rd ACM Workshop on Hot Topics in Networks. 195–204
2024
-
[22]
Monia Ghobadi, Ratul Mahajan, Amar Phanishayee, Nikhil Deva- nur, Janardhan Kulkarni, Gireeja Ranade, Pierre-Alexandre Blanche, Houman Rastegarfar, Madeleine Glick, and Daniel Kilper. 2016. Pro- jecToR: Agile Reconfigurable Data Center Interconnect. InProceed- ings of the 2016 ACM SIGCOMM Conference (SIGCOMM ’16). Asso- ciation for Computing Machinery, Ne...
arXiv 2016
-
[23]
Hamilton, Navendu Jain, Srikanth Kandula, Changhoon Kim, Parantap Lahiri, David A
Albert Greenberg, James R. Hamilton, Navendu Jain, Srikanth Kandula, Changhoon Kim, Parantap Lahiri, David A. Maltz, Parveen Patel, and Sudipta Sengupta. 2009. VL2: A Scalable and Flexible Data Center Network. InProceedings of the ACM SIGCOMM 2009 Conference on 13 Data Communication (SIGCOMM ’09). Association for Computing Ma- chinery, New York, NY, USA, ...
doi:10.1145/1592568 2009
-
[24]
Navid Hamedazimi, Zafar Qazi, Himanshu Gupta, Vyas Sekar, Samir R. Das, Jon P. Longtin, Himanshu Shah, and Ashish Tanwer. 2014. Fire- Fly: A Reconfigurable Wireless Data Center Fabric Using Free-Space Optics. InProceedings of the 2014 ACM Conference on SIGCOMM (SIG- COMM ’14). Association for Computing Machinery, New York, NY, USA, 319–330. https://doi.or...
arXiv 2014
-
[25]
Vipul Harsh, Sangeetha Abdu Jyothi, and P Brighten Godfrey. 2020. Spineless data centers. InProceedings of the 19th ACM Workshop on Hot Topics in Networks. 67–73
2020
-
[26]
Hewlett Packard Enterprise. 2021. HPE Cray EX Supercomputer Overview. https://www.hpe.com/psnow/doc/a50002546enw. (2021). Accessed: 2025-07-09
2021
-
[27]
2010.{ZooKeeper}: Wait-free coordination for internet-scale systems
Patrick Hunt, Mahadev Konar, Flavio P Junqueira, and Benjamin Reed. 2010.{ZooKeeper}: Wait-free coordination for internet-scale systems. In2010 USENIX Annual Technical Conference (USENIX ATC 10)
2010
-
[28]
Insu Jang, Zhenning Yang, Zhen Zhang, Xin Jin, and Mosharaf Chowd- hury. 2023. Oobleck: Resilient distributed training of large models using pipeline templates. InProceedings of the 29th Symposium on Operating Systems Principles. 382–395
2023
-
[29]
Norm Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, et al. 2023. Tpu v4: An optically reconfigurable supercom- puter for machine learning with hardware support for embeddings. In Proceedings of the 50th annual international symposium on computer architecture. 1–14
2023
-
[30]
Mehrdad Khani, Manya Ghobadi, Mohammad Alizadeh, Ziyi Zhu, Madeleine Glick, Keren Bergman, Amin Vahdat, Benjamin Klenk, and Eiman Ebrahimi. [n. d.]. SiP-ML: High-Bandwidth Optical Network Interconnects for Machine Learning Training. InProceedings of the 2021 ACM SIGCOMM 2021 Conference
2021
-
[31]
Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro
-
[32]
Abhishek Vijaya Kumar, Arjun Devraj, Darius Bunandar, and Rachee Singh. 2024. A case for server-scale photonic connectivity. InProceed- ings of the 23rd ACM Workshop on Hot Topics in Networks (HotNets ’24). Association for Computing Machinery, New York, NY, USA, 290–299. https://doi.org/10.1145/3696348.3696856
arXiv 2024
-
[33]
Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj, Darius Bunandar, and Rachee Singh. 2025. Morphlux: Transforming Torus Fabrics for Efficient Multi-tenant ML. (2025). arXiv:cs.NI/2508.03674 https://arxiv. org/abs/2508.03674
arXiv 2025
-
[34]
ChonLam Lao, Minlan Yu, Aditya Akella, Jiamin Cao, Yu Guan, Pengcheng Zhang, Zhilong Zheng, Yichi Xu, Ennan Zhai, Dennis Cai, et al. 2024. TrainMover: Efficient ML Training Live Migration with No Memory Overhead.arXiv e-prints(2024), arXiv–2412
2024
-
[35]
Cong Liang, Xiangli Song, Jing Cheng, Mowei Wang, Yashe Liu, Zhen- hua Liu, Shizhen Zhao, and Yong Cui. 2024. NegotiaToR: Towards A Simple Yet Effective On-demand Reconfigurable Datacenter Network. InProceedings of the ACM SIGCOMM 2024 Conference. 415–432
2024
-
[36]
Wanchao Liang, Tianyu Liu, Less Wright, Will Constable, Andrew Gu, Chien-Chin Huang, Iris Zhang, Wei Feng, Howard Huang, Junjie Wang, et al. 2024. TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training.arXiv preprint arXiv:2410.06511 (2024)
Pith/arXiv arXiv 2024
-
[38]
Xudong Liao, Yijun Sun, Han Tian, Xinchen Wan, Yilun Jin, Zilong Wang, Zhenghang Ren, Xinyang Huang, Wenxue Li, Kin Fai Tse, et al
-
[39]
Linux. [n. d.].ethtool(8) - Linux man page. https://linux.die.net/man/8/ ethtool
-
[40]
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al
-
[41]
InProceedings of the ACM SIGCOMM 2025 Conference
Mixnet: A runtime reconfigurable optical-electrical fabric for distributed mixture-of-experts training. InProceedings of the ACM SIGCOMM 2025 Conference. 554–574
2025
-
[42]
Lumentum Holdings Inc. 2025. Lumentum Optical Circuit Switch to Improve Next -Generation AI Data Center Scalabil- ity. https://www.lumentum.com/en/media-room/news-releases/ lumentum-optical-circuit-switch-improve-next-generation-ai-data-center. (26 March 2025). Accessed June 20, 2025
2025
-
[43]
William M Mellette, Rob McGuinness, Arjun Roy, Alex Forencich, George Papen, Alex C Snoeren, and George Porter. 2017. Rotornet: A scalable, low-complexity, optical datacenter network. InProceedings of the Conference of the ACM Special Interest Group on Data Communica- tion. 267–280
2017
-
[44]
NADDOD. 2025. NVIDIA Quantum-X800 XDR InfiniBand Switch, Q3400-RA. https://www.naddod.com/products/nvidia-networking/ 102612. (2025). Reseller price listing, Accessed: 2026-02-06
2025
-
[45]
Hao Liu, Matei Zaharia, and Pieter Abbeel. 2023. Ring attention with blockwise transformers for near-infinite context.arXiv preprint arXiv:2310.01889(2023)
Pith/arXiv arXiv 2023
-
[46]
nEye Systems. 2025. nEye: Dismantling Network Walls to Build a Sustainable AI Future. https://www.neye.ai/. (2025). Optical circuit switch platform for AI datacenter networking. Accessed: 2025-02-02
2025
-
[47]
2024.NVIDIA Firmware Tools (MFT) Docu- mentation
NVIDIA. 2024.NVIDIA Firmware Tools (MFT) Docu- mentation. https://docs.nvidia.com/networking/display/ nvidia-firmware-tools-mft-documentation-v4-32-0.0.pdf
2024
-
[48]
NVIDIA. 2025. Llama-3.1-405B DGXC Benchmarking Recipe. https://catalog.ngc.nvidia.com/orgs/nvidia/teams/ dgxc-benchmarking/resources/llama31-405b-dgxc-benchmarking-a. (2025). Version 24.11.1, modified January 29, 2025
2025
-
[49]
National Energy Research Scientific Computing Center (NERSC). 2025. Perlmutter Architecture — NERSC Documentation. https://docs.nersc. gov/systems/perlmutter/architecture/. (2025). Accessed: 2025-07-04
2025
-
[50]
NVIDIA Corporation. 2022. Doubling all -to-all Performance with NCCL 2.12: Introducing PXN (PCI X NVLink). NVIDIA Developer Blog. (Feb. 2022). https://developer.nvidia.com/blog/ doubling-all2all-performance-with-nvidia-collective-communication-library-2-12/ Describes PXN, which enables GPU-to-NIC communication via NVLink to optimize rail-aligned collectiv...
2022
-
[51]
NVIDIA Corporation. 2024. ConnectX -7 400G Adapters Datasheet. https://resources.nvidia. com/en-us-accelerated-networking-resource-library/ connectx-7-datasheet. (2024). Accessed: 2025-07-02
2024
-
[52]
NVIDIA Corporation. 2024. NVIDIA Q32xx and Q34xx XDR 800Gb/s InfiniBand Switch Systems User Manual. https://docs.nvidia.com/ networking/display/xdrswitcheshwum/specifications. (2024). Ac- cessed: 2026-02-06
2024
-
[53]
2020.NVIDIA Collective Communication Library (NCCL): Creating a Communicator
NVIDIA Corporation. 2020.NVIDIA Collective Communication Library (NCCL): Creating a Communicator. NVIDIA. https://docs.nvidia.com/ deeplearning/nccl/user-guide/docs/usage/communicators.html Ac- cessed July 6, 2025
2020
-
[54]
NVIDIA Corporation. 2025. Co -Packaged Silicon Photonics Network- ing Switches. Online; accessed July 5,2025. (2025). https://www. nvidia.com/en-us/networking/products/silicon-photonics/ Describes NVIDIA’s co-packaged optics (CPO) switches with integrated silicon photonics
2025
-
[55]
2025.NVIDIA Announces Spectrum -X Photonics, Co -Packaged Optics Networking Switches to Scale AI Factories to Millions of GPUs
NVIDIA Corporation. 2025.NVIDIA Announces Spectrum -X Photonics, Co -Packaged Optics Networking Switches to Scale AI Factories to Millions of GPUs. Press Release. NVIDIA Corpora- tion, Santa Clara, CA, USA. https://nvidianews.nvidia.com/news/ nvidia-spectrum-x-co-packaged-optics-networking-switches-ai-factories Unveiled at GTC 2025
2025
-
[56]
2025.NVIDIA Collective Communications Library (NCCL)
NVIDIA Corporation. 2025.NVIDIA Collective Communications Library (NCCL). NVIDIA Developer. https://developer.nvidia.com/nccl Version 2.x; MPI-compatible multi-GPU / multi-node collective communication library
2025
-
[57]
NVIDIA Corporation. 2025. ConnectX-6 Dx Firmware Download. https://network.nvidia.com/support/firmware/connectx6dx/. (2025). Accessed: 2026-02-06. 14
2025
-
[58]
2025.NVIDIA DGX SuperPOD
NVIDIA Corporation. 2025.NVIDIA DGX SuperPOD. NVIDIA. https: //www.nvidia.com/en-us/data-center/dgx-superpod/ Full-stack data center platform scaling to tens of thousands of GPUs; includes compute, networking, storage, and software
2025
-
[59]
NVIDIA Corporation. 2025. NVIDIA Ethernet Driver for Linux (mlnx_en). https://network.nvidia.com/products/ethernet-drivers/ linux/mlnx_en/. (2025). Accessed: 2026-02-06
2025
-
[60]
2025.NVIDIA HGX Platform
NVIDIA Corporation. 2025.NVIDIA HGX Platform. NVIDIA. https: //www.nvidia.com/en-us/data-center/hgx/ Reference architecture combining GPUs, NVLink/NVSwitch, networking, and AI/HPC soft- ware stack
2025
-
[61]
2025.NVIDIA DGX H200 Datasheet
NVIDIA Corporation. 2025.NVIDIA DGX H200 Datasheet. Datasheet. NVIDIA Corporation, Santa Clara, CA. https://resources.nvidia.com/ en-us-dgx-systems/dgx-h200-datasheet Includes specifications of the DGX H200 system, featuring 8×H200 GPUs, dual Xeon Platinum 8480C CPUs, 2 TB system memory, 30 TB NVMe SSD, and full NVIDIA AI Enterprise software stack
2025
-
[62]
Jeremie Eliahou Ontiveros, Dylan Patel, and Wei Zhou
-
[63]
Polatis. 2023. Polatis Series 6000n Optical Switch Datasheet. https://www.redhelix.com/wp-content/uploads/2023/11/Polatis_ 6000n_Data_Sheet-rhl.pdf. (2023). Datasheet, Accessed: 2026-02-06
2023
-
[64]
Polatis (a HUBER+SUHNER company). n.d.. Series 7000 - 384x384-port Software-Defined Optical Circuit Switch. https://www.polatis.com/ series-7000-384x384-port-software-controlled-optical-circuit-switch-sdn-enabled. asp. (n.d.). Accessed: 2025-07-01
2025
-
[65]
2025.Rail Optimized Topology Val- idation
NVIDIA Corporation. 2025.Rail Optimized Topology Val- idation. NVIDIA Networking, Santa Clara, CA. https: //docs.nvidia.com/networking/display/ibdiagnetusermanualv221/ Rail+Optimized+Topology+Validation Part of the ibdiagnet InfiniBand Fabric Diagnostic Tool User Manual; describes cabling validation and compute-fabric alignment in DGX SuperPOD rail-optimi...
2025
-
[66]
Penghui Qi, Xinyi Wan, Guangxing Huang, and Min Lin. 2023. Zero bubble pipeline parallelism.arXiv preprint arXiv:2401.10241(2023)
Pith/arXiv arXiv 2023
-
[67]
SemiAnal- ysis
xAI’s Colossus 2 - First Gigawatt Datacenter In The World, Unique RL Methodology, Capital Raise. SemiAnal- ysis. (Sept. 2025). https://newsletter.semianalysis.com/p/ xais-colossus-2-first-gigawatt-datacenter Accessed: 2026-01-23
2025
-
[68]
Saeed Rashidi, Srinivas Sridharan, Sudarshan Srinivasan, and Tushar Krishna. 2020. ASTRA-SIM: Enabling SW/HW Co-Design Exploration for Distributed DL Training Platforms. InIEEE International Sympo- sium on Performance Analysis of Systems and Software, ISPASS 2020, Boston, MA, USA, August 22-26, 2020. IEEE
2020
-
[69]
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He
-
[70]
Leon Poutievski, Omid Mashayekhi, Joon Ong, Arjun Singh, Mukar- ram Tariq, Rui Wang, Jianan Zhang, Virginia Beauregard, Patrick Con- ner, Steve Gribble, Rishi Kapoor, Stephen Kratzer, Nanfang Li, Hong Liu, Karthik Nagaraj, Jason Ornstein, Samir Sawhney, Ryohei Urata, Lorenzo Vicisano, Kevin Yasumura, Shidong Zhang, Junlan Zhou, and Amin Vahdat. 2022. Jupi...
2022
-
[71]
Aashaka Shah, Vijay Chidambaram, Meghan Cowan, Saeed Maleki, Madan Musuvathi, Todd Mytkowicz, Jacob Nelson, Olli Saarikivi, and Rachee Singh. 2023. TACCL: Guiding Collective Algorithm Synthe- sis using Communication Sketches. In20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). USENIX As- sociation, Boston, MA, 593–612. https:...
2023
-
[72]
Kun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao, Yichi Xu, Yu Guan, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao, et al. 2024. Alibaba hpn: A data center network for large language model training. In Proceedings of the ACM SIGCOMM 2024 Conference. 691–706
2024
-
[73]
Vishal Shrivastav, Asaf Valadarsky, Hitesh Ballani, Paolo Costa, Ki Suh Lee, Han Wang, Rachit Agarwal, and Hakim Weatherspoon. 2019. Shoal: A Network Architecture for Disaggregated Racks. In16th USENIX Symposium on Networked Systems Design and Implementa- tion (NSDI 19). USENIX Association, Boston, MA, 255–270. https: //www.usenix.org/conference/nsdi19/pr...
2019
-
[74]
Arjun Singh, Joon Ong, Amit Agarwal, Glen Anderson, Ashby Armis- tead, Roy Bannon, Seb Boving, Gaurav Desai, Bob Felderman, Paulie Germano, Anand Kanagala, Jeff Provost, Jason Simmons, Eiichi Tanda, Jim Wanderer, Urs Hölzle, Stephen Stuart, and Amin Vahdat. 2015. Jupiter Rising: A Decade of Clos Topologies and Centralized Control in Google’s Datacenter Ne...
arXiv 2015
-
[75]
Ankit Singla, P Brighten Godfrey, and Alexandra Kolla. 2014. High throughput data center topology design. In11th USENIX Symposium on Networked Systems Design and Implementation (NSDI 14). 29–41
2014
-
[76]
Peter Sanders, Jochen Speck, and Jesper Larsson Träff. 2009. Two-tree algorithms for full bandwidth broadcast, reduction and scan.Parallel Comput.35, 12 (2009), 581–594
2009
-
[77]
Srinivas Sridharan, Taekyung Heo, Louis Feng, Zhaodong Wang, Matt Bergeron, Wenyin Fu, Shengbao Zheng, Brian Coutinho, Saeed Rashidi, Changhai Man, and Tushar Krishna. 2023. Chakra: Advancing Perfor- mance Benchmarking and Co-design using Standardized Execution Traces.arXiv preprint arXiv:2305.14516(2023)
Pith/arXiv arXiv 2023
-
[78]
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi- billion parameter language models using model parallelism.arXiv preprint arXiv:1909.08053(2019)
Pith/arXiv arXiv 2019
-
[79]
Rajeev Thakur and William D Gropp. 2003. Improving the perfor- mance of collective operations in MPICH. InEuropean Parallel Virtual Machine/Message Passing Interface Users’ Group Meeting. Springer, 257– 267. 15
2003
-
[80]
Andersen, Michael Kaminsky, Konstantina Papagiannaki, T.S
Guohui Wang, David G. Andersen, Michael Kaminsky, Konstantina Papagiannaki, T.S. Eugene Ng, Michael Kozuch, and Michael Ryan
-
[81]
Guohui Wang, David G Andersen, Michael Kaminsky, Konstantina Papagiannaki, TS Eugene Ng, Michael Kozuch, and Michael Ryan. 2010. c-Through: Part-time optics in data centers. InProceedings of the ACM SIGCOMM 2010 Conference. 327–338
2010
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.