REVIEW 4 major objections 4 minor 1 cited by
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training
T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read By exposing flow- and chunk-level states inside collective communication libraries, Mycroft claims it can detect LLM-training anomalies within 15 seconds and find their root cause within 20 seconds.
desk verdict Solid design and real deployment, but the headline detection numbers rest on a sampling-coverage assumption that doesn't hold for very large hybrid-parallel jobs, and the detailed evaluation is missing from this version. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the two-level dependency trace inside a collective operation. Flow-level tracing records per-network-queue-pair progress, because topology and faults are defined per flow, not per operation; chunk-level tracing records how much data each hardware and software component (GPU SM copies, RDMA writes) moved in each time window, without logging every chunk event. The trigger in Algorithm 1 samples at most ten ranks and fires on a missing completion log or a halved throughput/doubled interval; the analysis in Algorithm 2 then builds the global state machine from recent logs and follows the first missing or late chunk to locate the faulty rank.
What would settle it
Inject a silent single-flow degradation on a rank outside the sampled set—for example, throttle one NIC queue pair by 30% without blocking completion—and observe whether Mycroft's trigger fires. If the sampled CollOps show no missing completion and no throughput halving for several detection windows, the anomaly goes undetected, contradicting the cascade premise.
Extended reading notes
Core claim
The central claim is that the collective communication layer is the right and sufficient place to observe LLM-training reliability, because failures and performance anomalies surface there as a stalled or slowed CollOp. Mycroft instruments the proxy critical path of the underlying collective communication library with two logs: a completion log recording finished operations, and a periodic real-time state log recording accumulated progress on GPU SM copies and RDMA writes. From these logs it reconstructs a short-window distributed state machine and follows chunk dependencies to the minimum-op or minimum-data rank—the rank whose operation or data is missing relative to its peers. That rank is
Load-bearing premise
Every significant reliability or performance problem that begins on an unsampled rank will quickly cascade to one of the sampled ranks, so a stall or a halving of throughput appears in a monitored operation within the detection window.
Editorial extensions
If this is right
- Operators can distinguish the faulty rank from downstream victims: the rank whose chunk or operation is first missing or late is the root cause, not the ranks blocked behind it.
- Always-on detection scales to tens of thousands of GPUs because only a small sampled set of ranks needs continuous monitoring, assuming anomalies cascade quickly to sampled operations.
- Integration with Python-stack dumps and ring-buffer CollOp traces bounds the problematic layer: if those look normal, the fault is inside the collective communication layer and Mycroft exposes it.
- Detection and root-cause times drop to seconds, so operators can act before a silent timeout escalates into a checkpoint loss or a full restart.
- Fault-injection tests on a 32-GPU testbed across seven fault classes demonstrate that Mycroft detects and localizes the injected faults; the paper reports lower overhead than kernel-level tracing baselines.
Reading between the lines
- A direct test of the cascade assumption would inject a fault on a deliberately unsampled rank, one that degrades only that rank's flows without stalling or halving throughput of sampled CollOps; if the trigger stays silent, the 15-second detection claim fails exactly where the paper's weakest assumption predicts.
- The same flow/chunk trace data could double as a performance-optimization telemetry source, since it records where time is spent on each network path, not just where it fails.
- Any collective library with a proxy-style control path could carry the same tracepoints, making Mycroft a template for communication-layer observability rather than a single-library patch.
- Because the trigger thresholds (50% throughput drop, doubled interval, 1-second straggler) are hand-tuned, the next natural step—which the paper leaves as future work—is automatic threshold adjustment per job, which would determine how robust the method is across heterogeneous models and hardware.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Mycroft is a distributed tracing and root-cause-analysis system for collective communication (CCL) in LLM training. It instruments NCCL with two lightweight tracepoints, samples up to 10 ranks for always-on monitoring, and uses trigger rules (stalled CollOps, throughput drop by half, interval doubling) to detect anomalies. Once triggered, it builds a dependency-driven state machine from flow- and chunk-level traces and identifies the suspected root cause. The paper reports a six-month production deployment at ByteDance, with claims of 90% anomaly detection within 15 seconds and 60% root-cause identification within 20 seconds, plus fault-injection experiments on a 32-GPU testbed. The core contribution is Coll-level observability and dependency-driven diagnosis, which the authors argue is missing from existing Op-, kernel-, and RDMA-level tools.
Significance. If the deployment and evaluation claims hold, Mycroft is a significant step toward making CCLs observable for reliability debugging: it is one of the first systems to exploit flow- and chunk-level dependencies inside NCCL, and the reported second-level detection/RCA latency would be substantially faster than prior tools. The paper also provides a credible implementation effort (1100 lines of NCCL C++ instrumentation, 4000 lines of Python backend) and describes how Mycroft integrates with py-spy and PyTorch Flight Recorder. These strengths are real and should be credited. However, the strongest quantitative claims rest on evaluation evidence that is largely absent from the submitted text, and the sampling design has an internal inconsistency. The paper's central idea is promising and likely defensible, but the current manuscript does not yet provide enough support for the headline numbers.
major comments (4)
- [Section 4.3 (Algorithm 1 and sampling paragraph)] The text states both 'Mycroft samples at least one rank per DP group to ensure coverage' and 'limits sampling to at most 10 ranks.' For a 10k-GPU hybrid-parallel job with TP=8, PP=16, there are on the order of 128 independent DP groups, so 10 samples cannot include one rank from each group. Thus the coverage guarantee is false for typical large topologies. The fallback 'a failure or performance issue quickly cascades to the whole cluster' is asserted, and the evidence promised in Section 7.2 is not present. Since the 90%-within-15s detection claim depends on detecting anomalies that may start in an unsampled DP group, this is a load-bearing coverage hole. Please either correct the sampling statement, add deployment data on sampling coverage, or provide explicit cascade-propagation experiments.
- [Section 7 (Evaluation)] The manuscript does not include the actual quantitative results promised in the abstract and introduction. After listing seven fault-injection categories, the text jumps to a threshold discussion; there are no reported detection rates, detection latencies, RCA accuracies, false-positive counts, or overhead comparisons against NVRx/XPUTIMER/Aegis. The headline numbers (90% within 15s, 60% within 20s) are therefore unverifiable from the submitted text. Please add the missing experimental setup and results, including per-fault-class detection/RCA accuracy and a precise definition of what counts as a correctly identified 'root cause.'
- [Algorithm 1 / Section 4.3; Section 9] There is a circularity concern in the evaluation of detection: an anomaly is defined as a sampled-rank stall, a throughput drop by half, or an interval doubling, and the injected faults are designed to create exactly those signals. The detection rate then partly reduces to checking whether the rule fires on the rule-defined condition. To support the claimed detection capability, the paper should report threshold sensitivity, false-positive rates on healthy workloads, and how production incidents were independently labeled. The Limitations paragraph in Section 9 acknowledges that 'distinguishing true anomalies from benign variations remains challenging,' but the abstract's strong 90% claim is not tempered or supported by this analysis.
- [Algorithm 2 / Section 5.1] The RCA procedure relies on undefined functions 'CheckMinOp', 'CheckMinData', and 'CheckRCTable'. The 60%-within-20s root-cause claim cannot be assessed without specifying these functions, the meaning of 'suspected root cause' (faulty node, NIC, GPU, link, etc.), and an evaluation of localization precision over the fault classes. Please define the algorithm formally and provide a confusion-matrix-style breakdown for RCA accuracy.
minor comments (4)
- [Section 4.3] The text uses 'sampled IP list' and 'sampled ranks' interchangeably; clarify whether sampling is at the IP/host level or per-GPU rank, since this affects the number of monitored CollOps and the interpretation of 'one rank per DP group.'
- [Section 2.4 / Table 2] The completion log and real-time state log are described only in prose. A concrete schema or an example trace record for each log type would improve reproducibility and help readers understand what 'chunk stuck time' and 'RDMA_done' mean exactly.
- [Section 6.1] For 10,000 GPUs the paper reports ~3 TB of trace data per day with a fixed 512MB per-host buffer. Please include the per-rank trace generation rate and explain how the buffer size was chosen, since this is central to the low-overhead claim.
- [Table 1] The Mycroft row lists 'Seconds' for real-time analysis, but the table has no column for deployment scale or production validation. Consider adding a footnote or column to clarify that the seconds-level figure refers to the headline abstract claim, not to a measured per-case result presented in this manuscript.
Circularity Check
Fault-injection detection partly reduces to trigger-rule satisfaction; core production deployment claim remains independent.
-
self definitional
[Section 4.3 (Algorithm 1 trigger rules) and Section 9 (threshold tuning)]
"Mycroft detects performance issues if its throughput drops by half or the operation interval doubles. ... we therefore use a 50% bandwidth reduction as the trigger."
The trigger predicate is itself the operational definition of a performance anomaly: throughput halves or operation interval doubles. The threshold is then explicitly set to a 50% bandwidth reduction, and the fault-injection evaluation creates exactly that condition. Consequently, for injected bandwidth faults, the reported detection result is forced by the trigger's own rule rather than being an independent prediction about communication faults. The detection-rate claim in the fault-injection experiments reduces to rule satisfaction; however, the root-cause localization and the production deployment metrics (90% within 15s, 60% within 20s) are not reducible to this threshold and provide independent content.
full rationale
The main potential circularity is in the trigger-based evaluation: Algorithm 1 defines an anomaly as a stall, a halved throughput, or a doubled interval, and Section 9 states the threshold is a 50% bandwidth reduction. Injecting a 50% bandwidth reduction and then observing detection is therefore partly a tautology. I did not count this as a higher score because the production results are based on real six-month deployment, not on injected faults, and the root-cause analysis via dependency state machines is an independent algorithmic contribution. The self-citation to Minder [4] for the cascade assumption is not load-bearing circularity here, since Minder is a separate peer-reviewed system with its own evidence, and the same claim is also attributed to external work [3,61]. The sampling coverage inconsistency (at most 10 ranks vs. one rank per DP group) is a correctness risk, not a circularity, so it is not scored in this pass. Overall, the core derivation is not circular, but the fault-injection detection metric is partially self-confirming.
Assumptions & free parameters
free parameters (4)
- Throughput-drop trigger threshold =
50% bandwidth reduction
- CollOp interval-doubling trigger =
2x interval
- Straggler latency threshold =
1 second
- Real-time state log interval =
100 ms
assumptions (3)
- domain assumption Anomalies propagate cluster-wide within a few hundred milliseconds, so a small sampled subset of ranks observes them.
- domain assumption NCCL proxy thread state and periodic chunks of aggregate progress are sufficient to infer underlying data transmission health without full CUDA kernel traces.
- ad hoc to paper The trigger thresholds (50% throughput drop, 2x interval, 1 second straggler) correspond to meaningful anomalies across the diverse workloads at ByteDance.
Cite this review
Pith. "Pith review of Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training." pith.science (2026). https://pith.science/paper/Q6FISYDF
@misc{pith2026250903018,
author = {Pith},
title = {Pith review of: Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training},
year = {2026},
howpublished = {\url{https://pith.science/paper/Q6FISYDF}},
note = {Machine review of arXiv:2509.03018}
}
read the original abstract
Reliability is essential for ensuring efficiency in LLM training. However, many real-world reliability issues remain difficult to resolve, resulting in wasted resources and degraded model performance. Unfortunately, today's collective communication libraries operate as black boxes, hiding critical information needed for effective root cause analysis. We propose Mycroft, a lightweight distributed tracing and root cause analysis system designed to address previously hidden reliability issues in collective communication. Mycroft's key idea is to trace collective communication states and leverage internal control and data dependencies to resolve reliability problems in LLM training. Mycroft has been deployed at ByteDance for over six months to debug collective communication related issues at runtime. It detected anomalies within 15 seconds in 90% of cases and identified the root cause within 20 seconds in 60% of cases. We also conducted extensive fault injection experiments to demonstrate Mycroft's capability and efficiency.
Figures
Forward citations
Cited by 1 Pith paper
-
MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training
MegaScale-Data is a distributed data loading system that disaggregates preprocessing and applies auto-partitioning to deliver 4.5x higher end-to-end training throughput and 13.5x lower CPU memory usage for multisource...
Reference graph
Works this paper leans on
-
[2]
2025. RCCL. h/t_tps://github.com/ROCm/rccl. (2025). Accessed: 2025- 04-17
work page 2025
-
[3]
Weihao Cui, Ji Zhang, Han Zhao, Chao Liu, Wenhao Zhang, Jian Sha, Quan Chen, Bingsheng He, and Minyi Guo. 2025. XPUTimer: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale. arXiv preprint arXiv:2502.05413 (2025)
arXiv 2025
-
[4]
Yangtao Deng, Xiang Shi, Zhuo Jiang, Xingjian Zhang, Lei Zhang, Zhang Zhang, Bo Li, Zuquan Song, Hang Zhu, Gaohong Liu, et al
-
[5]
In 22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25)
Minder: Faulty Machine Detection for Large-scale Distributed Model Training. In 22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25) . 505–521
-
[6]
Jianbo Dong, Bin Luo, Jun Zhang, Pengcheng Zhang, Fei Feng, Yikai Zhu, Ang Liu, Zian Chen, Yi Shi, Hairong Jiao, Gang Lu, Yu Guan, Ennan Zhai, Wencong Xiao, Hanyu Zhao, Man Yuan, Siran Yang, Xiang Li, Jiamang Wang, Rui Men, Jianwei Zhang, Chang Zhou, Dennis Cai, Yuan Xie, and Binzhang Fu. 2025. Enhancing Large-Scale AI Training Efficiency: The C4 Solution f...
arXiv 2025
-
[7]
Jianbo Dong, Kun Qian, Pengcheng Zhang, Zhilong Zheng, Liang Chen, Fei Feng, Yichi Xu, Yikai Zhu, Gang Lu, Xue Li, et al . 2025. Evolution of Aegis: Fault Diagnosis for AI Model Training Service in Production. In 22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25) . 865–881
work page 2025
-
[8]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Ka- dian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783 (2024)
arXiv 2024
-
[9]
Rodrigo Fonseca, George Porter, Randy H Katz, and Scott Shenker
Show all 76 references
-
[10]
Richard L Graham, Timothy S Woodall, and Jeffrey M Squyres. 2006. Open MPI: A Flexible High Performance MPI. InParallel Processing and Applied Mathematics: 6th International Conference, PPAM 2005, Poznań, Poland, September 11-14, 2005, Revised Selected Papers 6 . Springer, 228– 239
2006
-
[11]
Swapnil Gandhi, Mark Zhao, Athinagoras Skiadopoulos, and Christos Kozyrakis. 2024. ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation. In Proceedings of the ACM SIGOPS 30th Sympo- sium on Operating Systems Principles . 211–228
2024
-
[12]
Chuanxiong Guo, Lihua Yuan, Dong Xiang, Yingnong Dang, Ray Huang, Dave Maltz, Zhaoyi Liu, Vin Wang, Bin Pang, Hua Chen, et al. 2015. Pingmesh: A Large-scale System for Data Center Network Latency Measurement and Analysis. In Proceedings of the 2015 ACM Conference on Special In...
2015
-
[13]
Haryadi S Gunawi, Riza O Suminto, Russell Sears, Casey Golliher, Swaminathan Sundararaman, Xing Lin, Tim Emami, Weiguang Sheng, Nematollah Bidokhti, Caitie McCaffrey, et al . 2018. Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production Sys- tems. ACM Tr...
2018
-
[14]
Qinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang, Meng Zhang, Qiaoling Chen, Peng Sun, Dahua Lin, Xiaolin Wang, Yingwei Luo, et al. 2024. Characterization of Large Language Model Development in the Datacenter. In 21st USENIX Symposium on Networked Systems Design and Implement...
2024
-
[15]
Aaron Harlap, Henggang Cui, Wei Dai, Jinliang Wei, Gregory R Ganger, Phillip B Gibbons, Garth A Gibson, and Eric P Xing. 2016. Addressing the Straggler Problem for Iterative Convergent Parallel ML. In Pro- ceedings of the seventh ACM symposium on cloud computing . 98–111
2016
-
[16]
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al. 2019. GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism. Advances in neural information processing systems 32 (2019)
2019
-
[17]
Peng Huang, Chuanxiong Guo, Lidong Zhou, Jacob R Lorch, Yingnong Dang, Murali Chintalapati, and Randolph Yao. 2017. Gray Failure: The Achilles’ Heel of Cloud-Scale Systems. In Proceedings of the 16th Workshop on Hot Topics in Operating Systems . 150–155
2017
-
[18]
Insu Jang, Zhenning Yang, Zhen Zhang, Xin Jin, and Mosharaf Chowd- hury. 2023. Oobleck: Resilient Distributed Training of Large Models Using Pipeline Templates. In Proceedings of the 29th Symposium on Operating Systems Principles . 382–395
2023
-
[19]
IBM. 2025. Autopilot. h/t_tps://github.com/IBM/autopilot. (2025). Ac- cessed: 2025-04-17
2025
-
[20]
Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang, Yangrui Chen, Zhi Zhang, Yanghua Peng, Xiang Li, Cong Xie, Shibiao Nong, et al
-
[21]
Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, Junjie Qian, Wencong Xiao, and Fan Yang. 2019. Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads. In 2019 USENIX Annual Technical Conference (USENIX ATC 19) . 947–960
2019
-
[22]
Ao Li, Shan Lu, Zhuotao Liu, Suman Nath, Michael Leighton, Rohan Padhye, Diedi Hu, Bingchuan Tian, Vyas Sekar, Maomao Ding, et al
-
[23]
Haoyang Li, Fangcheng Fu, Hao Ge, Sheng Lin, Xuanyu Wang, Ji- awen Niu, Yujie Wang, Hailin Zhang, Xiaonan Nie, and Bin Cui. 2024. Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization. arXiv preprint arXiv:2410...
2024 arXiv
-
[24]
Apostolos Kokolis, Michael Kuchnik, John Hoffman, Adithya Kumar, Parth Malani, Faye Ma, Zachary DeVito, Shubho Sengupta, Kalyan Saladi, and Carole-Jean Wu. 2024. Revisiting Reliability in Large-Scale Machine Learning Research Clusters. arXiv preprint arXiv:2410.21680 (2024)
2024 arXiv
-
[25]
Kefei Liu, Zhuo Jiang, Jiao Zhang, Shixian Guo, Xuan Zhang, Yangyang Bai, Yongbin Dong, Feng Luo, Zhang Zhang, Lei Wang, et al . 2024. R- pingmesh: A Service-aware RoCE Network Monitoring and Diagnostic System. In Proceedings of the ACM SIGCOMM 2024 Conference . 554– 567
2024
-
[26]
In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24)
ExChain: Exception Dependency Analysis for Root Cause Diag- nosis. In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24). 2047–2062
-
[27]
Jonathan Mace, Ryan Roelke, and Rodrigo Fonseca. 2018. Pivot Tracing: Dynamic Causal Monitoring for Distributed Systems. ACM Transac- tions on Computer Systems (TOCS) 35, 4 (2018), 1–28
2018
-
[28]
Mingyu Liang, Wenyin Fu, Louis Feng, Zhongyi Lin, Pavani Panakanti, Shengbao Zheng, Srinivas Sridharan, and Christina Delimitrou. 2023. Mystique: Enabling Accurate and Scalable Generation of Production AI Benchmarks. In Proceedings of the 50th Annual International Sympo- sium ...
2023
-
[29]
Meta. 2025. Holistic Trace Analysis. h/t_tps://github.com/ facebookresearch/HolisticTraceAnalysis. (2025). Accessed: 2025-04-17
2025
-
[30]
Kefei Liu, Zhuo Jiang, Jiao Zhang, Haoran Wei, Xiaolong Zhong, Lizhuang Tan, Tian Pan, and Tao Huang. 2023. Hostping: Diagnosing Intra-host Network Bottlenecks in RDMA Servers. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). 15–29
2023
-
[31]
Meta. 2025. OPT-175B logbook. h/t_tps://github.com/ facebookresearch/metaseq/blob/main/projects/OPT/chronicles/ OPT175B_Logbook.pdf. (2025). Accessed: 2025-04-17
2025
-
[32]
Luo Mai, Guo Li, Marcel Wagenländer, Konstantinos Fertakis, Andrei- Octavian Brabete, and Peter Pietzuch. 2020. KungFu: Making Training in Distributed Machine Learning Adaptive. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . 937–954
2020
-
[33]
Microsoft. 2025. NCCL Profiling Kit (NPKit). h/t_tps://github.com/ microso/f_t/NPKit. (2025). Accessed: 2025-04-17
2025
-
[34]
Meta. 2025. Kineto. h/t_tps://github.com/pytorch/kineto. (2025). Ac- cessed: 2025-04-17
2025
-
[35]
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R Devanur, Gregory R Ganger, Phillip B Gibbons, and Matei Zaharia. 2019. PipeDream: Generalized Pipeline Parallelism for DNN Training. In Proceedings of the 27th ACM symposium on operating systems princip...
2019
-
[36]
Microsoft. 2025. MSCCL. h/t_tps://github.com/microso/f_t/msccl. (2025). Accessed: 2025-04-17
2025
-
[37]
Maxim Naumov, John Kim, Dheevatsa Mudigere, Srinivas Sridharan, Xiaodong Wang, Whitney Zhao, Serhat Yilmaz, Changkyu Kim, Hector Yuen, Mustafa Ozdal, et al . 2020. Deep Learning Training in Facebook Data Centers: Design of Scale-up and Scale-out Systems. arXiv preprint arXiv:2...
2020 arXiv
-
[38]
Jayashree Mohan, Amar Phanishayee, and Vijay Chidambaram. 2021. CheckFreq: Frequent, Fine-Grained DNN Checkpointing. In 19th USENIX Conference on File and Storage Technologies (FAST 21) . 203– 216
2021
-
[39]
NVIDIA. 2025. CUDA Profiling Tools Interface. h/t_tps:// developer.nvidia.com/cupti. (2025). Accessed: 2025-04-17
2025
-
[40]
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGres- ley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, et al . 2021. Effi- cient Large-scale Language Model Training on GPU Clusters Using Megatron-LM. In ...
2021
-
[41]
NVIDIA. 2025. NCCL Test. h/t_tps://github.com/NVIDIA/nccl-tests. (2025). Accessed: 2025-04-17
2025
-
[42]
NVIDIA. 2025. ConnectX-6. h/t_tps://www.nvidia.com/en-sg/ networking/ethernet/connectx-6/. (2025). Accessed: 2025-04-17
2025
-
[43]
NVIDIA. 2025. NVIDIA A100 Tensor Core GPU. h/t_tps:// www.nvidia.com/en-us/data-center/a100/ . (2025). Accessed: 2025-04- 17
2025
-
[44]
NVIDIA. 2025. Error Injection. h/t_tps://docs.nvidia.com/datacenter/ dcgm/latest/user-guide/dcgm-error-injection .html. (2025). Accessed: 2025-04-17
2025
-
[45]
NVIDIA. 2025. nvidia-resiliency-ext. h/t_tps://nvidia.github.io/nvidia- resiliency-ext/. (2025). Accessed: 2025-04-17
2025
-
[46]
NVIDIA. 2025. Nsight Systems. h/t_tps://developer.nvidia.com/nsight- systems. (2025). Accessed: 2025-04-17
2025
-
[47]
NVIDIA. 2025. perftest. h/t_tps://github.com/linux-rdma/per/f_test. (2025). Accessed: 2025-04-17
2025
-
[48]
NVIDIA. 2025. NVIDIA Collective Communications Library (NCCL). h/t_tps://developer.nvidia.com/nccl. (2025). Accessed: 2025-04-17
2025
-
[49]
PyTorch. 2025. Flight Recorder. h/t_tps://pytorch.org/tutorials/ prototype/flight_recorder_tutorial.html. (2025). Accessed: 2025-04- 17
2025
-
[50]
Nvidia. 2025. NVLink and NVSwitch. h/t_tps://www.nvidia.com/en- us/data-center/nvlink/. (2025). Accessed: 2025-04-17
2025
-
[51]
Penghui Qi, Xinyi Wan, Guangxing Huang, and Min Lin. 2024. Zero Bubble (almost) Pipeline Parallelism. In The Twelfth International Con- ference on Learning Representations
2024
-
[52]
PyTorch. 2025. Distributed communication package - torch.distributed. h/t_tps://pytorch.org/docs/stable/distributed.html. (2025). Accessed: 2025-04-17
2025
-
[53]
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-LM: Training Multi-billion Parameter Language Models Using Model Parallelism. arXiv preprint arXiv:1909.08053 (2019)
2019 arXiv
-
[54]
PyTorch. 2025. PyTorch Profiler. h/t_tps://pytorch.org/docs/stable/ profiler.html. (2025). Accessed: 2025-04-17
2025
-
[55]
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al . 2022. Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530b, a Large-Scale Generative Language ...
2022 arXiv
-
[56]
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He
-
[57]
John Thorpe, Pengzhan Zhao, Jonathan Eyolfson, Yifan Qiao, Zhihao Jia, Minjia Zhang, Ravi Netravali, and Guoqing Harry Xu. 2023. Bam- boo: Making Preemptible Instances Resilient for Affordable Training of Large DNNs. In 20th USENIX Symposium on Networked Systems Design and Impl...
2023
-
[58]
Borui Wan, Mingji Han, Yiyao Sheng, Yanghua Peng, Haibin Lin, Mofan Zhang, Zhichao Lai, Menghan Yu, Junda Zhang, Zuquan Song, et al . 2024. ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development. arXiv preprint arXiv:2407.20143 (2024)
2024 arXiv
-
[59]
Benjamin H Sigelman, Luiz André Barroso, Mike Burrows, Pat Stephen- son, Manoj Plakal, Donald Beaver, Saul Jaspan, and Chandan Shanbhag
-
[60]
Zhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang, Xinwei Fu, TS Eu- gene Ng, and Yida Wang. 2023. GEMINI: Fast Failure Recovery in Distributed Training with In-Memory Checkpoints. In Proceedings of the 29th Symposium on Operating Systems Principles . 364–381
2023
-
[61]
Tianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang, Wenchao Wu, Qinkai Duan, Guodong Yang, Jiamang Wang, Lin Qu, and Lip- ing Zhang. 2025. GREYHOUND: Hunting Fail-Slows in Hybrid- Parallel Training at Scale. In 2025 USENIX Annual Technical Conference (USENIX ATC 25). 731–747
2025
-
[62]
Srinivas Sridharan, Taekyung Heo, Louis Feng, Zhaodong Wang, Matt Bergeron, Wenyin Fu, Shengbao Zheng, Brian Coutinho, Saeed Rashidi, Changhai Man, et al . 2023. Chakra: Advancing Performance Bench- marking and Co-design using Standardized Execution Traces. arXiv preprint arXi...
2023 arXiv
-
[63]
Yifan Xiong, Yuting Jiang, Ziyue Yang, Lei Qu, Guoshuai Zhao, Shuguang Liu, Dong Zhong, Boris Pinzur, Jie Zhang, Yang Wang, et al. 2024. SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation. In 2024 USENIX Annual Technical Conference (USENIX ATC ...
2024
-
[64]
Zhenhe Yao, Changhua Pei, Wenxiao Chen, Hanzhang Wang, Liangfei Su, Huai Jiang, Zhe Xie, Xiaohui Nie, and Dan Pei. 2024. Chain-of- Event: Interpretable Root Cause Analysis for Microservices through Automatically Learning Weighted Event Causal Graph. In Companion Proceedings of...
2024
-
[65]
Shibo Wang, Jinliang Wei, Amit Sabne, Andy Davis, Berkin Ilbeyi, Blake Hechtman, Dehao Chen, Karthik Srinivasa Murthy, Marcello Maggioni, Qiao Zhang, et al . 2022. Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models. In Proceedings ...
2022
-
[66]
Lei Zhang, Zhiqiang Xie, Vaastav Anand, Ymir Vigfusson, and Jonathan Mace. 2023. The Benefit of Hindsight: Tracing Edge-Cases in Distributed Systems. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) . 321–339
2023
-
[67]
Zhen Zhang, Chaokun Chang, Haibin Lin, Yida Wang, Raman Arora, and Xin Jin. 2020. Is Network the Bottleneck of Distributed Training?. In Proceedings of the Workshop on Network Meets AI & ML . 8–13
2020
-
[68]
Zhiqiang Xie, Yujia Zheng, Lizi Ottens, Kun Zhang, Christos Kozyrakis, and Jonathan Mace. 2024. Cloud Atlas: Efficient Fault Localization for Cloud Systems using Language Models and Causal Insight. arXiv preprint arXiv:2407.08694 (2024)
2024 arXiv
-
[69]
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, et al. 2023. Pytorch FSDP: Experiences on Scaling Fully Sharded Data Parallel. arXiv preprint arXiv:2304.11277 (2023)
2023 arXiv
-
[70]
Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P Xing, et al . 2022. Alpa: Automating Inter-and Intra-Operator Parallelism for Distributed Deep Learning. In 16th USENIX Symposium on Operating Syste...
2022
-
[71]
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022. Orca: A Distributed Serving System for Transformer-Based Generative Models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) . 521–538
2022
-
[74]
Xiaoyang Zhao, Zhe Zhang, and Chuan Wu. 2024. AdapCC: Making Collective Communication in Distributed Machine Learning Adaptive. In 2024 IEEE 44th International Conference on Distributed Computing Systems (ICDCS). IEEE, 25–35
2024
-
[2007]
In 4th USENIX Symposium on Networked Systems Design & Implementation (NSDI 07)
X-Trace: A Pervasive Network Tracing Framework. In 4th USENIX Symposium on Networked Systems Design & Implementation (NSDI 07)
-
[2010]
Technical report, Google, Inc (2010)
Dapper, a Large-Scale Distributed Systems Tracing Infrastruc- ture. Technical report, Google, Inc (2010)
2010
-
[2020]
In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis
Zero: Memory Optimizations Toward Training Trillion Param- eter Models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 1–16
-
[2024]
In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24)
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs. In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) . 745–760
-
[2025]
h/t_tps://pypi.org/project/py-spy/
py-spy. h/t_tps://pypi.org/project/py-spy/. (2025). Accessed: 2025-04-17
2025
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.