REVIEW 3 major objections 5 minor 57 references
SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims that hangs, stragglers, and silent data corruption in LLM pre-training can all be localized online by a single rule: a rank whose behavior disagrees with the strict majority of its equivalent replicas is the outlier.
desk verdict A genuinely useful consensus-based failure-localization framework with honest small-scale evidence; the checkpoint-certification claim outruns its replay coverage and needs to be scoped down before publication. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Consensus Collective Communication (C3) abstraction is the central mechanism: a diagnostic collective that AllGathers compact evidence from each rank, finds the most frequent value for exact comparisons or a robust median for timing comparisons, and marks every rank whose evidence diverges from the strict majority as an outlier. It is supported by an out-of-band CPU observer that reads shared-memory progress and collective fingerprints after a hang, and by in-situ replay that reruns captured layers on live accelerators to expose stragglers and numerical corruption under production conditions. A coverage-compression rule for Mixture-of-Experts shapes keeps replay bounded, and a checkpoint gate promotes a candidate checkpoint to verified only after complete clean replay cycles.
What would settle it
Inject a fault that affects a strict majority of an equivalent peer group, for example a shared library that corrupts the same parameter on most replicas or a cluster-wide slowdown, and check whether SCOUT returns 'agree' with an empty outlier bitmap while the job remains unhealthy.
Extended reading notes
Core claim
SCOUT's central claim is that the three latent failure classes that plague synchronous LLM pre-training—hangs, stragglers, and silent data corruption—can all be localized by the same consensus rule: align equivalent replicas, gather compact evidence, and attribute divergence to any rank that a strict majority disagrees with. The paper argues that this converts redundancy already present in data-parallel and FSDP training into live diagnostic evidence, eliminating the need for absolute health thresholds, preselected golden ranks, or post-mortem reconstruction. For hangs, the evidence is a collective fingerprint and progress coordinate read by an out-of-band CPU observer; for stragglers, it is replay timing on controlled equivalent work; for SDC, it is a deterministic numerical signature. Clean replay coverage also certifies which checkpoint is numerically trustworthy for recovery.
Load-bearing premise
The load-bearing assumption is that healthy ranks hold a strict majority in every compared peer group and all report the same healthy value; if a bug or environmental fault affects a majority, SCOUT sees agreement and the failure escapes localization.
Editorial extensions
If this is right
- A hang that stalls a job can be attributed to a specific rank when its published collective fingerprint or progress coordinate diverges from the majority, while equal-progress stalls are reported as group-scoped for fabric diagnosis.
- A recurring straggler is identified by controlled replay timing: a rank that remains slower than equivalent peers on identical work, with computation and communication timed separately to distinguish a slow module from a slow collective.
- SDC is localized by deterministic numerical signatures: the rank whose output or gradient differs from the healthy majority is named, and the contaminated checkpoint is excluded from recovery.
- Clean replay coverage certifies a checkpoint: only after a complete recipe cycle passes without SDC can a candidate checkpoint be promoted to verified, so recovery avoids reintroducing corrupted state.
- SCOUT attaches to existing training stacks through public interfaces, so it can provide this diagnosis without modifying training loops or framework source.
Reading between the lines
- If the strict-majority rule holds at production scale, the same principle could be extended to other synchronous distributed workloads beyond LLM training, such as HPC simulations, wherever equivalent replicas exist.
- The checkpoint-certification logic implies a trade-off: the longer the verification cycle, the more healthy work is discarded on recovery, so tuning the recipe-catalog size and cadence against corruption probability is a natural next step the paper does not quantify.
- The assumption that healthy values are identical may break for numerical signatures under legitimate nondeterminism, so a testable extension is to make the comparison tolerate small numerical differences rather than requiring exact equality.
- A testable extension is to use SCOUT's per-surface outlier bitmaps to build a failure-type classifier that distinguishes compute, communication, and input stalls, which the current evaluation only partially demonstrates.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SCOUT, a runtime service that localizes latent failures in synchronous LLM pre-training by comparing each rank's progress, timing, and numerical signatures against those of equivalent replicas and reporting the outlier under a strict-majority rule. It introduces C3, a collective-comparison abstraction; an out-of-band CPU observer for hangs; in-situ replay for straggler and SDC diagnosis; an MoE shape-compression catalog; and a checkpoint gate that promotes a recovery checkpoint only after two clean recipe cycles. The evaluation uses a 16-A100 two-host testbed and fault injection across DDP, FSDP2, and HSDP, plus dedicated harnesses for MoE replay compression and kernel-role SDC. Results report perfect or near-perfect pass rates for the injected scenarios, with the paper explicitly stating that production-scale, end-to-end overhead, and RDMA/multi-rack validation remain future work.
Significance. SCOUT's core principle—using strict-majority consensus among data-parallel replicas as a live reference—is simple and well motivated by production evidence that faults are usually rank-local and rare. The paper's strengths are real: the implementation is open source; the fault-injection results are reported as exact counts with an honest statement of the one-seed, 16-GPU scope; the failure schemas (progress fingerprints for hangs, timed replay for stragglers, deterministic signatures for SDC) are cleanly separated; and the C3 abstraction returns an outlier bitmap while keeping the comparison rule local. The explicit limitation passages in Sections 5.3 and 11 are unusually candid. If the results hold, the framework would give operators a single runtime path to name a faulty rank or peer group for three failure classes and to gate recovery checkpoints; the missing overhead and scale measurements, however, mean the current evidence supports mechanism behavior on a small testbed rather than production deployment claims.
major comments (3)
- [Section 8, Algorithm 3, Table 3, Section 11] The checkpoint certification claim in the abstract and Section 8 is stronger than the rotating-coverage contract that SCOUT actually implements. Section 8 states that for dense models 'every in-memory checkpoint captured after an accepted SCOUT check is certified by that check' because shapes are static, but static shape does not imply that the values of un-replayed layers or un-replayed optimizer slices were verified. A persistent SDC in any un-sampled hidden layer, un-replayed optimizer slice, or un-qualified MoE shape is never observed, so the two-clean-cycle gate in Algorithm 3 can promote a numerically corrupt checkpoint as verified. Table 3's recovery cells inject faults only at the replayed surfaces (DDP parameter after backward, FSDP2 replay output, HSDP gradient shard), so the experiments demonstrate exclusion of faults at replayed surfaces, not certification of the whole checkpoint. Section 11 narrows the guarantee to 'Replay covers the layers, values, shapes, and communication paths exercised by its rotating checks'; the abstract and Section 8 should carry the same qualification, or the certification contract needs an explicit coverage argument.
- [Section 5.3 and Section 8] The strict-majority assumption is a load-bearing blind spot for the recovery claim. Section 5.3 explicitly concedes that 'a faulty value held by a strict majority is indistinguishable from the expected value,' and that common-mode corruption making every peer agree remains invisible. Yet the abstract and Section 8 present SCOUT as 'preventing recovery from selecting state corrupted by SDC' without that caveat. A deterministic software bug, a shared-library defect, or a cluster-wide environment event that corrupts a majority of replicas would yield an Agree verdict and a certified checkpoint. The paper should either state the healthy-majority assumption in the abstract and Section 8, or add and evaluate a temporal/majority-crossing detection path.
- [Sections 6.1.1, 10, and 11] The paper's practicality claim rests on an analytical overhead estimate, but no end-to-end measurement is provided. Section 6.1.1 estimates the amortized layer-equivalent replay overhead as V/(IN) ≈ 0.3% for V=3, I=20, N=50, while Section 11 states that the paper 'does not measure end-to-end throughput and resource overhead across replay cadences, or report recovery time and rollback distance.' Since Section 1 identifies minimizing the overhead of continuous diagnosis as a core design challenge, the absence of measured throughput overhead, recovery latency, multi-rack/RDMA validation, and error bars around the one-seed fault-injection counts leaves the production-deployment claim unverified. The revision should either add such measurements or explicitly scope the contributions to mechanism behavior on the small testbed.
minor comments (5)
- [Abstract and Section 1] There are spacing and formatting errors in the text, such as 'usesitsConsensus Collective Communication(C3)' and 'an O( 10,000)-GPU cluster'; these should be cleaned up.
- [Figures 2 and 3] The captions contain missing spaces in expressions like 'a4 ×2DP–FSDP mesh' and in the logical mesh notation; please format the dimension products consistently.
- [Section 9 and Section 10] Section 9 describes the GEMINI checkpoint coordinator, but Section 10 does not report a recovery experiment that explicitly exercises GEMINI's in-memory checkpoint path; please clarify whether the checkpoint-recovery tests used the GEMINI code path or a stand-in.
- [Section 6.2 and Table 4] The execution-path coverage preorder is evaluated only for uniform per-expert row counts, and Section 11 concedes that arbitrary heterogeneous routing vectors require separately qualified templates; this limitation should be stated at the point where the compression contribution is introduced.
- [Algorithm 3] The two-clean-cycle promotion logic is described in prose but would be clearer if the pseudocode included an explicit comment that the candidate promoted at the end of a cycle is the checkpoint captured at the previous cycle boundary.
Circularity Check
No significant circularity: SCOUT's majority-consensus verdicts are computed from independently collected runtime evidence, and the acknowledged replay-coverage limits bound the checkpoint-certification claim rather than closing a definitional loop.
full rationale
SCOUT's core derivation—localizing a failure to the rank whose progress, timing, or numerical signature disagrees with the strict-majority value among aligned peers—is computed directly from runtime evidence (Algorithm 1: AllGather the evidence, take the most frequent value for exact modes, median/MAD for timing modes, and set the outlier bitmap where evidence differs). No parameter is fitted to the testbed to make these verdicts pass; kappa, replay cadence, sampling interval, and recipe sizes are operator-configured or coverage-driven (§5.4, §6.1.1, §6.2). The only self-citation is GEMINI [46], an externally published checkpointing system; SCOUT builds on it as a component, and the paper does not use it to justify the consensus principle or to forbid alternatives, so it does not make the derivation circular. The nearest thing to a definitional loop is the checkpoint-certification claim: 'Clean replay coverage certifies checkpoint numerical integrity' is stronger than the coverage contract the paper itself concedes in Section 11 ('Replay covers the layers, values, shapes, and communication paths exercised by its rotating checks'), and the fault-injection cells in Table 3 only perturb surfaces SCOUT replays. That is a real coverage/overclaim concern, but it is not input-output circularity: the certification verdict is produced by executing those replay checks, not by assuming the conclusion. The healthy-majority limitation is stated explicitly in §5.3 ('a faulty value held by a strict majority is indistinguishable from the expected value'), which is an honest boundary rather than a hidden premise. I therefore find no significant circularity; score 0.
Assumptions & free parameters
free parameters (3)
- kappa (statistical consensus sensitivity multiplier) =
not specified (operator-configured)
- replay interval I and variant count V =
example V=3, I=20 used for overhead estimate
- optimizer replay slice size (64K elements) =
64K
assumptions (4)
- domain assumption Equivalent peers performing the same work provide a strict-majority healthy reference; failures are rare relative to healthy ranks.
- domain assumption At least three equivalent replicas exist per peer group for rank attribution; data parallelism or FSDP state sharding provides them.
- domain assumption Captured layer execution is deterministic given identical inputs and RNG state, with deterministic kernels in replay scope.
- domain assumption Framework public interfaces expose all relevant collectives and state transitions.
invented entities (1)
-
C3 (Consensus Collective Communication) abstraction
independent evidence
Cite this review
Pith. "Pith review of SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training." pith.science (2026). https://pith.science/paper/MEZ6FP32
@misc{pith2026260811034,
author = {Pith},
title = {Pith review of: SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training},
year = {2026},
howpublished = {\url{https://pith.science/paper/MEZ6FP32}},
note = {Machine review of arXiv:2608.11034}
}
read the original abstract
In LLM pre-training, synchronization propagates rank-local stalls, slowdowns, and numerical errors into job-wide symptoms, obscuring their origin. Existing diagnosis often relies on in-process monitors that cannot report after the trainer blocks or terminates, or on post-mortem logs that preserve only synchronized symptoms; offline health tests lose the workload and operating conditions that triggered the failure. We present SCOUT, a unified runtime failure-localization framework built on one design principle: identify outliers through strict-majority consensus among equivalent replicas. SCOUT aligns replica progress, timing, and numerical evidence, then uses its Consensus Collective Communication (C3) abstraction to identify ranks whose compact signatures disagree with their peers. An out-of-band CPU observer remains responsive when training hangs, whereas in-situ replay exercises recurring stragglers and silent data corruption (SDC) beside the live job with its model state, kernels, allocations, communication path, and thermal and memory pressure present. Collective fingerprints expose rank-local protocol divergence. Clean replay coverage certifies checkpoint numerical integrity, preventing recovery from selecting state corrupted by SDC. SCOUT integrates with PyTorch, TorchTitan, Megatron-Core, and DeepSpeed without training-loop or framework-source modifications. SCOUT is open source at https://github.com/LMResiliency/lm-resiliency.
Figures
Reference graph
Works this paper leans on
- [1]
-
[2]
DeepSeek-AI. 2026. DeepSeek-V4: Towards Highly Efficient Million-Token Con- text Intelligence. arXiv:2606.19348
arXiv 2026
-
[3]
Yangtao Deng, Xiang Shi, Zhuo Jiang, Xingjian Zhang, Lei Zhang, Zhang Zhang, Bo Li, Zuquan Song, Hang Zhu, Gaohong Liu, Fuliang Li, Shuguang Wang, Haibin Lin, Jianxi Ye, and Minlan Yu. 2025. Minder: Faulty Machine Detection for Large- Scale Distributed Model Training. In22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25). USENIX...
work page 2025
-
[5]
Jianbo Dong, Kun Qian, Pengcheng Zhang, Zhilong Zheng, Liang Chen, Fei Feng, Yichi Xu, Yikai Zhu, Gang Lu, Xue Li, Zhihui Ren, Zhicheng Wang, Bin Luo, Peng Zhang, Yang Liu, Yanqing Chen, Yu Guan, Weicheng Wang, Chaojie Yang, Yang Zhang, Man Yuan, Hanyu Zhao, Yong Li, Zihan Zhao, Shan Li, Xianlong Zeng, Zhiping Yao, Binzhang Fu, Ennan Zhai, Wei Lin, Chao W...
work page 2025
-
[6]
Assaf Eisenman, Kiran Kumar Matam, Steven Ingram, Dheevatsa Mudigere, Raghuraman Krishnamoorthi, Krishnakumar Nair, Misha Smelyanskiy, and Mu- rali Annavaram. 2022. Check-N-Run: A Checkpointing System for Training Deep Learning Recommendation Models. In19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22). USENIX Association, Rent...
work page 2022
-
[7]
Trevor Gale, Deepak Narayanan, Cliff Young, and Matei Zaharia. 2022. MegaBlocks: Efficient Sparse Training with Mixture-of-Experts. arXiv:2211.15841
arXiv 2022
-
[8]
GLM-5 Team. 2026. GLM-5: From Vibe Coding to Agentic Engineering. arXiv:2602.15763
arXiv 2026
-
[9]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. 2024. The Llama 3 Herd of Models. arXiv:2407.21783
arXiv 2024
Show all 57 references
-
[10]
Yu Guan, Zhiyu Yin, Haoyu Chen, Sheng Cheng, Chaojie Yang, Kun Qian, Tianyin Xu, Pengcheng Zhang, Yang Zhang, Hanyu Zhao, Yong Li, Dennis Cai, and Ennan Zhai. 2026. EROICA: Online Performance Troubleshooting for Large- Scale Model Training. In23rd USENIX Symposium on Networked...
2026
-
[11]
Le, Yonghui Wu, and Zhifeng Chen
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Xu Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen. 2019. GPipe: Efficient Training of Giant Neural Networks Using Pipeline Parallelism. InAdvances in Neural Information Proces...
2019
-
[12]
Boris Iglewicz and David C. Hoaglin. 1993.How to Detect and Handle Outliers. ASQC Basic References in Quality Control: Statistical Techniques, Vol. 16. ASQC Quality Press, Milwaukee, WI
1993
-
[13]
Yuxuan Jiang, Ziming Zhou, Boyu Xu, Beijie Liu, Runhui Xu, and Peng Huang
-
[14]
Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang, Yangrui Chen, Zhi Zhang, Yanghua Peng, Xiang Li, Cong Xie, Shibiao Nong, Yulu Jia, Sun He, Hongmin Chen, Zhihao Bai, Qi Hou, Shipeng Yan, Ding Zhou, Yiyao Sheng, Zhuo Jiang, Haohan Xu, Haoran Wei, Zhang Zhang, Pengfei Nie, Leqi...
2024
-
[15]
Chao Jin, Ziheng Jiang, Zhihao Bai, Zheng Zhong, Juncai Liu, Xiang Li, Ningxin Zheng, Xi Wang, Cong Xie, Qi Huang, Wen Heng, Yiyuan Ma, Wenlei Bao, Size Zheng, Yanghua Peng, Haibin Lin, Xuanzhe Liu, Xin Jin, and Xin Liu. 2026. MegaScale-MoE: Large-Scale Communication-Efficient...
2026
-
[16]
Kimi Team. 2026. Kimi K2.5: Visual Agentic Intelligence. arXiv:2602.02276
2026 arXiv
-
[17]
Leslie Lamport, Robert Shostak, and Marshall Pease. 1982. The Byzantine Gen- erals Problem.ACM Transactions on Programming Languages and Systems4, 3 (1982), 382–401. doi:10.1145/357172.357176
1982
-
[18]
ChonLam Lao, Jiaqi Gao, Jiamin Cao, Zhipeng Zhang, Pengcheng Zhang, Jiangfei Duan, Zhilong Zheng, Yu Guan, Yichi Xu, Yong Li, Zhengping Qian, Aditya Akella, Minlan Yu, Ennan Zhai, Dennis Cai, and Jingren Zhou. 2026. Train- Mover: An Interruption-Resilient Runtime for ML Traini...
2026
-
[19]
Kinman Lei, Liyan Zheng, Xiang Li, Hongmin Chen, Yun Zhang, Gaohong Liu, Zuquan Song, Zixuan Ma, Zhiyu Xue, Minghui Yu, Shuguang Wang, Wencong Xiao, Haibin Lin, Yuyang Jin, Jidong Zhai, Bo Liu, and Xin Liu. 2026. Safe- guarding LLM Training at Scale: Online SDC Detection and I...
2026
-
[20]
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. arXiv:2006.16668
2020 arXiv
-
[21]
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, and Soumith Chintala
-
[22]
Wanchao Liang, Tianyu Liu, Less Wright, Will Constable, Andrew Gu, Chien- Chin Huang, Iris Zhang, Wei Feng, Howard Huang, Junjie Wang, Sanket Puran- dare, Gokul Nadathur, and Stratos Idreos. 2025. TorchTitan: One-Stop PyTorch- Native Solution for Production-Ready LLM Pre-Train...
2025
-
[23]
Jinkun Lin, Ziheng Jiang, Zuquan Song, Sida Zhao, Menghan Yu, Zhanghan Wang, Chenyuan Wang, Zuocheng Shi, Xiang Shi, Wei Jia, Zherui Liu, Shuguang Wang, Haibin Lin, Xin Liu, Aurojit Panda, and Jinyang Li. 2025. Understanding Stragglers in Large Model Training Using What-if Ana...
2025
-
[24]
Linux Kernel Contributors. 2026. In-Field Scan. Linux Kernel Documentation. https://docs.kernel.org/arch/x86/ifs.html Accessed 2026-08-10
2026
-
[25]
Linux Kernel Contributors. 2026. Reliability, Availability and Serviceability (RAS). Linux Kernel Documentation. https://docs.kernel.org/admin-guide/RAS/main. html Accessed 2026-08-10
2026
-
[26]
Phillip Liu, Uttam Thakore, Junjie Wang, and Justin Yang. 2026. Flight Recorder: A New Lens for Understanding NCCL Watchdog Timeouts. PyTorch Blog. https://pytorch.org/blog/flight-recorder-a-new-lens-for-understanding- nccl-watchdog-timeouts/ Accessed 2026-08-03
2026
-
[28]
Dongning Ma, Fred Lin, Alban Desmaison, Joel Coburn, Daniel Moore, Sriram Sankar, and Xun Jiao. 2024. Dr. DNA: Combating Silent Data Corruptions in Deep Learning Using Distribution of Neuron Activations. InProceedings of the 29th ACM International Conference on Architectural S...
2024
-
[29]
Jeffrey Ma, Hengzhi Pei, Leonard Lausen, and George Karypis. 2025. Understand- ing Silent Data Corruption in LLM Training. arXiv:2502.12340
2025 arXiv
-
[30]
Meta AI. 2025. The Llama 4 Herd: The Beginning of a New Era of Natively Multimodal AI Innovation. Meta AI Blog. https://ai.meta.com/blog/llama-4- multimodal-intelligence/ Accessed 2026-08-03
2025
-
[31]
Smith, Pang Wei Koh, Amanpreet Singh, and Hannaneh Hajishirzi
Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Weijia Shi, Pete Walsh, Oyvind Tafjord, Nathan Lambert, Yuling Gu, Shane Arora, Akshita Bhagia, Dustin Schwenk, David Wadden, Alexander Wettig, Binyuan Hui, Tim Dettmers, Douwe Kiela, Ali F...
2024 arXiv
-
[32]
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGres- ley, Mostofa Patwary, Vijay Anand Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. 2021. Efficient Large-Scale Language Model Training on ...
2021
-
[33]
NVIDIA. 2026. CUTLASS Grouped Kernel Schedulers. CUTLASS Documenta- tion. https://github.com/NVIDIA/cutlass/blob/main/media/docs/cpp/grouped_ scheduler.md Accessed 2026-08-07
2026
-
[34]
NVIDIA Corporation. 2026. Health Monitoring. NVIDIA Data Center GPU Manager Documentation. https://docs.nvidia.com/datacenter/dcgm/latest/learn/ modules/health-monitoring.html Accessed 2026-08-03
2026
-
[35]
NVIDIA Corporation. 2026. Megatron-Core Tensor-Parallel Random-State API. NVIDIA Documentation. https://docs.nvidia.com/megatron-core/developer- guide/latest/apidocs/core/core.tensor_parallel.random.html Accessed 2026-08- 08
2026
-
[36]
NVIDIA Corporation. 2026. Megatron-Core User Guide. NVIDIA Documentation. https://docs.nvidia.com/megatron-core/developer-guide/latest/ Accessed 2026- 08-04
2026
-
[37]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Sys...
2019
-
[38]
PyTorch Contributors. 2026. Elastic Agent. PyTorch Documentation. https: //docs.pytorch.org/docs/stable/elastic/agent.html Accessed 2026-08-03
2026
-
[39]
PyTorch Contributors. 2026. Reproducibility. PyTorch Documentation. https: //docs.pytorch.org/docs/stable/notes/randomness.html Accessed 2026-08-08
2026
-
[40]
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models. InProceed- ings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, Atlanta, GA, US...
2020 arXiv
-
[41]
Arjun Roy, Hongyi Zeng, Jasmeet Bagga, and Alex C. Snoeren. 2017. Passive Realtime Datacenter Fault Detection and Localization. In14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). USENIX Association, Boston, MA, 595–612. https://www.usenix.org/con...
2017
-
[42]
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. arXiv:1909.08053 15
2019 arXiv
-
[43]
Cheng Tan, Ze Jin, Chuanxiong Guo, Tianrong Zhang, Haitao Wu, Karl Deng, Dongming Bi, and Dong Xiang. 2019. NetBouncer: Active Device and Link Failure Localization in Data Center Networks. In16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). USENIX...
2019
-
[44]
John Thorpe, Pengzhan Zhao, Jonathan Eyolfson, Yifan Qiao, Zhihao Jia, Minjia Zhang, Ravi Netravali, and Guoqing Harry Xu. 2023. Bamboo: Making Pre- emptible Instances Resilient for Affordable Training of Large DNNs. In20th USENIX Symposium on Networked Systems Design and Impl...
2023
-
[45]
Borui Wan, Gaohong Liu, Zuquan Song, Jun Wang, Yun Zhang, Guangming Sheng, Shuguang Wang, Houmin Wei, Chenyuan Wang, Weiqiang Lou, Xi Yang, Mofan Zhang, Kaihua Jiang, Cheng Ren, Xiaoyun Zhi, Menghan Yu, Zhe Nan, Zhuolin Zheng, Baoquan Zhong, Qinlong Wang, Huan Yu, Jinxin Chi, ...
2025
-
[46]
Zhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang, Xinwei Fu, T. S. Eugene Ng, and Yida Wang. 2023. GEMINI: Fast Failure Recovery in Distributed Training with In-Memory Checkpoints. InProceedings of the 29th ACM Symposium on Operating Systems Principles. ACM, New York, NY, USA, 3...
2023
-
[47]
Cunyang Wei and Abhinav Bhatele. 2026. Skew-aware Adaptive All-to-allv Algorithms for Dynamic Deep Learning Workloads. InProceedings of the 40th ACM International Conference on Supercomputing. ACM, Belfast, UK, 450–462. doi:10.1145/3797905.3800541
2026
-
[48]
Tianwen Wei, Bo Zhu, Liang Zhao, Cheng Cheng, Biye Li, Weiwei Lü, Peng Cheng, Jianhao Zhang, Xiaoyu Zhang, Liang Zeng, Xiaokun Wang, Yutuan Ma, Rui Hu, Shuicheng Yan, Han Fang, and Yahui Zhou. 2024. Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Langu...
2024 arXiv
-
[49]
Tianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang, Wenchao Wu, Qinkai Duan, Guodong Yang, Jiamang Wang, Lin Qu, and Liping Zhang. 2025. GREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at Scale. In2025 USENIX Annual Technical Conference (USENIX ATC 25). USENIX Association...
2025
-
[50]
Yifan Xiong, Yuting Jiang, Ziyue Yang, Lei Qu, Guoshuai Zhao, Shuguang Liu, Dong Zhong, Boris Pinzur, Jie Zhang, Yang Wang, Jithin Jose, Hossein Pourreza, Jeff Baxter, Kushal Datta, Prabhat Ram, Luke Melton, Joe Chau, Peng Cheng, Yongqiang Xiong, and Lidong Zhou. 2026. SuperBe...
2026 doi
-
[51]
Zhiyi Yao, Pengbo Hu, Congcong Miao, Xuya Jia, Zuning Liang, Yuedong Xu, Chunzhi He, Hao Lu, Mingzhuo Chen, Xiang Li, Zekun He, Yachen Wang, Xi- anneng Zou, and Junchen Jiang. 2025. Holmes: Localizing Irregularities in LLM Training with Mega-Scale GPU Clusters. In22nd USENIX S...
2025
-
[52]
Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Peng Zhang, Yuxiao Dong, and Jie Tang. 2023. GLM-130B: An Open Bilingual Pre-Trained Model. arXiv...
2023 arXiv
-
[53]
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2...
2022 arXiv
-
[54]
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews, and Shen Li. 2023. PyTorch FSDP: Experiences...
2023
-
[55]
Wenxin Zheng, Wenxiao Wang, Yun Zhang, Mingcong Han, Bin Xu, Jinyu Gu, Xingda Wei, Haibo Chen, Zuquan Song, Gaohong Liu, Yucheng Nie, Zhe Nan, Zhuolin Zheng, Huan Yu, Shuguang Wang, Ziming Zhou, Hang Zhu, Wencong Xiao, and Xin Liu. 2026. SDCs in the Wild: Characterizing and Di...
2026
-
[56]
Ziming Zhou, Yinjie Zhao, Hang Zhu, Wenxiao Wang, Zhihao Bai, Yun Zhang, Shuguang Wang, Haibin Lin, and Peng Huang. 2026. OpGuard: Bitwise Align- ment for Precise and General Debugging of Production LLM Training. In20th USENIX Symposium on Operating Systems Design and Implemen...
2026
-
[2020]
Proceedings of the VLDB Endowment13, 12 (2020), 3005–3018
PyTorch Distributed: Experiences on Accelerating Data Parallel Training. Proceedings of the VLDB Endowment13, 12 (2020), 3005–3018. doi:10.14778/ 3415478.3415530
2020
-
[2025]
In19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25)
Training with Confidence: Catching Silent Errors in Deep Learning Train- ing with Automated Proactive Checks. In19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25). USENIX Association, Boston, MA, 313–329. https://www.usenix.org/conference/osdi25/pre...
-
[3860]
doi:10.14778/3611540.3611569
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.