Pith. sign in

REVIEW 1 cited by

Understanding Silent Data Corruption in LLM Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12340 v1 pith:HZPTJ74N submitted 2025-02-17 cs.LG cs.DC

classification cs.LGcs.DC
keywords sdcstrainingnodesimpactcomputationdifferentunhealthycorruption
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As the scale of training large language models (LLMs) increases, one emergent failure is silent data corruption (SDC), where hardware produces incorrect computations without explicit failure signals. In this work, we are the first to investigate the impact of real-world SDCs on LLM training by comparing model training between healthy production nodes and unhealthy nodes exhibiting SDCs. With the help from a cloud computing platform, we access the unhealthy nodes that were swept out from production by automated fleet management. Using deterministic execution via XLA compiler and our proposed synchronization mechanisms, we isolate and analyze the impact of SDC errors on these nodes at three levels: at each submodule computation, at a single optimizer step, and at a training period. Our results reveal that the impact of SDCs on computation varies on different unhealthy nodes. Although in most cases the perturbations from SDCs on submodule computation and gradients are relatively small, SDCs can lead models to converge to different optima with different weights and even cause spikes in the training loss. Our analysis sheds light on further understanding and mitigating the impact of SDCs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training

    cs.DC 2026-08 conditional novelty 6.0 of 10

    SCOUT localizes hangs, stragglers, and silent data corruption in LLM pre-training through strict-majority consensus among equivalent replicas, with replay-based checkpoint certification.

Pith tools