GPU TEE encryption and authentication on each ring all-reduce step slows four-GPU DDP training by an average of 8.68x and up to about 42x; enlarging the DDP gradient bucket to 400 MB cuts most of the overhead but leaves a 3x gap.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Characterization of GPU TEE Overheads in Distributed Data Parallel ML Training
GPU TEE encryption and authentication on each ring all-reduce step slows four-GPU DDP training by an average of 8.68x and up to about 42x; enlarging the DDP gradient bucket to 400 MB cuts most of the overhead but leaves a 3x gap.