GPU TEE encryption and authentication on each ring all-reduce step slows four-GPU DDP training by an average of 8.68x and up to about 42x; enlarging the DDP gradient bucket to 400 MB cuts most of the overhead but leaves a 3x gap.
Arm security technology building a secure system using trustzone technology,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Characterization of GPU TEE Overheads in Distributed Data Parallel ML Training
GPU TEE encryption and authentication on each ring all-reduce step slows four-GPU DDP training by an average of 8.68x and up to about 42x; enlarging the DDP gradient bucket to 400 MB cuts most of the overhead but leaves a 3x gap.