Energy use in data-parallel neural network training grows roughly linearly with GPU hours, but the energy cost per GPU hour varies by model, hardware, and the number of samples and gradient updates per GPU hour.
On the sdes and scaling rules for adaptive gradient algorithms,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Energy Consumption in Parallel Neural Network Training
Energy use in data-parallel neural network training grows roughly linearly with GPU hours, but the energy cost per GPU hour varies by model, hardware, and the number of samples and gradient updates per GPU hour.