A parameter and activation memory model for DeepSeek-v3 training gives per-GPU footprints under PP/TP/EP and ZeRO, but with no empirical validation.
Reducing activation recomputation in large transformer models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.PF 1years
2025 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Memory Analysis on the Training Course of DeepSeek Models
A parameter and activation memory model for DeepSeek-v3 training gives per-GPU footprints under PP/TP/EP and ZeRO, but with no empirical validation.