A production-scale TPU training stack for Google Ads models combines shared input memoization, hybrid embedding partitioning, pipelining, RPC coalescing, and preemption holds to improve training throughput by 116% and cut training cost by 18% on five representative models.
Understanding training efficiency of deep learning recommendation models at scale,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Scalable Machine Learning Training Infrastructure for Online Ads Recommendation and Auction Scoring Modeling at Google
A production-scale TPU training stack for Google Ads models combines shared input memoization, hybrid embedding partitioning, pipelining, RPC coalescing, and preemption holds to improve training throughput by 116% and cut training cost by 18% on five representative models.