On 4xH100 nodes, FP16, pin_memory, and NHWC/DALI speed up image recognition, while LoRA is faster than DPO and QLoRA for LLM tuning, and PyTorch DataLoader loses scaling beyond 2 GPUs.
https://mitsloan.mit.edu/ideas- made-to-matter/ai-has-high-data-center-energy-costs-there-are-solutions
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Profiling and optimization of multi-card GPU machine learning jobs
On 4xH100 nodes, FP16, pin_memory, and NHWC/DALI speed up image recognition, while LoRA is faster than DPO and QLoRA for LLM tuning, and PyTorch DataLoader loses scaling beyond 2 GPUs.