An optimized GPU implementation of the HPG-MxP benchmark achieves a 1.6x speedup with mixed single-double precision GMRES on Frontier, with a full-system run at 17.23 petaflops.
The Design of Fast and Energy-Efficient Linear Solvers: On the Potential of Half-Precision Arithmetic and Iterative Refinement Techniques
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Scaling the memory wall using mixed-precision -- HPG-MxP on an exascale machine
An optimized GPU implementation of the HPG-MxP benchmark achieves a 1.6x speedup with mixed single-double precision GMRES on Frontier, with a full-system run at 17.23 petaflops.