REVIEW 3 major objections 6 minor 2 cited by
Mixed-precision numerics in scientific applications: survey and perspectives
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This survey argues that mixed-precision numerics can deliver up to 8x speedups in compute-intensive scientific workloads by aligning algorithms with hardware that now favors low-precision arithmetic.
desk verdict A competent and useful survey of mixed-precision numerics, but the headline 8x speedup claim rests on a best-case benchmark and is not representative of the application-level evidence the survey itself reports. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The workhorse is iterative refinement: solve a linear system in low precision, compute the residual in high precision, and correct the solution until the error reaches the desired level. Around that core, the survey organizes current methods into three classes: inner low precision, where approximate operations like preconditioners or multigrid smoothers run in low precision; mixed precision with refinement, where a high-precision correction step restores accuracy; and splitting schemes such as Ozaki splitting, which represent a high-precision matrix as a sum of several low-precision matrices and emulate high-precision matrix multiplication on tensor cores. The split count, and thus the speedup, depends on the dynamic range of the matrix entries, and the survey gives a threshold: emulation beats native double-precision GEMM only when low-precision throughput exceeds FP64 throughput by roughly 40–118x depending on the matrix and the splitting variant.
What would settle it
Measure the end-to-end speedup of a production scientific application that is known to be memory-bandwidth-limited when its solver is switched from double to single precision with iterative refinement. If the speedup stays at or below 2x while the hardware's FP16-to-FP64 throughput ratio is large, the claim that mixed-precision capabilities deliver 8x speedups in scientific workloads is falsified for that class of applications.
Extended reading notes
Core claim
The paper's central claim is that mixed-precision numerics can reshape computational science by aligning algorithms with the evolving hardware capability landscape. The authors support this by reviewing applications that have already adopted mixed-precision strategies and reporting the speedups they achieved, and by surveying algorithmic techniques—iterative refinement, splitting and emulation schemes, and adaptive precision solvers—that let a low-precision computation be corrected to full-precision accuracy. On a benchmark that isolates compute-bound dense factorization, the paper reports speedups of 9.50x and 8.31x over double-precision Linpack on the tested systems, and the survey argues that similar gains are available to production codes whose dominant motifs are dense matrix operations. For memory-bandwidth-limited applications, the paper itself sets the maximum speedup at 2x from double to single precision, since the gain comes only from moving fewer bits.
Load-bearing premise
The argument rests on the expectation that the throughput gap between low-precision and double-precision arithmetic on future hardware will continue to widen, and that benchmark speedups like the 8.31x figure carry over to production scientific codes; the survey itself notes that many applications are memory-bandwidth-limited, where the maximum speedup from double to single precision is only 2x.
Editorial extensions
If this is right
- Compute-bound applications built on dense matrix multiplication can expect large speedups from mixed-precision LU factorization with iterative refinement, with reported values ranging from 3x to 9.5x on current GPU systems.
- Memory-bandwidth-bound applications will see far smaller gains, at most about 2x from double to single precision, so the headline 8x claim does not extend to them.
- As the low-precision to FP64 throughput ratio grows, emulating double precision with splitting schemes becomes competitive and eventually faster than native FP64 GEMM, as already demonstrated on one current platform.
- Production scientific packages that have not yet adopted mixed-precision solvers, especially in CFD and quantum chemistry, can gain time-to-solution and energy savings by using existing libraries with mixed-precision iterative refinement.
- The survey recommends co-design among domain scientists, numerical analysts, and computer scientists because the right precision choice is domain- and problem-specific.
Reading between the lines
- The paper's 8x figure comes from an idealized compute-bound benchmark; a careful reader should treat it as an upper bound, because most real scientific applications are at least partly memory-bound and therefore capped near the 2x bit-width ratio.
- If hardware vendors continue shifting silicon toward low-precision tensor cores, the default scientific computing stack may eventually emulate FP64 arithmetic in software on low-precision units, making today's specialized splitting libraries a general-purpose foundation.
- A testable extension of the survey's argument: instrument a production implicit CFD solver with adaptive-precision sparse matrix-vector products inside GMRES, and measure whether speedups in the 1.1x–6x range reported for benchmark matrices appear in the full application.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a survey of mixed-precision numerics in scientific computing, organized around four areas: application domains (CFD, weather and climate, quantum chemistry, genomics, AI), numerical algorithms (iterative refinement, splitting and emulation schemes, Krylov solvers, multigrid, FFT, eigensolvers, ODE solvers), resource utilization, and library support. It argues that the widening throughput gap between low-precision and FP64 hardware, together with recent algorithmic advances, makes mixed-precision methods an important opportunity for scientific computation. The abstract claims performance improvements of 8x over double precision in extreme compute-intensive workloads and concludes that mixed-precision numerics can reshape computational science.
Significance. The survey is timely, broad, and useful as a roadmap: it consolidates a large and scattered literature, proposes a clear taxonomy of mixed-precision algorithm classes (low precision, MxP-inL, MxP-R, splitting), and maps applications to computational motifs and libraries. The authors are also generally transparent about which figures are measured and which are estimated, for example by flagging assumptions in Table 2. The main weakness is that the headline 8x speedup claim is anchored in the deliberately favorable HPL-MxP benchmark, while the application-level evidence collected in the same manuscript shows speedups mostly in the 1.1x-2x range; this tension needs to be addressed before the survey's central promise can be accepted as stated.
major comments (3)
- [Abstract and §3.2] The abstract's opening claim of '8x compared to double-precision in extreme compute-intensive workloads' is anchored in the HPL-MxP result of 8.31x on Frontier [71]. As the paper itself explains, HPL-MxP uses a strictly diagonal-dominant matrix, which removes pivoting and minimizes iterative-refinement iterations, and the survey's own Table 2 reports representative application speedups of 1.44x-4.80x while §1 notes that memory-bandwidth-limited applications cap at 2x when moving from FP64 to FP32. The 8x figure is therefore an extreme upper bound, not a typical outcome of the surveyed production codes. Since the concluding claim that mixed precision can 'reshape computational science' rests on broad applicability, the authors should either restrict the abstract to 'dense, compute-bound factorization workloads' or provide evidence on the fraction of scientific workloads that are compute-bound dense linear algebra.
- [§3.1.2 and §4] Several load-bearing quantitative claims are attributed to private communications: the Blackwell HPL 1.8x performance-per-watt gain and the 4-34x GEMM energy-efficiency figures in §4, and the LSMS 'acceptable accuracy' statement in §2.3. In addition, the description of Ozaki II as 'ground-breaking' is the authors' reading of arXiv preprint [68], and the manuscript's own hedge ('if reproduced across different real-world matrices') indicates that this result has not yet been independently confirmed. For a survey that aims to guide adoption, these claims should be explicitly labeled as unverified and should not be used in support of the forward-looking hardware argument in §7 unless a public source is available.
- [Table 2 and §2.6] The asterisked speedups in Table 2 are labeled as 'reasonable assumptions by the current authors,' but the table is captioned 'Representative speedups obtained from mixed-precision methods.' The derivation of these entries is not transparent: for example, the 1.44x aerodynamics figure comes from multiplying the 1.2x speedup reported by Walden et al. by an assumed further 1.2x, as described in §2.1. Because the table is used to support the survey's assessment that application-level gains are modest, each assumed entry should either be replaced with a directly reported speedup or accompanied by a footnote showing the calculation, so that readers can distinguish measured results from the authors' estimates.
minor comments (6)
- [Table 1] The 'Range' column entries such as '±10 ±308' appear to be missing superscripts; format as ±10^±308 and ±10^±38.
- [§2.5] The sentence ending in 'inverse problems.[]' contains an empty citation that should be filled or removed.
- [§2.3] The phrase 'preliminary results indicate that the results have acceptable accuracy [private communication]' should be marked as unpublished; as written, the bracket may be mistaken for a reference.
- [§2.2] The phrase 'they make a second-hand claim of speedups approaching 40%' should name the original source instead of describing it as second-hand.
- [§2.2 and §3.8] Equation (11) is referenced in §2.2 before it is defined in §3.8; add a forward reference or renumber the equations.
- [§3.10] The phrase 'these results should be taken with a pinch of salt' is too informal for a journal; rephrase as 'these results are preliminary and should be interpreted with caution.'
Circularity Check
Survey reports externally measured benchmark speedups; no derivation reduces to its inputs.
full rationale
This is a survey rather than a derivation, so the circularity patterns that require an equation-level reduction do not arise. The strongest claim, an '8x' speedup in extreme compute-intensive workloads, is anchored to the HPL-MxP benchmark result of 8.31x on Frontier reported in [71]. Although [71] shares authors with this survey, the result is a measured benchmark on a concrete system (Frontier) and is not defined in terms of any quantity the survey itself constructs; the survey also states the benchmark matrix is strictly diagonal-dominant precisely to remove pivoting and minimize refinement iterations, which is an explicit limitation rather than a circular justification. The HPG-MxP result [83] cited for GMRES-IR speedups is likewise an independently reproducible benchmark measurement. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no ansatz is smuggled via citation: the algorithmic material (iterative refinement, Ozaki splitting, multigrid, etc.) is described from the cited numerical-analysis literature with error analyses in the original sources. Table 2's speedups are explicitly marked with asterisks where the current authors made assumptions, and those assumptions are estimates from the cited studies, not outputs of a fitted model that is then re-labeled as a finding. Accordingly, the survey's conclusions are self-contained external evidence, and any concern about the breadth of the 8x claim is a matter of representativeness, not circularity.
Assumptions & free parameters
assumptions (3)
- standard math IEEE 754 floating-point formats and roundoff analysis as background (Table 1).
- domain assumption Hardware trend of widening FP16/FP8 vs FP64 throughput (Fig. 1) is extrapolated into the future.
- ad hoc to paper Private communications from hardware vendors and application teams are treated as credible evidence.
Cite this review
Pith. "Pith review of Mixed-precision numerics in scientific applications: survey and perspectives." pith.science (2026). https://pith.science/paper/PF5LJ4ZC
@misc{pith2026241219322,
author = {Pith},
title = {Pith review of: Mixed-precision numerics in scientific applications: survey and perspectives},
year = {2026},
howpublished = {\url{https://pith.science/paper/PF5LJ4ZC}},
note = {Machine review of arXiv:2412.19322}
}
read the original abstract
The explosive demand for artificial intelligence (AI) workloads has led to a significant increase in silicon area dedicated to lower-precision computations on recent high-performance computing hardware designs. However, mixed-precision capabilities, which can achieve performance improvements of 8x compared to double-precision in extreme compute-intensive workloads, remain largely untapped in most scientific applications. A growing number of efforts have shown that mixed-precision algorithmic innovations can deliver superior performance without sacrificing accuracy. These developments should prompt computational scientists to seriously consider whether their scientific modeling and simulation applications could benefit from the acceleration offered by new hardware and mixed-precision algorithms. In this survey, we (1) review progress across diverse scientific domains -- including fluid dynamics, weather and climate, quantum chemistry, and computational genomics -- that have begun adopting mixed-precision strategies; (2) examine state-of-the-art algorithmic techniques such as iterative refinement, splitting and emulation schemes, and adaptive precision solvers; (3) assess their implications for accuracy, performance, and resource utilization; and (4) survey the emerging software ecosystem that enables mixed-precision methods at scale. We conclude with perspectives and recommendations on cross-cutting opportunities, domain-specific challenges, and the role of co-design between application scientists, numerical analysts and computer scientists. Collectively, this survey underscores that mixed-precision numerics can reshape computational science by aligning algorithms with the evolving landscape of hardware capabilities.
Forward citations
Cited by 2 Pith papers
-
Scaling the memory wall using mixed-precision -- HPG-MxP on an exascale machine
An optimized GPU implementation of the HPG-MxP benchmark achieves a 1.6x speedup with mixed single-double precision GMRES on Frontier, with a full-system run at 17.23 petaflops.
-
Data Readiness for Scientific AI at Scale
Scientific data can be graded on a five-level readiness scale crossed with five processing stages, yielding a maturity matrix for AI training at supercomputer scale.
Reference graph
Works this paper leans on
-
[71]
Lu, H., Matheson, M., Oles, V., Ellis, A., Joubert, W., Wang, F.: Climbing the summit and pushing the frontier of mixed precision benchmarks at extreme scale. In: Proceed- ings of the International Conference on High Performance Computing, Networking, Storage and Analysis. SC ’22. IEEE Press, ??? (2022). https://doi.org/10.5555/3571885. 3571988
-
[68]
https://arxiv.org/abs/2504.08009v3
Ozaki, K., Uchino, Y., Imamura, T.: Ozaki Scheme II: A GEMM-oriented emulation of floating-point matrix multiplication using an integer modular technique (2025). https://arxiv.org/abs/2504.08009v3
arXiv 2025
-
[1]
Dover 46 Publications, New York (1986)
Hamming, R.W.: Numerical Methods for Scientists and Engineers, 2nd edn. Dover 46 Publications, New York (1986)
1986
-
[2]
ACM Computing Surveys (CSUR)23(1), 5–48 https://doi.org/10.1145/103162
Goldberg, D.: What every computer scientist should know about floating-point arith- metic. ACM Computing Surveys (CSUR)23(1), 5–48 https://doi.org/10.1145/103162. 103163
-
[3]
Physical Review E106(1), 015308 (2022)
Lehmann, M., Krause, M.J., Amati, G., Sega, M., Harting, J., Gekle, S.: Accuracy and performance of the lattice boltzmann method with 64-bit, 32-bit, and customized 16-bit number formats. Physical Review E106(1), 015308 (2022)
2022
-
[4]
Proceedings of the IEEE105(12), 2295–2329 (2017)
Sze, V., Chen, Y.-H., Yang, T.-J., Emer, J.S.: Efficient processing of deep neural networks: A tutorial and survey. Proceedings of the IEEE105(12), 2295–2329 (2017)
2017
-
[5]
In: Proceedings of the International Conference on High Performance Computing in Asia-Pacific Region, pp
Sakamoto, R., Kondo, M., Fujita, K., Ichimura, T., Nakajima, K.: The effectiveness of low-precision floating arithmetic on numerical codes: a case study on power consump- tion. In: Proceedings of the International Conference on High Performance Computing in Asia-Pacific Region, pp. 199–206 (2020)
2020
-
[6]
Computing in Science & Engineering 24(4), 12–22 (2022) https://doi.org/10.1109/MCSE.2022.3215477
Ltaief, H., Genton, M.G., Gratadour, D., Keyes, D.E., Ravasi, M.: Responsibly reckless matrix algorithms for HPC scientific applications. Computing in Science & Engineering 24(4), 12–22 (2022) https://doi.org/10.1109/MCSE.2022.3215477
arXiv 2022
Show all 132 references
-
[7]
https://www.nextplatform.com/2024/06/03/ amd-previews-turin-epyc-cpus-expands-instinct-gpu-roadmap/ Accessed 2024-10-30
Morgan, T.P.: AMD Previews “Turin” Epyc CPUs, Expands Instinct GPU roadmap. https://www.nextplatform.com/2024/06/03/ amd-previews-turin-epyc-cpus-expands-instinct-gpu-roadmap/ Accessed 2024-10-30
2024
-
[8]
https://www.nextplatform.com/2024/06/02/ nvidia-unfolds-gpu-interconnect-roadmaps-out-to-2027/ Accessed 2024-10-30
Morgan, T.P.: Nvidia Unfolds GPU, Interconnect Roadmaps Out To 2027. https://www.nextplatform.com/2024/06/02/ nvidia-unfolds-gpu-interconnect-roadmaps-out-to-2027/ Accessed 2024-10-30
2027
-
[9]
https://arxiv.org/ abs/2411.12090
Dongarra, J., Gunnels, J., Bayraktar, H., Haidar, A., Ernst, D.: Hardware Trends Impact- ing Floating-Point Computations In Scientific Applications (2024). https://arxiv.org/ abs/2411.12090
2024 arXiv
-
[10]
Cook, J.D.: What Is Bfloat16? https://www.johndcook.com/blog/2018/11/15/bfloat16/ Accessed 2024-09-05
2018
-
[11]
arXiv preprint arXiv:1905.12322 (2019)
Kalamkar, D., Mudigere, D., Mellempudi, N., Das, D., Banerjee, K., Avancha, S., Vooturi, D.T., Jammalamadaka, N., Huang, J., Yuen, H., et al.: A study of bfloat16 for deep learning training. arXiv preprint arXiv:1905.12322 (2019)
2019 arXiv
-
[12]
Kharya, P.: What Is the TensorFloat-32 Precision Format? https://blogs.nvidia.com/ blog/tensorfloat-32-precision-format/ Accessed 2024-09-05
2024
-
[13]
https://docs.nvidia.com/deeplearning/ transformer-engine/user-guide/examples/fp8 primer.html Accessed 2024-09-05
NVIDIA: Using FP8 with Transformer Engine. https://docs.nvidia.com/deeplearning/ transformer-engine/user-guide/examples/fp8 primer.html Accessed 2024-09-05
2024
-
[14]
Technical Report LLNL-TR-825909, Lawrence Livermore National Lab
Abdelfattah, A., Anzt, H., Ayala, A., Boman, E., Carson, E., Cayrols, S., Cojean, T., 47 Dongarra, J., Falgout, R., Gates, M., Gruetzmacher, T., Higham, N., Kruger, S., Li, X., Lindquist, N., Liu, Y., Loe, J., Luszczek, P., Nayak, P., Osei-Kuffuor, D., Pranesh, S., Rajamanicka...
2021
-
[15]
The International Journal of High Performance Computing Applications 35(4), 344–369 (2021) https://doi.org/10.1177/10943420211003313
Abdelfattah, A., Anzt, H., Boman, E.G., Carson, E., Cojean, T., Dongarra, J., Fox, A., Gates, M., Higham, N.J., Li, X.S., Loe, J., Luszczek, P., Pranesh, S., Rajamanickam, S., Ribizel, T., Smith, B.F., Swirydowicz, K., Thomas, S., Tomov, S., Tsai, Y.M., Yang, U.M.: A survey of...
2021 doi
-
[16]
Technical Report LLNL-SR-861087, Lawrence Livermore National Lab
Anzt, H.: xSDK-multiprecision final report for subcontract partner KIT. Technical Report LLNL-SR-861087, Lawrence Livermore National Lab. (LLNL), Livermore, CA (United States) (2024)
2024
-
[17]
Acta Numerica31, 347–414 (2022) https://doi.org/10.1017/S0962492922000022
Higham, N.J., Mary, T.: Mixed precision algorithms in numerical linear algebra. Acta Numerica31, 347–414 (2022) https://doi.org/10.1017/S0962492922000022
2022 doi
-
[18]
In: Bhatele, A., Hammond, J., Baboulin, M., Kruse, C
Budiardja, R.D., Berrill, M., Eisenbach, M., Jansen, G.R., Joubert, W., Nichols, S., Rogers, D.M., Tharrington, A., Bronson Messer, O.E.: Ready for the frontier: Prepar- ing applications for the world’s first exascale system. In: Bhatele, A., Hammond, J., Baboulin, M., Kruse, ...
2023 doi
-
[19]
Future Generation Computer Systems152, 1–16 (2024)
Brogi, F., Bn `a, S., Boga, G., Amati, G., Ongaro, T.E., Cerminara, M.: On floating point precision in computational fluid dynamics using openfoam. Future Generation Computer Systems152, 1–16 (2024)
2024
-
[20]
Physical Review Letters56(14), 1505–1508 (1986) https://doi.org/10.1103/ PhysRevLett.56.1505
Frisch, U., Hasslacher, B., Pomeau, Y.: Lattice-gas automata for the Navier-Stokes equation. Physical Review Letters56(14), 1505–1508 (1986) https://doi.org/10.1103/ PhysRevLett.56.1505
1986
-
[21]
Computational Geoscience 25, 871–895 (2021) https://doi.org/10.1007/s10596-020-10028-9
McClure, J.E., Li, Z., Berrill, M., Ramstad, T.: The LBPM software package for sim- ulating multiphase flow on digital images of porous rocks. Computational Geoscience 25, 871–895 (2021) https://doi.org/10.1007/s10596-020-10028-9
2021 doi
-
[22]
In: 2019 IEEE/ACM 9th Workshop on Irregular Applications: Architectures and Algorithms (IA3), pp
Walden, A., Nielsen, E., Diskin, B., Zubair, M.: A mixed precision multicolor point- implicit solver for unstructured grids on GPUs. In: 2019 IEEE/ACM 9th Workshop on Irregular Applications: Architectures and Algorithms (IA3), pp. 23–30 (2019). https://doi.org/10.1109/IA349570...
2019
-
[23]
Parallel Computing27(4), 337–362 (2001) https://doi.org/10.1016/ S0167-8191(00)00075-2
Gropp, W.D., Kaushik, D.K., Keyes, D.E., Smith, B.F.: High-performance paral- lel implicit cfd. Parallel Computing27(4), 337–362 (2001) https://doi.org/10.1016/ S0167-8191(00)00075-2 . Parallel computing in aerospace 48
2001
-
[24]
The Astrophysical Journal Supplement Series217(2), 24 (2015) https://doi.org/10.1088/0067-0049/217/2/24
Schneider, E.E., Robertson, B.E.: Cholla: A new massively parallel hydrodynamics code for astrophysical simulation. The Astrophysical Journal Supplement Series217(2), 24 (2015) https://doi.org/10.1088/0067-0049/217/2/24
2015 doi
-
[25]
Communications on Applied Mathematics and Computation5(1), 97– 115 (2021) https://doi.org/10.1007/s42967-021-00129-2
Field, S.E., Gottlieb, S., Grant, Z.J., Isherwood, L.F., Khanna, G.: A GPU-accelerated mixed-precision WENO method for extremal black hole and gravitational wave physics computations. Communications on Applied Mathematics and Computation5(1), 97– 115 (2021) https://doi.org/10....
2021 doi
-
[27]
https://arxiv.org/abs/2506.05150
Karp, M., Stanly, R., Mukha, T., Galimberti, L., Toosi, S., Song, H., Dalcin, L., Reza- eiravesh, S., Jansson, N., Markidis, S., Parsani, M., Bose, S., Lele, S., Schlatter, P.: Effects of lower floating-point precision on scale-resolving numerical simulations of turbulence (20...
2025
-
[28]
https://arxiv.org/abs/2505.07392
Wilfong, B., Radhakrishnan, A., Berre, H.L., Tselepidis, N., Dorschner, B., Budiardja, R., Cornille, B., Abbott, S., Sch ¨afer, F., Bryngelson, S.H.: Simulating many-engine spacecraft: Exceeding 100 trillion grid points via information geometric regularization and the MFC flow...
2025
-
[29]
Quarterly Journal of the Royal Meteorological Society146(729), 1590–1607 (2020) https://doi.org/10.1002/qj
Saffin, L., Hatfield, S., D¨ uben, P., Palmer, T.: Reduced-precision parametrization: lessons from an intermediate-complexity atmospheric model. Quarterly Journal of the Royal Meteorological Society146(729), 1590–1607 (2020) https://doi.org/10.1002/qj. 3754
2020 doi
-
[30]
Quarterly Journal of the Royal Meteorological Society147(741), 4358–4370 (2021) https://doi.org/10.1002/qj.4181
Lang, S.T.K., Dawson, A., Diamantakis, M., Dueben, P., Hatfield, S., Leutbecher, M., Palmer, T., Prates, F., Roberts, C.D., Sandu, I., Wedi, N.: More accuracy with less precision. Quarterly Journal of the Royal Meteorological Society147(741), 4358–4370 (2021) https://doi.org/1...
2021 doi
-
[31]
Journal of Advances in Modeling Earth Systems 14(9) (2022) https://doi.org/10.1029/2022MS003148
Ackmann, J., Dueben, P.D., Palmer, T., Smolarkiewicz, P.K.: Mixed-precision for linear solvers in global geophysical flows. Journal of Advances in Modeling Earth Systems 14(9) (2022) https://doi.org/10.1029/2022MS003148
2022 doi
-
[32]
Society for Industrial and Applied Mathematics, ??? (2003)
Saad, Y.: Iterative Methods for Sparse Linear Systems, 2nd edn. Society for Industrial and Applied Mathematics, ??? (2003). https://doi.org/10.1137/1.9780898718003
2003 doi
-
[33]
SIAM Journal on Scientific Computing45(1), 1–19 (2023) https://doi.org/10.1137/21M1465032 49
Fasi, M., Higham, N.J., Lopez, F., Mary, T., Mikaitis, M.: Matrix multiplication in multiword arithmetic: Error analysis and application to GPU tensor cores. SIAM Journal on Scientific Computing45(1), 1–19 (2023) https://doi.org/10.1137/21M1465032 49
2023 doi
-
[34]
Journal of Advances in Modeling Earth Systems14(2), 2021–002684 (2022) https://doi.org/10.1029/2021MS002684
Kl ¨ower, M., Hatfield, S., Croci, M., D¨ uben, P.D., Palmer, T.N.: Fluid simulations accel- erated with 16 bits: Approaching 4x speedup on A64FX by squeezing ShallowWaters.jl into float16. Journal of Advances in Modeling Earth Systems14(2), 2021–002684 (2022) https://doi.org/...
2022 doi
-
[35]
SIAM Journal on Scientific Computing14(4), 783–799 (1993)
Higham, N.J.: The accuracy of floating point summation. SIAM Journal on Scientific Computing14(4), 783–799 (1993)
1993
-
[36]
Journal of Chemical Theory and Computation 9(1), 213–221 (2013) https://doi.org/10.1021/ct300321a
Titov, A.V., Ufimtsev, I.S., Luehr, N., Martinez, T.J.: Generating efficient quantum chemistry codes for novel architectures. Journal of Chemical Theory and Computation 9(1), 213–221 (2013) https://doi.org/10.1021/ct300321a
2013 doi
-
[37]
In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
Das, S., Motamarri, P., Gavini, V., Turcksin, B., Li, Y.W., Leback, B.: Fast, scalable and accurate finite-element based ab initio calculations using mixed precision computing: 46 pflops simulation of a metallic dislocation system. In: Proceedings of the International Conferen...
2019
-
[38]
In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
Das, S., Kanungo, B., Subramanian, V., Panigrahi, G., Motamarri, P., Rogers, D., Zimmerman, P., Gavini, V.: Large-scale materials modeling at quantum accuracy: Ab initio simulations of quasicrystals and interacting extended defects in metallic alloys. In: Proceedings of the In...
2023
-
[39]
Journal of Chemical Theory and Computation20(24), 10826–10837 (2024) https://doi.org/10.1021/acs.jctc.4c00938
Dawson, W., Ozaki, K., Domke, J., Nakajima, T.: Reducing numerical precision requirements in quantum chemistry calculations. Journal of Chemical Theory and Computation20(24), 10826–10837 (2024) https://doi.org/10.1021/acs.jctc.4c00938 . PMID: 39644230
2024 doi
-
[40]
Computer Physics Communica- tions211, 2–7 (2017) https://doi.org/10.1016/j.cpc.2016.07.013
Eisenbach, M., Larkin, J., Lutjens, J., Rennich, S., Rogers, J.H.: Gpu acceleration of the locally selfconsistent multiple scattering code for first principles calculation of the ground state and statistical physics of materials. Computer Physics Communica- tions211, 2–7 (2017...
2017 doi
-
[41]
In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
Malaya, N., Messer, B., Glenski, J., Georgiadou, A., Lietz, J., Gottiparthi, K., Day, M., Chen, J., Rood, J., Esclapez, L., White III, J., Jansen, G.R., Curtis, N., Nichols, S., Kurzak, J., Chalmers, N., Freitag, C., Bauman, P., Fanfarillo, A., Budiardja, R.D., Papatheodore, T...
2023
-
[42]
Journal of Chemical Theory and Computation18(12), 7260–7271 (2022) https://doi.org/10.1021/acs.jctc.2c00632
Tian, Y., Xie, Z., Luo, Z., Ma, H.: Mixed-precision implementation of the density matrix renormalization group. Journal of Chemical Theory and Computation18(12), 7260–7271 (2022) https://doi.org/10.1021/acs.jctc.2c00632
2022 doi
-
[43]
In: SC18: International Conference for High Performance Computing, Networking, Storage and Analysis, pp
Joubert, W., Weighill, D., Kainer, D., Climer, S., Justice, A., Fagnan, K., Jacobson, D.: Attacking the opioid epidemic: Determining the epistatic and pleiotropic genetic architectures for chronic pain and opioid addiction. In: SC18: International Conference for High Performan...
2018
-
[44]
Parallel Computing 84, 15–23 (2019) https://doi.org/10.1016/j.parco.2019.02.003
Joubert, W., Nance, J., Climer, S., Weighill, D., Jacobson, D.: Parallel accelerated cus- tom correlation coefficient calculations for genomics applications. Parallel Computing 84, 15–23 (2019) https://doi.org/10.1016/j.parco.2019.02.003
2019 doi
-
[45]
Parallel Computing75, 130–145 (2018)
Joubert, W., Nance, J., Weighill, D., Jacobson, D.: Parallel accelerated vector similarity calculations for genomics applications. Parallel Computing75, 130–145 (2018)
2018
-
[46]
In: SC24: International Conference for High Performance Computing, Networking, Storage and Analysis, pp
Ltaief, H., Alomairy, R., Cao, Q., Ren, J., Slim, L., Kurth, T., Dorschner, B., Bougouffa, S., Abdelkhalak, R., Keyes, D.E.: Toward capturing genetic epistasis from multivariate genome-wide association studies using mixed-precision kernel ridge regression. In: SC24: Internatio...
2024 arXiv
-
[47]
Physics of Fluids35(5) (2023)
Bhushan, S., Burgreen, G.W., Brewer, W., Dettwiller, I.D.: Assessment of neural net- work augmented Reynolds averaged Navier Stokes turbulence model in extrapolation modes. Physics of Fluids35(5) (2023)
2023
-
[48]
Meena, M.G., Liousas, D., Simin, A.D., Kashi, A., Brewer, W.H., Riley, J.J., Bruyn Kops, S.M.: Machine-Learned Closure of URANS for Stably Stratified Turbu- lence: Connecting Physical Timescales & Data Hyperparameters of Deep Time-Series Models (2024)
2024
-
[49]
Geophysical Research Letters46(11), 6069–6079 (2019) https://doi.org/10.1029/2018GL081646
Pal, A., Mahajan, S., Norman, M.R.: Using deep neural networks as cost-effective sur- rogate models for super-parameterized E3SM radiative transfer. Geophysical Research Letters46(11), 6069–6079 (2019) https://doi.org/10.1029/2018GL081646
2019 doi
-
[50]
npj Computational Materials5(1), 51 (2019) https://doi.org/ 10.1038/s41524-019-0189-9
Nyshadham, C., Rupp, M., Bekker, B., Shapeev, A.V., Mueller, T., Rosenbrock, C.W., Cs´anyi, G., Wingate, D.W., Hart, G.L.: Machine-learned multi-system surrogate models for materials prediction. npj Computational Materials5(1), 51 (2019) https://doi.org/ 10.1038/s41524-019-0189-9
2019 doi
-
[51]
In: 2021 IEEE High Performance Extreme Computing Conference (HPEC), pp
Brewer, W., Geyer, C., Kleiner, D., Horne, C.: Streaming detection and classifica- tion performance of a power9 edge supercomputer. In: 2021 IEEE High Performance Extreme Computing Conference (HPEC), pp. 1–7 (2021). IEEE
2021
-
[52]
Tu, R., White, C., Kossaifi, J., Bonev, B., Kovachki, N., Pekhimenko, G., Azizzade- nesheli, K., Anandkumar, A.: Guaranteed Approximation Bounds for Mixed-Precision 51 Neural Operators (2024)
2024
-
[53]
arXiv preprint arXiv:2404.14712 (2024)
Wang, X., Tsaris, A., Liu, S., Choi, J.-Y., Fan, M., Zhang, W., Yin, J., Ashfaq, M., Lu, D., Balaprakash, P.: ORBIT: Oak Ridge base foundation model for earth system predictability. arXiv preprint arXiv:2404.14712 (2024)
2024 arXiv
-
[54]
https://arxiv.org/abs/2405.01004
Chakravarty, A.: Deep Learning Models in Speech Recognition: Measuring GPU Energy Consumption, Impact of Noise and Model Quantization for Edge Deployment (2024). https://arxiv.org/abs/2405.01004
2024 arXiv
-
[55]
https://arxiv.org/abs/2502
Kermani, A., Zeraatkar, E., Irani, H.: Energy-Efficient Transformer Inference: Opti- mization Strategies for Time Series Classification (2025). https://arxiv.org/abs/2502. 16627
2025
-
[56]
https://netlib.org/benchmark/hpl/ Accessed 2024-08-08
Petitet, A., Whaley, R.C., Dongarra, J., Cleary, A.: HPL - A Portable Implementation of the High-Performance Linpack Benchmark for Distributed-Memory Computers. https://netlib.org/benchmark/hpl/ Accessed 2024-08-08
2024
-
[57]
Parallel Computing111, 102870 (2022) https://doi.org/10
´Swirydowicz, K., Darve, E., Jones, W., Maack, J., Regev, S., Saunders, M.A., Thomas, S.J., Peleˇs, S.: Linear solvers for power grid optimization problems: A review of GPU- accelerated linear solvers. Parallel Computing111, 102870 (2022) https://doi.org/10. 1016/j.parco.2021.102870
2022
-
[58]
SIAM Journal of Numerical Analysis10(2) (1973)
George, A.: Nested dissection of a regular finite element mesh. SIAM Journal of Numerical Analysis10(2) (1973)
1973
-
[59]
SIAM Journal of Numerical Analysis17(6) (1980)
George, A.: An automatic one-way dissection algorithm for irregular finite element problems. SIAM Journal of Numerical Analysis17(6) (1980)
1980
-
[60]
Mathematics of Computation31(138) (1977)
Brandt, A.: Multi-level adaptive solutions to boundary-value problems. Mathematics of Computation31(138) (1977)
1977
-
[61]
Moler, C.B.: Iterative refinement in floating point. J. ACM14(2), 316–321 (1967) https://doi.org/10.1145/321386.321394
1967
-
[62]
In: Proceedings of the 36th ACM International Conference on Supercomputing
Ma, Z., Wang, H., Feng, G., Zhang, C., Xie, L., He, J., Chen, S., Zhai, J.: Efficiently emulating high-bitwidth computation with low-bitwidth hardware. In: Proceedings of the 36th ACM International Conference on Supercomputing. ICS ’22. Association for Computing Machinery, New...
2022 doi
-
[63]
Numerical Algorithms59, 95–118 (2012) https://doi.org/10.1007/s11075-011-9478-1
Ozaki, K., Ogita, T., Oishi, S., Rump, S.M.: Error-free transformations of matrix multi- plication by using fast routines of matrix multiplication and its applications. Numerical Algorithms59, 95–118 (2012) https://doi.org/10.1007/s11075-011-9478-1
2012 doi
-
[64]
In: International Conference on High Performance 52 Computing, pp
Mukunoki, D., Ozaki, K., Ogita, T., Imamura, T.: DGEMM using tensor cores, and its accurate and reproducible versions. In: International Conference on High Performance 52 Computing, pp. 230–248 (2020). https://doi.org/10.1007/978-3-030-50743-5 12 . Springer
2020 doi
-
[65]
The International Journal of High Performance Computing Applications38(4), 297–313 (2024) https://doi.org/10.1177/10943420241239588
Ootomo, H., Ozaki, K., Yokota, R.: DGEMM on integer matrix multiplication unit. The International Journal of High Performance Computing Applications38(4), 297–313 (2024) https://doi.org/10.1177/10943420241239588
2024 doi
-
[66]
https://arxiv.org/abs/2409.13313
Uchino, Y., Ozaki, K., Imamura, T.: Performance Enhancement of the Ozaki Scheme on Integer Matrix Multiplication Unit (2024). https://arxiv.org/abs/2409.13313
2024 arXiv
-
[67]
https://arxiv.org/ abs/2506.11277
Abdelfattah, A., Dongarra, J., Fasi, M., Mikaitis, M., Tisseur, F.: Analysis of Floating- Point Matrix Multiplication Computed via Integer Arithmetic (2025). https://arxiv.org/ abs/2506.11277
2025 arXiv
-
[69]
SIAM Journal on Scientific Computing 46(1), 30–56 (2024) https://doi.org/10.1137/22M1522619
Graillat, S., J ´ez´equel, F., Mary, T., Molina, R.: Adaptive precision sparse matrix–vector product and its application to Krylov solvers. SIAM Journal on Scientific Computing 46(1), 30–56 (2024) https://doi.org/10.1137/22M1522619
2024 doi
-
[70]
In: SC18: International Conference for High Performance Computing, Networking, Storage and Analysis, pp
Haidar, A., Tomov, S., Dongarra, J., Higham, N.J.: Harnessing GPU tensor cores for fast fp16 arithmetic to speed up mixed-precision iterative refinement solvers. In: SC18: International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 603–613 (2...
2018
-
[72]
In: Krzhizhanovskaya, V.V., Z´avodszky, G., Lees, M.H., Dongarra, J.J., Sloot, P.M.A., Bris- sos, S., Teixeira, J
Abdelfattah, A., Tomov, S., Dongarra, J.: Investigating the benefit of fp16-enabled mixed-precision solvers for symmetric positive definite matrices using gpus. In: Krzhizhanovskaya, V.V., Z´avodszky, G., Lees, M.H., Dongarra, J.J., Sloot, P.M.A., Bris- sos, S., Teixeira, J. (...
2020 doi
-
[73]
ACM Trans
Buttari, A., Dongarra, J., Kurzak, J., Luszczek, P., Tomov, S.: Using mixed precision for sparse matrix computations to enhance the performance while achieving 64-bit accuracy. ACM Trans. Math. Softw.34(4) (2008) https://doi.org/10.1145/1377596. 1377597
2008 doi
-
[74]
ACM Transactions on Mathematical Software38(1) (2011) https://doi.org/10.1145/2049662
Davis, T.A., Hu, Y.: The University of Florida sparse matrix collection. ACM Transactions on Mathematical Software38(1) (2011) https://doi.org/10.1145/2049662. 2049663 53
2011 doi
-
[75]
ACM Transactions on Mathematical Software49(1) (2023) https://doi.org/10.1145/3582493
Amestoy, P., Buttari, A., Higham, N.J., L ’Excellent, J.-I., Mary, T., Vieubl´e, B.: Combin- ing sparse approximate factorizations with mixed-precision iterative refinement. ACM Transactions on Mathematical Software49(1) (2023) https://doi.org/10.1145/3582493
2023 doi
-
[76]
PeerJ Computer Science8(e778) (2022) https://doi.org/10.7717/peerj-cs.778
Zounon, M., Higham, N.J., Lucas, C., Tisseur, F.: Performance impact of precision reduction in sparse linear systems solvers. PeerJ Computer Science8(e778) (2022) https://doi.org/10.7717/peerj-cs.778
2022 doi
-
[77]
Loe, J.A., Glusa, C.A., Yamazaki, I., Boman, E.G., Rajamanickam, S.: A Study of Mixed Precision Strategies for GMRES on GPUs (2021)
2021
-
[78]
The International Journal of High Performance Computing Applications37(2), 82–100 (2023) https://doi.org/10
Aliaga, J.I., Anzt, H., Gr¨ utzmacher, T., Quintana-Ort´ı, E.S., Tom´as, A.E.: Compressed basis gmres on high-performance graphics processing units. The International Journal of High Performance Computing Applications37(2), 82–100 (2023) https://doi.org/10. 1177/10943420221115140
2023
-
[79]
https://arxiv.org/abs/2103.09210
Carson, E., Gergelits, T.: Mixed Precision𝑠-step Lanczos and Conjugate Gradient Algorithms (2021). https://arxiv.org/abs/2103.09210
2021 arXiv
-
[80]
working paper or preprint (2025)
Jang, Y., Jolivet, P., Mary, T.: Mixed Precision Augmented GMRES. working paper or preprint (2025). https://hal.science/hal-05163845
2025
-
[81]
In: 2022 IEEE/ACM International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS), pp
Yamazaki, I., Glusa, C., Loe, J., Luszczek, P., Rajamanickam, S., Dongarra, J.: High-performance GMRES multi-precision benchmark: Design, performance, and challenges. In: 2022 IEEE/ACM International Workshop on Performance Modeling, Benchmarking and Simulation of High Performa...
2022
-
[82]
The Inter- national Journal of High Performance Computing Applications30(1), 3–10 (2016) https://doi.org/10.1177/1094342015593158
Dongarra, J., Heroux, M.A., Luszczek, P.: High-performance conjugate-gradient benchmark: A new metric for ranking high-performance computing systems. The Inter- national Journal of High Performance Computing Applications30(1), 3–10 (2016) https://doi.org/10.1177/1094342015593158
2016 doi
-
[83]
https://arxiv.org/abs/2507.11512
Kashi, A., Koukpaizan, N., Lu, H., Matheson, M., Oral, S., Wang, F.: Scaling the memory wall using mixed-precision – HPG-MxP on an exascale machine (2025). https://arxiv.org/abs/2507.11512
2025 arXiv
-
[84]
ACM Trans
Flegar, G., Anzt, H., Cojean, T., Quintana-Ort ´ı, E.S.: Adaptive precision block-jacobi for high performance preconditioning in the ginkgo linear algebra software. ACM Trans. Math. Softw.47(2) (2021) https://doi.org/10.1145/3441850
2021 doi
-
[85]
In: Sousa, L., Roma, N., Tom´as, P
G ¨obel, F., Gr¨ utzmacher, T., Ribizel, T., Anzt, H.: Mixed precision incomplete and factorized sparse approximate inverse preconditioning on GPUs. In: Sousa, L., Roma, N., Tom´as, P. (eds.) Euro-Par 2021: Parallel Processing, pp. 550–564. Springer, Cham (2021). https://doi.o...
2021 doi
-
[86]
In: 54 Wyrzykowski, R., Dongarra, J., Deelman, E., Karczewski, K
Tsai, Y.-H.M., Beams, N., Anzt, H.: Mixed precision algebraic multigrid on GPUs. In: 54 Wyrzykowski, R., Dongarra, J., Deelman, E., Karczewski, K. (eds.) Parallel Processing and Applied Mathematics, pp. 113–125. Springer, Cham (2023). https://doi.org/10. 1007/978-3-031-30442-2 9
2023
-
[87]
Future Generation Computer Systems149, 280–293 (2023) https://doi.org/10.1016/j
Tsai, Y.-H.M., Beams, N., Anzt, H.: Three-precision algebraic multigrid on GPUs. Future Generation Computer Systems149, 280–293 (2023) https://doi.org/10.1016/j. future.2023.07.024
2023 doi
-
[88]
PhD thesis, Karlsruher Institut f¨ ur Technologie (KIT) (2024)
Tsai, Y.-H.: Portable mixed precision algebraic multigrid on high performance GPUs. PhD thesis, Karlsruher Institut f¨ ur Technologie (KIT) (2024). https://doi.org/10.5445/ IR/1000168914
2024
-
[89]
In: 2018 IEEE 25th Inter- national Conference on High Performance Computing Workshops (HiPCW), pp
Sorna, A., Cheng, X., D’ Azevedo, E., Won, K., Tomov, S.: Optimizing the fast Fourier transform using mixed precision on tensor core hardware. In: 2018 IEEE 25th Inter- national Conference on High Performance Computing Workshops (HiPCW), pp. 3–7 (2018). https://doi.org/10.1109...
2018
-
[90]
In: 2021 IEEE International Conference on Cluster Computing (CLUSTER), pp
Li, B., Cheng, S., Lin, J.: tcFFT: A fast half-precision FFT library for NVIDIA tensor cores. In: 2021 IEEE International Conference on Cluster Computing (CLUSTER), pp. 1–11 (2021). https://doi.org/10.1109/Cluster48925.2021.00035
2021
-
[91]
ACM Trans
Zhao, Y., Liu, F., Ma, W., Li, H., Peng, Y., Wang, C.: MFFT: A GPU accelerated highly efficient mixed-precision large-scale FFT framework. ACM Trans. Archit. Code Optim. 20(3) (2023) https://doi.org/10.1145/3605148
2023 doi
-
[92]
SIAM Review64(1), 191–211 (2022) https://doi.org/10.1137/20M1342902
Kelley, C.T.: Newton’s method in mixed precision. SIAM Review64(1), 191–211 (2022) https://doi.org/10.1137/20M1342902
2022 doi
-
[93]
CAMPS, D., MACH, T., V ANDEBRIL, R., WATKINS, D.S.: ON POLE-SWAPPING ALGORITHMS FOR THE EIGENV ALUE PROBLEM, vol. 52, pp. 480–508. Kent State University, ??? (2020). https://doi.org/10.1553/etna vol52s480
2020 doi
-
[94]
In: 2022 IEEE/ACM Workshop on Latest Advances in Scalable Algorithms for Large- Scale Heterogeneous Systems (ScalAH), pp
Tsai, Y.M., Luszczek, P., Dongarra, J.: Mixed-precision algorithm for finding selected eigenvalues and eigenvectors of symmetric and hermitian matrices1. In: 2022 IEEE/ACM Workshop on Latest Advances in Scalable Algorithms for Large- Scale Heterogeneous Systems (ScalAH), pp. 4...
2022
-
[95]
Japan Journal of Industrial and Applied Mathematics36(2), 699–717 (2019) https://doi.org/10.1007/s13160-019-00360-8 55
Alvermann, A., Basermann, A., Bungartz, H.-J., Carbogno, C., Ernst, D., Fehske, H., Futamura, Y., Galgon, M., Hager, G., Huber, S., Huckle, T., Ida, A., Imakura, A., Kawai, M., K¨ocher, S., Kreutzer, M., Kus, P., Lang, B., Lederer, H., Manin, V., Marek, A., Nakajima, K., Nemec...
2019
-
[96]
https://arxiv.org/abs/2503.22652
Kodali, N., Ramakrishnan, K., Motamarri, P.: Residual-based Chebyshev filtered subspace iteration for sparse Hermitian eigenvalue problems tolerant to inexact matrix-vector products (2025). https://arxiv.org/abs/2503.22652
2025
-
[97]
Grant, Z.J.: Perturbed Runge–Kutta methods for mixed precision applications. J. Sci. Comput.92(6) (2022) https://doi.org/10.1007/s10915-022-01801-2
2022 doi
-
[98]
Springer, Cham, Switzerland (2021)
Butcher, J.C.: B-series: Algebraic Analysis of Numerical Methods, 1st edn. Springer, Cham, Switzerland (2021). https://doi.org/10.1007/978-3-030-70956-3 . Springer Series in Computational Mathematics
2021 doi
-
[99]
https://arxiv.org/abs/2212.11849
Burnett, B., Gottlieb, S., Grant, Z.J.: Stability Analysis and Performance Evaluation of Mixed-Precision Runge-Kutta Methods (2022). https://arxiv.org/abs/2212.11849
2022 arXiv
-
[100]
https://arxiv.org/abs/2412.16638
Dravins, I., Koch, M., Griehl, V., Kormann, K.: Performance evaluation of mixed- precision Runge-Kutta methods for the solution of partial differential equations (2024). https://arxiv.org/abs/2412.16638
2024 arXiv
-
[101]
Journal of Computational Physics464, 111349 (2022) https://doi.org/10.1016/j.jcp.2022.111349
Croci, M., Rosilho de Souza, G.: Mixed-precision explicit stabilized Runge–Kutta methods for single- and multi-scale differential equations. Journal of Computational Physics464, 111349 (2022) https://doi.org/10.1016/j.jcp.2022.111349
2022
-
[102]
In: 2023 IEEE High Performance Extreme Computing Conference (HPEC), pp
Balos, C.J., Roberts, S., Gardner, D.J.: Leveraging mixed precision in exponential time integration methods. In: 2023 IEEE High Performance Extreme Computing Conference (HPEC), pp. 1–8 (2023). https://doi.org/10.1109/HPEC58863.2023.10363489
2023
-
[103]
Computer Science - Research and Development25(3) (2010) https://doi.org/10.1007/s00450-010-0124-2
Anzt, H., Rocker, B., Heuveline, V.: Energy efficiency of mixed precision iterative refinement methods using hybrid hardware platforms. Computer Science - Research and Development25(3) (2010) https://doi.org/10.1007/s00450-010-0124-2
2010 doi
-
[104]
In: Shi, Y., Fu, H., Tian, Y., Krzhizhanovskaya, V.V., Lees, M.H., Dongarra, J., Sloot, P.M.A
Haidar, A., Abdelfattah, A., Zounon, M., Wu, P., Pranesh, S., Tomov, S., Dongarra, J.: The design of fast and energy-efficient linear solvers: On the potential of half- precision arithmetic and iterative refinement techniques. In: Shi, Y., Fu, H., Tian, Y., Krzhizhanovskaya, V...
2018
-
[105]
SLATE Working Notes 15, ICL-UT-20- 08, Innovative Computing Laboratory, University of Tennessee, Knoxville (July 2020)
Abdelfattah, A., Anzt, H., Boman, E., Carson, E., Cojean, T., Dongarra, J., Gates, M., Gruetzmacher, T., Higham, N.J., Li, S., Lindquist, N., Liu, Y., Loe, J., Luszczek, P., Nayak, P., Pranesh, S., Rajamanickam, S., Ribizel, T., Smith, B., Swirydowicz, K., Thomas, S., Tomov, S...
2020
-
[106]
Parallel Computing36(5-6), 232–240 (2010) https: 56 //doi.org/10.1016/j.parco.2009.12.005
Tomov, S., Dongarra, J., Baboulin, M.: Towards dense linear algebra for hybrid GPU accelerated manycore systems. Parallel Computing36(5-6), 232–240 (2010) https: 56 //doi.org/10.1016/j.parco.2009.12.005
2010 doi
-
[107]
Parallel Computing36(12), 645–654 (2010) https://doi.org/10.1016/j.parco.2010.06.001
Tomov, S., Nath, R., Dongarra, J.: Accelerating the reduction to upper Hessenberg, tridiagonal, and bidiagonal forms through hybrid GPU-based computing. Parallel Computing36(12), 645–654 (2010) https://doi.org/10.1016/j.parco.2010.06.001
2010 doi
-
[108]
In: Kunkel, J.M., Ludwig, T
Chow, E., Anzt, H., Dongarra, J.: Asynchronous iterative algorithm for comput- ing incomplete factorizations on GPUs. In: Kunkel, J.M., Ludwig, T. (eds.) High Performance Computing, pp. 1–16. Springer, Cham (2015). https://doi.org/10.1007/ 978-3-319-20119-1 1
2015
-
[109]
In: Proceedings of the 8th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Systems
Haidar, A., Wu, P., Tomov, S., Dongarra, J.: Investigating half precision arithmetic to accelerate dense linear system solvers. In: Proceedings of the 8th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Systems. ScalA ’17. Association for Computing Machinery...
2017 doi
-
[110]
In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
Gates, M., Kurzak, J., Charara, A., YarKhan, A., Dongarra, J.: Slate: design of a modern distributed and accelerated linear algebra library. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. SC ’19. Association fo...
2019
-
[111]
The International Journal of High Performance Computing Applications0(0) (2024) https://doi.org/10.1177/10943420241286531
Gates, M., Abdelfattah, A., Akbudak, K., Farhan, M.A., Alomairy, R., Bielich, D., Burgess, T., Cayrols, S., Lindquist, N., Sukkari, D., YarKhan, A.: Evolution of the slate linear algebra library. The International Journal of High Performance Computing Applications0(0) (2024) h...
2024 doi
-
[112]
ACM Transactions on Mathematical Software48(1), 2–1233 (2022) https://doi.org/10.1145/3480935
Anzt, H., Cojean, T., Flegar, G., G¨obel, F., Gr¨ utzmacher, T., Nayak, P., Ribizel, T., Tsai, Y.M., Quintana-Ort´ı, E.S.: Ginkgo: A Modern Linear Operator Algebra Framework for High Performance Computing. ACM Transactions on Mathematical Software48(1), 2–1233 (2022) https://d...
2022 doi
-
[113]
Journal of Numerical Mathematics 31(3), 231–246 (2023) https://doi.org/10.1515/jnma-2023-0089
Arndt, D., Bangerth, W., Bergbauer, M., Feder, M., Fehling, M., Heinz, J., Heister, T., Heltai, L., Kronbichler, M., Maier, M., Munch, P., Pelteret, J.-P., Turcksin, B., Wells, D., Zampini, S.: Thedeal.IIlibrary, version 9.5. Journal of Numerical Mathematics 31(3), 231–246 (20...
2023 doi
-
[114]
Computers & Mathematics with Applications81, 42–74 (2021) https://doi.org/ 10.1016/j.camwa.2020.06.009
Anderson, R., Andrej, J., Barker, A., Bramwell, J., Camier, J.-S., Cerveny, J., Dobrev, V., Dudouit, Y., Fisher, A., Kolev, T., Pazner, W., Stowell, M., Tomov, V., Akkerman, I., Dahm, J., Medina, D., Zampini, S.: MFEM: A modular finite element methods library. Computers & Math...
2021 doi
-
[115]
ACM Transactions on Mathematical Software (TOMS) (2022) https://doi.org/10.1145/ 57 3539801
Gardner, D.J., Reynolds, D.R., Woodward, C.S., Balos, C.J.: Enabling new flexibil- ity in the SUNDIALS suite of nonlinear and differential/algebraic equation solvers. ACM Transactions on Mathematical Software (TOMS) (2022) https://doi.org/10.1145/ 57 3539801
2022
-
[116]
Physics of Plasmas24(5), 054508 (2017) https://doi.org/10.1063/1
Hager, R., Lang, J., Chang, C.S., Ku, S., Chen, Y., Parker, S.E., Adams, M.F.: Verifica- tion of long wavelength electromagnetic modes with a gyrokinetic-fluid hybrid model in the XGC code. Physics of Plasmas24(5), 054508 (2017) https://doi.org/10.1063/1. 4983320
2017 doi
-
[117]
(2020 (acccessed May 22, 2020))
Team, T.: The Trilinos Project Website. (2020 (acccessed May 22, 2020)). https: //trilinos.github.io
2020
-
[118]
Scientific Programming 20(3), 243875 (2012) https://doi.org/10.3233/SPR-2012-0352
Bavier, E., Hoemmen, M., Rajamanickam, S., Thornquist, H.: Amesos2 and Belos: Direct and iterative solvers for large sparse linear systems. Scientific Programming 20(3), 243875 (2012) https://doi.org/10.3233/SPR-2012-0352
2012 doi
-
[119]
Scientific Programming20(2), 693861 (2012) https://doi.org/10.3233/ SPR-2012-0349
Baker, C.G., Heroux, M.A.: Tpetra, and the use of generic programming in scien- tific computing. Scientific Programming20(2), 693861 (2012) https://doi.org/10.3233/ SPR-2012-0349
2012
-
[120]
Rajamanickam, S., Acer, S., Berger-Vergiat, L., Dang, V., Ellingwood, N., Harvey, E., Kelley, B., Trott, C.R., Wilke, J., Yamazaki, I.: Kokkos Kernels: Performance Portable Sparse/Dense Linear Algebra and Graph Kernels (2021)
2021
-
[121]
IEEE Transactions on Parallel and Distributed Systems33(4), 805–817 (2022) https://doi
Trott, C.R., Lebrun-Grandi ´e, D., Arndt, D., Ciesko, J., Dang, V., Ellingwood, N., Gayatri, R., Harvey, E., Hollman, D.S., Ibanez, D., Liber, N., Madsen, J., Miles, J., Poliakoff, D., Powell, A., Rajamanickam, S., Simberg, M., Sunderland, D., Turcksin, B., Wilke, J.: Kokkos 3...
2022
-
[122]
Computing in Science & Engineering23(5), 10–18 (2021) https://doi.org/10.1109/ MCSE.2021.3098509
Trott, C., Berger-Vergiat, L., Poliakoff, D., Rajamanickam, S., Lebrun-Grandie, D., Madsen, J., Al Awar, N., Gligoric, M., Shipman, G., Womeldorff, G.: The kokkos ecosystem: Comprehensive performance portability for high performance computing. Computing in Science & Engineerin...
2021
-
[123]
In: 2022 IEEE International Parallel and Dis- tributed Processing Symposium (IPDPS), pp
Yamazaki, I., Carson, E., Kelley, B.: Mixed precision𝑠-step conjugate gradient with residual replacement on GPUs. In: 2022 IEEE International Parallel and Dis- tributed Processing Symposium (IPDPS), pp. 886–896 (2022). https://doi.org/10.1109/ IPDPS53621.2022.00091
2022
-
[124]
In: Arge, E., Bruaset, A.M., Langtangen, H.P
Balay, S., Gropp, W.D., McInnes, L.C., Smith, B.F.: Efficienct management of par- allelism in object oriented numerical software libraries. In: Arge, E., Bruaset, A.M., Langtangen, H.P. (eds.) Modern Software Tools in Scientific Computing, pp. 163–202. Birkhauser Press, ??? (1997)
1997
-
[125]
Parallel Computing108, 102831 (2021) https://doi.org/10.1016/j.parco.2021.102831
Mills, R.T., Adams, M.F., Balay, S., Brown, J., Dener, A., Knepley, M., Kruger, S.E., Morgan, H., Munson, T., Rupp, K., Smith, B.F., Zampini, S., Zhang, H., Zhang, 58 J.: Toward performance-portable PETSc for GPU-based exascale systems. Parallel Computing108, 102831 (2021) htt...
2021
-
[126]
ACM Trans
Hernandez, V., Roman, J.E., Vidal, V.: SLEPc: A scalable and flexible toolkit for the solution of eigenvalue problems. ACM Trans. Math. Software31(3), 351–362 (2005)
2005
-
[127]
ACM Trans
Roman, J.E., Alvarruiz, F., Campos, C., Dalcin, L., Jolivet, P., Lamas Davi ˜na, A.: Improvements to SLEPc in releases 3.14–3.18. ACM Trans. Math. Software49(3), 29–12911 (2023)
2023
-
[128]
In: Sloot, P.M.A., Tan, C.J.K., Dongarra, J.J., Hoekstra, A.G
Falgout, R.D., Yang, U.M.: hypre: a library of high performance preconditioners. In: Sloot, P.M.A., Tan, C.J.K., Dongarra, J.J., Hoekstra, A.G. (eds.) Lecture Notes in Computer Science, vol. 2331, pp. 632–641. Springer, ??? (2002). UCRL-JC-146175
2002
-
[129]
ACM Transactions on Mathematical Software (TOMS)31(3), 363–396 (2005) https://doi.org/10.1145/1089014.1089020
Hindmarsh, A.C., Brown, P.N., Grant, K.E., Lee, S.L., Serban, R., Shumaker, D.E., Woodward, C.S.: SUNDIALS: Suite of nonlinear and differential/algebraic equation solvers. ACM Transactions on Mathematical Software (TOMS)31(3), 363–396 (2005) https://doi.org/10.1145/1089014.1089020
2005
-
[130]
In: Computational Science – ICCS 2020: 20th International Conference, Amsterdam, The Netherlands, June 3–5, 2020, Proceedings, Part I, pp
Ayala, A., Tomov, S., Haidar, A., Dongarra, J.: heffte: Highly efficient fft for exascale. In: Computational Science – ICCS 2020: 20th International Conference, Amsterdam, The Netherlands, June 3–5, 2020, Proceedings, Part I, pp. 262–275. Springer, Berlin, Heidelberg (2020). h...
2020 doi
-
[131]
ICL Technical Report ICL-UT-22-04, Innovative Computing Laboratory, University of Tennessee, Knoxville (2022-05 2022)
Cayrols, S., Li, J., Bosilca, G., Tomov, S., Ayala, A., Dongarra, J.: Mixed precision and approximate 3D FFTs: Speed for accuracy trade-off with GPU-aware MPI and run- time data compression. ICL Technical Report ICL-UT-22-04, Innovative Computing Laboratory, University of Tenn...
2022
-
[132]
Accessed: 2025-07-30 (2021)
Li, B.: tcFFT. Accessed: 2025-07-30 (2021). https://github.com/rox906/tcFFT
2021
-
[133]
https: //arxiv.org/abs/2507.04647 59
Hoerold, F., Ivanov, I.R., Dhruv, A., Moses, W.S., Dubey, A., Wahib, M., Domke, J.: RAPTOR: Practical Numerical Profiling of Scientific Applications (2025). https: //arxiv.org/abs/2507.04647 59
2025
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.