REVIEW 3 major objections 4 minor 37 references
Reduced and mixed precision turbulent flow simulations using explicit finite difference schemes
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Storing only the accumulated flow variables in full precision, while keeping transient residual and work arrays in lower precision, delivers 1.3–4x speedups and up to 4x memory savings in compressible turbulent flow simulations without…
desk verdict Solid engineering result for mixed-precision finite-difference CFD, but the 'independent of Reynolds number' claim goes beyond the single Re=800 test. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is the structure of the explicit Runge-Kutta update, written in low-storage form as $\tilde{Q}_i = A_i \tilde{Q}_{i-1} + \Delta t R_{i-1}$ and $Q_i = Q_{i-1} + B_i \tilde{Q}_i$. Here $Q$ is the vector of conservative variables that accumulates over substeps and time levels, $R$ is the residual containing all flux and source terms, and $W$ denotes the many work arrays holding intermediate derivatives. Because $Q$ is the only quantity that integrates errors over thousands of steps, it must be kept in high precision, while the transient $R$ and $W$ arrays, which are recomputed every step, can tolerate lower precision. The paper encodes this allocation as a user-facing configuration in the OpenSBLI code generator, which then emits OPS kernels with per-dataset precision settings and explicit casting.
What would settle it
Run the same half-single or single-double configuration on a shock-containing flow, such as a compressible shock-tube or a wall-bounded turbulent channel, and compare the dissipation and mean profiles against double precision; if the mixed-precision error grows beyond the level of the higher component precision, or if the simulation becomes unstable, the claim of accuracy without compromise would be refuted. Alternatively, extend the Taylor-Green run past t=20 and check whether the absolute error in dissipation remains bounded at the mixed-precision level.
Extended reading notes
Core claim
The central discovery is that in explicit low-storage Runge-Kutta time advancement of the compressible Navier-Stokes equations, the arrays that need full precision are only the conservative variables Q that accumulate over time. The residuals R and work arrays W, which hold intermediate derivatives and are discarded each step, can be stored in a lower precision without affecting the long-time accuracy of the solution. For the Taylor-Green vortex at Re=800, M=0.5 on a $256^{3}$ grid, the paper shows that half-single (HPSP) and single-double (SPDP) mixtures produce solenoidal dissipation histories that overlap the double-precision reference, with absolute errors close to the higher of the two component precisions. Pure FP16, by contrast, produces clearly wrong flow states. The findings hold across variation of timestep and Mach number, and across different split-form formulations of the convection terms, with only minor differences between double and half-single precision.
Load-bearing premise
The conclusion that lowering the precision of residuals and work arrays does not compromise accuracy rests on assuming that only the accumulated conservative variables need full precision over long integration times, an assumption validated only on the smooth Taylor-Green vortex up to t=20 and not on flows with shocks, walls, or longer time horizons.
Editorial extensions
If this is right
- Mixed precision configurations like SPDP and HPSP can be used as drop-in replacements for full double precision in explicit finite-difference compressible solvers, yielding 1.3–4x speedups and up to 4x memory savings on GPU and CPU systems.
- The reduced memory footprint lets larger grid sizes fit on the same hardware, effectively enabling higher-resolution simulations or larger multi-block cases within fixed resources.
- Communication volume in MPI-based multi-GPU runs shrinks proportionally to the lowered precision of the communicated arrays, improving strong and weak scaling at scale.
- The accuracy benefit of split-form convective formulations such as KEEP and KGP is preserved under mixed precision, so numerical robustness of the discretization transfers to reduced-precision runs.
- Pure half precision is not viable for this class of problems, so precision selection must be deliberate; the framework's per-dataset control makes such selection practical.
Reading between the lines
- The same accumulator-in-high-precision, transients-in-low-precision principle should apply to other explicit time-marching schemes beyond Runge-Kutta, such as Adams-Bashforth or MacCormack-type methods, where the same one-way error flow from transient to accumulated state exists; this is a testable extension the paper does not make.
- The validation is confined to smooth, subsonic turbulence with periodic boundaries; shock-capturing schemes and wall-bounded flows introduce stencils and limiters that may amplify low-precision rounding, so the portability of the 1.3–4x speedup to those regimes is an open question.
- Adaptive precision, lowering precision only in regions where the local error budget allows it, analogous to adaptive mesh refinement, could extend the speedup beyond the uniform mixed-precision cases studied here, and the framework's per-dataset presets are a step toward that.
- The error accumulation argument is asymptotic; over very long integration times far beyond t=20, the difference between mixed and full precision may drift, so the 'without compromising accuracy' claim should be read as scoped to typical simulation horizons.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes an extension of the OPS library and the OpenSBLI code-generation framework to support mixed-precision storage for explicit finite-difference compressible flow solvers. The proposed strategy is to store the conservative variables Q and the low-storage Runge-Kutta arrays in a higher precision, while storing the residual arrays R and work arrays W in a lower precision. Accuracy is assessed on the compressible Taylor-Green vortex (TGSym) at Re=800, M=0.5, on a 256^3 mesh, comparing FP64, FP32, FP16, and mixed HPSP and SPDP configurations; only pure FP16 shows large errors. Performance measurements on an NVIDIA A100, an AMD MI250X, and Intel and AMD CPUs report speedups of roughly 1.3x to 4x, with memory savings up to 4x. The paper also compares several split-form convective schemes in the inviscid limit, finding that the mixed-precision HPSP results track the double-precision results closely.
Significance. The software contribution is concrete and useful: extending OPS/OpenSBLI with a simple interface for per-dataset precision allocation makes mixed precision practical for users of these frameworks, and the portability results across GPU and CPU platforms are valuable. The accuracy data for the Re=800 Taylor-Green vortex are convincing, and the performance measurements quantify the expected memory-bandwidth and memory-capacity benefits. The main limitation is the generality of the central claim: the paper explicitly states that accuracy is independent of Reynolds number, but all viscous simulations use Re=800. In the non-dimensional governing equations, the viscous contribution to the residual scales as 1/Re, so at higher Reynolds numbers the half-precision rounding error in the residual and work arrays can become larger than the physical viscous signal that the dissipation measure is intended to capture. Because this is a plausible failure mode that is not tested or analyzed, the claim 'without compromising numerical accuracy' is established only in the tested regime.
major comments (3)
- [Section 4.2, concluding paragraph; Section 1 bullet list] The paper claims that the accuracy of the mixed-precision results is independent of Reynolds number, but all viscous TGV simulations are performed at Re=800. In the non-dimensional equations used here, the viscous flux divergence in the residual is proportional to 1/Re while the inviscid part is order one; at Re=800, 1/Re = 1.25e-3 is already comparable to the FP16 unit roundoff of about 5e-4, and at higher Re the half-precision rounding error in R and W would exceed the viscous signal that controls the dissipation measure ϵ_S defined in Eq. (12). A Reynolds-number sweep (e.g., Re=1600 and Re=8000) or a quantitative roundoff analysis is needed to support the claimed generality.
- [Section 4.2, Figure 8] The inviscid split-form test omits viscous terms entirely, so it cannot expose a failure mode in which low-precision residuals overwhelm the physical viscous contribution at high Reynolds number. The agreement between DP and HPSP in the inviscid limit demonstrates split-form robustness for the convective terms, but it does not provide evidence about the Reynolds-number scaling of precision-induced error in the viscous residual.
- [Section 3.1, Eqs. (3)-(4)] The argument that R and W need not be stored in high precision because they are discarded at each step overlooks that the rounding error in R is multiplied by the timestep and added to Q at every substep, so it can accumulate over the 8000 iterations used here. The experimental results at Re=800 are reassuring, but a bound on the accumulated roundoff error would be needed to justify the algorithm for longer integration times and for regimes where the physical dissipation is weaker.
minor comments (4)
- [Section 4.3, Tables 1-5] Performance values are reported as single numbers with no variance; since each case was run five times, reporting the standard deviation or range would improve confidence in the speedup comparisons.
- [Section 4.1, Figure 2 caption] The caption contains the typo 'mutualy' and should read 'mutually perpendicular slices'.
- [Section 4.1, Figure 3] The vertical axis label for the dissipation curve is shown as 'S' but should be 'ϵ_S' to match Eq. (12).
- [General] There is no data or code availability statement; for a reproducibility-oriented paper in a computational journal, a link to the OPS/OpenSBLI code or to the TGSym setup used would be helpful.
Circularity Check
No significant circularity: the mixed-precision accuracy claim is validated against an external double-precision Taylor-Green reference, not derived from fitted inputs or self-cited premises.
full rationale
The paper's central claim is an empirical one: storing conservative variables Q in higher precision while storing residuals R and work arrays W in lower precision preserves accuracy relative to a double-precision run. This is tested directly on the compressible Taylor-Green vortex benchmark by comparing dissipation and kinetic energy against a fully double-precision simulation, which is an external reference rather than a quantity fitted by the method. The precision configurations are discrete design choices, not fitted parameters, and no quantity is predicted from data that was used to define the method. The argument that R and W do not accumulate long-term error because they are re-formed each Runge-Kutta substep is a structural observation about the algorithm, not a circular definition of the accuracy result; it is then checked by the experiments in Section 4.2. Self-citations appear in descriptions of the OpenSBLI and OPS frameworks (e.g., refs. [9], [23], [33]) and in setting up the Taylor-Green test case, but these support implementation details and benchmark provenance, not the mixed-precision accuracy conclusion itself. The paper does make a broader claim that accuracy is independent of Reynolds number and Mach number, while the only full Navier-Stokes Reynolds number tested is Re=800 (plus an inviscid case); this is a scope-of-evidence limitation rather than circular reasoning. Overall, no step in the derivation chain reduces to its own inputs, so the circularity burden is very low.
Assumptions & free parameters
assumptions (3)
- standard math IEEE 754 floating-point arithmetic behaves as specified for FP16, FP32, and FP64.
- domain assumption Taylor-Green vortex is a representative benchmark for compressible turbulent flow accuracy assessment.
- ad hoc to paper Keeping the accumulated state Q in higher precision while lowering residual and work array precision prevents significant error accumulation.
Cite this review
Pith. "Pith review of Reduced and mixed precision turbulent flow simulations using explicit finite difference schemes." pith.science (2026). https://pith.science/paper/EGG66FKM
@misc{pith2026250520911,
author = {Pith},
title = {Pith review of: Reduced and mixed precision turbulent flow simulations using explicit finite difference schemes},
year = {2026},
howpublished = {\url{https://pith.science/paper/EGG66FKM}},
note = {Machine review of arXiv:2505.20911}
}
read the original abstract
The use of reduced and mixed precision computing has gained increasing attention in high-performance computing (HPC) as a means to improve computational efficiency, particularly on modern hardware architectures like GPUs. In this work, we explore the application of mixed precision arithmetic in compressible turbulent flow simulations using explicit finite difference schemes. We extend the OPS and OpenSBLI frameworks to support customizable precision levels, enabling fine-grained control over precision allocation for different computational tasks. Through a series of numerical experiments on the Taylor-Green vortex benchmark, we demonstrate that mixed precision strategies, such as half-single and single-double combinations, can offer significant performance gains without compromising numerical accuracy. However, pure half-precision computations result in unacceptable accuracy loss, underscoring the need for careful precision selection. Our results show that mixed precision configurations can reduce memory usage and communication overhead, leading to notable speedups, particularly on multi-CPU and multi-GPU systems.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
J. Sun, G. D. Peterson, O. O. Storaasli, High-performance mixed- precision linear solver for fpgas, IEEE Transactions on Computers 57 (12) (2008) 1614–1623. doi:10.1109/TC.2008.89
-
[2]
A. Abdelfattah, S. Tomov, J. Dongarra, Towards half-precision compu- tation for complex matrices: A case study for mixed precision solvers on gpus, in: 2019 IEEE /ACM 10th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Systems (ScalA), 2019, pp. 17–24. doi:10.1109/ScalA49573.2019.00008
arXiv 2019
-
[3]
J. D. Hogg, J. A. Scott, A fast and robust mixed-precision solver for the solution of sparse symmetric linear systems, ACM Trans. Math. Softw. 37 (2) (apr 2010). doi:10.1145/1731022.1731027. URL https://doi.org/10.1145/1731022.1731027
arXiv 2010
-
[4]
N. J. Higham, T. Mary, Mixed precision algorithms in numerical lin- ear algebra, Acta Numerica 31 (2022) 347–414. doi:10.1017/ S0962492922000022
work page 2022
-
[5]
P. Luszczek, A. Abdelfattah, H. Anzt, A. Suzuki, S. Tomov, Batched sparse and mixed-precision linear algebra interface for e fficient use of gpu hardware accelerators in scientific applications, Future Generation Computer Systems 160 (2024) 359–374. doi:https://doi.org/10. 1016/j.future.2024.06.004. URL https://www.sciencedirect.com/science/article/pii/ S...
work page 2024
-
[6]
A. Abdelfattah, H. Anzt, A. Ayala, E. G. Boman, E. C. Carson, S. Cayrols, T. Cojean, J. J. Dongarra, R. Falgout, M. Gates, T. Grutzmacher, N. J. Higham, S. E. Kruger, S. Li, N. Lindquist, Y . Liu, J. A. Loe, P. Nayak, D. Osei-Kuffuor, S. Pranesh, S. Rajamanickam, T. Ribizel, B. B. Smith, K. Swirydowicz, S. J. Thomas, S. Tomov, Y . M. Tsai, I. Yamazaki, U....
-
[7]
M. Lehmann, M. J. Krause, G. Amati, M. Sega, J. Harting, S. Gekle, Accuracy and performance of the lattice boltzmann method with 64-bit, 32-bit, and customized 16-bit number formats, Phys. Rev. E 106 (2022) 015308. doi:10.1103/PhysRevE.106.015308. URL https://link.aps.org/doi/10.1103/PhysRevE.106. 015308
-
[8]
F. Brogi, S. Bn `a, G. Boga, G. Amati, T. Esposti Ongaro, M. Cermi- nara, On floating point precision in computational fluid dynamics us- ing openfoam, Future Generation Computer Systems 152 (2024) 1–16. doi:https://doi.org/10.1016/j.future.2023.10.006. URL https://www.sciencedirect.com/science/article/pii/ S0167739X23003813
Show all 37 references
-
[9]
D. J. Lusher, S. P. Jammy, N. D. Sandham, OpenSBLI: Automated code- generation for heterogeneous computing architectures applied to com- pressible fluid dynamics on structured grids, Computer Physics Com- munications 267 (2021) 108063. doi:https://doi.org/10.1016/j. cpc.2021.108063
2021
-
[10]
I. Z. Reguly, G. R. Mudalige, M. B. Giles, Loop Tiling in Large-Scale Stencil Codes at Run-Time with OPS, IEEE Transactions on Parallel and Distributed Systems 29 (4) (2018) 873–886. doi:10.1109/TPDS. 2017.2778161
2018
-
[11]
Ieee standard for floating-point arithmetic, IEEE Std 754-2019 (Revi- sion of IEEE 754-2008) (2019) 1–84 doi:10.1109/IEEESTD.2019. 8766229
2019 doi
-
[12]
Goubault, Static analyses of the precision of floating-point operations, in: P
E. Goubault, Static analyses of the precision of floating-point operations, in: P. Cousot (Ed.), Static Analysis, Springer Berlin Heidelberg, Berlin, Heidelberg, 2001, pp. 234–259
2001
-
[13]
F. Benz, A. Hildebrandt, S. Hack, A dynamic program analysis to find floating-point accuracy problems, SIGPLAN Not. 47 (6) (2012) 453–462. doi:10.1145/2345156.2254118. URL https://doi.org/10.1145/2345156.2254118
2012
-
[14]
Haidar, S
A. Haidar, S. Tomov, J. Dongarra, N. J. Higham, Harnessing gpu ten- sor cores for fast fp16 arithmetic to speed up mixed-precision iterative refinement solvers, in: SC18: International Conference for High Perfor- mance Computing, Networking, Storage and Analysis, 2018, pp. 603–
2018
-
[15]
J. Wan, W. Wang, Z. Zhang, Enhancing computational e fficiency in 3- d seismic modelling with half-precision floating-point numbers based on the curvilinear grid finite-di fference method, Geophysical Journal Inter- national (2024). URL https://api.semanticscholar.org/CorpusID...
2024
-
[16]
Luszczek, I
P. Luszczek, I. Yamazaki, J. Dongarra, Increasing accuracy of iterative refinement in limited floating-point arithmetic on half-precision acceler- ators, in: 2019 IEEE High Performance Extreme Computing Conference (HPEC), 2019, pp. 1–6. doi:10.1109/HPEC.2019.8916392
2019
-
[17]
com/content/dam/develop/external/us/en/documents/ bf16-hardware-numerics-definition-white-paper.pdf (November 2018)
Intel, Bfloat16 – hardware numerics definition, https://www.intel. com/content/dam/develop/external/us/en/documents/ bf16-hardware-numerics-definition-white-paper.pdf (November 2018)
2018
-
[18]
Tortorella, L
Y . Tortorella, L. Bertaccini, L. Benini, D. Rossi, F. Conti, Redmule: A mixed-precision matrix–matrix operation engine for flexible and energy- efficient on-chip linear algebra and tinyml training acceleration, Fu- ture Generation Computer Systems 149 (2023) 122–135. doi:http...
2023 doi
-
[19]
I. Ober, I. Ober, On Patterns of Multi-domain Interaction for Scien- tific Software Development focused on Separation of Concerns, Procedia Computer Science 108 (2017) 2298–2302
2017
-
[20]
C. T. Jacobs, S. P. Jammy, N. D. Sandham, OpenSBLI: A framework for the automated derivation and parallel execution of finite difference solvers on a range of computer architectures, Journal of Computational Science 18 (2017) 12 – 23
2017
-
[21]
D. J. Lusher, S. P. Jammy, N. D. Sandham, Shock-wave /boundary-layer interactions in the automatic source-code generation framework opensbli, Computers & Fluids 173 (2018) 17 – 21
2018
-
[22]
D. J. Lusher, G. N. Coleman, Numerical study of compressible wall- bounded turbulence – the e ffect of thermal wall conditions on the turbu- lent Prandtl number in the low-supersonic regime, International Journal of Computational Fluid Dynamics 36 (9) (2022) 797–815
2022
-
[23]
D. J. Lusher, A. Sansica, N. D. Sandham, J. Meng, B. Sikl ´osi, 14 A. Hashimoto, OpenSBLI v3.0: High-fidelity multi-block transonic aero- foil CFD simulations using domain specific languages on GPUs, Com- puter Physics Communications 307 (2025) 109406.doi:https://doi. org/10.1...
2025
-
[24]
D. J. Lusher, A. Sansica, A. Hashimoto, E ffect of Tripping and Domain Width on Transonic Buffet on Periodic NASA-CRM Airfoils, AIAA Jour- nal (2024) 1–20
2024
-
[25]
D. J. Lusher, A. Sansica, A. Hashimoto, Implicit large eddy simula- tions of three-dimensional turbulent transonic buffet on wide-span infinite wings, Journal of Fluid Mechanics 1007 (2025) A26. doi:10.1017/ jfm.2025.48
2025
-
[26]
Pirozzoli, Numerical Methods for High-Speed Flows, Annual Review of Fluid Mechanics 43 (1) (2011) 163–194
S. Pirozzoli, Numerical Methods for High-Speed Flows, Annual Review of Fluid Mechanics 43 (1) (2011) 163–194
2011
-
[27]
Coppola, F
G. Coppola, F. Capuano, L. de Luca, Discrete Energy-Conservation Prop- erties in the Numerical Simulation of the Navier–Stokes Equations, Ap- plied Mechanics Reviews 71 (1), 010803 (03 2019)
2019
-
[28]
Coppola, F
G. Coppola, F. Capuano, S. Pirozzoli, L. de Luca, Numerically stable for- mulations of convective terms for turbulent compressible flows, Journal of Computational Physics 382 (2019) 86–104
2019
-
[29]
W. J. Feiereisen, W. C. Reynolds, J. H. Ferziger, Numerical simulation of a compressible homogeneous, turbulent shear flow. ph.d. thesis, 1981. URL https://ntrs.nasa.gov/citations/19820003523
1981
-
[30]
Blaisdell, E
G. Blaisdell, E. Spyropoulos, J. Qin, The e ffect of the formulation of nonlinear terms on aliasing errors in spectral methods, Applied Numeri- cal Mathematics 21 (3) (1996) 207–219. doi:https://doi.org/10. 1016/0168-9274(96)00005-0
1996
-
[31]
D. J. Lusher, N. D. Sandham, Assessment of Low-Dissipative Shock- Capturing Schemes for the Compressible Taylor–Green V ortex, AIAA Journal 59 (2) (2020) 533–545. doi:10.2514/1.J059672
2020 doi
-
[32]
Schlatter, M
P. Schlatter, M. Karp, R. Stanly, H. Song, T. Mukha, L. Galimberti, S. Toosi, M. M¨unsch, L. Dalcin, S. Rezaeiravesh, N. Jansson, S. Markidis, M. Parsani, S. T. Bose, S. K. Lele, Impact of low floating-point precision on high-fidelity simulations of turbulence, in: 77th Annual...
2024
-
[33]
S. P. Jammy, C. T. Jacobs, N. D. Sandham, Performance evaluation of explicit finite di fference algorithms with varying amounts of computa- tional and memory intensity, Journal of Computational Science 36 (2019) 100565. doi:https://doi.org/10.1016/j.jocs.2016.10.015. URL https...
2019 doi
-
[34]
J. DeBonis, Solutions of the Taylor-Green V ortex Problem Using High- Resolution Explicit Finite Difference Methods, in: 51st AIAA Aerospace Sciences Meeting including the New Horizons Forum and Aerospace Ex- position, Aerospace Sciences Meetings, American Institute of Aeronau...
2013
-
[35]
Chapelier, D
J.-B. Chapelier, D. J. Lusher, W. Van Noordt, C. Wenzel, T. Gibis, P. Mossier, A. Beck, G. Lodato, C. Brehm, M. Ruggeri, C. Scalo, N. Sand- ham, Comparison of high-order numerical methodologies for the simula- tion of the supersonic Taylor–Green vortex flow, Physics of Fluids ...
2024 doi
-
[36]
Y . Kuya, K. Totani, S. Kawai, Kinetic energy and entropy preserving schemes for compressible flows by split convective forms, Journal of Computational Physics 375 (2018) 823–853. doi:https://doi.org/ 10.1016/j.jcp.2018.08.058. 15
2018 doi
-
[613]
doi:10.1109/SC.2018.00050
2018
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.