Pith. sign in

REVIEW 3 major objections 5 minor 3 cited by

Hardware Trends Impacting Floating-Point Computations In Scientific Applications

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper argues that AI hardware's low-precision engines can be emulated up to double-precision accuracy, making the fastest chips also the most efficient for science.

desk verdict A competent survey with one load-bearing preliminary benchmark table that needs accuracy validation before its emulation speedup claim is taken at face value. read the letter →

arxiv 2411.12090 v2 pith:6GT5Z7QQ submitted 2024-11-18 math.NA cs.NA

classification math.NAcs.NA
keywords floating-pointarithmeticmixed-precisioncomputingemulationreducedprecisionGPUiterativerefinementhigh-performanceenergyefficiency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Artificial intelligence has pushed hardware toward low-precision arithmetic, and this paper argues that scientific computing can capture that speed instead of being hurt by it. The historical path from software emulation to dedicated floating-point units has become a loop: modern chips emulate high-precision arithmetic on low-precision, integer-based matrix engines, and mixed-precision iterative refinement corrects the low-precision bulk work back to double-precision accuracy. The paper's most concrete evidence is a measurement where the HPL benchmark, using emulation with seven 8-bit integer slices, ran about twice as fast and used 60 to 70 percent less energy per operation than native double precision on a current GPU. A mixed-precision solver on slightly older hardware is reported as 4.4 times faster and 5.8 times more energy-efficient than its full-precision counterpart. The result, if it holds, means the next generation of scientific supercomputers may be built around AI-style reduced-precision hardware with emulation and mixed precision doing the accuracy work.

What carries the argument

The central mechanism is emulation of high-precision arithmetic from many low-precision 'slices,' combined with mixed-precision iterative refinement. In the $s=7$ case, double-precision data are represented and multiplied as seven 8-bit integer slices on the tensor-core matrix-multiply units that AI workloads use, so a chip's cheap integer throughput becomes FP64-grade arithmetic; the paper relies on a recently published scheme for making integer matrix-multiplication units deliver that emulation efficiently. The companion mechanism is iterative refinement: a low-precision factorization does the heavy lifting, and a higher-precision residual correction repeats until the solution meets double-precision accuracy. Stochastic rounding is also cited as an error-control tool that keeps low-precision steps from polluting the final result.

What would settle it

On one GPU, solve the same dense system with native FP64 and with INT8 $s=7$ emulation at matched problem size, and compare the residual norms of the computed solutions; if the emulated residual is meaningfully worse, the reported 2x speedup is a precision trade-off, not a pure emulation gain.

Watch

Extended reading notes

Core claim

The paper's central claim, stated on its own terms, is that the AI-driven adoption of reduced-precision floating-point types is not merely a challenge for scientific computing but an opportunity, because emulation and mixed-precision algorithms can convert abundant low-precision throughput into double-precision-quality results. The load-bearing demonstration is an HPL measurement on a B200-class GPU: with emulation using $s=7$ eight-bit integer data elements, the run reached about 68 TFLOP/s compared with 34.5 TFLOP/s for native FP64 at maximum performance, and about 53 TFLOP/s compared with 23 TFLOP/s at maximum efficiency, with energy efficiency improving by roughly 60 to 70 percent. For a 32,000-by-32,000 complex double-precision system, the paper's mixed-precision iterative refinement solver reached 124 TFLOP/s and 529 GFLOP/s/Watt on an H200 GPU, against 42.6 and 78 for native FP64. The broader assertion is that dynamically switching precision during a computation, and emulating high precision on low-precision hardware, will define how scientific applications stay accurate while riding the performance curve of AI hardware.

Load-bearing premise

The emulated INT8 HPL run is assumed to be as accurate as native FP64, but the paper reports only speed and power efficiency, not residuals or error.

Editorial extensions

If this is right

  • A machine's scientific throughput becomes tied to its low-precision throughput, so the correlation between the standard FP64 ranking and application-relevant performance weakens.
  • Dense solvers can be restructured into a low-precision bulk phase plus a high-precision correction phase, with the reported payoff of 4.4x speed and 5.8x energy efficiency on data-center GPUs.
  • Energy per useful operation, not peak FLOPS, becomes the binding design constraint as power budgets approach 40 megawatts.
  • Emulation becomes a deliberate design feature, not a stopgap, letting a single system cover both AI and double-precision scientific workloads.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the accuracy of the emulated HPL run is confirmed, the same seven-slice technique should transfer to other dense linear algebra kernels, giving near-2x speedups on existing AI accelerators.
  • The paper stops short of showing the emulated HPL run is as accurate as native FP64; a reader who wants the speedup should check residuals first.
  • For ill-conditioned systems, iterative refinement will need extra passes, so the emulation advantage should shrink as the condition number grows; that is a testable prediction.
  • If the historical widening of the gap between low-precision and high-precision throughput continues, the emulation speedup should grow across future hardware generations.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper is a perspective/review article on how hardware trends, particularly the adoption of reduced-precision floating-point formats driven by AI, are reshaping scientific computing. It surveys the historical evolution of floating-point support from software emulation and coprocessors to integrated FPUs and GPUs, discusses community benchmarks (HPL, HPCG, Green500, HPL-MxP), and synthesizes recent developments in mixed-precision algorithms and emulation techniques. The original contributions are quantitative: Table I charts throughput and memory bandwidth across four NVIDIA GPU generations; Table II reports performance and power efficiency for a mixed-precision iterative-refinement solver; Table III reports preliminary HPL results on a Blackwell B200 comparing native FP64 with INT8-slice emulation (s=7), claiming a 2.0-2.3x speedup and 60-70% power-efficiency improvement. The paper argues that these trends reflect a broader shift toward flexibility in precision, where emulation and mixed-precision allow lower-precision hardware to serve high-precision scientific workloads.

Significance. If the new measurements are trustworthy, the paper provides a valuable and credible synthesis of the current state and near-term trajectory of floating-point computing, with an authoritative author list. The survey of historical developments and benchmark evolution is generally accurate and well organized, and the emphasis on INT8-slice emulation as a means to leverage AI-oriented tensor hardware for FP64-class computation is a timely and potentially important observation. The paper would be strengthened as a reference point for the community if its quantitative claims were backed by reproducible methodology. The main significance currently rests on the two self-reported benchmark tables (II and III), which are exactly the parts that lack detail; this limits the paper's value as more than an opinionated literature review.

major comments (3)
  1. [Section IV-C, Table III] The central quantitative claim, that INT8 emulation with s=7 roughly doubles HPL performance and improves power efficiency by 60-70% versus native FP64 on a B200, assumes that both configurations are solving the same problem to the same numerical accuracy. The table and surrounding text report only TFLOP/s and GFLOP/s/Watt, with no residual, validation metric, or reference to the HPL acceptance criterion. The table is explicitly labeled 'preliminary' and cites reference [38], but the paper does not state whether that reference contains an accuracy validation, nor does Section IV-C provide any error analysis. As written, the comparison is only a throughput/efficiency comparison, not a capability-equivalent one. Please either provide the scaled residuals or other acceptance data demonstrating FP64-comparable accuracy, or revise the claim to be explicitly about raw speed without implying equivalent solution quality.
  2. [Section IV-B, Table II / Figure 4] The mixed-precision iterative-refinement results in Table II (and the A100 power curve in Figure 4) lack the experimental detail needed to assess the claimed 4.4x speedup and 5.8x power-efficiency gain. The table reports only the matrix size (32K complex) and the labels 'FP16+FP64 MxP', but does not specify the number of iterative-refinement iterations, the convergence threshold, the achieved residual, or whether the final solution is returned in FP64. Since this table is used as evidence that mixed-precision solvers preserve accuracy while improving efficiency, the authors should cite a public reproducible implementation, provide typical residual values, or state the accuracy target. Without this, the reader cannot tell whether the speedup is achieved while meeting the same numerical quality as the FP64 baseline.
  3. [Section XI-A, Figure 5] The dashed line in Figure 5, described as 'Tensor Core accelerated DGEMM performance using integer based emulation with 7 slices (INT8 data storage elements) [38]', is presented as if it were a direct point of comparison with the Bytes/FLOP curves of Table I, but the figure does not show the data points or the accuracy of the emulated DGEMM. If this line represents a single measurement or an extrapolation, that should be stated. The same accuracy-equivalence concern as in Table III applies here: without a statement about the numerical error of the emulated DGEMM relative to FP64, the comparison is misleading for readers who might infer that the emulated math is a drop-in replacement for native FP64 matrix multiply.
minor comments (5)
  1. [Figure 3 caption] The caption contains a typo: 'Fontier' should be 'Frontier'.
  2. [Throughout text and references] The name 'Volta' appears as 'V olta' in multiple places (e.g., Section IV-A, VI-B, and reference [50]); the spacing should be removed.
  3. [Table III] The variable 's' is used without definition. Define it as the number of integer slices used in the emulation scheme (presumably following [38]) and explain why s=7 was chosen.
  4. [Section V-A] The sentence 'as seen in Figure 3' for the precision/efficiency trade-off is confusing, because Figure 3 plots the HPL-MxP-to-HPL Rmax ratio over time, which is not directly a precision-versus-efficiency trade-off. Please re-reference or rephrase.
  5. [Table I / Section XI-A] Bytes/FLOP is an informative figure of merit, but the table should clarify whether the ratios are computed from the listed peak TFLOP/s values or from some measured values, and state the source or formula for the memory bandwidth numbers.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation: the paper's claims are contextual observations supported by benchmarks and vendor measurements, with only non-load-bearing self-citations.

full rationale

No load-bearing step reduces to the paper's own inputs. The central assertions—AI-driven reduced precision, mixed-precision gains, emulation benefits, and hardware/software co-evolution—are presented as review-level interpretations of external benchmarks (HPL, HPCG, Green500, TOP500, HPL-MxP) and of directly reported vendor measurements (Tables I-III, Figures 2-4). Table III reports a measured 2.0-2.3x HPL speedup for INT8 emulation versus native FP64; this is a benchmark ratio, not a quantity derived from an assumed accuracy equivalence, so it is not circular by construction. The absence of residual or validation data for the emulated run is a legitimate correctness/validation concern, but not a circularity. Self-authored citations ([5], [9], [10], [31]-[33]) appear as benchmark definitions and prior algorithm descriptions; the paper's trend conclusions do not rest on those citations as unverified premises. Therefore no specific circular step can be identified; the low score only notes the presence of self-citations and vendor-sourced measurements.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

As a review, the paper introduces no free parameters or invented entities. Its conclusions rest on external benchmarks and vendor-supplied measurements. The main domain assumptions are that these benchmarks are meaningful proxies and that the reported measurements are accurate and accuracy-preserving.

assumptions (3)
  • domain assumption Benchmark metrics (TOP500/HPL, Green500, HPCG, HPL-MxP) are meaningful proxies for real-world supercomputing performance.
    Section II introduces these benchmarks as 'standard tools' and uses them to support the paper's hardware-trend analysis.
  • domain assumption The emulated HPL computation using INT8 slices (s=7) achieves the same numerical accuracy as native FP64.
    Section IV-C and Table III compare only speed and efficiency, never reporting accuracy of the emulated solution; the speedup claim is valid only if accuracy is preserved.
  • domain assumption Vendor-supplied performance and power measurements (Tables I to III) are accurate and representative.
    Tables I and II cite NVIDIA white papers and product pages; Table III is labeled preliminary and 'data subject to change' with no measurement methodology.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Hardware Trends Impacting Floating-Point Computations In Scientific Applications." pith.science (2026). https://pith.science/paper/6GT5Z7QQ

@misc{pith2026241112090,
  author       = {Pith},
  title        = {Pith review of: Hardware Trends Impacting Floating-Point Computations In Scientific Applications},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6GT5Z7QQ}},
  note         = {Machine review of arXiv:2411.12090}
}
read the original abstract

The evolution of floating-point computation has been shaped by algorithmic advancements, architectural innovations, and the increasing computational demands of modern technologies, such as artificial intelligence (AI) and high-performance computing (HPC). This paper examines the historical progression of floating-point computation in scientific applications and contextualizes recent trends driven by AI, particularly the adoption of reduced-precision floating-point types. The challenges posed by these trends, including the trade-offs between performance, efficiency, and precision, are discussed, as are innovations in mixed-precision computing and emulation algorithms that offer solutions to these challenges. This paper also explores architectural shifts, including the role of specialized and general-purpose hardware, and how these trends will influence future advancements in scientific computing, energy efficiency, and system design.

Figures

Figures reproduced from arXiv: 2411.12090 by the authors.

Figure 1
Figure 1. Various floating-point (FP) representations used today in scientific [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Historical record of the advances made in transistor density and [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 4
Figure 4. Representative power consumption curves measured on an NVIDIA [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figures from the paper (1 more)
Figure 5
Figure 5. Figure 5: Comparison of Bytes/FLOP across four generations of GPUs for FMA [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Ascend to Science: Exploration of AI Chips for Scientific Computing

    cs.DC 2026-07 conditional novelty 5.0 of 10

    AI-oriented Ascend NPUs can run scientific workloads with FP32-like accuracy and competitive throughput when algorithms are reformulated and data movement is explicitly orchestrated.

  2. CHAMB-GA: A Containerized HPC Scalable Microservice-Based Framework for Genetic Algorithms

    cs.DC 2026-06 unverdicted novelty 5.0 of 10

    CHAMB-GA provides a microservice architecture with containers and a message broker to decouple genetic operations from fitness evaluations, enabling consistent scaling from small machines to over 3500 CPU cores on clo...

  3. Mixed-precision numerics in scientific applications: survey and perspectives

    cs.CE 2024-12 conditional novelty 3.0 of 10

    A survey of mixed-precision numerical methods across CFD, climate, chemistry, and genomics, reporting speedups up to 8x on benchmarks and recommending co-design to unlock them.

Reference graph

Works this paper leans on

70 extracted references · 56 canonical work pages · cited by 3 Pith papers

  1. [38]

    Performance enhancement of the ozaki scheme on integer matrix multiplication unit

    Y . Uchino, K. Ozaki, and T. Imamura, “Performance enhancement of the ozaki scheme on integer matrix multiplication unit.” arXiv:2409.13313 [cs.DC], Sept. 2024

  2. [1]

    Deep learning,

    Y . LeCun, Y . Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, p. 436, 2015

  3. [2]

    A study of bfloat16 for deep learning training

    D. Kalamkar, D. Mudigere, N. Mellempudi, D. Das, K. Banerjee, S. Avancha, D. T. V ooturi, N. Jammalamadaka, J. Huang, H. Yuen, J. Yang, J. Park, A. Heinecke, E. Georganas, S. Srinivasan, A. Kundu, M. Smelyanskiy, B. Kaul, and P. Dubey, “A study of bfloat16 for deep learning training.” arXiv:1905.12322 [cs.LG], May 2019

  4. [3]

    Fp8 formats for deep learning

    P. Micikevicius, D. Stosic, N. Burgess, M. Cornea, P. Dubey, R. Grisen- thwaite, S. Ha, A. Heinecke, P. Judd, J. Kamalu, N. Mellempudi, S. Oberman, M. Shoeybi, M. Siu, and H. Wu, “Fp8 formats for deep learning.” arXiv:2209.05433 [cs.LG], Sept. 2022. 8

  5. [4]

    Introducing the graph 500,

    R. C. Murphy, K. Pingali, J. D. Feo, and D. A. Bader, “Introducing the graph 500,” Cray User Group (CUG) , 2010

  6. [5]

    The linpack benchmark: Past, present, and future,

    J. J. Dongarra, “The linpack benchmark: Past, present, and future,” Concurrency and Computation: Practice and Experience , vol. 15, no. 9, pp. 803–820, 2003

  7. [6]

    Dongarra and P

    J. Dongarra and P. Luszczek, “Top500,” in Encyclopedia of Parallel Computing (D. Padua, ed.), pp. 2055–2057, Boston, MA: Springer US, 2011

  8. [7]

    Top 500. the list

    E. Strohmaier, J. Dongarra, H. Simon, and M. Meuer, “Top 500. the list..” https://top500.org, June 2024

Show all 70 references
  1. [8]

    The green500 list: Encouraging sustainable supercomputing,

    W. Feng and K. Cameron, “The green500 list: Encouraging sustainable supercomputing,” Computer, vol. 40, no. 12, pp. 50–55, 2007

  2. [9]

    A new benchmark for ranking high performance computing systems,

    J. J. Dongarra, M. A. Heroux, and P. Luszczek, “A new benchmark for ranking high performance computing systems,” Tech. Rep. UT-EECS- 13-736, University of Tennessee, 2013

  3. [10]

    Hpl- ai mixed-precision benchmark: The next frontier of supercomputing,

    J. J. Dongarra, P. Luszczek, S. Tomov, and M. A. Heroux, “Hpl- ai mixed-precision benchmark: The next frontier of supercomputing,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis , 2021

  4. [11]

    The state of the transistor in 3 charts

    IEEE Spectrum, “The state of the transistor in 3 charts.” https://spectr um.ieee.org/transistor-density. [Accessed 16-11-2024]

  5. [12]

    The IBM 701 Speedcoding system,

    J. W. Backus, “The IBM 701 Speedcoding system,” Journal of the ACM, vol. 1, no. 1, pp. 4–6, 1954

  6. [13]

    The intel 8087 numeric data processor,

    J. Palmer, “The intel 8087 numeric data processor,” in Proceedings of the 7th Annual Symposium on Computer Architecture, La Baule, France, May 6-8, 1980 (J. Lenfant, B. R. Borgerson, D. E. Atkins, K. B. Irani, D. Kinniment, and H. Aiso, eds.), pp. 174–181, ACM, 1980

  7. [14]

    The MC68881 floating-point copro- cessor,

    C. Huntsman and D. Cawthron, “The MC68881 floating-point copro- cessor,” IEEE Micro, vol. 3, pp. 44–54, Nov./Dec. 1983

  8. [15]

    Developing the WTL3170/3171 Sparc floating-point copro- cessors,

    M. Birman, A. Samuels, G. Chu, T. Chuk, L. Hu, J. McLeod, and J. Barnes, “Developing the WTL3170/3171 Sparc floating-point copro- cessors,” IEEE Micro, vol. 10, pp. 55–64, Jan./Feb. 1990

  9. [16]

    Heinrich, MIPS R4000 user’s manual

    J. Heinrich, MIPS R4000 user’s manual. USA: Prentice-Hall, Inc., 1993

  10. [17]

    W. A. Triebel, The 80386, 80486, and Pentium Microprocessors: Hard- ware, Software, and Interfacing. Simon & Schuster Trade, 1st ed., 1997

  11. [18]

    The 68040 processor. i. design and implementation,

    R. Edenfield, M. Gallup, W. Ledbetter, R. McGarity, E. Quintana, and R. Reininger, “The 68040 processor. i. design and implementation,” IEEE Micro, vol. 10, no. 1, pp. 66–78, 1990

  12. [19]

    Design considerations for the powerpc 601 microprocessor,

    M. T. Vaden, L. J. Merkel, C. R. Moore, T. M. Potter, and R. J. Reese, “Design considerations for the powerpc 601 microprocessor,” IBM Journal of Research and Development , vol. 38, no. 5, pp. 605– 620, 1994

  13. [20]

    Evolution of the graphics processing unit (gpu),

    W. J. Dally, S. W. Keckler, and D. B. Kirk, “Evolution of the graphics processing unit (gpu),” IEEE Micro, vol. 41, no. 6, pp. 42–51, 2021

  14. [21]

    Brook for gpus: stream computing on graphics hardware,

    I. Buck, T. Foley, D. Horn, J. Sugerman, K. Fatahalian, M. Houston, and P. Hanrahan, “Brook for gpus: stream computing on graphics hardware,” ACM Trans. Graph., vol. 23, p. 777–786, Aug. 2004

  15. [22]

    Scalable parallel programming with cuda.,

    J. Nickolls, I. Buck, M. Garland, and K. Skadron, “Scalable parallel programming with cuda.,” in SIGGRAPH Classes , pp. 16:1–16:14, ACM, 2008

  16. [23]

    Cuda: Scalable parallel programming for high-performance scientific computing,

    D. Luebke, “Cuda: Scalable parallel programming for high-performance scientific computing,” in 2008 5th IEEE International Symposium on Biomedical Imaging: From Nano to Macro , pp. 836–838, 2008

  17. [24]

    Accelerating molecular dynamics simulations using graphics processing units with cuda,

    W. Liu, B. Schmidt, G. V oss, and W. M ¨uller-Wittig, “Accelerating molecular dynamics simulations using graphics processing units with cuda,” Computer physics communications, vol. 179, no. 9, pp. 634–641, 2008

  18. [25]

    Gpu computing,

    J. D. Owens, M. Houston, D. Luebke, S. Green, J. E. Stone, and J. C. Phillips, “Gpu computing,” Proceedings of the IEEE , vol. 96, no. 5, pp. 879–899, 2012

  19. [26]

    Large-scale deep unsupervised learning using graphics processors,

    R. Raina, A. Madhavan, and A. Y . Ng, “Large-scale deep unsupervised learning using graphics processors,” in Proceedings of the 26th annual international conference on machine learning, pp. 873–880, ACM, 2009

  20. [27]

    Imagenet classifica- tion with deep convolutional neural networks,

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classifica- tion with deep convolutional neural networks,” in Advances in Neural Information Processing Systems (F. Pereira, C. Burges, L. Bottou, and K. Weinberger, eds.), vol. 25, Curran Associates, Inc., 2012

  21. [28]

    OCP 8-bit Floating Point Specification (OFP8),

    P. Micikevicius, S. Oberman, P. Dubey, M. Cornea, A. Rodriguez, I. Bratt, R. Grisenthwaite, N. Jouppi, C. Chou, A. Huffman, M. Schulte, R. Wittig, D. Jani, and S. Deng, “OCP 8-bit Floating Point Specification (OFP8),” Open Compute Project , 2023

  22. [29]

    Microscal- ing data formats for deep learning

    B. D. Rouhani, R. Zhao, A. More, M. Hall, A. Khodamoradi, S. Deng, D. Choudhary, M. Cornea, E. Dellinger, K. Denolf, S. Dusan, V . Elango, M. Golub, A. Heinecke, P. James-Roxby, D. Jani, G. Kolhe, M. Lang- hammer, A. Li, L. Melnick, M. Mesmakhosroshahi, A. Rodriguez, M. Schult...

  23. [30]

    Nvidia v100 gpu architecture

    NVIDIA Corporation, “Nvidia v100 gpu architecture.” https://images.n vidia.com/content/volta-architecture/pdf/volta-architecture-whitepaper. pdf, 2017. White paper

  24. [31]

    Accelerating scientific computations with mixed precision algorithms,

    M. Baboulin, A. Buttari, J. Dongarra, J. Kurzak, J. Langou, J. Lan- gou, P. Luszczek, and S. Tomov, “Accelerating scientific computations with mixed precision algorithms,” Computer Physics Communications , vol. 180, no. 12, pp. 2526–2533, 2009

  25. [32]

    A survey of numerical linear algebra methods utilizing mixed-precision arithmetic,

    A. Abdelfattah, H. Anzt, E. G. Boman, E. Carson, T. Cojean, J. Don- garra, A. Fox, M. Gates, N. J. Higham, X. S. Li, J. Loe, P. Luszczek, S. Pranesh, S. Rajamanickam, T. Ribizel, B. F. Smith, K. Swirydowicz, S. Thomas, S. Tomov, Y . M. Tsai, and U. M. Yang, “A survey of numeri...

  26. [33]

    The design of fast and energy-efficient linear solvers: On the potential of half-precision arithmetic and iterative refinement techniques,

    A. Haidar, A. Abdelfattah, M. Zounon, P. Wu, S. Pranesh, S. Tomov, and J. Dongarra, “The design of fast and energy-efficient linear solvers: On the potential of half-precision arithmetic and iterative refinement techniques,” in International Conference on Computational Science...

  27. [34]

    Deep learning with limited numerical precision,

    S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan, “Deep learning with limited numerical precision,” in Proceedings of the 32nd International Conference on Machine Learning (F. Bach and D. Blei, eds.), vol. 37 of Proceedings of Machine Learning Research , (Lille, Franc...

  28. [35]

    Solving lattice qcd systems of equations using mixed precision solvers on gpus,

    M. Clark, R. Babich, K. Barros, R. Brower, and C. Rebbi, “Solving lattice qcd systems of equations using mixed precision solvers on gpus,” Computer Physics Communications , vol. 181, no. 9, pp. 1517–1528, 2010

  29. [36]

    Stochastic rounding: implementation, error analysis and applications,

    M. Croci, M. Fasi, N. J. Higham, T. Mary, and M. Mikaitis, “Stochastic rounding: implementation, error analysis and applications,” Royal Soci- ety Open Science , vol. 9, no. 3, 2022

  30. [37]

    Recovering single precision accuracy from tensor cores while surpassing the FP32 theoretical peak performance,

    H. Ootomo and R. Yokota, “Recovering single precision accuracy from tensor cores while surpassing the FP32 theoretical peak performance,” Int. J. High Performance Computing Applications , vol. 36, p. 475–491, June 2022

  31. [39]

    Leveraging the bfloat16 artificial intelligence datatype for higher-precision computations

    G. Henry, P. T. P. Tang, and A. Heinecke, “Leveraging the bfloat16 artificial intelligence datatype for higher-precision computations.” arXiv:1904.06376 [cs.MS], Apr. 2019

  32. [40]

    Simulation intelligence: Towards a new generation of scientific methods

    A. Lavin, D. Krakauer, H. Zenil, J. Gottschlich, T. Mattson, J. Brehmer, A. Anandkumar, S. Choudry, K. Rocki, A. G. Baydin, C. Prunkl, B. Paige, O. Isayev, E. Peterson, P. L. McMahon, J. Macke, K. Cranmer, J. Zhang, H. Wainwright, A. Hanuka, M. Veloso, S. Assefa, S. Zheng, and...

  33. [41]

    Functionality and performance of nvlink with ibm power9 processors,

    IBM POWER9 NPU team, “Functionality and performance of nvlink with ibm power9 processors,” IBM J. Res. Dev. , vol. 62, p. 9:1–9:10, July 2018

  34. [42]

    Nvidia gh200 grace hopper superchip architec- ture

    NVIDIA Corporation, “Nvidia gh200 grace hopper superchip architec- ture.” https://resources.nvidia.com/en-us-grace-cpu/nvidia-grace-hopper, 2023

  35. [43]

    Porting hpc applications to amd instinct TM mi300a using unified memory and openmp

    S. Tandon, L. Grinberg, G.-T. Bercea, C. Bertolli, M. Olesen, S. Bn `a, and N. Malaya, “Porting hpc applications to amd instinct TM mi300a using unified memory and openmp.” arXiv.org:2405.00436 [cs.DC], May 2024

  36. [44]

    AMD CDNA Architecture

    AMD Corporation, “AMD CDNA Architecture.” https://www.amd.com/ en/technologies/cdna.html. [Accessed 5-12-2024]

  37. [45]

    Apple unveils m3, m3 pro, and m3 max, the most advanced chips for a personal computer

    “Apple unveils m3, m3 pro, and m3 max, the most advanced chips for a personal computer.” https://www.apple.com/newsroom/2023/10/apple -unveils-m3-m3-pro-and-m3-max-the-most-advanced-chips-for-a-per sonal-computer/. [Accessed 14-11-2024]

  38. [46]

    First impressions of the nvidia grace cpu superchip and nvidia grace hopper superchip for scientific workloads,

    N. A. Simakov, M. D. Jones, T. R. Furlani, E. Siegmann, and R. J. Harrison, “First impressions of the nvidia grace cpu superchip and nvidia grace hopper superchip for scientific workloads,” in Proceedings of the International Conference on High Performance Computing in Asia- P...

  39. [47]

    A survey of cpu-gpu heterogeneous computing techniques,

    S. Mittal and J. S. Vetter, “A survey of cpu-gpu heterogeneous computing techniques,” ACM Comput. Surv., vol. 47, July 2015

  40. [48]

    Programming model for a heterogeneous x86 platform,

    B. Saha, X. Zhou, H. Chen, Y . Gao, S. Yan, M. Rajagopalan, J. Fang, P. Zhang, R. Ronen, and A. Mendelson, “Programming model for a heterogeneous x86 platform,” in Proceedings of the 30th ACM SIGPLAN Conference on Programming Language Design and Implementation , PLDI ’09, (New...

  41. [49]

    Designing a unified program- ming model for heterogeneous machines,

    M. Garland, M. Kudlur, and Y . Zheng, “Designing a unified program- ming model for heterogeneous machines,” in SC ’12: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis , pp. 1–11, 2012

  42. [50]

    Dis- secting the nvidia volta gpu architecture via microbenchmarking

    Z. Jia, M. Maggioni, B. Staiger, and D. P. Scarpazza, “Dis- secting the nvidia volta gpu architecture via microbenchmarking.” arXiv:1804.06826 [cs.DC], Apr. 2018

  43. [51]

    Tensor cores

    NVIDIA Corporation, “Tensor cores.” https://www.nvidia.com/en-us/da ta-center/tensor-cores/, 2024

  44. [52]

    CUDA PTX ISA

    NVIDIA Corporation, “CUDA PTX ISA.” https://docs.nvidia.com/cuda /pdf/ptx isa 8.5.pdf, May 2024

  45. [53]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , NIPS’17, (Red Hook, NY , USA), p. 6000–6010, Curran...

  46. [54]

    Mixed precision training,

    P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, et al. , “Mixed precision training,” International Conference on Learning Representa- tions (ICLR), 2018

  47. [55]

    Simulating low precision floating-point arithmetic,

    N. J. Higham and S. Pranesh, “Simulating low precision floating-point arithmetic,” SIAM Journal on Scientific Computing , vol. 41, no. 5, pp. C585–C602, 2019

  48. [56]

    Improving weather forecast skill through reduced-precision data assimilation,

    S. Hatfield, A. Subramanian, T. Palmer, and P. D ¨uben, “Improving weather forecast skill through reduced-precision data assimilation,” Monthly Weather Review, vol. 146, no. 1, pp. 49 – 62, 2018

  49. [57]

    Pytorch: An imperative style, high- performance deep learning library

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K ¨opf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-...

  50. [58]

    Tensorflow: a system for large-scale machine learning,

    M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V . Vasudevan, P. Warden, M. Wicke, Y . Yu, and X. Zheng, “Tensorflow: a system for large-sca...

  51. [59]

    JAX: composable transformations of Python+NumPy pro- grams

    J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclau- rin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang, “JAX: composable transformations of Python+NumPy pro- grams.” http://github.com/jax-ml/jax, 2018

  52. [60]

    In-datacenter performance analysis of a tensor processing unit,

    N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P.-l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V . Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. ...

  53. [61]

    Gpt-4 technical report

    OpenAI, “Gpt-4 technical report.” https://arxiv.org/abs/2303.08774, 2024

  54. [62]

    Dynamic voltage and frequency scaling: the laws of diminishing returns,

    E. Le Sueur and G. Heiser, “Dynamic voltage and frequency scaling: the laws of diminishing returns,” in Proceedings of the 2010 International Conference on Power Aware Computing and Systems , HotPower’10, (USA), p. 1–8, USENIX Association, 2010

  55. [63]

    Intel® pentium® m processor power estimation, budgeting, optimization, and validation.,

    D. Genossar and N. Shamir, “Intel® pentium® m processor power estimation, budgeting, optimization, and validation.,” Intel Technology Journal, vol. 7, no. 2, 2003

  56. [64]

    Evaluating and modeling power consumption of multi-core processors,

    R. Basmadjian and H. de Meer, “Evaluating and modeling power consumption of multi-core processors,” in Proceedings of the 3rd International Conference on Future Energy Systems: Where Energy, Computing and Communication Meet , e-Energy ’12, (New York, NY , USA), Association for...

  57. [65]

    Beating floating point at its own game: Posit arithmetic,

    Gustafson and Yonemoto, “Beating floating point at its own game: Posit arithmetic,” Supercomput. Front. Innov.: Int. J. , vol. 4, p. 71–86, June 2017

  58. [66]

    The spinnaker 2 processing element architec- ture for hybrid digital neuromorphic computing

    S. H ¨oppner, Y . Yan, B. V ogginger, C. Liu, F. Kelber, A. Dixius, S. Scholze, J. Partzsch, M. Stolba, F. Neum ¨arker, G. Ellguth, S. Hart- mann, S. Schiefer, T. Hocker, D. Walter, G. Liu, M. Mikaitis, J. Garside, S. Furber, and C. Mayr, “The spinnaker 2 processing element ar...

  59. [67]

    Analog and digital, continuous and discrete,

    C. J. Maley, “Analog and digital, continuous and discrete,” Philosophical Studies, vol. 155, pp. 117–131, 2011

  60. [68]

    Hbm (high bandwidth memory) dram technology and architecture,

    H. Jun, J. Cho, K. Lee, H.-Y . Son, K. Kim, H. Jin, and K. Kim, “Hbm (high bandwidth memory) dram technology and architecture,” in 2017 IEEE International Memory Workshop (IMW) , pp. 1–4, 2017

  61. [69]

    Flashattention: Fast and memory-efficient exact attention with io-awareness

    T. Dao, D. Y . Fu, S. Ermon, A. Rudra, and C. R ´e, “Flashattention: Fast and memory-efficient exact attention with io-awareness.” https: //arxiv.org/abs/2205.14135, 2022

  62. [70]

    Flashattention-2: Faster attention with better parallelism and work partitioning

    T. Dao, “Flashattention-2: Faster attention with better parallelism and work partitioning.” https://arxiv.org/abs/2307.08691, 2023. 10

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.