Pith. sign in

REVIEW 2 major objections 4 minor 45 references

Adding MFMA Support to gem5

T0 review · 2 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper reports that adding Matrix Core Engine support to gem5 times AMD MI200 and MI300 MFMA instructions within 1.5% and 1.3% of real hardware, making cycle-level ML workload simulation practical.

desk verdict A solid, narrowly-scoped gem5 MFMA support paper with credible microbenchmark validation; the broad ML-workload claims outrun the evidence, but the core contribution deserves review. read the letter →

arxiv 2501.18113 v2 pith:V3WDOF6B submitted 2025-01-30 cs.AR

classification cs.AR
keywords gem5MFMAMatrixCoreEngineGPUsimulationAMDMI200MI300cycle-levelmachinelearningworkloads
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper adds Matrix Core Engine (MCE) support to the gem5 simulator for AMD MI200 and MI300 GPUs, so that Matrix Fused Multiply Add (MFMA) instructions, the matrix-multiply workhorses of modern ML libraries, are timed at cycle level rather than treated as ordinary vector operations. The authors validate the new timing model with hand-written assembly microbenchmarks on real MI200 and MI300 hardware and report mean absolute percentage errors of 1.5% and 1.3%, respectively, across a range of precisions and block sizes. If the validation holds, gem5 can simulate modern ML workloads that rely on MFMA instructions with timing within about one to two percent of real hardware, and researchers can perturb MCE latency through a new configuration flag to explore future GPU designs.

What carries the argument

The load-bearing mechanism is a scoreboard-based MFMA issue rule in gem5's compute-unit timing logic, backed by a per-instruction latency lookup table (mfma_cycles) derived from the MI300 ISA manual and validated by Equation 1: T_MFMA = (T_total - T_memtime - T_inst)/(N_MFMA - 1), where T_memtime = 40 and T_inst = 4 are overhead constants from prior calibration. The scoreboard's NRDY_MATRIX_CORE field per SIMD unit decides when the MCE is free; since the scalar pipeline that runs s_memtime is independent of the MCE pipeline, dependent MFMA chains are needed to make the timing instruction wait for matrix completion.

What would settle it

Run the identical dependent-MFMA timing kernels on an MI200 or MI300 while replacing the s_memtime overhead constants with directly measured values, for example by timing a single s_memtime around a known number of s_nop instructions, and compare the resulting per-MFMA latencies; a systematic shift larger than the reported MAPE would indicate Equation 1's constants are wrong. Separately, launch two independent MFMA streams from different wavefronts to the same MCE and measure whether throughput exceeds one MFMA per latency, which would falsify the non-pipelining assumption.

Watch

Extended reading notes

Core claim

On the paper's own terms: gem5's GPU model now includes MCEs as separate functional units, one per SIMD unit (four per compute unit), and uses the NRDY_MATRIX_CORE scoreboard field to prevent more than one MFMA on the same SIMD unit at a time. This design matches AMD's reported MCE throughput and mirrors the compiler's apparent assumption that MFMA instructions from a given wavefront are not pipelined in the MCE. The validation method times back-to-back dependent MFMA instructions with s_memtime and subtracts measured overheads (T_memtime = 40 cycles, T_inst = 4 cycles) to recover per-instruction latency; across the tested instructions, gem5 matches the ISA manual's expected cycle counts and real hardware within 1.5% (MI200) and 1.3% (MI300) mean absolute percentage error. The paper also adds a --mfma-scale parameter that scales MFMA latency to support what-if analysis of future MCE designs.

Load-bearing premise

The load-bearing premise is that the fixed timing overheads (T_memtime = 40 cycles, T_inst = 4 cycles) and the assumption that a wavefront's MFMA instructions never pipeline in the MCE correctly describe real MI200 and MI300 hardware; if either is wrong, the recovered per-instruction latencies, and therefore the claimed 1.5% and 1.3% accuracy, would be biased.

Editorial extensions

If this is right

  • Modern ML workloads that call MFMA-based libraries can now be run in gem5 with cycle-level timing rather than only functional simulation.
  • Researchers can model MCE improvements, faster or slower, by setting --mfma-scale, which multiplies the per-instruction MFMA latency and exposes how application runtime responds.
  • MI200 and MI300 differences, including new instructions, removed instructions, and changed latencies, are captured in the gem5 model, so cross-generation comparisons can be simulated.
  • Validation error decreases as more MFMA instructions are timed back-to-back, from 2.3% MAPE at 2 MFMAs to 0.4% at 5 on MI200, suggesting the timing methodology is limited by transient effects rather than by the core model.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the non-pipelining assumption is inferred from compiler-inserted NOPs rather than measured hardware behavior, the gem5 model may need revision if AMD's compiler or future hardware starts pipelining MFMA instructions from the same wavefront; a direct microbenchmark that varies the amount of independent work between dependent MFMAs would settle this.
  • The same interleaved s_memtime harness could serve as a portable benchmark for measuring MFMA latencies on future AMD GPUs and refreshing the mfma_cycles table, since it depends only on the ISA, not on gem5 internals.
  • The paper's limitation note implies that --mfma-scale results likely overstate how a real compiled workload would respond to faster or slower MCEs, because the compiler inserts NOPs and independent work based on the original latency; scaling studies therefore need compiler co-design to be realistic.
  • A similar scoreboard-plus-latency-table structure could be adapted to model NVIDIA TensorCore timing in gem5, although the validation harness would need different timing instructions than s_memtime.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper adds Matrix Core Engine (MCE) support for AMD MI200 and MI300 GPUs to the gem5 simulator, including functional and timing models for MFMA instructions. The authors validate their implementation with handwritten assembly microbenchmarks that interleave back-to-back, dependent MFMA instructions with s_memtime timing instructions, running the same kernels on real GPUs and in gem5. They compute per-MFMA latency using Equation (1), which subtracts a fixed 44-cycle overhead (Tmemtime=40, Tinst=4) and divides by N_MFMA−1. The reported results show 1.5% MAPE for MI200 and 1.3% MAPE for MI300 relative to real hardware, with gem5 values also matching the 'Expected' latencies from the ISA manuals. The paper also introduces a --mfma-scale configuration parameter for what-if analysis of MFMA latency. The manuscript claims that these changes enable running state-of-the-art ML workloads in gem5, but the validation is limited to microbenchmarks.

Significance. The work fills a real gap in gem5's GPU support: MCEs are a central feature of modern AMD accelerators, and their absence prevented faithful simulation of ML workloads. The paper's strengths are its validation against real MI200 and MI300 hardware, its integration into gem5's mainline public codebase, and its introduction of a configurable latency scaling knob. The reported accuracy (1.5%/1.3% MAPE) is compelling for the tested instruction subset, and the alignment of real-hardware measurements with the ISA manual's expected latencies lends credibility. However, the paper's central motivational claim—that this support 'enables running state-of-the-art ML workloads'—is not backed by any end-to-end workload experiment, and the validation methodology's dependence on fixed overhead constants is not discussed. These issues temper the otherwise solid engineering contribution.

major comments (2)
  1. [Abstract, Section I, Section IV-B] The abstract and introduction state that the changes 'enable running state-of-the-art ML workloads in gem5,' and Section IV-B asserts 'our MFMA support also works for these larger workloads,' yet no ML workload (e.g., PyTorch, TensorFlow, or any real DNN) is run or validated in the paper. The validation is entirely microbenchmark-based. Since the ability to run ML workloads is a primary stated motivation and a load-bearing claim of the contribution, the authors should either add an end-to-end demonstration (even a single representative model) or explicitly soften these claims to indicate that such support is expected but not yet validated.
  2. [Section IV-C, Eq. (1), Tables II-V] The validation hinges on Equation (1), which assumes a fixed 44-cycle overhead (Tmemtime=40, Tinst=4) taken from prior work. The paper does not analyze how sensitive the extracted latencies are to these constants, nor does it justify that the overhead is identical for padded and unpadded instruction sequences (the blue rows in Tables II-V required extra s_nop padding to avoid I-cache misses). The observed agreement of real-hardware measurements with the 'Expected' columns in Tables II and IV partially mitigates this concern, because incorrect overhead constants would systematically shift the extracted values away from the manual's expected latencies. Nevertheless, a brief sensitivity analysis (e.g., varying Tmemtime/Tinst over a plausible range and reporting the resulting MAPE) or a direct per-test justification of the overhead would substantially strengthen the central accuracy claim.
minor comments (4)
  1. [Listing 1] The code snippet in Listing 1 appears to have formatting issues (e.g., incomplete operands in the asm lines). Please ensure the listing is complete and compilable as presented.
  2. [Table I] Table I uses 'LI Instruction Cache' and 'LI Scalar Cache'; the 'LI' is likely intended to be 'L1' for consistency with the other row labels.
  3. [Section V-A] The statement 'we provide accurate timing models' should be qualified as 'for the tested MFMA instructions,' since the validation does not cover all MFMA variants (e.g., those using s_set_gpr_idx are explicitly unsupported).
  4. [Reproducibility] The paper says the gem5 changes are in 'mainline public support' but does not provide a commit hash, patch, or artifact link. Adding a specific revision or repository URL would improve reproducibility for researchers who want to use or extend the MFMA support.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity; the gem5 MFMA timings are sourced from AMD's ISA manual and validated against independently measured hardware, with only a minor self-cited scalar-timing constant in the measurement formula.

full rationale

The central validation chain is self-contained with respect to MFMA timing. Section III sets gem5's MFMA latencies from AMD's reported MCE operation counts and the MI300 ISA manual (Table 27), not from the paper's own measurements. Section IV-C extracts per-MFMA latencies on real hardware using Eq. 1 with T_memtime=40 and T_inst=4, constants adopted from prior work [35]-[37]; these are scalar-pipeline timing values, not MFMA latencies, and they are not fitted to the target results reported here. The MAPE values in Section V therefore compare an externally sourced gem5 model against independently extracted hardware timings (and against the same ISA manual), which is a consistency check rather than a derivation that reduces to its inputs. The self-citation of the two constants is real but minor: if those constants were inaccurate, the extracted hardware latencies could be biased, but that is a measurement-validity risk, not a circular reduction, since the gem5 MFMA model does not depend on Eq. 1. No step in the paper defines MFMA latency in terms of the validation target or fits a parameter to the reported data and then re-predicts it.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The central contribution is an implementation plus validation, so the axiom ledger is short. The main free parameters are two timing constants imported from prior work. The key domain assumptions concern MCE pipelining behavior and the timing methodology.

free parameters (2)
  • Tmemtime = 40 cycles
    Assumed latency of the s_memtime instruction, taken from prior work and used in Equation 1 to compute MFMA latency. It is a calibrated constant rather than a number fitted in this paper, but it affects every reported latency.
  • Tinst = 4 cycles
    Assumed overhead for the non-MFMA instructions in the timing loop, from prior work, used in Equation 1. Together with Tmemtime, it directly shifts the derived MFMA cycle counts.
assumptions (3)
  • domain assumption MFMA instructions from one wavefront cannot be pipelined in the MCE; the scoreboard blocks concurrent MFMAs on the same SIMD unit.
    Section III states this is based on how the AMD compiler appears to behave (inserting NOPs) rather than a direct hardware measurement. If real hardware pipelines MFMAs, the gem5 model would be conservative for hand-written code, though still accurate for compiler-generated code.
  • domain assumption The s_memtime timing instruction does not wait for the last MFMA in the sequence, so subtracting Tmemtime and Tinst isolates the MFMA latency.
    Section IV-C relies on this to derive Equation 1. If the scalar pipeline and MCE interaction differ, the measured latency would be biased.
  • domain assumption AMD's reported MCE operations per clock imply 1 MCE per SIMD unit and 4 per CU for MI200 and MI300.
    Section III uses this to configure the number of MCEs in the model. If the true MCE count differs, the concurrency behavior under mixed workloads would change.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Adding MFMA Support to gem5." pith.science (2026). https://pith.science/paper/V3WDOF6B

@misc{pith2026250118113,
  author       = {Pith},
  title        = {Pith review of: Adding MFMA Support to gem5},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V3WDOF6B}},
  note         = {Machine review of arXiv:2501.18113}
}
read the original abstract

In this work we have enhanced gem5's GPU model support to add Matrix Core Engines (MCEs). Specifically, on the AMD MI200 and MI300 GPUs that gem5 supports, these MCEs perform Matrix Fused Multiply Add (MFMA) instructions for a variety of precisions. By adding this support, our changes enable running state-of-the-art ML workloads in gem5, as well as examining how MCE optimizations impact the behavior of future systems.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

45 extracted references · 35 canonical work pages

  1. [1]

    Megatron-LM: Training Multi-Billion Parameter Lan guage Mod- els Using Model Parallelism,

    M. Shoeybi, M. Patwary, R. Puri, P . LeGresley, J. Casper, and B. Catan- zaro, “Megatron-LM: Training Multi-Billion Parameter Lan guage Mod- els Using Model Parallelism,” CoRR, vol. abs/1909.08053, 2019

  2. [2]

    AI and Memory Wall,

    A. Gholami, Z. Y ao, S. Kim, C. Hooper, M. W. Mahoney, and K. Keutzer, “AI and Memory Wall,” IEEE Micro , vol. 44, no. 03, pp. 33–39, May 2024. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/MM.2024.3373763

  3. [3]

    Pioneering Chiplet Technology and Design for the AMD EPYC™ and Ryzen™ Processor Families : Industrial Produc t,

    S. Naffziger, N. Beck, T. Burd, K. Lepak, G. H. Loh, M. Subr amony, and S. White, “Pioneering Chiplet Technology and Design for the AMD EPYC™ and Ryzen™ Processor Families : Industrial Produc t,” in ACM/IEEE 48th Annual International Symposium on Computer A rchi- tecture, ser. ISCA. New Y ork, NY , USA: Association for Computing Machinery, 2021, pp. 57–70

  4. [4]

    TPU v4: An Optically Reconfigur able Supercomputer for Machine Learning with Hardware Support f or Embeddings,

    N. Jouppi, G. Kurian, S. Li, P . Ma, R. Nagarajan, L. Nai, N. Patil, S. Subramanian, A. Swing, B. Towles, C. Y oung, X. Zh ou, Z. Zhou, and D. A. Patterson, “TPU v4: An Optically Reconfigur able Supercomputer for Machine Learning with Hardware Support f or Embeddings,” in Proceedings of the 50th Annual International Symposium on Computer Architecture , ser...

  5. [5]

    Ten Lessons f rom Three Generations Shaped Google’s TPUv4i,

    N. P . Jouppi, D. H. Y oon, M. Ashcraft, M. Gottscho, T. B. Ja blin, G. Kurian, J. Laudon, S. Li, P . Ma, X. Ma, T. Norrie, N. Patil, S. Prasad, C. Y oung, Z. Zhou, and D. Patterson, “Ten Lessons f rom Three Generations Shaped Google’s TPUv4i,” in Proceedings of the 48th Annual International Symposium on Computer Architect ure, ser. ISCA. Piscataway, NJ, ...

  6. [6]

    The gem5 simul ator,

    N. Binkert, B. Beckmann, G. Black, S. K. Reinhardt, A. Sai di, A. Basu, J. Hestness, D. R. Hower, T. Krishna, S. Sardashti, R. Sen, K. Sewell, M. Shoaib, N. V aish, M. D. Hill, and D. A. Wood, “The gem5 simul ator,” ACM SIGARCH Computer Architecture News , vol. 39, no. 2, pp. 1–7, 2011

  7. [7]

    The gem5 simulator: V ersion 20.0+,

    J. Lowe-Power, A. M. Ahmad, A. Akram, M. Alian, R. Amsling er, M. Andreozzi, A. Armejach, N. Asmussen, S. Bharadwaj, G. Bla ck, G. Bloom, B. R. Bruce, D. R. Carvalho, J. Castrillon, L. Chen, N. Derumigny, S. Diestelhorst, W. Elsasser, M. Fariborz, A. Farmahini- Farahani, P . Fotouhi, R. Gambord, J. Gandhi, D. Gope, T. Gras s, B. Hanindhito, A. Hansson, S....

  8. [8]

    L ost in Abstraction: Pitfalls of Analyzing GPUs at the Intermediat e Language Level,

    A. Gutierrez, B. M. Beckmann, A. Dutu, J. Gross, M. LeBean e, J. Kalamatianos, O. Kayiran, M. Poremba, B. Potter, S. Putho or, M. D. Sinclair, M. Wyse, J. Yin, X. Zhang, A. Jain, and T. Rogers, “L ost in Abstraction: Pitfalls of Analyzing GPUs at the Intermediat e Language Level,” in IEEE International Symposium on High Performance Com- puter Architecture...

Show all 45 references
  1. [9]

    Modeling Modern GPU Applic ations in gem5,

    K. Roarty and M. D. Sinclair, “Modeling Modern GPU Applic ations in gem5,” in 3rd gem5 Users’ W orkshop , June 2020

  2. [10]

    gem5 -SALAM: A System Architecture for LL VM-based Accelerator Modeling ,

    S. Rogers, J. Slycord, M. Baharani, and H. Tabkhi, “gem5 -SALAM: A System Architecture for LL VM-based Accelerator Modeling ,” in 53rd Annual IEEE/ACM International Symposium on Microarch itecture, 2020, pp. 471–482

  3. [11]

    Expan ding Hardware Accelerator System Design Space Exploration with gem5-SAL AMv2,

    Z. Spencer, S. Rogers, J. Slycord, and H. Tabkhi, “Expan ding Hardware Accelerator System Design Space Exploration with gem5-SAL AMv2,” Journal of Systems Architecture , vol. 154, p. 103211, 2024

  4. [12]

    Enablin g Multi-GPU Support in gem5,

    B. W. Y ogatama, M. D. Sinclair, and M. M. Swift, “Enablin g Multi-GPU Support in gem5,” in 3rd gem5 Users’ W orkshop , June 2020

  5. [13]

    On the Rise of AMD Matrix Cores: Performance, Power Efficiency, and Programmability,

    G. Schieffer, D. A. De Medeiros, J. Faj, A. Marathe, and I . Peng, “On the Rise of AMD Matrix Cores: Performance, Power Efficiency, and Programmability,” in IEEE International Symposium on Performance Analysis of Systems and Software , ser. ISPASS, 2024, pp. 132–143

  6. [14]

    V olta: Performa nce and Pro- grammability,

    J. Choquette, O. Giroux, and D. Foley, “V olta: Performa nce and Pro- grammability,” IEEE Micro, vol. 38, no. 2, pp. 42–52, 2018

  7. [15]

    Modeling Deep Le arning Accelerator Enabled GPUs,

    M. A. Raihan, N. Goli, and T. M. Aamodt, “Modeling Deep Le arning Accelerator Enabled GPUs,” in IEEE International Symposium on Performance Analysis of Systems and Software , ser. ISPASS, IEEE. Piscataway, NJ, USA: IEEE Press, 2019, pp. 79–92

  8. [16]

    A Research Retrospec tive on AMD’s Exascale Computing Journey,

    G. H. Loh, M. J. Schulte, M. Ignatowski, V . Adhinarayana n, S. Aga, D. Aguren, V . Agrawal, A. M. Aji, J. Alsop, P . Bauman, B. M. Beckmann, M. V . Beigi, S. Blagodurov, T. Boraten, M. Boyer, W . C. Brantley, N. Chalmers, S. Chen, K. Cheng, M. L. Chu, D. Cownie , N. Curtis, J...

  9. [17]

    Realizing the AMD Exascale Heterogeneous Processor Vision : Industry Product,

    A. Smith, G. H. Loh, M. J. Schulte, M. Ignatowski, S. Naff ziger, M. Mantor, N. Kalyanasundharam, V . Alla, N. Malaya, J. L. Gre athouse, E. Chapman, and R. Swaminathan, “Realizing the AMD Exascale Heterogeneous Processor Vision : Industry Product,” in 51st ACM/IEEE Annual Int...

  10. [18]

    Co- designing Accelerators and SoC Interfaces using gem5-Alad din,

    Y . S. Shao, S. L. Xi, V . Srinivasan, G.-Y . Wei, and D. Broo ks, “Co- designing Accelerators and SoC Interfaces using gem5-Alad din,” in 49th Annual IEEE/ACM International Symposium on Microarchitec ture, ser. MICRO, 2016, pp. 1–12

  11. [19]

    Enabling Repro ducible and Agile Full-System Simulation,

    B. R. Bruce, A. Akram, H. Nguyen, K. Roarty, M. Samani, M. Fariborz, T. Reddy, M. D. Sinclair, and J. Lowe-Power, “Enabling Repro ducible and Agile Full-System Simulation,” in IEEE International Symposium on Performance Analysis of Systems and Software , ser. ISPASS. Los Alami...

  12. [20]

    DNNMark: A Deep Neural Network Ben chmark Suite for GPUs,

    S. Dong and D. Kaeli, “DNNMark: A Deep Neural Network Ben chmark Suite for GPUs,” in Proceedings of the General Purpose GPUs , ser. GPGPU. New Y ork, NY , USA: ACM, 2017, pp. 63–72. [Online]. Available: http://doi.acm.org/10.1145/3038228.3038239

  13. [21]

    An Update to DeepBench with a Fo cus on Deep Learning Inference,

    S. Narang and G. Diamos, “An Update to DeepBench with a Fo cus on Deep Learning Inference,” https://svail.github.io/De epBench-update/, 2017

  14. [22]

    MIOpen: An Open Source Li brary For Deep Learning Primitives,

    J. Khan, P . Fultz, A. Tamazov, D. Lowell, C. Liu, M. Meles se, M. Nand- himandalam, K. Nasyrov, I. Perminov, T. Shah, V . Filippov, J . Zhang, J. Zhou, B. Natarajan, and M. Daga, “MIOpen: An Open Source Li brary For Deep Learning Primitives,” CoRR, vol. abs/1910.00078, 2019

  15. [23]

    rocBLAS Library,

    AMD, “rocBLAS Library,” https://rocm-documentation .readthedocs.io/en/latest/ROCm T 2024

  16. [24]

    PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation a nd Graph Compilation,

    J. Ansel, E. Y ang, H. He, N. Gimelshein, A. Jain, M. V ozne sensky, B. Bao, P . Bell, D. Berard, E. Burovski, G. Chauhan, A. Chourd ia, W. Constable, A. Desmaison, Z. DeVito, E. Ellison, W. Feng, J. Gong, M. Gschwind, B. Hirsh, S. Huang, K. Kalambarkar, L. Kirsch, M. Lazos, M...

  17. [25]

    Automatic different iation in PyTorch,

    A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Y ang, Z. D eVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic different iation in PyTorch,” in NeurIPS-W, 2017

  18. [26]

    TensorFlow: La rge- Scale Machine Learning on Heterogeneous Systems,

    M. Abadi, A. Agarwal, P . Barham, E. Brevdo, Z. Chen, C. Ci tro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfel low, A. Harp, G. Irving, M. Isard, Y . Jia, R. Jozefowicz, L. Kaiser , M. Kudlur, J. Levenberg, D. Man´ e, R. Monga, S. Moore, D. Murray, C. Ola...

  19. [27]

    Simulation Support for Fast and Accurate Large-Scale GPGPU & Accelerat or Workloads,

    V . Ramadas, M. Poremba, B. Beckmann, and M. D. Sinclair, “Simulation Support for Fast and Accurate Large-Scale GPGPU & Accelerat or Workloads,” in Third W orkshop on Open-Source Computer Architecture Research, ser. OSCAR, June 2024

  20. [28]

    Simulating Machine Lear ning Models at Scale,

    V . Ramadas and M. D. Sinclair, “Simulating Machine Lear ning Models at Scale,” in SRC TECHCON , September 2024

  21. [29]

    AMD CDNA™ 3 Architecture,

    AMD, “AMD CDNA™ 3 Architecture,” 2023. [Online]. Avail able: https://www.amd.com/content/dam/amd/en/documents/instinct-tech-docs/white-papers/amd-cdna-3-white-pape r.pdf

  22. [30]

    [Online]

    Advanced Micro Devices (AMD), AMD Instinct MI300 Instruction Set Architecture Reference Guide , June 2024. [Online]. Available: https://www.amd.com/content/dam/amd/en/documents/instinct-tech-docs/instruction-set-architectures/amd- instinct-mi300-cdna3-instruction-set-architecture.p df

  23. [31]

    NVIDIA Hopper H100 GPU: Scaling Perform ance,

    J. Choquette, “NVIDIA Hopper H100 GPU: Scaling Perform ance,” IEEE Micro, vol. 43, no. 03, pp. 9–17, May 2023. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/MM.2023.3256796

  24. [32]

    Cooperative War p Execution in Tensor Core for RISC-V GPGPU,

    A. Nada, G. M. Sarda, and E. Lenormand, “Cooperative War p Execution in Tensor Core for RISC-V GPGPU,” in IEEE 31st International Symposium on High Performance Computer Architecture , ser. HPCA, 2025

  25. [33]

    CPElide : Efficient Multi-Chiplet GPU Implicit Synchronization,

    P . Dalmia, R. Shashi Kumar, and M. D. Sinclair, “CPElide : Efficient Multi-Chiplet GPU Implicit Synchronization,” in Proceedings of 57th IEEE/ACM International Symposium on Microarchitecture, ser. MICRO. Los Alamitos, CA, USA: IEEE Computer Society, 2024

  26. [34]

    ROCm: Open Platform For Development, Discovery and Education around GPU Computing,

    AMD, “ROCm: Open Platform For Development, Discovery and Education around GPU Computing,” https://gpuopen.com/compute-product/rocm/, 2021

  27. [35]

    GAP: gem5 GPU Accuracy Profiler,

    C. Jamieson, A. Chandrashekar, I. McDougall, and M. D. S inclair, “GAP: gem5 GPU Accuracy Profiler,” in 4th gem5 Users’ W orkshop , June 2022

  28. [36]

    Closing the Gap: Improving the Accuracy of gem5’s GPU Models,

    V . Ramadas, D. Kouchekinia, N. Osuji, and M. D. Sinclair , “Closing the Gap: Improving the Accuracy of gem5’s GPU Models,” in 5th gem5 Users’ W orkshop, June 2023

  29. [37]

    Furthe r Closing the GAP: Improving the Accuracy of gem5’s GPU Models,

    V . Ramadas, D. Kouchekinia, and M. D. Sinclair, “Furthe r Closing the GAP: Improving the Accuracy of gem5’s GPU Models,” in 6th Young Architects’ W orkshop, ser. Y Arch, April 2024

  30. [38]

    Introducing AMD CDNA™ 2 Architecture,

    AMD, “Introducing AMD CDNA™ 2 Architecture,” https://www.amd.com/content/dam/amd/en/documents/instinct-business-docs/white-papers/amd-cdna2-white-p aper.pdf, 2022

  31. [39]

    AMD AI & HPC Fund: Accelerate Y our Research with AMD,

    ——, “AMD AI & HPC Fund: Accelerate Y our Research with AMD,” 2024. [Online]. Available: https://www.amd.com/en/corporate/hpc-fund.html

  32. [40]

    AMD lab notes,

    ——, “AMD lab notes,” https://github.com/amd/amd-lab -notes/ , 2024

  33. [41]

    Language Models are Few-Shot Learners,

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P . D hariwal, A. Neelakantan, P . Shyam, G. Sastry, A. Askell, S. Agarwal, A . Herbert- V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegle r, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray,...

  34. [42]

    Language Models are Unsupervised Multitask Learners,

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sut skever, “Language Models are Unsupervised Multitask Learners,” OpenAI Blog, vol. 1, no. 8, 2019

  35. [43]

    GPT-4 Technical Report,

    OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akk aya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anad kat, R. Avila, I. Babuschkin, S. Balaji, V . Balcom, P . Baltescu, H . Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-S hapiro, C. Bern...

  36. [44]

    MLPerf Infere nce Benchmark,

    V . J. Reddi, C. Cheng, D. Kanter, P . Mattson, G. Schmuell ing, C.-J. Wu, B. Anderson, M. Breughe, M. Charlebois, W. Chou, R. Chukk a, C. Coleman, S. Davis, P . Deng, G. Diamos, J. Duke, D. Fick, J. S . Gardner, I. Hubara, S. Idgunji, T. B. Jablin, J. Jiao, T. S. Jo hn, P . K...

  37. [45]

    MLPerf Training Benchmark,

    P . Mattson, C. Cheng, G. Diamos, C. Coleman, P . Micikevi cius, D. Patterson, H. Tang, G.-Y . Wei, P . Bailis, V . Bittorf, D. Br ooks, D. Chen, D. Dutta, U. Gupta, K. Hazelwood, A. Hock, X. Huang, D. Kang, D. Kanter, N. Kumar, J. Liao, D. Narayanan, T. Ogunte bi, G. Pekhimen...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.