REVIEW 2 major objections 4 minor 45 references
Adding MFMA Support to gem5
T0 review · 2 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper reports that adding Matrix Core Engine support to gem5 times AMD MI200 and MI300 MFMA instructions within 1.5% and 1.3% of real hardware, making cycle-level ML workload simulation practical.
desk verdict A solid, narrowly-scoped gem5 MFMA support paper with credible microbenchmark validation; the broad ML-workload claims outrun the evidence, but the core contribution deserves review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a scoreboard-based MFMA issue rule in gem5's compute-unit timing logic, backed by a per-instruction latency lookup table (mfma_cycles) derived from the MI300 ISA manual and validated by Equation 1: T_MFMA = (T_total - T_memtime - T_inst)/(N_MFMA - 1), where T_memtime = 40 and T_inst = 4 are overhead constants from prior calibration. The scoreboard's NRDY_MATRIX_CORE field per SIMD unit decides when the MCE is free; since the scalar pipeline that runs s_memtime is independent of the MCE pipeline, dependent MFMA chains are needed to make the timing instruction wait for matrix completion.
What would settle it
Run the identical dependent-MFMA timing kernels on an MI200 or MI300 while replacing the s_memtime overhead constants with directly measured values, for example by timing a single s_memtime around a known number of s_nop instructions, and compare the resulting per-MFMA latencies; a systematic shift larger than the reported MAPE would indicate Equation 1's constants are wrong. Separately, launch two independent MFMA streams from different wavefronts to the same MCE and measure whether throughput exceeds one MFMA per latency, which would falsify the non-pipelining assumption.
Extended reading notes
Core claim
On the paper's own terms: gem5's GPU model now includes MCEs as separate functional units, one per SIMD unit (four per compute unit), and uses the NRDY_MATRIX_CORE scoreboard field to prevent more than one MFMA on the same SIMD unit at a time. This design matches AMD's reported MCE throughput and mirrors the compiler's apparent assumption that MFMA instructions from a given wavefront are not pipelined in the MCE. The validation method times back-to-back dependent MFMA instructions with s_memtime and subtracts measured overheads (T_memtime = 40 cycles, T_inst = 4 cycles) to recover per-instruction latency; across the tested instructions, gem5 matches the ISA manual's expected cycle counts and real hardware within 1.5% (MI200) and 1.3% (MI300) mean absolute percentage error. The paper also adds a --mfma-scale parameter that scales MFMA latency to support what-if analysis of future MCE designs.
Load-bearing premise
The load-bearing premise is that the fixed timing overheads (T_memtime = 40 cycles, T_inst = 4 cycles) and the assumption that a wavefront's MFMA instructions never pipeline in the MCE correctly describe real MI200 and MI300 hardware; if either is wrong, the recovered per-instruction latencies, and therefore the claimed 1.5% and 1.3% accuracy, would be biased.
Editorial extensions
If this is right
- Modern ML workloads that call MFMA-based libraries can now be run in gem5 with cycle-level timing rather than only functional simulation.
- Researchers can model MCE improvements, faster or slower, by setting --mfma-scale, which multiplies the per-instruction MFMA latency and exposes how application runtime responds.
- MI200 and MI300 differences, including new instructions, removed instructions, and changed latencies, are captured in the gem5 model, so cross-generation comparisons can be simulated.
- Validation error decreases as more MFMA instructions are timed back-to-back, from 2.3% MAPE at 2 MFMAs to 0.4% at 5 on MI200, suggesting the timing methodology is limited by transient effects rather than by the core model.
Reading between the lines
- Because the non-pipelining assumption is inferred from compiler-inserted NOPs rather than measured hardware behavior, the gem5 model may need revision if AMD's compiler or future hardware starts pipelining MFMA instructions from the same wavefront; a direct microbenchmark that varies the amount of independent work between dependent MFMAs would settle this.
- The same interleaved s_memtime harness could serve as a portable benchmark for measuring MFMA latencies on future AMD GPUs and refreshing the mfma_cycles table, since it depends only on the ISA, not on gem5 internals.
- The paper's limitation note implies that --mfma-scale results likely overstate how a real compiled workload would respond to faster or slower MCEs, because the compiler inserts NOPs and independent work based on the original latency; scaling studies therefore need compiler co-design to be realistic.
- A similar scoreboard-plus-latency-table structure could be adapted to model NVIDIA TensorCore timing in gem5, although the validation harness would need different timing instructions than s_memtime.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper adds Matrix Core Engine (MCE) support for AMD MI200 and MI300 GPUs to the gem5 simulator, including functional and timing models for MFMA instructions. The authors validate their implementation with handwritten assembly microbenchmarks that interleave back-to-back, dependent MFMA instructions with s_memtime timing instructions, running the same kernels on real GPUs and in gem5. They compute per-MFMA latency using Equation (1), which subtracts a fixed 44-cycle overhead (Tmemtime=40, Tinst=4) and divides by N_MFMA−1. The reported results show 1.5% MAPE for MI200 and 1.3% MAPE for MI300 relative to real hardware, with gem5 values also matching the 'Expected' latencies from the ISA manuals. The paper also introduces a --mfma-scale configuration parameter for what-if analysis of MFMA latency. The manuscript claims that these changes enable running state-of-the-art ML workloads in gem5, but the validation is limited to microbenchmarks.
Significance. The work fills a real gap in gem5's GPU support: MCEs are a central feature of modern AMD accelerators, and their absence prevented faithful simulation of ML workloads. The paper's strengths are its validation against real MI200 and MI300 hardware, its integration into gem5's mainline public codebase, and its introduction of a configurable latency scaling knob. The reported accuracy (1.5%/1.3% MAPE) is compelling for the tested instruction subset, and the alignment of real-hardware measurements with the ISA manual's expected latencies lends credibility. However, the paper's central motivational claim—that this support 'enables running state-of-the-art ML workloads'—is not backed by any end-to-end workload experiment, and the validation methodology's dependence on fixed overhead constants is not discussed. These issues temper the otherwise solid engineering contribution.
major comments (2)
- [Abstract, Section I, Section IV-B] The abstract and introduction state that the changes 'enable running state-of-the-art ML workloads in gem5,' and Section IV-B asserts 'our MFMA support also works for these larger workloads,' yet no ML workload (e.g., PyTorch, TensorFlow, or any real DNN) is run or validated in the paper. The validation is entirely microbenchmark-based. Since the ability to run ML workloads is a primary stated motivation and a load-bearing claim of the contribution, the authors should either add an end-to-end demonstration (even a single representative model) or explicitly soften these claims to indicate that such support is expected but not yet validated.
- [Section IV-C, Eq. (1), Tables II-V] The validation hinges on Equation (1), which assumes a fixed 44-cycle overhead (Tmemtime=40, Tinst=4) taken from prior work. The paper does not analyze how sensitive the extracted latencies are to these constants, nor does it justify that the overhead is identical for padded and unpadded instruction sequences (the blue rows in Tables II-V required extra s_nop padding to avoid I-cache misses). The observed agreement of real-hardware measurements with the 'Expected' columns in Tables II and IV partially mitigates this concern, because incorrect overhead constants would systematically shift the extracted values away from the manual's expected latencies. Nevertheless, a brief sensitivity analysis (e.g., varying Tmemtime/Tinst over a plausible range and reporting the resulting MAPE) or a direct per-test justification of the overhead would substantially strengthen the central accuracy claim.
minor comments (4)
- [Listing 1] The code snippet in Listing 1 appears to have formatting issues (e.g., incomplete operands in the asm lines). Please ensure the listing is complete and compilable as presented.
- [Table I] Table I uses 'LI Instruction Cache' and 'LI Scalar Cache'; the 'LI' is likely intended to be 'L1' for consistency with the other row labels.
- [Section V-A] The statement 'we provide accurate timing models' should be qualified as 'for the tested MFMA instructions,' since the validation does not cover all MFMA variants (e.g., those using s_set_gpr_idx are explicitly unsupported).
- [Reproducibility] The paper says the gem5 changes are in 'mainline public support' but does not provide a commit hash, patch, or artifact link. Adding a specific revision or repository URL would improve reproducibility for researchers who want to use or extend the MFMA support.
Circularity Check
No significant circularity; the gem5 MFMA timings are sourced from AMD's ISA manual and validated against independently measured hardware, with only a minor self-cited scalar-timing constant in the measurement formula.
full rationale
The central validation chain is self-contained with respect to MFMA timing. Section III sets gem5's MFMA latencies from AMD's reported MCE operation counts and the MI300 ISA manual (Table 27), not from the paper's own measurements. Section IV-C extracts per-MFMA latencies on real hardware using Eq. 1 with T_memtime=40 and T_inst=4, constants adopted from prior work [35]-[37]; these are scalar-pipeline timing values, not MFMA latencies, and they are not fitted to the target results reported here. The MAPE values in Section V therefore compare an externally sourced gem5 model against independently extracted hardware timings (and against the same ISA manual), which is a consistency check rather than a derivation that reduces to its inputs. The self-citation of the two constants is real but minor: if those constants were inaccurate, the extracted hardware latencies could be biased, but that is a measurement-validity risk, not a circular reduction, since the gem5 MFMA model does not depend on Eq. 1. No step in the paper defines MFMA latency in terms of the validation target or fits a parameter to the reported data and then re-predicts it.
Assumptions & free parameters
free parameters (2)
- Tmemtime =
40 cycles
- Tinst =
4 cycles
assumptions (3)
- domain assumption MFMA instructions from one wavefront cannot be pipelined in the MCE; the scoreboard blocks concurrent MFMAs on the same SIMD unit.
- domain assumption The s_memtime timing instruction does not wait for the last MFMA in the sequence, so subtracting Tmemtime and Tinst isolates the MFMA latency.
- domain assumption AMD's reported MCE operations per clock imply 1 MCE per SIMD unit and 4 per CU for MI200 and MI300.
Cite this review
Pith. "Pith review of Adding MFMA Support to gem5." pith.science (2026). https://pith.science/paper/V3WDOF6B
@misc{pith2026250118113,
author = {Pith},
title = {Pith review of: Adding MFMA Support to gem5},
year = {2026},
howpublished = {\url{https://pith.science/paper/V3WDOF6B}},
note = {Machine review of arXiv:2501.18113}
}
read the original abstract
In this work we have enhanced gem5's GPU model support to add Matrix Core Engines (MCEs). Specifically, on the AMD MI200 and MI300 GPUs that gem5 supports, these MCEs perform Matrix Fused Multiply Add (MFMA) instructions for a variety of precisions. By adding this support, our changes enable running state-of-the-art ML workloads in gem5, as well as examining how MCE optimizations impact the behavior of future systems.
Reference graph
Works this paper leans on
-
[1]
Megatron-LM: Training Multi-Billion Parameter Lan guage Mod- els Using Model Parallelism,
M. Shoeybi, M. Patwary, R. Puri, P . LeGresley, J. Casper, and B. Catan- zaro, “Megatron-LM: Training Multi-Billion Parameter Lan guage Mod- els Using Model Parallelism,” CoRR, vol. abs/1909.08053, 2019
arXiv 1909
-
[2]
A. Gholami, Z. Y ao, S. Kim, C. Hooper, M. W. Mahoney, and K. Keutzer, “AI and Memory Wall,” IEEE Micro , vol. 44, no. 03, pp. 33–39, May 2024. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/MM.2024.3373763
arXiv 2024
-
[3]
S. Naffziger, N. Beck, T. Burd, K. Lepak, G. H. Loh, M. Subr amony, and S. White, “Pioneering Chiplet Technology and Design for the AMD EPYC™ and Ryzen™ Processor Families : Industrial Produc t,” in ACM/IEEE 48th Annual International Symposium on Computer A rchi- tecture, ser. ISCA. New Y ork, NY , USA: Association for Computing Machinery, 2021, pp. 57–70
work page 2021
-
[4]
N. Jouppi, G. Kurian, S. Li, P . Ma, R. Nagarajan, L. Nai, N. Patil, S. Subramanian, A. Swing, B. Towles, C. Y oung, X. Zh ou, Z. Zhou, and D. A. Patterson, “TPU v4: An Optically Reconfigur able Supercomputer for Machine Learning with Hardware Support f or Embeddings,” in Proceedings of the 50th Annual International Symposium on Computer Architecture , ser...
arXiv 2023
-
[5]
Ten Lessons f rom Three Generations Shaped Google’s TPUv4i,
N. P . Jouppi, D. H. Y oon, M. Ashcraft, M. Gottscho, T. B. Ja blin, G. Kurian, J. Laudon, S. Li, P . Ma, X. Ma, T. Norrie, N. Patil, S. Prasad, C. Y oung, Z. Zhou, and D. Patterson, “Ten Lessons f rom Three Generations Shaped Google’s TPUv4i,” in Proceedings of the 48th Annual International Symposium on Computer Architect ure, ser. ISCA. Piscataway, NJ, ...
arXiv 2021
-
[6]
N. Binkert, B. Beckmann, G. Black, S. K. Reinhardt, A. Sai di, A. Basu, J. Hestness, D. R. Hower, T. Krishna, S. Sardashti, R. Sen, K. Sewell, M. Shoaib, N. V aish, M. D. Hill, and D. A. Wood, “The gem5 simul ator,” ACM SIGARCH Computer Architecture News , vol. 39, no. 2, pp. 1–7, 2011
work page 2011
-
[7]
The gem5 simulator: V ersion 20.0+,
J. Lowe-Power, A. M. Ahmad, A. Akram, M. Alian, R. Amsling er, M. Andreozzi, A. Armejach, N. Asmussen, S. Bharadwaj, G. Bla ck, G. Bloom, B. R. Bruce, D. R. Carvalho, J. Castrillon, L. Chen, N. Derumigny, S. Diestelhorst, W. Elsasser, M. Fariborz, A. Farmahini- Farahani, P . Fotouhi, R. Gambord, J. Gandhi, D. Gope, T. Gras s, B. Hanindhito, A. Hansson, S....
work page 2020
-
[8]
L ost in Abstraction: Pitfalls of Analyzing GPUs at the Intermediat e Language Level,
A. Gutierrez, B. M. Beckmann, A. Dutu, J. Gross, M. LeBean e, J. Kalamatianos, O. Kayiran, M. Poremba, B. Potter, S. Putho or, M. D. Sinclair, M. Wyse, J. Yin, X. Zhang, A. Jain, and T. Rogers, “L ost in Abstraction: Pitfalls of Analyzing GPUs at the Intermediat e Language Level,” in IEEE International Symposium on High Performance Com- puter Architecture...
work page 2018
Show all 45 references
-
[9]
Modeling Modern GPU Applic ations in gem5,
K. Roarty and M. D. Sinclair, “Modeling Modern GPU Applic ations in gem5,” in 3rd gem5 Users’ W orkshop , June 2020
2020
-
[10]
gem5 -SALAM: A System Architecture for LL VM-based Accelerator Modeling ,
S. Rogers, J. Slycord, M. Baharani, and H. Tabkhi, “gem5 -SALAM: A System Architecture for LL VM-based Accelerator Modeling ,” in 53rd Annual IEEE/ACM International Symposium on Microarch itecture, 2020, pp. 471–482
2020
-
[11]
Expan ding Hardware Accelerator System Design Space Exploration with gem5-SAL AMv2,
Z. Spencer, S. Rogers, J. Slycord, and H. Tabkhi, “Expan ding Hardware Accelerator System Design Space Exploration with gem5-SAL AMv2,” Journal of Systems Architecture , vol. 154, p. 103211, 2024
2024
-
[12]
Enablin g Multi-GPU Support in gem5,
B. W. Y ogatama, M. D. Sinclair, and M. M. Swift, “Enablin g Multi-GPU Support in gem5,” in 3rd gem5 Users’ W orkshop , June 2020
2020
-
[13]
On the Rise of AMD Matrix Cores: Performance, Power Efficiency, and Programmability,
G. Schieffer, D. A. De Medeiros, J. Faj, A. Marathe, and I . Peng, “On the Rise of AMD Matrix Cores: Performance, Power Efficiency, and Programmability,” in IEEE International Symposium on Performance Analysis of Systems and Software , ser. ISPASS, 2024, pp. 132–143
2024
-
[14]
V olta: Performa nce and Pro- grammability,
J. Choquette, O. Giroux, and D. Foley, “V olta: Performa nce and Pro- grammability,” IEEE Micro, vol. 38, no. 2, pp. 42–52, 2018
2018
-
[15]
Modeling Deep Le arning Accelerator Enabled GPUs,
M. A. Raihan, N. Goli, and T. M. Aamodt, “Modeling Deep Le arning Accelerator Enabled GPUs,” in IEEE International Symposium on Performance Analysis of Systems and Software , ser. ISPASS, IEEE. Piscataway, NJ, USA: IEEE Press, 2019, pp. 79–92
2019
-
[16]
A Research Retrospec tive on AMD’s Exascale Computing Journey,
G. H. Loh, M. J. Schulte, M. Ignatowski, V . Adhinarayana n, S. Aga, D. Aguren, V . Agrawal, A. M. Aji, J. Alsop, P . Bauman, B. M. Beckmann, M. V . Beigi, S. Blagodurov, T. Boraten, M. Boyer, W . C. Brantley, N. Chalmers, S. Chen, K. Cheng, M. L. Chu, D. Cownie , N. Curtis, J...
-
[17]
Realizing the AMD Exascale Heterogeneous Processor Vision : Industry Product,
A. Smith, G. H. Loh, M. J. Schulte, M. Ignatowski, S. Naff ziger, M. Mantor, N. Kalyanasundharam, V . Alla, N. Malaya, J. L. Gre athouse, E. Chapman, and R. Swaminathan, “Realizing the AMD Exascale Heterogeneous Processor Vision : Industry Product,” in 51st ACM/IEEE Annual Int...
2024
-
[18]
Co- designing Accelerators and SoC Interfaces using gem5-Alad din,
Y . S. Shao, S. L. Xi, V . Srinivasan, G.-Y . Wei, and D. Broo ks, “Co- designing Accelerators and SoC Interfaces using gem5-Alad din,” in 49th Annual IEEE/ACM International Symposium on Microarchitec ture, ser. MICRO, 2016, pp. 1–12
2016
-
[19]
Enabling Repro ducible and Agile Full-System Simulation,
B. R. Bruce, A. Akram, H. Nguyen, K. Roarty, M. Samani, M. Fariborz, T. Reddy, M. D. Sinclair, and J. Lowe-Power, “Enabling Repro ducible and Agile Full-System Simulation,” in IEEE International Symposium on Performance Analysis of Systems and Software , ser. ISPASS. Los Alami...
2021
-
[20]
DNNMark: A Deep Neural Network Ben chmark Suite for GPUs,
S. Dong and D. Kaeli, “DNNMark: A Deep Neural Network Ben chmark Suite for GPUs,” in Proceedings of the General Purpose GPUs , ser. GPGPU. New Y ork, NY , USA: ACM, 2017, pp. 63–72. [Online]. Available: http://doi.acm.org/10.1145/3038228.3038239
2017
-
[21]
An Update to DeepBench with a Fo cus on Deep Learning Inference,
S. Narang and G. Diamos, “An Update to DeepBench with a Fo cus on Deep Learning Inference,” https://svail.github.io/De epBench-update/, 2017
2017
-
[22]
MIOpen: An Open Source Li brary For Deep Learning Primitives,
J. Khan, P . Fultz, A. Tamazov, D. Lowell, C. Liu, M. Meles se, M. Nand- himandalam, K. Nasyrov, I. Perminov, T. Shah, V . Filippov, J . Zhang, J. Zhou, B. Natarajan, and M. Daga, “MIOpen: An Open Source Li brary For Deep Learning Primitives,” CoRR, vol. abs/1910.00078, 2019
1910 arXiv
-
[23]
rocBLAS Library,
AMD, “rocBLAS Library,” https://rocm-documentation .readthedocs.io/en/latest/ROCm T 2024
2024
-
[24]
PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation a nd Graph Compilation,
J. Ansel, E. Y ang, H. He, N. Gimelshein, A. Jain, M. V ozne sensky, B. Bao, P . Bell, D. Berard, E. Burovski, G. Chauhan, A. Chourd ia, W. Constable, A. Desmaison, Z. DeVito, E. Ellison, W. Feng, J. Gong, M. Gschwind, B. Hirsh, S. Huang, K. Kalambarkar, L. Kirsch, M. Lazos, M...
2024
-
[25]
Automatic different iation in PyTorch,
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Y ang, Z. D eVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic different iation in PyTorch,” in NeurIPS-W, 2017
2017
-
[26]
TensorFlow: La rge- Scale Machine Learning on Heterogeneous Systems,
M. Abadi, A. Agarwal, P . Barham, E. Brevdo, Z. Chen, C. Ci tro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfel low, A. Harp, G. Irving, M. Isard, Y . Jia, R. Jozefowicz, L. Kaiser , M. Kudlur, J. Levenberg, D. Man´ e, R. Monga, S. Moore, D. Murray, C. Ola...
2015
-
[27]
Simulation Support for Fast and Accurate Large-Scale GPGPU & Accelerat or Workloads,
V . Ramadas, M. Poremba, B. Beckmann, and M. D. Sinclair, “Simulation Support for Fast and Accurate Large-Scale GPGPU & Accelerat or Workloads,” in Third W orkshop on Open-Source Computer Architecture Research, ser. OSCAR, June 2024
2024
-
[28]
Simulating Machine Lear ning Models at Scale,
V . Ramadas and M. D. Sinclair, “Simulating Machine Lear ning Models at Scale,” in SRC TECHCON , September 2024
2024
-
[29]
AMD CDNA™ 3 Architecture,
AMD, “AMD CDNA™ 3 Architecture,” 2023. [Online]. Avail able: https://www.amd.com/content/dam/amd/en/documents/instinct-tech-docs/white-papers/amd-cdna-3-white-pape r.pdf
2023
-
[30]
[Online]
Advanced Micro Devices (AMD), AMD Instinct MI300 Instruction Set Architecture Reference Guide , June 2024. [Online]. Available: https://www.amd.com/content/dam/amd/en/documents/instinct-tech-docs/instruction-set-architectures/amd- instinct-mi300-cdna3-instruction-set-architecture.p df
2024
-
[31]
NVIDIA Hopper H100 GPU: Scaling Perform ance,
J. Choquette, “NVIDIA Hopper H100 GPU: Scaling Perform ance,” IEEE Micro, vol. 43, no. 03, pp. 9–17, May 2023. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/MM.2023.3256796
2023
-
[32]
Cooperative War p Execution in Tensor Core for RISC-V GPGPU,
A. Nada, G. M. Sarda, and E. Lenormand, “Cooperative War p Execution in Tensor Core for RISC-V GPGPU,” in IEEE 31st International Symposium on High Performance Computer Architecture , ser. HPCA, 2025
2025
-
[33]
CPElide : Efficient Multi-Chiplet GPU Implicit Synchronization,
P . Dalmia, R. Shashi Kumar, and M. D. Sinclair, “CPElide : Efficient Multi-Chiplet GPU Implicit Synchronization,” in Proceedings of 57th IEEE/ACM International Symposium on Microarchitecture, ser. MICRO. Los Alamitos, CA, USA: IEEE Computer Society, 2024
2024
-
[34]
ROCm: Open Platform For Development, Discovery and Education around GPU Computing,
AMD, “ROCm: Open Platform For Development, Discovery and Education around GPU Computing,” https://gpuopen.com/compute-product/rocm/, 2021
2021
-
[35]
GAP: gem5 GPU Accuracy Profiler,
C. Jamieson, A. Chandrashekar, I. McDougall, and M. D. S inclair, “GAP: gem5 GPU Accuracy Profiler,” in 4th gem5 Users’ W orkshop , June 2022
2022
-
[36]
Closing the Gap: Improving the Accuracy of gem5’s GPU Models,
V . Ramadas, D. Kouchekinia, N. Osuji, and M. D. Sinclair , “Closing the Gap: Improving the Accuracy of gem5’s GPU Models,” in 5th gem5 Users’ W orkshop, June 2023
2023
-
[37]
Furthe r Closing the GAP: Improving the Accuracy of gem5’s GPU Models,
V . Ramadas, D. Kouchekinia, and M. D. Sinclair, “Furthe r Closing the GAP: Improving the Accuracy of gem5’s GPU Models,” in 6th Young Architects’ W orkshop, ser. Y Arch, April 2024
2024
-
[38]
Introducing AMD CDNA™ 2 Architecture,
AMD, “Introducing AMD CDNA™ 2 Architecture,” https://www.amd.com/content/dam/amd/en/documents/instinct-business-docs/white-papers/amd-cdna2-white-p aper.pdf, 2022
2022
-
[39]
AMD AI & HPC Fund: Accelerate Y our Research with AMD,
——, “AMD AI & HPC Fund: Accelerate Y our Research with AMD,” 2024. [Online]. Available: https://www.amd.com/en/corporate/hpc-fund.html
2024
-
[40]
AMD lab notes,
——, “AMD lab notes,” https://github.com/amd/amd-lab -notes/ , 2024
2024
-
[41]
Language Models are Few-Shot Learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P . D hariwal, A. Neelakantan, P . Shyam, G. Sastry, A. Askell, S. Agarwal, A . Herbert- V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegle r, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray,...
2020
-
[42]
Language Models are Unsupervised Multitask Learners,
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sut skever, “Language Models are Unsupervised Multitask Learners,” OpenAI Blog, vol. 1, no. 8, 2019
2019
-
[43]
GPT-4 Technical Report,
OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akk aya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anad kat, R. Avila, I. Babuschkin, S. Balaji, V . Balcom, P . Baltescu, H . Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-S hapiro, C. Bern...
-
[44]
MLPerf Infere nce Benchmark,
V . J. Reddi, C. Cheng, D. Kanter, P . Mattson, G. Schmuell ing, C.-J. Wu, B. Anderson, M. Breughe, M. Charlebois, W. Chou, R. Chukk a, C. Coleman, S. Davis, P . Deng, G. Diamos, J. Duke, D. Fick, J. S . Gardner, I. Hubara, S. Idgunji, T. B. Jablin, J. Jiao, T. S. Jo hn, P . K...
2020
-
[45]
MLPerf Training Benchmark,
P . Mattson, C. Cheng, G. Diamos, C. Coleman, P . Micikevi cius, D. Patterson, H. Tang, G.-Y . Wei, P . Bailis, V . Bittorf, D. Br ooks, D. Chen, D. Dutta, U. Gupta, K. Hazelwood, A. Hock, X. Huang, D. Kang, D. Kanter, N. Kumar, J. Liao, D. Narayanan, T. Ogunte bi, G. Pekhimen...
2020
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.