REVIEW 5 minor 81 references
How to keep pushing ML accelerator performance? Know your rooflines!
T0 review · 0 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read An enhanced two-line roofline model, one for throughput and one for energy efficiency, can unify how ML accelerator designers reason about compute, memory, and data movement, and guide which optimization to pursue.
desk verdict A competent, well-organized survey that repackages the throughput and energy rooflines into a two-line framework for ML accelerators; useful for designers, but not a new research result. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the per-memory-level arithmetic intensity $AI_{L_i}=N_{op}/N_{L_i}$, the number of operations per byte fetched from memory level $L_i$. It carries the argument because both rooflines are functions of this quantity: the throughput roofline of Eq. (4) takes a minimum over the per-level bandwidth products and the compute peak, while the energy roofline of Eq. (5) sums each level's energy per access divided by the per-level intensity. Data reuse, quantization, sparsity, and memory-proximity techniques all enter the model as changes in these intensities or in the roofline parameters, which is what makes the framework unifying.
What would settle it
A concrete test: on a real accelerator, run a sparse tensor kernel and measure its per-level memory traffic, peak bandwidths, and peak MAC rate, then compute the predicted operating points from Eqs. (4) and (5). If the measured throughput and energy efficiency lie far below both rooflines and the gap cannot be explained by the utilization factors described in Section II-B, the claim that the two rooflines bound and explain accelerator efficiency for that workload is refuted. A softer check is to take two architectures with identical rooflines but different dataflow flexibility and show that workload-level efficiency differs in a way the rooflines cannot express.
Extended reading notes
Core claim
The paper's central claim is that attainable throughput is $P_{TP}=f_{\mathrm{clk}}\min(AI_{L_n}B_{L_n},\dots,AI_{L_1}B_{L_1},A_{op})$ and attainable energy efficiency is $P_E=1/(E_{op}+\sum_i E_{L_i}/AI_{L_i})$, where $AI_{L_i}=N_{op}/N_{L_i}$ is the arithmetic intensity toward memory level $L_i$. Plotted together, these two expressions are claimed to explain where each execution regime falls and what action improves it. A distinctive consequence is that the throughput roofline has a sharp knee while the energy roofline is curved, and the two knees can sit at different arithmetic intensities, so an accelerator can be compute-bound for throughput and memory-bound for energy at the same operating point. The paper then reads the major efficiency techniques through this lens, treating sparsity as an intensity-lowering, utilization-changing effect and in-memory computing as a way to raise the compute roofline while eliminating the first-level memory diagonal, at the cost of new storage-compute coupling and utilization losses.
Load-bearing premise
The load-bearing premise is that a workload is well summarized by its arithmetic intensity at each memory level and that compute and memory latency overlap perfectly, so the minimum in the throughput equation captures latency; sparse and irregular workloads violate this, and the paper itself notes that they fall below the roofline and need hardware-specific utilization factors.
Editorial extensions
If this is right
- A designer can classify any proposed optimization as raising the compute roofline, raising a memory roofline, moving the operating point to higher arithmetic intensity, or improving utilization, and can choose the category that addresses the actual bottleneck.
- Because the throughput and energy knees depend on different parameters, improving peak TOPS and improving TOPS/W can require different changes to the same architecture, and a system can be compute-bound for one while memory-bound for the other.
- Sparsity, especially unstructured sparsity, can lower effective arithmetic intensity and push a workload away from the roofline even when it reduces total operations; the paper concludes that roofline position alone is not enough to judge sparse hardware, and end-to-end energy and latency must be considered.
- Near- and in-memory computing raise the compute roofline and reduce data-movement energy, but they couple storage and compute in ways that cause spatial and temporal utilization losses, so the full benefit depends on adding a second memory level and carefully mapping workloads.
- Larger parallel compute arrays raise the roofline but shift the knee to higher arithmetic intensity, making the system more memory-bound unless data reuse keeps the workload's intensity high.
Reading between the lines
- A natural extension the paper does not formalize is a utilization-aware design rule: choose compute parallelism and memory bandwidth so that both knees sit near the arithmetic intensity of the target workload mix; the paper's area-allocation discussion points that way but stops short of a closed-form rule.
- The framework suggests a reporting standard for published accelerators: give the throughput and energy rooflines together with the achieved utilization at the workload point, since identical rooflines can hide large efficiency differences; this is measurable and would make survey comparisons fairer.
- For sparse workloads, defining a sparsity-aware arithmetic intensity that counts only non-zero operations and effective bytes with indices might restore the roofline's predictive power; the paper notes the gap but does not propose such a metric.
- Treating the inter-chip network in multi-chip LLM systems as an additional memory level with its own bandwidth and energy per byte would extend the framework to scale-out architectures, which the paper mentions only as future outlook.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript proposes an enhanced roofline modeling framework for ML accelerators, comprising a throughput roofline (Eq. 4) and an energy-efficiency roofline (Eq. 5) that account for per-level arithmetic intensities across a multi-level memory hierarchy. After deriving the two roofline equations from an additive energy model (Eq. 1) and a max-latency model (Eq. 2), the paper surveys five hardware techniques—parallelism, spatial/temporal data reuse, quantization, sparsity, and near/in-memory computing—and illustrates how each affects the roofline position and the workload operating point. It closes with trade-offs between parallelism and utilization, programmability and specialization, and an outlook on emerging technologies.
Significance. The paper's value lies in its organized synthesis and explicit two-roofline formulation. The equations are correct under the stated additive-energy and max-latency assumptions, and the paper is unusually candid about the model's limits, particularly in Section III-D and Section IV-B, where it acknowledges that sparse/irregular workloads fall below the roofline due to utilization losses and that roofline analysis alone is not sufficient for end-to-end speedup or energy decisions. No data fitting or fabricated results are involved; the illustrative parameters are clearly labeled as assumptions. These strengths make the paper a useful reference for practitioners and students, though it does not present new experimental data. The central claim is appropriately scoped as a framework for understanding rather than for prediction, so the absence of an integrated utilization factor is a limitation but not a fatal one.
minor comments (5)
- [Section III-F1, Eq. (6)] Equation (6) is self-referential as written: D_y appears on both sides, and the expression is dimensionally inconsistent with the surrounding prose. The text in Section III-F2 correctly states that the dynamic range in bits is the sum of input bits, weight bits, and log2(P_R). Please correct Eq. (6) to a bits-based form such as B_y = B_x + B_w + log2(P_R), or provide the equivalent linear-scale expression.
- [Abstract and title] The abstract and title contain the typo "rooline" instead of "roofline"; this should be corrected throughout.
- [Section I and Section V] The introduction states "nearly 100 × performance improvement every 24 months, maintained over the past 8 years" (which would imply an enormous cumulative factor), while the conclusion states "roughly 1000 × increase in throughput and energy efficiency." These two statements are inconsistent; please reconcile them or clarify the time scales being referenced.
- [Figure 3 caption] The caption's expression "AI=AIL3/16=AIL2=AIL1 ∗ 16" is ambiguous and potentially misleading. Please define the assumed relationships among the per-level arithmetic intensities explicitly, for instance by stating that each successive level differs by a factor of 16.
- [Section III-A] There is a typo in the first paragraph: "paralelization" should be "parallelization." Similar minor typographical errors appear elsewhere, such as "rooline" in the abstract and "the Samsung's" in Section III-A; a careful proofreading pass is recommended.
Circularity Check
No circularity: the two-roofline equations are algebraic restatements of the paper's own latency and energy definitions, and the central framing rests on external roofline literature.
full rationale
The paper's central equations are derived directly from its own stated definitions. Equation (4) is obtained by substituting the arithmetic-intensity definition AILi = Nop/NLi into the latency expression Ltask = (1/fclk)*max(NLn/BLn,...,NL1/BL1,Nop/Aop) and writing PTP = Nop/Ltask; Equation (5) is obtained by substituting the same definition into the energy expression Etask = Nop*Eop + sum(NLi*ELi) and writing PE = Nop/Etask. These are algebraic identities given the definitions, not empirical predictions, and no parameter is fitted to data and then renamed as a prediction. The paper makes no claim to predict measured accelerator performance from first principles; it presents an organizing framework whose validity rests on the external roofline model of Williams et al. [3] and the energy roofline of Choi et al. [9]. The self-citations that are present, such as [55] for the in-memory-compute dynamic-range trade-off, are explicitly attributed as prior work ('Simplifying the more detailed analysis in [55]') and are used as surveyed examples rather than as load-bearing justification for the paper's roofline framework. The paper even concedes in Section III-D that 'Roofline analysis alone, albeit useful, is not sufficient' for sparse workloads, which is a limitation of scope rather than a circular step. No load-bearing step in the derivation chain reduces by construction to its own inputs, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (5)
- standard math The classic roofline model of Williams et al. [3] and its energy variant by Choi et al. [9] are valid starting points.
- domain assumption Energy and latency of ML accelerators are dominated by MAC operations and data movement, each with constant per-byte/per-op costs (Eq. 1).
- domain assumption Arithmetic intensity varies monotonically across memory levels (AIL3 > AIL2 > AIL1) for typical ML workloads.
- domain assumption Compute and memory transfers overlap such that latency is the max of the individual times (Eq. 4).
- domain assumption The dynamic-range versus parallelism trade-off in IMC is as described in [55] by Verma et al.
Cite this review
Pith. "Pith review of How to keep pushing ML accelerator performance? Know your rooflines!." pith.science (2026). https://pith.science/paper/D72YGMIR
@misc{pith2026250516346,
author = {Pith},
title = {Pith review of: How to keep pushing ML accelerator performance? Know your rooflines!},
year = {2026},
howpublished = {\url{https://pith.science/paper/D72YGMIR}},
note = {Machine review of arXiv:2505.16346}
}
read the original abstract
The rapidly growing importance of Machine Learning (ML) applications, coupled with their ever-increasing model size and inference energy footprint, has created a strong need for specialized ML hardware architectures. Numerous ML accelerators have been explored and implemented, primarily to increase task-level throughput per unit area and reduce task-level energy consumption. This paper surveys key trends toward these objectives for more efficient ML accelerators and provides a unifying framework to understand how compute and memory technologies/architectures interact to enhance system-level efficiency and performance. To achieve this, the paper introduces an enhanced version of the roofline model and applies it to ML accelerators as an effective tool for understanding where various execution regimes fall within roofline bounds and how to maximize performance and efficiency under the rooline. Key concepts are illustrated with examples from state-of-the-art designs, with a view towards open research opportunities to further advance accelerator performance.
Figures
Figures from the paper (15 more)
Reference graph
Works this paper leans on
-
[55]
In-memory computing: Advances and prospects,
N. Verma, H. Jia, H. Valavi, Y . Tang, M. Ozatay, L.-Y . Chen, B. Zhang, and P. Deaville, “In-memory computing: Advances and prospects,”IEEE Solid-State Circuits Magazine , vol. 11, no. 3, pp. 43–55, 2019
2019
-
[1]
Visualizing size of large language models,
G. Anil, “Visualizing size of large language models,” 2023, accessed: 2024-10-10. [Online]. Available: https://medium.com/@georgeanil/ visualizing-size-of-large-language-models-ec576caa5557
2023
-
[2]
Trends in deep learning hardware,
B. Dally, “Trends in deep learning hardware,” 2024, talk, presented by NVIDIA. PREPRINT OF ARTICLE PUBLISHED IN JOURNAL OF SOLID STATE CIRCUITS 17
work page 2024
-
[3]
Roofline: an insightful visual performance model for multicore architectures,
S. Williams, A. Waterman, and D. Patterson, “Roofline: an insightful visual performance model for multicore architectures,” Communications of the ACM , vol. 52, no. 4, pp. 65–76, 2009
work page 2009
-
[4]
Eyeriss: A apatial architecture for energy-efficient dataflow for convolutional neural networks,
Y .-H. Chen, J. Emer, and V . Sze, “Eyeriss: A apatial architecture for energy-efficient dataflow for convolutional neural networks,” in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), 2016, pp. 367–379
work page 2016
-
[5]
Roofline performance analysis of dnn architectures on cpu and gpu systems,
H. Prashanth and M. Rao, “Roofline performance analysis of dnn architectures on cpu and gpu systems,” in 2024 25th International Symposium on Quality Electronic Design (ISQED) . IEEE, 2024, pp. 1–8
work page 2024
-
[6]
S. W. Williams, Book: The roofline model . University of California, 2010
work page 2010
-
[8]
Understanding reuse, performance, and hardware cost of dnn dataflow: A data-centric approach,
H. Kwon, P. Chatarasi, M. Pellauer, A. Parashar, V . Sarkar, and T. Kr- ishna, “Understanding reuse, performance, and hardware cost of dnn dataflow: A data-centric approach,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , 2019, pp. 754–768
work page 2019
Show all 81 references
-
[9]
A roofline model of energy,
J. W. Choi, D. Bedard, R. Fowler, and R. Vuduc, “A roofline model of energy,” in 2013 IEEE 27th International Symposium on Parallel and Distributed Processing. IEEE, 2013, pp. 661–672
2013
-
[10]
Symphony: Orchestrating sparse and dense tensors with hierarchical heterogeneous processing,
M. Pellauer, J. Clemons, V . Balaji, N. Crago, A. Jaleel, D. Lee, M. O’Connor, A. Parashar, S. Treichler, P.-A. Tsai et al. , “Symphony: Orchestrating sparse and dense tensors with hierarchical heterogeneous processing,” ACM Transactions on Computer Systems , vol. 41, no. 1-4,...
2023
-
[11]
Lots of questions on Google’s “Trillium
T. P. Morgan, “Lots of questions on Google’s “Trillium” TPU v6, a few answers,” Oct 2024. [Online]. Available: https://www.nextplatform.com/2024/06/10/ lots-of-questions-on-googles-trillium-tpu-v6-a-few-answers/
2024
-
[12]
Envi- sion: A 0.26-to-10tops/w subword-parallel dynamic-voltage-accuracy- frequency-scalable convolutional neural network processor in 28nm fdsoi,
B. Moons, R. Uytterhoeven, W. Dehaene, and M. Verhelst, “Envi- sion: A 0.26-to-10tops/w subword-parallel dynamic-voltage-accuracy- frequency-scalable convolutional neural network processor in 28nm fdsoi,” in 2017 IEEE International Solid-State Circuits Conference (ISSCC). IEEE...
2017
-
[13]
9.5 a 6k-mac feature-map-sparsity-aware neural processing unit in 5nm flagship mobile soc,
J.-S. Park, J.-W. Jang, H. Lee, D. Lee, S. Lee, H. Jung, S. Lee, S. Kwon, K. Jeong, J.-H. Song et al. , “9.5 a 6k-mac feature-map-sparsity-aware neural processing unit in 5nm flagship mobile soc,” in 2021 IEEE International Solid-State Circuits Conference (ISSCC) , vol. 64. IE...
2021
-
[14]
Compute solution for tesla’s full self-driving computer,
E. Talpes, D. D. Sarma, G. Venkataramanan, P. Bannon, B. McGee, B. Floering, A. Jalote, C. Hsiong, S. Arora, A. Gorti et al., “Compute solution for tesla’s full self-driving computer,” IEEE Micro , vol. 40, no. 2, pp. 25–35, 2020
2020
-
[15]
7.2 a 12nm programmable convolution-efficient neural- processing-unit chip achieving 825tops,
Y . Jiao, L. Han, R. Jin, Y .-J. Su, C. Ho, L. Yin, Y . Li, L. Chen, Z. Chen, L. Liu et al. , “7.2 a 12nm programmable convolution-efficient neural- processing-unit chip achieving 825tops,” in 2020 IEEE International Solid-State Circuits Conference-(ISSCC) . IEEE, 2020, pp. 136–140
2020
-
[16]
Groq rocks neural networks,
L. Gwennap, “Groq rocks neural networks,” Microprocessor Report, Tech. Rep., jan, 2020
2020
-
[17]
9.1 a 7nm 4-core ai chip with 25.6tflops hybrid fp8 training, 102.4tops int4 inference and workload-aware throttling,
A. Agrawal, S. K. Lee, J. Silberman, M. Ziegler, M. Kang, S. Venkatara- mani, N. Cao, B. Fleischer, M. Guillorn, M. Cohen, S. Mueller, J. Oh, M. Lutz, J. Jung, S. Koswatta, C. Zhou, V . Zalani, J. Bonanno, R. Casat- uta, C.-Y . Chen, J. Choi, H. Haynie, A. Herbert, R. Jain, M....
2021
-
[18]
16.7 a 40-310tops/w sram-based all-digital up to 4b in-memory computing multi-tiled nn accelerator in fd-soi 18nm for deep-learning edge applications,
G. Desoli, N. Chawla, T. Boesch, M. Avodhyawasi, H. Rawat, H. Chawla, V . Abhijith, P. Zambotti, A. Sharma, C. Cappetta, M. Rossi, A. De Vita, and F. Girardi, “16.7 a 40-310tops/w sram-based all-digital up to 4b in-memory computing multi-tiled nn accelerator in fd-soi 18nm for...
2023
-
[19]
Charm: Composing heterogeneous accelerators for matrix multiply on versal acap architecture,
J. Zhuang, J. Lau, H. Ye, Z. Yang, Y . Du, J. Lo, K. Denolf, S. Neuendorffer, A. Jones, J. Hu, D. Chen, J. Cong, and P. Zhou, “Charm: Composing heterogeneous accelerators for matrix multiply on versal acap architecture,” in Proceedings of the 2023 ACM/SIGDA International Sympo...
2023
-
[20]
Davinci: A scalable architecture for neural network computing,
H. Liao, J. Tu, J. Xia, and X. Zhou, “Davinci: A scalable architecture for neural network computing,” in 2019 IEEE Hot Chips 31 Symposium (HCS). IEEE Computer Society, 2019, pp. 1–44
2019
-
[21]
Nvidia tensor core programmability, performance & precision,
S. Markidis, S. W. Der Chien, E. Laure, I. B. Peng, and J. S. Vetter, “Nvidia tensor core programmability, performance & precision,” in 2018 IEEE international parallel and distributed processing symposium workshops (IPDPSW). IEEE, 2018, pp. 522–531
2018
-
[22]
A charge domain sram compute-in-memory macro with c-2c ladder- based 8-bit mac unit in 22-nm finfet process for edge inference,
H. Wang, R. Liu, R. Dorrance, D. Dasalukunte, D. Lake, and B. Carlton, “A charge domain sram compute-in-memory macro with c-2c ladder- based 8-bit mac unit in 22-nm finfet process for edge inference,” IEEE Journal of Solid-State Circuits , vol. 58, no. 4, pp. 1037–1050, 2023
2023
-
[23]
A 22 nm, 1540 top/s/w, 12.1 top/s/mm 2 in-memory analog matrix-vector-multiplier for dnn acceleration,
I. A. Papistas, S. Cosemans, B. Rooseleer, J. Doevenspeck, M.-H. Na, A. Mallik, P. Debacker, and D. Verkest, “A 22 nm, 1540 top/s/w, 12.1 top/s/mm 2 in-memory analog matrix-vector-multiplier for dnn acceleration,” in 2021 IEEE Custom Integrated Circuits Conference (CICC). IEEE...
2021
-
[24]
A 64-tile 2.4- mb in-memory-computing cnn accelerator employing charge-domain compute,
H. Valavi, P. J. Ramadge, E. Nestler, and N. Verma, “A 64-tile 2.4- mb in-memory-computing cnn accelerator employing charge-domain compute,” IEEE Journal of Solid-State Circuits, vol. 54, no. 6, pp. 1789– 1799, 2019
2019
-
[25]
Compute Solution for Tesla’s Full Self-Driving Computer,
E. Talpes, D. D. Sarma, G. Venkataramanan, P. Bannon, B. McGee, B. Floering, A. Jalote, C. Hsiong, S. Arora, A. Gorti, and G. S. Sachdev, “Compute Solution for Tesla’s Full Self-Driving Computer,” IEEE Micro, vol. 40, no. 2, pp. 25–35, 2020
2020
-
[26]
Hardware for deep learning,
B. Dally, “Hardware for deep learning,” in IEEE Hot Chips Symposium (HCS), vol. 35. IEEE, 2023, pp. 1–58
2023
-
[27]
Lincoln ai computing survey (laics) update,
A. Reuther, P. Michaleas, M. Jones, V . Gadepally, S. Samsi, and J. Kepner, “Lincoln ai computing survey (laics) update,” in 2023 IEEE High Performance Extreme Computing Conference (HPEC) . IEEE, 2023, pp. 1–7
2023
-
[28]
Neural network accelerator comparison
K. Guo, W. Li, K. Zhong, Z. Zhu, S. Zeng, T. Xie, S. Han, Y . Xie, P. Debacker, M. Verhelst, and Y . Wang, “Neural network accelerator comparison.” [Online]. Available: https://nicsefc.ee.tsinghua. edu.cn/project.html
-
[29]
Llm inference unveiled: Survey and roofline model insights,
Z. Yuan, Y . Shang, Y . Zhou, Z. Dong, Z. Zhou, C. Xue, B. Wu, Z. Li, Q. Gu, Y . J. Lee et al. , “Llm inference unveiled: Survey and roofline model insights,” arXiv preprint arXiv:2402.16363 , 2024
2024 arXiv
-
[30]
Minifloats on risc-v cores: Isa extensions with mixed- precision short dot products,
L. Bertaccini, G. Paulin, M. Cavalcante, T. Fischer, S. Mach, and L. Benini, “Minifloats on risc-v cores: Isa extensions with mixed- precision short dot products,” IEEE Transactions on Emerging Topics in Computing, 2024
2024
-
[31]
Cutie: Beyond petaop/s/w ternary dnn inference acceleration with better-than-binary energy efficiency,
M. Scherer, G. Rutishauser, L. Cavigelli, and L. Benini, “Cutie: Beyond petaop/s/w ternary dnn inference acceleration with better-than-binary energy efficiency,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 41, no. 4, pp. 1020–1033, 2021
2021
-
[32]
Binareye: An always-on energy-accuracy-scalable binary cnn processor with all memory on chip in 28nm cmos,
B. Moons, D. Bankman, L. Yang, B. Murmann, and M. Verhelst, “Binareye: An always-on energy-accuracy-scalable binary cnn processor with all memory on chip in 28nm cmos,” in 2018 IEEE Custom Integrated Circuits Conference (CICC) . IEEE, 2018, pp. 1–4
2018
-
[33]
Bitnet: Scaling 1-bit transformers for large language models,
H. Wang, S. Ma, L. Dong, S. Huang, H. Wang, L. Ma, F. Yang, R. Wang, Y . Wu, and F. Wei, “Bitnet: Scaling 1-bit transformers for large language models,” arXiv preprint arXiv:2310.11453 , 2023
2023 arXiv
-
[34]
A 3 tops/w risc-v parallel cluster for inference of fine-grain mixed-precision quantized neural networks,
A. Nadalini, G. Rutishauser, A. Burrello, N. Bruschi, A. Garofalo, L. Benini, F. Conti, and D. Rossi, “A 3 tops/w risc-v parallel cluster for inference of fine-grain mixed-precision quantized neural networks,” in 2023 IEEE Computer Society Annual Symposium on VLSI (ISVLSI) . I...
2023
-
[35]
Marsellus: A heterogeneous risc-v ai-iot end-node soc with 2–8 b dnn acceleration and 30%-boost adaptive body biasing,
F. Conti, G. Paulin, A. Garofalo, D. Rossi, A. Di Mauro, G. Rutishauser, G. Ottavi, M. Eggiman, H. Okuhara, and L. Benini, “Marsellus: A heterogeneous risc-v ai-iot end-node soc with 2–8 b dnn acceleration and 30%-boost adaptive body biasing,” IEEE Journal of Solid-State Circu...
2023
-
[36]
Microscaling data formats for deep learning,
B. D. Rouhani, R. Zhao, A. More, M. Hall, A. Khodamoradi, S. Deng, D. Choudhary, M. Cornea, E. Dellinger, K. Denolf et al., “Microscaling data formats for deep learning,” arXiv preprint arXiv:2310.10537, 2023
2023 arXiv
-
[37]
Nvidia blackwell platform: Advancing generative ai and accelerated computing,
A. Tirumala and R. Wong, “Nvidia blackwell platform: Advancing generative ai and accelerated computing,” in 2024 IEEE Hot Chips 36 Symposium (HCS), 2024, pp. 1–33
2024
-
[38]
Siracusa: A 16 nm heterogenous risc-v soc for extended reality with at-mram neural engine,
A. S. Prasad, M. Scherer, F. Conti, D. Rossi, A. Di Mauro, M. Eggimann, J. T. G ´omez, Z. Li, S. S. Sarwar, Z. Wang et al. , “Siracusa: A 16 nm heterogenous risc-v soc for extended reality with at-mram neural engine,” IEEE Journal of Solid-State Circuits , 2024
2024
-
[39]
Onyx: A 12nm 756 gops/w coarse-grained reconfigurable array for accelerating dense and sparse applications,
K. Koul, M. Strange, J. Melchert, A. Carsello, Y . Mei, O. Hsu, T. Kong, P.-H. Chen, H. Ke, K. Zhang et al. , “Onyx: A 12nm 756 gops/w coarse-grained reconfigurable array for accelerating dense and sparse applications,” in 2024 IEEE Symposium on VLSI Technology and Circuits (V...
2024
-
[40]
Learning n: m fine-grained structured sparse neural networks from scratch,
A. Zhou, Y . Ma, J. Zhu, J. Liu, Z. Zhang, K. Yuan, W. Sun, and H. Li, “Learning n: m fine-grained structured sparse neural networks from scratch,” arXiv preprint arXiv:2102.04010 , 2021
2021 arXiv
-
[41]
3.2 the a100 datacenter gpu and ampere architecture,
J. Choquette, E. Lee, R. Krashinsky, V . Balan, and B. Khailany, “3.2 the a100 datacenter gpu and ampere architecture,” in 2021 IEEE International Solid-State Circuits Conference (ISSCC) , vol. 64. IEEE, 2021, pp. 48–50
2021
-
[42]
Venom: A vectorized n: M format for unleashing the power of sparse tensor cores,
R. L. Castro, A. Ivanov, D. Andrade, T. Ben-Nun, B. B. Fraguela, and T. Hoefler, “Venom: A vectorized n: M format for unleashing the power of sparse tensor cores,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis ,...
2023
-
[43]
Occamy: A 432-core dual-chiplet dual-hbm2e 768-dp-gflop/s risc-v system for 8- to-64-bit dense and sparse computing in 12-nm finfet,
P. Scheffler, T. Benz, V . Potocnik, T. Fischer, L. Colagrande, N. Wistoff, Y . Zhang, L. Bertaccini, G. Ottavi, M. Eggimann, M. Cavalcante, G. Paulin, F. K. G ¨urkaynak, D. Rossi, and L. Benini, “Occamy: A 432-core dual-chiplet dual-hbm2e 768-dp-gflop/s risc-v system for 8- t...
2025
-
[44]
Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing,
G. Heo, S. Lee, J. Cho, H. Choi, S. Lee, H. Ham, G. Kim, D. Mahajan, and J. Park, “Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating...
2024
-
[45]
Inclusive-pim: Hardware-software co-design for broad acceleration on commercial pim architectures,
J. Alsop, S. Aga, M. Ibrahim, M. Islam, A. Mccrabb, and N. Jayasena, “Inclusive-pim: Hardware-software co-design for broad acceleration on commercial pim architectures,” 2024. [Online]. Available: https://arxiv.org/abs/2309.07984
2024 arXiv
-
[46]
In-memory computation of a machine-learning classifier in a standard 6t sram array,
J. Zhang, Z. Wang, and N. Verma, “In-memory computation of a machine-learning classifier in a standard 6t sram array,” IEEE Journal of Solid-State Circuits , vol. 52, no. 4, pp. 915–924, 2017
2017
-
[47]
An energy-efficient memory-based high-throughput vlsi architecture for convolutional networks,
M. Kang, S. K. Gonugondla, M.-S. Keel, and N. R. Shanbhag, “An energy-efficient memory-based high-throughput vlsi architecture for convolutional networks,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 1037– 1041
2015
-
[48]
Fast, energy-efficient, robust, and reproducible mixed-signal neuromorphic classifier based on embedded nor flash memory technology,
X. Guo, F. M. Bayat, M. Bavandpour, M. Klachko, M. R. Mahmoodi, M. Prezioso, K. K. Likharev, and D. B. Strukov, “Fast, energy-efficient, robust, and reproducible mixed-signal neuromorphic classifier based on embedded nor flash memory technology,” in 2017 IEEE International Ele...
2017
-
[49]
Analog in-memory subthreshold deep neural network accelerator,
L. Fick, D. Blaauw, D. Sylvester, S. Skrzyniarz, M. Parikh, and D. Fick, “Analog in-memory subthreshold deep neural network accelerator,” in 2017 IEEE Custom Integrated Circuits Conference (CICC) , 2017, pp. 1–4
2017
-
[50]
A 5-nm 254-tops/w 221-tops/mm2 fully-digital computing-in-memory macro supporting wide-range dynamic-voltage- frequency scaling and simultaneous mac and write operations,
H. Fujiwara, H. Mori, W.-C. Zhao, M.-C. Chuang, R. Naous, C.-K. Chuang, T. Hashizume, D. Sun, C.-F. Lee, K. Akarvardar, S. Adham, T.- L. Chou, M. E. Sinangil, Y . Wang, Y .-D. Chih, Y .-H. Chen, H.-J. Liao, and T.-Y . J. Chang, “A 5-nm 254-tops/w 221-tops/mm2 fully-digital com...
2022
-
[51]
16.4 an 89tops/w and 16.3tops/mm2 all-digital sram-based full-precision compute-in memory macro in 22nm for machine-learning edge applications,
Y .-D. Chih, P.-H. Lee, H. Fujiwara, Y .-C. Shih, C.-F. Lee, R. Naous, Y .-L. Chen, C.-P. Lo, C.-H. Lu, H. Mori, W.-C. Zhao, D. Sun, M. E. Sinangil, Y .-H. Chen, T.-L. Chou, K. Akarvardar, H.-J. Liao, Y . Wang, M.-F. Chang, and T.-Y . J. Chang, “16.4 an 89tops/w and 16.3tops/m...
2021
-
[52]
A maximally row- parallel mram in-memory-computing macro addressing readout circuit sensitivity and area,
P. Deaville, B. Zhang, L.-Y . Chen, and N. Verma, “A maximally row- parallel mram in-memory-computing macro addressing readout circuit sensitivity and area,” in ESSCIRC 2021 - IEEE 47th European Solid State Circuits Conference (ESSCIRC) , 2021, pp. 75–78
2021
-
[53]
A programmable heterogeneous microprocessor based on bit-scalable in-memory comput- ing,
H. Jia, H. Valavi, Y . Tang, J. Zhang, and N. Verma, “A programmable heterogeneous microprocessor based on bit-scalable in-memory comput- ing,” IEEE Journal of Solid-State Circuits, vol. 55, no. 9, pp. 2609–2621, 2020
2020
-
[54]
A crossbar array of magnetoresistive memory devices for in-memory computing,
S. Jung, H. Lee, S. Myung, H. Kim, S. K. Yoon, S.-W. Kwon, Y . Ju, M. Kim, W. Yi, S. Han, B. Kwon, B. Seo, K. Lee, G.-H. Koh, K. Lee, Y . Song, C. Choi, D. Ham, and S. J. Kim, “A crossbar array of magnetoresistive memory devices for in-memory computing,” Nature, vol. 601, no. ...
2022 doi
-
[56]
14.2 a compute sram with bit-serial integer/floating-point operations for programmable in-memory vector acceleration,
J. Wang, X. Wang, C. Eckert, A. Subramaniyan, R. Das, D. Blaauw, and D. Sylvester, “14.2 a compute sram with bit-serial integer/floating-point operations for programmable in-memory vector acceleration,” in 2019 IEEE International Solid-State Circuits Conference - (ISSCC), 2019...
2019
-
[57]
A 40nm 64kb 26.56tops/w 2.37mb/mm2rram binary/compute-in-memory macro with 4.23x im- provement in density and > 75% use of sensing dynamic range,
S. D. Spetalnick, M. Chang, B. Crafton, W.-S. Khwa, Y .-D. Chih, M.-F. Chang, and A. Raychowdhury, “A 40nm 64kb 26.56tops/w 2.37mb/mm2rram binary/compute-in-memory macro with 4.23x im- provement in density and > 75% use of sensing dynamic range,” in2022 IEEE International Soli...
2022
-
[58]
Funda- mental limits on energy-delay-accuracy of in-memory architectures in inference applications,
S. K. Gonugondla, C. Sakr, H. Dbouk, and N. R. Shanbhag, “Funda- mental limits on energy-delay-accuracy of in-memory architectures in inference applications,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 41, no. 10, pp. 3188–3201, 2022
2022
-
[59]
11.3 metis aipu: A 12nm 15tops/w 209.6tops soc for cost- and energy-efficient inference at the edge,
P. A. Hager, B. Moons, S. Cosemans, I. A. Papistas, B. Rooseleer, J. V . Loon, R. Uytterhoeven, F. Zaruba, S. Koumousi, M. Stanisavljevic, S. Mach, S. Mutsaards, R. K. Aljameh, G. H. Khov, B. Machiels, C. Olar, A. Psarras, S. Geursen, J. Vermeeren, Y . Lu, A. Maringanti, D. Am...
2024
-
[60]
Benchmarking in-memory computing architectures,
N. R. Shanbhag and S. K. Roy, “Benchmarking in-memory computing architectures,” IEEE Open Journal of the Solid-State Circuits Society , vol. 2, pp. 288–300, 2022
2022
-
[61]
A 22nm 128-kb mram row/column-parallel in-memory computing macro with memory- resistance boosting and multi-column adc readout,
P. Deaville, B. Zhang, and N. Verma, “A 22nm 128-kb mram row/column-parallel in-memory computing macro with memory- resistance boosting and multi-column adc readout,” in 2022 IEEE Symposium on VLSI Technology and Circuits (VLSI Technology and Circuits), 2022, pp. 268–269
2022
-
[62]
A 64-core mixed-signal in-memory compute chip based on phase-change memory for deep neural network inference,
M. Le Gallo, R. Khaddam-Aljameh, M. Stanisavljevic, A. Vasilopoulos, B. Kersting, M. Dazzi, G. Karunaratne, M. Br ¨andli, A. Singh, S. M. M ¨uller, J. B ¨uchel, X. Timoneda, V . Joshi, M. J. Rasch, U. Egger, A. Garofalo, A. Petropoulos, T. Antonakopoulos, K. Brew, S. Choi, I. ...
2023
-
[63]
An n40 256k×44 embedded rram macro with sl-precharge sa and low-voltage current limiter to improve read and write performance,
C.-C. Chou, Z.-J. Lin, P.-L. Tseng, C.-F. Li, C.-Y . Chang, W.-C. Chen, Y .-D. Chih, and T.-Y . J. Chang, “An n40 256k×44 embedded rram macro with sl-precharge sa and low-voltage current limiter to improve read and write performance,” in 2018 IEEE International Solid-State Cir...
2018
-
[64]
Cmos- embedded stt-mram arrays in 2x nm nodes for gp-mcu applications,
D. Shum, D. Houssameddine, S. T. Woo, Y . S. You, J. Wong, K. W. Wong, C. C. Wang, K. H. Lee, K. Yamane, V . B. Naik, C. S. Seet, T. Tahmasebi, C. Hai, H. W. Yang, N. Thiyagarajah, R. Chao, J. W. Ting, N. L. Chung, T. Ling, T. H. Chan, S. Y . Siah, R. Nair, S. Deshpande, R. Wh...
2017
-
[65]
A switched-capacitor sram in-memory computing macro with high-precision, high-efficiency differential archi- tecture,
J. Lee, B. Zhang, and N. Verma, “A switched-capacitor sram in-memory computing macro with high-precision, high-efficiency differential archi- tecture,” in 2024 European Conference on Solid-State Circuits , 2024
2024
-
[66]
Scalable and Programmable Neural Network Inference Accelerator Based on In-Memory Computing,
H. Jia, M. Ozatay, Y . Tang, H. Valavi, R. Pathak, J. Lee, and N. Verma, “Scalable and Programmable Neural Network Inference Accelerator Based on In-Memory Computing,” IEEE Journal of Solid-State Circuits, vol. 57, no. 1, pp. 198–211, 2022
2022
-
[67]
Interstellar: Using Halide’s Scheduling Language to Analyze DNN Accelerators,
X. Yang, M. Gao, Q. Liu, J. Setter, J. Pu, A. Nayak, S. Bell, K. Cao, H. Ha, P. Raina, C. Kozyrakis, and M. Horowitz, “Interstellar: Using Halide’s Scheduling Language to Analyze DNN Accelerators,” in Proceedings of the Twenty-Fifth International Conference on Architectural Su...
2020
-
[68]
MAESTRO: A Data-Centric Approach to Understand Reuse, Performance, and Hardware Cost of DNN Mappings,
H. Kwon, P. Chatarasi, V . Sarkar, T. Krishna, M. Pellauer, and A. Parashar, “MAESTRO: A Data-Centric Approach to Understand Reuse, Performance, and Hardware Cost of DNN Mappings,” IEEE Micro, vol. 40, no. 3, pp. 20–29, 2020
2020
-
[69]
Timeloop: A Systematic Approach to DNN Accelerator Evaluation,
A. Parashar, P. Raina, Y . S. Shao, Y .-H. Chen, V . A. Ying, A. Mukkara, R. Venkatesan, B. Khailany, S. W. Keckler, and J. Emer, “Timeloop: A Systematic Approach to DNN Accelerator Evaluation,” in 2019 IEEE International Symposium on Performance Analysis of Systems and Softwa...
2019
-
[70]
ZigZag: Enlarging Joint Architecture-Mapping Design Space Exploration for DNN Accelerators,
L. Mei, P. Houshmand, V . Jain, S. Giraldo, and M. Verhelst, “ZigZag: Enlarging Joint Architecture-Mapping Design Space Exploration for DNN Accelerators,” IEEE Transactions on Computers , vol. 70, no. 8, pp. 1160–1174, 2021
2021
-
[71]
CoSA: Scheduling by constrained op- timization for spatial accelerators,
Q. Huang, A. Kalaiah, M. Kang, J. Demmel, G. Dinh, J. Wawrzynek, T. Norell, and Y . S. Shao, “CoSA: Scheduling by constrained op- timization for spatial accelerators,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) , 2021, pp. 554–566
2021
-
[72]
Mind Mappings: Enabling Efficient Algorithm-Accelerator Mapping Space Search,
K. Hegde, P.-A. Tsai, S. Huang, V . Chandra, A. Parashar, and C. W. Fletcher, “Mind Mappings: Enabling Efficient Algorithm-Accelerator Mapping Space Search,” in Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operatin...
2021
-
[73]
GAMMA: Automating the HW Mapping of DNN Models on Accelerators via Genetic Algorithm,
S.-C. Kao and T. Krishna, “GAMMA: Automating the HW Mapping of DNN Models on Accelerators via Genetic Algorithm,” in Proceedings of the 39th International Conference on Computer-Aided Design , ser. ICCAD ’20. New York, NY , USA: Association for Computing Machinery, 2020. [Onli...
2020 doi
-
[74]
Stream: Design space exploration of layer-fused dnns on hetero- geneous dataflow accelerators,
A. Symons, L. Mei, S. Colleman, P. Houshmand, S. Karl, and M. Ver- helst, “Stream: Design space exploration of layer-fused dnns on hetero- geneous dataflow accelerators,” IEEE Transactions on Computers, 2024
2024
-
[75]
The groq software-defined scale-out tensor streaming multiprocessor : From chips-to-systems architectural overview,
D. Abts, J. Kim, G. Kimmell, M. Boyd, K. Kang, S. Parmar, A. Ling, A. Bitar, I. Ahmed, and J. Ross, “The groq software-defined scale-out tensor streaming multiprocessor : From chips-to-systems architectural overview,” in 2022 IEEE Hot Chips 34 Symposium (HCS) , 2022, pp. 1–69
2022
-
[76]
Application specific instruction processor based implementation of a gnss receiver on an fpga,
K. G ¨otz and T. Noll, “Application specific instruction processor based implementation of a gnss receiver on an fpga,” in Design & Test in Europe Conference, 2006, pp. 58–63
2006
-
[77]
How flexible is your com- puting system?
S. Huang, L. Waeijen, and H. Corporaal, “How flexible is your com- puting system?” ACM Transactions on Embedded Computing Systems (TECS), vol. 21, no. 4, pp. 1–41, 2022
2022
-
[78]
Tandem processor: Grappling with emerging operators in neural networks,
S. Ghodrati, S. Kinzer, H. Xu, R. Mahapatra, Y . Kim, B. H. Ahn, D. K. Wang, L. Karthikeyan, A. Yazdanbakhsh, J. Park et al., “Tandem processor: Grappling with emerging operators in neural networks,” in Proceedings of the 29th ACM International Conference on Architectural Supp...
2024
-
[79]
Mec: memory-efficient convolution for deep neural network,
M. M. Cho and D. Brand, “Mec: memory-efficient convolution for deep neural network,” in ICML-34: Proceedings of the International Conference on Machine Learning - Volume 70 . JMLR.org, 2017, p. 815–824
2017
-
[80]
A formalism of dnn accelerator flexibility,
S.-C. Kao, H. Kwon, M. Pellauer, A. Parashar, and T. Krishna, “A formalism of dnn accelerator flexibility,” Proceedings of the ACM on Measurement and Analysis of Computing Systems , vol. 6, no. 2, pp. 1–23, 2022
2022
-
[81]
Mlir: Scaling compiler infrastructure for domain specific computation,
C. Lattner, M. Amini, U. Bondhugula, A. Cohen, A. Davis, J. Pien- aar, R. Riddle, T. Shpeisman, N. Vasilache, and O. Zinenko, “Mlir: Scaling compiler infrastructure for domain specific computation,” in 2021 IEEE/ACM International Symposium on Code Generation and Optimization (...
2021
-
[82]
The hardware lottery,
S. Hooker, “The hardware lottery,” Communications of the ACM, vol. 64, no. 12, pp. 58–65, 2021. Marian Verhelst Marian Verhelst is a professor at the MICAS labs of KU Leuven and a research director at imec. Her research focuses on embedded machine learning, hardware accelerato...
2021
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.