REVIEW 3 major objections 4 minor 92 references
DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline Model
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A near-core decompression accelerator claims up to 4x faster compressed matrix multiplies and 1.6–2.6x faster LLM token generation on CPUs.
desk verdict A solid simulation-based architecture paper with a genuinely useful model, held back by an uncalibrated simulator and a uniform-sparsity assumption that may inflate the speedups. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Roof-Surface model: a three-dimensional performance bound whose equation is $\mathit{TPS} = \min\{\mathit{MBW}\cdot \mathit{AI}_{\mathit{XM}}, \mathit{VOS}\cdot \mathit{AI}_{\mathit{XV}}, \mathit{MOS}\}$, with $\mathit{AI}_{\mathit{XM}}$ the matrix operations per memory byte and $\mathit{AI}_{\mathit{XV}}$ the matrix operations per vector operation. The model identifies which of the three resources bounds each compressed kernel and is used to size DECA's parameters: $W$ output elements per vector operation and $L$ lookup tables. DECA itself is a three-stage pipeline, with dequantization via lookup tables, expansion via a bitmask-controlled crossbar fed by parallel-prefix-sum indices, and scaling, plus two loaders and two output registers for hardware double buffering. TEPL is the remaining mechanism: a new instruction that updates a loader's metadata, triggers the tile fetch, and completes only when the decompressed tile is in a tile register, while a TEPL queue and two ports let it execute speculatively out of order, removing per-iteration fences.
What would settle it
Run the software-only compressed GEMM kernels on a real server with the same core and HBM and check whether profiling shows the vector/SIMD units, not memory or the matrix engine, saturated; then measure a DECA prototype on the same kernels to see whether the speedup reaches the predicted 4x and whether next-token latency improves by 1.6–2.6x.
Extended reading notes
Core claim
The central claim is that the limiting resource for compressed GEMMs on a CPU with an in-core tile-matrix engine is neither memory bandwidth nor matrix throughput, but the core's vector decompression path, and that a dedicated near-core unit can relieve it cheaply. The paper formalizes this with the Roof-Surface equation, giving tiles per second as $\mathit{TPS} = \min\{\mathit{MBW}\cdot \mathit{AI}_{\mathit{XM}}, \mathit{VOS}\cdot \mathit{AI}_{\mathit{XV}}, \mathit{MOS}\}$, where the three rates are memory, vector, and matrix throughput. DECA performs dequantization through lookup tables, de-sparsification through bitmask-driven expansion, and optional group scaling, producing ready-to-use BF16 tiles; TEPL fuses the metadata write and the tile load into one instruction that is issued speculatively and out of order, hiding the core-to-accelerator communication. In the paper's simulation of a 56-core HBM server, compressed GEMMs run up to 4x faster than with optimized software kernels, and next-token time for Llama2-70B and OPT-66B falls by 1.6x–2.6x, with DECA's total area estimated below 0.2% of the die.
Load-bearing premise
The reported speedups depend on the paper's cycle-level simulator reproducing a real 56-core server's core, matrix engine, L2, and HBM behavior accurately, and the paper reports no calibration against physical hardware.
Editorial extensions
If this is right
- On HBM systems, compressed GEMM kernels that were vector-bound move to the memory- or matrix-bound region, so DECA reaches near the performance ceiling predicted by the Roof-Surface model.
- The Roof-Surface model dimensions the accelerator: $\{W=32, L=8\}$ achieves near-saturated performance, while an overprovisioned $\{W=64, L=64\}$ design adds less than 3% performance, preventing overbuilding.
- Next-token latency for Llama2-70B and OPT-66B drops by 1.6x–2.6x versus software-only decompression and by 2.5x–5.0x versus the uncompressed BF16 model.
- Sixteen DECA-augmented cores outperform 56 conventional cores on the compressed GEMM workload, so the freed cores can be power-gated or repurposed for other work.
- TEPL's out-of-order invocation is decisive at low densities, roughly doubling performance at 5% weight density relative to store-and-fence invocation of the accelerator.
Reading between the lines
- A direct testable extension is to apply the same three-rate equation to GPUs, whose tensor cores, SIMT lanes, and HBM form the same matrix/vector/memory triangle; the paper notes that a decompression unit inspired by DECA could attach to a data-movement engine.
- Because the binomial bubble model assumes nonzeros are spread uniformly through a tile, real pruned LLMs often cluster nonzeros, so DECA's throughput on actual weights should be measured rather than predicted from density alone.
- Because TEPL is a generic near-core invocation mechanism, the authors' design implies a reusable architectural slot: future accelerators that preprocess tile-shaped data, not just decompression, could use the same queue, ports, and speculative-issue discipline.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a three-part solution for compressed LLM inference on CPUs equipped with in-core matrix engines (AMX/TMUL) and HBM. First, it introduces the Roof-Surface model, a 3D analytical performance model that expresses GeMM throughput as the minimum of memory, vector, and matrix processing rates. Second, it designs DECA, a near-core accelerator that dequantizes and de-sparsifies weight tiles off the core, with a W/L design space explored through the Roof-Surface model. Third, it introduces TEPL, an ISA extension that allows out-of-order, speculative invocation of DECA from the core. The evaluation, performed in a modified Sniper-based simulator of a 56-core Sapphire Rapids-class system with HBM, reports up to 4x speedups over libxsmm for compressed GeMMs and 1.6x–2.6x speedups for Llama2-70B and OPT-66B next-token generation, with DECA performance close to the Roof-Surface prediction.
Significance. If the simulated results transfer to real hardware, the paper makes a useful contribution: the Roof-Surface model provides clear insight into why software decompression is vector-bound with HBM, and DECA plus TEPL offers a plausible near-core design that offloads decompression without burdening the core's superscalar width. The work is notable for explicitly separating memory, vector, and matrix bounds, for using the model to dimension the accelerator, and for validating the W/L choice against three simulated configurations (underprovisioned, best, overprovisioned). The TEPL mechanism, with speculative invocation and squash handling, is an interesting and reusable idea. At the same time, the central quantitative claims rest entirely on an unvalidated internal simulator and on a uniform-sparsity assumption in the analytical model, so the current evidence is not yet sufficient to establish the claimed speedups on real LLM workloads.
major comments (3)
- [Section 6.2, binomial bubble formula] The bubble formula that feeds the Roof-Surface model and the design-space exploration assumes nonzeros are independently and uniformly distributed in each W-element window. Real LLM pruning methods such as SparseGPT produce correlated, often clustered masks, not i.i.d. Bernoulli masks. Under clustered sparsity, the probability that a window contains more than Lq nonzeros is higher than the binomial prediction, so the Dequantization stage will inject more bubbles than the formula estimates. Since the evaluation in Sections 8 and 9 describes sparse workloads only by aggregate density and gives no evidence that the simulated matrices have realistic clustered structure, both the model-based DSE (W=32, L=8) and the simulated speedups in Figures 12–13 and Table 4 are likely based on the same optimistic uniform-sparsity assumption. This is load-bearing for the paper's central claim that DECA hides decompression overhead and approaches the Roof-Surface bound. Please add experiments with actual SparseGPT/AWQ-pruned weight matrices, or at least with synthetic clustered sparsity patterns, and report the resulting bubble counts, the selected W/L point, and the end-to-end speedups.
- [Section 8, simulator methodology] All quantitative claims—the libxsmm baseline, the 4x compressed-GeMM speedup, and the 1.6x–2.6x LLM speedups—are produced by an internal Sniper-based simulator extended with DECA and TEPL, but the manuscript reports no calibration or validation against a physical Sapphire Rapids system. In particular, the absolute performance of the optimized software baseline depends on how faithfully the simulator models AMX/TMUL throughput, AVX instruction mix, memory latency, L2/HBM bandwidth, and out-of-order execution. Without at least a calibration study against real SPR measurements for the software baseline, or a sensitivity analysis over the key simulator parameters, the numerical speedups should be regarded as preliminary. This is a major concern because the paper's headline results are simulation-only.
- [Section 9.2, design-space validation] The validation of the Roof-Surface DSE uses only three simulated configurations (W=8,L=4, W=32,L=8, W=64,L=64) and reports only two relative comparisons (2x faster than underprovisioned; less than 3% faster than overprovisioned). This supports the qualitative ordering but does not validate the model's absolute prediction that the chosen point is on the boundary between the VEC-bound and non-VEC-bound regions. A more direct check would be to compare the model-predicted AIXV and TPS values against the simulated values for each of the three points, and to repeat the comparison for workloads that are not generated under the model's own uniform-sparsity assumption. As written, the agreement between the model and the simulator may partly reflect the shared assumption rather than an independent confirmation.
minor comments (4)
- [Section 7] There is a typographical error in the sentence 'such an increase is unnecessary for our kernels. : most of them become bound by memory after escaping the vector-bound region.' The stray colon should be removed.
- [Section 6.2] The summation formula for expected bubbles is under-specified: the meaning of the upper limit W/Lq - 1, the event for each k, and the relationship between the binomial CDF and the number of bubbles should be stated explicitly with a short derivation, since this formula is central to the DSE.
- [Section 8] The simulator setup would be much easier to assess if it included a summary table of processor parameters: core width, ROB size, OoO window, load/store queue sizes, branch predictor, cache sizes and latencies, AMX/TMUL timing, memory latency, HBM bandwidth model, and the cycle counts latencies of the DECA pipeline stages. Currently these details are omitted.
- [Table 4] In the Llama2-70B rows, the software-only (SW) latency for BF8_20% and BF8_5% is reported as 98.1 ms for both, despite very different compression factors; please clarify whether this is a typo, an artifact of the measurement, or an expected outcome (for example, because both are bound by the same decompression overhead).
Circularity Check
No significant circularity: the paper's claimed speedups come from a simulator, and the analytical model is used for design-space exploration, not as the source of the reported results.
full rationale
The central speedup claims (up to 4x compressed GeMM, 1.6x-2.6x LLM next-token) are produced by a modified Sniper simulator extended with DECA and TEPL, as described in Section 8, not by substituting the Roof-Surface equation back into the evaluation. The Roof-Surface model is used to motivate DECA and to choose the W and L parameters in Section 6.2 and Section 9.2, but the reported performance in Figures 12-13 and Table 4 is simulated. For example, Section 9.2 states 'To dimension DECA, we pick the smallest {W,L} pair for which the predicted performance saturates' and then separately says 'We simulate the performance of these pairs to validate the model's accuracy.' This is a design-space exploration followed by an independent simulation, not a fitted parameter renamed as a prediction. The binomial uniform-sparsity bubble formula in Section 6.2 is an analytical assumption for the DSE; even if the simulator's synthetic sparse tiles share that assumption, this is a modeling and generality limitation, not a circular reduction. No equation in the paper sets the simulated DECA throughput equal to the model's predicted throughput by construction. Self-citations (TPP [17], SPADE [20], HotTiles [19]) appear in methodology and related-work comparisons and are not load-bearing; the libxsmm software baseline is an external Intel kernel. The lack of calibration against physical Sapphire Rapids hardware is a validation concern, but it is outside the circularity definition. Therefore no step meets the quote-and-reduce bar, and the derivation is self-contained as far as circularity is concerned.
Assumptions & free parameters
free parameters (2)
- DECA vector width W =
32
- DECA LUT count L =
8
assumptions (4)
- domain assumption Nonzeros are uniformly distributed across each tile for the bubble-count calculation.
- domain assumption The Sniper-based simulator extended with AMX, DECA, and TEPL faithfully models Sapphire Rapids behavior.
- domain assumption DECA's two loaders, double-buffered TOut registers, and TEPL completely hide core-to-DECA communication latency.
- domain assumption Performance is bounded by the slowest of three independent throughput rates in Equation 1, with memory latency as a secondary effect.
invented entities (2)
-
DECA PE
-
TEPL instruction
Cite this review
Pith. "Pith review of DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline Model." pith.science (2026). https://pith.science/paper/XHB7KX7X
@misc{pith2026250519349,
author = {Pith},
title = {Pith review of: DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/XHB7KX7X}},
note = {Machine review of arXiv:2505.19349}
}
read the original abstract
To alleviate the memory bandwidth bottleneck in Large Language Model (LLM) inference workloads, weight matrices are stored in memory in quantized and sparsified formats. Hence, before tiles of these matrices can be processed by in-core generalized matrix multiplication (GeMM) hardware engines, they need to be dequantized and de-sparsified. This is currently performed in software with vector operations. Unfortunately, this approach delivers only modest performance. Moreover, it is hard to understand how to improve the system, as the overall GeMM performance depends on the interaction between memory resources, vector units, and hardware matrix engines. To improve the performance of LLM inference in advanced platforms equipped with in-core GeMM engines and HBM, this paper makes three main contributions. First, it develops an analytical performance model with a 3D visual representation that provides insights into how memory resources, vector units, and hardware matrix engines interact to deliver compressed GeMM performance. Second, it proposes DECA, a new near-core ML-model decompression accelerator. DECA offloads tile de-sparsification and dequantization from the CPU, producing ready-to-use tiles for in-core GeMM engines. Third, it introduces a new ISA extension that enables out-of-order invocation of the near-core accelerator. With this extension, accelerator and core computations can interleave and overlap with high-performance. Our evaluation shows that, in a simulated 56-core Xeon 4 server with HBM, DECA accelerates the execution of compressed GeMMs by up to 4x over the use of optimized Intel software kernels. Further, DECA reduces the next-token generation time of Llama2-70B and OPT-66B by 1.6x-2.6x.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
Fryman, Wim Heirman, Jason Howard, Ibrahim Hur, Samkit Jain, Marek M
Sriram Aananthakrishnan, Shamsul Abedin, Vincent Cavé, Fabio Checconi, Kristof Du Bois, Stijn Eyerman, Joshua B. Fryman, Wim Heirman, Jason Howard, Ibrahim Hur, Samkit Jain, Marek M. Landowski, Kevin Ma, Jarrod A. Nelson, Robert Pawlowski, Fabrizio Petrini, Sebastian Szkoda, Sanjaya Tayal, Jesmin Ja- han Tithi, and Yves Vandriessche. 2023. The Intel Progr...
arXiv 2023
-
[2]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Bal- tescu, Haim ing Bao, Mo Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny B...
arXiv 2023
-
[3]
Matthew Joseph Adiletta, Jesmin Jahan Tithi, Emmanouil-Ioannis Farsarakis, Gerasimos Gerogiannis, Robert Adolf, Robert Benke, Sidharth Kashyap, Samuel Hsia, Kartik Lakhotia, Fabrizio Petrini, et al. 2023. Characterizing the scalability of graph convolutional networks on intel ® piuma. In 2023 IEEE International Symposium on Performance Analysis of Systems...
2023
-
[4]
Rajeev Balasubramonian, Andrew B Kahng, Naveen Muralimanohar, Ali Shafiee, and Vaishnav Srinivas. 2017. CACTI 7: New tools for interconnect exploration in innovative off-chip memories. ACM Transactions on Architecture and Code Optimization (TACO) 14, 2 (2017), 1–25
2017
-
[5]
Puneeth Bhat, José Moreira, and Satish Kumar Sadasivam. 2021. Matrix-multiply Assist Best Practices Guide . Technical Report. IBM, Tech. Rep., 2021.[Online]. Available: https://www.redbooks.ibm.com
2021
-
[6]
Arijit Biswas and Sailesh Kottapalli. 2021. Next-Gen Intel Xeon CPU-Sapphire Rapids. In Hot Chips, Vol. 33
2021
-
[7]
Rouhani Bita Darvish, Garegrat Nitin, Savell Tom, More Ankit, Han Kyung- Nam, Zhao Ritchie, Hall Mathew, Klar Jasmine, Chung Eric, Yu Yuan, Schulte Michael, Wittig Ralph, Bratt Ian, Stephens Nigel, Milanovic Jelena, Brothers John, Dubey Pradeep, Cornea Marius, Heinecke Alexander, Rodriguez Andres, Langhammer Martin, Deng Summer, Naumov Maxim, Micikevicius...
2023
-
[8]
Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag
Show all 92 references
-
[9]
Nafea Bshara. 2024. AWS Trainium: The Journey for Designing and Optimization Full Stack ML Hardware. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 (La Jolla, CA, USA)(ASPLOS ’24). Asso...
2024
-
[10]
Cagla Cakir, Ron Ho, Jon Lexau, and Ken Mai. 2015. Modeling and design of high- radix on-chip crossbar switches. InProceedings of the 9th International Symposium on Networks-on-Chip. 1–8
2015
-
[11]
Carlson, Wim Heirman, Stijn Eyerman, Ibrahim Hur, and Lieven Eeck- hout
Trevor E. Carlson, Wim Heirman, Stijn Eyerman, Ibrahim Hur, and Lieven Eeck- hout. 2014. An Evaluation of High-Level Mechanistic Core Models. ACM Trans- actions on Architecture and Code Optimization (TACO) 11, 3, Article 28 (Aug. 2014), 25 pages
2014
-
[12]
Xinyu Chen, Yao Chen, Feng Cheng, Hongshi Tan, Bingsheng He, and Weng-Fai Wong. 2022. ReGraph: Scaling graph processing on HBM-enabled FPGAs with heterogeneous pipelines. In 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1342–1358
2022
-
[13]
João PL de Carvalho, José E Moreira, and José Nelson Amaral. 2022. Compiling for the IBM matrix engine for enterprise workloads. IEEE Micro 42, 5 (2022), 34–40
2022
-
[14]
Lei Deng, Guoqi Li, Song Han, Luping Shi, and Yuan Xie. 2020. Model compression and hardware acceleration for neural networks: A comprehensive survey. Proc. IEEE 108, 4 (2020), 485–532
2020
-
[15]
Haozheng Fan, Hao Zhou, Guangtai Huang, Parameswaran Raman, Xinwei Fu, Gaurav Gupta, Dhananjay Ram, Yida Wang, and Jun Huan. 2024. HLAT: High- quality Large Language Model Pre-trained on AWS Trainium. arXiv preprint arXiv:2404.10630 (2024)
2024 arXiv
-
[16]
Elias Frantar and Dan Alistarh. 2023. SparseGPT: Massive Language Models Can Be Accurately Pruned in One-shot. In International Conference on Machine Learning. PMLR, 10323–10337
2023
-
[17]
Evangelos Georganas, Dhiraj Kalamkar, Sasikanth Avancha, Menachem Adelman, Cristina Anderson, Alexander Breuer, Jeremy Bruestle, Narendra Chaudhary, Abhisek Kundu, Denise Kutnick, Frank Laub, Vasimuddin Md, Sanchit Misra, Ramanarayan Mohanty, Hans Pabst, Barukh Ziv, and Alexan...
2021
-
[18]
Evangelos Georganas, Dhiraj Kalamkar, Kirill Voronin, Abhisek Kundu, Antonio Noack, Hans Pabst, Alexander Breuer, and Alexander Heinecke. 2023. Harnessing Deep Learning and HPC Kernels via High-Level Loop and Tensor Abstractions on CPU Architectures. arXiv preprint arXiv:2304....
2023 arXiv
-
[19]
Gerasimos Gerogiannis, Sriram Aananthakrishnan, Josep Torrellas, and Ibrahim Hur. 2024. HotTiles: Accelerating SpMM with Heterogeneous Accelerator Archi- tectures. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 1012–1028
2024
-
[20]
Gerasimos Gerogiannis, Serif Yesil, Damitha Lenadora, Dingyuan Cao, Charith Mendis, and Josep Torrellas. 2023. SPADE: A Flexible and Scalable Accelerator for SpMM and SDDMM. InProceedings of the 50th Annual International Symposium on Computer Architecture (Orlando, FL, USA) (I...
2023
-
[21]
Soroush Ghodrati, Sean Kinzer, Hanyang Xu, Rohan Mahapatra, Yoonsung Kim, Byung Hoon Ahn, Dong Kai Wang, Lavanya Karthikeyan, Amir Yazdanbakhsh, Jongse Park, Nam Sung Kim, and Hadi Esmaeilzadeh. 2024. Tandem processor: Grappling with emerging operators in neural networks. In P...
2024
-
[22]
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer. 2022. A survey of quantization methods for efficient neural network inference. In Low-Power Computer Vision. Chapman and Hall/CRC, 291–326
2022
-
[23]
Ashish Gondimalla, Noah Chesnut, Mithuna Thottethodi, and TN Vijaykumar
-
[24]
Zhangxiaowen Gong, Houxiang Ji, Christopher W Fletcher, Christopher J Hughes, Sara Baghsorkhi, and Josep Torrellas. 2020. Save: Sparsity-aware vector engine for accelerating dnn training and inference on CPUs. In 2020 53rd Annual IEEE/ACM International Symposium on Microarchit...
2020
-
[25]
Zhangxiaowen Gong, Houxiang Ji, Yao Yao, Christopher W Fletcher, Christopher J Hughes, and Josep Torrellas. 2022. Graphite: optimizing graph neural networks on CPUs through cooperative software-hardware techniques. In Proceedings of the 49th Annual International Symposium on C...
2022
-
[26]
Oh, Yeonhong Park, Yoonho Song, Jung-Hun Park, Sanghee Lee, Kyoung Park, Jae W
Tae Jun Ham, Sung Jun Jung, Seonghak Kim, Young H. Oh, Yeonhong Park, Yoonho Song, Jung-Hun Park, Sanghee Lee, Kyoung Park, Jae W. Lee, and Deog- Kyoon Jeong. 2020. Aˆ 3: Accelerating attention mechanisms in neural networks with approximation. In 2020 IEEE International Sympos...
2020
-
[27]
Tae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim, Hyunji Choi, Sung Jun Jung, and Jae W Lee. 2021. ELSA: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Arch...
2021
-
[28]
Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A Horowitz, and William J Dally. 2016. EIE: Efficient inference engine on compressed deep neural network. ACM SIGARCH Computer Architecture News 44, 3 (2016), 243– 254
2016
-
[29]
Song Han, Huizi Mao, and William J Dally. 2015. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149 (2015)
2015 arXiv
-
[30]
Simla Burcu Harma, Ayan Chakraborty, Elizaveta Kostenok, Danila Mishin, Dongho Ha, Babak Falsafi, Martin Jaggi, Ming Liu, Yunho Oh, Suvinay Sub- ramanian, and Amir Yazdanbakhsh. 2024. Effective Interplay between Sparsity and Quantization: From Theory to Practice. arXiv preprin...
2024 arXiv
-
[31]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
2016
-
[32]
Fletcher
Kartik Hegde, Hadi Asghari-Moghaddam, Michael Pellauer, Neal Crago, Aamer Jaleel, Edgar Solomonik, Joel Emer, and Christopher W. Fletcher. 2019. ExTensor: An Accelerator for Sparse Tensor Algebra. In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchi...
2019
-
[33]
Alexander Heinecke, Greg Henry, Maxwell Hutchinson, and Hans Pabst. 2016. LIBXSMM: accelerating small matrix multiplications by runtime code genera- tion. In SC’16: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . I...
2016
-
[34]
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste
-
[35]
Intel. 2022. Accelerate Artificial Intelligence (AI) Workloads with Intel Advanced Matrix Extensions (Intel AMX). https://www.intel.com/content/dam/www/ central-libraries/us/en/documents/2022-12/accelerate-ai-with-amx-sb.pdf
2022
-
[36]
Intel. 2024. Intel® 64 and IA-32 Architectures Optimization Reference Manual
2024
-
[37]
Jaeyong Jang, Yulhwa Kim, Juheun Lee, and Jae-Joon Kim. 2024. FIGNA: Inte- ger Unit-Based Accelerator Design for FP-INT GEMM Preserving Numerical Accuracy. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 760–773
2024
-
[38]
Geonhwa Jeong, Sana Damani, Abhimanyu Rajeshkumar Bambhaniya, Eric Qin, Christopher J Hughes, Sreenivas Subramoney, Hyesoon Kim, and Tushar Kr- ishna. 2023. Vegeta: Vertically-integrated extensions for sparse/dense gemm tile acceleration on cpus. In 2023 IEEE International Sym...
2023
-
[39]
Geonhwa Jeong, Eric Qin, Ananda Samajdar, Christopher J Hughes, Sreenivas Sub- ramoney, Hyesoon Kim, and Tushar Krishna. 2021. Rasa: Efficient register-aware systolic array matrix engine for cpu. In 2021 58th ACM/IEEE Design Automation Conference (DAC). IEEE, 253–258
2021
-
[40]
Norm Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, Clifford Young, Xiang Zhou, Zongwei Zhou, and David A Patterson. 2023. TPU v4: An optically reconfigurable supercomputer for machine learn...
2023
-
[41]
Christoforos Kachris. 2024. A Survey on Hardware Accelerators for Large Lan- guage Models. arXiv preprint arXiv:2401.09890 (2024)
2024 arXiv
-
[42]
Dhiraj Kalamkar, Dheevatsa Mudigere, Naveen Mellempudi, Dipankar Das, Kunal Banerjee, Sasikanth Avancha, Dharma Teja Vooturi, Nataraj Jammalamadaka, Jianyu Huang, Hector Yuen, Jiyan Yang, Jongsoo Park, Alexander Heinecke, Evangelos Georganas, Sudarshan Srinivasan, Abhisek Kund...
2019 arXiv
-
[43]
Dinesh Kalla, Nathan Smith, Fnu Samaah, and Sivaraju Kuraku. 2023. Study and analysis of chat GPT and its impact on different fields of study. International journal of innovative science and research technology 8, 3 (2023)
2023
-
[44]
Jeonghoon Kim, Jung Hyun Lee, Sungdong Kim, Joonsuk Park, Kang Min Yoo, Se Jung Kwon, and Dongsoo Lee. 2024. Memory-efficient fine-tuning of com- pressed large language models via sub-4-bit integer quantization. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[45]
Yann LeCun, John Denker, and Sara Solla. 1989. Optimal brain damage.Advances in neural information processing systems 2 (1989)
1989
-
[46]
Tailin Liang, John Glossner, Lei Wang, Shaobo Shi, and Xiaotong Zhang. 2021. Pruning and quantization for deep neural network acceleration: A survey. Neu- rocomputing 461 (2021), 370–403
2021
-
[47]
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2023. Awq: Activation-aware weight quantization for LLM compression and acceleration. arXiv preprint arXiv:2306.00978 (2023)
2023 arXiv
-
[48]
Liyang Liu, Shilong Zhang, Zhanghui Kuang, Aojun Zhou, Jing-Hao Xue, Xin- jiang Wang, Yimin Chen, Wenming Yang, Qingmin Liao, and Wayne Zhang
-
[49]
Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li, Tao Wang, and Yun Liang. 2021. Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture. In MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture. 977–991
2021
-
[50]
Liqiang Lu, Jiaming Xie, Ruirui Huang, Jiansong Zhang, Wei Lin, and Yun Liang
-
[51]
Weile Luo, Ruibo Fan, Zeyu Li, Dayou Du, Qiang Wang, and Xiaowen Chu. 2024. Benchmarking and dissecting the nvidia hopper gpu architecture. arXiv preprint arXiv:2402.13499 (2024)
2024 arXiv
-
[52]
In International Conference on Machine Learning
Group fisher pruning for practical network compression. In International Conference on Machine Learning . PMLR, 7021–7032
-
[53]
Munch, Carleton L
Nevine Nassif, Ashley O. Munch, Carleton L. Molnar, Gerald Pasdast, Sitara- man V. Lyer, Zibing Yang, Oscar Mendoza, Mark Huddart, Srikrishnan Venkatara- man, Sireesha Kandula, Rafi Marom, Alexandra M. Kern, Bill Bowhill, David R. Mulvihill, Srikanth Nimmagadda, Varma Kalidind...
2022
-
[54]
Thomas Norrie, Nishant Patil, Doe Hyun Yoon, George Kurian, Sheng Li, James Laudon, Cliff Young, Norman Jouppi, and David Patterson. 2021. The Design Process for Google’s Training Chips: TPUv2 and TPUv3. IEEE Micro 41, 2 (2021), 56–63. https://doi.org/10.1109/MM.2021.3058217
2021
-
[55]
In 2019 IEEE 27th Annual International Symposium on Field- Programmable Custom Computing Machines (FCCM)
An efficient hardware accelerator for sparse convolutional neural net- works on FPGAs. In 2019 IEEE 27th Annual International Symposium on Field- Programmable Custom Computing Machines (FCCM) . IEEE, 17–25
2019
-
[56]
Marcelo Orenes-Vera, Aninda Manocha, Jonathan Balkind, Fei Gao, Juan L Aragón, David Wentzlaff, and Margaret Martonosi. 2022. Tiny but mighty: de- signing and realizing scalable latency tolerance for manycore socs. InProceedings of the 49th Annual International Symposium on Co...
2022
-
[57]
Stefano Markidis, Steven Wei Der Chien, Erwin Laure, Ivy Bo Peng, and Jeffrey S Vetter. 2018. Nvidia tensor core programmability, performance & precision. In 2018 IEEE international parallel and distributed processing symposium workshops (IPDPSW). IEEE, 522–531
2018
-
[58]
Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rang- harajan Venkatesan, Brucek Khailany, Joel Emer, Stephen W Keckler, and William J Dally. 2017. SCNN: An accelerator for compressed-sparse convo- lutional neural networks. ACM SIGARCH computer architecture ...
2017
-
[59]
Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri, Aashaka Shah, Saeed Maleki, and Ricardo Bianchini. 2023. Splitwise: Efficient generative LLM inference using phase splitting. arXiv preprint arXiv:2311.18677 (2023)
2023 arXiv
-
[60]
NVIDIA. 2024. NVIDIA Blackwell Architecture Technical Brief. Retrieved 2024 from https://resources.nvidia.com/en-us-blackwell-architecture
2024
-
[61]
Alexandra Peste, Eugenia Iofinova, Adrian Vladu, and Dan Alistarh. 2021. Ac/dc: Alternating compressed/decompressed training of deep neural networks. Ad- vances in neural information processing systems 34 (2021), 8557–8570
2021
-
[62]
Subbarao Palacharla, Norman P Jouppi, and James E Smith. 1997. Complexity- effective superscalar processors. In Proceedings of the 24th annual international symposium on Computer architecture . 206–218. 13
1997
-
[63]
Marco Siracusa, Víctor Soria-Pardos, Francesco Sgherzi, Joshua Randall, Douglas J Joseph, Miquel Moretó Planas, and Adrià Armejach. 2023. A Tensor Marshaling Unit for Sparse Tensor Algebra on General-Purpose Processors. In Proceedings of the 56th Annual IEEE/ACM International ...
2023
-
[64]
Nitish Srivastava, Hanchen Jin, Shaden Smith, Hongbo Rong, David Albonesi, and Zhiru Zhang. 2020. Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor Computations. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). 689–702. http...
2020
-
[65]
Christodoulos Peltekis, Vasileios Titopoulos, Chrysostomos Nicopoulos, and Giorgos Dimitrakopoulos. 2024. DeMM: A Decoupled Matrix Multiplication Engine Supporting Relaxed Structured Sparsity. IEEE Computer Architecture Letters (2024)
2024
-
[66]
Qidong Su, Christina Giannoula, and Gennady Pekhimenko. 2023. The synergy of speculative decoding and batching in serving large language models. arXiv preprint arXiv:2310.18813 (2023)
2023 arXiv
-
[67]
Sungju Ryu, Hyungjun Kim, Wooseok Yi, Eunhwan Kim, Yulhwa Kim, Taesu Kim, and Jae-Joon Kim. 2022. Bitblade: Energy-efficient variable bit-precision hardware accelerator for quantized neural networks. IEEE Journal of Solid-State Circuits 57, 6 (2022), 1924–1935
2022
-
[68]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)
2017
-
[69]
Hanrui Wang, Zhekai Zhang, and Song Han. 2021. Spatten: Efficient sparse atten- tion architecture with cascade token and head pruning. In2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 97–110
2021
-
[70]
Aaron Stillmaker and Bevan Baas. 2017. Scaling equations for the accurate pre- diction of CMOS device performance from 180 nm to 7 nm. Integration, the VLSI Journal 58 (2017), 74–81. http://vcl.ece.ucdavis.edu/pubs/2017.02.VLSIintegration. TechScale/
2017
-
[71]
Wikipedia. 2024. Sapphire Rapids Die Configurations. https://en.wikipedia.org/ wiki/Sapphire_Rapids#Die_configurations
2024
-
[72]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yas- mine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucu- rull, David Esiobu, Jude Fernandes, Jeremy...
2023 arXiv
-
[73]
Haojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang, Zhongzhu Zhou, Xiafei Qiu, Yong Li, Wei Lin, and Shuaiwen Leon Song. 2023. Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Un- structured Sparsity. Proceedings of the VLDB Endowmen...
2023
-
[74]
Yifan Yang, Joel S Emer, and Daniel Sanchez. 2024. Trapezoid: A Versatile Ac- celerator for Dense and Sparse Matrix Multiplications. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). IEEE, 931–945
2024
-
[75]
Xiuying Wei, Yunchen Zhang, Yuhang Li, Xiangguo Zhang, Ruihao Gong, Jinyang Guo, and Xianglong Liu. 2023. Outlier suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling. arXiv preprint arXiv:2304.09145 (2023)
2023 arXiv
-
[76]
Zhihang Yuan, Yuzhang Shang, Yang Zhou, Zhen Dong, Zhe Zhou, Chenhao Xue, Bingzhe Wu, Zhikai Li, Qingyi Gu, Yong Jae Lee, Yan Yan, Beidi Chen, Guangyu Sun, and Kurt Keutzer. 2024. LLM Inference Unveiled: Survey and Roofline Model Insights. arXiv preprint arXiv:2402.16363 (2024)
2024 arXiv
-
[77]
Finn Wilkinson and Simon McIntosh-Smith. 2022. An Initial Evaluation of Arm’s Scalable Matrix Extension. In 2022 IEEE/ACM International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS). IEEE, 135–140
2022
-
[78]
Hao Zhang, Dongdong Chen, and Seok-Bum Ko. 2019. New flexible multiple- precision multiply-accumulate unit for deep neural network training and infer- ence. IEEE Trans. Comput. 69, 1 (2019), 26–38
2019
-
[79]
Haopeng Zhang, Xiao Liu, and Jiawei Zhang. 2023. Summit: Iterative text sum- marization via chatGPT. arXiv preprint arXiv:2305.14835 (2023)
2023 arXiv
-
[80]
Binwei Yao, Ming Jiang, Diyi Yang, and Junjie Hu. 2023. Empowering LLM-based machine translation with cultural awareness. arXiv preprint arXiv:2305.14328 (2023)
2023 arXiv
-
[81]
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2...
2022 arXiv
-
[82]
Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong
-
[83]
Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. 2024. Atom: Low- bit quantization for efficient and accurate LLM serving. Proceedings of Machine Learning and Systems 6 (2024), 196–209
2024
-
[84]
Pengyuan Zhou, Lin Wang, Zhi Liu, Yanbin Hao, Pan Hui, Sasu Tarkoma, and Jussi Kangasharju. 2024. A survey on generative AI and LLM for video generation, understanding, and streaming. arXiv preprint arXiv:2404.16038 (2024)
2024 arXiv
-
[85]
Xunyu Zhu, Jian Li, Yong Liu, Can Ma, and Weiping Wang. 2023. A survey on model compression for large language models. arXiv preprint arXiv:2308.07633 (2023)
2023 arXiv
-
[86]
Shijin Zhang, Zidong Du, Lei Zhang, Huiying Lan, Shaoli Liu, Ling Li, Qi Guo, Tianshi Chen, and Yunji Chen. 2016. Cambricon-X: An accelerator for sparse neural networks. In 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1–12
2016
-
[88]
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong W...
2023 arXiv
-
[92]
Zeyu Zhu, Fanrong Li, Gang Li, Zejian Liu, Zitao Mo, Qinghao Hu, Xiaoyao Liang, and Jian Cheng. 2024. MEGA: A Memory-Efficient GNN Accelerator Ex- ploiting Degree-Aware Mixed-Precision Quantization. In 2024 IEEE International Symposium on High-Performance Computer Architecture...
2024
-
[2015]
In Proceedings of the 2015 ACM/SIGDA international symposium on field-programmable gate arrays
Optimizing FPGA-based accelerator design for deep convolutional neural networks. In Proceedings of the 2015 ACM/SIGDA international symposium on field-programmable gate arrays. 161–170
2015
-
[2019]
In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microar- chitecture
SparTen: A sparse tensor accelerator for convolutional neural networks. In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microar- chitecture. 151–165
-
[2020]
What is the state of neural network pruning? Proceedings of machine learning and systems 2 (2020), 129–146
2020
-
[2021]
Journal of Machine Learning Research 22, 241 (2021), 1–124
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. Journal of Machine Learning Research 22, 241 (2021), 1–124
2021
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.