Pith. sign in

REVIEW 3 major objections 4 minor 92 references

DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline Model

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A near-core decompression accelerator claims up to 4x faster compressed matrix multiplies and 1.6–2.6x faster LLM token generation on CPUs.

desk verdict A solid simulation-based architecture paper with a genuinely useful model, held back by an uncalibrated simulator and a uniform-sparsity assumption that may inflate the speedups. read the letter →

arxiv 2505.19349 v2 pith:XHB7KX7X submitted 2025-05-25 cs.AR

classification cs.AR
keywords LLMinferencemodelcompressionquantizationunstructuredsparsitynear-coreacceleratordecompressionrooflinematrixmultiplication
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

LLM inference on modern server CPUs is limited by memory bandwidth, so weight matrices are stored quantized and sparse. The paper argues that the usual two-dimensional roofline model hides the real bottleneck: on high-bandwidth-memory machines, the vector instructions that dequantize and de-sparsify tiles saturate the CPU's SIMD units before memory or the matrix engine does. It proposes a three-resource performance model, called the Roof-Surface, whose minimum-of-three rates predicts when this vector-bound regime applies. To remove the bottleneck, it designs DECA, a small near-core unit that converts compressed tiles into the dense BF16 tiles the in-core matrix engine expects, and TEPL, an instruction extension that lets the core invoke DECA out of order so decompression overlaps with computation. The paper reports up to 4x speedups on compressed matrix multiplies and 1.6–2.6x faster next-token generation for two large open models.

What carries the argument

The load-bearing mechanism is the Roof-Surface model: a three-dimensional performance bound whose equation is $\mathit{TPS} = \min\{\mathit{MBW}\cdot \mathit{AI}_{\mathit{XM}}, \mathit{VOS}\cdot \mathit{AI}_{\mathit{XV}}, \mathit{MOS}\}$, with $\mathit{AI}_{\mathit{XM}}$ the matrix operations per memory byte and $\mathit{AI}_{\mathit{XV}}$ the matrix operations per vector operation. The model identifies which of the three resources bounds each compressed kernel and is used to size DECA's parameters: $W$ output elements per vector operation and $L$ lookup tables. DECA itself is a three-stage pipeline, with dequantization via lookup tables, expansion via a bitmask-controlled crossbar fed by parallel-prefix-sum indices, and scaling, plus two loaders and two output registers for hardware double buffering. TEPL is the remaining mechanism: a new instruction that updates a loader's metadata, triggers the tile fetch, and completes only when the decompressed tile is in a tile register, while a TEPL queue and two ports let it execute speculatively out of order, removing per-iteration fences.

What would settle it

Run the software-only compressed GEMM kernels on a real server with the same core and HBM and check whether profiling shows the vector/SIMD units, not memory or the matrix engine, saturated; then measure a DECA prototype on the same kernels to see whether the speedup reaches the predicted 4x and whether next-token latency improves by 1.6–2.6x.

Watch

Extended reading notes

Core claim

The central claim is that the limiting resource for compressed GEMMs on a CPU with an in-core tile-matrix engine is neither memory bandwidth nor matrix throughput, but the core's vector decompression path, and that a dedicated near-core unit can relieve it cheaply. The paper formalizes this with the Roof-Surface equation, giving tiles per second as $\mathit{TPS} = \min\{\mathit{MBW}\cdot \mathit{AI}_{\mathit{XM}}, \mathit{VOS}\cdot \mathit{AI}_{\mathit{XV}}, \mathit{MOS}\}$, where the three rates are memory, vector, and matrix throughput. DECA performs dequantization through lookup tables, de-sparsification through bitmask-driven expansion, and optional group scaling, producing ready-to-use BF16 tiles; TEPL fuses the metadata write and the tile load into one instruction that is issued speculatively and out of order, hiding the core-to-accelerator communication. In the paper's simulation of a 56-core HBM server, compressed GEMMs run up to 4x faster than with optimized software kernels, and next-token time for Llama2-70B and OPT-66B falls by 1.6x–2.6x, with DECA's total area estimated below 0.2% of the die.

Load-bearing premise

The reported speedups depend on the paper's cycle-level simulator reproducing a real 56-core server's core, matrix engine, L2, and HBM behavior accurately, and the paper reports no calibration against physical hardware.

Editorial extensions

If this is right

  • On HBM systems, compressed GEMM kernels that were vector-bound move to the memory- or matrix-bound region, so DECA reaches near the performance ceiling predicted by the Roof-Surface model.
  • The Roof-Surface model dimensions the accelerator: $\{W=32, L=8\}$ achieves near-saturated performance, while an overprovisioned $\{W=64, L=64\}$ design adds less than 3% performance, preventing overbuilding.
  • Next-token latency for Llama2-70B and OPT-66B drops by 1.6x–2.6x versus software-only decompression and by 2.5x–5.0x versus the uncompressed BF16 model.
  • Sixteen DECA-augmented cores outperform 56 conventional cores on the compressed GEMM workload, so the freed cores can be power-gated or repurposed for other work.
  • TEPL's out-of-order invocation is decisive at low densities, roughly doubling performance at 5% weight density relative to store-and-fence invocation of the accelerator.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct testable extension is to apply the same three-rate equation to GPUs, whose tensor cores, SIMT lanes, and HBM form the same matrix/vector/memory triangle; the paper notes that a decompression unit inspired by DECA could attach to a data-movement engine.
  • Because the binomial bubble model assumes nonzeros are spread uniformly through a tile, real pruned LLMs often cluster nonzeros, so DECA's throughput on actual weights should be measured rather than predicted from density alone.
  • Because TEPL is a generic near-core invocation mechanism, the authors' design implies a reusable architectural slot: future accelerators that preprocess tile-shaped data, not just decompression, could use the same queue, ports, and speculative-issue discipline.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a three-part solution for compressed LLM inference on CPUs equipped with in-core matrix engines (AMX/TMUL) and HBM. First, it introduces the Roof-Surface model, a 3D analytical performance model that expresses GeMM throughput as the minimum of memory, vector, and matrix processing rates. Second, it designs DECA, a near-core accelerator that dequantizes and de-sparsifies weight tiles off the core, with a W/L design space explored through the Roof-Surface model. Third, it introduces TEPL, an ISA extension that allows out-of-order, speculative invocation of DECA from the core. The evaluation, performed in a modified Sniper-based simulator of a 56-core Sapphire Rapids-class system with HBM, reports up to 4x speedups over libxsmm for compressed GeMMs and 1.6x–2.6x speedups for Llama2-70B and OPT-66B next-token generation, with DECA performance close to the Roof-Surface prediction.

Significance. If the simulated results transfer to real hardware, the paper makes a useful contribution: the Roof-Surface model provides clear insight into why software decompression is vector-bound with HBM, and DECA plus TEPL offers a plausible near-core design that offloads decompression without burdening the core's superscalar width. The work is notable for explicitly separating memory, vector, and matrix bounds, for using the model to dimension the accelerator, and for validating the W/L choice against three simulated configurations (underprovisioned, best, overprovisioned). The TEPL mechanism, with speculative invocation and squash handling, is an interesting and reusable idea. At the same time, the central quantitative claims rest entirely on an unvalidated internal simulator and on a uniform-sparsity assumption in the analytical model, so the current evidence is not yet sufficient to establish the claimed speedups on real LLM workloads.

major comments (3)
  1. [Section 6.2, binomial bubble formula] The bubble formula that feeds the Roof-Surface model and the design-space exploration assumes nonzeros are independently and uniformly distributed in each W-element window. Real LLM pruning methods such as SparseGPT produce correlated, often clustered masks, not i.i.d. Bernoulli masks. Under clustered sparsity, the probability that a window contains more than Lq nonzeros is higher than the binomial prediction, so the Dequantization stage will inject more bubbles than the formula estimates. Since the evaluation in Sections 8 and 9 describes sparse workloads only by aggregate density and gives no evidence that the simulated matrices have realistic clustered structure, both the model-based DSE (W=32, L=8) and the simulated speedups in Figures 12–13 and Table 4 are likely based on the same optimistic uniform-sparsity assumption. This is load-bearing for the paper's central claim that DECA hides decompression overhead and approaches the Roof-Surface bound. Please add experiments with actual SparseGPT/AWQ-pruned weight matrices, or at least with synthetic clustered sparsity patterns, and report the resulting bubble counts, the selected W/L point, and the end-to-end speedups.
  2. [Section 8, simulator methodology] All quantitative claims—the libxsmm baseline, the 4x compressed-GeMM speedup, and the 1.6x–2.6x LLM speedups—are produced by an internal Sniper-based simulator extended with DECA and TEPL, but the manuscript reports no calibration or validation against a physical Sapphire Rapids system. In particular, the absolute performance of the optimized software baseline depends on how faithfully the simulator models AMX/TMUL throughput, AVX instruction mix, memory latency, L2/HBM bandwidth, and out-of-order execution. Without at least a calibration study against real SPR measurements for the software baseline, or a sensitivity analysis over the key simulator parameters, the numerical speedups should be regarded as preliminary. This is a major concern because the paper's headline results are simulation-only.
  3. [Section 9.2, design-space validation] The validation of the Roof-Surface DSE uses only three simulated configurations (W=8,L=4, W=32,L=8, W=64,L=64) and reports only two relative comparisons (2x faster than underprovisioned; less than 3% faster than overprovisioned). This supports the qualitative ordering but does not validate the model's absolute prediction that the chosen point is on the boundary between the VEC-bound and non-VEC-bound regions. A more direct check would be to compare the model-predicted AIXV and TPS values against the simulated values for each of the three points, and to repeat the comparison for workloads that are not generated under the model's own uniform-sparsity assumption. As written, the agreement between the model and the simulator may partly reflect the shared assumption rather than an independent confirmation.
minor comments (4)
  1. [Section 7] There is a typographical error in the sentence 'such an increase is unnecessary for our kernels. : most of them become bound by memory after escaping the vector-bound region.' The stray colon should be removed.
  2. [Section 6.2] The summation formula for expected bubbles is under-specified: the meaning of the upper limit W/Lq - 1, the event for each k, and the relationship between the binomial CDF and the number of bubbles should be stated explicitly with a short derivation, since this formula is central to the DSE.
  3. [Section 8] The simulator setup would be much easier to assess if it included a summary table of processor parameters: core width, ROB size, OoO window, load/store queue sizes, branch predictor, cache sizes and latencies, AMX/TMUL timing, memory latency, HBM bandwidth model, and the cycle counts latencies of the DECA pipeline stages. Currently these details are omitted.
  4. [Table 4] In the Llama2-70B rows, the software-only (SW) latency for BF8_20% and BF8_5% is reported as 98.1 ms for both, despite very different compression factors; please clarify whether this is a typo, an artifact of the measurement, or an expected outcome (for example, because both are bound by the same decompression overhead).

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's claimed speedups come from a simulator, and the analytical model is used for design-space exploration, not as the source of the reported results.

full rationale

The central speedup claims (up to 4x compressed GeMM, 1.6x-2.6x LLM next-token) are produced by a modified Sniper simulator extended with DECA and TEPL, as described in Section 8, not by substituting the Roof-Surface equation back into the evaluation. The Roof-Surface model is used to motivate DECA and to choose the W and L parameters in Section 6.2 and Section 9.2, but the reported performance in Figures 12-13 and Table 4 is simulated. For example, Section 9.2 states 'To dimension DECA, we pick the smallest {W,L} pair for which the predicted performance saturates' and then separately says 'We simulate the performance of these pairs to validate the model's accuracy.' This is a design-space exploration followed by an independent simulation, not a fitted parameter renamed as a prediction. The binomial uniform-sparsity bubble formula in Section 6.2 is an analytical assumption for the DSE; even if the simulator's synthetic sparse tiles share that assumption, this is a modeling and generality limitation, not a circular reduction. No equation in the paper sets the simulated DECA throughput equal to the model's predicted throughput by construction. Self-citations (TPP [17], SPADE [20], HotTiles [19]) appear in methodology and related-work comparisons and are not load-bearing; the libxsmm software baseline is an external Intel kernel. The lack of calibration against physical Sapphire Rapids hardware is a validation concern, but it is outside the circularity definition. Therefore no step meets the quote-and-reduce bar, and the derivation is self-contained as far as circularity is concerned.

Assumptions & free parameters 2 free parameters · 4 assumptions · 2 invented entities

The design parameters W=32 and L=8 are chosen by hand through the Roof-Surface DSE, the bubble-count calculation assumes uniform sparsity, and the entire evaluation sits on an internal simulator not calibrated to real hardware. The paper introduces DECA and TEPL as new entities with simulation-only support; no constants are fitted to measured data, and no item is circularly derived from the target performance claim.

free parameters (2)
  • DECA vector width W = 32
    Chosen in Section 6.2 via Roof-Surface design space exploration as the smallest width that pushes the evaluated kernels out of the VEC-bound region; reported speedups depend on this choice.
  • DECA LUT count L = 8
    Chosen together with W in Section 6.2 to keep area low while saturating predicted performance; the paper validates the W,L choices against underprovisioned and overprovisioned alternatives.
assumptions (4)
  • domain assumption Nonzeros are uniformly distributed across each tile for the bubble-count calculation.
    Section 6.2 assumes the number of nonzeros in a W-element window follows a binomial distribution with density d; real pruned models can have non-uniform or structured sparsity, which would change the expected number of pipeline bubbles and could alter the chosen W,L.
  • domain assumption The Sniper-based simulator extended with AMX, DECA, and TEPL faithfully models Sapphire Rapids behavior.
    Section 8 uses an internal Intel simulator with full AMX support; no calibration against a physical SPR system is reported, so simulated absolute latencies, bandwidths, and core timing are unverified.
  • domain assumption DECA's two loaders, double-buffered TOut registers, and TEPL completely hide core-to-DECA communication latency.
    Sections 5.2 and 5.3 rely on hardware double buffering and speculative TEPL issue to overlap decompression with AMX GeMM; the claimed near-optimal performance assumes no residual exposed latency beyond what the simulator models.
  • domain assumption Performance is bounded by the slowest of three independent throughput rates in Equation 1, with memory latency as a secondary effect.
    The Roof-Surface model treats MEM, VEC, and MTX rates as independent ceilings; the paper itself notes BF16_30% sits below the surface because of memory latency, so the model is approximate rather than exact.
invented entities (2)
  • DECA PE
    purpose: Near-core processing element that dequantizes, de-sparsifies, and scales compressed weight tiles before feeding the in-core AMX/TMUL engine.
    Simulation-only: no fabricated silicon, emulator, or real-hardware measurements are provided; its timing, area, and power come from the internal simulator and CACTI-based estimates.
  • TEPL instruction
    purpose: ISA extension that invokes a DECA loader out of order and speculatively, combining a metadata write and a tile load to remove fences and hide latency.
    Evaluated only in the modified Sniper simulator; no ISA specification, hardware implementation, or independent measurement outside the paper is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline Model." pith.science (2026). https://pith.science/paper/XHB7KX7X

@misc{pith2026250519349,
  author       = {Pith},
  title        = {Pith review of: DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XHB7KX7X}},
  note         = {Machine review of arXiv:2505.19349}
}
read the original abstract

To alleviate the memory bandwidth bottleneck in Large Language Model (LLM) inference workloads, weight matrices are stored in memory in quantized and sparsified formats. Hence, before tiles of these matrices can be processed by in-core generalized matrix multiplication (GeMM) hardware engines, they need to be dequantized and de-sparsified. This is currently performed in software with vector operations. Unfortunately, this approach delivers only modest performance. Moreover, it is hard to understand how to improve the system, as the overall GeMM performance depends on the interaction between memory resources, vector units, and hardware matrix engines. To improve the performance of LLM inference in advanced platforms equipped with in-core GeMM engines and HBM, this paper makes three main contributions. First, it develops an analytical performance model with a 3D visual representation that provides insights into how memory resources, vector units, and hardware matrix engines interact to deliver compressed GeMM performance. Second, it proposes DECA, a new near-core ML-model decompression accelerator. DECA offloads tile de-sparsification and dequantization from the CPU, producing ready-to-use tiles for in-core GeMM engines. Third, it introduces a new ISA extension that enables out-of-order invocation of the near-core accelerator. With this extension, accelerator and core computations can interleave and overlap with high-performance. Our evaluation shows that, in a simulated 56-core Xeon 4 server with HBM, DECA accelerates the execution of compressed GeMMs by up to 4x over the use of optimized Intel software kernels. Further, DECA reduces the next-token generation time of Llama2-70B and OPT-66B by 1.6x-2.6x.

Figures

Figures reproduced from arXiv: 2505.19349 by the authors.

Figure 2
Figure 2. Libxsmm com￾pressed GeMM kernel pseudocode. To achieve high performance in compressed GeMM kernels and hide the decompression overhead, Intel recently introduced a soft￾ware solution integrated in the Libxsmm framework [33] ( [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Traditional rooflines for a GeMM with N=4. [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 4
Figure 4. (a) The 3D Roof-Surface model. (b) The optimal [PITH_FULL_IMAGE:figures/full_fig_p004_4.png] view at source ↗
Figures from the paper (11 more)
Figure 6
Figure 6. Figure 6: 2D BORD for HBM with 4x VOS. Finally, [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]
Figure 5
Figure 5. Figure 5: 2D bounding-region diagrams (BORD). Figure 5a shows the BORD for HBM SPR. The figure shows the equations of the lines that separate the three regions. They are: 𝑦 = (𝑀𝐵𝑊 /𝑉𝑂𝑆) ∗ 𝑥, 𝑥 = 𝑀𝑂𝑆/𝑀𝐵𝑊 , and 𝑦 = 𝑀𝑂𝑆/𝑉𝑂𝑆. It also shows the positions of the different compressed G…
Figure 7
Figure 7. Figure 7: DECA placement next to a core. DECA shares the L2 TLB with the core like prior work [20, 25, 63] and, therefore, uses the virtual space of the CPU core. A DECA can potentially be used by multiple processes. One approach is to save and restore the DECA state on context …
Figure 8
Figure 8. Figure 8: DECA-CPU core cooperative tile processing. [PITH_FULL_IMAGE:figures/full_fig_p006_8.png]
Figure 11
Figure 11. Figure 11: displays the DECA PE microarchitecture. To understand it, we describe its multiple components. Bitmask Q Out Tile Reg Scale Factor Q POPCNT LDQ PF SQQ LDQ POPCNT Parallel Prefix Sum Bitmask Queue Scale Factor Queue next window head SD Reg Wnd . . . DD Reg TOut Reg Spa…
Figure 14
Figure 14. Figure 14: TFLOPS across all compressions for DDR and N=4. [PITH_FULL_IMAGE:figures/full_fig_p010_14.png]
Figure 12
Figure 12. Figure 12: Compressed GeMM speedup for DDR and N=1. [PITH_FULL_IMAGE:figures/full_fig_p010_12.png]
Figure 13
Figure 13. Figure 13: Compressed GeMM speedup for HBM and N=1. [PITH_FULL_IMAGE:figures/full_fig_p010_13.png]
Figure 15
Figure 15. Figure 15: DECA vs traditional vector scaling for HBM & N=1. [PITH_FULL_IMAGE:figures/full_fig_p010_15.png]
Figure 16
Figure 16. Figure 16: HBM BORDs with no DECA and with different [PITH_FULL_IMAGE:figures/full_fig_p011_16.png]
Figure 17
Figure 17. Figure 17: DECA integration features for HBM and N=4. [PITH_FULL_IMAGE:figures/full_fig_p011_17.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

92 extracted references · 50 canonical work pages

  1. [1]

    Fryman, Wim Heirman, Jason Howard, Ibrahim Hur, Samkit Jain, Marek M

    Sriram Aananthakrishnan, Shamsul Abedin, Vincent Cavé, Fabio Checconi, Kristof Du Bois, Stijn Eyerman, Joshua B. Fryman, Wim Heirman, Jason Howard, Ibrahim Hur, Samkit Jain, Marek M. Landowski, Kevin Ma, Jarrod A. Nelson, Robert Pawlowski, Fabrizio Petrini, Sebastian Szkoda, Sanjaya Tayal, Jesmin Ja- han Tithi, and Yves Vandriessche. 2023. The Intel Progr...

  2. [2]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Bal- tescu, Haim ing Bao, Mo Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny B...

  3. [3]

    Matthew Joseph Adiletta, Jesmin Jahan Tithi, Emmanouil-Ioannis Farsarakis, Gerasimos Gerogiannis, Robert Adolf, Robert Benke, Sidharth Kashyap, Samuel Hsia, Kartik Lakhotia, Fabrizio Petrini, et al. 2023. Characterizing the scalability of graph convolutional networks on intel ® piuma. In 2023 IEEE International Symposium on Performance Analysis of Systems...

  4. [4]

    Rajeev Balasubramonian, Andrew B Kahng, Naveen Muralimanohar, Ali Shafiee, and Vaishnav Srinivas. 2017. CACTI 7: New tools for interconnect exploration in innovative off-chip memories. ACM Transactions on Architecture and Code Optimization (TACO) 14, 2 (2017), 1–25

  5. [5]

    Puneeth Bhat, José Moreira, and Satish Kumar Sadasivam. 2021. Matrix-multiply Assist Best Practices Guide . Technical Report. IBM, Tech. Rep., 2021.[Online]. Available: https://www.redbooks.ibm.com

  6. [6]

    Arijit Biswas and Sailesh Kottapalli. 2021. Next-Gen Intel Xeon CPU-Sapphire Rapids. In Hot Chips, Vol. 33

  7. [7]

    Rouhani Bita Darvish, Garegrat Nitin, Savell Tom, More Ankit, Han Kyung- Nam, Zhao Ritchie, Hall Mathew, Klar Jasmine, Chung Eric, Yu Yuan, Schulte Michael, Wittig Ralph, Bratt Ian, Stephens Nigel, Milanovic Jelena, Brothers John, Dubey Pradeep, Cornea Marius, Heinecke Alexander, Rodriguez Andres, Langhammer Martin, Deng Summer, Naumov Maxim, Micikevicius...

  8. [8]

    Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag

Show all 92 references
  1. [9]

    Nafea Bshara. 2024. AWS Trainium: The Journey for Designing and Optimization Full Stack ML Hardware. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 (La Jolla, CA, USA)(ASPLOS ’24). Asso...

  2. [10]

    Cagla Cakir, Ron Ho, Jon Lexau, and Ken Mai. 2015. Modeling and design of high- radix on-chip crossbar switches. InProceedings of the 9th International Symposium on Networks-on-Chip. 1–8

  3. [11]

    Carlson, Wim Heirman, Stijn Eyerman, Ibrahim Hur, and Lieven Eeck- hout

    Trevor E. Carlson, Wim Heirman, Stijn Eyerman, Ibrahim Hur, and Lieven Eeck- hout. 2014. An Evaluation of High-Level Mechanistic Core Models. ACM Trans- actions on Architecture and Code Optimization (TACO) 11, 3, Article 28 (Aug. 2014), 25 pages

  4. [12]

    Xinyu Chen, Yao Chen, Feng Cheng, Hongshi Tan, Bingsheng He, and Weng-Fai Wong. 2022. ReGraph: Scaling graph processing on HBM-enabled FPGAs with heterogeneous pipelines. In 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1342–1358

  5. [13]

    João PL de Carvalho, José E Moreira, and José Nelson Amaral. 2022. Compiling for the IBM matrix engine for enterprise workloads. IEEE Micro 42, 5 (2022), 34–40

  6. [14]

    Lei Deng, Guoqi Li, Song Han, Luping Shi, and Yuan Xie. 2020. Model compression and hardware acceleration for neural networks: A comprehensive survey. Proc. IEEE 108, 4 (2020), 485–532

  7. [15]

    Haozheng Fan, Hao Zhou, Guangtai Huang, Parameswaran Raman, Xinwei Fu, Gaurav Gupta, Dhananjay Ram, Yida Wang, and Jun Huan. 2024. HLAT: High- quality Large Language Model Pre-trained on AWS Trainium. arXiv preprint arXiv:2404.10630 (2024)

  8. [16]

    Elias Frantar and Dan Alistarh. 2023. SparseGPT: Massive Language Models Can Be Accurately Pruned in One-shot. In International Conference on Machine Learning. PMLR, 10323–10337

  9. [17]

    Evangelos Georganas, Dhiraj Kalamkar, Sasikanth Avancha, Menachem Adelman, Cristina Anderson, Alexander Breuer, Jeremy Bruestle, Narendra Chaudhary, Abhisek Kundu, Denise Kutnick, Frank Laub, Vasimuddin Md, Sanchit Misra, Ramanarayan Mohanty, Hans Pabst, Barukh Ziv, and Alexan...

  10. [18]

    Evangelos Georganas, Dhiraj Kalamkar, Kirill Voronin, Abhisek Kundu, Antonio Noack, Hans Pabst, Alexander Breuer, and Alexander Heinecke. 2023. Harnessing Deep Learning and HPC Kernels via High-Level Loop and Tensor Abstractions on CPU Architectures. arXiv preprint arXiv:2304....

  11. [19]

    Gerasimos Gerogiannis, Sriram Aananthakrishnan, Josep Torrellas, and Ibrahim Hur. 2024. HotTiles: Accelerating SpMM with Heterogeneous Accelerator Archi- tectures. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 1012–1028

  12. [20]

    Gerasimos Gerogiannis, Serif Yesil, Damitha Lenadora, Dingyuan Cao, Charith Mendis, and Josep Torrellas. 2023. SPADE: A Flexible and Scalable Accelerator for SpMM and SDDMM. InProceedings of the 50th Annual International Symposium on Computer Architecture (Orlando, FL, USA) (I...

  13. [21]

    Soroush Ghodrati, Sean Kinzer, Hanyang Xu, Rohan Mahapatra, Yoonsung Kim, Byung Hoon Ahn, Dong Kai Wang, Lavanya Karthikeyan, Amir Yazdanbakhsh, Jongse Park, Nam Sung Kim, and Hadi Esmaeilzadeh. 2024. Tandem processor: Grappling with emerging operators in neural networks. In P...

  14. [22]

    Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer. 2022. A survey of quantization methods for efficient neural network inference. In Low-Power Computer Vision. Chapman and Hall/CRC, 291–326

  15. [23]

    Ashish Gondimalla, Noah Chesnut, Mithuna Thottethodi, and TN Vijaykumar

  16. [24]

    Zhangxiaowen Gong, Houxiang Ji, Christopher W Fletcher, Christopher J Hughes, Sara Baghsorkhi, and Josep Torrellas. 2020. Save: Sparsity-aware vector engine for accelerating dnn training and inference on CPUs. In 2020 53rd Annual IEEE/ACM International Symposium on Microarchit...

  17. [25]

    Zhangxiaowen Gong, Houxiang Ji, Yao Yao, Christopher W Fletcher, Christopher J Hughes, and Josep Torrellas. 2022. Graphite: optimizing graph neural networks on CPUs through cooperative software-hardware techniques. In Proceedings of the 49th Annual International Symposium on C...

  18. [26]

    Oh, Yeonhong Park, Yoonho Song, Jung-Hun Park, Sanghee Lee, Kyoung Park, Jae W

    Tae Jun Ham, Sung Jun Jung, Seonghak Kim, Young H. Oh, Yeonhong Park, Yoonho Song, Jung-Hun Park, Sanghee Lee, Kyoung Park, Jae W. Lee, and Deog- Kyoon Jeong. 2020. Aˆ 3: Accelerating attention mechanisms in neural networks with approximation. In 2020 IEEE International Sympos...

  19. [27]

    Tae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim, Hyunji Choi, Sung Jun Jung, and Jae W Lee. 2021. ELSA: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Arch...

  20. [28]

    Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A Horowitz, and William J Dally. 2016. EIE: Efficient inference engine on compressed deep neural network. ACM SIGARCH Computer Architecture News 44, 3 (2016), 243– 254

  21. [29]

    Song Han, Huizi Mao, and William J Dally. 2015. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149 (2015)

  22. [30]

    Simla Burcu Harma, Ayan Chakraborty, Elizaveta Kostenok, Danila Mishin, Dongho Ha, Babak Falsafi, Martin Jaggi, Ming Liu, Yunho Oh, Suvinay Sub- ramanian, and Amir Yazdanbakhsh. 2024. Effective Interplay between Sparsity and Quantization: From Theory to Practice. arXiv preprin...

  23. [31]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778

  24. [32]

    Fletcher

    Kartik Hegde, Hadi Asghari-Moghaddam, Michael Pellauer, Neal Crago, Aamer Jaleel, Edgar Solomonik, Joel Emer, and Christopher W. Fletcher. 2019. ExTensor: An Accelerator for Sparse Tensor Algebra. In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchi...

  25. [33]

    Alexander Heinecke, Greg Henry, Maxwell Hutchinson, and Hans Pabst. 2016. LIBXSMM: accelerating small matrix multiplications by runtime code genera- tion. In SC’16: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . I...

  26. [34]

    Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste

  27. [35]

    Intel. 2022. Accelerate Artificial Intelligence (AI) Workloads with Intel Advanced Matrix Extensions (Intel AMX). https://www.intel.com/content/dam/www/ central-libraries/us/en/documents/2022-12/accelerate-ai-with-amx-sb.pdf

  28. [36]

    Intel. 2024. Intel® 64 and IA-32 Architectures Optimization Reference Manual

  29. [37]

    Jaeyong Jang, Yulhwa Kim, Juheun Lee, and Jae-Joon Kim. 2024. FIGNA: Inte- ger Unit-Based Accelerator Design for FP-INT GEMM Preserving Numerical Accuracy. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 760–773

  30. [38]

    Geonhwa Jeong, Sana Damani, Abhimanyu Rajeshkumar Bambhaniya, Eric Qin, Christopher J Hughes, Sreenivas Subramoney, Hyesoon Kim, and Tushar Kr- ishna. 2023. Vegeta: Vertically-integrated extensions for sparse/dense gemm tile acceleration on cpus. In 2023 IEEE International Sym...

  31. [39]

    Geonhwa Jeong, Eric Qin, Ananda Samajdar, Christopher J Hughes, Sreenivas Sub- ramoney, Hyesoon Kim, and Tushar Krishna. 2021. Rasa: Efficient register-aware systolic array matrix engine for cpu. In 2021 58th ACM/IEEE Design Automation Conference (DAC). IEEE, 253–258

  32. [40]

    Norm Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, Clifford Young, Xiang Zhou, Zongwei Zhou, and David A Patterson. 2023. TPU v4: An optically reconfigurable supercomputer for machine learn...

  33. [41]

    Christoforos Kachris. 2024. A Survey on Hardware Accelerators for Large Lan- guage Models. arXiv preprint arXiv:2401.09890 (2024)

  34. [42]

    Dhiraj Kalamkar, Dheevatsa Mudigere, Naveen Mellempudi, Dipankar Das, Kunal Banerjee, Sasikanth Avancha, Dharma Teja Vooturi, Nataraj Jammalamadaka, Jianyu Huang, Hector Yuen, Jiyan Yang, Jongsoo Park, Alexander Heinecke, Evangelos Georganas, Sudarshan Srinivasan, Abhisek Kund...

  35. [43]

    Dinesh Kalla, Nathan Smith, Fnu Samaah, and Sivaraju Kuraku. 2023. Study and analysis of chat GPT and its impact on different fields of study. International journal of innovative science and research technology 8, 3 (2023)

  36. [44]

    Jeonghoon Kim, Jung Hyun Lee, Sungdong Kim, Joonsuk Park, Kang Min Yoo, Se Jung Kwon, and Dongsoo Lee. 2024. Memory-efficient fine-tuning of com- pressed large language models via sub-4-bit integer quantization. Advances in Neural Information Processing Systems 36 (2024)

  37. [45]

    Yann LeCun, John Denker, and Sara Solla. 1989. Optimal brain damage.Advances in neural information processing systems 2 (1989)

  38. [46]

    Tailin Liang, John Glossner, Lei Wang, Shaobo Shi, and Xiaotong Zhang. 2021. Pruning and quantization for deep neural network acceleration: A survey. Neu- rocomputing 461 (2021), 370–403

  39. [47]

    Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2023. Awq: Activation-aware weight quantization for LLM compression and acceleration. arXiv preprint arXiv:2306.00978 (2023)

  40. [48]

    Liyang Liu, Shilong Zhang, Zhanghui Kuang, Aojun Zhou, Jing-Hao Xue, Xin- jiang Wang, Yimin Chen, Wenming Yang, Qingmin Liao, and Wayne Zhang

  41. [49]

    Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li, Tao Wang, and Yun Liang. 2021. Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture. In MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture. 977–991

  42. [50]

    Liqiang Lu, Jiaming Xie, Ruirui Huang, Jiansong Zhang, Wei Lin, and Yun Liang

  43. [51]

    Weile Luo, Ruibo Fan, Zeyu Li, Dayou Du, Qiang Wang, and Xiaowen Chu. 2024. Benchmarking and dissecting the nvidia hopper gpu architecture. arXiv preprint arXiv:2402.13499 (2024)

  44. [52]

    In International Conference on Machine Learning

    Group fisher pruning for practical network compression. In International Conference on Machine Learning . PMLR, 7021–7032

  45. [53]

    Munch, Carleton L

    Nevine Nassif, Ashley O. Munch, Carleton L. Molnar, Gerald Pasdast, Sitara- man V. Lyer, Zibing Yang, Oscar Mendoza, Mark Huddart, Srikrishnan Venkatara- man, Sireesha Kandula, Rafi Marom, Alexandra M. Kern, Bill Bowhill, David R. Mulvihill, Srikanth Nimmagadda, Varma Kalidind...

  46. [54]

    Thomas Norrie, Nishant Patil, Doe Hyun Yoon, George Kurian, Sheng Li, James Laudon, Cliff Young, Norman Jouppi, and David Patterson. 2021. The Design Process for Google’s Training Chips: TPUv2 and TPUv3. IEEE Micro 41, 2 (2021), 56–63. https://doi.org/10.1109/MM.2021.3058217

  47. [55]

    In 2019 IEEE 27th Annual International Symposium on Field- Programmable Custom Computing Machines (FCCM)

    An efficient hardware accelerator for sparse convolutional neural net- works on FPGAs. In 2019 IEEE 27th Annual International Symposium on Field- Programmable Custom Computing Machines (FCCM) . IEEE, 17–25

  48. [56]

    Marcelo Orenes-Vera, Aninda Manocha, Jonathan Balkind, Fei Gao, Juan L Aragón, David Wentzlaff, and Margaret Martonosi. 2022. Tiny but mighty: de- signing and realizing scalable latency tolerance for manycore socs. InProceedings of the 49th Annual International Symposium on Co...

  49. [57]

    Stefano Markidis, Steven Wei Der Chien, Erwin Laure, Ivy Bo Peng, and Jeffrey S Vetter. 2018. Nvidia tensor core programmability, performance & precision. In 2018 IEEE international parallel and distributed processing symposium workshops (IPDPSW). IEEE, 522–531

  50. [58]

    Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rang- harajan Venkatesan, Brucek Khailany, Joel Emer, Stephen W Keckler, and William J Dally. 2017. SCNN: An accelerator for compressed-sparse convo- lutional neural networks. ACM SIGARCH computer architecture ...

  51. [59]

    Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri, Aashaka Shah, Saeed Maleki, and Ricardo Bianchini. 2023. Splitwise: Efficient generative LLM inference using phase splitting. arXiv preprint arXiv:2311.18677 (2023)

  52. [60]

    NVIDIA. 2024. NVIDIA Blackwell Architecture Technical Brief. Retrieved 2024 from https://resources.nvidia.com/en-us-blackwell-architecture

  53. [61]

    Alexandra Peste, Eugenia Iofinova, Adrian Vladu, and Dan Alistarh. 2021. Ac/dc: Alternating compressed/decompressed training of deep neural networks. Ad- vances in neural information processing systems 34 (2021), 8557–8570

  54. [62]

    Subbarao Palacharla, Norman P Jouppi, and James E Smith. 1997. Complexity- effective superscalar processors. In Proceedings of the 24th annual international symposium on Computer architecture . 206–218. 13

  55. [63]

    Marco Siracusa, Víctor Soria-Pardos, Francesco Sgherzi, Joshua Randall, Douglas J Joseph, Miquel Moretó Planas, and Adrià Armejach. 2023. A Tensor Marshaling Unit for Sparse Tensor Algebra on General-Purpose Processors. In Proceedings of the 56th Annual IEEE/ACM International ...

  56. [64]

    Nitish Srivastava, Hanchen Jin, Shaden Smith, Hongbo Rong, David Albonesi, and Zhiru Zhang. 2020. Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor Computations. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). 689–702. http...

  57. [65]

    Christodoulos Peltekis, Vasileios Titopoulos, Chrysostomos Nicopoulos, and Giorgos Dimitrakopoulos. 2024. DeMM: A Decoupled Matrix Multiplication Engine Supporting Relaxed Structured Sparsity. IEEE Computer Architecture Letters (2024)

  58. [66]

    Qidong Su, Christina Giannoula, and Gennady Pekhimenko. 2023. The synergy of speculative decoding and batching in serving large language models. arXiv preprint arXiv:2310.18813 (2023)

  59. [67]

    Sungju Ryu, Hyungjun Kim, Wooseok Yi, Eunhwan Kim, Yulhwa Kim, Taesu Kim, and Jae-Joon Kim. 2022. Bitblade: Energy-efficient variable bit-precision hardware accelerator for quantized neural networks. IEEE Journal of Solid-State Circuits 57, 6 (2022), 1924–1935

  60. [68]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)

  61. [69]

    Hanrui Wang, Zhekai Zhang, and Song Han. 2021. Spatten: Efficient sparse atten- tion architecture with cascade token and head pruning. In2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 97–110

  62. [70]

    Aaron Stillmaker and Bevan Baas. 2017. Scaling equations for the accurate pre- diction of CMOS device performance from 180 nm to 7 nm. Integration, the VLSI Journal 58 (2017), 74–81. http://vcl.ece.ucdavis.edu/pubs/2017.02.VLSIintegration. TechScale/

  63. [71]

    Wikipedia. 2024. Sapphire Rapids Die Configurations. https://en.wikipedia.org/ wiki/Sapphire_Rapids#Die_configurations

  64. [72]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yas- mine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucu- rull, David Esiobu, Jude Fernandes, Jeremy...

  65. [73]

    Haojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang, Zhongzhu Zhou, Xiafei Qiu, Yong Li, Wei Lin, and Shuaiwen Leon Song. 2023. Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Un- structured Sparsity. Proceedings of the VLDB Endowmen...

  66. [74]

    Yifan Yang, Joel S Emer, and Daniel Sanchez. 2024. Trapezoid: A Versatile Ac- celerator for Dense and Sparse Matrix Multiplications. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). IEEE, 931–945

  67. [75]

    Xiuying Wei, Yunchen Zhang, Yuhang Li, Xiangguo Zhang, Ruihao Gong, Jinyang Guo, and Xianglong Liu. 2023. Outlier suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling. arXiv preprint arXiv:2304.09145 (2023)

  68. [76]

    Zhihang Yuan, Yuzhang Shang, Yang Zhou, Zhen Dong, Zhe Zhou, Chenhao Xue, Bingzhe Wu, Zhikai Li, Qingyi Gu, Yong Jae Lee, Yan Yan, Beidi Chen, Guangyu Sun, and Kurt Keutzer. 2024. LLM Inference Unveiled: Survey and Roofline Model Insights. arXiv preprint arXiv:2402.16363 (2024)

  69. [77]

    Finn Wilkinson and Simon McIntosh-Smith. 2022. An Initial Evaluation of Arm’s Scalable Matrix Extension. In 2022 IEEE/ACM International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS). IEEE, 135–140

  70. [78]

    Hao Zhang, Dongdong Chen, and Seok-Bum Ko. 2019. New flexible multiple- precision multiply-accumulate unit for deep neural network training and infer- ence. IEEE Trans. Comput. 69, 1 (2019), 26–38

  71. [79]

    Haopeng Zhang, Xiao Liu, and Jiawei Zhang. 2023. Summit: Iterative text sum- marization via chatGPT. arXiv preprint arXiv:2305.14835 (2023)

  72. [80]

    Binwei Yao, Ming Jiang, Diyi Yang, and Junjie Hu. 2023. Empowering LLM-based machine translation with cultural awareness. arXiv preprint arXiv:2305.14328 (2023)

  73. [81]

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2...

  74. [82]

    Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong

  75. [83]

    Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. 2024. Atom: Low- bit quantization for efficient and accurate LLM serving. Proceedings of Machine Learning and Systems 6 (2024), 196–209

  76. [84]

    Pengyuan Zhou, Lin Wang, Zhi Liu, Yanbin Hao, Pan Hui, Sasu Tarkoma, and Jussi Kangasharju. 2024. A survey on generative AI and LLM for video generation, understanding, and streaming. arXiv preprint arXiv:2404.16038 (2024)

  77. [85]

    Xunyu Zhu, Jian Li, Yong Liu, Can Ma, and Weiping Wang. 2023. A survey on model compression for large language models. arXiv preprint arXiv:2308.07633 (2023)

  78. [86]

    Shijin Zhang, Zidong Du, Lei Zhang, Huiying Lan, Shaoli Liu, Ling Li, Qi Guo, Tianshi Chen, and Yunji Chen. 2016. Cambricon-X: An accelerator for sparse neural networks. In 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1–12

  79. [88]

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong W...

  80. [92]

    Zeyu Zhu, Fanrong Li, Gang Li, Zejian Liu, Zitao Mo, Qinghao Hu, Xiaoyao Liang, and Jian Cheng. 2024. MEGA: A Memory-Efficient GNN Accelerator Ex- ploiting Degree-Aware Mixed-Precision Quantization. In 2024 IEEE International Symposium on High-Performance Computer Architecture...

  81. [2015]

    In Proceedings of the 2015 ACM/SIGDA international symposium on field-programmable gate arrays

    Optimizing FPGA-based accelerator design for deep convolutional neural networks. In Proceedings of the 2015 ACM/SIGDA international symposium on field-programmable gate arrays. 161–170

  82. [2019]

    In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microar- chitecture

    SparTen: A sparse tensor accelerator for convolutional neural networks. In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microar- chitecture. 151–165

  83. [2020]

    What is the state of neural network pruning? Proceedings of machine learning and systems 2 (2020), 129–146

  84. [2021]

    Journal of Machine Learning Research 22, 241 (2021), 1–124

    Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. Journal of Machine Learning Research 22, 241 (2021), 1–124

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.