Pith. sign in

REVIEW 2 major objections 4 minor 94 references

Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration

T0 review · 2 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read This paper claims that adding small hard blocks that understand structured sparsity to FPGA fabric lets one slice handle dense, 2:4, 1:3, and 1:4 sparse GEMMs at full utilization, yielding up to 5x frequency and 3.52x DNN speedups.

desk verdict SST slices are a solid, well-evaluated architecture proposal putting structured sparsity into in-fabric FPGA tensor blocks; all claims are modeled, not silicon, but the central design is defensible and the baseline-speedup concern does not hold up. read the letter →

arxiv 2502.03763 v1 pith:EKXMGI6F submitted 2025-02-06 cs.AR

classification cs.AR
keywords FPGAarchitecturestructuredsparsitysystolicarraysparseprocessingelementGEMMacceleratorin-fabrichardblockdeeplearninghardware2:4
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper's goal is to put structured sparsity support directly into the FPGA fabric as a hard in-fabric block, rather than leaving sparse GEMM to soft logic or dense tensor blocks. It proposes systolic sparse tensor (SST) slices: small 4x4 systolic arrays of sparse processing elements that can operate dense, 2:4 (50%), 1:3 (66.7%), or 1:4 (75%) sparse, switching per layer. Across all modes each processing element keeps doing one multiply-accumulate per cycle, so sparsity translates almost linearly into fewer cycles for the same GEMM. If the modeled numbers hold, an FPGA with these slices would run sparse GEMMs at up to 5x the clock frequency of a CLB+DSP implementation while using up to 10.9x less area, and would accelerate sparse DeiT/ConvNeXt models by up to 3.52x over dense in-fabric acceleration.

What carries the argument

The load-bearing mechanism is the sparse processing element (SPE) datapath: a MAC unit whose A-operand port is fed by an index-selected lane of a multi-register B buffer. In dense mode the SPE is a normal systolic PE; in 2:4, 1:3, and 1:4 modes the compressed A values and their 2-bit indices travel through pipeline registers while B values are held in groups of four (or three) and muxed by the index, so the MAC never stalls. Around this core, the SST slice is a 4x4 output-stationary systolic array with a six-element buffer that converts diagonal PE completions into column-wise writes, and with dedicated vertical wires that chain slices in a column, cutting global-routing pressure. The sparsity-level signal (dense/2:4/1:3/1:4) and the d_type signal (int8/bfloat16) are dynamic controls, so one accelerator can change its operation per layer.

What would settle it

Build a 7nm test chip with one column of SST slices and run a 40x40 sparse GEMM against a CLB+DSP implementation on the same die; if the slice's measured clock frequency or per-tile area differs from the modeled values by more than a few percent, or if 1:3 mode does not sustain one MAC per cycle per SPE, the central claims are wrong.

Watch

Extended reading notes

Core claim

The central discovery is that a single in-fabric block, the SST slice, can absorb the hardware cost of multiple structured-sparsity patterns while keeping the systolic dataflow intact. Each 4x4 slice is built from sparse processing elements that, in sparse modes, read the non-zero entries of matrix A plus 2-bit location indices, load four (or three) B values into registers, and use a 4:1 multiplexer to pick the B value that pairs with each A value. This makes a K-deep reduction finish in K/s cycles for sparsity ratio s (s=2, 3, or 4) with every SPE busy every cycle, and it lets dense QKV-style GEMMs run without added idle cycles. The slice also includes a six-element buffer to extract output columns evenly from a diagonal-finishing systolic array, and vertical dedicated wires between slices that reduce routing wirelength. The paper positions this as the first structured-sparsity support inside FPGA fabric, reporting up to 5x higher frequency and 10.9x lower area than traditional FPGA implementations, with small area overhead over dense-only in-fabric slices.

Load-bearing premise

The 5x frequency and 10.9x area claims come from a modeled 7nm FPGA, not a fabricated chip, so they stand or fall on whether that model's delay and area numbers reflect real silicon.

Editorial extensions

If this is right

  • FPGA vendors could add a hard block like the SST slice to their fabric and give sparse DNN accelerators the same frequency and area benefits that dense in-fabric tensor blocks gave dense workloads.
  • Because the sparsity mode is set dynamically, one GEMM accelerator can mix dense, 2:4, 1:3, and 1:4 across layers, letting each DNN layer run at its best accuracy-speed tradeoff without reconfiguration.
  • Sparse DeiT and ConvNeXt models would see 1.88x to 3.52x inference speedups over dense in-fabric acceleration at the reported accuracy levels, with weight-memory reductions up to 3.5x from the compressed format.
  • Vertical dedicated wires between slices reduce routing wirelength by 15 to 31 percent, so scaling to larger systolic arrays is cheaper than with global-routing-only tensor slices.
  • Non-AI FPGA workloads lose less than 1 percent in maximum frequency, so adding these blocks would not degrade general-purpose FPGA use.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the 1:3 mode is truly free in hardware, the same mux-based SPE design could be adopted in non-FPGA tensor units (GPU or CPU matrix cores) that currently support only 2:4 sparsity, though the paper does not extend its claim beyond FPGAs.
  • The dynamic sparsity-level control signal suggests a runtime-adaptive scheme where a deployed accelerator chooses a sparsity level per layer based on input statistics or accuracy targets; the paper only evaluates static layer-wise assignments.
  • The index-based compressed format could be generalized to other N:M ratios such as 2:6 or 3:8 with wider index fields; whether the area overhead stays as low as for 2:4, 1:3, and 1:4 is a testable design question the paper does not address.
  • A concrete test would be to prototype the SPE datapath in soft logic on a commercial FPGA and measure whether 1:3 mode really adds zero LUTs and flip-flops over the 2:4 plus 1:4 datapath, since the claim that 1:3 is free depends on the fourth mux input being simply unused.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper proposes adding in-fabric 2D systolic 'systolic sparse tensor' (SST) slices to FPGAs for accelerating GEMMs with structured sparsity levels dense, 2:4, 1:3, and 1:4. Each SST slice is a 4x4 array of sparse processing elements that use an index-based compressed format for the sparse weight operand and multi-bank B buffers to feed four B values per cycle; a dedicated vertical interconnect reduces routing wirelength. The authors implement the slices in an ASAP7 7nm standard-cell flow and use COFFE/VTR to model a 7nm FPGA enriched with SST columns. They report up to 5x higher frequency and 10.9x lower area against CLB+DSP soft-logic implementations, low area overhead of 13.3% against dense in-fabric slices, and up to 3.52x speedup on DeiT and ConvNeXt with roughly 1% accuracy loss.

Significance. If the modeling is accepted, this is a useful and well-scoped contribution to FPGA architecture exploration: it is, to my knowledge, the first study to integrate multi-level structured sparsity into in-fabric systolic blocks, and it provides a complete parameterized design and evaluation flow (ASAP7 standard cells, COFFE, VTR). The compression arithmetic in Table 1 checks out, the 1:3 mode is a sensible reuse of the 2:4/1:4 hardware, and the comparisons against SDT_GIO and CLB+DSP are external rather than constructed, so the central speedup is not forced by definition. The reported frequency and area trends are internally consistent. The main caveats are the fairness of the area-overhead baseline and the opacity of the DNN speedup model, both of which are fixable in revision.

major comments (2)
  1. [Sec. 4.4 and Table 4] The dense SDT_GIO baseline is provisioned with four B banks per column to match the sparse SST design, although each dense SPE can consume at most one B value per cycle. This makes the cycle-count speedup comparisons fair (the baseline is compute-bound, not I/O-starved), but it understates the area cost of the sparse block relative to a realistic dense in-fabric design, which would use one B bank per column as in Fig. 4a. Because the abstract's 'minimal area increase of up to 13.3%' is a headline contribution, please report the area overhead and area efficiency against both the four-bank baseline and the natural one-bank dense baseline, and restate the corresponding claims.
  2. [Sec. 4.5 and Table 5] The speedup estimation model is described in a single sentence; no equations are given. The table mixes uniform and layer-wise sparsity and includes layers that must remain dense (e.g., QK^T/AV attention GEMMs), but the reader cannot reproduce how the native SA size, zero padding, DRAM bandwidth, and dense-layer fractions combine to yield the reported speedups. Please provide the model equations and a per-layer cycle/bandwidth breakdown for at least the DeiT-B [dense, 1:4] case so that the 3.52x headline can be independently verified.
minor comments (4)
  1. [Abstract vs. Table 5] The abstract says 'up to 3.52x speedup', but Table 5 shows 3.63x for ConvNeXt-S uniform 1:4; clarify that the 3.52x figure is the maximum under the stated ~1% accuracy-degradation constraint, or update the abstract to match the full table.
  2. [Sec. 4.5] The phrase 'QKV computation' is imprecise: for ViT, the QKV projections are weight GEMMs that could be pruned, while the QK^T and AV attention GEMMs have activation-only operands and must be dense; please use the correct terminology.
  3. [Fig. 3] The text says multiplexing logic is omitted for clarity, but the SPE diagrams would be easier to follow if the 4:1 selection path for the B values were drawn or annotated at least once.
  4. [Fig. 9a] The label 'blfoat16' in the figure legend is a typo and should read 'bfloat16'.

Circularity Check

0 steps flagged · score 0.0 of 10

No meaningful circularity: the sparse-mode speedups are direct datapath properties, and the DNN-level speedups are explicitly modeled extrapolations from the measured GEMM designs against external baselines.

full rationale

The paper does not fit any parameter and rename it as a prediction. The 2x/3x/4x sparse speedups are structural consequences of the SPE design described in Sec. 3.3: in 1:4 mode each SPE ingests four B values per cycle and completes an output in K/4 cycles. Table 1 labels these as designed speedups, not as measured discoveries. The DNN speedups in Table 5 are explicitly stated to be an analytical model based on the Sec. 4.4 GEMM implementations, so they are extrapolations from the same hardware design rather than independent predictions that secretly re-import their own inputs. The GEMM comparisons include external and semi-external baselines (CLB+DSP soft logic, SDT_GIO dense slices derived from prior work, and Versal AIE-ML), so the central architecture claims are not forced by a self-citation chain. The self-citations ([13,14,36,71,72]) support background, the SDT_GIO baseline construction, and the layer-wise sparsity search; none of them carries the central result by invoking an unverified uniqueness theorem or an ansatz. The four-B-bank SDT_GIO baseline is an area/bandwidth fairness choice; it does not shorten the dense cycle count, so the sparse/dense speedup is not an artifact of I/O starvation. The main caveat is the unmeasured COFFE/ASAP7 7nm modeling, which is an external-validity risk, not a circularity. Accordingly, no circularity steps are identified.

Assumptions & free parameters 0 free parameters · 4 assumptions · 2 invented entities

The paper's central claims rest on the fidelity of the COFFE/VTR 7nm FPGA model, the assumption that GEMM dominates DNN runtime, the assumption that N:M sparsity is deployable without unacceptable accuracy loss, and the representativeness of the Versal AIE-ML optimized kernels used for comparison. No silicon measurement or third-party artifact independently confirms the SST slice's frequency, area, or DNN speedup.

assumptions (4)
  • domain assumption COFFE/VTR modeled 7nm FPGA with ASAP7 PDK accurately represents real FPGA delay and area.
    All frequency and area results in Tables 3-4 and Figs. 5-7 come from this model; there is no silicon validation.
  • domain assumption GEMM accounts for more than 90% of DNN execution time and other operations can be overlapped.
    Used to justify the analytical speedup model in Sec. 4.5; citations [4,79] provide the basis.
  • domain assumption N:M structured sparsity can be applied to DNN layers with accuracy loss acceptable for deployment.
    The paper demonstrates this for DeiT-S/B and ConvNeXt-S on ImageNet-1K, but the central claim that SSTs accelerate a wide variety of DNNs depends on this generalizing.
  • domain assumption The Versal AIE-ML codes used for comparison are representative of its achievable performance.
    Comparison in Fig. 9 uses AMD's optimized kernels, not an independently tuned implementation.
invented entities (2)
  • SST slice (systolic sparse tensor slice)
    purpose: In-fabric FPGA block for dense and structured-sparse GEMM.
    The block exists only in COFFE/VTR simulation; no fabricated test chip independently demonstrates its frequency, area, or power.
  • 1:3 structured sparsity level
    purpose: A sparsity pattern with one nonzero per three consecutive elements, designed to fill the gap between 2:4 and 1:4.
    The paper's accuracy experiments on DeiT/ConvNeXt are the only evidence for its utility; no prior literature confirms it as a standard pattern.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration." pith.science (2026). https://pith.science/paper/EKXMGI6F

@misc{pith2026250203763,
  author       = {Pith},
  title        = {Pith review of: Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EKXMGI6F}},
  note         = {Machine review of arXiv:2502.03763}
}
read the original abstract

FPGA architectures have recently been enhanced to meet the substantial computational demands of modern deep neural networks (DNNs). To this end, both FPGA vendors and academic researchers have proposed in-fabric blocks that perform efficient tensor computations. However, these blocks are primarily optimized for dense computation, while most DNNs exhibit sparsity. To address this limitation, we propose incorporating structured sparsity support into FPGA architectures. We architect 2D systolic in-fabric blocks, named systolic sparse tensor (SST) slices, that support multiple degrees of sparsity to efficiently accelerate a wide variety of DNNs. SSTs support dense operation, 2:4 (50%) and 1:4 (75%) sparsity, as well as a new 1:3 (66.7%) sparsity level to further increase flexibility. When demonstrating on general matrix multiplication (GEMM) accelerators, which are the heart of most current DNN accelerators, our sparse SST-based designs attain up to 5x higher FPGA frequency and 10.9x lower area, compared to traditional FPGAs. Moreover, evaluation of the proposed SSTs on state-of-the-art sparse ViT and CNN models exhibits up to 3.52x speedup with minimal area increase of up to 13.3%, compared to dense in-fabric acceleration.

Figures

Figures reproduced from arXiv: 2502.03763 by the authors.

Figure 1
Figure 1. 50% unstructured sparse matrix (a) and 2:4 (50%) [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Systolic Sparse Tensor slice architecture. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Sparsity modes in Systolic Sparse Elements of the SST slices (multiplexing logic omitted for clarity). [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: 2D systolic GEMM design: dense implementation (a) and dynamic configuration of all supported sparsity modes (b). [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Routing wirelength comparison of dense GEMM mapped to SST and SDT_GIO slices, for various SA sizes. aforementioned global routing I/Os, the SST slices need to access 6 SBs, thus spanning 6 CLB (logic) tiles. However, since more out￾puts are needed for the SDT_GIOs (pri…
Figure 8
Figure 8. Figure 8: SST-based GEMM implemented in VTR (SA size: [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 7
Figure 7. Figure 7: Total area comparison of various sparse GEMM implementations utilizing SSTs vs. SDT_GIOs vs. CLBs+DSPs. all SA sizes. We note that the lowest area increase, i.e., 10.2% and 13.3% for int8 and bfloat16, respectively, occurs for the highest, 40x40 SA size. This is mainly…
Figure 9
Figure 9. Figure 9: Compute utilization (a) and compression ratio (b) [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

94 extracted references · 60 canonical work pages

  1. [1]

    ASAP7 PDK and Cell Libraries

    2023. ASAP7 PDK and Cell Libraries. https://github.com/The-OpenROAD- Project/asap7

  2. [2]

    Achronix. 2024. Machine Learning Processor. https://www.achronix.com/ machine-learning-processor

  3. [3]

    Achronix. 2024. Speedster7t FPGAs: Product Brief. https://www.achronix.com/ sites/default/files/docs/Speedster7t_Product_Brief_PB033.pdf

  4. [4]

    Robert Adolf, Saketh Rama, Brandon Reagen, Gu-Yeon Wei, and David Brooks

  5. [5]

    AMD. 2023. AI Engine API User Guide. https://www.xilinx.com/htmldocs/ xilinx2023_2/aiengine_api/aie_api/doc/group__group__mmul.html

  6. [6]

    AMD. 2023. AI Engine-ML Programming. https://github.com/Xilinx/Vitis- Tutorials/blob/2023.2/AI_Engine_Development/AIE-ML/Design_Tutorials/01- AIE-ML-programming-and-optimization/ComputeOptimization.md

  7. [7]

    AMD. 2023. AMD CDNA 3 Architecture. https://www.amd.com/content/ dam/amd/en/documents/instinct-tech-docs/white-papers/amd-cdna-3-white- paper.pdf

  8. [8]

    AMD. 2024. AI Engine-ML Kernel and Graph Programming Guide (UG1603). https://docs.amd.com/r/en-US/ug1603-ai-engine-ml-kernel-graph/ Overview?tocId=UXKuGgdu4z0LghCWyYalPw

Show all 94 references
  1. [9]

    AMD. 2024. AMD Versal™ AI Edge Series VEK280 Evaluation Kit. https://www. xilinx.com/products/boards-and-kits/vek280.html

  2. [10]

    AMD. 2024. VCK5000 Versal Development Card. https://www.xilinx.com/ products/boards-and-kits/vck5000.html

  3. [11]

    AMD. 2024. Versal Adaptive SoC AIE-ML Architecture Manual (AM020). https: //docs.amd.com/r/en-US/am020-versal-aie-ml/Overview

  4. [12]

    AMD. 2024. Versal AI Core Series Data Sheet: DC and AC Switching Charac- teristics (DS957). https://docs.amd.com/r/en-US/ds957-versal-ai-core/DSP58- Switching-Characteristics

  5. [13]

    Aman Arora, Moinak Ghosh, Samidh Mehta, Vaughn Betz, and Lizy K. John

  6. [15]

    Andrew Boutros, Aman Arora, and Vaughn Betz. 2024. Field-Programmable Gate Array Architecture for Deep Learning: Survey & Future Directions. arXiv:2404.10076 [cs.AR] https://arxiv.org/abs/2404.10076

  7. [16]

    Andrew Boutros and Vaughn Betz. 2021. FPGA Architecture: Principles and Progression. IEEE Circuits and Systems Magazine 21, 2 (2021), 4–29. https: //doi.org/10.1109/MCAS.2021.3071607

  8. [18]

    Lei Cai, Jing Wang, Lianfeng Yu, Bonan Yan, Yaoyu Tao, and Yuchao Yang. 2023. Accelerating Neural-ODE Inference on FPGAs with Two-Stage Structured Prun- ing and History-based Stepsize Search. In Proceedings of the 2023 ACM/SIGDA International Symposium on Field Programmable Ga...

  9. [19]

    Aiken Cairncross, Basile Henry, Chris Chalmers, Douglas Reid, Jonny Shipton, Jon Fowler, Liz Corrigan, and Mike Ashby. 2023. AI Benchmarking on Achronix Speedster® 7t FPGAs. (2023)

  10. [20]

    Shijie Cao, Chen Zhang, Zhuliang Yao, Wencong Xiao, Lanshun Nie, Dechen Zhan, Yunxin Liu, Ming Wu, and Lintao Zhang. 2019. Efficient and Effective Sparse LSTM on FPGA with Bank-Balanced Sparsity. In Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmabl...

  11. [21]

    Yu-Hsin Chen, Tien-Ju Yang, Joel Emer, and Vivienne Sze. 2019. Eyeriss v2: A Flexible Accelerator for Emerging Deep Neural Networks on Mobile Devices. IEEE Journal on Emerging and Selected Topics in Circuits and Systems 9, 2 (2019), 292–308. https://doi.org/10.1109/JETCAS.2019.2910232

  12. [22]

    Charles Chiasson and Vaughn Betz. 2013. COFFE: Fully-automated transistor siz- ing for FPGAs. In2013 International Conference on Field-Programmable Technology (FPT). 34–41. https://doi.org/10.1109/FPT.2013.6718327

  13. [23]

    Clark, Vinay Vashishtha, Lucian Shifren, Aditya Gujja, Saurabh Sinha, Brian Cline, Chandarasekaran Ramamurthy, and Greg Yeric

    Lawrence T. Clark, Vinay Vashishtha, Lucian Shifren, Aditya Gujja, Saurabh Sinha, Brian Cline, Chandarasekaran Ramamurthy, and Greg Yeric. 2016. ASAP7: A 7-nm finFET predictive process design kit. Microelectronics Journal 53 (2016), 105–115. https://doi.org/10.1016/j.mejo.2016.04.006

  14. [24]

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A Large-Scale Hierarchical Image Database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition . Ieee, 248–255

  15. [25]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In North American Chapter of the Association for Computational Linguistics . https: //api.semanticscholar.org/CorpusID:52967399

  16. [26]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2020. An Image is Worth 16x16 Words: Transformers for Image Recogn...

  17. [27]

    Chao Fang, Aojun Zhou, and Zhongfeng Wang. 2022. An Algorithm–Hardware Co-Optimized Framework for Accelerating N:M Sparse Transformers. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 30, 11 (2022), 1573–

  18. [28]

    Chung, and Greg Stitt

    Jeremy Fowers, Kalin Ovtcharov, Karin Strauss, Eric S. Chung, and Greg Stitt. 2014. A High Memory Bandwidth FPGA Accelerator for Sparse Matrix-Vector Multipli- cation. In 2014 IEEE 22nd Annual International Symposium on Field-Programmable Custom Computing Machines. 36–43. http...

  19. [29]

    Ashish Gondimalla, Noah Chesnut, Mithuna Thottethodi, and T. N. Vijaykumar

  20. [31]

    Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, Zhaohui Yang, Yiman Zhang, and Dacheng Tao. 2023. A Survey on Vision Transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 1 (2023...

  21. [32]

    Horowitz, and William J

    Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A. Horowitz, and William J. Dally. 2016. EIE: Efficient Inference Engine On Compressed Deep Neural Network (ISCA ’16). IEEE Press, 243–254. https://doi.org/10.1109/ISCA. 2016.30

  22. [33]

    Song Han, Jeff Pool, John Tran, and William J. Dally. 2015. Learning both weights and connections for efficient neural networks. In Proceedings of the 28th Interna- tional Conference on Neural Information Processing Systems - Volume 1 (Montreal, Canada) (NIPS’15). MIT Press, C...

  23. [35]

    Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste

  24. [36]

    Ning-Chi Huang, Chi-Chih Chang, Wei-Cheng Lin, Endri Taka, Diana Marculescu, and Kai-Chiang Wu. 2024. ELSA: Exploiting Layer-wise N:M Sparsity for Vision Transformer Acceleration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Works...

  25. [37]

    Sitao Huang, Carl Pearson, Rakesh Nagi, Jinjun Xiong, Deming Chen, and Wen- mei Hwu. 2019. Accelerating Sparse Deep Neural Networks on FPGAs. In 2019 IEEE High Performance Extreme Computing Conference (HPEC) . 1–7. https://doi. org/10.1109/HPEC.2019.8916419

  26. [38]

    Intel. 2020. Intel Agilex Variable Precision DSP Blocks User Guide. https://www.intel.com/content/dam/altera-www/global/en_US/pdfs/literature/ hb/agilex/ug-ag-dsp.pdf

  27. [39]

    Intel. 2023. Intel Stratix 10 Embedded Memory User Guide. https: //www.intel.com/content/www/us/en/docs/programmable/683423/23- 2/embedded-memory-configurations.html

  28. [40]

    Intel. 2024. Agilex ™ 5 FPGAs and SoCs Device Data Sheet. https://www.intel. com/content/www/us/en/docs/programmable/813918/current/agilex-5-fpgas- and-socs-device-data-sheet.html

  29. [41]

    Intel. 2024. Agilex ™ 5 FPGAs: Enhanced DSP with AI Tensor Block. https://www.intel.com/content/www/us/en/content-details/776602/agilex-5- fpgas-enhanced-dsp-with-ai-tensor-block.html

  30. [42]

    Intel. 2024. Embedded Memory User Guide: Agilex ™ 5 FPGAs and SoCs. https://www.intel.com/content/www/us/en/docs/programmable/813901/ 24-1/embedded-memory-configurations.html

  31. [43]

    Intel. 2024. Embedded Memory User Guide: Agilex ™ 7 FPGAs and SoCs. https://www.intel.com/content/www/us/en/docs/programmable/683241/ 24-2/embedded-memory-configurations.html. FPGA ’25, February 27–March 1, 2025, Monterey, CA, USA Endri Taka et al

  32. [44]

    Saidul Islam, Hanae Elmekki, Ahmed Elsebai, Jamal Bentahar, Nagat Drawel, Gaith Rjoub, and Witold Pedrycz. 2024. A Comprehensive Survey on Applications of Transformers for Deep Learning Tasks. Expert Systems with Applications 241 (2024), 122666. https://doi.org/10.1016/j.eswa....

  33. [45]

    Abhishek Kumar Jain, Sharan Kumar, Aashish Tripathi, and Dinesh Gaitonde

  34. [46]

    Abhishek Kumar Jain, Hossein Omidian, Henri Fraisse, Mansimran Benipal, Lisa Liu, and Dinesh Gaitonde. 2020. A Domain-Specific Architecture for Accel- erating Sparse Matrix Vector Multiplication on FPGAs. In 2020 30th Interna- tional Conference on Field-Programmable Logic and ...

  35. [47]

    Hughes, Sreenivas Subramoney, Hyesoon Kim, and Tushar Krishna

    Geonhwa Jeong, Sana Damani, Abhimanyu Rajeshkumar Bambhaniya, Eric Qin, Christopher J. Hughes, Sreenivas Subramoney, Hyesoon Kim, and Tushar Krishna

  36. [48]

    Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B

    Norman P. Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B. Jablin, George Kurian, James Laudon, Sheng Li, Peter Ma, Xiaoyu Ma, Thomas Norrie, Nishant Patil, Sushma Prasad, Cliff Young, Zongwei Zhou, and David Patterson. 2021. Ten Lessons From Three Generations...

  37. [49]

    Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazi...

  38. [50]

    Kung. 1982. Why systolic architectures? Computer 15, 1 (1982), 37–46. https: //doi.org/10.1109/MC.1982.1653825

  39. [51]

    In 2021 IEEE High Performance Extreme Computing Conference (HPEC)

    Sparse Deep Neural Network Acceleration on HBM-Enabled FPGA Platform. In 2021 IEEE High Performance Extreme Computing Conference (HPEC). 1–7. https: //doi.org/10.1109/HPEC49654.2021.9622804

  40. [52]

    Yun Liang, Liqiang Lu, Yicheng Jin, Jiaming Xie, Ruirui Huang, Jiansong Zhang, and Wei Lin. 2022. An Efficient Hardware Design for Accelerating Sparse CNNs With NAS-Based Models. IEEE Transactions on Computer-Aided Design of Inte- grated Circuits and Systems 41, 3 (2022), 597–...

  41. [53]

    Linqiao Liu and Stephen Brown. 2021. Leveraging Fine-grained Structured Sparsity for CNN Inference on Systolic Array Architectures. In 2021 31st Interna- tional Conference on Field-Programmable Logic and Applications (FPL) . 301–305. https://doi.org/10.1109/FPL53798.2021.00060

  42. [54]

    Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. 2022. A ConvNet for the 2020s. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 11966–11976. https://doi.org/ 10.1109/CVPR52688.2022.01167

  43. [55]

    Whatmough, and Matthew Mattina

    Zhi-Gang Liu, Paul N. Whatmough, and Matthew Mattina. 2020. Sparse Systolic Tensor Array for Efficient CNN Hardware Acceleration. arXiv:2009.02381 [cs.AR] https://arxiv.org/abs/2009.02381

  44. [56]

    Whatmough, and Matthew Mattina

    Zhi-Gang Liu, Paul N. Whatmough, and Matthew Mattina. 2020. Systolic Tensor Array: An Efficient Structured-Sparse GEMM Accelerator for Mobile CNN Inference. IEEE Computer Architecture Letters 19, 1 (2020), 34–37. https: //doi.org/10.1109/LCA.2020.2979965

  45. [57]

    Whatmough, Yuhao Zhu, and Matthew Mattina

    Zhi-Gang Liu, Paul N. Whatmough, Yuhao Zhu, and Matthew Mattina. 2022. S2TA: Exploiting Structured Sparsity for Energy-Efficient Mobile CNN Acceleration. In 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 573–586. https://doi.org/10.1109/HPC...

  46. [58]

    Martin Langhammer, Eriko Nurvitadhi, Bogdan Pasca, and Sergey Gribok. 2021. Stratix 10 NX Architecture and Applications. In The 2021 ACM/SIGDA Inter- national Symposium on Field-Programmable Gate Arrays (Virtual Event, USA) (FPGA ’21). Association for Computing Machinery, New ...

  47. [59]

    Liqiang Lu, Jiaming Xie, Ruirui Huang, Jiansong Zhang, Wei Lin, and Yun Liang. 2019. An Efficient Hardware Accelerator for Sparse Convolutional Neu- ral Networks on FPGAs. In 2019 IEEE 27th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM) ....

  48. [60]

    Asit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius. 2021. Accelerating Sparse Deep Neural Networks. https://arxiv.org/abs/2104.08378

  49. [61]

    Murray, Oleg Petelin, Sheng Zhong, Jia Min Wang, Mohamed Eldafrawy, Jean-Philippe Legault, Eugene Sha, Aaron G

    Kevin E. Murray, Oleg Petelin, Sheng Zhong, Jia Min Wang, Mohamed Eldafrawy, Jean-Philippe Legault, Eugene Sha, Aaron G. Graham, Jean Wu, Matthew J. P. Walker, Hanqing Zeng, Panagiotis Patros, Jason Luu, Kenneth B. Kent, and Vaughn Betz. 2020. VTR 8: High-performance CAD and C...

  50. [62]

    Thomas Norrie, Nishant Patil, Doe Hyun Yoon, George Kurian, Sheng Li, James Laudon, Cliff Young, Norman Jouppi, and David Patterson. 2021. The Design Process for Google’s Training Chips: TPUv2 and TPUv3. IEEE Micro 41, 2 (2021), 56–63. https://doi.org/10.1109/MM.2021.3058217

  51. [63]

    NVIDIA. 2020. NVIDIA A100 Tensor Core GPU Architecture. https://images.nvidia.com/aem-dam/en-zz/Solutions/data-center/nvidia- ampere-architecture-whitepaper.pdf

  52. [64]

    NVIDIA. 2023. Confidential Compute on NVIDIA Hopper H100. https://images. nvidia.com/aem-dam/en-zz/Solutions/data-center/HCC-Whitepaper-v1.0.pdf

  53. [65]

    Liqiang Lu and Yun Liang. 2018. SpWA: An Efficient Sparse Winograd Convo- lutional Neural Networks Accelerator on FPGAs. In 2018 55th ACM/ESDA/IEEE Design Automation Conference (DAC). 1–6. https://doi.org/10.1109/DAC.2018. 8465842

  54. [66]

    SeyedRamin Rasoulinezhad, Hao Zhou, Lingli Wang, and Philip H.W. Leong

  55. [67]

    Ananda Samajdar, Jan Moritz Joseph, Yuhao Zhu, Paul Whatmough, Matthew Mattina, and Tushar Krishna. 2020. A Systematic Methodology for Characterizing Scalability of DNN Accelerators using SCALE-Sim. In 2020 IEEE International Symposium on Performance Analysis of Systems and So...

  56. [68]

    Gil Shomron, Tal Horowitz, and Uri Weiser. 2019. SMT-SA: Simultaneous Mul- tithreading in Systolic Arrays. IEEE Computer Architecture Letters 18, 2 (2019), 99–102. https://doi.org/10.1109/LCA.2019.2924007

  57. [69]

    Stillmaker and B

    A. Stillmaker and B. Baas. 2017. Scaling equations for the accurate prediction of CMOS device performance from 180 nm to 7 nm. Integration, the VLSI Jour- nal 58 (2017), 74–81. http://vcl.ece.ucdavis.edu/pubs/2017.02.VLSIintegration. TechScale/

  58. [70]

    Xiao Sun, Naigang Wang, Chia-Yu Chen, Jiamin Ni, Ankur Agrawal, Xiaodong Cui, Swagath Venkataramani, Kaoutar El Maghraoui, Vijayalakshmi (Viji) Srini- vasan, and Kailash Gopalakrishnan. 2020. Ultra-Low Precision 4-bit Training of Deep Neural Networks. In Advances in Neural Inf...

  59. [71]

    Endri Taka, Aman Arora, Kai-Chiang Wu, and Diana Marculescu. 2023. MaxEVA: Maximizing the Efficiency of Matrix Multiplication on Versal AI Engine. In 2023 International Conference on Field Programmable Technology (ICFPT) . 96–105. https://doi.org/10.1109/ICFPT59805.2023.00016

  60. [72]

    Subhankar Pal, Jonathan Beaumont, Dong-Hyeon Park, Aporva Amarnath, Siy- ing Feng, Chaitali Chakrabarti, Hun-Seok Kim, David Blaauw, Trevor Mudge, and Ronald Dreslinski. 2018. OuterSPACE: An Outer Product Based Sparse Ma- trix Multiplication Accelerator. In 2018 IEEE Internati...

  61. [73]

    Titopoulos, K

    V. Titopoulos, K. Alexandridis, C. Peltekis, C. Nicopoulos, and G. Dimitrakopoulos

  62. [74]

    In 2019 IEEE 27th Annual International Symposium on Field- Programmable Custom Computing Machines (FCCM)

    PIR-DSP: An FPGA DSP Block Architecture for Multi-precision Deep Neural Networks. In 2019 IEEE 27th Annual International Symposium on Field- Programmable Custom Computing Machines (FCCM) . 35–44. https://doi.org/10. 1109/FCCM.2019.00015

  63. [75]

    Fengbin Tu, Yiqi Wang, Ling Liang, Yufei Ding, Leibo Liu, Shaojun Wei, Shouyi Yin, and Yuan Xie. 2023. SDP: Co-Designing Algorithm, Dataflow, and Archi- tecture for In-SRAM Sparse NN Acceleration. IEEE Transactions on Computer- Aided Design of Integrated Circuits and Systems 4...

  64. [76]

    Vinay Vashishtha, Manoj Vangala, and Lawrence T. Clark. 2017. ASAP7 predictive design kit development and cell design technology co-optimization: Invited paper. In 2017 IEEE/ACM International Conference on Computer-Aided Design (ICCAD) . 992–998. https://doi.org/10.1109/ICCAD....

  65. [77]

    Gomez, Łukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, ...

  66. [78]

    VTR. 2023. Verilog-to-Routing Documentation. https://docs.verilogtorouting. org/en/latest/vtr/benchmarks/

  67. [79]

    Yu Emma Wang, Carole-Jean Wu, Xiaodong Wang, Kim Hazelwood, and David Brooks. 2021. Exploiting Parallelism Opportunities with Deep Learning Frame- works. ACM Trans. Archit. Code Optim. 18, 1, Article 9 (Dec 2021), 23 pages. https://doi.org/10.1145/3431388

  68. [81]

    Yannan Nellie Wu, Po-An Tsai, Saurav Muralidharan, Angshuman Parashar, Vivi- enne Sze, and Joel Emer. 2023. HighLight: Efficient and Flexible DNN Acceleration with Hierarchical Structured Sparsity. InProceedings of the 56th Annual IEEE/ACM International Symposium on Microarchi...

  69. [82]

    Xiaoru Xie, Jun Lin, Zhongfeng Wang, and Jinghe Wei. 2021. An Efficient and Flexible Accelerator Design for Sparse Convolutional Neural Networks. IEEE Transactions on Circuits and Systems I: Regular Papers 68, 7 (2021), 2936–2949. https://doi.org/10.1109/TCSI.2021.3074300

  70. [83]

    Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. 2021. Training Data-efficient Image Trans- formers & Distillation Through Attention. In Proceedings of the 38th Interna- tional Conference on Machine Learning (Proceedings of...

  71. [84]

    Sadegh Yazdanshenas and Vaughn Betz. 2019. COFFE 2: Automatic Modelling and Optimization of Complex and Heterogeneous FPGA Architectures. ACM Trans. Reconfigurable Technol. Syst. 12, 1, Article 3 (jan 2019), 27 pages. https: //doi.org/10.1145/3301298

  72. [85]

    Wenhua Ye, Xu Zhou, Joey Zhou, Cen Chen, and Kenli Li. 2023. Accelerating Attention Mechanism on FPGAs based on Efficient Reconfigurable Systolic Array. ACM Trans. Embed. Comput. Syst. 22, 6, Article 93 (nov 2023), 22 pages. https: //doi.org/10.1145/3549937

  73. [86]

    Zhihang Yuan, Chenhao Xue, Yiqi Chen, Qiang Wu, and Guangyu Sun. 2022. PTQ4ViT: Post-training Quantization for Vision Transformers with Twin Uniform Quantization. In Computer Vision – ECCV 2022: 17th European Conference, Tel A viv, Israel, October 23–27, 2022, Proceedings, Par...

  74. [87]

    Shijin Zhang, Zidong Du, Lei Zhang, Huiying Lan, Shaoli Liu, Ling Li, Qi Guo, Tianshi Chen, and Yunji Chen. 2016. Cambricon-X: An accelerator for sparse neural networks. In 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). 1–12. https://doi.org/10...

  75. [88]

    Yuxin Zhang, Mingbao Lin, ZhiHang Lin, Yiting Luo, Ke Li, Fei Chao, Yongjian Wu, and Rongrong Ji. 2022. Learning Best Combination for Efficient N:M Spar- sity. In Advances in Neural Information Processing Systems , S. Koyejo, S. Mo- hamed, A. Agarwal, D. Belgrave, K. Cho, and ...

  76. [89]

    Xuechao Wei, Cody Hao Yu, Peng Zhang, Youxiang Chen, Yuxin Wang, Han Hu, Yun Liang, and Jason Cong. 2017. Automated systolic array architecture synthesis for high throughput CNN inference on FPGAs. In 2017 54th ACM/EDAC/IEEE Design Automation Conference (DAC) . 1–6. https://do...

  77. [90]

    Aojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu, Zhijie Zhang, Kun Yuan, Wenxiu Sun, and Hongsheng Li. 2021. Learning N:M Fine-grained Structured Sparse Neural Networks From Scratch. arXiv:2102.04010 [cs.CV] https://arxiv.org/abs/ 2102.04010

  78. [91]

    Maohua Zhu, Tao Zhang, Zhenyu Gu, and Yuan Xie. 2019. Sparse Tensor Core: Algorithm and Hardware Co-Design for Vector-wise Sparse Neural Networks on Modern GPUs. In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture (Columbus, OH, USA) (MICRO ...

  79. [92]

    Shuxin Yang, Chenchen Ding, Mingqiang Huang, Kai Li, Chenghao Li, Zikun Wei, Sixiao Huang, Jingyao Dong, Liuyang Zhang, and Hao Yu. 2024. LAMPS: A Layer-wised Mixed-Precision-and-Sparsity Accelerator for NAS-Optimized CNNs on FPGA. In 2024 IEEE 32nd Annual International Sympos...

  80. [98]

    Zhekai Zhang, Hanrui Wang, Song Han, and William J. Dally. 2020. SpArch: Efficient Architecture for Sparse Matrix Multiplication. 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) (2020), 261–274. https://api.semanticscholar.org/CorpusID:211205022

  81. [1586]

    https://doi.org/10.1109/TVLSI.2022.3197282

  82. [2016]

    In 2016 IEEE International Symposium on Workload Characterization (IISWC)

    Fathom: Reference workloads for modern deep learning methods. In 2016 IEEE International Symposium on Workload Characterization (IISWC). IEEE, 1–10

  83. [2021]

    Sparsity in Deep Learning: Pruning and Growth for Efficient Inference and Training in Neural Networks. J. Mach. Learn. Res. 22, 1, Article 241 (jan 2021), 124 pages

  84. [2022]

    ACM Trans

    Tensor Slices: FPGA Building Blocks For The Deep Learning Era. ACM Trans. Reconfigurable Technol. Syst. 15, 4, Article 46 (aug 2022), 34 pages. https: //doi.org/10.1145/3529650

  85. [2023]

    In2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA)

    VEGETA: Vertically-Integrated Extensions for Sparse/Dense GEMM Tile Acceleration on CPUs. In2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 259–272. https://doi.org/10.1109/HPCA56546. 2023.10071058

  86. [2024]

    In 2024 Design, Automation & Test in Europe Con- ference & Exhibition (DATE)

    IndexMAC: A Custom RISC-V Vector Instruction to Accelerate Structured- Sparse Matrix Multiplications. In 2024 Design, Automation & Test in Europe Con- ference & Exhibition (DATE) . 1–6. https://doi.org/10.23919/DATE58400.2024. 10546747

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.