REVIEW 2 major objections 4 minor 94 references
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
T0 review · 2 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read This paper claims that adding small hard blocks that understand structured sparsity to FPGA fabric lets one slice handle dense, 2:4, 1:3, and 1:4 sparse GEMMs at full utilization, yielding up to 5x frequency and 3.52x DNN speedups.
desk verdict SST slices are a solid, well-evaluated architecture proposal putting structured sparsity into in-fabric FPGA tensor blocks; all claims are modeled, not silicon, but the central design is defensible and the baseline-speedup concern does not hold up. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the sparse processing element (SPE) datapath: a MAC unit whose A-operand port is fed by an index-selected lane of a multi-register B buffer. In dense mode the SPE is a normal systolic PE; in 2:4, 1:3, and 1:4 modes the compressed A values and their 2-bit indices travel through pipeline registers while B values are held in groups of four (or three) and muxed by the index, so the MAC never stalls. Around this core, the SST slice is a 4x4 output-stationary systolic array with a six-element buffer that converts diagonal PE completions into column-wise writes, and with dedicated vertical wires that chain slices in a column, cutting global-routing pressure. The sparsity-level signal (dense/2:4/1:3/1:4) and the d_type signal (int8/bfloat16) are dynamic controls, so one accelerator can change its operation per layer.
What would settle it
Build a 7nm test chip with one column of SST slices and run a 40x40 sparse GEMM against a CLB+DSP implementation on the same die; if the slice's measured clock frequency or per-tile area differs from the modeled values by more than a few percent, or if 1:3 mode does not sustain one MAC per cycle per SPE, the central claims are wrong.
Extended reading notes
Core claim
The central discovery is that a single in-fabric block, the SST slice, can absorb the hardware cost of multiple structured-sparsity patterns while keeping the systolic dataflow intact. Each 4x4 slice is built from sparse processing elements that, in sparse modes, read the non-zero entries of matrix A plus 2-bit location indices, load four (or three) B values into registers, and use a 4:1 multiplexer to pick the B value that pairs with each A value. This makes a K-deep reduction finish in K/s cycles for sparsity ratio s (s=2, 3, or 4) with every SPE busy every cycle, and it lets dense QKV-style GEMMs run without added idle cycles. The slice also includes a six-element buffer to extract output columns evenly from a diagonal-finishing systolic array, and vertical dedicated wires between slices that reduce routing wirelength. The paper positions this as the first structured-sparsity support inside FPGA fabric, reporting up to 5x higher frequency and 10.9x lower area than traditional FPGA implementations, with small area overhead over dense-only in-fabric slices.
Load-bearing premise
The 5x frequency and 10.9x area claims come from a modeled 7nm FPGA, not a fabricated chip, so they stand or fall on whether that model's delay and area numbers reflect real silicon.
Editorial extensions
If this is right
- FPGA vendors could add a hard block like the SST slice to their fabric and give sparse DNN accelerators the same frequency and area benefits that dense in-fabric tensor blocks gave dense workloads.
- Because the sparsity mode is set dynamically, one GEMM accelerator can mix dense, 2:4, 1:3, and 1:4 across layers, letting each DNN layer run at its best accuracy-speed tradeoff without reconfiguration.
- Sparse DeiT and ConvNeXt models would see 1.88x to 3.52x inference speedups over dense in-fabric acceleration at the reported accuracy levels, with weight-memory reductions up to 3.5x from the compressed format.
- Vertical dedicated wires between slices reduce routing wirelength by 15 to 31 percent, so scaling to larger systolic arrays is cheaper than with global-routing-only tensor slices.
- Non-AI FPGA workloads lose less than 1 percent in maximum frequency, so adding these blocks would not degrade general-purpose FPGA use.
Reading between the lines
- If the 1:3 mode is truly free in hardware, the same mux-based SPE design could be adopted in non-FPGA tensor units (GPU or CPU matrix cores) that currently support only 2:4 sparsity, though the paper does not extend its claim beyond FPGAs.
- The dynamic sparsity-level control signal suggests a runtime-adaptive scheme where a deployed accelerator chooses a sparsity level per layer based on input statistics or accuracy targets; the paper only evaluates static layer-wise assignments.
- The index-based compressed format could be generalized to other N:M ratios such as 2:6 or 3:8 with wider index fields; whether the area overhead stays as low as for 2:4, 1:3, and 1:4 is a testable design question the paper does not address.
- A concrete test would be to prototype the SPE datapath in soft logic on a commercial FPGA and measure whether 1:3 mode really adds zero LUTs and flip-flops over the 2:4 plus 1:4 datapath, since the claim that 1:3 is free depends on the fourth mux input being simply unused.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes adding in-fabric 2D systolic 'systolic sparse tensor' (SST) slices to FPGAs for accelerating GEMMs with structured sparsity levels dense, 2:4, 1:3, and 1:4. Each SST slice is a 4x4 array of sparse processing elements that use an index-based compressed format for the sparse weight operand and multi-bank B buffers to feed four B values per cycle; a dedicated vertical interconnect reduces routing wirelength. The authors implement the slices in an ASAP7 7nm standard-cell flow and use COFFE/VTR to model a 7nm FPGA enriched with SST columns. They report up to 5x higher frequency and 10.9x lower area against CLB+DSP soft-logic implementations, low area overhead of 13.3% against dense in-fabric slices, and up to 3.52x speedup on DeiT and ConvNeXt with roughly 1% accuracy loss.
Significance. If the modeling is accepted, this is a useful and well-scoped contribution to FPGA architecture exploration: it is, to my knowledge, the first study to integrate multi-level structured sparsity into in-fabric systolic blocks, and it provides a complete parameterized design and evaluation flow (ASAP7 standard cells, COFFE, VTR). The compression arithmetic in Table 1 checks out, the 1:3 mode is a sensible reuse of the 2:4/1:4 hardware, and the comparisons against SDT_GIO and CLB+DSP are external rather than constructed, so the central speedup is not forced by definition. The reported frequency and area trends are internally consistent. The main caveats are the fairness of the area-overhead baseline and the opacity of the DNN speedup model, both of which are fixable in revision.
major comments (2)
- [Sec. 4.4 and Table 4] The dense SDT_GIO baseline is provisioned with four B banks per column to match the sparse SST design, although each dense SPE can consume at most one B value per cycle. This makes the cycle-count speedup comparisons fair (the baseline is compute-bound, not I/O-starved), but it understates the area cost of the sparse block relative to a realistic dense in-fabric design, which would use one B bank per column as in Fig. 4a. Because the abstract's 'minimal area increase of up to 13.3%' is a headline contribution, please report the area overhead and area efficiency against both the four-bank baseline and the natural one-bank dense baseline, and restate the corresponding claims.
- [Sec. 4.5 and Table 5] The speedup estimation model is described in a single sentence; no equations are given. The table mixes uniform and layer-wise sparsity and includes layers that must remain dense (e.g., QK^T/AV attention GEMMs), but the reader cannot reproduce how the native SA size, zero padding, DRAM bandwidth, and dense-layer fractions combine to yield the reported speedups. Please provide the model equations and a per-layer cycle/bandwidth breakdown for at least the DeiT-B [dense, 1:4] case so that the 3.52x headline can be independently verified.
minor comments (4)
- [Abstract vs. Table 5] The abstract says 'up to 3.52x speedup', but Table 5 shows 3.63x for ConvNeXt-S uniform 1:4; clarify that the 3.52x figure is the maximum under the stated ~1% accuracy-degradation constraint, or update the abstract to match the full table.
- [Sec. 4.5] The phrase 'QKV computation' is imprecise: for ViT, the QKV projections are weight GEMMs that could be pruned, while the QK^T and AV attention GEMMs have activation-only operands and must be dense; please use the correct terminology.
- [Fig. 3] The text says multiplexing logic is omitted for clarity, but the SPE diagrams would be easier to follow if the 4:1 selection path for the B values were drawn or annotated at least once.
- [Fig. 9a] The label 'blfoat16' in the figure legend is a typo and should read 'bfloat16'.
Circularity Check
No meaningful circularity: the sparse-mode speedups are direct datapath properties, and the DNN-level speedups are explicitly modeled extrapolations from the measured GEMM designs against external baselines.
full rationale
The paper does not fit any parameter and rename it as a prediction. The 2x/3x/4x sparse speedups are structural consequences of the SPE design described in Sec. 3.3: in 1:4 mode each SPE ingests four B values per cycle and completes an output in K/4 cycles. Table 1 labels these as designed speedups, not as measured discoveries. The DNN speedups in Table 5 are explicitly stated to be an analytical model based on the Sec. 4.4 GEMM implementations, so they are extrapolations from the same hardware design rather than independent predictions that secretly re-import their own inputs. The GEMM comparisons include external and semi-external baselines (CLB+DSP soft logic, SDT_GIO dense slices derived from prior work, and Versal AIE-ML), so the central architecture claims are not forced by a self-citation chain. The self-citations ([13,14,36,71,72]) support background, the SDT_GIO baseline construction, and the layer-wise sparsity search; none of them carries the central result by invoking an unverified uniqueness theorem or an ansatz. The four-B-bank SDT_GIO baseline is an area/bandwidth fairness choice; it does not shorten the dense cycle count, so the sparse/dense speedup is not an artifact of I/O starvation. The main caveat is the unmeasured COFFE/ASAP7 7nm modeling, which is an external-validity risk, not a circularity. Accordingly, no circularity steps are identified.
Assumptions & free parameters
assumptions (4)
- domain assumption COFFE/VTR modeled 7nm FPGA with ASAP7 PDK accurately represents real FPGA delay and area.
- domain assumption GEMM accounts for more than 90% of DNN execution time and other operations can be overlapped.
- domain assumption N:M structured sparsity can be applied to DNN layers with accuracy loss acceptable for deployment.
- domain assumption The Versal AIE-ML codes used for comparison are representative of its achievable performance.
invented entities (2)
-
SST slice (systolic sparse tensor slice)
-
1:3 structured sparsity level
Cite this review
Pith. "Pith review of Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration." pith.science (2026). https://pith.science/paper/EKXMGI6F
@misc{pith2026250203763,
author = {Pith},
title = {Pith review of: Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration},
year = {2026},
howpublished = {\url{https://pith.science/paper/EKXMGI6F}},
note = {Machine review of arXiv:2502.03763}
}
read the original abstract
FPGA architectures have recently been enhanced to meet the substantial computational demands of modern deep neural networks (DNNs). To this end, both FPGA vendors and academic researchers have proposed in-fabric blocks that perform efficient tensor computations. However, these blocks are primarily optimized for dense computation, while most DNNs exhibit sparsity. To address this limitation, we propose incorporating structured sparsity support into FPGA architectures. We architect 2D systolic in-fabric blocks, named systolic sparse tensor (SST) slices, that support multiple degrees of sparsity to efficiently accelerate a wide variety of DNNs. SSTs support dense operation, 2:4 (50%) and 1:4 (75%) sparsity, as well as a new 1:3 (66.7%) sparsity level to further increase flexibility. When demonstrating on general matrix multiplication (GEMM) accelerators, which are the heart of most current DNN accelerators, our sparse SST-based designs attain up to 5x higher FPGA frequency and 10.9x lower area, compared to traditional FPGAs. Moreover, evaluation of the proposed SSTs on state-of-the-art sparse ViT and CNN models exhibits up to 3.52x speedup with minimal area increase of up to 13.3%, compared to dense in-fabric acceleration.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
ASAP7 PDK and Cell Libraries
2023. ASAP7 PDK and Cell Libraries. https://github.com/The-OpenROAD- Project/asap7
2023
-
[2]
Achronix. 2024. Machine Learning Processor. https://www.achronix.com/ machine-learning-processor
2024
-
[3]
Achronix. 2024. Speedster7t FPGAs: Product Brief. https://www.achronix.com/ sites/default/files/docs/Speedster7t_Product_Brief_PB033.pdf
2024
-
[4]
Robert Adolf, Saketh Rama, Brandon Reagen, Gu-Yeon Wei, and David Brooks
-
[5]
AMD. 2023. AI Engine API User Guide. https://www.xilinx.com/htmldocs/ xilinx2023_2/aiengine_api/aie_api/doc/group__group__mmul.html
2023
-
[6]
AMD. 2023. AI Engine-ML Programming. https://github.com/Xilinx/Vitis- Tutorials/blob/2023.2/AI_Engine_Development/AIE-ML/Design_Tutorials/01- AIE-ML-programming-and-optimization/ComputeOptimization.md
2023
-
[7]
AMD. 2023. AMD CDNA 3 Architecture. https://www.amd.com/content/ dam/amd/en/documents/instinct-tech-docs/white-papers/amd-cdna-3-white- paper.pdf
2023
-
[8]
AMD. 2024. AI Engine-ML Kernel and Graph Programming Guide (UG1603). https://docs.amd.com/r/en-US/ug1603-ai-engine-ml-kernel-graph/ Overview?tocId=UXKuGgdu4z0LghCWyYalPw
2024
Show all 94 references
-
[9]
AMD. 2024. AMD Versal™ AI Edge Series VEK280 Evaluation Kit. https://www. xilinx.com/products/boards-and-kits/vek280.html
2024
-
[10]
AMD. 2024. VCK5000 Versal Development Card. https://www.xilinx.com/ products/boards-and-kits/vck5000.html
2024
-
[11]
AMD. 2024. Versal Adaptive SoC AIE-ML Architecture Manual (AM020). https: //docs.amd.com/r/en-US/am020-versal-aie-ml/Overview
2024
-
[12]
AMD. 2024. Versal AI Core Series Data Sheet: DC and AC Switching Charac- teristics (DS957). https://docs.amd.com/r/en-US/ds957-versal-ai-core/DSP58- Switching-Characteristics
2024
-
[13]
Aman Arora, Moinak Ghosh, Samidh Mehta, Vaughn Betz, and Lizy K. John
-
[15]
Andrew Boutros, Aman Arora, and Vaughn Betz. 2024. Field-Programmable Gate Array Architecture for Deep Learning: Survey & Future Directions. arXiv:2404.10076 [cs.AR] https://arxiv.org/abs/2404.10076
2024
-
[16]
Andrew Boutros and Vaughn Betz. 2021. FPGA Architecture: Principles and Progression. IEEE Circuits and Systems Magazine 21, 2 (2021), 4–29. https: //doi.org/10.1109/MCAS.2021.3071607
2021
-
[18]
Lei Cai, Jing Wang, Lianfeng Yu, Bonan Yan, Yaoyu Tao, and Yuchao Yang. 2023. Accelerating Neural-ODE Inference on FPGAs with Two-Stage Structured Prun- ing and History-based Stepsize Search. In Proceedings of the 2023 ACM/SIGDA International Symposium on Field Programmable Ga...
2023
-
[19]
Aiken Cairncross, Basile Henry, Chris Chalmers, Douglas Reid, Jonny Shipton, Jon Fowler, Liz Corrigan, and Mike Ashby. 2023. AI Benchmarking on Achronix Speedster® 7t FPGAs. (2023)
2023
-
[20]
Shijie Cao, Chen Zhang, Zhuliang Yao, Wencong Xiao, Lanshun Nie, Dechen Zhan, Yunxin Liu, Ming Wu, and Lintao Zhang. 2019. Efficient and Effective Sparse LSTM on FPGA with Bank-Balanced Sparsity. In Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmabl...
2019
-
[21]
Yu-Hsin Chen, Tien-Ju Yang, Joel Emer, and Vivienne Sze. 2019. Eyeriss v2: A Flexible Accelerator for Emerging Deep Neural Networks on Mobile Devices. IEEE Journal on Emerging and Selected Topics in Circuits and Systems 9, 2 (2019), 292–308. https://doi.org/10.1109/JETCAS.2019.2910232
2019
-
[22]
Charles Chiasson and Vaughn Betz. 2013. COFFE: Fully-automated transistor siz- ing for FPGAs. In2013 International Conference on Field-Programmable Technology (FPT). 34–41. https://doi.org/10.1109/FPT.2013.6718327
2013
-
[23]
Clark, Vinay Vashishtha, Lucian Shifren, Aditya Gujja, Saurabh Sinha, Brian Cline, Chandarasekaran Ramamurthy, and Greg Yeric
Lawrence T. Clark, Vinay Vashishtha, Lucian Shifren, Aditya Gujja, Saurabh Sinha, Brian Cline, Chandarasekaran Ramamurthy, and Greg Yeric. 2016. ASAP7: A 7-nm finFET predictive process design kit. Microelectronics Journal 53 (2016), 105–115. https://doi.org/10.1016/j.mejo.2016.04.006
2016 doi
-
[24]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A Large-Scale Hierarchical Image Database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition . Ieee, 248–255
2009
-
[25]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In North American Chapter of the Association for Computational Linguistics . https: //api.semanticscholar.org/CorpusID:52967399
2019
-
[26]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2020. An Image is Worth 16x16 Words: Transformers for Image Recogn...
2020 arXiv
-
[27]
Chao Fang, Aojun Zhou, and Zhongfeng Wang. 2022. An Algorithm–Hardware Co-Optimized Framework for Accelerating N:M Sparse Transformers. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 30, 11 (2022), 1573–
2022
-
[28]
Chung, and Greg Stitt
Jeremy Fowers, Kalin Ovtcharov, Karin Strauss, Eric S. Chung, and Greg Stitt. 2014. A High Memory Bandwidth FPGA Accelerator for Sparse Matrix-Vector Multipli- cation. In 2014 IEEE 22nd Annual International Symposium on Field-Programmable Custom Computing Machines. 36–43. http...
2014 doi
-
[29]
Ashish Gondimalla, Noah Chesnut, Mithuna Thottethodi, and T. N. Vijaykumar
-
[31]
Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, Zhaohui Yang, Yiman Zhang, and Dacheng Tao. 2023. A Survey on Vision Transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 1 (2023...
2023
-
[32]
Horowitz, and William J
Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A. Horowitz, and William J. Dally. 2016. EIE: Efficient Inference Engine On Compressed Deep Neural Network (ISCA ’16). IEEE Press, 243–254. https://doi.org/10.1109/ISCA. 2016.30
2016 doi
-
[33]
Song Han, Jeff Pool, John Tran, and William J. Dally. 2015. Learning both weights and connections for efficient neural networks. In Proceedings of the 28th Interna- tional Conference on Neural Information Processing Systems - Volume 1 (Montreal, Canada) (NIPS’15). MIT Press, C...
2015
-
[35]
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste
-
[36]
Ning-Chi Huang, Chi-Chih Chang, Wei-Cheng Lin, Endri Taka, Diana Marculescu, and Kai-Chiang Wu. 2024. ELSA: Exploiting Layer-wise N:M Sparsity for Vision Transformer Acceleration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Works...
2024
-
[37]
Sitao Huang, Carl Pearson, Rakesh Nagi, Jinjun Xiong, Deming Chen, and Wen- mei Hwu. 2019. Accelerating Sparse Deep Neural Networks on FPGAs. In 2019 IEEE High Performance Extreme Computing Conference (HPEC) . 1–7. https://doi. org/10.1109/HPEC.2019.8916419
2019
-
[38]
Intel. 2020. Intel Agilex Variable Precision DSP Blocks User Guide. https://www.intel.com/content/dam/altera-www/global/en_US/pdfs/literature/ hb/agilex/ug-ag-dsp.pdf
2020
-
[39]
Intel. 2023. Intel Stratix 10 Embedded Memory User Guide. https: //www.intel.com/content/www/us/en/docs/programmable/683423/23- 2/embedded-memory-configurations.html
2023
-
[40]
Intel. 2024. Agilex ™ 5 FPGAs and SoCs Device Data Sheet. https://www.intel. com/content/www/us/en/docs/programmable/813918/current/agilex-5-fpgas- and-socs-device-data-sheet.html
2024
-
[41]
Intel. 2024. Agilex ™ 5 FPGAs: Enhanced DSP with AI Tensor Block. https://www.intel.com/content/www/us/en/content-details/776602/agilex-5- fpgas-enhanced-dsp-with-ai-tensor-block.html
2024
-
[42]
Intel. 2024. Embedded Memory User Guide: Agilex ™ 5 FPGAs and SoCs. https://www.intel.com/content/www/us/en/docs/programmable/813901/ 24-1/embedded-memory-configurations.html
2024
-
[43]
Intel. 2024. Embedded Memory User Guide: Agilex ™ 7 FPGAs and SoCs. https://www.intel.com/content/www/us/en/docs/programmable/683241/ 24-2/embedded-memory-configurations.html. FPGA ’25, February 27–March 1, 2025, Monterey, CA, USA Endri Taka et al
2024
-
[44]
Saidul Islam, Hanae Elmekki, Ahmed Elsebai, Jamal Bentahar, Nagat Drawel, Gaith Rjoub, and Witold Pedrycz. 2024. A Comprehensive Survey on Applications of Transformers for Deep Learning Tasks. Expert Systems with Applications 241 (2024), 122666. https://doi.org/10.1016/j.eswa....
2024
-
[45]
Abhishek Kumar Jain, Sharan Kumar, Aashish Tripathi, and Dinesh Gaitonde
-
[46]
Abhishek Kumar Jain, Hossein Omidian, Henri Fraisse, Mansimran Benipal, Lisa Liu, and Dinesh Gaitonde. 2020. A Domain-Specific Architecture for Accel- erating Sparse Matrix Vector Multiplication on FPGAs. In 2020 30th Interna- tional Conference on Field-Programmable Logic and ...
2020
-
[47]
Hughes, Sreenivas Subramoney, Hyesoon Kim, and Tushar Krishna
Geonhwa Jeong, Sana Damani, Abhimanyu Rajeshkumar Bambhaniya, Eric Qin, Christopher J. Hughes, Sreenivas Subramoney, Hyesoon Kim, and Tushar Krishna
-
[48]
Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B
Norman P. Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B. Jablin, George Kurian, James Laudon, Sheng Li, Peter Ma, Xiaoyu Ma, Thomas Norrie, Nishant Patil, Sushma Prasad, Cliff Young, Zongwei Zhou, and David Patterson. 2021. Ten Lessons From Three Generations...
2021
-
[49]
Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazi...
2017
-
[50]
Kung. 1982. Why systolic architectures? Computer 15, 1 (1982), 37–46. https: //doi.org/10.1109/MC.1982.1653825
1982
-
[51]
In 2021 IEEE High Performance Extreme Computing Conference (HPEC)
Sparse Deep Neural Network Acceleration on HBM-Enabled FPGA Platform. In 2021 IEEE High Performance Extreme Computing Conference (HPEC). 1–7. https: //doi.org/10.1109/HPEC49654.2021.9622804
2021
-
[52]
Yun Liang, Liqiang Lu, Yicheng Jin, Jiaming Xie, Ruirui Huang, Jiansong Zhang, and Wei Lin. 2022. An Efficient Hardware Design for Accelerating Sparse CNNs With NAS-Based Models. IEEE Transactions on Computer-Aided Design of Inte- grated Circuits and Systems 41, 3 (2022), 597–...
2022
-
[53]
Linqiao Liu and Stephen Brown. 2021. Leveraging Fine-grained Structured Sparsity for CNN Inference on Systolic Array Architectures. In 2021 31st Interna- tional Conference on Field-Programmable Logic and Applications (FPL) . 301–305. https://doi.org/10.1109/FPL53798.2021.00060
2021
-
[54]
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. 2022. A ConvNet for the 2020s. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 11966–11976. https://doi.org/ 10.1109/CVPR52688.2022.01167
2022
-
[55]
Whatmough, and Matthew Mattina
Zhi-Gang Liu, Paul N. Whatmough, and Matthew Mattina. 2020. Sparse Systolic Tensor Array for Efficient CNN Hardware Acceleration. arXiv:2009.02381 [cs.AR] https://arxiv.org/abs/2009.02381
2020 arXiv
-
[56]
Whatmough, and Matthew Mattina
Zhi-Gang Liu, Paul N. Whatmough, and Matthew Mattina. 2020. Systolic Tensor Array: An Efficient Structured-Sparse GEMM Accelerator for Mobile CNN Inference. IEEE Computer Architecture Letters 19, 1 (2020), 34–37. https: //doi.org/10.1109/LCA.2020.2979965
2020
-
[57]
Whatmough, Yuhao Zhu, and Matthew Mattina
Zhi-Gang Liu, Paul N. Whatmough, Yuhao Zhu, and Matthew Mattina. 2022. S2TA: Exploiting Structured Sparsity for Energy-Efficient Mobile CNN Acceleration. In 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 573–586. https://doi.org/10.1109/HPC...
2022
-
[58]
Martin Langhammer, Eriko Nurvitadhi, Bogdan Pasca, and Sergey Gribok. 2021. Stratix 10 NX Architecture and Applications. In The 2021 ACM/SIGDA Inter- national Symposium on Field-Programmable Gate Arrays (Virtual Event, USA) (FPGA ’21). Association for Computing Machinery, New ...
2021
-
[59]
Liqiang Lu, Jiaming Xie, Ruirui Huang, Jiansong Zhang, Wei Lin, and Yun Liang. 2019. An Efficient Hardware Accelerator for Sparse Convolutional Neu- ral Networks on FPGAs. In 2019 IEEE 27th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM) ....
2019
-
[60]
Asit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius. 2021. Accelerating Sparse Deep Neural Networks. https://arxiv.org/abs/2104.08378
2021 arXiv
-
[61]
Murray, Oleg Petelin, Sheng Zhong, Jia Min Wang, Mohamed Eldafrawy, Jean-Philippe Legault, Eugene Sha, Aaron G
Kevin E. Murray, Oleg Petelin, Sheng Zhong, Jia Min Wang, Mohamed Eldafrawy, Jean-Philippe Legault, Eugene Sha, Aaron G. Graham, Jean Wu, Matthew J. P. Walker, Hanqing Zeng, Panagiotis Patros, Jason Luu, Kenneth B. Kent, and Vaughn Betz. 2020. VTR 8: High-performance CAD and C...
2020 doi
-
[62]
Thomas Norrie, Nishant Patil, Doe Hyun Yoon, George Kurian, Sheng Li, James Laudon, Cliff Young, Norman Jouppi, and David Patterson. 2021. The Design Process for Google’s Training Chips: TPUv2 and TPUv3. IEEE Micro 41, 2 (2021), 56–63. https://doi.org/10.1109/MM.2021.3058217
2021
-
[63]
NVIDIA. 2020. NVIDIA A100 Tensor Core GPU Architecture. https://images.nvidia.com/aem-dam/en-zz/Solutions/data-center/nvidia- ampere-architecture-whitepaper.pdf
2020
-
[64]
NVIDIA. 2023. Confidential Compute on NVIDIA Hopper H100. https://images. nvidia.com/aem-dam/en-zz/Solutions/data-center/HCC-Whitepaper-v1.0.pdf
2023
-
[65]
Liqiang Lu and Yun Liang. 2018. SpWA: An Efficient Sparse Winograd Convo- lutional Neural Networks Accelerator on FPGAs. In 2018 55th ACM/ESDA/IEEE Design Automation Conference (DAC). 1–6. https://doi.org/10.1109/DAC.2018. 8465842
2018 doi
-
[66]
SeyedRamin Rasoulinezhad, Hao Zhou, Lingli Wang, and Philip H.W. Leong
-
[67]
Ananda Samajdar, Jan Moritz Joseph, Yuhao Zhu, Paul Whatmough, Matthew Mattina, and Tushar Krishna. 2020. A Systematic Methodology for Characterizing Scalability of DNN Accelerators using SCALE-Sim. In 2020 IEEE International Symposium on Performance Analysis of Systems and So...
2020
-
[68]
Gil Shomron, Tal Horowitz, and Uri Weiser. 2019. SMT-SA: Simultaneous Mul- tithreading in Systolic Arrays. IEEE Computer Architecture Letters 18, 2 (2019), 99–102. https://doi.org/10.1109/LCA.2019.2924007
2019
-
[69]
Stillmaker and B
A. Stillmaker and B. Baas. 2017. Scaling equations for the accurate prediction of CMOS device performance from 180 nm to 7 nm. Integration, the VLSI Jour- nal 58 (2017), 74–81. http://vcl.ece.ucdavis.edu/pubs/2017.02.VLSIintegration. TechScale/
2017
-
[70]
Xiao Sun, Naigang Wang, Chia-Yu Chen, Jiamin Ni, Ankur Agrawal, Xiaodong Cui, Swagath Venkataramani, Kaoutar El Maghraoui, Vijayalakshmi (Viji) Srini- vasan, and Kailash Gopalakrishnan. 2020. Ultra-Low Precision 4-bit Training of Deep Neural Networks. In Advances in Neural Inf...
2020
-
[71]
Endri Taka, Aman Arora, Kai-Chiang Wu, and Diana Marculescu. 2023. MaxEVA: Maximizing the Efficiency of Matrix Multiplication on Versal AI Engine. In 2023 International Conference on Field Programmable Technology (ICFPT) . 96–105. https://doi.org/10.1109/ICFPT59805.2023.00016
2023
-
[72]
Subhankar Pal, Jonathan Beaumont, Dong-Hyeon Park, Aporva Amarnath, Siy- ing Feng, Chaitali Chakrabarti, Hun-Seok Kim, David Blaauw, Trevor Mudge, and Ronald Dreslinski. 2018. OuterSPACE: An Outer Product Based Sparse Ma- trix Multiplication Accelerator. In 2018 IEEE Internati...
2018
-
[73]
Titopoulos, K
V. Titopoulos, K. Alexandridis, C. Peltekis, C. Nicopoulos, and G. Dimitrakopoulos
-
[74]
In 2019 IEEE 27th Annual International Symposium on Field- Programmable Custom Computing Machines (FCCM)
PIR-DSP: An FPGA DSP Block Architecture for Multi-precision Deep Neural Networks. In 2019 IEEE 27th Annual International Symposium on Field- Programmable Custom Computing Machines (FCCM) . 35–44. https://doi.org/10. 1109/FCCM.2019.00015
2019
-
[75]
Fengbin Tu, Yiqi Wang, Ling Liang, Yufei Ding, Leibo Liu, Shaojun Wei, Shouyi Yin, and Yuan Xie. 2023. SDP: Co-Designing Algorithm, Dataflow, and Archi- tecture for In-SRAM Sparse NN Acceleration. IEEE Transactions on Computer- Aided Design of Integrated Circuits and Systems 4...
2023
-
[76]
Vinay Vashishtha, Manoj Vangala, and Lawrence T. Clark. 2017. ASAP7 predictive design kit development and cell design technology co-optimization: Invited paper. In 2017 IEEE/ACM International Conference on Computer-Aided Design (ICCAD) . 992–998. https://doi.org/10.1109/ICCAD....
2017
-
[77]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, ...
2017
-
[78]
VTR. 2023. Verilog-to-Routing Documentation. https://docs.verilogtorouting. org/en/latest/vtr/benchmarks/
2023
-
[79]
Yu Emma Wang, Carole-Jean Wu, Xiaodong Wang, Kim Hazelwood, and David Brooks. 2021. Exploiting Parallelism Opportunities with Deep Learning Frame- works. ACM Trans. Archit. Code Optim. 18, 1, Article 9 (Dec 2021), 23 pages. https://doi.org/10.1145/3431388
2021 doi
-
[81]
Yannan Nellie Wu, Po-An Tsai, Saurav Muralidharan, Angshuman Parashar, Vivi- enne Sze, and Joel Emer. 2023. HighLight: Efficient and Flexible DNN Acceleration with Hierarchical Structured Sparsity. InProceedings of the 56th Annual IEEE/ACM International Symposium on Microarchi...
2023
-
[82]
Xiaoru Xie, Jun Lin, Zhongfeng Wang, and Jinghe Wei. 2021. An Efficient and Flexible Accelerator Design for Sparse Convolutional Neural Networks. IEEE Transactions on Circuits and Systems I: Regular Papers 68, 7 (2021), 2936–2949. https://doi.org/10.1109/TCSI.2021.3074300
2021
-
[83]
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. 2021. Training Data-efficient Image Trans- formers & Distillation Through Attention. In Proceedings of the 38th Interna- tional Conference on Machine Learning (Proceedings of...
2021
-
[84]
Sadegh Yazdanshenas and Vaughn Betz. 2019. COFFE 2: Automatic Modelling and Optimization of Complex and Heterogeneous FPGA Architectures. ACM Trans. Reconfigurable Technol. Syst. 12, 1, Article 3 (jan 2019), 27 pages. https: //doi.org/10.1145/3301298
2019 doi
-
[85]
Wenhua Ye, Xu Zhou, Joey Zhou, Cen Chen, and Kenli Li. 2023. Accelerating Attention Mechanism on FPGAs based on Efficient Reconfigurable Systolic Array. ACM Trans. Embed. Comput. Syst. 22, 6, Article 93 (nov 2023), 22 pages. https: //doi.org/10.1145/3549937
2023 doi
-
[86]
Zhihang Yuan, Chenhao Xue, Yiqi Chen, Qiang Wu, and Guangyu Sun. 2022. PTQ4ViT: Post-training Quantization for Vision Transformers with Twin Uniform Quantization. In Computer Vision – ECCV 2022: 17th European Conference, Tel A viv, Israel, October 23–27, 2022, Proceedings, Par...
2022 doi
-
[87]
Shijin Zhang, Zidong Du, Lei Zhang, Huiying Lan, Shaoli Liu, Ling Li, Qi Guo, Tianshi Chen, and Yunji Chen. 2016. Cambricon-X: An accelerator for sparse neural networks. In 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). 1–12. https://doi.org/10...
2016
-
[88]
Yuxin Zhang, Mingbao Lin, ZhiHang Lin, Yiting Luo, Ke Li, Fei Chao, Yongjian Wu, and Rongrong Ji. 2022. Learning Best Combination for Efficient N:M Spar- sity. In Advances in Neural Information Processing Systems , S. Koyejo, S. Mo- hamed, A. Agarwal, D. Belgrave, K. Cho, and ...
2022
-
[89]
Xuechao Wei, Cody Hao Yu, Peng Zhang, Youxiang Chen, Yuxin Wang, Han Hu, Yun Liang, and Jason Cong. 2017. Automated systolic array architecture synthesis for high throughput CNN inference on FPGAs. In 2017 54th ACM/EDAC/IEEE Design Automation Conference (DAC) . 1–6. https://do...
2017 doi
-
[90]
Aojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu, Zhijie Zhang, Kun Yuan, Wenxiu Sun, and Hongsheng Li. 2021. Learning N:M Fine-grained Structured Sparse Neural Networks From Scratch. arXiv:2102.04010 [cs.CV] https://arxiv.org/abs/ 2102.04010
2021 arXiv
-
[91]
Maohua Zhu, Tao Zhang, Zhenyu Gu, and Yuan Xie. 2019. Sparse Tensor Core: Algorithm and Hardware Co-Design for Vector-wise Sparse Neural Networks on Modern GPUs. In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture (Columbus, OH, USA) (MICRO ...
2019
-
[92]
Shuxin Yang, Chenchen Ding, Mingqiang Huang, Kai Li, Chenghao Li, Zikun Wei, Sixiao Huang, Jingyao Dong, Liuyang Zhang, and Hao Yu. 2024. LAMPS: A Layer-wised Mixed-Precision-and-Sparsity Accelerator for NAS-Optimized CNNs on FPGA. In 2024 IEEE 32nd Annual International Sympos...
2024
-
[98]
Zhekai Zhang, Hanrui Wang, Song Han, and William J. Dally. 2020. SpArch: Efficient Architecture for Sparse Matrix Multiplication. 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) (2020), 261–274. https://api.semanticscholar.org/CorpusID:211205022
2020
-
[1586]
https://doi.org/10.1109/TVLSI.2022.3197282
2022
-
[2016]
In 2016 IEEE International Symposium on Workload Characterization (IISWC)
Fathom: Reference workloads for modern deep learning methods. In 2016 IEEE International Symposium on Workload Characterization (IISWC). IEEE, 1–10
2016
-
[2021]
Sparsity in Deep Learning: Pruning and Growth for Efficient Inference and Training in Neural Networks. J. Mach. Learn. Res. 22, 1, Article 241 (jan 2021), 124 pages
2021
-
[2022]
ACM Trans
Tensor Slices: FPGA Building Blocks For The Deep Learning Era. ACM Trans. Reconfigurable Technol. Syst. 15, 4, Article 46 (aug 2022), 34 pages. https: //doi.org/10.1145/3529650
2022 doi
-
[2023]
In2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
VEGETA: Vertically-Integrated Extensions for Sparse/Dense GEMM Tile Acceleration on CPUs. In2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 259–272. https://doi.org/10.1109/HPCA56546. 2023.10071058
2023
-
[2024]
In 2024 Design, Automation & Test in Europe Con- ference & Exhibition (DATE)
IndexMAC: A Custom RISC-V Vector Instruction to Accelerate Structured- Sparse Matrix Multiplications. In 2024 Design, Automation & Test in Europe Con- ference & Exhibition (DATE) . 1–6. https://doi.org/10.23919/DATE58400.2024. 10546747
2024
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.