REVIEW 3 major objections 6 minor 121 references
Reconfigurable Stream Network Architecture
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read An ISA-level network abstraction treats the datapath as a circuit-switched network of stateful FUs; on VCK190 it cuts BERT latency 6.1x and lifts throughput 2.4x-3.2x over the prior art.
desk verdict Genuine prototype with reproducible numbers; the ISA-generality claim outruns the evidence because datapath construction is manual, but this is a solid, citable systems paper. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the reconfigurable stream network: a circuit-switched network whose nodes are stateful functional units (FUs), each with a uOP decoder, input/output ports, and kernel logic, and whose edges are latency-insensitive streams. Programming is triggering a path; a 32-bit packet with opcode, mask, window size, and reuse count encodes repeated uOP sequences, so one byte of instruction can drive up to 1.6 GFLOPs of computation. The load-bearing mechanism is partial path reprogramming: only FUs whose dataflow changes receive new instructions, so switching between mapping styles (one large GEMM, pipelined small GEMMs, fused non-MM ops) is cheap, and the DDR FU's explicit load/store interleaving keeps off-chip bandwidth busy during phase transitions.
What would settle it
Run an unseen transformer layer containing an operator not in the hand-built functional-unit set, or a model with data-dependent control flow such as dynamic masking, on the same RSN-XNN bitstream: if it executes without datapath modification and keeps close to the reported 6.1x latency gain, the abstraction generalizes; if it stalls or requires a rebuilt FU network, the compile-time path model is limited to statically known, hand-covered layers.
Extended reading notes
Core claim
RSN's central claim is that a circuit-switched network of stateful functional units, with latency-insensitive streams between them, is a sufficient and efficient ISA abstraction for DNN computation: each computation is launched by triggering a path, and software sees the compute and communication latency of every unit so it can overlap and fuse phases. The paper further claims that this abstraction is the first on FPGAs to combine dynamic layer fusion with fine-grained bandwidth mapping, and the prototype measurements back that up with a 6.1x latency reduction, 2.4x-3.2x throughput gains, and near-peak GEMM throughput of 6.78 TFLOPS on VCK190.
Load-bearing premise
The architecture presumes that DNN workloads are deterministic and have low control entropy, so compile-time path programming can replace runtime scheduling; it also relies on a designer-built union datapath, since automatic datapath generation is explicitly out of scope.
Editorial extensions
If this is right
- Overlay accelerators no longer need to serialize at layer granularity: the same bitstream can dynamically switch between a single fused GEMM and a pipeline of dependent small GEMMs, which is how attention layers avoid off-chip round-trips.
- Instruction-level control overhead can be made negligible, with one byte of instruction driving up to 1.6 GFLOPs, so the bottleneck becomes the datapath rather than the decoder.
- Heterogeneous units such as AIE arrays and FPGA fabric can be virtualized behind one FU interface, letting software treat the whole device as a network without knowing each node's implementation.
- Fine-grained load/store interleaving, not just double buffering, can hide phase-transition stalls and keep a single DDR channel nearly fully utilized.
- Energy efficiency gains relative to GPUs come from a 2.6x-2.8x reduction in off-chip DRAM traffic, achieved through on-chip reuse and pipelined execution that keeps intermediates on-chip.
Reading between the lines
- The paper leaves implicit that its manual union-datapath construction is a natural next target for automation; testing whether a compiler-generated datapath preserves the 6.1x and 2.4x-3.2x numbers on models the designers did not hand-tune would settle how general the abstraction really is.
- The explicit bandwidth-interleaving mechanism points toward an instruction-level memory-scheduling policy; comparing RSN-XNN against the same datapath with a hardware memory controller scheduler would isolate how much of the speedup comes from instruction-level interleaving rather than raw bandwidth.
- If the network abstraction is extended to other streaming-intensive domains such as scientific computing, the same path-triggering model should apply to kernels with statically known loop nests, but the FU set and union datapath would need to be derived automatically for those workloads.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces the Reconfigurable Stream Network (RSN), an ISA abstraction that models a DNN accelerator's datapath as a circuit-switched network of stateful functional units (FUs) with streaming edges, where programming a computation corresponds to triggering a path through the network. The authors implement a proof-of-concept design, RSN-XNN, on the AMD Versal VCK190 (combining AIEs and FPGA fabric), using a fixed set of FUs (MMEs, MemA/B/C, MeshA/B, DDR, LPDDR) controlled by a multi-level decoder. They report a measured latency of 17.98 ms for the first encoder of BERT-Large, a 6.1x latency reduction and 2.4x–3.2x throughput improvement over CHARM, an AIE GEMM throughput up to 6.78 TFLOPS (59% of the 8 TFLOPS peak), and a 2.1x FP32 energy-efficiency advantage over an A100 at the same 7 nm node. The artifact is open-source and includes the expected 17.98 ms result for reproducibility.
Significance. If the measured results are reproducible, the paper demonstrates a useful orchestration mechanism for heterogeneous AIE+FPGA systems: the stream-network abstraction achieves low instruction overhead (1.4 MB/s instruction rate, 1.6 GFLOPs per instruction byte) and enables dynamic switching between mapping types (single-layer, pipelined, fused) on a fixed bitstream. The main strengths are the concrete prototype and the open artifact: latency is measured on board, the expected value is stated in the artifact appendix, and the CHARM comparison can be checked against a public repository. The main limitation is that the datapath is hand-constructed for transformer/MLP workloads; no evidence is provided that the 'program by triggering a path' model extends to DNN layer types outside the pre-built FU set. This is a substantial caveat, but the paper can be revised to scope the contribution more precisely.
major comments (3)
- [§4.2, §4.5, §1] The central claim of a programmable ISA is not yet supported by the evidence. §4.2 describes datapath generation as a designer-led 'union datapath' construction, and §4.5 states that 'automatic generation of the datapath from arbitrary input code is beyond the scope of this paper.' All four evaluated models (BERT, ViT, NCF, MLP) consist only of the hand-built FU types (GEMM, softmax, GELU, LayerNorm, transpose), so the measured 6.1x and 2.4x–3.2x results do not demonstrate that a computation outside this set, such as a strided convolution or an LSTM with data-dependent gates, can be expressed by triggering paths. The abstract's claim that 'programming a computation corresponds to triggering a path' requires the path to exist in hardware; for an arbitrary DNN the path does not exist. Please either add a case study that requires a new FU type or datapath reconfiguration, provide an expressiveness analysis of the FU set for a defined DNN domain, or revise the central claims to describe RSN as a model-family-specific overlay with a fixed FU library.
- [§3.3] The paper states that 'comprehensive deadlock prevention is more complex and beyond the scope of this paper' and reports only that setting FIFO depths to six is deadlock-free in the implementation. Because RSN is presented as a general execution model in which arbitrary paths can be triggered, the correctness contract is incomplete: a programmer has no stated condition (e.g., acyclicity of the stream graph, or a buffer-sizing rule) to guarantee that a given uOP sequence does not deadlock. Please provide a formal deadlock-avoidance condition for the supported program class, or explicitly restrict the programming model to acyclic stream graphs and state this restriction in the abstraction definition in §3.1.
- [§5.7] The bandwidth sensitivity analysis simulates different off-chip bandwidths by changing the amount of data moved off-chip and padding the remainder on-chip. This alters the access pattern and does not reproduce the timing behavior of a real bandwidth change (DRAM bank conflicts, refresh, AXI arbitration), so the conclusion that 'the current use of bandwidth is already highly efficient' and the associated 78.6% utilization figure are not directly validated. Please either implement a real bandwidth change (e.g., by clock or interconnect configuration) or explicitly label this as a first-order estimate and discuss its limitations. This does not affect the directly measured latency/throughput results but weakens the paper's bandwidth-efficiency argument.
minor comments (6)
- [Table 7] The header row of Table 7 spells the design name as 'RSD-XNN'; this should be 'RSN-XNN'.
- [§4.1] In the second sentence of Section 4.1, 'for the LHS operands from MeshB FUs' should read 'for the RHS operands from MeshB FUs'.
- [§5.6, Table 10] The GPU comparison reports single-point latency and power numbers (vendor reports for T4/V100/A100 and one Colab session for L4) without variance or measurement repetitions; please report the number of runs and the observed spread, and clarify whether the VCK190 power in Table 10 is the BEAM measurement from Section 5 or the Vivado estimate from Table 4.
- [§4.2] The 'first-order formula-based calculation' used for model segmentation is mentioned but never specified; please include the formula (or pseudocode) so that the segmentation decision can be reproduced.
- [§3.3] The deadlock-free FIFO depth of six is reported for the tested implementation only; a short discussion of how the required FIFO depth scales with the number of FUs, stream depths, and instruction windows would help readers apply the result to other RSN configurations.
- [§1] The phrase 'dynamic layer fusion' in the contributions list and abstract is not formally defined; please clarify that it refers to runtime switching between mapping types such as pipelined, fused, and layer-by-layer execution.
Circularity Check
No significant circularity: the measured prototype results are self-contained and open to external checking, and the manual datapath construction is a scope limitation rather than a circular step.
full rationale
The paper's central numerical claims—6.1x latency reduction, 2.4x–3.2x throughput improvement, and 2.1x FP32 energy efficiency versus A100—are based on on-board measurements of the RSN-XNN prototype, not on parameters fitted to the outcomes being predicted. The evaluation section reports measured latencies, resource utilization, power estimates, and instruction counts, and the artifact appendix provides a bootable SD card image, reference outputs, and scripts for reproducing the reported 17.98 ms BERT-Large encoder latency. No equation in the paper reduces a predicted quantity to a fitted input or to the definition of another claimed result. The RSN abstraction is introduced conceptually and then instantiated in hardware; the flexibility and low instruction overhead are supported by direct measurements such as the 1.4 MB/s instruction processing rate and the 1.6 GFLOPs-per-byte compute-to-instruction ratio. The manual datapath generation process described in Section 4.2 and the statement in Section 4.5 that 'Exploring the automatic generation of the datapath from arbitrary input code is beyond the scope of this paper' are genuine scope limitations that constrain the generality of the claimed abstraction, but they do not make the reported measurements circular. The comparison baseline CHARM shares a co-author with the present paper, but it is a published prior implementation used as a benchmark, not an unverified self-citation that the argument depends on. Overall, the derivation chain is self-contained and empirically grounded, so the circularity score is 0.
Assumptions & free parameters
free parameters (3)
- MM tiling parameters (LHS 768x128, RHS 128x1024, OUT 768x1024) =
768x128, 128x1024, 768x1024
- AIE tile grouping (4x4x4, 64 tiles per MME) =
64 tiles per MME, 6 groups, 384 tiles
- Decoder FIFO depth between uOP and mOP decoders =
6
assumptions (5)
- domain assumption DNN execution is deterministic and has low control information entropy, so compile-time scheduling without runtime speculation is sufficient.
- domain assumption Latency-insensitive streaming with matching send/receive counts is a correct execution model; a receiver with fewer sends blocks indefinitely, a sender with excess sends blocks when the channel is full.
- domain assumption The roofline model adequately estimates latency for the four mapping types in Table 3.
- ad hoc to paper Simulating bandwidth variations by padding data on-chip preserves the latency model.
- domain assumption Vivado power estimates (over-estimated in absolute terms) provide valid ratios for energy comparisons.
invented entities (2)
-
RSN functional unit (FU) abstraction
-
RSN instruction packet with window size and reuse
Cite this review
Pith. "Pith review of Reconfigurable Stream Network Architecture." pith.science (2026). https://pith.science/paper/2PCMGJHU
@misc{pith2026241117966,
author = {Pith},
title = {Pith review of: Reconfigurable Stream Network Architecture},
year = {2026},
howpublished = {\url{https://pith.science/paper/2PCMGJHU}},
note = {Machine review of arXiv:2411.17966}
}
read the original abstract
As AI systems grow increasingly specialized and complex, managing hardware heterogeneity becomes a pressing challenge. How can we efficiently coordinate and synchronize heterogeneous hardware resources to achieve high utilization? How can we minimize the friction of transitioning between diverse computation phases, reducing costly stalls from initialization, pipeline setup, or drain? Our insight is that a network abstraction at the ISA level naturally unifies heterogeneous resource orchestration and phase transitions. This paper presents a Reconfigurable Stream Network Architecture (RSN), a novel ISA abstraction designed for the DNN domain. RSN models the datapath as a circuit-switched network with stateful functional units as nodes and data streaming on the edges. Programming a computation corresponds to triggering a path. Software is explicitly exposed to the compute and communication latency of each functional unit, enabling precise control over data movement for optimizations such as compute-communication overlap and layer fusion. As nodes in a network naturally differ, the RSN abstraction can efficiently virtualize heterogeneous hardware resources by separating control from the data plane, enabling low instruction-level intervention. We build a proof-of-concept design RSN-XNN on VCK190, a heterogeneous platform with FPGA fabric and AI engines. Compared to the SOTA solution on this platform, it reduces latency by 6.1x and improves throughput by 2.4x-3.2x. Compared to the T4 GPU with the same FP32 performance, it matches latency with only 18% of the memory bandwidth. Compared to the A100 GPU at the same 7nm process node, it achieves 2.1x higher energy efficiency in FP32.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Mohamed S. Abdelfattah, David Han, Andrew Bitar, Roberto DiCecco, Shane O’Connell, Nitika Shanker, Joseph Chu, Ian Prins, Joshua Fender, Andrew C. Ling, and Gordon R. Chiu. 2018. DLA: Compiler and FPGA Overlay for Neural Network Inference Acceleration. In 2018 28th International Conference on Field Reconfigurable Stream Network Architecture ISCA ’25, June...
arXiv 2018
- [2]
-
[3]
Dennis Abts, Garrin Kimmell, Andrew Ling, John Kim, Matt Boyd, Andrew Bitar, Sahil Parmar, Ibrahim Ahmed, Roberto DiCecco, David Han, John Thompson, Michael Bye, Jennifer Hwang, Jeremy Fowers, Peter Lillian, Ashwin Murthy, Elyas Mehtabuddin, Chetan Tekur, Thomas Sohmers, Kris Kang, Stephen Maresh, and Jonathan Ross. 2022. A software-defined tensor streami...
arXiv 2022
-
[4]
Sagheer Ahmad, Sridhar Subramanian, Vamsi Boppana, Shankar Lakka, Fu-Hing Ho, Tomai Knopp, Juanjo Noguera, Gaurav Singh, and Ralph Wittig. 2019. Xilinx First 7nm Device: Versal AI Core (VC1902). In2019 IEEE Hot Chips 31 Symposium (HCS). 1–28. https://doi.org/10.1109/HOTCHIPS.2019.8875639
arXiv 2019
- [5]
-
[6]
AMD. 2021. Versal AI Core Series VCK190 Evaluation Kit . Available: https: //www.xilinx.com/products/boards-and-kits/vck190.html, Accessed: 16-August- 2024
2021
-
[7]
AMD. 2023. DPU IP Details and System Integration. Available: https:// xilinx.github.io/Vitis-AI/3.5/html/docs/workflow-system-integration, Accessed: 5-November-2024
2023
-
[8]
AMD. 2023. Vitis Unified Software Platform 2023.2 . https://www.xilinx.com/ products/design-tools/vitis.html Software
2023
Show all 121 references
-
[9]
AMD. 2024. Versal ACAP Package Pinout Documentation: Mechan- ical - VC1802 and VC1902. https://docs.amd.com/r/en-US/am013- versal-pkg-pinout/VIVA1596-Mechanical-VC1802-and-VC1902 Avail- able: https://docs.amd.com/r/en-US/am013-versal-pkg-pinout/VIVA1596- Mechanical-VC1802-and-...
2024
-
[10]
AMD. 2024. Vivado Design Suite User Guide. Available: https://docs.amd.com/ r/en-US/ug906-vivado-design-analysis/Report-Power, Accessed: 5-November- 2024
2024
-
[11]
Giovanni Ansaloni, Paolo Bonzini, and Laura Pozzi. 2011. EGRA: A Coarse Grained Reconfigurable Architectural Template.IEEE Transactions on Very Large Scale Integration (VLSI) Systems 19, 6 (2011), 1062–1074. https://doi.org/10.1109/ TVLSI.2010.2044667
2011
-
[12]
Oguzhan Atak and Abdullah Atalar. 2013. BilRC: An Execution Triggered Coarse Grained Reconfigurable Architecture. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 21, 7 (2013), 1285–1298. https://doi.org/10.1109/ TVLSI.2012.2207748
2013
-
[13]
Yueyin Bai, Hao Zhou, Keqing Zhao, Hongji Wang, Jianli Chen, Jun Yu, and Kun Wang. 2023. FET-OPU: A Flexible and Efficient FPGA-Based Overlay Processor for Transformer Networks. In2023 IEEE/ACM International Conference on Computer Aided Design (ICCAD) . 1–9. https://doi.org/10...
2023
-
[14]
Thilini Kaushalya Bandara, Dhananjaya Wijerathne, Tulika Mitra, and Li-Shiuan Peh. 2022. REVAMP: a systematic framework for heterogeneous CGRA realiza- tion. In Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operatin...
2022
-
[15]
Suhail Basalama and Jason Cong. 2025. Stream-HLS: Towards Automatic Dataflow Acceleration. In Proceedings of the 2025 ACM/SIGDA International Symposium on Field Programmable Gate Arrays (Monterey, CA, USA) (FPGA ’25). Association for Computing Machinery, New York, NY, USA, 103...
2025
-
[16]
Suhail Basalama, Atefeh Sohrabizadeh, Jie Wang, and Jason Cong. 2022. A Versatile Systolic Array for Transposed and Dilated Convolution on FPGA. In 2022 IEEE 30th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM). 1–2. https://doi.org/10.110...
2022
-
[17]
Suhail Basalama, Atefeh Sohrabizadeh, Jie Wang, Licheng Guo, and Jason Cong
-
[18]
Suhail Basalama, Jie Wang, and Jason Cong. 2023. A Comprehensive Automated Exploration Framework for Systolic Array Designs. In 2023 60th ACM/IEEE Design Automation Conference (DAC). 1–6. https://doi.org/10.1109/DAC56929. 2023.10248016
2023
-
[19]
Riadh Ben Abdelhamid, Yoshiki Yamaguchi, and Taisuke Boku. 2019. MITRACA: A Next-Gen Heterogeneous Architecture. In 2019 IEEE 13th International Sym- posium on Embedded Multicore/Many-core Systems-on-Chip (MCSoC) . 304–311. https://doi.org/10.1109/MCSoC.2019.00050
2019
-
[20]
Maciej Besta, Marc Fischer, Tal Ben-Nun, Dimitri Stanojevic, Johannes De Fine Licht, and Torsten Hoefler. 2020. Substream-Centric Maximum Matchings on FPGA. ACM Trans. Reconfigurable Technol. Syst. 13, 2, Article 8 (April 2020), 33 pages. https://doi.org/10.1145/3377871
2020 doi
-
[21]
Andrew Boutros, Aman Arora, and Vaughn Betz. 2024. Field-Programmable Gate Array Architecture for Deep Learning: Survey & Future Directions. arXiv:2404.10076 [cs.AR]
2024
-
[22]
Hoe, Vaughn Betz, and Martin Langhammer
Andrew Boutros, Eriko Nurvitadhi, Rui Ma, Sergey Gribok, Zhipeng Zhao, James C. Hoe, Vaughn Betz, and Martin Langhammer. 2020. Beyond Peak Per- formance: Comparing the Real Performance of AI-Optimized FPGAs and GPUs. In 2020 International Conference on Field-Programmable Techn...
2020 arXiv
-
[23]
Jingwei Cai, Yuchen Wei, Zuotong Wu, Sen Peng, and Kaisheng Ma. 2023. core.cpp in SET-ISCA2023. Available: https://github.com/SET-Scheduling- Project/SET-ISCA2023/blob/master/src/core.cpp, Accessed: 5-November-2024
2023
-
[24]
Jingwei Cai, Yuchen Wei, Zuotong Wu, Sen Peng, and Kaisheng Ma. 2023. Inter- layer Scheduling Space Definition and Exploration for Tiled Accelerators. In Proceedings of the 50th Annual International Symposium on Computer Architecture (Orlando, FL, USA) (ISCA ’23). Association ...
2023
-
[25]
Jingwei Cai, Zuotong Wu, Sen Peng, Yuchen Wei, Zhanhong Tan, Guim- ing Shi, Mingyu Gao, and Kaisheng Ma. 2024. core.cpp in GEMINI- HPCA2024. Available: https://github.com/SET-Scheduling-Project/GEMINI- HPCA2024/blob/master/src/core.cpp, Accessed: 5-November-2024
2024
-
[26]
Jingwei Cai, Zuotong Wu, Sen Peng, Yuchen Wei, Zhanhong Tan, Guiming Shi, Mingyu Gao, and Kaisheng Ma. 2024. Gemini: Mapping and Architecture Co- exploration for Large-scale DNN Chiplet Accelerators. In2024 IEEE International Symposium on High-Performance Computer Architecture...
2024
-
[28]
Hongzheng Chen, Jiahao Zhang, Yixiao Du, Shaojie Xiang, Zichao Yue, Niansong Zhang, Yaohui Cai, and Zhiru Zhang. 2024. Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference. ACM Transactions on Reconfigurable Technology and Systems (...
2024 doi
-
[29]
Alexander Chin, Kuang Ping Niu, Matthew Walker, Shizhang Yin, Alexander Mertens, Jongeun Lee, and Jason H
S. Alexander Chin, Kuang Ping Niu, Matthew Walker, Shizhang Yin, Alexander Mertens, Jongeun Lee, and Jason H. Anderson. 2018. Architecture Exploration of Standard-Cell and FPGA-Overlay CGRAs Using the Open-Source CGRA-ME Framework. In Proceedings of the 2018 International Symp...
2018
-
[30]
Dally, Ujval J
William J. Dally, Ujval J. Kapasi, Brucek Khailany, Jung Ho Ahn, and Abhishek Das. 2004. Stream Processors: Progammability and Efficiency: Will this new kid on the block muscle out ASIC and DSP? Queue 2, 1 (March 2004), 52–62. https://doi.org/10.1145/984458.984486
2004
-
[31]
Xiaodong Deng, Shijie Wang, Tianyi Gao, Jing Liu, Longjun Liu, and Nanning Zheng. 2024. AMA: An Analytical Approach to Maximizing the Efficiency of Deep Learning on Versal AI Engine. In 2024 34th International Conference on Field-Programmable Logic and Applications (FPL) . 227...
2024
-
[32]
Reinhardt, Adrian M
Jeremy Fowers, Kalin Ovtcharov, Michael Papamichael, Todd Massengill, Ming Liu, Daniel Lo, Shlomi Alkalay, Michael Haselman, Logan Adams, Mahdi Ghandi, Stephen Heil, Prerak Patel, Adam Sapek, Gabriel Weisz, Lisa Woods, Sitaram Lanka, Steven K. Reinhardt, Adrian M. Caulfield, E...
2018
-
[33]
Mingyu Gao, Xuan Yang, Jing Pu, Mark Horowitz, and Christos Kozyrakis
-
[34]
Mingyu Gao, Xuan Yang, Jing Pu, Mark Horowitz, and Christos Kozyrakis. 2019. TANGRAM: Optimized Coarse-Grained Dataflow for Scalable NN Accelerators. In Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating S...
2019
-
[35]
Graham Gobieski, Ahmet Oguz Atli, Kenneth Mai, Brandon Lucia, and Nathan Beckmann. 2021. Snafu: An Ultra-Low-Power, Energy-Minimal CGRA- Generation Framework and Architecture. In 2021 ACM/IEEE 48th Annual In- ternational Symposium on Computer Architecture (ISCA) . 1027–1040. h...
2021
-
[36]
Venkatraman Govindaraju, Chen-Han Ho, Tony Nowatzki, Jatin Chhugani, Na- dathur Satish, Karthikeyan Sankaralingam, and Changkyu Kim. 2012. DySER: Unifying Functionality and Parallelism Specialization for Energy-Efficient Com- puting. IEEE Micro 32, 5 (2012), 38–51. https://doi...
2012 doi
-
[37]
Yijin Guan, Hao Liang, Ningyi Xu, Wenqiang Wang, Shaoshuai Shi, Xi Chen, Guangyu Sun, Wei Zhang, and Jason Cong. 2017. FP-DNN: An Automated Framework for Mapping Deep Neural Networks onto FPGAs with RTL-HLS Hybrid Templates. In 2017 IEEE 25th Annual International Symposium on ...
2017 doi
-
[38]
Licheng Guo, Yuze Chi, Jason Lau, Linghao Song, Xingyu Tian, Moazin Khatti, Weikang Qiao, Jie Wang, Ecenur Ustun, Zhenman Fang, Zhiru Zhang, and Jason Cong. 2023. TAPA: A Scalable Task-parallel Dataflow Programming Framework for Modern FPGAs with Co-optimization of HLS and Phy...
2023 doi
-
[39]
Licheng Guo, Pongstorn Maidee, Yun Zhou, Chris Lavin, Eddie Hung, Wuxi Li, Jason Lau, Weikang Qiao, Yuze Chi, Linghao Song, Yuanlong Xiao, Alireza Kaviani, Zhiru Zhang, and Jason Cong. 2023. RapidStream 2.0: Automated Parallel Implementation of Latency–Insensitive FPGA Designs...
2023 doi
-
[40]
Zibo Guo, Kai Liu, Wei Liu, Xiaoyao Sun, Chongyang Ding, and Shangrong Li. 2024. An Overlay Accelerator of DeepLab CNN for Spacecraft Image Seg- mentation on FPGA. Remote Sensing 16 (03 2024), 894. https://doi.org/10.3390/ rs16050894
2024
-
[41]
Mathew Hall and Vaughn Betz. 2020. HPIPE: Heterogeneous Layer-Pipelined and Sparse-Aware CNN Inference for FPGAs. In Proceedings of the 2020 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (Seaside, CA, USA) (FPGA ’20). Association for Computing Machinery, ...
2020
-
[43]
Zifan He, Anderson Truong, Yingqi Cao, and Jason Cong. 2025. InTAR: Inter- Task Auto-Reconfigurable Accelerator Design for High Data Volume Variation in DNNs. arXiv preprint arXiv:2502.08807 (2025)
2025 arXiv
-
[44]
Seongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee, Minsub Kim, Dong- soo Lee, and Joo-Young Kim. 2022. DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation. In 2022 IEEE Hot Chips 34 Symposium (HCS). 1–17. https://doi.org/10.1109/HCS55...
2022
-
[45]
Hugging Face
Inc. Hugging Face. 2023. BERT-LARGE model implementation in PyTorch. https://huggingface.co/transformers/model_doc/bert.html Software
2023
-
[46]
Suyeon Hur, Seongmin Na, Dongup Kwon, Joonsung Kim, Andrew Boutros, Eriko Nurvitadhi, and Jangwoo Kim. 2023. A Fast and Flexible FPGA-based Accelerator for Natural Language Processing Neural Networks. ACM Trans. Archit. Code Optim. 20, 1, Article 11 (Feb 2023), 24 pages. https...
2023
-
[47]
Intel. 2020. INT8 VS. FP32 Performance Comparision. Avail- able: https://intelkevinputnam.github.io/openvino-docs/pages/openvino_docs_ performance_int8_vs_fp32.html, Accessed: 5-November-2024
2020
-
[48]
Intel. 2019. Intel Deep Learning Boost. https://www.intel.com/content/ dam/www/central-libraries/us/en/documents/2022-09/xeon-accelerated-ai- product-brief.pdf
2019
-
[49]
Lana Josipović, Radhika Ghosal, and Paolo Ienne. 2018. Dynamically Sched- uled High-level Synthesis. In Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (Monterey, CALIFORNIA, USA) (FPGA ’18). Association for Computing Machinery, New ...
2018
-
[50]
Lana Josipovic, Andrea Guerrieri, and Paolo Ienne. 2021. Synthesizing General- Purpose Code Into Dynamically Scheduled Circuits. IEEE Circuits and Systems Magazine 21, 2 (2021), 97–118. https://doi.org/10.1109/MCAS.2021.3071631
2021
-
[52]
Kapasi, W.J
U.J. Kapasi, W.J. Dally, S. Rixner, J.D. Owens, and B. Khailany. 2002. The Imagine Stream Processor. In Proceedings. IEEE International Conference on Computer Design: VLSI in Computers and Processors . 282–288. https://doi.org/10.1109/ ICCD.2002.1106783
2002 arXiv
-
[53]
Kapasi, S
U.J. Kapasi, S. Rixner, W.J. Dally, B. Khailany, Jung Ho Ahn, P. Mattson, and J.D. Owens. 2003. Programmable stream processors. Computer 36, 8 (2003), 54–62. https://doi.org/10.1109/MC.2003.1220582
2003 arXiv
-
[54]
Manupa Karunaratne, Aditi Kulkarni Mohite, Tulika Mitra, and Li-Shiuan Peh
-
[55]
Hamza Khan, Asma Khan, Zainab Khan, Lun Bin Huang, Kun Wang, and Lei He
-
[56]
Konijnen- burg, Soojung Ryu, and Jeongwook Kim
Changmoo Kim, Moo-Kyoung Chung, Yeon-Gon Cho, Mario H. Konijnen- burg, Soojung Ryu, and Jeongwook Kim. 2012. ULP-SRP: Ultra low power Samsung Reconfigurable Processor for biomedical applications. 2012 Interna- tional Conference on Field-Programmable Technology (2012), 329–334....
2012
-
[57]
Mahapatra, and Kiyoung Choi
Yoonjin Kim, Rabi N. Mahapatra, and Kiyoung Choi. 2010. Design Space Ex- ploration for Efficient Resource Utilization in Coarse-Grained Reconfigurable Architecture. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 18, 10 (2010), 1471–1482. https://doi.org/10.11...
2010
-
[58]
Kalhan Koul, Jackson Melchert, Kavya Sreedhar, Leonard Truong, Gedeon Nyen- gele, Keyi Zhang, Qiaoyi Liu, Jeff Setter, Po-Han Chen, Yuchen Mei, Maxwell Strange, Ross Daly, Caleb Donovick, Alex Carsello, Taeyoung Kong, Kathleen Feng, Dillon Huff, Ankita Nayak, Rajsekhar Setalur...
2023
-
[59]
Ronny Krashinsky, Christopher Batten, Mark Hampton, Steve Gerding, Brian Pharris, Jared Casper, and Krste Asanovic. 2004. The Vector-Thread Architec- ture. In Proceedings of the 31st Annual International Symposium on Computer Architecture (München, Germany) (ISCA ’04). IEEE Co...
2004
-
[60]
Hyoukjun Kwon, Prasanth Chatarasi, Vivek Sarkar, Tushar Krishna, Michael Pellauer, and Angshuman Parashar. 2020. AHWAccelerator.hpp in Mae- stro. Available: https://github.com/maestro-project/maestro/blob/master/cost- model/include/abstract-hardware-model/AHW_Accelerator.hpp, ...
2020
-
[61]
Hyoukjun Kwon, Prasanth Chatarasi, Vivek Sarkar, Tushar Krishna, Michael Pellauer, and Angshuman Parashar. 2020. MAESTRO: A Data-Centric Approach to Understand Reuse, Performance, and Hardware Cost of DNN Mappings.IEEE Micro 40, 3 (2020), 20–29
2020
-
[62]
Martin Langhammer, Eriko Nurvitadhi, Bogdan Pasca, and Sergey Gribok. 2021. Stratix 10 NX Architecture and Applications. In The 2021 ACM/SIGDA Inter- national Symposium on Field-Programmable Gate Arrays (Virtual Event, USA) (FPGA ’21). Association for Computing Machinery, New ...
2021
-
[63]
Jason Lau, Yuanlong Xiao, Yutong Xie, Yuze Chi, Linghao Song, Shaojie Xiang, Michael Lo, Zhiru Zhang, Jason Cong, and Licheng Guo. 2025. RapidStream IR: Infrastructure for FPGA High-Level Physical Synthesis. InProceedings of the 43rd IEEE/ACM International Conference on Comput...
2025
-
[64]
Yunsup Lee, Rimas Avizienis, Alex Bishara, Richard Xia, Derek Lockhart, Christo- pher Batten, and Krste Asanović. 2011. Exploring the tradeoffs between pro- grammability and efficiency in data-parallel accelerators. In 2011 38th Annual International Symposium on Computer Archi...
2011
-
[65]
Bingbing Li, Santosh Pandey, Haowen Fang, Yanjun Lyv, Ji Li, Jieyang Chen, Mimi Xie, Lipeng Wan, Hang Liu, and Caiwen Ding. 2020. FTRANS: energy- efficient acceleration of transformers using FPGA. In Proceedings of the ACM/IEEE International Symposium on Low Power Electronics ...
2020
-
[66]
Jingyuan Li, Yunhui Qiu, Guowei Zhu, Qilong Zhu, Wenbo Yin, and Lingli Wang
-
[67]
Sihao Liu, Jian Weng, Dylan Kupsh, Atefeh Sohrabizadeh, Zhengrong Wang, Licheng Guo, Jiuyang Liu, Maxim Zhulin, Rishabh Mani, Lucheng Zhang, Ja- son Cong, and Tony Nowatzki. 2022. OverGen: Improving FPGA Usability Reconfigurable Stream Network Architecture ISCA ’25, June 21–25...
2022
-
[68]
Rui Ma, Jia-Ching Hsu, Tian Tan, Eriko Nurvitadhi, David Sheffield, Rob Pelt, Martin Langhammer, Jaewoong Sim, Aravind Dasu, and Derek Chiou. 2019. Specializing FGPU for Persistent Deep Learning. In 2019 29th International Conference on Field Programmable Logic and Application...
2019
-
[69]
Bingfeng Mei, Serge Vernalde, Diederik Verkest, Hugo De Man, and Rudy Lauw- ereins. 2003. ADRES: An Architecture with Tightly Coupled VLIW Processor and Coarse-Grained Reconfigurable Matrix. In International Conference on Field- Programmable Logic and Applications. https://api...
2003
-
[70]
Thierry Moreau, Tianqi Chen, Luis Vega, Jared Roesch, Eddie Yan, Lianmin Zheng, Josh Fromm, Ziheng Jiang, Luis Ceze, Carlos Guestrin, and Arvind Kr- ishnamurthy. 2019. A Hardware–Software Blueprint for Flexible Deep Learning Specialization. IEEE Micro 39, 5 (2019), 8–16. https...
2019 doi
-
[71]
In 2023 IEEE International Symposium on Circuits and Systems (ISCAS)
THRAM: A Template-based Heterogeneous CGRA Modeling Framework Supporting Fast DSE. In 2023 IEEE International Symposium on Circuits and Systems (ISCAS). 1–5. https://doi.org/10.1109/ISCAS46773.2023.10182204
2023
-
[72]
Chris Nicol. 2017. A Coarse Grain Reconfigurable Array ( CGRA ) for Statically Scheduled Data Flow Computing. https://api.semanticscholar.org/CorpusID: 199394670
2017
-
[74]
NVIDIA. 2017. NVIDIA Tesla V100 GPU Architecture. (2017). Available: https://images.nvidia.com/content/volta-architecture/pdf/volta-architecture- whitepaper.pdf, Accessed: 21-November-2024
2017
-
[75]
NVIDIA. 2018. NVIDIA T4 Tensor Core GPU. https://resources.nvidia.com/en- us-gpu-resources/t4-tensor-core-datas?lx=CPwSfP Available: https://resources. nvidia.com/en-us-gpu-resources/t4-tensor-core-datas?lx=CPwSfP, Accessed: 21-November-2024
2018
-
[76]
Rene Mueller, Jens Teubner, and Gustavo Alonso. 2009. Streams on wires: a query compiler for FPGAs. Proc. VLDB Endow. 2, 1 (Aug. 2009), 229–240. https://doi.org/10.14778/1687627.1687654
2009
-
[77]
NVIDIA. 2024. DeepLearningExamples: BERT Language Modeling with Tensor- Flow 2. Available: https://github.com/NVIDIA/DeepLearningExamples/tree/ master/TensorFlow2/LanguageModeling/BERT, Accessed: 16-August-2024
2024
-
[78]
NVIDIA. 2024. NVIDIA H100 TENSOR CORE GPU. Available: https://resources. nvidia.com/en-us-tensor-core/nvidia-tensor-core-gpu-datasheet, Accessed: 20- Feburary-2025
2024
-
[79]
NVIDIA. 2024. NVIDIA L4 TENSOR CORE GPU. Available: https:// resources.nvidia.com/en-us-data-center-overview/l4-gpu-datasheet, Accessed: 21-November-2024
2024
-
[80]
Nvidia. 2025. Nsight Compute CLI. Available: https://docs.nvidia.com/nsight- compute/NsightComputeCli/index.html, Accessed: 14-February-2025
2025
-
[81]
NVIDIA. 2021. NVIDIA A100 TENSOR CORE GPU. Avail- able: https://www.nvidia.com/content/dam/en-zz/Solutions/Data- Center/a100/pdf/nvidia-a100-datasheet-us-nvidia-1758950-r4-web.pdf, Accessed: 21-November-2024
2021
-
[82]
Angshuman Parashar, Michael Pellauer, Michael Adler, Bushra Ahsan, Neal Crago, Daniel Lustig, Vladimir Pavlov, Antonia Zhai, Mohit Gambhir, Aamer Jaleel, Randy Allmon, Rachid Rayess, Stephen Maresh, and Joel Emer. 2013. Triggered instructions: a control paradigm for spatially-...
2013
-
[83]
Artur Podobas, Kentaro Sano, and Satoshi Matsuoka. 2020. A Survey on Coarse- Grained Reconfigurable Architectures From a Performance Perspective. IEEE Access 8 (2020), 146719–146743. https://doi.org/10.1109/ACCESS.2020.3012084
2020
-
[84]
Shah, Zhengyu Chen, Kaizhao Liang, Swayambhoo Jain, Urmish Thakker, Dawei Huang, Sumti Jairath, Kevin J
Raghu Prabhakar, Ram Sivaramakrishnan, Darshan Gandhi, Yun Du, Mingran Wang, Xiangyu Song, Kejie Zhang, Tianren Gao, Angela Wang, Xiaoyan Li, Yongning Sheng, Joshua Brot, Denis Sokolov, Apurv Vivek, Calvin Leung, Arjun Sabnis, Jiayu Bai, Tuowen Zhao, Mark Gottscho, David Jacks...
-
[86]
Oliveira, Michael Canesche, Lucas Reis, José Augusto Miranda Nacif, and Ricardo S
Westerley C. Oliveira, Michael Canesche, Lucas Reis, José Augusto Miranda Nacif, and Ricardo S. Ferreira. 2022. Heterogeneous reconfigurable architectures for machine learning dataflows. Concurrency and Computation: Practice and Experience 35 (2022). https://api.semanticschola...
2022
-
[87]
Karthikeyan Sankaralingam, Tony Nowatzki, Vinay Gangadhar, Preyas Shah, Michael Davies, William Galliher, Ziliang Guo, Jitu Khare, Deepak Vijay, Poly Palamuttam, Maghawan Punde, Alex Tan, Vijay Thiruvengadam, Rongyi Wang, and Shunmiao Xu. 2022. The Mozart reuse exposed dataflo...
2022
-
[88]
Colin Schmidt. 2021. Extending Temporal-Vector Microarchitectures for Two- Dimensional Computations. Ph. D. Dissertation. University of California, Berke- ley, USA. https://www.escholarship.org/uc/item/2mr167rk
2021
-
[89]
Yongming Shen, Michael Ferdman, and Peter Milder. 2017. Maximizing CNN Accelerator Efficiency Through Resource Partitioning.SIGARCH Comput. Archit. News 45, 2 (2017), 535–547. https://doi.org/10.1145/3140659.3080221
2017
-
[90]
Singh, Ming-Hau Lee, Guangming Lu, F.J
H. Singh, Ming-Hau Lee, Guangming Lu, F.J. Kurdahi, N. Bagherzadeh, and E.M. Chaves Filho. 2000. MorphoSys: an integrated reconfigurable system for data-parallel and computation-intensive applications. IEEE Trans. Comput. 49, 5 (2000), 465–481. https://doi.org/10.1109/12.859540
2000 doi
-
[91]
James E. Smith. 1982. Decoupled access/execute computer architectures. In Proceedings of the 9th Annual Symposium on Computer Architecture (Austin, Texas, USA) (ISCA ’82). IEEE Computer Society Press, Washington, DC, USA, 112–119
1982
-
[92]
Richard M. Russell. 1978. The CRAY-1 computer system. Commun. ACM 21, 1 (Jan. 1978), 63–72. https://doi.org/10.1145/359327.359336
1978
-
[93]
Atefeh Sohrabizadeh, Yuze Chi, and Jason Cong. 2022. StreamGCN: Accelerating Graph Convolutional Networks with Streaming Processing. In2022 IEEE Custom Integrated Circuits Conference (CICC) . 1–8. https://doi.org/10.1109/CICC53496. 2022.9772832
2022
-
[94]
Linghao Song, Yuze Chi, Atefeh Sohrabizadeh, Young-kyu Choi, Jason Lau, and Jason Cong. 2022. Sextans: A Streaming Accelerator for General-Purpose Sparse- Matrix Dense-Matrix Multiplication. In Proceedings of the 2022 ACM/SIGDA International Symposium on Field-Programmable Gat...
2022
-
[95]
Nambiar, Anh Tuan Do, Thilini Kaushalya Bandara, Aditi Kulkarni Mohite, and Bo Wang
Lingzhi Su, Wang Ling Goh, Jingjing Lan, Vishnu P. Nambiar, Anh Tuan Do, Thilini Kaushalya Bandara, Aditi Kulkarni Mohite, and Bo Wang. 2022. An Energy-Efficient Processing Element Design for Coarse-Grained Reconfigurable Architecture on FPGA. In 2022 11th International Confer...
2022
-
[96]
Endri Taka, Aman Arora, Kai Chiang Wu, and Diana Marculescu. 2023. MaxEVA: Maximizing the Efficiency of Matrix Multiplication on Versal AI Engine. In2023 International Conference on Field Programmable Technology (ICFPT). IEEE, 96–105. https://doi.org/10.1109/ICFPT59805.2023.00016
2023
-
[97]
Barker, and Antonino Tumeo
Cheng Tan, Chenhao Xie, Ang Li, Kevin J. Barker, and Antonino Tumeo. 2021. AURORA: Automated Refinement of Coarse-Grained Reconfigurable Accelera- tors. In 2021 Design, Automation & Test in Europe Conference & Exhibition (DATE). 1388–1393. https://doi.org/10.23919/DATE51398.20...
2021
-
[98]
Hayden Kwok-Hay So and Cheng Liu. 2016. FPGA Overlays. Springer Interna- tional Publishing, Cham, 285–305. https://doi.org/10.1007/978-3-319-26408- 0_16
2016 doi
-
[99]
Dani Voitsechov and Yoav Etsion. 2014. Single-graph multiple flows: energy efficient design alternative for GPGPUs. In Proceeding of the 41st Annual Inter- national Symposium on Computer Architecuture (Minneapolis, Minnesota, USA) (ISCA ’14). IEEE Press, 205–216
2014
-
[100]
Dani Voitsechov and Yoav Etsion. 2015. Control flow coalescing on a hybrid dataflow/von Neumann GPGPU. In 2015 48th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . 216–227. https://doi.org/10.1145/ 2830772.2830817
2015
-
[101]
Dani Voitsechov, Oron Port, and Yoav Etsion. 2018. Inter-Thread Communication in Multithreaded, Reconfigurable Coarse-Grain Arrays. In 2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). 42–54. https: //doi.org/10.1109/MICRO.2018.00013
2018
-
[102]
Bo Wang, Manupa Karunarathne, Aditi Kulkarni, Tulika Mitra, and Li-Shiuan Peh. 2019. HyCUBE: A 0.9V 26.4 MOPS/mW, 290 pJ/op, Power Efficient Accel- erator for IoT Applications. In 2019 IEEE Asian Solid-State Circuits Conference (A-SSCC). 133–136. https://doi.org/10.1109/A-SSCC...
2019
-
[103]
Teng Wang, Lei Gong, Chao Wang, Yang Yang, Yingxue Gao, Xuehai Zhou, and Huaping Chen. 2022. ViA: A Novel Vision-Transformer Accelerator Based on FPGA. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 41, 11 (2022), 4088–4099. https://doi.org/10.11...
2022
-
[104]
Swagath Venkataramani, Ashish Ranjan, Subarno Banerjee, Dipankar Das, Sasikanth Avancha, Ashok Jagannathan, Ajaya Durg, Dheemanth Nagaraj, Bharat Kaul, Pradeep Dubey, and Anand Raghunathan. 2017. SCALEDEEP: A scalable compute architecture for learning and evaluating deep netwo...
2017
-
[105]
Xuechao Wei, Yun Liang, Tao Wang, Songwu Lu, and Jason Cong. 2017. Through- put optimization for streaming applications on CPU-FPGA heterogeneous sys- tems. In 2017 22nd Asia and South Pacific Design Automation Conference (ASP- DAC). 488–493. https://doi.org/10.1109/ASPDAC.201...
2017
-
[106]
Xuechao Wei, Cody Hao Yu, Peng Zhang, Youxiang Chen, Yuxin Wang, Han Hu, Yun Liang, and Jason Cong. 2017. Automated Systolic Array Architecture Synthesis for High Throughput CNN Inference on FPGAs. In Proceedings of the 54th Annual Design Automation Conference 2017 (Austin, TX...
2017
-
[107]
Jian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang, Preyas Shah, and Tony Nowatzki. 2020. DSAGEN: Synthesizing Programmable Spatial Accelerators. In 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). 268–281. https://doi.org/10.1109/ISCA45697.2020.00032
2020
-
[108]
Jian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu, and Tony Nowatzki. 2020. A Hybrid Systolic-Dataflow Architecture for Inductive Matrix Algorithms. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). 703–716. https://doi.org/10.1109/HPCA475...
2020
-
[109]
Emer, and Daniel Sanchez
Yifan Yang, Joel S. Emer, and Daniel Sanchez. 2023. ISOSceles: Accelerating Sparse CNNs through Inter-Layer Pipelining. In 2023 IEEE International Sym- posium on High-Performance Computer Architecture (HPCA) . 598–610. https: //doi.org/10.1109/HPCA56546.2023.10071080
2023
-
[111]
Yunxuan Yu, Chen Wu, Tiandong Zhao, Kun Wang, and Lei He. 2020. OPU: An FPGA-Based Overlay Processor for Convolutional Neural Networks. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 28, 1 (2020), 35–47. https://doi.org/10.1109/TVLSI.2019.2939726
2020
-
[113]
Zhang, H
B. Zhang, H. Zeng, and V. K. Prasanna. 2023. GraphAGILE: An FPGA-Based Overlay Accelerator for Low-Latency GNN Inference. IEEE Transactions on Parallel & Distributed Systems 34, 09 (2023), 2580–2597. https://doi.org/10.1109/ TPDS.2023.3287883
2023
-
[114]
Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong. 2015. Optimizing FPGA-based Accelerator Design for Deep Convo- lutional Neural Networks. In Proceedings of the 2015 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (Monterey, Califo...
2015
-
[115]
Chen Zhang, Guangyu Sun, Zhenman Fang, Peipei Zhou, Peichen Pan, and Jason Cong. 2019. Caffeine: Toward Uniformed Representation and Acceleration for Deep Convolutional Neural Networks. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 38, 11 (2019)...
2019
-
[116]
Hanchen Ye, Xiaofan Zhang, Zhize Huang, Gengsheng Chen, and Deming Chen
-
[117]
Xiaofan Zhang, Hanchen Ye, Junsong Wang, Yonghua Lin, Jinjun Xiong, Wen- mei Hwu, and Deming Chen. 2020. DNNExplorer: a framework for modeling and exploring a novel paradigm of FPGA-based DNN accelerator. InProceedings of the 39th International Conference on Computer-Aided Des...
2020
-
[118]
Shixuan Zheng, Xianjue Zhang, Leibo Liu, Shaojun Wei, and Shouyi Yin. 2022. Atomic Dataflow based Graph-Level Workload Orchestration for Scalable DNN Accelerators. In 2022 IEEE International Symposium on High-Performance Com- puter Architecture (HPCA). 475–489. https://doi.org...
2022
-
[119]
Jinming Zhuang, Jason Lau, Hanchen Ye, Zhuoping Yang, Yubo Du, Jack Lo, Kristof Denolf, Stephen Neuendorffer, Alex Jones, Jingtong Hu, Deming Chen, Jason Cong, and Peipei Zhou. 2023. CHARM: Composing Heterogeneous Accel- erators for Matrix Multiply on Versal ACAP Architecture....
2023
-
[120]
Jones, Jingtong Hu, Yiyu Shi, and Peipei Zhou
Jinming Zhuang, Zhuoping Yang, Shixin Ji, Heng Huang, Alex K. Jones, Jingtong Hu, Yiyu Shi, and Peipei Zhou. 2024. SSR: Spatial Sequential Hybrid Architecture for Latency Throughput Tradeoff in Transformer Acceleration. In Proceedings of the 2024 ACM/SIGDA International Sympos...
2024
-
[121]
Zhipeng Zhao, Joseph Melber, Siddharth Sahay, Shashank Obla, Eriko Nurvi- tadhi, James C. Hoe. 2023. Exploiting the Common Case When Accelerating Input-Dependent Stream Processing by FPGA. In IEEE Transactions on Comput- ers (TC ’23). IEEE. https://doi.org/10.1109/TC.2022.3200576
2023
-
[123]
Xiaofan Zhang, Junsong Wang, Chao Zhu, Yonghua Lin, Jinjun Xiong, Wen-mei Hwu, and Deming Chen. 2018. DNNBuilder: an Automated Tool for Building High-Performance DNN Hardware Accelerators for FPGAs. In 2018 IEEE/ACM International Conference on Computer-Aided Design (ICCAD) . 1...
2018
-
[130]
No Optimize
Zhipeng Zhao, Nirav Atre, Hugo Sadok, Siddharth Sahay, Shashank Obla, James C. Hoe, Justine Sherry. 2022. Pigasus 2.0: making the pigasus IDS robust to attacks and different workloads. In In Proceedings of the SIGCOMM ’22 Poster and Demo Sessions (SIGCOMM ’22) . Association fo...
2022
-
[2019]
Available: https://github.com/stanford- mast/nn_dataflow/tree/master/nn_dataflow/core, Accessed: 5-November-2024
core directory in Tangram. Available: https://github.com/stanford- mast/nn_dataflow/tree/master/nn_dataflow/core, Accessed: 5-November-2024
2024
-
[2020]
In Proceedings of the 57th ACM/EDAC/IEEE Design Automation Conference (Virtual Event, USA) (DAC ’20)
HybridDNN: a framework for high-performance hybrid DNN accelerator design and implementation. In Proceedings of the 57th ACM/EDAC/IEEE Design Automation Conference (Virtual Event, USA) (DAC ’20). IEEE Press, Article 129, 6 pages
-
[2021]
In The 2021 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (Virtual Event, USA) (FPGA ’21)
NPE: An FPGA-based Overlay Processor for Natural Language Processing. In The 2021 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (Virtual Event, USA) (FPGA ’21). Association for Computing Machinery, New York, NY, USA, 227. https://doi.org/10.1145/3431920.3439477
2021
-
[2023]
ACM Trans
FlexCNN: An End-to-end Framework for Composing CNN Accelerators on FPGA. ACM Trans. Reconfigurable Technol. Syst. 16, 2, Article 23 (March 2023), 32 pages. https://doi.org/10.1145/3570928
2023 doi
-
[2024]
In 2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO)
SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts . In 2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE Computer Society, Los Alamitos, CA, USA, 1353–1366. https://doi.org/10.1109/MICRO61859.2024.00100
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.