Pith. sign in

REVIEW 3 major objections 6 minor 121 references

Reconfigurable Stream Network Architecture

T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read An ISA-level network abstraction treats the datapath as a circuit-switched network of stateful FUs; on VCK190 it cuts BERT latency 6.1x and lifts throughput 2.4x-3.2x over the prior art.

desk verdict Genuine prototype with reproducible numbers; the ISA-generality claim outruns the evidence because datapath construction is manual, but this is a solid, citable systems paper. read the letter →

arxiv 2411.17966 v3 pith:2PCMGJHU submitted 2024-11-27 cs.AR

classification cs.AR
keywords reconfigurablestreamnetworkFPGAoverlayAIengineVersalVCK190dataflowarchitecturelayerfusionbandwidthmappingDNNinference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the right way to program heterogeneous DNN hardware is to expose the datapath as a reconfigurable network: stateful functional units as nodes, streams as edges, and a computation as a triggered path through the network. The authors claim this abstraction unifies resource orchestration across FPGAs and AI engines, removes layer-granular serialization in overlays, and makes phase transitions nearly stall-free by letting software interleave loads, compute, and stores at fine grain. On the Versal VCK190 platform, their RSN-XNN prototype is reported to cut BERT encoder latency 6.1x and improve throughput 2.4x-3.2x against the state of the art, and to reach 2.1x better FP32 energy efficiency than an A100 at the same process node. A sympathetic reader would care because the result suggests that low-entropy, deterministic DNN control can be encoded at compile time as sparse path-triggering instructions, avoiding both bitstream reconfiguration and heavyweight runtime scheduling.

What carries the argument

The central object is the reconfigurable stream network: a circuit-switched network whose nodes are stateful functional units (FUs), each with a uOP decoder, input/output ports, and kernel logic, and whose edges are latency-insensitive streams. Programming is triggering a path; a 32-bit packet with opcode, mask, window size, and reuse count encodes repeated uOP sequences, so one byte of instruction can drive up to 1.6 GFLOPs of computation. The load-bearing mechanism is partial path reprogramming: only FUs whose dataflow changes receive new instructions, so switching between mapping styles (one large GEMM, pipelined small GEMMs, fused non-MM ops) is cheap, and the DDR FU's explicit load/store interleaving keeps off-chip bandwidth busy during phase transitions.

What would settle it

Run an unseen transformer layer containing an operator not in the hand-built functional-unit set, or a model with data-dependent control flow such as dynamic masking, on the same RSN-XNN bitstream: if it executes without datapath modification and keeps close to the reported 6.1x latency gain, the abstraction generalizes; if it stalls or requires a rebuilt FU network, the compile-time path model is limited to statically known, hand-covered layers.

Watch

Extended reading notes

Core claim

RSN's central claim is that a circuit-switched network of stateful functional units, with latency-insensitive streams between them, is a sufficient and efficient ISA abstraction for DNN computation: each computation is launched by triggering a path, and software sees the compute and communication latency of every unit so it can overlap and fuse phases. The paper further claims that this abstraction is the first on FPGAs to combine dynamic layer fusion with fine-grained bandwidth mapping, and the prototype measurements back that up with a 6.1x latency reduction, 2.4x-3.2x throughput gains, and near-peak GEMM throughput of 6.78 TFLOPS on VCK190.

Load-bearing premise

The architecture presumes that DNN workloads are deterministic and have low control entropy, so compile-time path programming can replace runtime scheduling; it also relies on a designer-built union datapath, since automatic datapath generation is explicitly out of scope.

Editorial extensions

If this is right

  • Overlay accelerators no longer need to serialize at layer granularity: the same bitstream can dynamically switch between a single fused GEMM and a pipeline of dependent small GEMMs, which is how attention layers avoid off-chip round-trips.
  • Instruction-level control overhead can be made negligible, with one byte of instruction driving up to 1.6 GFLOPs, so the bottleneck becomes the datapath rather than the decoder.
  • Heterogeneous units such as AIE arrays and FPGA fabric can be virtualized behind one FU interface, letting software treat the whole device as a network without knowing each node's implementation.
  • Fine-grained load/store interleaving, not just double buffering, can hide phase-transition stalls and keep a single DDR channel nearly fully utilized.
  • Energy efficiency gains relative to GPUs come from a 2.6x-2.8x reduction in off-chip DRAM traffic, achieved through on-chip reuse and pipelined execution that keeps intermediates on-chip.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves implicit that its manual union-datapath construction is a natural next target for automation; testing whether a compiler-generated datapath preserves the 6.1x and 2.4x-3.2x numbers on models the designers did not hand-tune would settle how general the abstraction really is.
  • The explicit bandwidth-interleaving mechanism points toward an instruction-level memory-scheduling policy; comparing RSN-XNN against the same datapath with a hardware memory controller scheduler would isolate how much of the speedup comes from instruction-level interleaving rather than raw bandwidth.
  • If the network abstraction is extended to other streaming-intensive domains such as scientific computing, the same path-triggering model should apply to kernels with statically known loop nests, but the FU set and union datapath would need to be derived automatically for those workloads.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The manuscript introduces the Reconfigurable Stream Network (RSN), an ISA abstraction that models a DNN accelerator's datapath as a circuit-switched network of stateful functional units (FUs) with streaming edges, where programming a computation corresponds to triggering a path through the network. The authors implement a proof-of-concept design, RSN-XNN, on the AMD Versal VCK190 (combining AIEs and FPGA fabric), using a fixed set of FUs (MMEs, MemA/B/C, MeshA/B, DDR, LPDDR) controlled by a multi-level decoder. They report a measured latency of 17.98 ms for the first encoder of BERT-Large, a 6.1x latency reduction and 2.4x–3.2x throughput improvement over CHARM, an AIE GEMM throughput up to 6.78 TFLOPS (59% of the 8 TFLOPS peak), and a 2.1x FP32 energy-efficiency advantage over an A100 at the same 7 nm node. The artifact is open-source and includes the expected 17.98 ms result for reproducibility.

Significance. If the measured results are reproducible, the paper demonstrates a useful orchestration mechanism for heterogeneous AIE+FPGA systems: the stream-network abstraction achieves low instruction overhead (1.4 MB/s instruction rate, 1.6 GFLOPs per instruction byte) and enables dynamic switching between mapping types (single-layer, pipelined, fused) on a fixed bitstream. The main strengths are the concrete prototype and the open artifact: latency is measured on board, the expected value is stated in the artifact appendix, and the CHARM comparison can be checked against a public repository. The main limitation is that the datapath is hand-constructed for transformer/MLP workloads; no evidence is provided that the 'program by triggering a path' model extends to DNN layer types outside the pre-built FU set. This is a substantial caveat, but the paper can be revised to scope the contribution more precisely.

major comments (3)
  1. [§4.2, §4.5, §1] The central claim of a programmable ISA is not yet supported by the evidence. §4.2 describes datapath generation as a designer-led 'union datapath' construction, and §4.5 states that 'automatic generation of the datapath from arbitrary input code is beyond the scope of this paper.' All four evaluated models (BERT, ViT, NCF, MLP) consist only of the hand-built FU types (GEMM, softmax, GELU, LayerNorm, transpose), so the measured 6.1x and 2.4x–3.2x results do not demonstrate that a computation outside this set, such as a strided convolution or an LSTM with data-dependent gates, can be expressed by triggering paths. The abstract's claim that 'programming a computation corresponds to triggering a path' requires the path to exist in hardware; for an arbitrary DNN the path does not exist. Please either add a case study that requires a new FU type or datapath reconfiguration, provide an expressiveness analysis of the FU set for a defined DNN domain, or revise the central claims to describe RSN as a model-family-specific overlay with a fixed FU library.
  2. [§3.3] The paper states that 'comprehensive deadlock prevention is more complex and beyond the scope of this paper' and reports only that setting FIFO depths to six is deadlock-free in the implementation. Because RSN is presented as a general execution model in which arbitrary paths can be triggered, the correctness contract is incomplete: a programmer has no stated condition (e.g., acyclicity of the stream graph, or a buffer-sizing rule) to guarantee that a given uOP sequence does not deadlock. Please provide a formal deadlock-avoidance condition for the supported program class, or explicitly restrict the programming model to acyclic stream graphs and state this restriction in the abstraction definition in §3.1.
  3. [§5.7] The bandwidth sensitivity analysis simulates different off-chip bandwidths by changing the amount of data moved off-chip and padding the remainder on-chip. This alters the access pattern and does not reproduce the timing behavior of a real bandwidth change (DRAM bank conflicts, refresh, AXI arbitration), so the conclusion that 'the current use of bandwidth is already highly efficient' and the associated 78.6% utilization figure are not directly validated. Please either implement a real bandwidth change (e.g., by clock or interconnect configuration) or explicitly label this as a first-order estimate and discuss its limitations. This does not affect the directly measured latency/throughput results but weakens the paper's bandwidth-efficiency argument.
minor comments (6)
  1. [Table 7] The header row of Table 7 spells the design name as 'RSD-XNN'; this should be 'RSN-XNN'.
  2. [§4.1] In the second sentence of Section 4.1, 'for the LHS operands from MeshB FUs' should read 'for the RHS operands from MeshB FUs'.
  3. [§5.6, Table 10] The GPU comparison reports single-point latency and power numbers (vendor reports for T4/V100/A100 and one Colab session for L4) without variance or measurement repetitions; please report the number of runs and the observed spread, and clarify whether the VCK190 power in Table 10 is the BEAM measurement from Section 5 or the Vivado estimate from Table 4.
  4. [§4.2] The 'first-order formula-based calculation' used for model segmentation is mentioned but never specified; please include the formula (or pseudocode) so that the segmentation decision can be reproduced.
  5. [§3.3] The deadlock-free FIFO depth of six is reported for the tested implementation only; a short discussion of how the required FIFO depth scales with the number of FUs, stream depths, and instruction windows would help readers apply the result to other RSN configurations.
  6. [§1] The phrase 'dynamic layer fusion' in the contributions list and abstract is not formally defined; please clarify that it refers to runtime switching between mapping types such as pipelined, fused, and layer-by-layer execution.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the measured prototype results are self-contained and open to external checking, and the manual datapath construction is a scope limitation rather than a circular step.

full rationale

The paper's central numerical claims—6.1x latency reduction, 2.4x–3.2x throughput improvement, and 2.1x FP32 energy efficiency versus A100—are based on on-board measurements of the RSN-XNN prototype, not on parameters fitted to the outcomes being predicted. The evaluation section reports measured latencies, resource utilization, power estimates, and instruction counts, and the artifact appendix provides a bootable SD card image, reference outputs, and scripts for reproducing the reported 17.98 ms BERT-Large encoder latency. No equation in the paper reduces a predicted quantity to a fitted input or to the definition of another claimed result. The RSN abstraction is introduced conceptually and then instantiated in hardware; the flexibility and low instruction overhead are supported by direct measurements such as the 1.4 MB/s instruction processing rate and the 1.6 GFLOPs-per-byte compute-to-instruction ratio. The manual datapath generation process described in Section 4.2 and the statement in Section 4.5 that 'Exploring the automatic generation of the datapath from arbitrary input code is beyond the scope of this paper' are genuine scope limitations that constrain the generality of the claimed abstraction, but they do not make the reported measurements circular. The comparison baseline CHARM shares a co-author with the present paper, but it is a published prior implementation used as a benchmark, not an unverified self-citation that the argument depends on. Overall, the derivation chain is self-contained and empirically grounded, so the circularity score is 0.

Assumptions & free parameters 3 free parameters · 5 assumptions · 2 invented entities

The results rest on hand-tuned tiling, stream-grouping, and FIFO depth choices; on workload assumptions of determinism and low entropy; and on standard performance-modeling assumptions. The RSN FU and instruction packet are new design constructs whose only evidence so far is the authors' own artifact.

free parameters (3)
  • MM tiling parameters (LHS 768x128, RHS 128x1024, OUT 768x1024) = 768x128, 128x1024, 768x1024
    Chosen to achieve 768x RHS reuse and 1024x LHS reuse, directly enabling the 4.7 TFLOPS result in Table 6b; these are hand-tuned to the VCK190's bandwidth/compute ratio and not derived from first principles.
  • AIE tile grouping (4x4x4, 64 tiles per MME) = 64 tiles per MME, 6 groups, 384 tiles
    Hand-chosen grouping to fit within the 234/156 AIE-to-PL stream limits while using 96% of AIE tiles; this configuration is a design decision affecting the achieved GEMM throughput.
  • Decoder FIFO depth between uOP and mOP decoders = 6
    The paper asserts this depth is deadlock-free in their implementation without proof; increasing it would change area, decreasing it risks deadlock, so the value is an ad hoc design parameter for the deadlock-freedom claim.
assumptions (5)
  • domain assumption DNN execution is deterministic and has low control information entropy, so compile-time scheduling without runtime speculation is sufficient.
    Stated in §1 ('the deterministic nature of DNN execution allows for compile-time analysis of data dependencies') and §2.1; underpins the entire network-of-FUs abstraction and rules out data-dependent dynamic behavior.
  • domain assumption Latency-insensitive streaming with matching send/receive counts is a correct execution model; a receiver with fewer sends blocks indefinitely, a sender with excess sends blocks when the channel is full.
    Defined in §3.1; this is the synchronization contract of RSN. It assumes the programmer/compiler can statically match stream counts, which is not proven for all programs.
  • domain assumption The roofline model adequately estimates latency for the four mapping types in Table 3.
    Used in §4.3 to justify Type D pipelining for attention layers; standard performance modeling assumption.
  • ad hoc to paper Simulating bandwidth variations by padding data on-chip preserves the latency model.
    Table 11's sensitivity analysis pads data on-chip to emulate lower/higher bandwidth, which changes data movement patterns and may not faithfully represent real bandwidth changes.
  • domain assumption Vivado power estimates (over-estimated in absolute terms) provide valid ratios for energy comparisons.
    The paper states Table 4 numbers are 'over-estimated in absolute terms' but uses them to argue decoder overhead <0.08%; the validity of these ratios is not cross-checked against on-board measurement.
invented entities (2)
  • RSN functional unit (FU) abstraction
    purpose: Virtualize heterogeneous AIE/PL/memory resources as stateful network nodes with uOP control.
    The abstraction is the paper's proposal; its only implementation is the authors' RSN-XNN, and no third-party reproduction is reported.
  • RSN instruction packet with window size and reuse
    purpose: Compress repeated uOP sequences to reduce instruction overhead.
    The 1.6 GFLOPs/byte figure is measured on the authors' prototype; no external benchmark validates the packet format.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Reconfigurable Stream Network Architecture." pith.science (2026). https://pith.science/paper/2PCMGJHU

@misc{pith2026241117966,
  author       = {Pith},
  title        = {Pith review of: Reconfigurable Stream Network Architecture},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2PCMGJHU}},
  note         = {Machine review of arXiv:2411.17966}
}
read the original abstract

As AI systems grow increasingly specialized and complex, managing hardware heterogeneity becomes a pressing challenge. How can we efficiently coordinate and synchronize heterogeneous hardware resources to achieve high utilization? How can we minimize the friction of transitioning between diverse computation phases, reducing costly stalls from initialization, pipeline setup, or drain? Our insight is that a network abstraction at the ISA level naturally unifies heterogeneous resource orchestration and phase transitions. This paper presents a Reconfigurable Stream Network Architecture (RSN), a novel ISA abstraction designed for the DNN domain. RSN models the datapath as a circuit-switched network with stateful functional units as nodes and data streaming on the edges. Programming a computation corresponds to triggering a path. Software is explicitly exposed to the compute and communication latency of each functional unit, enabling precise control over data movement for optimizations such as compute-communication overlap and layer fusion. As nodes in a network naturally differ, the RSN abstraction can efficiently virtualize heterogeneous hardware resources by separating control from the data plane, enabling low instruction-level intervention. We build a proof-of-concept design RSN-XNN on VCK190, a heterogeneous platform with FPGA fabric and AI engines. Compared to the SOTA solution on this platform, it reduces latency by 6.1x and improves throughput by 2.4x-3.2x. Compared to the T4 GPU with the same FP32 performance, it matches latency with only 18% of the memory bandwidth. Compared to the A100 GPU at the same 7nm process node, it achieves 2.1x higher energy efficiency in FP32.

Figures

Figures reproduced from arXiv: 2411.17966 by the authors.

Figure 1
Figure 1. Reconfigurable Stream Network Overview Motivated by these observed challenges, we ask the question: What is the right abstraction for bridging software with highly het￾erogeneous hardware? Ideally, it should meet two key requirements: • Flexibility: Computation and bandwidth must be flexibly allocated to support different phases, such as prolog, steady state, and epilog within a layer, as well as varying operator ty… view at source ↗
Figure 3
Figure 3. Four Mapping Types and Their Disadvantages [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 4
Figure 4. Functional Unit and Datapath Abstraction [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figures from the paper (10 more)
Figure 6
Figure 6. Figure 6: Comparison of RSN and Baseline Datapath must ensure that the number of sends from the producer kernel exactly matches the number of receives in the consumer kernels. If the sends are fewer than the receives, the receiving kernel will block indefinitely; if the sends ex…
Figure 7
Figure 7. Figure 7: A Flexible Datapath Supporting Dynamic Two Se [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: Instruction Decoder: Fuse uOP Streams into 1 RSN [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 10
Figure 10. Figure 10: RSN-XNN Datapath and Example Application [PITH_FULL_IMAGE:figures/full_fig_p008_10.png]
Figure 11
Figure 11. Figure 11: Pipeline Non-MMs and Their Adjacent MMs [PITH_FULL_IMAGE:figures/full_fig_p010_11.png]
Figure 12
Figure 12. Figure 12: Three Ways to Map Load and Store Operations to [PITH_FULL_IMAGE:figures/full_fig_p010_12.png]
Figure 14
Figure 14. Figure 14: Device View of the Routed Design [PITH_FULL_IMAGE:figures/full_fig_p011_14.png]
Figure 15
Figure 15. Figure 15: Power Estimation Summary (Total 98.66 W) [PITH_FULL_IMAGE:figures/full_fig_p011_15.png]
Figure 17
Figure 17. Figure 17: Reuse of AIE to/from PL Streams [PITH_FULL_IMAGE:figures/full_fig_p012_17.png]
Figure 18
Figure 18. Figure 18: Achieved Latency/Throughput VS CHARM [119] [PITH_FULL_IMAGE:figures/full_fig_p012_18.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

121 extracted references · 26 canonical work pages

  1. [1]

    Abdelfattah, David Han, Andrew Bitar, Roberto DiCecco, Shane O’Connell, Nitika Shanker, Joseph Chu, Ian Prins, Joshua Fender, Andrew C

    Mohamed S. Abdelfattah, David Han, Andrew Bitar, Roberto DiCecco, Shane O’Connell, Nitika Shanker, Joseph Chu, Ian Prins, Joshua Fender, Andrew C. Ling, and Gordon R. Chiu. 2018. DLA: Compiler and FPGA Overlay for Neural Network Inference Acceleration. In 2018 28th International Conference on Field Reconfigurable Stream Network Architecture ISCA ’25, June...

  2. [2]

    Riadh Ben Abdelhamid, Yoshiki Yamaguchi, and Taisuke Boku. 2021. A Highly- Efficient and Tightly-Connected Many-Core Overlay Architecture. IEEE Access 9 (2021), 65277–65292. https://doi.org/10.1109/ACCESS.2021.3074171

  3. [3]

    Dennis Abts, Garrin Kimmell, Andrew Ling, John Kim, Matt Boyd, Andrew Bitar, Sahil Parmar, Ibrahim Ahmed, Roberto DiCecco, David Han, John Thompson, Michael Bye, Jennifer Hwang, Jeremy Fowers, Peter Lillian, Ashwin Murthy, Elyas Mehtabuddin, Chetan Tekur, Thomas Sohmers, Kris Kang, Stephen Maresh, and Jonathan Ross. 2022. A software-defined tensor streami...

  4. [4]

    Sagheer Ahmad, Sridhar Subramanian, Vamsi Boppana, Shankar Lakka, Fu-Hing Ho, Tomai Knopp, Juanjo Noguera, Gaurav Singh, and Ralph Wittig. 2019. Xilinx First 7nm Device: Versal AI Core (VC1902). In2019 IEEE Hot Chips 31 Symposium (HCS). 1–28. https://doi.org/10.1109/HOTCHIPS.2019.8875639

  5. [5]

    AMD. 2021. BEAM Tools. Available: https://xilinx-wiki.atlassian.net/ wiki/spaces/A/pages/973078551/BEAM+Tool+for+VCK190+Evaluation+Kit, Ac- cessed: 16-August-2024

  6. [6]

    AMD. 2021. Versal AI Core Series VCK190 Evaluation Kit . Available: https: //www.xilinx.com/products/boards-and-kits/vck190.html, Accessed: 16-August- 2024

  7. [7]

    AMD. 2023. DPU IP Details and System Integration. Available: https:// xilinx.github.io/Vitis-AI/3.5/html/docs/workflow-system-integration, Accessed: 5-November-2024

  8. [8]

    AMD. 2023. Vitis Unified Software Platform 2023.2 . https://www.xilinx.com/ products/design-tools/vitis.html Software

Show all 121 references
  1. [9]

    AMD. 2024. Versal ACAP Package Pinout Documentation: Mechan- ical - VC1802 and VC1902. https://docs.amd.com/r/en-US/am013- versal-pkg-pinout/VIVA1596-Mechanical-VC1802-and-VC1902 Avail- able: https://docs.amd.com/r/en-US/am013-versal-pkg-pinout/VIVA1596- Mechanical-VC1802-and-...

  2. [10]

    AMD. 2024. Vivado Design Suite User Guide. Available: https://docs.amd.com/ r/en-US/ug906-vivado-design-analysis/Report-Power, Accessed: 5-November- 2024

  3. [11]

    Giovanni Ansaloni, Paolo Bonzini, and Laura Pozzi. 2011. EGRA: A Coarse Grained Reconfigurable Architectural Template.IEEE Transactions on Very Large Scale Integration (VLSI) Systems 19, 6 (2011), 1062–1074. https://doi.org/10.1109/ TVLSI.2010.2044667

  4. [12]

    Oguzhan Atak and Abdullah Atalar. 2013. BilRC: An Execution Triggered Coarse Grained Reconfigurable Architecture. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 21, 7 (2013), 1285–1298. https://doi.org/10.1109/ TVLSI.2012.2207748

  5. [13]

    Yueyin Bai, Hao Zhou, Keqing Zhao, Hongji Wang, Jianli Chen, Jun Yu, and Kun Wang. 2023. FET-OPU: A Flexible and Efficient FPGA-Based Overlay Processor for Transformer Networks. In2023 IEEE/ACM International Conference on Computer Aided Design (ICCAD) . 1–9. https://doi.org/10...

  6. [14]

    Thilini Kaushalya Bandara, Dhananjaya Wijerathne, Tulika Mitra, and Li-Shiuan Peh. 2022. REVAMP: a systematic framework for heterogeneous CGRA realiza- tion. In Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operatin...

  7. [15]

    Suhail Basalama and Jason Cong. 2025. Stream-HLS: Towards Automatic Dataflow Acceleration. In Proceedings of the 2025 ACM/SIGDA International Symposium on Field Programmable Gate Arrays (Monterey, CA, USA) (FPGA ’25). Association for Computing Machinery, New York, NY, USA, 103...

  8. [16]

    Suhail Basalama, Atefeh Sohrabizadeh, Jie Wang, and Jason Cong. 2022. A Versatile Systolic Array for Transposed and Dilated Convolution on FPGA. In 2022 IEEE 30th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM). 1–2. https://doi.org/10.110...

  9. [17]

    Suhail Basalama, Atefeh Sohrabizadeh, Jie Wang, Licheng Guo, and Jason Cong

  10. [18]

    Suhail Basalama, Jie Wang, and Jason Cong. 2023. A Comprehensive Automated Exploration Framework for Systolic Array Designs. In 2023 60th ACM/IEEE Design Automation Conference (DAC). 1–6. https://doi.org/10.1109/DAC56929. 2023.10248016

  11. [19]

    Riadh Ben Abdelhamid, Yoshiki Yamaguchi, and Taisuke Boku. 2019. MITRACA: A Next-Gen Heterogeneous Architecture. In 2019 IEEE 13th International Sym- posium on Embedded Multicore/Many-core Systems-on-Chip (MCSoC) . 304–311. https://doi.org/10.1109/MCSoC.2019.00050

  12. [20]

    Maciej Besta, Marc Fischer, Tal Ben-Nun, Dimitri Stanojevic, Johannes De Fine Licht, and Torsten Hoefler. 2020. Substream-Centric Maximum Matchings on FPGA. ACM Trans. Reconfigurable Technol. Syst. 13, 2, Article 8 (April 2020), 33 pages. https://doi.org/10.1145/3377871

  13. [21]

    Andrew Boutros, Aman Arora, and Vaughn Betz. 2024. Field-Programmable Gate Array Architecture for Deep Learning: Survey & Future Directions. arXiv:2404.10076 [cs.AR]

  14. [22]

    Hoe, Vaughn Betz, and Martin Langhammer

    Andrew Boutros, Eriko Nurvitadhi, Rui Ma, Sergey Gribok, Zhipeng Zhao, James C. Hoe, Vaughn Betz, and Martin Langhammer. 2020. Beyond Peak Per- formance: Comparing the Real Performance of AI-Optimized FPGAs and GPUs. In 2020 International Conference on Field-Programmable Techn...

  15. [23]

    Jingwei Cai, Yuchen Wei, Zuotong Wu, Sen Peng, and Kaisheng Ma. 2023. core.cpp in SET-ISCA2023. Available: https://github.com/SET-Scheduling- Project/SET-ISCA2023/blob/master/src/core.cpp, Accessed: 5-November-2024

  16. [24]

    Jingwei Cai, Yuchen Wei, Zuotong Wu, Sen Peng, and Kaisheng Ma. 2023. Inter- layer Scheduling Space Definition and Exploration for Tiled Accelerators. In Proceedings of the 50th Annual International Symposium on Computer Architecture (Orlando, FL, USA) (ISCA ’23). Association ...

  17. [25]

    Jingwei Cai, Zuotong Wu, Sen Peng, Yuchen Wei, Zhanhong Tan, Guim- ing Shi, Mingyu Gao, and Kaisheng Ma. 2024. core.cpp in GEMINI- HPCA2024. Available: https://github.com/SET-Scheduling-Project/GEMINI- HPCA2024/blob/master/src/core.cpp, Accessed: 5-November-2024

  18. [26]

    Jingwei Cai, Zuotong Wu, Sen Peng, Yuchen Wei, Zhanhong Tan, Guiming Shi, Mingyu Gao, and Kaisheng Ma. 2024. Gemini: Mapping and Architecture Co- exploration for Large-scale DNN Chiplet Accelerators. In2024 IEEE International Symposium on High-Performance Computer Architecture...

  19. [28]

    Hongzheng Chen, Jiahao Zhang, Yixiao Du, Shaojie Xiang, Zichao Yue, Niansong Zhang, Yaohui Cai, and Zhiru Zhang. 2024. Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference. ACM Transactions on Reconfigurable Technology and Systems (...

  20. [29]

    Alexander Chin, Kuang Ping Niu, Matthew Walker, Shizhang Yin, Alexander Mertens, Jongeun Lee, and Jason H

    S. Alexander Chin, Kuang Ping Niu, Matthew Walker, Shizhang Yin, Alexander Mertens, Jongeun Lee, and Jason H. Anderson. 2018. Architecture Exploration of Standard-Cell and FPGA-Overlay CGRAs Using the Open-Source CGRA-ME Framework. In Proceedings of the 2018 International Symp...

  21. [30]

    Dally, Ujval J

    William J. Dally, Ujval J. Kapasi, Brucek Khailany, Jung Ho Ahn, and Abhishek Das. 2004. Stream Processors: Progammability and Efficiency: Will this new kid on the block muscle out ASIC and DSP? Queue 2, 1 (March 2004), 52–62. https://doi.org/10.1145/984458.984486

  22. [31]

    Xiaodong Deng, Shijie Wang, Tianyi Gao, Jing Liu, Longjun Liu, and Nanning Zheng. 2024. AMA: An Analytical Approach to Maximizing the Efficiency of Deep Learning on Versal AI Engine. In 2024 34th International Conference on Field-Programmable Logic and Applications (FPL) . 227...

  23. [32]

    Reinhardt, Adrian M

    Jeremy Fowers, Kalin Ovtcharov, Michael Papamichael, Todd Massengill, Ming Liu, Daniel Lo, Shlomi Alkalay, Michael Haselman, Logan Adams, Mahdi Ghandi, Stephen Heil, Prerak Patel, Adam Sapek, Gabriel Weisz, Lisa Woods, Sitaram Lanka, Steven K. Reinhardt, Adrian M. Caulfield, E...

  24. [33]

    Mingyu Gao, Xuan Yang, Jing Pu, Mark Horowitz, and Christos Kozyrakis

  25. [34]

    Mingyu Gao, Xuan Yang, Jing Pu, Mark Horowitz, and Christos Kozyrakis. 2019. TANGRAM: Optimized Coarse-Grained Dataflow for Scalable NN Accelerators. In Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating S...

  26. [35]

    Graham Gobieski, Ahmet Oguz Atli, Kenneth Mai, Brandon Lucia, and Nathan Beckmann. 2021. Snafu: An Ultra-Low-Power, Energy-Minimal CGRA- Generation Framework and Architecture. In 2021 ACM/IEEE 48th Annual In- ternational Symposium on Computer Architecture (ISCA) . 1027–1040. h...

  27. [36]

    Venkatraman Govindaraju, Chen-Han Ho, Tony Nowatzki, Jatin Chhugani, Na- dathur Satish, Karthikeyan Sankaralingam, and Changkyu Kim. 2012. DySER: Unifying Functionality and Parallelism Specialization for Energy-Efficient Com- puting. IEEE Micro 32, 5 (2012), 38–51. https://doi...

  28. [37]

    Yijin Guan, Hao Liang, Ningyi Xu, Wenqiang Wang, Shaoshuai Shi, Xi Chen, Guangyu Sun, Wei Zhang, and Jason Cong. 2017. FP-DNN: An Automated Framework for Mapping Deep Neural Networks onto FPGAs with RTL-HLS Hybrid Templates. In 2017 IEEE 25th Annual International Symposium on ...

  29. [38]

    Licheng Guo, Yuze Chi, Jason Lau, Linghao Song, Xingyu Tian, Moazin Khatti, Weikang Qiao, Jie Wang, Ecenur Ustun, Zhenman Fang, Zhiru Zhang, and Jason Cong. 2023. TAPA: A Scalable Task-parallel Dataflow Programming Framework for Modern FPGAs with Co-optimization of HLS and Phy...

  30. [39]

    Licheng Guo, Pongstorn Maidee, Yun Zhou, Chris Lavin, Eddie Hung, Wuxi Li, Jason Lau, Weikang Qiao, Yuze Chi, Linghao Song, Yuanlong Xiao, Alireza Kaviani, Zhiru Zhang, and Jason Cong. 2023. RapidStream 2.0: Automated Parallel Implementation of Latency–Insensitive FPGA Designs...

  31. [40]

    Zibo Guo, Kai Liu, Wei Liu, Xiaoyao Sun, Chongyang Ding, and Shangrong Li. 2024. An Overlay Accelerator of DeepLab CNN for Spacecraft Image Seg- mentation on FPGA. Remote Sensing 16 (03 2024), 894. https://doi.org/10.3390/ rs16050894

  32. [41]

    Mathew Hall and Vaughn Betz. 2020. HPIPE: Heterogeneous Layer-Pipelined and Sparse-Aware CNN Inference for FPGAs. In Proceedings of the 2020 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (Seaside, CA, USA) (FPGA ’20). Association for Computing Machinery, ...

  33. [43]

    Zifan He, Anderson Truong, Yingqi Cao, and Jason Cong. 2025. InTAR: Inter- Task Auto-Reconfigurable Accelerator Design for High Data Volume Variation in DNNs. arXiv preprint arXiv:2502.08807 (2025)

  34. [44]

    Seongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee, Minsub Kim, Dong- soo Lee, and Joo-Young Kim. 2022. DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation. In 2022 IEEE Hot Chips 34 Symposium (HCS). 1–17. https://doi.org/10.1109/HCS55...

  35. [45]

    Hugging Face

    Inc. Hugging Face. 2023. BERT-LARGE model implementation in PyTorch. https://huggingface.co/transformers/model_doc/bert.html Software

  36. [46]

    Suyeon Hur, Seongmin Na, Dongup Kwon, Joonsung Kim, Andrew Boutros, Eriko Nurvitadhi, and Jangwoo Kim. 2023. A Fast and Flexible FPGA-based Accelerator for Natural Language Processing Neural Networks. ACM Trans. Archit. Code Optim. 20, 1, Article 11 (Feb 2023), 24 pages. https...

  37. [47]

    Intel. 2020. INT8 VS. FP32 Performance Comparision. Avail- able: https://intelkevinputnam.github.io/openvino-docs/pages/openvino_docs_ performance_int8_vs_fp32.html, Accessed: 5-November-2024

  38. [48]

    Intel. 2019. Intel Deep Learning Boost. https://www.intel.com/content/ dam/www/central-libraries/us/en/documents/2022-09/xeon-accelerated-ai- product-brief.pdf

  39. [49]

    Lana Josipović, Radhika Ghosal, and Paolo Ienne. 2018. Dynamically Sched- uled High-level Synthesis. In Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (Monterey, CALIFORNIA, USA) (FPGA ’18). Association for Computing Machinery, New ...

  40. [50]

    Lana Josipovic, Andrea Guerrieri, and Paolo Ienne. 2021. Synthesizing General- Purpose Code Into Dynamically Scheduled Circuits. IEEE Circuits and Systems Magazine 21, 2 (2021), 97–118. https://doi.org/10.1109/MCAS.2021.3071631

  41. [52]

    Kapasi, W.J

    U.J. Kapasi, W.J. Dally, S. Rixner, J.D. Owens, and B. Khailany. 2002. The Imagine Stream Processor. In Proceedings. IEEE International Conference on Computer Design: VLSI in Computers and Processors . 282–288. https://doi.org/10.1109/ ICCD.2002.1106783

  42. [53]

    Kapasi, S

    U.J. Kapasi, S. Rixner, W.J. Dally, B. Khailany, Jung Ho Ahn, P. Mattson, and J.D. Owens. 2003. Programmable stream processors. Computer 36, 8 (2003), 54–62. https://doi.org/10.1109/MC.2003.1220582

  43. [54]

    Manupa Karunaratne, Aditi Kulkarni Mohite, Tulika Mitra, and Li-Shiuan Peh

  44. [55]

    Hamza Khan, Asma Khan, Zainab Khan, Lun Bin Huang, Kun Wang, and Lei He

  45. [56]

    Konijnen- burg, Soojung Ryu, and Jeongwook Kim

    Changmoo Kim, Moo-Kyoung Chung, Yeon-Gon Cho, Mario H. Konijnen- burg, Soojung Ryu, and Jeongwook Kim. 2012. ULP-SRP: Ultra low power Samsung Reconfigurable Processor for biomedical applications. 2012 Interna- tional Conference on Field-Programmable Technology (2012), 329–334....

  46. [57]

    Mahapatra, and Kiyoung Choi

    Yoonjin Kim, Rabi N. Mahapatra, and Kiyoung Choi. 2010. Design Space Ex- ploration for Efficient Resource Utilization in Coarse-Grained Reconfigurable Architecture. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 18, 10 (2010), 1471–1482. https://doi.org/10.11...

  47. [58]

    Kalhan Koul, Jackson Melchert, Kavya Sreedhar, Leonard Truong, Gedeon Nyen- gele, Keyi Zhang, Qiaoyi Liu, Jeff Setter, Po-Han Chen, Yuchen Mei, Maxwell Strange, Ross Daly, Caleb Donovick, Alex Carsello, Taeyoung Kong, Kathleen Feng, Dillon Huff, Ankita Nayak, Rajsekhar Setalur...

  48. [59]

    Ronny Krashinsky, Christopher Batten, Mark Hampton, Steve Gerding, Brian Pharris, Jared Casper, and Krste Asanovic. 2004. The Vector-Thread Architec- ture. In Proceedings of the 31st Annual International Symposium on Computer Architecture (München, Germany) (ISCA ’04). IEEE Co...

  49. [60]

    Hyoukjun Kwon, Prasanth Chatarasi, Vivek Sarkar, Tushar Krishna, Michael Pellauer, and Angshuman Parashar. 2020. AHWAccelerator.hpp in Mae- stro. Available: https://github.com/maestro-project/maestro/blob/master/cost- model/include/abstract-hardware-model/AHW_Accelerator.hpp, ...

  50. [61]

    Hyoukjun Kwon, Prasanth Chatarasi, Vivek Sarkar, Tushar Krishna, Michael Pellauer, and Angshuman Parashar. 2020. MAESTRO: A Data-Centric Approach to Understand Reuse, Performance, and Hardware Cost of DNN Mappings.IEEE Micro 40, 3 (2020), 20–29

  51. [62]

    Martin Langhammer, Eriko Nurvitadhi, Bogdan Pasca, and Sergey Gribok. 2021. Stratix 10 NX Architecture and Applications. In The 2021 ACM/SIGDA Inter- national Symposium on Field-Programmable Gate Arrays (Virtual Event, USA) (FPGA ’21). Association for Computing Machinery, New ...

  52. [63]

    Jason Lau, Yuanlong Xiao, Yutong Xie, Yuze Chi, Linghao Song, Shaojie Xiang, Michael Lo, Zhiru Zhang, Jason Cong, and Licheng Guo. 2025. RapidStream IR: Infrastructure for FPGA High-Level Physical Synthesis. InProceedings of the 43rd IEEE/ACM International Conference on Comput...

  53. [64]

    Yunsup Lee, Rimas Avizienis, Alex Bishara, Richard Xia, Derek Lockhart, Christo- pher Batten, and Krste Asanović. 2011. Exploring the tradeoffs between pro- grammability and efficiency in data-parallel accelerators. In 2011 38th Annual International Symposium on Computer Archi...

  54. [65]

    Bingbing Li, Santosh Pandey, Haowen Fang, Yanjun Lyv, Ji Li, Jieyang Chen, Mimi Xie, Lipeng Wan, Hang Liu, and Caiwen Ding. 2020. FTRANS: energy- efficient acceleration of transformers using FPGA. In Proceedings of the ACM/IEEE International Symposium on Low Power Electronics ...

  55. [66]

    Jingyuan Li, Yunhui Qiu, Guowei Zhu, Qilong Zhu, Wenbo Yin, and Lingli Wang

  56. [67]

    Sihao Liu, Jian Weng, Dylan Kupsh, Atefeh Sohrabizadeh, Zhengrong Wang, Licheng Guo, Jiuyang Liu, Maxim Zhulin, Rishabh Mani, Lucheng Zhang, Ja- son Cong, and Tony Nowatzki. 2022. OverGen: Improving FPGA Usability Reconfigurable Stream Network Architecture ISCA ’25, June 21–25...

  57. [68]

    Rui Ma, Jia-Ching Hsu, Tian Tan, Eriko Nurvitadhi, David Sheffield, Rob Pelt, Martin Langhammer, Jaewoong Sim, Aravind Dasu, and Derek Chiou. 2019. Specializing FGPU for Persistent Deep Learning. In 2019 29th International Conference on Field Programmable Logic and Application...

  58. [69]

    Bingfeng Mei, Serge Vernalde, Diederik Verkest, Hugo De Man, and Rudy Lauw- ereins. 2003. ADRES: An Architecture with Tightly Coupled VLIW Processor and Coarse-Grained Reconfigurable Matrix. In International Conference on Field- Programmable Logic and Applications. https://api...

  59. [70]

    Thierry Moreau, Tianqi Chen, Luis Vega, Jared Roesch, Eddie Yan, Lianmin Zheng, Josh Fromm, Ziheng Jiang, Luis Ceze, Carlos Guestrin, and Arvind Kr- ishnamurthy. 2019. A Hardware–Software Blueprint for Flexible Deep Learning Specialization. IEEE Micro 39, 5 (2019), 8–16. https...

  60. [71]

    In 2023 IEEE International Symposium on Circuits and Systems (ISCAS)

    THRAM: A Template-based Heterogeneous CGRA Modeling Framework Supporting Fast DSE. In 2023 IEEE International Symposium on Circuits and Systems (ISCAS). 1–5. https://doi.org/10.1109/ISCAS46773.2023.10182204

  61. [72]

    Chris Nicol. 2017. A Coarse Grain Reconfigurable Array ( CGRA ) for Statically Scheduled Data Flow Computing. https://api.semanticscholar.org/CorpusID: 199394670

  62. [74]

    NVIDIA. 2017. NVIDIA Tesla V100 GPU Architecture. (2017). Available: https://images.nvidia.com/content/volta-architecture/pdf/volta-architecture- whitepaper.pdf, Accessed: 21-November-2024

  63. [75]

    NVIDIA. 2018. NVIDIA T4 Tensor Core GPU. https://resources.nvidia.com/en- us-gpu-resources/t4-tensor-core-datas?lx=CPwSfP Available: https://resources. nvidia.com/en-us-gpu-resources/t4-tensor-core-datas?lx=CPwSfP, Accessed: 21-November-2024

  64. [76]

    Rene Mueller, Jens Teubner, and Gustavo Alonso. 2009. Streams on wires: a query compiler for FPGAs. Proc. VLDB Endow. 2, 1 (Aug. 2009), 229–240. https://doi.org/10.14778/1687627.1687654

  65. [77]

    NVIDIA. 2024. DeepLearningExamples: BERT Language Modeling with Tensor- Flow 2. Available: https://github.com/NVIDIA/DeepLearningExamples/tree/ master/TensorFlow2/LanguageModeling/BERT, Accessed: 16-August-2024

  66. [78]

    NVIDIA. 2024. NVIDIA H100 TENSOR CORE GPU. Available: https://resources. nvidia.com/en-us-tensor-core/nvidia-tensor-core-gpu-datasheet, Accessed: 20- Feburary-2025

  67. [79]

    NVIDIA. 2024. NVIDIA L4 TENSOR CORE GPU. Available: https:// resources.nvidia.com/en-us-data-center-overview/l4-gpu-datasheet, Accessed: 21-November-2024

  68. [80]

    Nvidia. 2025. Nsight Compute CLI. Available: https://docs.nvidia.com/nsight- compute/NsightComputeCli/index.html, Accessed: 14-February-2025

  69. [81]

    NVIDIA. 2021. NVIDIA A100 TENSOR CORE GPU. Avail- able: https://www.nvidia.com/content/dam/en-zz/Solutions/Data- Center/a100/pdf/nvidia-a100-datasheet-us-nvidia-1758950-r4-web.pdf, Accessed: 21-November-2024

  70. [82]

    Angshuman Parashar, Michael Pellauer, Michael Adler, Bushra Ahsan, Neal Crago, Daniel Lustig, Vladimir Pavlov, Antonia Zhai, Mohit Gambhir, Aamer Jaleel, Randy Allmon, Rachid Rayess, Stephen Maresh, and Joel Emer. 2013. Triggered instructions: a control paradigm for spatially-...

  71. [83]

    Artur Podobas, Kentaro Sano, and Satoshi Matsuoka. 2020. A Survey on Coarse- Grained Reconfigurable Architectures From a Performance Perspective. IEEE Access 8 (2020), 146719–146743. https://doi.org/10.1109/ACCESS.2020.3012084

  72. [84]

    Shah, Zhengyu Chen, Kaizhao Liang, Swayambhoo Jain, Urmish Thakker, Dawei Huang, Sumti Jairath, Kevin J

    Raghu Prabhakar, Ram Sivaramakrishnan, Darshan Gandhi, Yun Du, Mingran Wang, Xiangyu Song, Kejie Zhang, Tianren Gao, Angela Wang, Xiaoyan Li, Yongning Sheng, Joshua Brot, Denis Sokolov, Apurv Vivek, Calvin Leung, Arjun Sabnis, Jiayu Bai, Tuowen Zhao, Mark Gottscho, David Jacks...

  73. [86]

    Oliveira, Michael Canesche, Lucas Reis, José Augusto Miranda Nacif, and Ricardo S

    Westerley C. Oliveira, Michael Canesche, Lucas Reis, José Augusto Miranda Nacif, and Ricardo S. Ferreira. 2022. Heterogeneous reconfigurable architectures for machine learning dataflows. Concurrency and Computation: Practice and Experience 35 (2022). https://api.semanticschola...

  74. [87]

    Karthikeyan Sankaralingam, Tony Nowatzki, Vinay Gangadhar, Preyas Shah, Michael Davies, William Galliher, Ziliang Guo, Jitu Khare, Deepak Vijay, Poly Palamuttam, Maghawan Punde, Alex Tan, Vijay Thiruvengadam, Rongyi Wang, and Shunmiao Xu. 2022. The Mozart reuse exposed dataflo...

  75. [88]

    Colin Schmidt. 2021. Extending Temporal-Vector Microarchitectures for Two- Dimensional Computations. Ph. D. Dissertation. University of California, Berke- ley, USA. https://www.escholarship.org/uc/item/2mr167rk

  76. [89]

    Yongming Shen, Michael Ferdman, and Peter Milder. 2017. Maximizing CNN Accelerator Efficiency Through Resource Partitioning.SIGARCH Comput. Archit. News 45, 2 (2017), 535–547. https://doi.org/10.1145/3140659.3080221

  77. [90]

    Singh, Ming-Hau Lee, Guangming Lu, F.J

    H. Singh, Ming-Hau Lee, Guangming Lu, F.J. Kurdahi, N. Bagherzadeh, and E.M. Chaves Filho. 2000. MorphoSys: an integrated reconfigurable system for data-parallel and computation-intensive applications. IEEE Trans. Comput. 49, 5 (2000), 465–481. https://doi.org/10.1109/12.859540

  78. [91]

    James E. Smith. 1982. Decoupled access/execute computer architectures. In Proceedings of the 9th Annual Symposium on Computer Architecture (Austin, Texas, USA) (ISCA ’82). IEEE Computer Society Press, Washington, DC, USA, 112–119

  79. [92]

    Richard M. Russell. 1978. The CRAY-1 computer system. Commun. ACM 21, 1 (Jan. 1978), 63–72. https://doi.org/10.1145/359327.359336

  80. [93]

    Atefeh Sohrabizadeh, Yuze Chi, and Jason Cong. 2022. StreamGCN: Accelerating Graph Convolutional Networks with Streaming Processing. In2022 IEEE Custom Integrated Circuits Conference (CICC) . 1–8. https://doi.org/10.1109/CICC53496. 2022.9772832

  81. [94]

    Linghao Song, Yuze Chi, Atefeh Sohrabizadeh, Young-kyu Choi, Jason Lau, and Jason Cong. 2022. Sextans: A Streaming Accelerator for General-Purpose Sparse- Matrix Dense-Matrix Multiplication. In Proceedings of the 2022 ACM/SIGDA International Symposium on Field-Programmable Gat...

  82. [95]

    Nambiar, Anh Tuan Do, Thilini Kaushalya Bandara, Aditi Kulkarni Mohite, and Bo Wang

    Lingzhi Su, Wang Ling Goh, Jingjing Lan, Vishnu P. Nambiar, Anh Tuan Do, Thilini Kaushalya Bandara, Aditi Kulkarni Mohite, and Bo Wang. 2022. An Energy-Efficient Processing Element Design for Coarse-Grained Reconfigurable Architecture on FPGA. In 2022 11th International Confer...

  83. [96]

    Endri Taka, Aman Arora, Kai Chiang Wu, and Diana Marculescu. 2023. MaxEVA: Maximizing the Efficiency of Matrix Multiplication on Versal AI Engine. In2023 International Conference on Field Programmable Technology (ICFPT). IEEE, 96–105. https://doi.org/10.1109/ICFPT59805.2023.00016

  84. [97]

    Barker, and Antonino Tumeo

    Cheng Tan, Chenhao Xie, Ang Li, Kevin J. Barker, and Antonino Tumeo. 2021. AURORA: Automated Refinement of Coarse-Grained Reconfigurable Accelera- tors. In 2021 Design, Automation & Test in Europe Conference & Exhibition (DATE). 1388–1393. https://doi.org/10.23919/DATE51398.20...

  85. [98]

    Hayden Kwok-Hay So and Cheng Liu. 2016. FPGA Overlays. Springer Interna- tional Publishing, Cham, 285–305. https://doi.org/10.1007/978-3-319-26408- 0_16

  86. [99]

    Dani Voitsechov and Yoav Etsion. 2014. Single-graph multiple flows: energy efficient design alternative for GPGPUs. In Proceeding of the 41st Annual Inter- national Symposium on Computer Architecuture (Minneapolis, Minnesota, USA) (ISCA ’14). IEEE Press, 205–216

  87. [100]

    Dani Voitsechov and Yoav Etsion. 2015. Control flow coalescing on a hybrid dataflow/von Neumann GPGPU. In 2015 48th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . 216–227. https://doi.org/10.1145/ 2830772.2830817

  88. [101]

    Dani Voitsechov, Oron Port, and Yoav Etsion. 2018. Inter-Thread Communication in Multithreaded, Reconfigurable Coarse-Grain Arrays. In 2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). 42–54. https: //doi.org/10.1109/MICRO.2018.00013

  89. [102]

    Bo Wang, Manupa Karunarathne, Aditi Kulkarni, Tulika Mitra, and Li-Shiuan Peh. 2019. HyCUBE: A 0.9V 26.4 MOPS/mW, 290 pJ/op, Power Efficient Accel- erator for IoT Applications. In 2019 IEEE Asian Solid-State Circuits Conference (A-SSCC). 133–136. https://doi.org/10.1109/A-SSCC...

  90. [103]

    Teng Wang, Lei Gong, Chao Wang, Yang Yang, Yingxue Gao, Xuehai Zhou, and Huaping Chen. 2022. ViA: A Novel Vision-Transformer Accelerator Based on FPGA. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 41, 11 (2022), 4088–4099. https://doi.org/10.11...

  91. [104]

    Swagath Venkataramani, Ashish Ranjan, Subarno Banerjee, Dipankar Das, Sasikanth Avancha, Ashok Jagannathan, Ajaya Durg, Dheemanth Nagaraj, Bharat Kaul, Pradeep Dubey, and Anand Raghunathan. 2017. SCALEDEEP: A scalable compute architecture for learning and evaluating deep netwo...

  92. [105]

    Xuechao Wei, Yun Liang, Tao Wang, Songwu Lu, and Jason Cong. 2017. Through- put optimization for streaming applications on CPU-FPGA heterogeneous sys- tems. In 2017 22nd Asia and South Pacific Design Automation Conference (ASP- DAC). 488–493. https://doi.org/10.1109/ASPDAC.201...

  93. [106]

    Xuechao Wei, Cody Hao Yu, Peng Zhang, Youxiang Chen, Yuxin Wang, Han Hu, Yun Liang, and Jason Cong. 2017. Automated Systolic Array Architecture Synthesis for High Throughput CNN Inference on FPGAs. In Proceedings of the 54th Annual Design Automation Conference 2017 (Austin, TX...

  94. [107]

    Jian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang, Preyas Shah, and Tony Nowatzki. 2020. DSAGEN: Synthesizing Programmable Spatial Accelerators. In 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). 268–281. https://doi.org/10.1109/ISCA45697.2020.00032

  95. [108]

    Jian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu, and Tony Nowatzki. 2020. A Hybrid Systolic-Dataflow Architecture for Inductive Matrix Algorithms. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). 703–716. https://doi.org/10.1109/HPCA475...

  96. [109]

    Emer, and Daniel Sanchez

    Yifan Yang, Joel S. Emer, and Daniel Sanchez. 2023. ISOSceles: Accelerating Sparse CNNs through Inter-Layer Pipelining. In 2023 IEEE International Sym- posium on High-Performance Computer Architecture (HPCA) . 598–610. https: //doi.org/10.1109/HPCA56546.2023.10071080

  97. [111]

    Yunxuan Yu, Chen Wu, Tiandong Zhao, Kun Wang, and Lei He. 2020. OPU: An FPGA-Based Overlay Processor for Convolutional Neural Networks. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 28, 1 (2020), 35–47. https://doi.org/10.1109/TVLSI.2019.2939726

  98. [113]

    Zhang, H

    B. Zhang, H. Zeng, and V. K. Prasanna. 2023. GraphAGILE: An FPGA-Based Overlay Accelerator for Low-Latency GNN Inference. IEEE Transactions on Parallel & Distributed Systems 34, 09 (2023), 2580–2597. https://doi.org/10.1109/ TPDS.2023.3287883

  99. [114]

    Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong. 2015. Optimizing FPGA-based Accelerator Design for Deep Convo- lutional Neural Networks. In Proceedings of the 2015 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (Monterey, Califo...

  100. [115]

    Chen Zhang, Guangyu Sun, Zhenman Fang, Peipei Zhou, Peichen Pan, and Jason Cong. 2019. Caffeine: Toward Uniformed Representation and Acceleration for Deep Convolutional Neural Networks. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 38, 11 (2019)...

  101. [116]

    Hanchen Ye, Xiaofan Zhang, Zhize Huang, Gengsheng Chen, and Deming Chen

  102. [117]

    Xiaofan Zhang, Hanchen Ye, Junsong Wang, Yonghua Lin, Jinjun Xiong, Wen- mei Hwu, and Deming Chen. 2020. DNNExplorer: a framework for modeling and exploring a novel paradigm of FPGA-based DNN accelerator. InProceedings of the 39th International Conference on Computer-Aided Des...

  103. [118]

    Shixuan Zheng, Xianjue Zhang, Leibo Liu, Shaojun Wei, and Shouyi Yin. 2022. Atomic Dataflow based Graph-Level Workload Orchestration for Scalable DNN Accelerators. In 2022 IEEE International Symposium on High-Performance Com- puter Architecture (HPCA). 475–489. https://doi.org...

  104. [119]

    Jinming Zhuang, Jason Lau, Hanchen Ye, Zhuoping Yang, Yubo Du, Jack Lo, Kristof Denolf, Stephen Neuendorffer, Alex Jones, Jingtong Hu, Deming Chen, Jason Cong, and Peipei Zhou. 2023. CHARM: Composing Heterogeneous Accel- erators for Matrix Multiply on Versal ACAP Architecture....

  105. [120]

    Jones, Jingtong Hu, Yiyu Shi, and Peipei Zhou

    Jinming Zhuang, Zhuoping Yang, Shixin Ji, Heng Huang, Alex K. Jones, Jingtong Hu, Yiyu Shi, and Peipei Zhou. 2024. SSR: Spatial Sequential Hybrid Architecture for Latency Throughput Tradeoff in Transformer Acceleration. In Proceedings of the 2024 ACM/SIGDA International Sympos...

  106. [121]

    Zhipeng Zhao, Joseph Melber, Siddharth Sahay, Shashank Obla, Eriko Nurvi- tadhi, James C. Hoe. 2023. Exploiting the Common Case When Accelerating Input-Dependent Stream Processing by FPGA. In IEEE Transactions on Comput- ers (TC ’23). IEEE. https://doi.org/10.1109/TC.2022.3200576

  107. [123]

    Xiaofan Zhang, Junsong Wang, Chao Zhu, Yonghua Lin, Jinjun Xiong, Wen-mei Hwu, and Deming Chen. 2018. DNNBuilder: an Automated Tool for Building High-Performance DNN Hardware Accelerators for FPGAs. In 2018 IEEE/ACM International Conference on Computer-Aided Design (ICCAD) . 1...

  108. [130]

    No Optimize

    Zhipeng Zhao, Nirav Atre, Hugo Sadok, Siddharth Sahay, Shashank Obla, James C. Hoe, Justine Sherry. 2022. Pigasus 2.0: making the pigasus IDS robust to attacks and different workloads. In In Proceedings of the SIGCOMM ’22 Poster and Demo Sessions (SIGCOMM ’22) . Association fo...

  109. [2019]

    Available: https://github.com/stanford- mast/nn_dataflow/tree/master/nn_dataflow/core, Accessed: 5-November-2024

    core directory in Tangram. Available: https://github.com/stanford- mast/nn_dataflow/tree/master/nn_dataflow/core, Accessed: 5-November-2024

  110. [2020]

    In Proceedings of the 57th ACM/EDAC/IEEE Design Automation Conference (Virtual Event, USA) (DAC ’20)

    HybridDNN: a framework for high-performance hybrid DNN accelerator design and implementation. In Proceedings of the 57th ACM/EDAC/IEEE Design Automation Conference (Virtual Event, USA) (DAC ’20). IEEE Press, Article 129, 6 pages

  111. [2021]

    In The 2021 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (Virtual Event, USA) (FPGA ’21)

    NPE: An FPGA-based Overlay Processor for Natural Language Processing. In The 2021 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (Virtual Event, USA) (FPGA ’21). Association for Computing Machinery, New York, NY, USA, 227. https://doi.org/10.1145/3431920.3439477

  112. [2023]

    ACM Trans

    FlexCNN: An End-to-end Framework for Composing CNN Accelerators on FPGA. ACM Trans. Reconfigurable Technol. Syst. 16, 2, Article 23 (March 2023), 32 pages. https://doi.org/10.1145/3570928

  113. [2024]

    In 2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO)

    SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts . In 2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE Computer Society, Los Alamitos, CA, USA, 1353–1366. https://doi.org/10.1109/MICRO61859.2024.00100

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.