REVIEW 3 major objections 6 minor 1 cited by
A-Graph: A Unified Graph Representation for At-Will Simulation across System Stacks
T0 review · 3 major / 6 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read A-Graph claims that representing a system as a weighted directed acyclic graph of events lets designers simulate performance and cost at any granularity across application, software, architecture, and circuit stacks.
desk verdict A credible unified-graph framing for pre-RTL DSE with real EDA validation, but the 'at-will' claim overreaches: temporal behavior lives in user code, granularity is chosen post hoc, and superconducting validation partly reproduces prior trends. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the weighted directed acyclic graph (WDAG): nodes are events (workload events, module events, and subevents), edges carry weights that count parent-to-child invocations, and acyclicity guarantees a deterministic topological order for metric aggregation. Three aggregation patterns—module, summation, and specified (sequential versus parallel)—define how metrics propagate upward, and scope-based metric retrieval lets the same graph answer questions about a single module, a tagged group such as a processing element, an event, or the whole workload. The paper also introduces a constraint-graph-based front end that generates and sweeps design points from user code, and a
What would settle it
Run Archx on a fixed CMOS GEMM systolic array with leaf nodes set at processing-element level and at submodule level, then compare reported area and dynamic energy against a full place-and-route EDA flow for the same design. If the submodule version underestimates dynamic energy by more than about 30% for a 32x32 array while the processing-element version matches, the claim that users can simulate accurately at any chosen granularity is not self-fulfilling without explicit guidance on choosing leaf granularity.
Extended reading notes
Core claim
The paper's central claim is that at-will simulation is achievable: a user-defined weighted DAG of events, spanning application, software, architecture, and circuit, lets a designer simulate performance and cost at any granularity with any metric, and is agnostic to technology, architecture, and application. Events can be as coarse as a full workload or as fine as an individual register; edge weights count how many times a subevent is invoked by its parent, and metrics are computed by topologically traversing the graph and aggregating leaf-node values according to three patterns: module (direct leaf metric), summation (additive across edges), and specified (sequential sum or parallel max). B
Load-bearing premise
The central claim collapses if a user's performance model for an event does not correctly encode the spatial and temporal dependencies inside that event, or if the chosen leaf-module granularity's precomputed circuit data stops being valid when composed; the paper concedes in Section VII.B that Archx relies on user expertise to maintain proper spatial and temporal relationships, and its systolic-array study shows errors up to about 30% when the granularity is too fine.
Editorial extensions
If this is right
- The same A-Graph specification can describe and simulate CMOS and superconducting designs, so early design-space exploration can compare fundamentally different technologies before committing to a design flow.
- Users can define new metrics, such as area in Josephson Junctions instead of square millimeters, by registering a metric name, unit, and aggregation pattern; no simulator rewrite is needed.
- Because metric retrieval is scope-based, the same design point can be analyzed at the module, processing-element, event, or workload level, giving designers hierarchical visibility into where cost comes from.
- The reported simulation speedup over full EDA flows is large (up to roughly 10^5 times faster for pure simulation), which makes broad sweeps of the design space practical before RTL.
- Accuracy of the composed graph depends on leaf-module granularity: the systolic-array study shows that choosing submodules instead of whole processing elements can underestimate dynamic energy by roughly 20–30% at larger array sizes, so practical use requires careful granularity choice or an ensemble of module databases.
Reading between the lines
- If A-Graph were widely adopted, a natural next step would be shared, validated libraries of event decompositions and module databases, letting cross-stack co-design become a matter of composing vetted nodes rather than hand-tuning simulators.
- The metric-aggregation mechanism is general enough that it could host non-hardware metrics—energy-delay product, thermal budget, cost per wafer, carbon footprint—as long as they are additive or max-composed along the same dependency graph; the paper gestures at this but does not develop it.
- A testable extension is automatic granularity selection: because the paper shows that leaf-node granularity changes accuracy by up to about 30%, a wrapper could search over granularity by predicting wiring and fanout effects, turning at-will granularity into well-chosen granularity.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes A-Graph, a weighted directed acyclic graph intended to unify application, software, architecture, and circuit abstractions into a single representation, together with Archx, a Python framework that implements A-Graph for design space exploration. A-Graph uses event nodes, weighted dependency edges, and a small set of aggregation patterns (module, summation, specified sequential/parallel) to compute metrics such as area, power, energy, cycle count, and runtime. Archx adds a front-end programming interface, automatic sweeping under user constraints, and scope-based metric retrieval. The authors validate the approach through five case studies: CMOS FFT and systolic arrays and a TNN column against a full EDA flow, and superconducting FIR and CNN arrays against published results. The central claim is that A-Graph enables 'at-will simulation' with high accuracy across arbitrary technologies, architectures, applications, and granularities.
Significance. If the claims were fully supported, this would be a useful contribution: a single pre-RTL representation spanning multiple system stacks, with pattern-based metric aggregation and flexible user-defined metrics, is more general than domain-specific simulators such as Aladdin, Accelergy, or DSAGEN. The CMOS validation against a genuine EDA flow (synthesis and place-and-route) is a strength, and the two superconducting case studies demonstrate the intended technology flexibility. The paper also has a concrete artifact, Archx, and the scope-based retrieval idea is genuinely convenient for hierarchical analysis. However, the conceptual and empirical support for the strongest claims is incomplete: the graph itself does not encode temporal dependencies, the accuracy of the examples depends on post-hoc granularity choices, and the superconducting FIR validation is not independent. These issues are addressable in revision but currently prevent me from recommending acceptance.
major comments (3)
- [Sections III.A.2, III.C, IV.B.1, VII.B] The central claim that cross-stack changes can be captured 'by updating nodes and edges properly' is not supported by the design as described. Section III.A.2 states that 'weighted edges do not reveal temporal dependencies between events' and that temporal dependencies are 'captured inside the performance models of the parent event.' Section IV.B.1 confirms that these performance models are arbitrary user-written Python functions that compute edge weights once before graph traversal, and Section III.C provides only three static aggregation patterns (module, summation, specified). Thus dynamic behavior such as pipeline stalls, backpressure, data-dependent timing, and contention is pre-collapsed into scalar values by the user's model rather than represented in the WDAG. The paper's own limitation section (VII.B) concedes that Archx must 'rely on user expertise to maintain proper spatial an
- [Section VI.B, Figures 8 and 10] The claimed 'high accuracy' is strongly dependent on the user's choice of leaf-module granularity, and there is no principled way to choose it a priori. In the FFT study (Figure 8), the PE granularity caps errors at 15%, while PE-Sub granularity gives errors that grow with array size. In the systolic array study (Figure 10), the PE implementation is better at small array sizes but PE-Sub is better at larger, with dynamic energy errors reaching -30.8%. The paper suggests using 'an ensemble of module databases' but does not provide a method for selecting or combining granularities without access to ground-truth EDA results. Since 'simulate at any granularity, thus accurately, at their own will' is a core claim, this user-dependence must be addressed directly, for example by automatic granularity selection, error bounds as a function of granularity, or a concrete criterion based on wiring c
- [Section VI.C.1 and Table IV] The FIR superconducting case study is not an independent validation. The text explicitly says 'Without access to the original throughput and area values, we reproduce their reported trends' (Section VI.C.1), so Figure 12 demonstrates only that the Archx model can be tuned to reproduce trends, not that it predicts unknown values. More concerning, Table IV reports leakage power of 8.4 mW with a relative error of 1.81e-14, exactly matching the [27] baseline; this is effectively fitting to the target. Please either obtain the original data and report true prediction errors, or relabel the FIR study as a qualitative/functional reproduction and base the cross-technology 'high accuracy' claim on the CNN case study, where actual baseline numbers are available. As written, the superconducting validation overstates the evidence.
minor comments (6)
- [Listing 1, line 17] The code sample has unbalanced brackets/parentheses: `param_value=[[2, 2], [4, 4], sweep=True)` is not valid Python. This appears to be a typo and should be fixed.
- [Listing 2, lines 14-15] The performance model snippet shows a malformed dictionary entry (`'runtime': 2}}` and an apparent brace mismatch). Please ensure the listings compile or use ellipses consistently.
- [Section VII.B] The text says 'Like Alladin [65]' but the correct name is 'Aladdin.' Also, the author affiliation city is spelled 'Pittsburg' in the byline; it should be 'Pittsburgh.'
- [References] References [77] and [78] appear to be duplicate versions of the same TNNGen paper. Please check the bibliography and cite each source once.
- [Figure 7] The x-axis is hard to read: groups of array sizes are listed without separators, and the legend entries (A-Sub, A-PE, EDA-F, EDA-Sub, EDA-PE) are not clearly tied to the bar groups. A table or clearer axis labels would help.
- [Section II.C] There is a typo: 'V on Neumann' should be 'von Neumann.' Please proofread the manuscript for similar minor errors.
Circularity Check
Superconducting validation reduces to fitting self-authored baselines; CMOS comparisons are independent.
-
fitted input called prediction
[Section VI.C.1, 'Power study', Table IV]
"TABLE IV: FIR power results. Metric Archx Baseline Relative error (%) Dynamic power (µW) 8.125 8.4 -3.27 Leakage power (mW) 8.4 8.4 1.81×10 −14 ... yielding nearly identical results to the baseline."
The Archx leakage power is reported as 8.4 mW with a relative error of 1.81e-14 against the baseline, i.e., identical to the reference value to machine precision. Since the paper explicitly states it had no access to the original throughput and area values yet reproduces the baseline trends, the leakage-power 'prediction' appears to be the baseline value itself, not an independent derivation. The validation for FIR power therefore reduces to fitting the target and renaming it as a prediction.
-
fitted input called prediction
[Section VI.C.2, 'Convolutional Neural Network', Table V]
"Our results match the baseline across all metrics other than area. In place of JTL chains, common for long connections, the baseline utilizes passive transmission lines (PTLs) as overhead. Including this overhead reduces the error from 15% to 3%, emphasizing the flexibility."
The area prediction is corrected by adding a PTL-overhead component whose existence and magnitude are taken from the baseline being validated against. The post-hoc inclusion of this overhead to reduce the error from 15% to 3% means the final reported accuracy is achieved by adding the target's known design feature, rather than by prediction from the graph alone.
1 more flagged steps
-
self citation load bearing
[Section VI.C, superconducting validation; references [27], [28]]
"We validate superconducting through a finite impulse response (FIR) tag array [27] and a convolutional neural network (CNN) PE array [28]. ... Our results yield nearly identical lines for the U-SFQ 32- and 256-tap FIR arrays (U 32, U 256) from [27]."
The ground truth for the superconducting case studies is the authors' own prior work ([27], [28]), and the goal is to reproduce that work. Because the validation target is provided by the same research group, the demonstration that A-Graph 'generalizes' to superconducting reduces to matching self-produced numbers, with no independent external benchmark. This is load-bearing because the only evidence for superconducting accuracy is agreement with these self-citations.
full rationale
The A-Graph framework itself is not inherently circular: the CSS case studies (FFT, systolic array, TNN) are validated against an external full EDA flow (Cadence Genus/Innovus), and the composition errors (e.g., PE-Sub granularity up to 30%) are honestly reported. The central graph/aggregation mechanism is well-specified and neither self-definitional nor a renaming of a known result. However, the superconducting validation contains clear circular elements. In the FIR power study, the reported leakage power equals the baseline to 1.81e-14 relative error, and the text admits that only trends were reproduced without access to original values—so the 'prediction' is effectively the baseline value itself. In the CNN study, the area error is reduced from 15% to 3% only after adding the baseline's PTL overhead, a post-hoc fit to the target. Both superconducting ground truths are the authors' own prior publications, making the validation self-referential. These issues are confined to the superconducting case studies; the CMOS benchmarks provide independent evidence that the framework can compose precomputed module data with reasonable accuracy. Overall, partial circularity in the validation of the superconducting claims, but the central framework retains independent content.
Assumptions & free parameters
free parameters (3)
- Leaf-module granularity per case study (PE vs PE-Sub) =
chosen per configuration; e.g., FFT uses PE, systolic switches preference by array size
- Performance-model edge factors (e.g., cycle_count factor=2) =
2 in Listing 2
- Superconducting FIR leakage power baseline =
8.4 mW (identical to [27] baseline)
assumptions (5)
- domain assumption A weighted DAG with multiplicative edge weights and three aggregation patterns (module, summation, specified) is sufficient to represent the performance and cost of arbitrary systems.
- domain assumption Temporal dependencies can be fully captured inside parent-event performance models rather than in graph edges.
- domain assumption Precomputed circuit-level database entries remain valid when composed into larger designs.
- domain assumption External baselines (Cadence EDA, WRSPICE, prior U-SFQ papers) are correct ground truth for validation.
- standard math The graph-tool library and Python front-end correctly implement graph construction and traversal.
Cite this review
Pith. "Pith review of A-Graph: A Unified Graph Representation for At-Will Simulation across System Stacks." pith.science (2026). https://pith.science/paper/EMLNNQ5A
@misc{pith2026260204847,
author = {Pith},
title = {Pith review of: A-Graph: A Unified Graph Representation for At-Will Simulation across System Stacks},
year = {2026},
howpublished = {\url{https://pith.science/paper/EMLNNQ5A}},
note = {Machine review of arXiv:2602.04847}
}
read the original abstract
As computer systems continue to diversify across technologies, architectures, applications, and beyond, the relevant design space has become larger and more complex. Given such trends, design space exploration (DSE) at early stages is critical to ensure agile development towards optimal performance and cost. Industry-grade EDA tools directly take in RTL code and report accurate results, but do not perform DSE. Recent works have attempted to explore the design space via simulation. However, most of these works are domain-specific and constrain the space that users are allowed to explore, offering limited flexibility between technologies, architecture, and applications. Moreover, they often demand high domain expertise to ensure high accuracy. To enable simulation that is agnostic to technology, architecture, and application at any granularity, we introduce Architecture-Graph (Agraph), a graph that unifies the system representation surrounding any arbitrary application, software, architecture, and circuit. Such a unified representation distinguishes Agraph from prior works, which focus on a single stack, allowing users to freely explore the design space across system stacks. To fully unleash the potential of Agraph, we further present Archx, a framework that implements Agraph. Archx is user-friendly in two ways. First, Archx has an easy-to-use programming interface to automatically generate and sweep design points under user constraints, boosting the programmability. Second, Archx adopts scope-based metric retrieval to analyze and understand each design point at any user-preferred hierarchy, enhancing the explainability. We conduct case studies that demonstrate Agraph's generalization across technologies, architecture, and applications with high simulation accuracy. Overall, we argue that Agraph and Archx serve as a foundation to simulate both performance and cost at will.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 1 Pith paper
-
Lottery BP: Unlocking Quantum Error Decoding at Scale
Lottery BP adds randomness to belief propagation decoding and uses syndrome voting to achieve far higher accuracy on topological quantum codes while reducing reliance on expensive global decoders.
Reference graph
Works this paper leans on
-
[28]
Toward practical superconducting accelerators for machine learning using u-sfq,
P. Gonzalez-Guerrero, K. Huch, N. Patra, T. Popovici, and G. Michelogiannakis, “Toward practical superconducting accelerators for machine learning using u-sfq,”J. Emerg. Technol. Comput. Syst., vol. 20, no. 2, Jun. 2024. [Online]. Available: https://doi.org/10.1145/3653073
-
[27]
Temporal and sfq pulse-streams encoding for area-efficient superconducting accelerators,
P. Gonzalez-Guerrero, M. G. Bautista, D. Lyles, and G. Michelogiannakis, “Temporal and sfq pulse-streams encoding for area-efficient superconducting accelerators,” inProceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, ser. ASPLOS ’22. New York, NY , USA: Association for Computing M...
arXiv 2022
-
[1]
3-d memristor crossbars for ana- log and neuromorphic computing applications,
G. C. Adam, B. D. Hoskins, M. Prezioso, F. Merrikh-Bayat, B. Chakrabarti, and D. B. Strukov, “3-d memristor crossbars for ana- log and neuromorphic computing applications,”IEEE Transactions on Electron Devices, vol. 64, no. 1, pp. 312–318, 2016
2016
-
[2]
Universal photonic artificial intelligence acceleration,
S. R. Ahmed, R. Baghdadi, M. Bernadskiy, N. Bowman, R. Braid, J. Carr, C. Chen, P. Ciccarella, M. Cole, J. Cookeet al., “Universal photonic artificial intelligence acceleration,”Nature, vol. 640, no. 8058, pp. 368–374, 2025
2025
-
[3]
A manufacturable platform for photonic quantum computing,
K. Alexander, A. Bahgat, A. Benyamini, D. Black, D. Bonneau, S. Bur- gos, B. Burridge, G. Campbell, G. Catalano, A. Ceballoset al., “A manufacturable platform for photonic quantum computing,”Nature, vol. 641, no. 8064, pp. 876–883, 2025
2025
-
[4]
Control flow analysis,
F. E. Allen, “Control flow analysis,”ACM Sigplan Notices, vol. 5, no. 7, pp. 1–19, 1970
1970
-
[5]
Theory of superconducting tc,
P. B. Allen and B. Mitrovi ´c, “Theory of superconducting tc,”Solid state physics, vol. 37, pp. 1–92, 1983
1983
-
[6]
Resparc: A reconfigurable and energy-efficient architecture with memristive crossbars for deep spiking neural networks,
A. Ankit, A. Sengupta, P. Panda, and K. Roy, “Resparc: A reconfigurable and energy-efficient architecture with memristive crossbars for deep spiking neural networks,” inProceedings of the 54th Annual Design Automation Conference 2017, 2017, pp. 1–6
2017
Show all 90 references
-
[7]
Human action recognition with a large-scale brain-inspired photonic computer,
P. Antonik, N. Marsal, D. Brunner, and D. Rontani, “Human action recognition with a large-scale brain-inspired photonic computer,”Nature Machine Intelligence, vol. 1, no. 11, pp. 530–537, 2019
2019
-
[8]
CACTI 7: New Tools for Interconnect Exploration in Innovative Off-Chip Memories,
R. Balasubramonian, A. B. Kahng, N. Muralimanohar, A. Shafiee, and V . Srinivas, “CACTI 7: New Tools for Interconnect Exploration in Innovative Off-Chip Memories,”Transactions on Architecture and Code Optimization, 2017
2017
-
[9]
The gem5 simulator,
N. Binkert, B. Beckmann, G. Black, S. K. Reinhardt, A. Saidi, A. Basu, J. Hestness, D. R. Hower, T. Krishna, S. Sardashtiet al., “The gem5 simulator,”ACM SIGARCH computer architecture news, vol. 39, no. 2, pp. 1–7, 2011
2011
-
[10]
Quantum simulations with trapped ions,
R. Blatt and C. F. Roos, “Quantum simulations with trapped ions,” Nature Physics, vol. 8, no. 4, pp. 277–284, 2012
2012
-
[11]
Bfloat16 processing for neural networks,
N. Burgess, J. Milanovic, N. Stephens, K. Monachopoulos, and D. Mansell, “Bfloat16 processing for neural networks,”IEEE Xplore, 2019
2019
-
[12]
New application of superconductors: High sensitivity cryogenic light detectors,
L. Cardani, F. Bellini, N. Casali, M. Castellano, I. Colantoni, A. Cop- polecchia, C. Cosmelli, A. Cruciani, A. D’Addabbo, S. Di Domizio et al., “New application of superconductors: High sensitivity cryogenic light detectors,”Nuclear Instruments and Methods in Physics Research...
2017
-
[13]
Phoenixsim: A simulator for physical-layer analysis of chip-scale photonic interconnection networks,
J. Chan, G. Hendry, A. Biberman, K. Bergman, and L. P. Carloni, “Phoenixsim: A simulator for physical-layer analysis of chip-scale photonic interconnection networks,” in2010 Design, Automation & Test in Europe Conference & Exhibition (DATE 2010). IEEE, 2010, pp. 691–696
2010
-
[14]
A ferroelectric memristor,
A. Chanthbouala, V . Garcia, R. O. Cherifi, K. Bouzehouane, S. Fusil, X. Moya, S. Xavier, H. Yamada, C. Deranlot, N. D. Mathuret al., “A ferroelectric memristor,”Nature materials, vol. 11, no. 10, pp. 860–864, 2012
2012
-
[15]
Unsupervised clustering of time series signals using neuromorphic energy-efficient temporal neural networks,
S. Chaudhari, H. Nair, J. M. Moura, and J. P. Shen, “Unsupervised clustering of time series signals using neuromorphic energy-efficient temporal neural networks,” inICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, ...
2021
-
[16]
Rapid sin- gle flux quantum t-flip flop operating up to 770 ghz,
W. Chen, A. Rylyakov, V . Patel, J. Lukens, and K. Likharev, “Rapid sin- gle flux quantum t-flip flop operating up to 770 ghz,”IEEE Transactions on Applied Superconductivity, vol. 9, no. 2, pp. 3212–3215, 1999
1999
-
[17]
Memristor-the missing circuit element,
L. Chua, “Memristor-the missing circuit element,”IEEE Transactions on circuit theory, vol. 18, no. 5, pp. 507–519, 2003
2003
-
[18]
Superconducting quantum bits,
J. Clarke and F. K. Wilhelm, “Superconducting quantum bits,”Nature, vol. 453, no. 7198, pp. 1031–1042, 2008
2008
-
[19]
What is the fast fourier transform?
W. Cochran, J. Cooley, D. Favin, H. Helms, R. Kaenel, W. Lang, G. Maling, D. Nelson, C. Rader, and P. Welch, “What is the fast fourier transform?”Proceedings of the IEEE, vol. 55, no. 10, pp. 1664–1674, 1967
1967
-
[20]
Charm: A language for closed-form high-level architecture modeling,
W. Cui, Y . Ding, D. Dangwal, A. Holmes, J. McMahan, A. Javadi- Abhari, G. Tzimpragos, F. Chong, and T. Sherwood, “Charm: A language for closed-form high-level architecture modeling,” in2018 ACM/IEEE 45th Annual International Symposium on Computer Archi- tecture (ISCA), 2018, ...
2018
-
[21]
Superneuro: A fast and scalable simulator for neuromorphic computing,
P. Date, C. Gunaratne, S. R. Kulkarni, R. Patton, M. Coletti, and T. Potok, “Superneuro: A fast and scalable simulator for neuromorphic computing,” inProceedings of the 2023 International Conference on Neuromorphic Systems, ser. ICONS ’23. New York, NY , USA: Association for C...
2023
-
[22]
Josim—superconductor spice simulator,
J. A. Delport, K. Jackman, P. Le Roux, and C. J. Fourie, “Josim—superconductor spice simulator,”IEEE Transactions on Applied Superconductivity, vol. 29, no. 5, pp. 1–5, 2019
2019
-
[23]
High-efficiency single- photon source above the loss-tolerant threshold for efficient linear optical quantum computing,
X. Ding, Y .-P. Guo, M.-C. Xu, R.-Z. Liu, G.-Y . Zou, J.-Y . Zhao, Z.-X. Ge, Q.-H. Zhang, H.-L. Liu, L.-J. Wanget al., “High-efficiency single- photon source above the loss-tolerant threshold for efficient linear optical quantum computing,”Nature Photonics, pp. 1–5, 2025
2025
-
[24]
Partial coherence enhances parallelized photonic computing,
B. Dong, F. Br ¨uckerhoff-Pl¨uckelmann, L. Meyer, J. Dijkstra, I. Bente, D. Wendland, A. Varri, S. Aggarwal, N. Farmakidis, M. Wanget al., “Partial coherence enhances parallelized photonic computing,”Nature, vol. 632, no. 8023, pp. 55–62, 2024
2024
-
[25]
Photonic computing using the modified signed-digit number representation,
B. L. Drake, R. P. Bocker, M. E. Lasher, R. H. Patterson, and W. J. Miceli, “Photonic computing using the modified signed-digit number representation,”Optical Engineering, vol. 25, no. 1, pp. 38–43, 1986
1986
-
[26]
Spiking neural networks,
S. Ghosh-Dastidar and H. Adeli, “Spiking neural networks,”Interna- tional journal of neural systems, vol. 19, no. 04, pp. 295–308, 2009
2009
-
[29]
Noise injection adaption: End-to-end reram crossbar non-ideal effect adaption for neural network mapping,
Z. He, J. Lin, R. Ewetz, J.-S. Yuan, and D. Fan, “Noise injection adaption: End-to-end reram crossbar non-ideal effect adaption for neural network mapping,” inProceedings of the 56th Annual Design Automation Conference 2019, ser. DAC ’19. New York, NY , USA: Association for Co...
2019
-
[30]
A new golden age for computer architecture,
J. L. Hennessy and D. A. Patterson, “A new golden age for computer architecture,”Commun. ACM, vol. 62, no. 2, p. 48–60, Jan. 2019. [Online]. Available: https://doi.org/10.1145/3282307
2019 doi
-
[31]
Nvidia 12 simnet™: An ai-accelerated multi-physics simulation framework,
O. Hennigh, S. Narasimhan, M. A. Nabian, A. Subramaniam, K. Tangsali, Z. Fang, M. Rietmann, W. Byeon, and S. Choudhry, “Nvidia 12 simnet™: An ai-accelerated multi-physics simulation framework,” in International conference on computational science. Springer, 2021, pp. 447–461
2021
-
[32]
Ultra- low-power superconductor logic,
Q. P. Herr, A. Y . Herr, O. T. Oberg, and A. G. Ioannidis, “Ultra- low-power superconductor logic,”Journal of Applied Physics, vol. 109, no. 10, May 2011. [Online]. Available: http://dx.doi.org/10.1063/ 1.3585849
2011
-
[33]
Cosa: scheduling by ¡u¿c¡/u¿onstrained ¡u¿o¡/u¿ptimization for ¡u¿s¡/u¿patial ¡u¿a¡/u¿ccelerators,
Q. Huang, M. Kang, G. Dinh, T. Norell, A. Kalaiah, J. Demmel, J. Wawrzynek, and Y . S. Shao, “Cosa: scheduling by ¡u¿c¡/u¿onstrained ¡u¿o¡/u¿ptimization for ¡u¿s¡/u¿patial ¡u¿a¡/u¿ccelerators,” in Proceedings of the 48th Annual International Symposium on Computer Architecture,...
2021
-
[34]
Nangate Open Cell Library: 45nm open cell standard cell library,
N. Inc., “Nangate Open Cell Library: 45nm open cell standard cell library,” 2009, available at https://www.nangate.com
2009
-
[35]
Demonstration of a neutral atom controlled- not quantum gate,
L. Isenhower, E. Urban, X. Zhang, A. Gill, T. Henage, T. A. Johnson, T. Walker, and M. Saffman, “Demonstration of a neutral atom controlled- not quantum gate,”Physical review letters, vol. 104, no. 1, p. 010503, 2010
2010
-
[36]
Highly accurate protein structure prediction with alphafold,
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. ˇZ´ıdek, A. Potapenkoet al., “Highly accurate protein structure prediction with alphafold,”nature, vol. 596, no. 7873, pp. 583–589, 2021
2021
-
[37]
Advanced SFQ5ee process at MIT Lincoln Laboratory,
A. F. Kirichenko, S. Sarwana, and D. Gupta, “Advanced SFQ5ee process at MIT Lincoln Laboratory,” inProceedings of the IEEE International Superconductive Electronics Conference (ISEC), 2003, pp. K4–1
2003
-
[38]
Understanding reuse, performance, and hardware cost of DNN dataflow: A data-centric approach,
H. Kwon, P. Chatarasi, M. Pellauer, A. Parashar, V . Sarkar, and T. Kr- ishna, “Understanding reuse, performance, and hardware cost of DNN dataflow: A data-centric approach,” inProceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture, MICRO. ACM, 20...
2019
-
[39]
Sustain- able memristors from shiitake mycelium for high-frequency bioelectron- ics,
J. LaRocco, Q. Tahmina, R. Petreaca, J. Simonis, and J. Hill, “Sustain- able memristors from shiitake mycelium for high-frequency bioelectron- ics,”PLoS One, vol. 20, no. 10, p. e0328965, 2025
2025
-
[40]
Analogue signal and image processing with large memristor crossbars,
C. Li, M. Hu, Y . Li, H. Jiang, N. Ge, E. Montgomery, J. Zhang, W. Song, N. D ´avila, C. E. Graveset al., “Analogue signal and image processing with large memristor crossbars,”Nature electronics, vol. 1, no. 1, pp. 52–59, 2018
2018
-
[41]
Rsfq logic/memory family: a new josephson-junction technology for sub-terahertz-clock-frequency digital systems,
K. Likharev and V . Semenov, “Rsfq logic/memory family: a new josephson-junction technology for sub-terahertz-clock-frequency digital systems,”IEEE Transactions on Applied Superconductivity, vol. 1, no. 1, pp. 3–28, 1991
1991
-
[42]
All-optical machine learning using diffractive deep neural networks,
X. Lin, Y . Rivenson, N. T. Yardimci, M. Veli, Y . Luo, M. Jarrahi, and A. Ozcan, “All-optical machine learning using diffractive deep neural networks,”Science, vol. 361, no. 6406, pp. 1004–1008, 2018
2018
-
[43]
Catwalk: Unary top-k for efficient ramp-no-leak neuron design for temporal neural networks,
D. Lister, P. Vellaisamy, J. P. Shen, and D. Wu, “Catwalk: Unary top-k for efficient ramp-no-leak neuron design for temporal neural networks,” in2025 IEEE Computer Society Annual Symposium on VLSI (ISVLSI), vol. 1, 2025, pp. 1–6
2025
-
[44]
A spiking neuromorphic design with resistive crossbar,
C. Liu, B. Yan, C. Yang, L. Song, Z. Li, B. Liu, Y . Chen, H. Li, Q. Wu, and H. Jiang, “A spiking neuromorphic design with resistive crossbar,” in Proceedings of the 52nd Annual Design Automation Conference, 2015, pp. 1–6
2015
-
[45]
Overgen: Improving fpga usability through domain-specific overlay generation,
S. Liu, J. Weng, D. Kupsh, A. Sohrabizadeh, Z. Wang, L. Guo, J. Liu, M. Zhulin, R. Mani, L. Zhang, J. Cong, and T. Nowatzki, “Overgen: Improving fpga usability through domain-specific overlay generation,” in2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO...
2022
-
[46]
Camj: Enabling system- level energy modeling and architectural exploration for in-sensor visual computing,
T. Ma, Y . Feng, X. Zhang, and Y . Zhu, “Camj: Enabling system- level energy modeling and architectural exploration for in-sensor visual computing,” inProceedings of the 50th Annual International Symposium on Computer Architecture, 2023, pp. 1–14
2023
-
[47]
Uncon- ventional compute methods and future challenges for superconducting digital computing,
G. Michelogiannakis, A. Butko, P. Gonzalez-Guerrero, D. Vasudevan, M. Gay Bautista-Jurney, C. Grace, P. Zarkos, and J. Shalf, “Uncon- ventional compute methods and future challenges for superconducting digital computing,”Frontiers in Materials, vol. 12, p. 1618615, 2025
2025
-
[48]
Energy-efficient single flux quantum technology,
O. A. Mukhanov, “Energy-efficient single flux quantum technology,” IEEE Transactions on Applied Superconductivity, vol. 21, no. 3, pp. 760–769, 2011
2011
-
[49]
ASAP7: A 7-nm finfet predictive process design kit,
S. Naffziger, S.-C. Yang, D. J. Carlberg, B. Sarkar, S. K. Sinha, G. Yeric, and M. Haycock, “ASAP7: A 7-nm finfet predictive process design kit,” inProceedings of the 54th Annual Design Automation Conference (DAC). ACM, 2017, pp. 1–6
2017
-
[50]
Direct cmos implementation of neuromorphic temporal neural networks for sensory processing,
H. Nair, J. P. Shen, and J. E. Smith, “Direct cmos implementation of neuromorphic temporal neural networks for sensory processing,”arXiv preprint arXiv:2009.00457, 2020
2009 arXiv
-
[51]
Tnn7: A custom macro suite for implementing highly optimized designs of neuromorphic tnns,
H. Nair, P. Vellaisamy, S. Bhasuthkar, and J. P. Shen, “Tnn7: A custom macro suite for implementing highly optimized designs of neuromorphic tnns,” in2022 IEEE Computer Society Annual Symposium on VLSI (ISVLSI). IEEE, 2022, pp. 152–157
2022
-
[52]
Ros´e: A hardware-software co-simulation infrastructure enabling pre-silicon full-stack robotics soc evaluation,
D. Nikiforov, S. C. Dong, C. L. Zhang, S. Kim, B. Nikolic, and Y . S. Shao, “Ros´e: A hardware-software co-simulation infrastructure enabling pre-silicon full-stack robotics soc evaluation,” inProceedings of the 50th Annual International Symposium on Computer Architecture, 202...
2023
-
[53]
Optical quantum computing,
J. L. O’brien, “Optical quantum computing,”Science, vol. 318, no. 5856, pp. 1567–1570, 2007
2007
-
[54]
Carat: Unlocking Value-Level Parallelism for Multiplier-Free GEMMs,
Z. Pan, J. San Miguel, and D. Wu, “Carat: Unlocking Value-Level Parallelism for Multiplier-Free GEMMs,” inInternational Conference on Architectural Support for Programming Languages and Operating Systems, 2024
2024
-
[55]
Timeloop: A systematic approach to dnn accelerator evaluation,
A. Parashar, P. Raina, Y . S. Shao, Y .-H. Chen, V . A. Ying, A. Mukkara, R. Venkatesan, B. Khailany, S. W. Keckler, and J. Emer, “Timeloop: A systematic approach to dnn accelerator evaluation,” in2019 IEEE International Symposium on Performance Analysis of Systems and Softwar...
2019
-
[56]
The graph-tool python library,
T. P. Peixoto, “The graph-tool python library,”figshare, 2014. [Online]. Available: http://figshare.com/articles/graph tool/1164194
2014
-
[57]
Demonstration of the trapped-ion quantum ccd computer architecture,
J. M. Pino, J. M. Dreiling, C. Figgatt, J. P. Gaebler, S. A. Moses, M. Allman, C. Baldwin, M. Foss-Feig, D. Hayes, K. Mayeret al., “Demonstration of the trapped-ion quantum ccd computer architecture,” Nature, vol. 592, no. 7853, pp. 209–213, 2021
2021
-
[58]
Exploring neuromorphic computing based on spiking neural networks: Algorithms to hardware,
N. Rathi, I. Chakraborty, A. Kosta, A. Sengupta, A. Ankit, P. Panda, and K. Roy, “Exploring neuromorphic computing based on spiking neural networks: Algorithms to hardware,”ACM Computing Surveys, vol. 55, no. 12, pp. 1–49, 2023
2023
-
[59]
In-memory computing on a photonic platform,
C. R ´ıos, N. Youngblood, Z. Cheng, M. Le Gallo, W. H. Pernice, C. D. Wright, A. Sebastian, and H. Bhaskaran, “In-memory computing on a photonic platform,”Science advances, vol. 5, no. 2, p. eaau5759, 2019
2019
-
[60]
The structural simulation toolkit,
A. F. Rodrigues, K. S. Hemmert, B. W. Barrett, C. Kersey, R. Oldfield, M. Weston, R. Risen, J. Cook, P. Rosenfeld, E. Cooper-Baliset al., “The structural simulation toolkit,”ACM SIGMETRICS Performance Evaluation Review, vol. 38, no. 4, pp. 37–42, 2011
2011
-
[61]
Dramsim2: A cycle accurate memory system simulator,
P. Rosenfeld, E. Cooper-Balis, and B. Jacob, “Dramsim2: A cycle accurate memory system simulator,”IEEE computer architecture letters, vol. 10, no. 1, pp. 16–19, 2011
2011
-
[62]
Why i am optimistic about the silicon-photonic route to quantum computing,
T. Rudolph, “Why i am optimistic about the silicon-photonic route to quantum computing,”APL photonics, vol. 2, no. 3, 2017
2017
-
[63]
Savitskii,Superconducting materials
E. Savitskii,Superconducting materials. Springer Science & Business Media, 2012
2012
-
[64]
Neutral atom quantum register,
D. Schrader, I. Dotsenko, M. Khudaverdyan, Y . Miroshnychenko, A. Rauschenbeutel, and D. Meschede, “Neutral atom quantum register,” Physical Review Letters, vol. 93, no. 15, p. 150501, 2004
2004
-
[65]
Aladdin: A pre-rtl, power-performance accelerator simulator enabling large design space exploration of customized architectures,
Y . S. Shao, B. Reagen, G.-Y . Wei, and D. Brooks, “Aladdin: A pre-rtl, power-performance accelerator simulator enabling large design space exploration of customized architectures,” in2014 ACM/IEEE 41st International Symposium on Computer Architecture (ISCA), 2014, pp. 97–108
2014
-
[66]
Photonics for artificial intelligence and neuromorphic computing,
B. J. Shastri, A. N. Tait, T. Ferreira de Lima, W. H. Pernice, H. Bhaskaran, C. D. Wright, and P. R. Prucnal, “Photonics for artificial intelligence and neuromorphic computing,”Nature Photonics, vol. 15, no. 2, pp. 102–114, 2021
2021
-
[67]
Space-time computing with temporal neural networks,
J. E. Smith, “Space-time computing with temporal neural networks,” Synthesis Lectures on Computer Architecture, vol. 12, no. 2, pp. i–215, 2017
2017
-
[68]
Space-time algebra: A model for neocortical computation,
——, “Space-time algebra: A model for neocortical computation,” in 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). IEEE, 2018, pp. 289–300
2018
-
[69]
A temporal neural network architecture for online learning,
——, “A temporal neural network architecture for online learning,” arXiv preprint arXiv:2011.13844, 2020
2011 arXiv
-
[70]
A macrocolumn architecture implemented with temporal (spik- ing) neurons,
——, “A macrocolumn architecture implemented with temporal (spik- ing) neurons,”arXiv preprint arXiv:2207.05081, 2022
2022 arXiv
-
[71]
Neuromorphic online clustering and classification,
——, “Neuromorphic online clustering and classification,”arXiv preprint arXiv:2310.17797, 2023
2023 arXiv
-
[72]
Superconducting sensors and methods in geophysical 13 applications,
R. Stolz, M. Schmelz, V . Zakosarenko, C. Foley, K. Tanabe, X. Xie, and R. Fagaly, “Superconducting sensors and methods in geophysical 13 applications,”Superconductor Science and Technology, vol. 34, no. 3, p. 033001, 2021
2021
-
[73]
A case for superconducting accelerators,
S. S. Tannu, P. Das, M. L. Lewis, R. Krick, D. M. Carmean, and M. K. Qureshi, “A case for superconducting accelerators,” inProceedings of the 16th ACM International Conference on Computing Frontiers, ser. CF ’19. New York, NY , USA: Association for Computing Machinery, 2019, p...
2019
-
[74]
Memristor-based neural networks,
A. Thomas, “Memristor-based neural networks,”Journal of Physics D: Applied Physics, vol. 46, no. 9, p. 093001, 2013
2013
-
[75]
A computational temporal logic for superconducting accelerators,
G. Tzimpragos, D. Vasudevan, N. Tsiskaridze, G. Michelogiannakis, A. Madhavan, J. V olk, J. Shalf, and T. Sherwood, “A computational temporal logic for superconducting accelerators,” inProceedings of the Twenty-Fifth International Conference on Architectural Support for Progra...
2020
-
[76]
Advances in photonic reservoir computing,
G. Van der Sande, D. Brunner, and M. C. Soriano, “Advances in photonic reservoir computing,”Nanophotonics, vol. 6, no. 3, pp. 561–576, 2017
2017
-
[77]
Tnngen: Automated design of neuromorphic sensory processing units for time-series clustering,
P. Vellaisamy, H. Nair, V . Ratnakaram, D. Gupta, and J. Paul Shen, “Tnngen: Automated design of neuromorphic sensory processing units for time-series clustering,”IEEE Transactions on Circuits and Systems II: Express Briefs, vol. 71, no. 5, pp. 2519–2523, 2024
2024
-
[78]
Tnngen: Automated design of neuromorphic sensory processing units for time-series clustering,
P. Vellaisamy, H. Nair, V . Ratnakaram, D. Gupta, and J. P. Shen, “Tnngen: Automated design of neuromorphic sensory processing units for time-series clustering,”IEEE Transactions on Circuits and Systems II: Express Briefs, 2024
2024
-
[79]
Dsagen: Synthesizing programmable spatial accelerators,
J. Weng, S. Liu, V . Dadu, Z. Wang, P. Shah, and T. Nowatzki, “Dsagen: Synthesizing programmable spatial accelerators,” in2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA), 2020, pp. 268–281
2020
-
[80]
How we found the missing memristor,
R. S. Williams, “How we found the missing memristor,”IEEE spectrum, vol. 45, no. 12, pp. 28–35, 2008
2008
-
[81]
Astra-sim2.0: Modeling hierarchical networks and disaggregated systems for large-model training at scale,
W. Won, T. Heo, S. Rashidi, S. Sridharan, S. Srinivasan, and T. Kr- ishna, “Astra-sim2.0: Modeling hierarchical networks and disaggregated systems for large-model training at scale,” in2023 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), ...
2023
-
[82]
uGEMM: Unary Computing Architecture for GEMM Applications,
D. Wu, J. Li, R. Yin, H. Hsiao, Y . Kim, and J. S. Miguel, “uGEMM: Unary Computing Architecture for GEMM Applications,” inInterna- tional Symposium on Computer Architecture, 2020
2020
-
[83]
uSystolic: Byte-Crawling Unary Systolic Array,
D. Wu and J. S. Miguel, “uSystolic: Byte-Crawling Unary Systolic Array,” inInternational Symposium on High-Performance Computer Architecture, 2022
2022
-
[84]
Accelergy: An architecture- level energy estimation methodology for accelerator designs,
Y . N. Wu, J. S. Emer, and V . Sze, “Accelergy: An architecture- level energy estimation methodology for accelerator designs,” in2019 IEEE/ACM International Conference on Computer-Aided Design (IC- CAD), 2019, pp. 1–8
2019
-
[85]
Fast, robust, and transferable prediction for hardware logic synthesis,
C. Xu, P. Sharma, T. Wang, and L. W. Wills, “Fast, robust, and transferable prediction for hardware logic synthesis,” in2023 56th IEEE/ACM International Symposium on Microarchitecture (MICRO), 2023, pp. 167–179
2023
-
[86]
Pimsim: A flexible and detailed processing-in-memory simulator,
S. Xu, X. Chen, Y . Wang, Y . Han, X. Qian, and X. Li, “Pimsim: A flexible and detailed processing-in-memory simulator,”IEEE Computer Architecture Letters, vol. 18, no. 1, pp. 6–9, 2018
2018
-
[87]
A memristor device model,
C. Yakopcic, T. M. Taha, G. Subramanyam, R. E. Pino, and S. Rogers, “A memristor device model,”IEEE electron device letters, vol. 32, no. 10, pp. 1436–1438, 2011
2011
-
[88]
Spiking neural networks and their applications: A review,
K. Yamazaki, V .-K. V o-Ho, D. Bulsara, and N. Le, “Spiking neural networks and their applications: A review,”Brain sciences, vol. 12, no. 7, p. 863, 2022
2022
-
[89]
Fully hardware-implemented memristor convolutional neural network,
P. Yao, H. Wu, B. Gao, J. Tang, Q. Zhang, W. Zhang, J. J. Yang, and H. Qian, “Fully hardware-implemented memristor convolutional neural network,”Nature, vol. 577, no. 7792, pp. 641–646, 2020
2020
-
[90]
Sata: Sparsity-aware training accelerator for spiking neural networks,
R. Yin, A. Moitra, A. Bhattacharjee, Y . Kim, and P. Panda, “Sata: Sparsity-aware training accelerator for spiking neural networks,”IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 2022. 14
2022
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.