REVIEW 3 major objections 6 minor 58 references
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
T0 review · 3 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This paper claims that on a commercial compute-in-SRAM device, three data-movement optimizations make retrieval-augmented-generation (RAG) retrieval 4.8x–6.6x faster than an optimized CPU and as fast as an NVIDIA A6000 GPU while using 54.4x
desk verdict First credible end-to-end evaluation of a commercial compute-in-SRAM device; measured Phoenix/APU data are solid, but the RAG speedup and energy claims depend on simulated HBM and should be read as provisional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The operative object is the APU's vector register (VR): 32,768 16-bit elements stored bit-sliced across 16 banks with a bit processor per column. Because every element of a VR is processed in parallel, element-wise (inter-VR) operations are cheap, while moving data within a single VR—shifts, subgroup reductions, scattered stores—is serialized through global latch lines and scales with shift distance or group size. The analytical framework converts each data-movement and computation operation into a cycle cost, and the three optimizations are exactly the transformations that move work from the expensive intra-VR class to the cheap inter-VR class and replace PIO with DMA or smaller lookup tabl
What would settle it
Connect a GSI APU to real HBM2e-class memory (or a cycle-accurate emulator of it) and run the paper's optimized exact-nearest-neighbor inner-product kernel on the 10/50/200GB corpora, measuring retrieval latency with the same cycle counters. If the measured latency diverges from the Ramulator 2 + DRAMPower prediction by more than the ~6% modeling error reported for Phoenix, or if the 4.8x–6.6x CPU speedup and GPU parity do not reproduce, the central claim fails.
Extended reading notes
Core claim
On a commercial compute-in-SRAM accelerator, realistic workloads are governed less by peak TOPS than by where data is placed and how it moves: reductions within one vector register ('intra-VR') cost roughly an order of magnitude more than element-wise operations spanning registers ('inter-VR'), PIO transfers are far costlier than DMA, and lookup bandwidth grows with table size. The three proposed optimizations—communication-aware reduction mapping (turning spatial intra-VR reductions into temporal inter-VR ones), coalesced DMA (merging duplicate transfers via subgroup copies), and broadcast-friendly data layouts (shrinking lookup table spans)—compound to take a 1024x1024 binary matrix multip
Load-bearing premise
The load-bearing premise is that the simulated HBM2e off-chip memory behaves like a real 380–420 GB/s HBM; the on-chip parts are measured, but the headline RAG retrieval and energy results combine those measurements with modeled memory timing.
Editorial extensions
If this is right
- Exact nearest-neighbor RAG becomes practical without the accuracy loss of approximate search: corpora of 10–200GB can be scanned at GPU-level latency, so LLM answers can cite from full corpora rather than a pruned index.
- For compute-in-SRAM devices generally, the paper implies that data-layout and data-movement optimization can be worth more than increasing raw TOPS, and future designs that support cheap intra-VR reduction and flexible DMA would inherit these gains.
- The 54.4x–117.9x energy gap suggests that energy-hungry retrieval serving could be offloaded to compute-in-SRAM devices, with system-level savings in power and cooling infrastructure beyond the chip's own 60W TDP.
- The analytical framework predicts Phoenix benchmark latencies within an average of 97.3%, meaning developers can evaluate candidate mappings before programming the device, shortening the optimization loop.
- The workload characterization shows compute-in-SRAM is not a universal accelerator: applications with frequent intra-VR operations, such as histogram, matrix multiply, and reverse index, gain less, so target selection matters.
Reading between the lines
- An extension the paper leaves implicit: the RAG speedup is sensitive to achieved off-chip bandwidth, so a useful next experiment is a sensitivity sweep of retrieval latency against HBM bandwidth and latency to see where the APU loses its CPU-beating and GPU-matching advantage.
- The energy breakdown, where static power is about 71% of the 200GB retrieval energy, suggests that the largest energy win for in-memory retrieval may come from faster time-to-sleep and lower standby power, not from per-operation compute efficiency; power gating could multiply the reported advantage.
- The binary matrix multiplication recipe—scalar-vector product mapping, spatial unrolling, DMA coalescing, and broadcast-friendly layout—likely transfers to other bit-parallel kernels, such as binarized transformers, which share the same popcount-accumulate structure.
- The same exact-nearest-neighbor inner-product kernel underlies recommendation, semantic caching, and agent-memory retrieval, so the measured 4.8x–6.6x retrieval speedup is a plausible template for those workloads too.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a measurement-based study of the GSI APU, a commercial compute-in-SRAM accelerator with 2M bit processors. The contributions are: (i) a latency/energy characterization of Phoenix workloads and binary matrix multiplication on the real device; (ii) an analytical latency model built from measured data-movement and computation costs; (iii) three optimizations—communication-aware reduction mapping, coalesced DMA, and broadcast-friendly data layouts; and (iv) an end-to-end retrieval-augmented generation (RAG) comparison against CPU and GPU. The on-chip RAG components are measured on the APU, but the off-chip memory system is replaced by a Ramulator 2 / DRAMPower simulation of HBM2e. This simulated HBM2e directly supports the headline claims of 4.8x–6.6x retrieval acceleration and 1.1x–1.8x end-to-end speedup over CPU, as well as the claim of matching an NVIDIA A6000.
Significance. If the claims are sustained, this is a valuable data point: it is one of the first realistic-workload evaluations on a commercial compute-in-SRAM device, with unusually detailed cycle-level measurements and on-board energy telemetry. The analytical model's maximum 6.2% error on eight Phoenix benchmarks is a genuine validation, and the 1024x1024 binary matrix multiplication case study (226.3 ms to 12.0 ms) is concrete and well explained. However, the most prominent RAG claims are not measured on the device as it exists: the real 23.8 GB/s DDR4 is replaced by a simulated 380–420 GB/s HBM2e, and the simulation is not validated against the APU's actual DMA behavior. The energy comparison in the same section is also ambiguous about which memory system was measured. These issues are load-bearing for the abstract's headline numbers, but they are fixable with sensitivity analysis and clarification, so the underlying contribution remains credible.
major comments (3)
- [Section 5.3.1, Table 8, Fig. 14] The RAG retrieval speedups reported in the abstract and §5.3.3 (6.3x/4.8x/6.6x retrieval; 1.05x/1.15x/1.75x end-to-end) depend on replacing the device's 23.8 GB/s DDR4 with a Ramulator 2 / DRAMPower-simulated HBM2e. In Table 8, the 'Load Embedding' row is the only component using the simulated memory: at 200 GB, 2.4 GB of embeddings completes in 6.1 ms, corresponding to ~393 GB/s. At the device's real DDR4 bandwidth, the same transfer would take ~101 ms, so the optimized 200 GB retrieval would be approximately 179 ms rather than 84.2 ms. This would lower the CPU retrieval speedup from ~6.6x to ~3.1x and the end-to-end gain from ~1.75x to ~1.5x. The paper should either report RAG results at the actual DDR4 bandwidth as the primary numbers or provide validation (e.g., a bandwidth sensitivity sweep, measured DMA efficiency, and a description of how Ramulator timings were coupled to the meas
- [Section 5.3.5, Figure 15] The 54.4x–117.9x energy-efficiency claim is presented as applying to the optimized APU RAG system, but the energy measurements use the on-board UCD9090/ISL8273M telemetry of the actual Leda-E board (Section 5), which contains DDR4, while the latency results in the same evaluation use the simulated HBM2e. The paper does not state whether the reported 'APU energy' includes the simulated HBM's energy (e.g., via a DRAMPower HBM power model) or only the measured DDR4 system. If it is the latter, the energy claim is for a different configuration than the latency claim; if it is the former, the HBM energy model must be described. This distinction matters for the headline comparison with the A6000 and should be clarified in a revision.
- [Section 3.3, Eq. (1)] The reduction latency model is a cubic polynomial in log2(subgroup size) with eight experimentally fitted coefficients (alpha_i, beta_i). The paper does not report the fitting procedure, the number of measurements, residuals, or confidence intervals. The aggregate validation in Table 7 does not isolate Eq. (1) from the rest of the framework, so a good aggregate error does not establish that the polynomial form generalizes. Since the paper advertises the framework as a flexible tool for design-space exploration, the model needs a more rigorous fit/validation narrative: data points used, residuals per subgroup size, and ideally an out-of-sample check (e.g., leave-one-benchmark-out). The current presentation is under-specified for assessing how much of the 97.3% accuracy comes from the polynomial form versus the measured constants in Table 4.
minor comments (6)
- [Section 5.3.1] The text mentions 'two GPUs (one dedicated to generation and the other to retrieval)' but the subsequent comparison is unclear about the memory bandwidth and model of the retrieval GPU. Please state whether the retrieval GPU is the same A6000 model or a different one.
- [Table 4] The column header 'Analytical Meas.' is ambiguous; clarify that the two columns are modeled and measured cycles, and state the units explicitly.
- [Eq. (10)] The printed formula appears to be missing a multiplication operator ('...· 𝑀 ⌊𝑙/𝑁⌋ ·𝐾'); check the rendering.
- [Figure 13] The legend defines Opt1/2/3 only in the text; add the definitions to the caption for readability.
- [Section 2.1.1] The phrase '16-bit IEEE floating point' should be 'IEEE binary16' to avoid ambiguity; the following sentence mentions a 'custom GSI floating point format with a 6-bit exponent and a 9-bit mantissa,' which is non-standard and should be flagged as such.
- [Section 5.3.1] The simulated HBM2e configuration ('16 GB, 2 ranks, 8 channels, 1.6 GHz') is described in one sentence; please provide the key timing parameters (tCK, tRCD, tCL, tRP) or cite the Ramulator 2 configuration file so the results are reproducible.
Circularity Check
No circular derivation: headline results are measured on hardware (plus explicitly simulated HBM), not derived from the fitted analytical framework; self-citations are not load-bearing.
full rationale
The paper's central claims (Phoenix speedups, RAG speedups, energy ratios) come from direct measurement on the GSI APU or from combining measured on-chip times with Ramulator-2/DRAMPower-simulated HBM2e timings; they are not obtained by plugging the conclusion back into the analytical model. The framework's constants (Tables 4–5) are empirically measured microbenchmark latencies, and Table 7 composes those latencies to predict Phoenix runtimes with ≤6.2% error. That is a compositional sanity check on the same device, not a self-validating derivation. No parameter is fitted to the RAG result and then renamed as a prediction: the RAG speedups are hardware measurements, and the load-embedding component is simulated under an explicitly disclosed assumption. Table 8 states: 'Load embedding latency reflects simulated HBM2e performance; all other values are measured on GSI APU hardware.' The abstract similarly says: 'The shared off-chip memory bandwidth is modeled using a simulated HBM, while all other components are measured on the real compute-in-SRAM device.' This makes the 4.8–6.6x retrieval and 1.1–1.8x end-to-end numbers conditional on the HBM simulation being representative — a validity/correctness caveat, not a circularity. Self-citations to prior APU work by the same group ([18], [19], [33]) are background examples and are not load-bearing; the paper imports no uniqueness theorem or ansatz from them. The score of 2 reflects the minor self-citations and the disclosed simulation condition, not an identified circular step.
Assumptions & free parameters
free parameters (4)
- Reduction polynomial coefficients alpha_i, beta_i (i=0..3) =
8 values, not all listed; fitted to microbenchmarked subgroup reduction latencies
- DMA bandwidth and init overhead constants =
e.g., 0.19d+411, 0.63d+548, 22272 cycles for L4->L1
- Lookup table coefficient C and init overhead =
7.15 cycles/entry, 629 cycles init
- Shift cost constant =
373 cycles per element for shift_e(k)
assumptions (4)
- domain assumption The GSI APU is a representative instance of the general-purpose compute-in-SRAM abstraction in Figure 1.
- domain assumption The simulated HBM2e memory system, built with Ramulator 2 and DRAMPower, accurately models off-chip memory for the RAG evaluation.
- domain assumption nvidia-smi GPU power readings and on-board APU point-of-load power monitors measure comparable device-level energy.
- ad hoc to paper Subgroup reduction latency follows the cubic-in-log-subgroup-size polynomial form of Equation 1.
Cite this review
Pith. "Pith review of Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device." pith.science (2026). https://pith.science/paper/HBB3N446
@misc{pith2026250905451,
author = {Pith},
title = {Pith review of: Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device},
year = {2026},
howpublished = {\url{https://pith.science/paper/HBB3N446}},
note = {Machine review of arXiv:2509.05451}
}
abstract
Compute-in-SRAM architectures offer a promising approach to achieving higher performance and energy efficiency across a range of data-intensive applications. However, prior evaluations have largely relied on simulators or small prototypes, limiting the understanding of their real-world potential. In this work, we present a comprehensive performance and energy characterization of a commercial compute-in-SRAM device, the GSI APU, under realistic workloads. We compare the GSI APU against established architectures, including CPUs and GPUs, to quantify its energy efficiency and performance potential. We introduce an analytical framework for general-purpose compute-in-SRAM devices that reveals fundamental optimization principles by modeling performance trade-offs, thereby guiding program optimizations. Exploiting the fine-grained parallelism of tightly integrated memory-compute architectures requires careful data management. We address this by proposing three optimizations: communication-aware reduction mapping, coalesced DMA, and broadcast-friendly data layouts. When applied to retrieval-augmented generation (RAG) over large corpora (10GB--200GB), these optimizations enable our compute-in-SRAM system to accelerate retrieval by 4.8$\times$--6.6$\times$ over an optimized CPU baseline, improving end-to-end RAG latency by 1.1$\times$--1.8$\times$. The shared off-chip memory bandwidth is modeled using a simulated HBM, while all other components are measured on the real compute-in-SRAM device. Critically, this system matches the performance of an NVIDIA A6000 GPU for RAG while being significantly more energy-efficient (54.4$\times$-117.9$\times$ reduction). These findings validate the viability of compute-in-SRAM for complex, real-world applications and provide guidance for advancing the technology.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
Shaizeen Aga, Supreet Jeloka, Arun Subramaniyan, Satish Narayanasamy, David Blaauw, and Reetuparna Das. 2017. Compute Caches. In2017 IEEE International Symposium on High Performance Computer Architecture (HPCA)(Austin, TX, USA). IEEE Computer Society, Los Alamitos, CA, USA, 481–492. https://doi.org/ 10.1109/HPCA.2017.21
-
[2]
Amogh Agrawal, Akhilesh Jaiswal, Chankyu Lee, and Kaushik Roy. 2018. X- SRAM: Enabling in-memory Boolean computations in CMOS static random access memories.IEEE Transactions on Circuits and Systems I: Regular Papers65, 12 (2018), 4219–4232
work page 2018
-
[3]
Amogh Agrawal, Akhilesh Jaiswal, Deboleena Roy, Bing Han, Gopalakrishnan Srinivasan, Aayush Ankit, and Kaushik Roy. 2019. Xcel-RAM: Accelerating binary neural networks in high-throughput SRAM compute arrays.IEEE Transactions on Circuits and Systems I: Regular Papers66, 8 (2019), 3064–3076
work page 2019
-
[4]
Junwhan Ahn, Sungpack Hong, Sungjoo Yoo, Onur Mutlu, and Kiyoung Choi
-
[5]
Junwhan Ahn, Sungjoo Yoo, Onur Mutlu, and Kiyoung Choi. 2015. PIM-enabled instructions: A low-overhead, locality-aware processing-in-memory architecture. ACM SIGARCH Computer Architecture News43, 3S (2015), 336–348
work page 2015
-
[6]
Khalid Al-Hawaj, Olalekan Afuye, Shady Agwa, Alyssa Apsel, and Christopher Batten. 2020. Towards a Reconfigurable Bit-Serial/Bit-Parallel Vector Accelerator Using In-Situ Processing-in-SRAM. In2020 IEEE International Symposium on Circuits and Systems (ISCAS)(Seville, Spain (Virtual)). Institute of Electrical and Electronics Engineers, Piscataway, NJ, USA,...
arXiv 2020
-
[7]
Khalid Al-Hawaj, Tuan Ta, Nick Cebry, Shady Agwa, Olalekan Afuye, Eric Hall, Courtney Golden, Alyssa B. Apsel, and Christopher Batten. 2023. EVE: Ephemeral Vector Engines. In2023 IEEE International Symposium on High-Performance Com- puter Architecture (HPCA)(Montreal, QC, Canada). IEEE Computer Society, Los Alamitos, CA, USA, 691–704. https://doi.org/10.1...
arXiv 2023
-
[8]
Daniel Bankman, Lita Yang, Bert Moons, Marian Verhelst, and Boris Murmann
Show all 58 references
-
[9]
Chandrakasan
Avishek Biswas and Anantha P. Chandrakasan. 2018. Conv-RAM: An Energy- Efficient SRAM with Embedded Convolution Computation for Low-Power CNN- Based Machine Learning Applications. In2018 IEEE International Solid-State Circuits Conference (ISSCC)(San Francisco, CA, USA). Instit...
2018
-
[10]
Brockman, Shyamkumar Thoziyoor, Shannon K
Jay B. Brockman, Shyamkumar Thoziyoor, Shannon K. Kuntz, and Peter M. Kogge
-
[11]
Martínez
Helena Caminal, Kailin Yang, Srivatsa Srinivasa, Akshay Krishna Ramanathan, Khalid Al-Hawaj, Tianshu Wu, Vijaykrishnan Narayanan, Christopher Batten, and José F. Martínez. 2021. CAPE: A Content-Addressable Processing Engine. In 2021 IEEE International Symposium on High-Perform...
2021
-
[12]
2024.DRAM- Power: Open-source DRAM Power & Energy Estimation Tool
Karthik Chandrasekar, Christian Weis, Yonghui Li, Sven Goossens, Matthias Jung, Omar Naji, Benny Akesson, Norbert Wehn, and Kees Goossens. 2024.DRAM- Power: Open-source DRAM Power & Energy Estimation Tool. TU Kaiserslautern, Microelectronic Systems Design (MSD) Research Group....
2024
-
[13]
Johannes de Fine Licht, Maciej Besta, Simon Meierhans, and Torsten Hoefler. 2020. Transformations of high-level synthesis codes for high-performance computing. IEEE Transactions on Parallel and Distributed Systems32, 5 (2020), 1014–1029
2020
- [14]
-
[15]
Caxton C. Foster. 1976.Content Addressable Parallel Processors. Van Nostrand Reinhold Company, New York, NY, USA
1976
- [16]
-
[17]
Joseph Gebis, Sam Williams, David Patterson, and Christos Kozyrakis. 2004. VIRAM1: A Media-Oriented Vector Processor with Embedded DRAM. InDAC ’04: Proceedings of the 41st Annual Design Automation Conference – Student Design Contest(San Diego, CA, USA). Association for Computi...
2004
-
[18]
Courtney Golden, Dan Ilan, Nicholas Cebry, and Christopher Batten. 2023. Ac- celerating Seed Location Filtering in DNA Read Mapping Using a Commercial Compute-in-SRAM Architecture. InProceedings of the 5th Workshop on Accelerator Architecture in Computational Biology and Bioin...
2023 arXiv
-
[19]
Courtney Golden, Dan Ilan, Caroline Huang, Niansong Zhang, Zhiru Zhang, and Christopher Batten. 2024. Supporting a Virtual Vector Instruction Set on a Commercial Compute-in-SRAM Accelerator.IEEE Computer Architecture Letters 23, 1 (Jan 2024), 29–32. https://doi.org/10.1109/LCA...
2024
- [20]
-
[21]
Friedman
Qing Guo, Xiaochen Guo, Ravi Patel, Engin Ipek, and Eby G. Friedman. 2013. AC-DIMM: Associative Computing with STT-MRAM. InISCA ’13: Proceedings of the 40th Annual International Symposium on Computer Architecture(Tel-Aviv, Israel). Association for Computing Machinery, New York...
2013
-
[22]
2020.In-Memory Acceleration for Big Data
Linley Gwennap. 2020.In-Memory Acceleration for Big Data. Technical Report. The Linley Group. https://gsitechnology.com/wp-content/uploads/2023/01/GSIT- Gemini-WP-Final-Linley.pdf
2020
-
[23]
Bastian Hagedorn, Bin Fan, Hanfeng Chen, Cris Cecka, Michael Garland, and Vinod Grover. 2023. Graphene: An IR for Optimized Tensor Computations on GPUs. InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Syst...
2023
-
[24]
Mohsen Imani, Saransh Gupta, Yeseong Kim, and Tajana Rosing. 2019. FloatPIM: In-Memory Acceleration of Deep Neural Network Training with High Precision. InISCA ’19: Proceedings of the 46th International Symposium on Computer Archi- tecture(Phoenix, AZ, USA). Association for Co...
2019
-
[25]
Ishida, T
M. Ishida, T. Kawakami, A. Tsuji, N. Kawamoto, M. Motoyoshi, and N. Ouchi. 1998. A Novel 6T-SRAM Cell Technology Designed with Rectangular Patterns Scalable Beyond 0.18𝜇m Generation and Desirable for Ultra High Speed Operation. In 1998 IEEE International Electron Devices Meeti...
1998
-
[26]
Supreet Jeloka, Naveen Bharathwaj Akesh, Dennis Sylvester, and David Blaauw
-
[27]
Zhewei Jiang, Shihui Yin, Jae-Sun Seo, and Mingoo Seok. 2020. C3SRAM: An in- memory-computing SRAM macro based on robust capacitive coupling computing mechanism.IEEE Journal of Solid-State Circuits55, 7 (2020), 1888–1897
2020
-
[28]
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for C...
2019
-
[29]
Niemier, and X
Ann Franchesca Laguna, Arman Kazemi, Michael T. Niemier, and X. Sharon Hu. 2021. In-Memory Computing Based Accelerator for Transformer Networks for Long Sequences. InProceedings of the 2021 Design, Automation & Test in Europe Conference & Exhibition (DATE)(Grenoble, France (Vi...
2021
-
[30]
Sharon Hu
Ann Franchesca Laguna, Mohammed Mehdi Sharifi, Arman Kazemi, Xunzhao Yin, Michael Niemier, and X. Sharon Hu. 2022. Hardware-Software Co-Design of an In-Memory Transformer Network Accelerator.Frontiers in Electronics3, Article 847069 (Apr 2022), 21 pages. https://doi.org/10.338...
2022
-
[31]
Phuoc-Hoan Charles Le and Xinlin Li. 2023. BinaryViT: Pushing Binary Vision Transformers Towards Convolutional Models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (Vancouver, BC, Canada). IEEE Computer Society, Los Alam...
2023
-
[32]
Eunyoung Lee, Taeyoung Han, Donguk Seo, Gicheol Shin, Jaerok Kim, Seonho Kim, Soyoun Jeong, Johnny Rhe, Jaehyun Park, Jong Hwan Ko, et al . 2021. A charge-domain scalable-weight in-memory computing macro with dual-SRAM ar- chitecture for precision-scalable DNN accelerators.IEE...
2021
-
[33]
Kaitlyn Lee, Brian Donnelly, Tomer Sery, Dan Ilan, Bertrand Cambou, and Michael Gowanlock. 2023. Evaluating Accelerators for a High-Throughput Hash-Based Security Protocol. InProceedings of the 52nd International Conference on Parallel Processing Workshops (ICPP-W ’23)(Salt La...
2023
-
[34]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing S...
2020
-
[35]
Nisa Bostanci, Ataberk Olgun, A
Haocong Luo, Yahya Can Tu, F. Nisa Bostanci, Ataberk Olgun, A. Giray Yaglikci, and Onur Mutlu. 2024. Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator.IEEE Computer Architecture Letters23, 1 (2024), 112–116. https://doi.org/10.1109/LCA.2023.3333759
2024
-
[36]
Chong, and Timothy Sherwood
Mark Oskin, Frederic T. Chong, and Timothy Sherwood. 1998. Active Pages: A Computation Model for Intelligent Memory. InProceedings of the 25th Annual International Symposium on Computer Architecture (ISCA ’98)(Barcelona, Spain). IEEE Computer Society, Los Alamitos, CA, USA, 19...
1998
-
[37]
David Patterson, Thomas Anderson, Neal Cardwell, Richard Fromm, Kimberly Keeton, Christoforos Kozyrakis, Randi Thomas, and Katherine Yelick. 1997. A case for intelligent RAM.IEEE micro17, 2 (1997), 34–44
1997
-
[38]
David Patterson, Thomas Anderson, Neal Cardwell, Richard Fromm, Kimber- ley Keeton, Christoforos Kozyrakis, Randi Thomas, and Katherine Yelick. 1997. Intelligent RAM (IRAM): Chips that Remember and Compute. In1997 IEEE Inter- national Solid-State Circuits Conference (ISSCC) Di...
1997
-
[39]
Jerry Potter, Johnnie Baker, Stephen Scott, Arvind Bansal, Chokchai Leangsuksun, and Chandra Asthagiri. 1994. ASC: an associative-computing paradigm.Computer 27, 11 (1994), 19–25
1994
-
[40]
Derrick Quinn, Mohammad Nouri, Neel Patel, John Salihu, Alireza Salemi, Sukhan Lee, Hamed Zamani, and Mohammad Alian. 2025. Accelerating Retrieval- Augmented Generation. InProceedings of the 30th ACM International Confer- ence on Architectural Support for Programming Languages...
2025
-
[41]
Colby Ranger, Ramanan Raghuraman, Arun Penmetsa, Gary Bradski, and Chris- tos Kozyrakis. 2007. Evaluating MapReduce for Multi-core and Multiprocessor Systems. In2007 IEEE 13th International Symposium on High Performance Com- puter Architecture (HPCA)(Phoenix, AZ, USA). IEEE Co...
2007
-
[42]
Alireza Salemi and Hamed Zamani. 2024. Evaluating Retrieval Quality in Retrieval-Augmented Generation. InSIGIR ’24: Proceedings of the 47th Inter- national ACM SIGIR Conference on Research and Development in Information Retrieval(Washington, DC, USA). Association for Computing...
2024
-
[43]
Gordon E. Sayre. 1976. Staran: An Associative Approach to Multiprocessor Architecture. InComputer Architecture: Workshop of the Gesellschaft für In- formatik, Erlangen, May 22–23, 1975. Springer, Berlin, Heidelberg, 199–221. https://doi.org/10.1007/978-3-642-66400-7_9
1976 doi
-
[44]
In-Memory Acceleration for Big Data
GSI Technology. 2023. GSI Technology’s Gemini-I®APU Showcased in “In-Memory Acceleration for Big Data”. https://ir.gsitechnology.com/news- releases/news-release-details/gsi-technologys-gemini-ir-apu-showcased- memory-acceleration-big
2023
-
[45]
Fengbin Tu, Zihan Wu, Yiqi Wang, Ling Liang, Liu Liu, Yufei Ding, Leibo Liu, Shaojun Wei, Yuan Xie, and Shouyi Yin. 2023. TranCIM: Full-Digital Bitline- Transpose CIM-based Sparse Transformer Accelerator With Pipeline/Parallel Reconfigurable Modes.IEEE Journal of Solid-State C...
2023
- [46]
-
[47]
Yue Zha and Jing Li. 2020. Hyper-AP: Enhancing Associative Processing Through a Full-Stack Optimization. In2020 ACM/IEEE 47th Annual Interna- tional Symposium on Computer Architecture (ISCA)(Valencia, Spain (Virtual)). Institute of Electrical and Electronics Engineers, Piscata...
2020
-
[48]
Bo Zhang, Shihui Yin, Minkyu Kim, Jyotishman Saikia, Soonwan Kwon, Sung- meen Myung, Hyunsoo Kim, Sang Joon Kim, Jae-Sun Seo, and Mingoo Seok
-
[49]
Jintao Zhang, Zhuo Wang, and Naveen Verma. 2017. In-memory computation of a machine-learning classifier in a standard 6T SRAM array.IEEE Journal of Solid-State Circuits52, 4 (2017), 915–924
2017
-
[50]
Yichi Zhang, Junhao Pan, Xinheng Liu, Hongzheng Chen, Deming Chen, and Zhiru Zhang. 2021. FracBNN: Accurate and FPGA-Efficient Binary Neural Networks with Fractional Activations. InFPGA ’21: Proceedings of the 2021 ACM/SIGDA International Symposium on Field-Programmable Gate A...
2021
-
[51]
Yichi Zhang, Zhiru Zhang, and Lukasz Lew. 2022. PokeBNN: A Binary Pursuit of Lightweight Accuracy. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(New Orleans, LA, USA). IEEE Computer So- ciety, Los Alamitos, CA, USA, 12475–12485. htt...
2022
-
[1809]
https://doi.org/10.1109/JSSC.2022.3213542
2022
-
[2004]
InWMPI ’04: Proceedings of the 3rd Workshop on Memory Performance Issues, in conjunction with the 31st International Symposium on Computer Architecture(Munich, Germany)
A Low Cost, Multithreaded Processing-in-Memory System. InWMPI ’04: Proceedings of the 3rd Workshop on Memory Performance Issues, in conjunction with the 31st International Symposium on Computer Architecture(Munich, Germany). Association for Computing Machinery, New York, NY, U...
-
[2015]
InISCA ’15: Proceedings of the 42nd Annual International Symposium on Computer Architecture(Portland, OR, USA)
A Scalable Processing-in-Memory Accelerator for Parallel Graph Processing. InISCA ’15: Proceedings of the 42nd Annual International Symposium on Computer Architecture(Portland, OR, USA). Association for Computing Machinery, New York, NY, USA, 105–117. https://doi.org/10.1145/2...
-
[2016]
A 28 nm configurable memory (TCAM/BCAM/SRAM) using push-rule 6T bit cell enabling logic-in-memory.IEEE Journal of Solid-State Circuits51, 4 (2016), 1009–1021
2016
-
[2018]
An Always-On 3.8𝜇J 86% CIFAR-10 mixed-signal binary CNN processor with all memory on chip in 28-nm CMOS.IEEE Journal of Solid-State Circuits54, 1 (2018), 158–172
2018
-
[2023]
https://doi.org/10.1109/JSSC.2022.3211290
PIMCA: A Programmable In-Memory Computing Accelerator for Energy- Efficient DNN Inference.IEEE Journal of Solid-State Circuits58, 5 (May 2023), 1436–1449. https://doi.org/10.1109/JSSC.2022.3211290
2023
-
[4673]
https://doi.org/10.1109/CVPRW59228.2023.00492
2023
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.