REVIEW 3 major objections 6 minor 145 references
DX100: A Programmable Data Access Accelerator for Indirection
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A shared programmable accelerator that reorders bulk tiles of 16K indirect accesses raises DRAM row-buffer hit rate 2.7x and delivers 2.6x speedups on irregular workloads.
desk verdict Credible simulation-backed accelerator with an unevaluated compiler and a load-bearing coherency assumption. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Indirect Access unit's three-table pipeline operating at tile granularity. A Row Table with BCAM slices per DRAM bank records up to 64 target rows and 8 columns per row; a Word Table keeps a linked list of tile iterations that hit the same DRAM column, enabling coalescing; and a Request Generator arbitrates among row-table slices so that channels and bank groups are interleaved. This converts an arbitrary index stream into grouped row accesses, giving the accelerator a reordering window (16K accesses) far larger than the memory controller's 32-entry request buffer. The same structure supports conditional accesses, multiple indirection levels, and stream, ALU, and range-fuser units for address calculation and loop fusion.
What would settle it
Run a workload where one core writes to an indirect array between DX100's fill stage and its request stage, or bypass the compiler's no-alias check, and observe stale values or incorrect RMW results; a clean alternative is to force a low-sparsity index set at a 1K tile size and check whether DX100's bandwidth utilization still exceeds the baseline by the claimed 3.9x average.
Extended reading notes
Core claim
On its own terms, DX100's central claim is that indirect memory traffic can be treated as a bulk data-movement problem rather than a latency-hiding problem. The accelerator fetches index tiles into a scratchpad, evaluates loop conditions with an ALU unit, computes addresses, and then issues the resulting loads, stores, and RMWs from a 16K-element window that the memory controller never sees. Its Indirect Access unit uses a Row Table to keep addresses that belong to the same DRAM row together, a Word Table to link repeated accesses to the same column so they are coalesced, and a Request Generator to interleave the resulting requests across channels and bank groups. This ordering, the paper argues, is what raises request-buffer occupancy by 12.1x and row-buffer hit rate by 2.7x on average, converting time wasted on precharge and activate into useful data movement. The paper also claims a 3.6x reduction in core instructions because address arithmetic and condition checks move off the core.
Load-bearing premise
Everything in the evaluation rests on DX100 being able to prove and maintain exclusive write access to the indirect arrays for the duration of an offloaded kernel: the interface snoops coherency directories when a tile is filled and issues the actual requests later, and the compiler's alias analysis must show there are no cross-iteration dependencies and no core or other-DX100 stores to those arrays in between, or the reordered loads, stores, and RMWs can corrupt data and the claimed speedups vanish.
Editorial extensions
If this is right
- Across the 12 evaluated workloads, DX100 reports a 2.6x geometric mean speedup over a four-core baseline and a 2.0x speedup over a state-of-the-art indirect prefetcher, with bandwidth utilization 3.9x higher and row-buffer hit rate 2.7x higher.
- The reordering benefit is largely independent of index order: in synthetic all-miss experiments DX100 keeps DRAM bandwidth utilization between 82% and 85% while the baseline falls to 27-65% as row-buffer hits, channel interleaving, and bank-group interleaving degrade.
- Tile size is a first-order knob: increasing the tile from 1K to 32K elements raises the reported speedup from 1.7x to 2.9x, mainly through more coalescing and a 27% higher row-buffer hit rate.
- By acting as the sole writer to indirect regions, DX100 removes the need for fine-grained atomicity: its RMW microbenchmark is 17.8x faster than an atomic baseline and its scatter microbenchmark is 6.6x faster than a single-core baseline.
- The shared design scales: moving from four to eight cores with one or two DX100 instances preserves 2.5-2.7x speedups under a coarse-grained region-based coherence protocol.
Reading between the lines
- Inference: the measured speedups assume the compiler's alias analysis and the coherency snooping at fill time reliably guarantee that no other writer touches indirect arrays between fill and request. On a real system this exclusivity needs explicit enforcement; otherwise the same reordering that brings the bandwidth gains also introduces stale-read and lost-update risks.
- Inference: the row-table reordering technique suggests a cheaper variant where the memory controller itself grows its scheduling window using similar address clustering on detected indirect streams, potentially capturing part of the 2.6x benefit without a dedicated accelerator.
- Inference: bulk indirect reordering transfers naturally to GPUs and sparse-tensor accelerators, where gather and scatter streams also arrive in large windows; the paper's 16K-tile granularity gives a concrete design point for those systems.
- Inference: a direct comparison against software reordering methods (batch-and-reorder runtime systems) would isolate how much of the gain is the hardware window versus the compile-time hoisting; the paper does not report such a comparison.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DX100, a shared, memory-mapped programmable accelerator that offloads bulk indirect loads, stores, and read-modify-write (RMW) operations from CPU cores. DX100 fetches a tile of indices (up to 16K elements), reorders the resulting accesses with per-bank row and word tables, coalesces duplicate accesses, and interleaves requests across DRAM channels and bank groups, aiming to improve DRAM row-buffer hit rate and bandwidth utilization. The evaluation uses gem5 plus Ramulator2 on a 4-core Skylake-like configuration across 12 benchmarks, and reports a geometric-mean speedup of 2.6x over a multicore baseline with a 2MB-larger LLC, 2.0x over the DMP indirect prefetcher, 3.9x higher bandwidth utilization, and 2.7x higher row-buffer hit rate. The paper also presents an 8-instruction ISA, an MLIR-based compiler pass pipeline, RTL synthesis results for area and power, and a public artifact appendix.
Significance. If the correctness contract can be enforced, the contribution is significant: the central idea of using a large reordering window near the memory controller to improve DRAM bandwidth utilization for irregular accesses is well motivated and clearly explained. The evaluation is unusually thorough for the systems community: full-system simulation with a detailed DRAM model, baseline LLC compensation for the accelerator's scratchpad, reproduction of DMP using its public artifact, and a public artifact with Docker support and expected results. The row-buffer hit rate and request-buffer occupancy analyses in Figures 10(b) and 10(c) provide concrete mechanisms for the reported speedups. The main risk is correctness under concurrency, not circularity: there are no fitted parameters or ad-hoc equations, and the reported benefits are empirical simulation results. The key unresolved issue is the unenforced exclusive-writer contract for indirect arrays, which the RMW results rely on.
major comments (3)
- [Section 3.6/4.2/6.6] The correctness of the bypass path relies on the assertion in Section 3.6 that no core can modify a cache line between the fill-time snoop (which sets the H bit) and the later request issue. In the evaluated single-instance, multi-core configuration, this exclusivity is not enforced by the hardware: the Coherency Agent tracks only scratchpad cache lines, the Controller scoreboard tracks only scratchpad tile hazards (Section 3.5), and the region-lock mechanism in Section 6.6 is described only for multiple DX100 instances. The compiler legality condition in Section 4.2 ("no core stores to the memory regions accessed by DX100 within the loop body") does not cover concurrent DX100 instructions from different cores that target overlapping memory regions, which is precisely the situation in the RMW kernels PR, BC, and UME in Table 1, where multiple iterations or threads may update the same indirect location. The paper should either (a) specify and model a hardware mechanism that enforces the exclusive-writer contract for a single DX100 instance, such as memory-region overlap tracking or serialization of overlapping instructions, including its cost in traffic and latency, or (b) demonstrate that the benchmark implementations already provide this exclusivity through software barriers or locks, and quantify the resulting overhead. Without this, the measured speedups, including the 17.8x RMW microbenchmark gain in Section 6.1, rest on an unverified correctness assumption.
- [Section 4.2/6] The MLIR compiler is a stated contribution and the abstract claims DX100 can be used "without significant programming efforts," but no evaluation of the compiler appears in Section 6. All reported benchmark numbers are produced with hand-inserted APIs, and Section 4.1 concedes that the manual method is a fallback "due to compiler limitations (memory dependence analysis and code pattern detection)." The paper should report the compiler's success rate on the 12 benchmarks or a representative subset, the correctness of its alias-analysis-based legality checks, and the performance of compiler-generated code relative to hand-inserted APIs. If the compiler is intended as future work, the abstract and introduction should be revised to avoid claiming a complete automatic transformation path.
- [Section 6.3/7] The experimental comparison is limited to DMP, an indirect prefetcher, but the closest prior art for a programmable data access accelerator is arguably the general-purpose fetcher units SpZip [130] and Terminus [65], which are discussed in Section 7 but not evaluated. The claim that fetchers "provide insufficient visibility into future memory accesses" is asserted, not measured. The paper should either add a configuration-matched comparison against SpZip and Terminus, or provide a quantitative argument, for example from the row-buffer hit rate and request-buffer occupancy data, why those designs cannot achieve the reported 2.6x and 2.0x improvements. Without this, the positioning of DX100 relative to prior fetcher work is not fully supported.
minor comments (6)
- [Section 5] The UME workload list contains a typo: "GZ, GP, GZI, and GZI" should presumably be "GZP, GZZ, GZPI, and GZZI" as used in Table 1 and Figure 9.
- [Table 1] The UME GZZ row shows "A[B[[i]]]" with a double opening bracket; this should be "A[B[i]]."
- [Section 6.3] The DMP comparison uses a 4-core configuration with a 10MB LLC, whereas DMP was originally validated on a single-core configuration with a 256KB L2, as the paper itself notes. A sensitivity experiment that also reports DMP's speedup in a configuration closer to its original target would help calibrate the meaning of the 2.0x claim.
- [Section 6.2] The statement that "only 54 instructions in a 224-entry ROB are loads and stores across all cycles of all benchmarks" is unclear; it should specify whether this is an average occupancy, a peak, or a fraction of the ROB, and how it was measured.
- [Section 3.6] The paper assumes huge pages and a 256-entry DX100 TLB for address translation, but it does not state whether TLB misses and page-table walks are modeled in the gem5 simulations; if they are not modeled, the overhead of PTE transfers and TLB misses should be discussed.
- [Section 6.4] The tile-size sensitivity study in Figure 13 shows increasing speedup up to 32K elements, but the paper does not explain why 16K is chosen as the evaluated tile size beyond the scratchpad configuration; a sentence on the area/capacity tradeoff would be helpful.
Circularity Check
No circularity: DX100's 2.6×/2.0× speedups are simulator measurements driven by a concrete microarchitecture, not by fitted parameters or self-cited constraints.
full rationale
The paper's central claims are empirical: DX100 is implemented in gem5 with Ramulator2, and performance, bandwidth utilization, row-buffer hit rate, request-buffer occupancy, instruction reduction, and MPKI are measured outputs under the configurations of Table 3. There is no fitted parameter that is later renamed a prediction; tile size, scratchpad capacity, and table sizes are fixed design choices, and Section 6.4 reports a sensitivity sweep rather than tuning to a target. The comparison against DMP uses the authors' public gem5 artifact [23], so the 2.0× result rests on an externally reproducible baseline rather than on the present authors' own claimed theorems. The few self-citations ([24, 36, 96, 109, 112, 113, 119]) appear in motivating context or as benchmark methodology and are not load-bearing for the central result. The correctness caveat about exclusive write access to indirect arrays (Sections 3.6 and 4.2) is a real implementation risk: fill-time snoops can become stale if a core writes before request issue, and the compiler legality pass is not evaluated on the 12 benchmarks. That makes the speedups conditional on an unverified invariant, but it is not circularity: the result does not reduce by construction to an input, and the paper explicitly states the assumption rather than defining the outcome through it. No circular step can be quoted from the paper.
Assumptions & free parameters
free parameters (3)
- Tile size =
16K 4B words
- Row Table capacity =
64 rows per bank slice, 8 columns per row
- Scratchpad size =
2MB (32 tiles of 16K words)
assumptions (4)
- domain assumption Offloaded indirect accesses have no cross-iteration data dependencies, and no core stores to the accessed arrays within the loop body.
- domain assumption RMW operations are restricted to associative and commutative operations (ADD, MAX, MIN) so reordering is safe.
- domain assumption Snooping coherency directories at fill time plus exclusive write access keeps indirect accesses coherent while bypassing the LLC.
- domain assumption Huge pages and the 256-entry TLB keep address translation off the critical path.
invented entities (1)
-
DX100 accelerator
Cite this review
Pith. "Pith review of DX100: A Programmable Data Access Accelerator for Indirection." pith.science (2026). https://pith.science/paper/BGNFTWYM
@misc{pith2026250523073,
author = {Pith},
title = {Pith review of: DX100: A Programmable Data Access Accelerator for Indirection},
year = {2026},
howpublished = {\url{https://pith.science/paper/BGNFTWYM}},
note = {Machine review of arXiv:2505.23073}
}
read the original abstract
Indirect memory accesses frequently appear in applications where memory bandwidth is a critical bottleneck. Prior indirect memory access proposals, such as indirect prefetchers, runahead execution, fetchers, and decoupled access/execute architectures, primarily focus on improving memory access latency by loading data ahead of computation but still rely on the DRAM controllers to reorder memory requests and enhance memory bandwidth utilization. DRAM controllers have limited visibility to future memory accesses due to the small capacity of request buffers and the restricted memory-level parallelism of conventional core and memory systems. We introduce DX100, a programmable data access accelerator for indirect memory accesses. DX100 is shared across cores to offload bulk indirect memory accesses and associated address calculation operations. DX100 reorders, interleaves, and coalesces memory requests to improve DRAM row-buffer hit rate and memory bandwidth utilization. DX100 provides a general-purpose ISA to support diverse access types, loop patterns, conditional accesses, and address calculations. To support this accelerator without significant programming efforts, we discuss a set of MLIR compiler passes that automatically transform legacy code to utilize DX100. Experimental evaluations on 12 benchmarks spanning scientific computing, database, and graph applications show that DX100 achieves performance improvements of 2.6x over a multicore baseline and 2.0x over the state-of-the-art indirect prefetcher.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[130]
Yifan Yang, Joel S. Emer, and Daniel Sanchez. 2021. SpZip: Architectural Support for Effective Data Compression In Irregular Applications. In2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). 1069–1082. doi:10.1109/ISCA52012.2021.00087
arXiv 2021
-
[65]
Hyun Ryong Lee and Daniel Sanchez. 2024. Terminus: A Programmable Acceler- ator for Read and Update Operations on Sparse Data Structures. InProceedings of the 57th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO-57)(Austin, Texas, USA)(MICRO ’24)
2024
-
[1]
Junwhan Ahn, Sungpack Hong, Sungjoo Yoo, Onur Mutlu, and Kiyoung Choi
-
[2]
Sam Ainsworth and Timothy M. Jones. 2017. Software prefetching for indi- rect memory accesses. In2017 IEEE/ACM International Symposium on Code Generation and Optimization (CGO). 305–317. doi:10.1109/CGO.2017.7863749
arXiv 2017
-
[3]
Sam Ainsworth and Timothy M Jones. 2018. An event-triggered programmable prefetcher for irregular workloads.ACM Sigplan Notices53, 2 (2018), 578–592
2018
-
[5]
Daehyeon Baek, Soojin Hwang, Taekyung Heo, Daehoon Kim, and Jaehyuk Huh
-
[6]
Jean-Loup Baer and Tien-Fu Chen. 1991. An effective on-chip preloading scheme to reduce data access penalty. InProceedings of the 1991 ACM/IEEE conference on Supercomputing. 176–186
1991
-
[7]
David H Bailey, Eric Barszcz, John T Barton, David S Browning, Robert L Carter, Leonardo Dagum, Rod A Fatoohi, Paul O Frederickson, Thomas A Lasinski, Rob S Schreiber, et al. 1991. The NAS parallel benchmarks—summary and preliminary results. InProceedings of the 1991 ACM/IEEE Conference on Supercomputing. 158–165
1991
Show all 145 references
-
[8]
Mohammad Bakhshalipour, Pejman Lotfi-Kamran, and Hamid Sarbazi-Azad
-
[9]
Mohammad Bakhshalipour, Mehran Shakerinava, Pejman Lotfi-Kamran, and Hamid Sarbazi-Azad. 2019. Bingo spatial data prefetcher. In2019 IEEE Interna- tional Symposium on High Performance Computer Architecture (HPCA). IEEE, 399–411
2019
-
[10]
Vignesh Balaji and Brandon Lucia. 2022. Improving Locality of Irregular Updates with Hardware Assisted Propagation Blocking.. InHPCA. 543–557
2022
-
[11]
Tamer Özsu
Cagri Balkesen, Jens Teubner, Gustavo Alonso, and M. Tamer Özsu. 2013. Main- memory hash joins on multi-core CPUs: Tuning to the underlying hardware. In 2013 IEEE 29th International Conference on Data Engineering (ICDE). 362–373. doi:10.1109/ICDE.2013.6544839
2013
-
[12]
Ronald Barber, Guy Lohman, Ippokratis Pandis, Vijayshankar Raman, Richard Sidle, Gopi Attaluri, Naresh Chainani, Sam Lightstone, and David Sharpe. 2014. Memory-efficient hash joins.Proceedings of the VLDB Endowment8, 4 (2014), 353–364
2014
-
[14]
Abanti Basak, Shuangchen Li, Xing Hu, Sang Min Oh, Xinfeng Xie, Li Zhao, Xiaowei Jiang, and Yuan Xie. 2019. Analysis and Optimization of the Mem- ory Hierarchy for Graph Processing Workloads. In2019 IEEE International Symposium on High Performance Computer Architecture (HPCA)....
2019
-
[15]
Scott Beamer, Krste Asanovic, and David Patterson. 2012. Direction-optimizing Breadth-First Search. InSC ’12: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis. 1–10. doi:10. 1109/SC.2012.50
2012
-
[16]
Scott Beamer, Krste Asanović, and David Patterson. 2017. Reducing pagerank communication via propagation blocking. In2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 820–831
2017
-
[17]
Scott Beamer, Krste Asanović, and David Patterson. 2017. The GAP Benchmark Suite. arXiv:1508.03619 [cs.DC] https://arxiv.org/abs/1508.03619
2017 arXiv
-
[18]
Michael Bekerman, Stephan Jourdan, Ronny Ronen, Gilad Kirshenboim, Lihu Rappoport, Adi Yoaz, and Uri Weiser. 1999. Correlated load-address predictors. ACM SIGARCH Computer Architecture News27, 2 (1999), 54–63
1999
-
[19]
Rahul Bera, Anant V Nori, Onur Mutlu, and Sreenivas Subramoney. 2019. Dspatch: Dual spatial pattern prefetcher. InProceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture. 531–544
2019
-
[20]
Marsha J Berger and Joseph Oliger. 1984. Adaptive mesh refinement for hy- perbolic partial differential equations.J. Comput. Phys.53, 3 (1984), 484–512. doi:10.1016/0021-9991(84)90073-1
1984 doi
-
[21]
Reinhardt, Ali Saidi, Arkaprava Basu, Joel Hestness, Derek R
Nathan Binkert, Bradford Beckmann, Gabriel Black, Steven K. Reinhardt, Ali Saidi, Arkaprava Basu, Joel Hestness, Derek R. Hower, Tushar Krishna, Somayeh Sardashti, Rathijit Sen, Korey Sewell, Muhammad Shoaib, Nilay Vaish, Mark D. Hill, and David A. Wood. 2011. The gem5 simulat...
2011
-
[22]
Ulrik Brandes. 2001. A faster algorithm for betweenness centrality*.The Journal of Mathematical Sociology25, 2 (2001), 163–177. doi:10.1080/0022250X.2001. 9990249
2001 doi
-
[23]
2024.A gem5 experimental repo in order to explore Data-dependent Access (DDA).https: //github.com/xjtuiair-cag/gem5_dda/tree/dmp-paper
Institute of AI & Robotics of Xi’an Jiaotong Univiersity CAG group. 2024.A gem5 experimental repo in order to explore Data-dependent Access (DDA).https: //github.com/xjtuiair-cag/gem5_dda/tree/dmp-paper
2024
-
[24]
Yuhan Chen, Alireza Khadem, Xin He, Nishil Talati, Tanvir Ahmed Khan, and Trevor Mudge. 2023. PEDAL: A Power Efficient GCN Accelerator with Multiple DAtafLows. In2023 Design, Automation & Test in Europe Conference & Exhibition (DATE). 1–6. doi:10.23919/DATE56975.2023.10137240
2023
-
[25]
Yunji Chen, Tao Luo, Shaoli Liu, Shijin Zhang, Liqiang He, Jia Wang, Ling Li, Tianshi Chen, Zhiwei Xu, Ninghui Sun, et al . 2014. Dadiannao: A machine- learning supercomputer. In2014 47th Annual IEEE/ACM International Symposium on Microarchitecture. IEEE, 609–622
2014
-
[26]
Emer, and Vivienne Sze
Yu-Hsin Chen, Tushar Krishna, Joel S. Emer, and Vivienne Sze. 2017. Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks.IEEE Journal of Solid-State Circuits52, 1 (2017), 127–138. doi:10.1109/ JSSC.2016.2616357
2017
-
[27]
Yuan Chou. 2007. Low-cost epoch-based correlation prefetching for commercial applications. In40th Annual IEEE/ACM International Symposium on Microarchi- tecture (MICRO 2007). IEEE, 301–313
2007
-
[28]
Vidushi Dadu, Jian Weng, Sihao Liu, and Tony Nowatzki. 2019. Towards general purpose acceleration by exploiting common data-dependence forms. InProceed- ings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture. 924–939
2019
-
[29]
James Dundas and Trevor Mudge. 1997. Improving data cache performance by pre-executing instructions under a cache miss. InProceedings of the 11th Interna- tional Conference on Supercomputing(Vienna, Austria)(ICS ’97). Association for Computing Machinery, New York, NY, USA, 68–...
1997
-
[30]
R D Falgout, J E Jones, and U M Yang. 2004. The Design and Implementation of hypre, a Library of Parallel High Performance Preconditioners.Lecture Notes in Computational Science and Engineering51 (7 2004). https://www.osti.gov/ biblio/875356
2004
-
[31]
Yu Feng, Gunnar Hammonds, Yiming Gan, and Yuhao Zhu. 2022. Crescent: taming memory irregularities for accelerating deep point cloud analytics. In Proceedings of the 49th Annual International Symposium on Computer Architecture. 962–977. 13 ISCA ’25, June 21–25, 2025, Tokyo, Jap...
2022
-
[32]
Bruce Fleischer, Sunil Shukla, Matthew Ziegler, Joel Silberman, Jinwook Oh, Vijavalakshmi Srinivasan, Jungwook Choi, Silvia Mueller, Ankur Agrawal, Tina Babinsky, et al. 2018. A scalable multi-TeraOPS deep learning processor core for AI trainina and inference. In2018 IEEE symp...
2018
-
[33]
Gelin Fu, Tian Xia, Zhongpei Luo, Ruiyang Chen, Wenzhe Zhao, and Pengju Ren. 2024. Differential-Matching Prefetcher for Indirect Memory Access. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 439–453. doi:10.1109/HPCA57654.2024.00040
2024
-
[34]
Michael Gittings, Robert Weaver, Michael Clover, Thomas Betlach, Nelson Byrne, Robert Coker, Edward Dendy, Robert Hueckstaedt, Kim New, W Rob Oakes, Dale Ranta, and Ryan Stefan. 2008. The RAGE radiation-hydrodynamic code. Computational Science & Discovery1, 1 (Nov. 2008), 0150...
2008 doi
-
[35]
John W. Grove. 2019. Eulerian Applications Project - xRage Introduction and Overview. (5 2019). doi:10.2172/1532688
2019 doi
-
[36]
Yufeng Gu, Arun Subramaniyan, Tim Dunn, Alireza Khadem, Kuan-Yu Chen, Somnath Paul, Md Vasimuddin, Sanchit Misra, David Blaauw, Satish Narayanasamy, and Reetuparna Das. 2023. GenDP: A Framework of Dynamic Programming Acceleration for Genome Sequencing Analysis. InProceedings o...
2023
-
[37]
Tae Jun Ham, Lisa Wu, Narayanan Sundaram, Nadathur Satish, and Margaret Martonosi. 2016. Graphicionado: A high-performance and energy-efficient accelerator for graph analytics. In2016 49th annual IEEE/ACM international symposium on microarchitecture (MICRO). IEEE, 1–13
2016
-
[38]
Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A Horowitz, and William J Dally. 2016. EIE: Efficient inference engine on compressed deep neural network.ACM SIGARCH Computer Architecture News44, 3 (2016), 243– 254
2016
-
[39]
Milad Hashemi, Khubaib, Eiman Ebrahimi, Onur Mutlu, and Yale N. Patt. 2016. Accelerating Dependent Cache Misses with an Enhanced Memory Controller. In 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA). 444–455. doi:10.1109/ISCA.2016.46
2016 doi
-
[40]
Xin He, Subhankar Pal, Aporva Amarnath, Siying Feng, Dong-Hyeon Park, Austin Rovinski, Haojie Ye, Yuhan Chen, Ronald Dreslinski, and Trevor Mudge
-
[41]
Kartik Hegde, Hadi Asghari-Moghaddam, Michael Pellauer, Neal Crago, Aamer Jaleel, Edgar Solomonik, Joel Emer, and Christopher W Fletcher. 2019. Extensor: An accelerator for sparse tensor algebra. InProceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchite...
2019
-
[42]
Paul Henning and USDOE National Nuclear Security Administration. 2023. Ume: Unstructured Mesh Explorations. doi:10.11578/dc.20230602.5
2023 doi
-
[43]
2013.HPCG Bench- mark Technical Specification
Michael Allen Heroux, Jack Dongarra, and Piotr Luszczek. 2013.HPCG Bench- mark Technical Specification. Technical Report. Sandia National Lab. (SNL-NM), Albuquerque, NM (United States). doi:10.2172/1113870
2013 doi
-
[44]
Byungchul Hong, Gwangsun Kim, Jung Ho Ahn, Yongkee Kwon, Hongsik Kim, and John Kim. 2016. Accelerating linked-list traversal through near- data processing. InProceedings of the 2016 International Conference on Parallel Architectures and Compilation. 113–124
2016
-
[45]
Mark Horowitz. 2014. 1.1 Computing’s energy problem (and what we can do about it). In2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC). 10–14. doi:10.1109/ISSCC.2014.6757323
2014
-
[46]
Zhigang Hu, Margaret Martonosi, and Stefanos Kaxiras. 2003. TCP: Tag corre- lating prefetchers. InThe Ninth International Symposium on High-Performance Computer Architecture, 2003. HPCA-9 2003. Proceedings.IEEE, 317–326
2003
-
[47]
Ibrahim Hur and Calvin Lin. 2004. Adaptive history-based memory schedulers. In37th International Symposium on Microarchitecture (MICRO-37’04). IEEE, 343– 354
2004
-
[48]
2025.Intel®64 and IA-32 Architectures Software Developer Manu- als
Intel. 2025.Intel®64 and IA-32 Architectures Software Developer Manu- als. https://www.intel.com/content/www/us/en/developer/articles/technical/ intel-sdm.html
2025
-
[49]
Engin Ipek, Onur Mutlu, José F Martínez, and Rich Caruana. 2008. Self- optimizing memory controllers: A reinforcement learning approach.ACM SIGARCH Computer Architecture News36, 3 (2008), 39–50
2008
-
[50]
Akanksha Jain and Calvin Lin. 2013. Linearizing irregular memory accesses for improved correlated prefetching. InProceedings of the 46th Annual IEEE/ACM International Symposium on Microarchitecture. 247–259
2013
-
[51]
Saba Jamilan, Tanvir Ahmed Khan, Grant Ayers, Baris Kasikci, and Heiner Litz
-
[52]
Supreet Jeloka, Naveen Bharathwaj Akesh, Dennis Sylvester, and David Blaauw
-
[53]
Doug Joseph and Dirk Grunwald. 1997. Prefetching using markov predictors. In Proceedings of the 24th annual international symposium on Computer architecture. 252–263
1997
-
[54]
Liu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks, Vikas Chandra, Utku Diril, Amin Firoozshahian, Kim Hazelwood, Bill Jia, Hsien-Hsin S Lee, et al. 2020. Recnmp: Accelerating personalized recommendation with near- memory processing. In2020 ACM/IEEE 47th Annual Internatio...
2020
-
[55]
Muneeb Khan and Erik Hagersten. 2014. Resource conscious prefetching for irregular applications in multicores. In2014 International Conference on Embed- ded Computer Systems: Architectures, Modeling, and Simulation (SAMOS XIV). IEEE, 34–43
2014
-
[56]
Lee, Eric Sedlar, Anthony D
Changkyu Kim, Tim Kaldewey, Victor W. Lee, Eric Sedlar, Anthony D. Nguyen, Nadathur Satish, Jatin Chhugani, Andrea Di Blas, and Pradeep Dubey. 2009. Sort vs. Hash revisited: fast join implementation on modern multi-core CPUs. Proc. VLDB Endow.2, 2 (Aug. 2009), 1378–1389. doi:1...
2009
-
[57]
Vladimir Kiriansky, Yunming Zhang, and Saman Amarasinghe. 2016. Optimizing indirect memory references with milk. InProceedings of the 2016 International Conference on Parallel Architectures and Compilation. 299–312
2016
-
[58]
Onur Kocberber, Boris Grot, Javier Picorel, Babak Falsafi, Kevin Lim, and Parthasarathy Ranganathan. 2013. Meet the walkers: accelerating index traver- sals for in-memory databases. InProceedings of the 46th Annual IEEE/ACM International Symposium on Microarchitecture(Davis, C...
2013
-
[59]
Akhilesh Kumar, Don Soltis, Irma Esmer, Adi Yoaz, and Sailesh Kottapalli. 2017. The new Intel Xeon scalable processor (formerly skylake-SP). InIEEE Hot Chips Symposium (HCS)
2017
-
[60]
Snehasish Kumar, Arrvindh Shriraman, Vijayalakshmi Srinivasan, Dan Lin, and Jordon Phillips. 2014. SQRL: Hardware accelerator for collecting software data structures. InProceedings of the 23rd international conference on Parallel architectures and compilation. 475–476
2014
-
[61]
Rossbach, and Em- mett Witchel
Youngjin Kwon, Hangchen Yu, Simon Peter, Christopher J. Rossbach, and Em- mett Witchel. 2016. Coordinated and Efficient Huge Page Management with Ingens. In12th USENIX Symposium on Operating Systems Design and Imple- mentation (OSDI 16). USENIX Association, Savannah, GA, 705–7...
2016
-
[62]
Nagesh B Lakshminarayana and Hyesoon Kim. 2014. Spare register aware prefetching for graph algorithms on GPUs. In2014 IEEE 20th international symposium on high performance computer architecture (HPCA). IEEE, 614–625
2014
-
[63]
Chris Lattner, Mehdi Amini, Uday Bondhugula, Albert Cohen, Andy Davis, Jacques Pienaar, River Riddle, Tatiana Shpeisman, Nicolas Vasilache, and Olek- sandr Zinenko. 2021. MLIR: Scaling compiler infrastructure for domain specific computation. In2021 IEEE/ACM International Sympo...
2021
-
[64]
Patrick Lavin, Jeffrey Young, Richard Vuduc, Jason Riedy, Aaron Vose, and Daniel Ernst. 2021. Evaluating Gather and Scatter Performance on CPUs and GPUs. In Proceedings of the International Symposium on Memory Systems(Washington, DC, USA)(MEMSYS ’20). Association for Computing...
2021
-
[66]
Ruipeng Li and Ulrike M. Yang. 2023. AMG2023. [Computer Software] https: //doi.org/10.11578/dc.20230413.1. doi:10.11578/dc.20230413.1
2023 doi
-
[67]
Xinyu Li, Lei Liu, Shengjie Yang, Lu Peng, and Jiefan Qiu. 2019. Thinking about A New Mechanism for Huge Page Management. InProceedings of the 10th ACM SIGOPS Asia-Pacific Workshop on Systems(Hangzhou, China)(APSys ’19). Association for Computing Machinery, New York, NY, USA, ...
2019
-
[68]
Mikko H Lipasti, William J Schmidt, Steven R Kunkel, and Robert R Roediger
-
[69]
MTP Liska, Koushik Chatterjee, D Issa, Doosoo Yoon, N Kaaz, A Tchekhovskoy, D Van Eijnatten, G Musoke, C Hesp, V Rohoza, et al . 2022. H-AMR: A New GPU-accelerated GRMHD Code for Exascale Computing with 3D Adaptive Mesh Refinement and Local Adaptive Time Stepping.The Astrophys...
2022
-
[70]
Los Alamos National Laboratory (LANL). 2024. ATS-5: The Fifth Advanced Technology System in the Advanced Simulation and Computing Program. https: //mission.lanl.gov/advanced-simulation-and-computing/platforms/ats-5/. Ac- cessed: 2024-11-15
2024
-
[71]
Nisa Bostancı, Ataberk Olgun, A
Haocong Luo, Yahya Can Tuğrul, F. Nisa Bostancı, Ataberk Olgun, A. Giray Yağlıkçı, , and Onur Mutlu. 2023. Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator. 14 DX100: A Programmable Data Access Accelerator for Indirection ISCA ’25, June 21–25, 2025, Tokyo, Japan
2023
-
[72]
Manegold, P
S. Manegold, P. Boncz, and M. Kersten. 2002. Optimizing main-memory join on modern hardware.IEEE Transactions on Knowledge and Data Engineering14, 4 (2002), 709–730. doi:10.1109/TKDE.2002.1019210
2002 arXiv
-
[73]
Sally A McKee, William A Wulf, James H Aylor, Robert H Klenke, Maximo H Salinas, Sung I Hong, and Dee AB Weikle. 2000. Dynamic access ordering for streamed computations.IEEE Trans. Comput.49, 11 (2000), 1255–1271
2000
-
[74]
Meyer and P
U. Meyer and P. Sanders. 2003. Δ-stepping: a parallelizable shortest path algo- rithm.J. Algorithms49, 1 (Oct. 2003), 114–152. doi:10.1016/S0196-6774(03)00076- 2
2003 doi
-
[75]
Theodore Michailidis, Alex Delis, and Mema Roussopoulos. 2019. MEGA: over- coming traditional problems with OS huge page management. InProceedings of the 12th ACM International Conference on Systems and Storage(Haifa, Is- rael)(SYSTOR ’19). Association for Computing Machinery,...
2019
-
[76]
Pierre Michaud. 2016. Best-offset hardware prefetching. In2016 IEEE Interna- tional Symposium on High Performance Computer Architecture (HPCA). IEEE, 469–480
2016
-
[77]
William S Moses, Lorenzo Chelini, Ruizhe Zhao, and Oleksandr Zinenko. 2021. Polygeist: Raising C to polyhedral MLIR. In2021 30th International Conference on Parallel Architectures and Compilation Techniques (PACT). IEEE, 45–59
2021
-
[78]
Moshovos
A. Moshovos. 2005. RegionScout: exploiting coarse grain sharing in snoop-based coherence. In32nd International Symposium on Computer Architecture (ISCA’05). 234–245. doi:10.1109/ISCA.2005.42
2005 doi
-
[79]
Todd Mowry and Anoop Gupta. 1991. Tolerating latency through software- controlled prefetching in shared-memory multiprocessors.Journal of parallel and Distributed Computing12, 2 (1991), 87–106
1991
-
[80]
Mowry, Monica S
Todd C. Mowry, Monica S. Lam, and Anoop Gupta. 1992. Design and evaluation of a compiler algorithm for prefetching. InProceedings of the Fifth International Conference on Architectural Support for Programming Languages and Operating Systems(Boston, Massachusetts, USA)(ASPLOS V...
1992
-
[81]
Anurag Mukkara, Nathan Beckmann, Maleen Abeydeera, Xiaosong Ma, and Daniel Sanchez. 2018. Exploiting locality in graph analytics through hardware- accelerated traversal scheduling. In2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1–14
2018
-
[82]
Anurag Mukkara, Nathan Beckmann, and Daniel Sanchez. 2019. PHI: Architec- tural support for synchronization-and bandwidth-efficient commutative scatter updates. InProceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture. 1009–1022
2019
-
[83]
Francisco Muñoz-Martínez, Raveesh Garg, Michael Pellauer, José L Abellán, Manuel E Acacio, and Tushar Krishna. 2023. Flexagon: A multi-dataflow sparse- sparse matrix multiplication accelerator for efficient dnn processing. InPro- ceedings of the 28th ACM International Conferen...
2023
-
[84]
Onur Mutlu, Hyesoon Kim, and Y.N. Patt. 2005. Techniques for efficient pro- cessing in runahead execution engines. In32nd International Symposium on Computer Architecture (ISCA’05). 370–381. doi:10.1109/ISCA.2005.49
2005 doi
-
[85]
Onur Mutlu and Thomas Moscibroda. 2008. Parallelism-aware batch scheduling: Enhancing both performance and fairness of shared DRAM systems.ACM SIGARCH Computer Architecture News36, 3 (2008), 63–74
2008
-
[86]
Mutlu, J
O. Mutlu, J. Stark, C. Wilkerson, and Y.N. Patt. 2003. Runahead execution: an alternative to very large instruction windows for out-of-order processors. In The Ninth International Symposium on High-Performance Computer Architecture,
2003
-
[87]
Ajeya Naithani, Sam Ainsworth, Timothy M Jones, and Lieven Eeckhout. 2021. Vector runahead. In2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 195–208
2021
-
[88]
Ajeya Naithani, Josué Feliu, Almutaz Adileh, and Lieven Eeckhout. 2020. Precise runahead execution. In2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 397–410
2020
-
[89]
Ajeya Naithani, Jaime Roelandts, Sam Ainsworth, Timothy M Jones, and Lieven Eeckhout. 2023. Decoupled vector runahead. InProceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture. 17–31
2023
-
[90]
Ajeya Naithani, Jaime Roelandts, Sam Ainsworth, Timothy M Jones, and Lieven Eeckhout. 2024. Decoupled Vector Runahead for Prefetching Nested Memory- Access Chains.IEEE Micro(2024)
2024
-
[91]
Karthik Nilakant, Valentin Dalibard, Amitabha Roy, and Eiko Yoneki. 2014. PrefEdge: SSD prefetcher for large-scale graph traversal. InProceedings of Inter- national Conference on Systems and Storage. 1–12
2014
-
[92]
Dimin Niu, Shuangchen Li, Yuhao Wang, Wei Han, Zhe Zhang, Yijin Guan, Tianchan Guan, Fei Sun, Fei Xue, Lide Duan, et al. 2022. 184QPS/W 64Mb/mm 2 3D logic-to-DRAM hybrid bonding with process-near-memory engine for rec- ommendation system. In2022 IEEE International Solid-State ...
2022
-
[93]
Marcelo Orenes-Vera, Aninda Manocha, Jonathan Balkind, Fei Gao, Juan L Aragón, David Wentzlaff, and Margaret Martonosi. 2022. Tiny but mighty: designing and realizing scalable latency tolerance for manycore SoCs. InPro- ceedings of the 49th Annual International Symposium on Co...
2022
-
[94]
Subhankar Pal, Jonathan Beaumont, Dong-Hyeon Park, Aporva Amarnath, Siy- ing Feng, Chaitali Chakrabarti, Hun-Seok Kim, David Blaauw, Trevor Mudge, and Ronald Dreslinski. 2018. Outerspace: An outer product based sparse ma- trix multiplication accelerator. In2018 IEEE Internatio...
2018
-
[95]
Seshadhri, and Nishil Talati
Yunjie Pan, Omkar Bhalerao, C. Seshadhri, and Nishil Talati. 2024. Accurate and Fast Estimation of Temporal Motifs using Path Sampling. InProceedings of the International Conference on Data Mining (ICDM 2024). IEEE
2024
-
[96]
Yunjie Pan, Jiecao Yu, Andrew Lukefahr, Reetuparna Das, and Scott Mahlke
-
[97]
Gopinath
Ashish Panwar, Aravinda Prasad, and K. Gopinath. 2018. Making Huge Pages Actually Useful.SIGPLAN Not.53, 2 (March 2018), 679–692. doi:10.1145/3296957. 3173203
2018 doi
-
[98]
Irma Esmer Papazian, Sailesh Kottapalli, Jeff Baxter, Jeff Chamberlain, Geetha Vedaraman, and Brian Morris. 2015. Ivy Bridge server: A converged design. IEEE Micro35, 2 (2015), 16–25
2015
-
[99]
Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rang- harajan Venkatesan, Brucek Khailany, Joel Emer, Stephen W Keckler, and William J Dally. 2017. SCNN: An accelerator for compressed-sparse convo- lutional neural networks.ACM SIGARCH computer architecture n...
2017
-
[100]
Leeor Peled, Shie Mannor, Uri Weiser, and Yoav Etsion. 2015. Semantic locality and context-based prefetching using reinforcement learning. InProceedings of the 42nd Annual International Symposium on Computer Architecture. 285–297
2015
-
[101]
Yue Peng, Bailin Deng, Juyong Zhang, Fanyu Geng, Wenjie Qin, and Ligang Liu
-
[102]
Scott Rixner. 2004. Memory Controller Optimizations for Web Servers. In 37th International Symposium on Microarchitecture (MICRO-37’04). 355–366. doi:10.1109/MICRO.2004.22
2004 doi
-
[103]
Dally, Ujval J
Scott Rixner, William J. Dally, Ujval J. Kapasi, Peter Mattson, and John D. Owens
-
[104]
Jaime Roelandts, Ajeya Naithani, Sam Ainsworth, Timothy M Jones, and Lieven Eeckhout. 2024. Scalar Vector Runahead. (2024)
2024
-
[105]
Alexander Rucker, Matthew Vilim, Tian Zhao, Yaqi Zhang, Raghu Prabhakar, and Kunle Olukotun. 2021. Capstan: A vector RDA for sparsity. InMICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture. 1022–1035
2021
-
[106]
Schwedock, Piratach Yoovidhya, Jennifer Seibert, and Nathan Beck- mann
Brian C. Schwedock, Piratach Yoovidhya, Jennifer Seibert, and Nathan Beck- mann. 2022. täk ¯o: a polymorphic cache hierarchy for general-purpose optimiza- tion of data movement. InProceedings of the 49th Annual International Sympo- sium on Computer Architecture(New York, New Y...
2022
-
[107]
Vivek Seshadri, Thomas Mullins, Amirali Boroumand, Onur Mutlu, Phillip B Gibbons, Michael A Kozuch, and Todd C Mowry. 2015. Gather-scatter DRAM: In-DRAM address translation to improve the spatial locality of non-unit strided accesses. InProceedings of the 48th International Sy...
2015
-
[108]
Jun Shao and Brian T Davis. 2007. A burst scheduling access reordering mecha- nism. In2007 IEEE 13th International Symposium on High Performance Computer Architecture. IEEE, 285–294
2007
-
[109]
Kevin Sheridan, Christopher Scott, Jered Dominguez-Trujillo, Agustin Vaca Valverde, Patrick Lavin, Galen Shipman, Richard Vuduc, and Jeffrey Young
-
[110]
Manjunath Shevgoor, Sahil Koladiya, Rajeev Balasubramonian, Chris Wilker- son, Seth H Pugsley, and Zeshan Chishti. 2015. Efficiently prefetching complex address patterns. InProceedings of the 48th International Symposium on Microar- chitecture. 141–152
2015
-
[111]
ACM Transactions on Graphics (TOG)37, 4 (2018), 1–14
Anderson acceleration for geometry optimization and physics simulation. ACM Transactions on Graphics (TOG)37, 4 (2018), 1–14
2018
-
[112]
Galen M Shipman, Jered Dominguez-Trujillo, Kevin Sheridan, and Sriram Swami- narayan. 2022. Assessing the Memory Wall in Complex Codes. In2022 IEEE/ACM Workshop on Memory Centric High Performance Computing (MCHPC). IEEE, 30– 35
2022
-
[113]
Galen M Shipman, Jason Pruet, David Daniel, Josh Dolence, Gary Grider, Brian M Haines, Aimee Hungerford, Stephen Poole, Tim Randles, Sriram Swaminarayan, et al. 2022. The future of HPC in nuclear security.IEEE Internet Computing27, 1 (2022), 16–23
2022
-
[114]
Marco Siracusa, Víctor Soria-Pardos, Francesco Sgherzi, Joshua Randall, Dou- glas J Joseph, Miquel Moretó Planas, and Adrià Armejach. 2023. A Tensor Marshaling Unit for Sparse Tensor Algebra on General-Purpose Processors. In 15 ISCA ’25, June 21–25, 2025, Tokyo, Japan Khadem e...
2023
-
[115]
James E. Smith. 1982. Decoupled access/execute computer architectures. In Proceedings of the 9th Annual Symposium on Computer Architecture(Austin, Texas, USA)(ISCA ’82). IEEE Computer Society Press, Washington, DC, USA, 112–119
1982
-
[116]
2011.A primer on memory consistency and cache coherence
Daniel Sorin, Mark Hill, and David Wood. 2011.A primer on memory consistency and cache coherence. Morgan & Claypool Publishers
2011
-
[117]
Sriseshan Srikanth, Anirudh Jain, Thomas M Conte, Erik P Debenedictis, and Jeanine Cook. 2021. SortCache: intelligent cache management for accelerating sparse data workloads.ACM Transactions on Architecture and Code Optimization (TACO)18, 4 (2021), 1–24
2021
-
[118]
Aaron Stillmaker and Bevan Baas. 2017. Scaling equations for the accurate prediction of CMOS device performance from 180nm to 7nm.Integration58 (2017), 74–81. doi:10.1016/j.vlsi.2017.02.002
2017 doi
-
[119]
Nishil Talati, Kyle May, Armand Behroozi, Yichen Yang, Kuba Kaszyk, Christos Vasiladiotis, Tarunesh Verma, Lu Li, Brandon Nguyen, Jiawen Sun, John Magnus Morton, Agreen Ahmadi, Todd Austin, Michael O’Boyle, Scott Mahlke, Trevor Mudge, and Ronald Dreslinski. 2021. Prodigy: Impr...
2021
-
[120]
Tam, Harry Muljono, Min Huang, Sitaraman Iyer, Kalapi Royneogi, Nagmohan Satti, Rizwan Qureshi, Wei Chen, Tom Wang, Hubert Hsieh, Sujal Vora, and Eddie Wang
Simon M. Tam, Harry Muljono, Min Huang, Sitaraman Iyer, Kalapi Royneogi, Nagmohan Satti, Rizwan Qureshi, Wei Chen, Tom Wang, Hubert Hsieh, Sujal Vora, and Eddie Wang. 2018. SkyLake-SP: A 14nm 28-Core xeon®processor. In 2018 IEEE International Solid-State Circuits Conference - ...
2018
-
[121]
Kai Troester and Ravi Bhargava. 2023. AMD Next Generation “Zen 4” Core and 4th Gen AMD EPYC™9004 Server CPU. In2023 IEEE Hot Chips 35 Symposium (HCS). IEEE Computer Society, 1–25
2023
-
[122]
Matthew Vilim, Alexander Rucker, and Kunle Olukotun. 2021. Aurochs: An architecture for dataflow threads. In2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 402–415
2021
-
[123]
Yossi Shiloach and Uzi Vishkin. 1982. An O(logn) parallel connectivity algorithm. Journal of Algorithms3, 1 (1982), 57–67. doi:10.1016/0196-6774(82)90008-6
1982 doi
-
[124]
Thomas F Wenisch, Michael Ferdman, Anastasia Ailamaki, Babak Falsafi, and Andreas Moshovos. 2009. Practical off-chip meta-data for temporal memory streaming. In2009 IEEE 15th International Symposium on High Performance Computer Architecture. IEEE, 79–90
2009
-
[125]
2025.Skylake (client) - Microarchitectures - Intel
WikiChip. 2025.Skylake (client) - Microarchitectures - Intel. https://en.wikichip. org/wiki/intel/microarchitectures/skylake_(client)
2025
-
[126]
Hao Wu, Krishnendra Nathella, Joseph Pusdesris, Dam Sunwoo, Akanksha Jain, and Calvin Lin. 2019. Temporal prefetching without the off-chip meta- data. InProceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture. 996–1008
2019
-
[127]
Xinfeng Xie, Zheng Liang, Peng Gu, Abanti Basak, Lei Deng, Ling Liang, Xing Hu, and Yuan Xie. 2021. SpaceA: Sparse matrix vector multiplication on processing-in-memory accelerator. In2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 570–583
2021
-
[128]
Xing and A
W. Xing and A. Ghorbani. 2004. Weighted PageRank algorithm. InProceedings. Second Annual Conference on Communication Networks and Services Research, 2004.305–314. doi:10.1109/DNSR.2004.1344743
2004 arXiv
-
[129]
Mingyu Yan, Xing Hu, Shuangchen Li, Abanti Basak, Han Li, Xin Ma, Itir Akgun, Yujing Feng, Peng Gu, Lei Deng, et al. 2019. Alleviating irregularity in graph analytics acceleration: A hardware/software co-design approach. InProceedings of the 52nd Annual IEEE/ACM International ...
2019
-
[131]
Hughes, Nadathur Satish, and Srinivas Devadas
Xiangyao Yu, Christopher J. Hughes, Nadathur Satish, and Srinivas Devadas
-
[132]
Chao Zhang, Maximilian Bremer, Cy Chan, John Shalf, and Xiaochen Guo. 2022. ASA: A ccelerating S parse A ccumulation in Column-wise SpGEMM.ACM Transactions on Architecture and Code Optimization (TACO)19, 4 (2022), 1–24
2022
-
[133]
Guowei Zhang and Daniel Sanchez. 2019. Leveraging caches to accelerate hash tables and memoization. InProceedings of the 52nd annual IEEE/ACM international symposium on microarchitecture. 440–452
2019
-
[134]
Shijin Zhang, Zidong Du, Lei Zhang, Huiying Lan, Shaoli Liu, Ling Li, Qi Guo, Tianshi Chen, and Yunji Chen. 2016. Cambricon-X: An accelerator for sparse neural networks. In2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1–12. 16 DX100: A P...
2016
-
[135]
Shu-Ting Wang, Hanyang Xu, Amin Mamandipoor, Rohan Mahapatra, Byung Hoon Ahn, Soroush Ghodrati, Krishnan Kailas, Mohammad Alian, and Hadi Esmaeilzadeh. 2024. Data motion acceleration: Chaining cross-domain multi accelerators. In2024 IEEE International Symposium on High-Perform...
2024
-
[144]
InProceedings of the 48th International Symposium on Microarchitecture(Waikiki, Hawaii)(MICRO-48)
IMP: indirect memory prefetcher. InProceedings of the 48th International Symposium on Microarchitecture(Waikiki, Hawaii)(MICRO-48). Association for Computing Machinery, New York, NY, USA, 178–190. doi:10.1145/2830772. 2830807
-
[1995]
InProceedings of the 28th annual international symposium on Microarchitecture
SPAID: Software prefetching in pointer-and call-intensive environments. InProceedings of the 28th annual international symposium on Microarchitecture. IEEE, 231–236
-
[2000]
InProceedings of the 27th Annual International Symposium on Computer Architecture(Vancouver, British Columbia, Canada) (ISCA ’00)
Memory access scheduling. InProceedings of the 27th Annual International Symposium on Computer Architecture(Vancouver, British Columbia, Canada) (ISCA ’00). Association for Computing Machinery, New York, NY, USA, 128–138. doi:10.1145/339647.339668
- [2003]
-
[2015]
InProceedings of the 42nd Annual International Symposium on Computer Architecture
A scalable processing-in-memory accelerator for parallel graph process- ing. InProceedings of the 42nd Annual International Symposium on Computer Architecture. 105–117
-
[2016]
doi:10.1109/JSSC.2016.2515510
A 28 nm Configurable Memory (TCAM/BCAM/SRAM) Using Push-Rule 6T Bit Cell Enabling Logic-in-Memory.IEEE Journal of Solid-State Circuits51, 4 (2016), 1009–1021. doi:10.1109/JSSC.2016.2515510
2016
-
[2018]
In2018 IEEE International Symposium on High Performance Computer Architecture (HPCA)
Domino temporal data prefetcher. In2018 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 131–142
-
[2020]
InProceedings of the 34th ACM international conference on supercomputing
Sparse-TPU: Adapting systolic arrays for sparse matrices. InProceedings of the 34th ACM international conference on supercomputing. 1–12
-
[2021]
In2021 30th International Conference on Parallel Architectures and Compilation Techniques (PACT)
Innersp: A memory efficient sparse matrix multiplication accelerator with locality-aware inner product processing. In2021 30th International Conference on Parallel Architectures and Compilation Techniques (PACT). IEEE, 116–128
-
[2022]
InProceedings of the Seventeenth European Conference on Computer Systems(Rennes, France)(EuroSys ’22)
APT-GET: profile-guided timely software prefetching. InProceedings of the Seventeenth European Conference on Computer Systems(Rennes, France)(EuroSys ’22). Association for Computing Machinery, New York, NY, USA, 747–764. doi:10. 1145/3492321.3519583
-
[2023]
BitSET: Bit-Serial Early Termination for Computation Reduction in Con- volutional Neural Networks.ACM Trans. Embed. Comput. Syst.22, 5s, Article 98 (Sept. 2023), 24 pages. doi:10.1145/3609093
2023 doi
-
[2024]
InProceedings of the International Symposium on Memory Systems(Wash- ington D.C., USA)
A Workflow for the Synthesis of Irregular Memory Access Microbench- marks. InProceedings of the International Symposium on Memory Systems(Wash- ington D.C., USA)
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.