REVIEW 3 major objections 3 minor 300 references
New Tools, Programming Models, and System Support for Processing-in-Memory Architectures
T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The dissertation seeks to establish that coordinated hardware and software support—benchmarks, compiler passes, a data-aware runtime, and a pattern-based programming model—can make processing-in-memory fast, efficient, and programmable.
desk verdict A solid, honest dissertation that compiles four already-published PIM systems papers; DAMOV and DaPPA are real, MIMDRAM/Proteus gains are simulation-bound and need in-DRAM primitive validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument rides on two analog in-DRAM primitives and three new ways of organizing them. All processing-using-DRAM work here builds on RowClone's row copy (an ACT-ACT-PRE command sequence that copies one row to another) and Ambit's triple-row activation, which makes the sense amplifier compute the majority of three rows in one step. MIMDRAM's central mechanism is per-mat instruction control inside a subarray, plus an intra-mat interconnect: different row segments run different PUD operations at the same time, and vector-to-scalar reduction happens natively. Proteus's central mechanism is scattering the bits of each word across separate subarrays (one bit per subarray), so the independent m
What would settle it
Prototype the MIMDRAM/Proteus subarray modifications (per-mat control, intra-subarray interconnect, native reduction columns, one-bit-per-subarray mapping) on an FPGA-controlled DDR4 testbed and measure whether triple-row activation and row-copy commands stay reliable across many chips, temperatures, and voltages. If the analog operations fail at modeled margins—or fabrication shows the added logic costs more than the claimed 1.11%/1.6% die area—the quantified gains (13.2x/173x performance, 582.4x/272x performance-per-watt) collapse. A cheaper check: confirm on real systems that operands in th
Extended reading notes
Core claim
The thesis statement is that end-to-end design of hardware and software support—benchmark suites and workload analysis, execution and programming models, compiler passes, and data-aware runtime mechanisms—can exploit the inherent parallelism of PIM architectures, ease their adoption, and enable large (factors or orders of magnitude) improvements in performance and energy efficiency. The evidence, on twelve real applications and 495 multi-programmed mixes, plus real UPMEM hardware for DaPPA: MIMDRAM reaches 13.2x and 173x the performance of the CPU and of SIMDRAM (the prior state-of-the-art processing-using-DRAM framework), with 582.4x and 272x the performance-per-watt; Proteus reaches 17x th
Load-bearing premise
The load-bearing assumption is that the simulated DRAM designs behave in a real chip as modeled: triple-row activation and row-copy operations must stay reliable under MIMDRAM's and Proteus's new access patterns, and the added per-subarray logic (per-mat control, intra-subarray interconnect, reduction columns) must be fabricable in a DRAM process at the claimed 1.11% and 1.6% die-area cost.
Editorial extensions
If this is right
- PUD no longer requires whole-row SIMD granularity: MIMDRAM shows that a single DRAM subarray can run several different operations at once, pointing toward efficient PIM execution of multi-programmed and irregular workloads.
- The two main PUD inefficiencies—rigid resource granularity and high per-operation latency—are separable and each is addressable: fine-grained control (MIMDRAM) for utilization, and bit-level parallelism plus carry-free arithmetic (Proteus) for latency.
- Dynamic precision reduction means PUD performance tracks the information content of data rather than the declared data type; since applications routinely over-provision bit widths, much of the work on leading zeros can be skipped at runtime.
- Pattern-based programming is sufficient to capture efficient PIM code: programs written abstractly can beat hand-tuned PIM implementations, suggesting the programming model is where PIM usability is won or lost, not only the hardware.
- If the thesis holds, PIM viability for new workloads can be predicted by DAMOV-style classification of their data-movement bottlenecks instead of ad-hoc profiling and manual porting.
Reading between the lines
- The quantified MIMDRAM and Proteus claims rest on simulation of analog DRAM behavior; the natural test is prototyping the new subarray structures (per-mat control, intra-subarray interconnect, bit-scattered layouts) on an FPGA-controlled DRAM testbed and measuring triple-row-activation and row-copy reliability across temperature, voltage, and chip variation—a step the dissertation's own discussion
- MIMDRAM and Proteus attack orthogonal inefficiencies (resource granularity versus operation latency) but are evaluated separately against the same SIMDRAM baseline; a combined substrate could compound their gains, which the dissertation does not itself build.
- The narrow-value insight generalizes beyond bit-serial PUD: processing-near-memory cores that execute fixed-width operations (such as UPMEM's DPUs) could also skip work on leading non-informative bits, so a precision-aware variant of DaPPA's generated code is a plausible follow-on.
- DAMOV's bottleneck classes could feed an automatic dispatch layer that chooses CPU, GPU, or PIM for each function at runtime; the dissertation uses the classes to explain suitability but does not construct that selector.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The dissertation proposes four contributions intended to ease adoption of processing-in-memory (PIM) systems: (1) DAMOV, a workload characterization methodology and data-movement benchmark suite built from a profiling study of 345 applications; (2) MIMDRAM, a hardware/software co-designed processing-using-DRAM (PUD) substrate that enables multiple-instruction, multiple-data (MIMD) execution inside DRAM subarrays; (3) Proteus, a data-aware runtime framework for PUD that dynamically reduces bit precision, exploits subarray-level parallelism, and selects data representation and arithmetic algorithms; and (4) DaPPA, a data-parallel programming framework targeting processing-near-memory (PNM) systems, validated on a real UPMEM system. The central thesis (Section 1.3) is that end-to-end design of hardware and software support for PIM can exploit the inherent parallelism of PIM architectures and deliver large (orders-of-magnitude) performance and energy-efficiency improvements. Headline results include MIMDRAM delivering 13.2x/0.22x/173x the performance and 582.4x/13612x/272x the performance-per-Watt of CPU/GPU/SIMDRAM across twelve applications; Proteus delivering 17x/7.3x/10.2x the performance per mm2 and 90.3x/21x/8.1x lower energy than CPU/GPU/SIMDRAM; and DaPPA improving end-to-end performance by 2.1x with a 94% reduction in lines of code on real UPMEM hardware.
Significance. If the results hold, the dissertation would be a substantial contribution: it combines a widely reusable benchmark suite and characterization methodology (DAMOV) with two PUD execution substrates, a runtime engine, and a programming framework, covering much of the software/hardware stack for PIM. Strong points that deserve credit: DAMOV is open-sourced and has already been used by other work; DaPPA is evaluated on a real, commercially available UPMEM PIM system; MIMDRAM and DAMOV are released as open-source artifacts; the text is structurally honest, including a self-flagged limitations section for DAMOV (Section 4.3.6) and sensitivity analyses for MIMDRAM (Section 5.6.4). The main weakness is that the order-of-magnitude PUD claims for MIMDRAM and Proteus rest on simulated DRAM designs whose core primitives are not validated at circuit level or on silicon. This asymmetry is significant: DAMOV and DaPPA are grounded in real systems, but MIMDRAM and Proteus carry the strongest quantitative claims in the thesis. The central thesis is therefore only partially established as stated; the PUD-specific claims are plausible design-space proposals but not yet demonstrated hardware results.
major comments (3)
- [§5.2, §5.5, §5.6.4–5.6.5] The headline MIMDRAM results (13.2x/0.22x/173x performance; 582.4x/13612x/272x performance-per-Watt vs. CPU/GPU/SIMDRAM) depend on new DRAM subarray mechanisms described in Section 5.2: per-mat instruction control, an intra-mat interconnect, concurrent activation of multiple mats, triple-row activation, and native vector reduction. The evaluation in Sections 5.5–5.6 uses DDR4 datasheet timings and analytic energy/area models, not circuit-level or silicon validation of these specific primitives. The sensitivity analysis in Section 5.6.4 varies subarray/bank counts, but it does not vary the timing, energy, or area of the underlying in-DRAM operations, which are the load-bearing assumptions. A concrete, proportionate request: report a sensitivity sweep over TRA/AAP/row-copy latencies and energies, and state the maximum primitive degradation under which the claimed improvements remain above
- [§6.3.1, §6.4, §6.5.5] Proteus similarly rests on unvalidated hardware assumptions: subarray-level parallelism (SALP) with bit-scattered operands across multiple subarrays, concurrent execution of in-DRAM primitives belonging to a single PUD operation, custom sense-amplifier/control logic, and redundant-binary representation with associated microprograms. The methodology (Section 6.4) uses analytic DRAM energy models and datasheet timing; the area analysis (Section 6.5.5) estimates a 1.6% DRAM-die overhead but does not validate manufacturability or the timing of the modified sense-amplifier/control paths. The reported 17x/7.3x/10.2x performance-per-mm2 and 90.3x/21x/8.1x energy reductions are therefore contingent on untested primitive costs. As with MIMDRAM, a sensitivity study on SALP activation limits and primitive timing/energy is needed before these claims can be taken as quantitative predictions.
- [§5.6.3, §6.5, and §5.4.2/§6.5.2] The main evaluation does not clearly separate PUD execution time from data-mapping, transposition, and representation-format-conversion overheads. Sections 5.4.2 and 6.5.2 describe these overheads, but the headline figures (e.g., Figure 5.8 and Figures 6.10–6.11) should state explicitly whether they are included in the reported end-to-end numbers. If they are excluded, the system-level benefits are overstated; if included, the paper should provide a breakdown so readers can see how much of the gain comes from the new PUD primitives versus the surrounding software/runtime support. This is important for calibrating the thesis statement's 'orders of magnitude' claim.
minor comments (3)
- [§1.4.3] Typo: 'Based on these two three key ideas ideas' should read 'Based on these three key ideas.' The same passage lists 'three limitations' and then says 'To solve the third limitation,' which is consistent, but the sentence structure should be cleaned up.
- [§2.2 and §4.3.6] The terminology shifts from 'processing-using-memory (PUM)' to 'processing-using-DRAM (PUD)' without an explicit statement that PUD is the DRAM-specific instance of PUM. A one-sentence clarification at first use would prevent confusion. Also, Section 4.3.6 is a genuine strength: it honestly lists limitations of the DAMOV methodology (e.g., profiling overhead and representativeness), and I do not consider those limitations disqualifying.
- [§7.5.1] Lines-of-code reduction is used as the primary programmability metric. This is standard and useful, but it is a weak proxy for programmer effort. Consider supplementing with a qualitative description of the programming steps eliminated (e.g., explicit data movement, alignment handling, DPU count management) or with a small user study, if available.
Circularity Check
No circularity found; the PUD gains are simulation outcomes built on published primitives and the real-system contributions are measured, so no claim reduces to its own inputs.
full rationale
The dissertation's derivation chain is self-contained with respect to the circularity patterns. DAMOV is a workload characterization and benchmark suite built from large-scale profiling (77K functions, 345 applications) and is validated by case studies; its conclusions do not presuppose MIMDRAM/Proteus/DaPPA. MIMDRAM's quantified claims (e.g., '13.2×/0.22×/173× the performance ... of the CPU/GPU/SIMDRAM baselines') are produced by simulating the proposed per-mat control, intra-subarray interconnect, and compiler passes under DDR4 datasheet timings and Ambit/RowClone operation semantics; those inputs do not contain the output speedups, so the result is not a fitted input called prediction. Proteus similarly derives its '17×, 7.3×, and 10.2× the performance per mm2' and energy reductions from a runtime cost model and evaluated microprograms, not from the definition of Proteus itself. DaPPA's '2.1×' end-to-end improvement and '94%' LOC reduction are measured on a real UPMEM system against PrIM hand-tuned code. The self-citations to SIMDRAM, PrIM, and the SAFARI ecosystem serve as published baselines and foundations, not as an unverified uniqueness theorem or a forced choice; SIMDRAM is an externally published architecture and PrIM is a published benchmark suite. The main weakness of the PUD chapters—that MIMDRAM/Proteus gains depend on unvalidated in-DRAM primitive timing, energy, and area models—is a validation/silicon-risk concern, not circularity: the simulation assumptions are stated and are not the same as the claimed results. Accordingly, no step reduces by construction to its own input.
Assumptions & free parameters
free parameters (4)
- MIMDRAM resource configuration (banks and subarrays allocated to PUD) =
16 banks, 64 subarrays per bank
- Proteus cost model thresholds (data mapping and arithmetic algorithm selection) =
not stated numerically
- DAMOV classification thresholds (MPKI boundary, locality clustering parameters) =
MPKI > 10 treated as high (borrowed from prior work)
- PUD primitive timing and energy parameters (TRA/AP/AAP latencies) =
from DDR4 datasheet timings in cited models
assumptions (5)
- domain assumption Triple row activation in DRAM produces a majority-of-three, and dual-contact cells produce NOT operations.
- domain assumption RowClone AAP semantics (ACT-ACT-PRE copies a row) remain valid in the proposed per-mat and fine-grained access designs.
- ad hoc to paper Proposed DRAM modifications are manufacturable within DRAM process constraints at stated area cost (1.11% MIMDRAM, 1.6% Proteus DRAM-die overhead).
- standard math RBR arithmetic limits carry propagation to two places, giving precision-independent addition latency.
- domain assumption Workload characteristics of the 345 applications (77K functions) from 37 benchmark suites are representative of 'modern workloads' at large.
invented entities (4)
-
MIMDRAM per-mat control and intra-mat interconnect (MIMD execution inside a DRAM subarray, native vector reduction)
independent evidence
-
Proteus parallelism-aware microprogram library, dynamic bit-precision engine, microprogram select unit
-
DaPPA data-parallel pattern APIs, dataflow programming interface, dynamic template-based compiler
independent evidence
-
DAMOV four data-movement metrics and six bottleneck classes
independent evidence
Cite this review
Pith. "Pith review of New Tools, Programming Models, and System Support for Processing-in-Memory Architectures." pith.science (2026). https://pith.science/paper/OO2L5BZC
@misc{pith2026250819868,
author = {Pith},
title = {Pith review of: New Tools, Programming Models, and System Support for Processing-in-Memory Architectures},
year = {2026},
howpublished = {\url{https://pith.science/paper/OO2L5BZC}},
note = {Machine review of arXiv:2508.19868}
}
read the original abstract
Our goal in this dissertation is to provide tools, programming models, and system support for PIM architectures (with a focus on DRAM-based solutions), to ease the adoption of PIM in current and future systems. To this end, we make at least four new major contributions. First, we introduce DAMOV, the first rigorous methodology to characterize memory-related data movement bottlenecks in modern workloads, and the first data movement benchmark suite. Second, we introduce MIMDRAM, a new hardware/software co-designed substrate that addresses the major current programmability and flexibility limitations of the bulk bitwise execution model of processing-using-DRAM (PUD) architectures. MIMDRAM enables the allocation and control of only the needed computing resources inside DRAM for PUD computing. Third, we introduce Proteus, the first hardware framework that addresses the high execution latency of bulk bitwise PUD operations in state-of-the-art PUD architectures by implementing a data-aware runtime engine for PUD. Proteus reduces the latency of PUD operations in three different ways: (i) Proteus concurrently executes independent in-DRAM primitives belong to a single PUD operation across DRAM arrays. (ii) Proteus dynamically reduces the bit-precision (and consequentially the latency and energy consumption) of PUD operations by exploiting narrow values (i.e., values with many leading zeros or ones). (iii) Proteus chooses and uses the most appropriate data representation and arithmetic algorithm implementation for a given PUD instruction transparently to the programmer. Fourth, we introduce DaPPA (data-parallel processing-in-memory architecture), a new programming framework that eases programmability for general-purpose PNM architectures by allowing the programmer to write efficient PIM-friendly code without the need to manage hardware resources explicitly.
Figures
Figures from the paper (53 more)
Reference graph
Works this paper leans on
-
[1]
Understanding and Improving the Latency of DRAM-Based Memory Systems,
Kevin K Chang, “Understanding and Improving the Latency of DRAM-Based Memory Systems, ” Ph.D. dissertation, Carnegie Mellon University, 2017
2017
-
[2]
A Signed Binary mMultiplication Technique,
Andrew D Booth, “A Signed Binary mMultiplication Technique, ”The Quarterly Journal of Mechanics and Applied Mathematics , 1951
1951
-
[3]
Multiplication of Many-Digital Numbers by Automatic Computers,
Anatolii Alekseevich Karatsuba and Yu P Ofman, “Multiplication of Many-Digital Numbers by Automatic Computers, ” inUSSR Academy of Sciences , 1962
1962
-
[4]
Phoenix Rebirth: Scalable MapReduce on a Large-Scale Shared-Memory System,
R. M. Yoo, A. Romano, and C. Kozyrakis, “Phoenix Rebirth: Scalable MapReduce on a Large-Scale Shared-Memory System, ” inIISWC, 2009
2009
-
[5]
SPEC CPU2017 Benchmarks,
Standard Performance Evaluation Corp., “SPEC CPU2017 Benchmarks, ” http://www. spec.org/cpu2017/
-
[6]
Benchmarking a New Paradigm: Experimental Analysis and Characterization of a Real Processing-in-Memory System,
Juan Gómez-Luna, Izzat El Hajj, Ivan Fernandez, Christina Giannoula, Geraldo F Oliveira, and Onur Mutlu, “Benchmarking a New Paradigm: Experimental Analysis and Characterization of a Real Processing-in-Memory System, ”IEEE Access, 2022
2022
-
[7]
Memory Scaling: A Systems Architecture Perspective,
Onur Mutlu, “Memory Scaling: A Systems Architecture Perspective, ” in IMW, 2013
2013
-
[8]
Research Problems and Opportunities in Memory Systems,
Onur Mutlu and Lavanya Subramanian, “Research Problems and Opportunities in Memory Systems, ”SUPERFRI, 2014
2014
Show all 300 references
-
[9]
The Tail at Scale,
Jeffrey Dean and Luiz André Barroso, “The Tail at Scale, ”CACM, 2013
2013
-
[10]
Profiling a Warehouse-Scale Computer,
Svilen Kanev, Juan Pablo Darago, Kim Hazelwood, Parthasarathy Ranganathan, Tipp Moseley, Gu-Yeon Wei, and David Brooks, “Profiling a Warehouse-Scale Computer, ” in ISCA, 2015
2015
-
[11]
Clearing the Clouds: A Study of Emerging Scale-Out Workloads on Modern Hardware,
Michael Ferdman, Almutaz Adileh, Onur Kocberber, Stavros Volos, Mohammad Al- isafaee, Djordje Jevdjic, Cansu Kaynak, Adrian Daniel Popescu, Anastasia Ailamaki, and Babak Falsafi, “Clearing the Clouds: A Study of Emerging Scale-Out Workloads on Modern Hardware, ” inASPLOS, 2012
2012
-
[12]
BigDataBench: A Big Data Bench- mark Suite from Internet Services,
Lei Wang, Jianfeng Zhan, Chunjie Luo, Yuqing Zhu, Qiang Yang, Yongqiang He, Wan- ling Gao, Zhen Jia, Yingjie Shi, Shujie Zhang et al., “BigDataBench: A Big Data Bench- mark Suite from Internet Services, ” inHPCA, 2014. 199 BIBLIOGRAPHY 200
2014
-
[13]
Google Workloads for Consumer Devices: Mitigating Data Movement Bottlenecks,
Amirali Boroumand, Saugata Ghose, Youngsok Kim, Rachata Ausavarungnirun, Eric Shiu, Rahul Thakur, Daehyun Kim, Aki Kuusela, Allan Knies, Parthasarathy Ran- ganathan et al., “Google Workloads for Consumer Devices: Mitigating Data Movement Bottlenecks, ” inASPLOS, 2018
2018
-
[14]
Google Neural Network Models for Edge Devices: Analyzing and Mitigating Machine Learning Inference Bottlenecks,
Amirali Boroumand, Saugata Ghose, Berkin Akin, Ravi Narayanaswami, Geraldo F Oliveira, Xiaoyu Ma, Eric Shiu, and Onur Mutlu, “Google Neural Network Models for Edge Devices: Analyzing and Mitigating Machine Learning Inference Bottlenecks, ” in PACT, 2021
2021
-
[15]
Google Neural Network Models for Edge Devices: Analyzing and Mitigating Machine Learning Inference Bottlenecks,
Amirali Boroumand, Saugata Ghose, Berkin Akin, Ravi Narayanaswami, Geraldo F Oliveira, Xiaoyu Ma, Eric Shiu, and Onur Mutlu, “Google Neural Network Models for Edge Devices: Analyzing and Mitigating Machine Learning Inference Bottlenecks, ” arXiv:2109.14320 [cs.AR], 2021
2021 arXiv
-
[16]
En- abling Practical Processing in and Near Memory for Data-Intensive Computing,
Onur Mutlu, Saugata Ghose, Juan Gómez-Luna, and Rachata Ausavarungnirun, “En- abling Practical Processing in and Near Memory for Data-Intensive Computing, ” in DAC, 2019
2019
-
[17]
Pro- cessing Data Where It Makes Sense: Enabling In-Memory Computation,
Onur Mutlu, Saugata Ghose, Juan Gómez-Luna, and Rachata Ausavarungnirun, “Pro- cessing Data Where It Makes Sense: Enabling In-Memory Computation, ”MicPro, 2019
2019
-
[18]
Intelligent Architectures for Intelligent Machines,
Onur Mutlu, “Intelligent Architectures for Intelligent Machines, ” in VLSI-DAT, 2020
2020
-
[19]
Processing-in-Memory: A Workload-Driven Perspective,
Saugata Ghose, Amirali Boroumand, Jeremie S Kim, Juan Gómez-Luna, and Onur Mutlu, “Processing-in-Memory: A Workload-Driven Perspective, ”IBM JRD, 2019
2019
-
[20]
Accelerating Genome Analysis: A Primer on an Ongoing Journey,
Mohammed Alser, Zülal Bingöl, Damla Senol Cali, Jeremie Kim, Saugata Ghose, Can Alkan, and Onur Mutlu, “Accelerating Genome Analysis: A Primer on an Ongoing Journey, ”IEEE Micro, 2020
2020
-
[21]
GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis,
Damla Senol Cali, Gurpreet S Kalsi, Zülal Bingöl, Can Firtina, Lavanya Subrama- nian, Jeremie S Kim, Rachata Ausavarungnirun, Mohammed Alser, Juan Gomez-Luna, Amirali Boroumand et al., “GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework fo...
2020
-
[22]
EDEN: Enabling Energy-Efficient, High- Performance Deep Neural Network Inference Using Approximate DRAM,
Skanda Koppula, Lois Orosa, A Giray Yağlıkçı, Roknoddin Azizi, Taha Shahroodi, Kon- stantinos Kanellopoulos, and Onur Mutlu, “EDEN: Enabling Energy-Efficient, High- Performance Deep Neural Network Inference Using Approximate DRAM, ” inMICRO, 2019
2019
-
[23]
SMASH: Co-Designing Software Compression and Hardware- Accelerated Indexing for Efficient Sparse Matrix Operations,
Konstantinos Kanellopoulos, Nandita Vijaykumar, Christina Giannoula, Roknoddin Azizi, Skanda Koppula, Nika Mansouri Ghiasi, Taha Shahroodi, Juan Gomez-Luna, and Onur Mutlu, “SMASH: Co-Designing Software Compression and Hardware- Accelerated Indexing for Efficient Sparse Matrix...
2019
-
[24]
DAMOV: A New Methodology and Benchmark Suite for Evaluating Data Movement Bottlenecks,
Geraldo F. Oliveira, Juan Gómez-Luna, Lois Orosa, Saugata Ghose, Nandita Vijayku- mar, Ivan Fernandez, Mohammad Sadrosadati, and Onur Mutlu, “DAMOV: A New Methodology and Benchmark Suite for Evaluating Data Movement Bottlenecks, ”IEEE Access, 2021. BIBLIOGRAPHY 201
2021
-
[25]
DAMOV: A New Methodology and Benchmark Suite for Evaluating Data Movement Bottlenecks,
Geraldo F. Oliveira, Juan Gómez-Luna, Lois Orosa, Saugata Ghose, Nandita Vijayku- mar, Ivan Fernandez, Mohammad Sadrosadati, and Onur Mutlu, “DAMOV: A New Methodology and Benchmark Suite for Evaluating Data Movement Bottlenecks, ” arXiv:2105.03725 [cs.AR], 2021
2021 arXiv
-
[26]
DAMOV: A New Methodology and Benchmark Suite for Evalu- ating Data Movement Bottlenecks,
Geraldo F. Oliveira, “DAMOV: A New Methodology and Benchmark Suite for Evalu- ating Data Movement Bottlenecks, ” https://people.inf.ethz.ch/omutlu/pub/DAMO V-Bottleneck-Analysis-and-DataMovement-Benchmarks_arxiv21-talk.pptx, video available at https://youtu.be/GWideVyo0nM, 202...
2021
-
[27]
Processing Data Where It Makes Sense: Enabling In-Memory Computation,
O. Mutlu, “Processing Data Where It Makes Sense: Enabling In-Memory Computation, ” https://people.inf.ethz.ch/omutlu/pub/onur-MST-Keynote-EnablingInMemoryCom putation-October-27-2017-unrolled-FINAL.pptx, 2017, Keynote talk at MST
2017
-
[28]
Processing Data Where It Makes Sense in Modern Computing Systems: Enabling In-Memory Computation,
O. Mutlu, “Processing Data Where It Makes Sense in Modern Computing Systems: Enabling In-Memory Computation, ” https://people.inf.ethz.ch/omutlu/pub/onur-G WU-EnablingInMemoryComputation-February-15-2019-unrolled-FINAL.pptx, video available at https://www.youtube.com/watch?v=o...
2019
-
[29]
Processing Data Where It Makes Sense in Modern Computing Systems: Enabling In-Memory Computation,
O. Mutlu, “Processing Data Where It Makes Sense in Modern Computing Systems: Enabling In-Memory Computation, ” https://people.inf.ethz.ch/omutlu/pub/onur-ISS CC2019-talk.pptx, 2019, Invited Talk at ISSCC Special Forum on "Intelligence at the Edge: How Can We Make Machine Learn...
2019
-
[30]
Processing Data Where It Makes Sense in Modern Computing Systems: Enabling In-Memory Computation,
O. Mutlu, “Processing Data Where It Makes Sense in Modern Computing Systems: Enabling In-Memory Computation, ” https://people.inf.ethz.ch/omutlu/pub/onur-G LSVLSI-KeynoteTalk-EnablingInMemoryComputation-May-10-2019-unrolled. pptx, 2019, Keynote Talk at 29th ACM Great Lakes Sym...
2019
-
[31]
Processing Data Where It Makes Sense in Modern Computing Systems: Enabling In-Memory Computation,
O. Mutlu, “Processing Data Where It Makes Sense in Modern Computing Systems: Enabling In-Memory Computation, ” https://people.inf.ethz.ch/omutlu/pub/onur-A PPT-Keynote-EnablingInMemoryComputation-August-16-2019-unrolled.pptx, video available at https://www.youtube.com/watch?v=...
2019
-
[32]
Processing Data Where It Makes Sense in Modern Computing Systems: Enabling In-Memory Computation,
O. Mutlu, “Processing Data Where It Makes Sense in Modern Computing Systems: Enabling In-Memory Computation, ” https://people.inf.ethz.ch/omutlu/pub/onur-I CCD-Keynote-EnablingInMemoryComputation-November-19-2019-unrolled.pptx, video available at https://www.youtube.com/watch?...
2019
-
[33]
Evaluating the Memory System Behavior of Smartphone Workloads,
G Narancic, Patrick Judd, D Wu, Islam Atta, M Elnacouzi, Jason Zebchuk, Jorge Alberi- cio, N Enright Jerger, Andreas Moshovos, K Kutulakos et al., “Evaluating the Memory System Behavior of Smartphone Workloads, ” inSAMOS, 2014. BIBLIOGRAPHY 202
2014
-
[34]
Memory Hierarchy for Web Search,
Grant Ayers, Jung Ho Ahn, Christos Kozyrakis, and Parthasarathy Ranganathan, “Memory Hierarchy for Web Search, ” inHPCA, 2018
2018
-
[35]
Understanding Big Data Analytics Workloads on Modern Processors,
Zhen Jia, Jianfeng Zhan, Lei Wang, Chunjie Luo, Wanling Gao, Yi Jin, Rui Han, and Lixin Zhang, “Understanding Big Data Analytics Workloads on Modern Processors, ” TPDS, 2016
2016
-
[36]
Optimizing database architecture for the new bottleneck: Memory access,
S. Manegold, P. A. Boncz, and M. L. Kersten, “Optimizing database architecture for the new bottleneck: Memory access, ”The VLDB Journal, vol. 9, no. 3, December 2000
2000
-
[37]
AI and Memory Wall,
Amir Gholami, Zhewei Yao, Sehoon Kim, Coleman Hooper, Michael W Mahoney, and Kurt Keutzer, “AI and Memory Wall, ”IEEE Micro, 2024
2024
-
[38]
Quantifying Per- formance Bottlenecks of Stencil Computations Using the Execution-Cache-Memory Model,
Holger Stengel, Jan Treibig, Georg Hager, and Gerhard Wellein, “Quantifying Per- formance Bottlenecks of Stencil Computations Using the Execution-Cache-Memory Model, ” inISC, 2015
2015
-
[39]
Memory System Characterization of Deep Learning Workloads,
Zeshan Chishti and Berkin Akin, “Memory System Characterization of Deep Learning Workloads, ” inMEMSYS, 2019
2019
-
[40]
The Declining Effectiveness of Dynamic Caching for General-Purpose Microprocessors,
Douglas C Burger, James R Goodman, and Alain Kagi, “The Declining Effectiveness of Dynamic Caching for General-Purpose Microprocessors, ” University of Wisconsin- Madison, Tech. Rep. 1261, 1995
1995
-
[41]
The Ar- chitectural Implications of Facebook’s DNN-based Personalized Recommendation,
Udit Gupta, Carole-Jean Wu, Xiaodong Wang, Maxim Naumov, Brandon Reagen, David Brooks, Bradford Cottel, Kim Hazelwood, Mark Hempstead, Bill Jia et al. , “The Ar- chitectural Implications of Facebook’s DNN-based Personalized Recommendation, ” in HPCA, 2020
2020
-
[42]
SoftSKU: Optimizing Server Architectures for Microservice Diversity @Scale,
Akshitha Sriraman, Abhishek Dhanotia, and Thomas F Wenisch, “SoftSKU: Optimizing Server Architectures for Microservice Diversity @Scale, ” inISCA, 2019
2019
-
[43]
Understanding Data Storage and Ingestion for Large-Scale Deep Recommendation Model Training: Indus- trial Product,
Mark Zhao, Niket Agarwal, Aarti Basant, Buğra Gedik, Satadru Pan, Mustafa Ozdal, Rakesh Komuravelli, Jerry Pan, Tianshu Bao, Haowei Lu et al., “Understanding Data Storage and Ingestion for Large-Scale Deep Recommendation Model Training: Indus- trial Product, ” inISCA, 2022
2022
-
[44]
{INSIDER}: Designing{In-Storage} Com- puting System for Emerging{High-Performance} Drive,
Zhenyuan Ruan, Tong He, and Jason Cong, “{INSIDER}: Designing{In-Storage} Com- puting System for Emerging{High-Performance} Drive, ” inATC, 2019
2019
-
[45]
The Architectural Implications of Cloud Microser- vices,
Yu Gan and Christina Delimitrou, “The Architectural Implications of Cloud Microser- vices, ”CAL, 2018
2018
-
[46]
Accelerometer: Understanding Accelera- tion Opportunities for Data Center Overheads at Hyperscale,
Akshitha Sriraman and Abhishek Dhanotia, “Accelerometer: Understanding Accelera- tion Opportunities for Data Center Overheads at Hyperscale, ” inASPLOS, 2020
2020
-
[47]
RAMBDA: RDMA-Driven Acceleration Framework for Memory-Intensive 𝜇s-Scale Datacenter Applications,
Yifan Yuan, Jinghan Huang, Yan Sun, Tianchen Wang, Jacob Nelson, Dan RK Ports, Yipeng Wang, Ren Wang, Charlie Tai, and Nam Sung Kim, “RAMBDA: RDMA-Driven Acceleration Framework for Memory-Intensive 𝜇s-Scale Datacenter Applications, ” in HPCA, 2023. BIBLIOGRAPHY 203
2023
-
[48]
Amdahl’s Law for Tail Latency,
Christina Delimitrou and Christos Kozyrakis, “Amdahl’s Law for Tail Latency, ”CACM, 2018
2018
-
[49]
Cross-Stack Workload Characterization of Deep Recommendation Systems,
Samuel Hsia, Udit Gupta, Mark Wilkening, Carole-Jean Wu, Gu-Yeon Wei, and David Brooks, “Cross-Stack Workload Characterization of Deep Recommendation Systems, ” in IISWC, 2020
2020
-
[50]
vbench: Benchmarking Video Transcoding in the Cloud,
Andrea Lottarini, Alex Ramirez, Joel Coburn, Martha A Kim, Parthasarathy Ran- ganathan, Daniel Stodolsky, and Mark Wachsler, “vbench: Benchmarking Video Transcoding in the Cloud, ” inASPLOS, 2018
2018
-
[51]
Characterizing Job Microarchitectural Profiles at Scale: Dataset and Analysis,
Kangjin Wang, Ying Li, Cheng Wang, Tong Jia, Kingsum Chow, Yang Wen, Yaoyong Dou, Guoyao Xu, Chuanjia Hou, Jie Yao et al., “Characterizing Job Microarchitectural Profiles at Scale: Dataset and Analysis, ” inICPP, 2022
2022
-
[52]
Missing the Forest for the Trees: End-to-End AI Application Performance in Edge Data Centers,
Daniel Richins, Dharmisha Doshi, Matthew Blackmore, Aswathy Thulaseedharan Nair, Neha Pathapati, Ankit Patel, Brainard Daguman, Daniel Dobrijalowski, Ramesh Il- likkal, Kevin Long et al., “Missing the Forest for the Trees: End-to-End AI Application Performance in Edge Data Cen...
2020
-
[53]
Reflections on the Memory Wall,
Sally A McKee, “Reflections on the Memory Wall, ” inCF, 2004
2004
-
[54]
Accelerating Depen- dent Cache Misses with an Enhanced Memory Controller,
Milad Hashemi, Eiman Ebrahimi, Onur Mutlu, Yale N Patt et al., “Accelerating Depen- dent Cache Misses with an Enhanced Memory Controller, ” inISCA, 2016
2016
-
[55]
Continuous Runahead: Transparent Hardware Acceleration for Memory Intensive Workloads,
Milad Hashemi, Onur Mutlu, and Yale N Patt, “Continuous Runahead: Transparent Hardware Acceleration for Memory Intensive Workloads, ” inMICRO, 2016
2016
-
[56]
A Scal- able Processing-in-Memory Accelerator for Parallel Graph Processing,
Junwhan Ahn, Sungpack Hong, Sungjoo Yoo, Onur Mutlu, and Kiyoung Choi, “A Scal- able Processing-in-Memory Accelerator for Parallel Graph Processing, ” inISCA, 2015
2015
-
[57]
PIM-Enabled Instruc- tions: A Low-Overhead, Locality-Aware Processing-in-Memory Architecture,
Junwhan Ahn, Sungjoo Yoo, Onur Mutlu, and Kiyoung Choi, “PIM-Enabled Instruc- tions: A Low-Overhead, Locality-Aware Processing-in-Memory Architecture, ” inISCA, 2015
2015
-
[58]
GPUs and the Future of Parallel Computing,
Stephen W Keckler, William J Dally, Brucek Khailany, Michael Garland, and David Glasco, “GPUs and the Future of Parallel Computing, ”Micro, IEEE, 2011
2011
-
[59]
Slave Memories and Dynamic Storage Allocation,
M. V. Wilkes, “Slave Memories and Dynamic Storage Allocation, ”TEC, 1965
1965
-
[60]
Norman P Jouppi, 1990
1990
-
[61]
A Case for Direct-Mapped Caches,
Mark D Hill, “A Case for Direct-Mapped Caches, ” Computer, 1988
1988
-
[62]
High-Bandwidth Data Memory Systems for Su- perscalar Processors,
Gurindar S Sohi and Manoj Franklin, “High-Bandwidth Data Memory Systems for Su- perscalar Processors, ” inASPLOS, 1991
1991
-
[63]
The Microarchitecture of the Intel Pentium 4 Processor on 90nm Technology,
Darrell Boggs, Aravindh Baktha, Jason Hawkins, Deborah T Marr, J Alan Miller, Patrice Roussel, Ronak Singhal, Bret Toll, and KS Venkatraman, “The Microarchitecture of the Intel Pentium 4 Processor on 90nm Technology, ”Intel Technology Journal, 2004
2004
-
[64]
Tuning the Pentium Pro Microarchitecture,
David B Papworth, “Tuning the Pentium Pro Microarchitecture, ”IEEE Micro, 1996. BIBLIOGRAPHY 204
1996
-
[65]
POWER7: IBM’s Next-Generation Server Processor,
Ron Kalla, Balaram Sinharoy, William J Starke, and Michael Floyd, “POWER7: IBM’s Next-Generation Server Processor, ”IEEE Micro, 2010
2010
-
[66]
IBM POWER7 Multicore Server Processor,
Balaram Sinharoy, Ron Kalla, William J Starke, Hung Q Le, Robert Cargnoni, James A Van Norstrand, Bruce J Ronchetti, Jeffrey Stuecheli, Jens Leenstra, Guy L Guthrieet al., “IBM POWER7 Multicore Server Processor, ”IJRD, 2011
2011
-
[67]
IBM POWER8 Processor Core Microarchitecture,
Balaram Sinharoy, JA Van Norstrand, Richard J Eickemeyer, Hung Q Le, Jens Leen- stra, Dung Q Nguyen, B Konigsburg, K Ward, MD Brown, José E Moreira et al., “IBM POWER8 Processor Core Microarchitecture, ”IJRD, 2015
2015
-
[68]
Developing the AMD-K5 Architecture,
Dave Christie, “Developing the AMD-K5 Architecture, ”IEEE Micro, 1996
1996
-
[69]
Content-Sensitive Data Prefetching,
Robert Cooksey, “Content-Sensitive Data Prefetching, ” Ph.D. dissertation, 2002
2002
-
[70]
A Hardware-Based Cache Pollution Filtering Mechanism for Aggressive Prefetches,
Xiaotong Zhuang and Hsien-Hsin S. Lee, “A Hardware-Based Cache Pollution Filtering Mechanism for Aggressive Prefetches, ” inICPP-32, 2003
2003
-
[71]
Effective Stream-Based and Execution-Based Data Prefetching,
Sorin Iacobovici, Lawrence Spracklen, Sudarshan Kadambi, Yuan Chou, and Santosh G. Abraham, “Effective Stream-Based and Execution-Based Data Prefetching, ” in ICS, 2004
2004
-
[72]
AC/DC: An Adaptive Data Cache Prefetcher,
Kyle J.Nesbit, Ashutosh S. Dhodapkar, and James E.Smith, “AC/DC: An Adaptive Data Cache Prefetcher, ” inPACT, 2004
2004
-
[73]
Memory Prefetching Using Adaptive Stream Detection,
Ibrahim Hur and Calvin Lin, “Memory Prefetching Using Adaptive Stream Detection, ” in MICRO, 2006
2006
-
[74]
Pythia: A Customizable Hardware Prefetching Framework using Online Reinforcement Learning,
Rahul Bera, Konstantinos Kanellopoulos, Anant Nori, Taha Shahroodi, Sreenivas Sub- ramoney, and Onur Mutlu, “Pythia: A Customizable Hardware Prefetching Framework using Online Reinforcement Learning, ” inMICRO, 2021
2021
-
[75]
Access Map Pattern Matching for Data Cache Prefetch,
Yasuo Ishii, Mary Inaba, and Kei Hiraki, “Access Map Pattern Matching for Data Cache Prefetch, ” inISC, 2009
2009
-
[76]
Stride Directed Prefetching in Scalar Processors,
John WC Fu, Janak H Patel, and Bob L Janssens, “Stride Directed Prefetching in Scalar Processors, ” inMICRO, 1992
1992
-
[77]
Effective Hardware-Based Data Prefetching for High-Performance Processors,
Tien-Fu Chen and Jean-Loup Baer, “Effective Hardware-Based Data Prefetching for High-Performance Processors, ”TC, 1995
1995
-
[78]
Domino Temporal Data Prefetcher,
Mohammad Bakhshalipour, Pejman Lotfi-Kamran, and Hamid Sarbazi-Azad, “Domino Temporal Data Prefetcher, ” inHPCA, 2018
2018
-
[79]
DSPatch: Dual Spatial Pattern Prefetcher,
Rahul Bera, Anant V Nori, Onur Mutlu, and Sreenivas Subramoney, “DSPatch: Dual Spatial Pattern Prefetcher, ” inMICRO, 2019
2019
-
[80]
Speculative Execution via Address Prediction and Data Prefetching,
José González and Antonio González, “Speculative Execution via Address Prediction and Data Prefetching, ” inICS, 1997
1997
-
[81]
Coordinated Control of Multiple Prefetchers in Multi-Core Systems,
Eiman Ebrahimi, Onur Mutlu, Chang Joo Lee, and Yale N Patt, “Coordinated Control of Multiple Prefetchers in Multi-Core Systems, ” inMICRO, 2009. BIBLIOGRAPHY 205
2009
-
[82]
A Study of Integrated Prefetch- ing and Caching Strategies,
Pei Cao, Edward W Felten, Anna R Karlin, and Kai Li, “A Study of Integrated Prefetch- ing and Caching Strategies, ” inSIGMETRICS, 1995
1995
-
[83]
Generalized Correlation-Based Hardware Prefetching,
Mark J. Charney and Anthony P. Reeves, “Generalized Correlation-Based Hardware Prefetching, ” Cornell Univ., Tech. Rep., 1995
1995
-
[84]
Effective Jump-Pointer Prefetching for Linked Data Structures,
Amir Roth and Gurindar S Sohi, “Effective Jump-Pointer Prefetching for Linked Data Structures, ” inISCA, 1999
1999
-
[85]
Data Prefetching by Dependence Graph Precomputation,
Murali Annavaram, Jignesh M Patel, and Edward S Davidson, “Data Prefetching by Dependence Graph Precomputation, ” inISCA, 2001
2001
-
[86]
Graph Prefetching Using Data Structure Knowledge,
Sam Ainsworth and Timothy M Jones, “Graph Prefetching Using Data Structure Knowledge, ” inICS, 2016
2016
-
[87]
Gretch: A Hardware Prefetcher for Graph Analytics,
Anirudh Mohan Kaushik, Gennady Pekhimenko, and Hiren Patel, “Gretch: A Hardware Prefetcher for Graph Analytics, ”TACO, 2021
2021
-
[88]
PrefEdge: SSD Prefetcher for Large-Scale Graph Traversal,
Karthik Nilakant, Valentin Dalibard, Amitabha Roy, and Eiko Yoneki, “PrefEdge: SSD Prefetcher for Large-Scale Graph Traversal, ” inSYSTOR, 2014
2014
-
[89]
Correlation-Based Hardware Prefetching,
Mark Charney, “Correlation-Based Hardware Prefetching, ” Ph.D. dissertation, Cornell University, 1995
1995
-
[90]
A Stateless, Content-Directed Data Prefetching Mechanism,
Robert Cooksey, Stephan Jourdan, and Dirk Grunwald, “A Stateless, Content-Directed Data Prefetching Mechanism, ” inASPLOS, 2002
2002
-
[91]
Prefetching Using Markov Predictors,
Doug Joseph and Dirk Grunwald, “Prefetching Using Markov Predictors, ” inISCA, 1997
1997
-
[92]
Feedback Directed Prefetching: Improving the Performance and Bandwidth-Efficiency of Hardware Prefetchers,
Santhosh Srinath, Onur Mutlu, Hyesoon Kim, and Yale N Patt, “Feedback Directed Prefetching: Improving the Performance and Bandwidth-Efficiency of Hardware Prefetchers, ” inHPCA, 2007
2007
-
[93]
Techniques for Bandwidth-Efficient Prefetching of Linked Data Structures in Hybrid Prefetching Systems,
Eiman Ebrahimi, Onur Mutlu, and Yale N Patt, “Techniques for Bandwidth-Efficient Prefetching of Linked Data Structures in Hybrid Prefetching Systems, ” inHPCA, 2009
2009
-
[94]
Speculative Precomputation: Long-Range Prefetching of Delinquent Loads,
Jamison D Collins, Hong Wang, Dean M Tullsen, Christopher Hughes, Yong-Fong Lee, Dan Lavery, and John P Shen, “Speculative Precomputation: Long-Range Prefetching of Delinquent Loads, ” inISCA, 2001
2001
-
[95]
LTRF: Enabling High-Capacity Register Files for GPUs via Hardware/Software Cooperative Register Prefetching,
Mohammad Sadrosadati, Amirhossein Mirhosseini, Seyed Borna Ehsani, Hamid Sarbazi-Azad, Mario Drumond, Babak Falsafi, Rachata Ausavarungnirun, and Onur Mutlu, “LTRF: Enabling High-Capacity Register Files for GPUs via Hardware/Software Cooperative Register Prefetching, ” inASPLOS, 2018
2018
-
[96]
Dependence Based Prefetching for Linked Data Structures,
Amir Roth, Andreas Moshovos, and Gurindar S. Sohi, “Dependence Based Prefetching for Linked Data Structures, ” inASPLOS, 1998
1998
-
[97]
SPAID: Software Prefetching in Pointer- and Call-Intensive Environments,
Mikko H. Lipasti, William J. Schmidt, Steven R. Kunkel, and Robert R. Roediger, “SPAID: Software Prefetching in Pointer- and Call-Intensive Environments, ” inMICRO, 1995. BIBLIOGRAPHY 206
1995
-
[98]
A Prefetching Technique for Irregular Accesses to Linked Data Structures,
Magnus Karlsson, Fredrik Dahlgren, and Per Stenström, “A Prefetching Technique for Irregular Accesses to Linked Data Structures, ” inHPCA, 2000
2000
-
[99]
Pointer Cache Assisted Prefetching,
Jamison D. Collins, Suleyman Sair, Brad Calder, and Dean M. Tullsen, “Pointer Cache Assisted Prefetching, ” inMICRO, 2002
2002
-
[100]
TCP: Tag Correlating Prefetchers,
Zhigang Hu, Margaret Martonosi, and Stefanos Kaxiras, “TCP: Tag Correlating Prefetchers, ” inHPCA, 2003
2003
-
[101]
IMP: Indirect Memory Prefetcher,
Xiangyao Yu, Christopher J. Hughes, Nadathur Satish, and Srinivas Devadas, “IMP: Indirect Memory Prefetcher, ” inMICRO, 2015
2015
-
[102]
Sequential Program Prefetching in Memory Hierarchies,
Alan Jay Smith, “Sequential Program Prefetching in Memory Hierarchies, ” Computer, 1978
1978
-
[103]
Lockup-Free Instruction Fetch/Prefetch Cache Organization,
David Kroft, “Lockup-Free Instruction Fetch/Prefetch Cache Organization, ” in ISCA, 1981
1981
-
[104]
Data Prefetching in Shared Memory Multipro- cessors,
R. L. Lee, P.-C. Yew, and D. H. Lawrie, “Data Prefetching in Shared Memory Multipro- cessors, ” inICPP, 1987
1987
-
[105]
An Efficient Architecture for Loop Based Data Prefetching,
William Y. Chen, Roger A. Bringmann, Scott A. Mahlke, Richard E. Hank, and James E. Sicolo, “An Efficient Architecture for Loop Based Data Prefetching, ” inMICRO, 1992
1992
-
[106]
Prefetching in Supercomputer Instruction Caches,
James E. Smith and W.-C. Hsu, “Prefetching in Supercomputer Instruction Caches, ” in SC, 1992
1992
-
[107]
Sequential Hardware Prefetch- ing in Shared-Memory Multiprocessors,
Fredrik Dahlgren, Michel Dubois, and Per Stenström, “Sequential Hardware Prefetch- ing in Shared-Memory Multiprocessors, ”IPDS, 1995
1995
-
[108]
Wrong-Path Instruction Prefetching,
Jim Pierce and Trevor Mudge, “Wrong-Path Instruction Prefetching, ” inMICRO, 1996
1996
-
[109]
Instruction Prefetching of System Codes With Layout Optimized for Reduced Cache Misses,
Chun Xia and Josep Torrellas, “Instruction Prefetching of System Codes With Layout Optimized for Reduced Cache Misses, ” inISCA, 1996
1996
-
[110]
Adaptive Data Prefetching Using Cache Information,
Ando Ki and Alan E.Knowles, “Adaptive Data Prefetching Using Cache Information, ” in SC, 1997
1997
-
[111]
Cooperative Prefetching: Compiler and Hard- ware Support for Effective Instruction Prefetching in Modern Processors,
Chi-Keung Luk and Todd C. Mowry, “Cooperative Prefetching: Compiler and Hard- ware Support for Effective Instruction Prefetching in Modern Processors, ” in MICRO, 1998
1998
-
[112]
Hardware-Only Stream Prefetching and Dynamic Access Ordering,
C. Zhang and S. A. McKee, “Hardware-Only Stream Prefetching and Dynamic Access Ordering, ” inICS, 2000
2000
-
[113]
Runahead Execution: An Effective Alternative to Large Instruction Windows,
Onur Mutlu, Jared Stark, Chris Wilkerson, and Yale N Patt, “Runahead Execution: An Effective Alternative to Large Instruction Windows, ”IEEE Micro, 2003
2003
-
[114]
Efficient Runahead Execution: Power- Efficient Memory Latency Tolerance,
Onur Mutlu, Hyesoon Kim, and Yale N Patt, “Efficient Runahead Execution: Power- Efficient Memory Latency Tolerance, ”IEEE Micro, 2006. BIBLIOGRAPHY 207
2006
-
[115]
Techniques for Efficient Processing in Runahead Execution Engines,
Onur Mutlu, Hyesoon Kim, and Yale N Patt, “Techniques for Efficient Processing in Runahead Execution Engines, ” inISCA, 2005
2005
-
[116]
Runahead Execution: An Alternative to Very Large Instruction Windows for Out-of-Order Processors,
Onur Mutlu, Jared Stark, Chris Wilkerson, and Yale N Patt, “Runahead Execution: An Alternative to Very Large Instruction Windows for Out-of-Order Processors, ” inHPCA, 2003
2003
-
[117]
Critical Issues Regarding HPS, a High Performance Microarchitecture,
Y. N. Patt, S. W. Melvin, W.-M. Hwu, and M. C. Shebanow, “Critical Issues Regarding HPS, a High Performance Microarchitecture, ” inMICRO, 1985
1985
-
[118]
HPS, a New Microarchitecture: Rationale and Introduction,
Y. N. Patt, W.-M. Hwu, and M. Shebanow, “HPS, a New Microarchitecture: Rationale and Introduction, ” inMICRO, 1985
1985
-
[119]
An Efficient Algorithm for Exploiting Multiple Arithmetic Units,
R. M. Tomasulo, “An Efficient Algorithm for Exploiting Multiple Arithmetic Units, ”IBM JRD, 1967
1967
-
[120]
IBM Power5 Chip: A Dual-Core Multithreaded Processor,
Ron Kalla, Balaram Sinharoy, and Joel M Tendler, “IBM Power5 Chip: A Dual-Core Multithreaded Processor, ”IEEE Micro, 2004
2004
-
[121]
POWER4 System Microarchitecture,
Joel Tendler, Steve Dodson, Steve Fields, Hung Le, and Balaram Sinharoy, “POWER4 System Microarchitecture, ”IBM JRD, 2001
2001
-
[122]
The MIPS R10000 Superscalar Microprocessor,
Kenneth C. Yeager, “The MIPS R10000 Superscalar Microprocessor, ”IEEE Micro, 1996
1996
-
[123]
The Microarchitecture of the Pentium 4 Processor,
Glenn Hinton, Dave Sager, Mike Upton, Darrell Boggs, Doug Carmean, Alan Kyker, and Patrice Roussel, “The Microarchitecture of the Pentium 4 Processor, ” Intel Technology Journal, 2001
2001
-
[124]
The Alpha 21264 Microprocessor,
Richard E. Kessler, “The Alpha 21264 Microprocessor, ”IEEE Micro, 1999
1999
-
[125]
Parallel Operation in the Control Data 6600,
James E Thornton, “Parallel Operation in the Control Data 6600, ” in AFIPS, 1964
1964
-
[126]
A Pipelined, Shared Resource MIMD Computer,
Burton J Smith, “A Pipelined, Shared Resource MIMD Computer, ” in ICPP, 1978
1978
-
[127]
Niagara: A 32-Way Multithreaded SPARC Processor,
Poonacha Kongetira, Kathirgamar Aingaran, and Kunle Olukotun, “Niagara: A 32-Way Multithreaded SPARC Processor, ”IEEE Micro, 2005
2005
-
[128]
Architecture and Applications of the HEP Multiprocessor Computer System,
Burton J Smith, “Architecture and Applications of the HEP Multiprocessor Computer System, ” inReal-Time Signal Processing IV, 1982
1982
-
[129]
Multithreading: A Revisionist View of Dataflow Architectures,
Gregory M. Papadopoulos and Kenneth R. Traub, “Multithreading: A Revisionist View of Dataflow Architectures, ” inISCA, 1991
1991
-
[130]
Simultaneous Multithreading: Maximizing On-Chip Parallelism,
Dean M. Tullsen, Susan J. Eggers, and Henry M. Levy, “Simultaneous Multithreading: Maximizing On-Chip Parallelism, ” inISCA, 1995
1995
-
[131]
A Dynamic Multithreading Processor,
Haitham Akkary and Michael A. Driscoll, “A Dynamic Multithreading Processor, ” in MICRO, 1998
1998
-
[132]
Speculative Data-Driven Multithreading,
Amir Roth and Gurindar S. Sohi, “Speculative Data-Driven Multithreading, ” in HPCA, 2001. BIBLIOGRAPHY 208
2001
-
[133]
Chip Multithreading: Opportunities and Challenges,
Lawrence Spracklen and Santosh G. Abraham, “Chip Multithreading: Opportunities and Challenges, ” inHPCA, 2005
2005
-
[134]
Efficient performance evaluation of memory hierarchy for highly multithreaded graphics pro- cessors,
Sara S. Baghsorkhi, Isaac Gelado, Matthieu Delahaye, and Wen-mei W. Hwu, “Efficient performance evaluation of memory hierarchy for highly multithreaded graphics pro- cessors, ” inPPoPP, 2012
2012
-
[135]
MLP-Aware Runahead Threads in a Simultaneous Multithreading Processor,
Kenzo Van Craeynest, Stijn Eyerman, and Lieven Eeckhout, “MLP-Aware Runahead Threads in a Simultaneous Multithreading Processor, ” inHiPEAC, 2009
2009
-
[136]
Bottleneck Identification and Scheduling in Multithreaded Applications,
José A Joao, M Aater Suleman, Onur Mutlu, and Yale N Patt, “Bottleneck Identification and Scheduling in Multithreaded Applications, ” inASPLOS, 2012
2012
-
[137]
Single-ISA Heterogeneous Multi-Core Architectures for Multithreaded Workload Performance,
Rakesh Kumar, Dean M Tullsen, Parthasarathy Ranganathan, Norman P Jouppi, and Keith I Farkas, “Single-ISA Heterogeneous Multi-Core Architectures for Multithreaded Workload Performance, ” inISCA, 2004
2004
-
[138]
Utility-Based Acceleration of Multithreaded Applications on Asymmetric CMPs,
José A Joao, M Aater Suleman, Onur Mutlu, and Yale N Patt, “Utility-Based Acceleration of Multithreaded Applications on Asymmetric CMPs, ” inISCA, 2013
2013
-
[139]
Tolerating Memory Latency Through Software-Controlled Pre- Execution in Simultaneous Multithreading Processors,
Chi-Keung Luk, “Tolerating Memory Latency Through Software-Controlled Pre- Execution in Simultaneous Multithreading Processors, ” inISCA, 2001
2001
-
[140]
A Family of 45nm IA Processors,
R. Kumar and G. Hinton, “A Family of 45nm IA Processors, ” in ISSCC, 2009
2009
-
[141]
A 48- Core IA-32 Message-Passing Processor with DVFS in 45nm CMOS,
J. Howard, S. Dighe, Y. Hoskote, S. Vangal, D. Finan, G. Ruhl, D. Jenkins, H. Wilson, N. Borkar, G. Schrom, F. Pailet, S. Jain, T. Jacob, S. Yada, S. Marella, P. Salihundam, V. Er- raguntla, M. Konow, M. Riepen, G. Droege, J. Lindemann, M. Gries, T. Apel, K. Henriss, T. Lund-L...
2010
-
[142]
An x86-64 Core Implemented in 32nm SOI CMOS,
R. Jotwani, S. Sundaram, S. Kosonocky, A. Schaefer, V. Andrade, G. Constant, A. Novak, and S. Naffziger, “An x86-64 Core Implemented in 32nm SOI CMOS, ” inISSCC, 2010
2010
-
[143]
5.5 Steamroller: An x86-64 Core Implemented in 28nm Bulk CMOS,
K. Gillespie, H. R. Fair, C. Henrion, R. Jotwani, S. Kosonocky, R. S. Orefice, D. A. Priore, J. White, and K. Wilcox, “5.5 Steamroller: An x86-64 Core Implemented in 28nm Bulk CMOS, ” inISSCC, 2014
2014
-
[144]
3.2 Zen: A Next- Generation High-Performance x86 Core,
T. Singh, S. Rangarajan, D. John, C. Henrion, S. Southard, H. McIntyre, A. Novak, S. Kosonocky, R. Jotwani, A. Schaefer, E. Chang, J. Bell, and M. Co, “3.2 Zen: A Next- Generation High-Performance x86 Core, ” inISSCC, 2017
2017
-
[145]
Memory-Centric Computing,
Onur Mutlu, “Memory-Centric Computing, ” arXiv:2305.20000 [cs.AR], 2023
2023 arXiv
-
[146]
Field-Effect Transistor Memory,
Robert H. Dennard, “Field-Effect Transistor Memory, ” U.S. Patent 3,387,286, 1968
1968
-
[147]
Design of Ion-Implanted MOSFET’s with Very Small Physical Dimensions,
Robert H Dennard, Fritz H Gaensslen, Hwa-Nien Yu, V Leo Rideout, Ernest Bassous, and Andre R LeBlanc, “Design of Ion-Implanted MOSFET’s with Very Small Physical Dimensions, ”JSSC, 1974. BIBLIOGRAPHY 209
1974
-
[148]
IBM’s Robert H. Dennard and the Chip That Changed the World,
John Markoff, “IBM’s Robert H. Dennard and the Chip That Changed the World, ” https: //www.ibm.com/blogs/think/2019/11/ibms-robert-h-dennard-and-the-chip-that-chan ged-the-world/, 2019
2019
-
[149]
Memory Lane,
Nature Electronics, “Memory Lane, ” 2018
2018
-
[150]
Adaptive Scheduling for Systems with Asymmetric Memory Hierarchies,
Po-An Tsai, Changping Chen, and Daniel Sanchez, “Adaptive Scheduling for Systems with Asymmetric Memory Hierarchies, ” inMICRO, 2018
2018
-
[151]
It’s the Memory, Stupid!
Richard Sites, “It’s the Memory, Stupid!” MPR, 1996
1996
-
[152]
Demystifying Complex Workload–DRAM Interactions: An Experimental Study,
S. Ghose, T. Li, N. Hajinazar, D. Senol Cali, and O. Mutlu, “Demystifying Complex Workload–DRAM Interactions: An Experimental Study, ” inSIGMETRICS, 2020
2020
-
[153]
Gather-Scatter DRAM: In-DRAM Address Translation to Improve the Spatial Locality of Non-Unit Strided Accesses,
Vivek Seshadri, Thomas Mullins, Amirali Boroumand, Onur Mutlu, Phillip B Gibbons, Michael A Kozuch, and Todd C Mowry, “Gather-Scatter DRAM: In-DRAM Address Translation to Improve the Spatial Locality of Non-Unit Strided Accesses, ” inMICRO, 2015
2015
-
[154]
GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Comput- ing Frameworks,
Lifeng Nai, Ramyad Hadidi, Jaewoong Sim, Hyojong Kim, Pranith Kumar, and Hye- soon Kim, “GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Comput- ing Frameworks, ” inHPCA, 2017
2017
-
[155]
Accelerating Pointer Chasing in 3D-Stacked Mem- ory: Challenges, Mechanisms, Evaluation,
Kevin Hsieh, Samira Khan, Nandita Vijaykumar, Kevin K Chang, Amirali Boroumand, Saugata Ghose, and Onur Mutlu, “Accelerating Pointer Chasing in 3D-Stacked Mem- ory: Challenges, Mechanisms, Evaluation, ” inICCD, 2016
2016
-
[156]
GRIM-Filter: Fast Seed Location Filtering in DNA Read Mapping Using Processing-in-Memory Tech- nologies,
Jeremie S Kim, Damla Senol Cali, Hongyi Xin, Donghyuk Lee, Saugata Ghose, Mo- hammed Alser, Hasan Hassan, Oguz Ergin, Can Alkan, and Onur Mutlu, “GRIM-Filter: Fast Seed Location Filtering in DNA Read Mapping Using Processing-in-Memory Tech- nologies, ”BMC Genomics, 2018
2018
-
[157]
Practical Mechanisms for Reducing Processor-Memory Data Movement in Modern Workloads,
Amirali Boroumand, “Practical Mechanisms for Reducing Processor-Memory Data Movement in Modern Workloads, ” Ph.D. dissertation, Carnegie Mellon University, 2020
2020
-
[158]
Co-Architecting Controllers and DRAM to Enhance DRAM Process Scaling,
Uksong Kang, Hak-Soo Yu, Churoo Park, Hongzhong Zheng, John Halbert, Kuljit Bains, S Jang, and Joo Sun Choi, “Co-Architecting Controllers and DRAM to Enhance DRAM Process Scaling, ” inThe Memory Forum, 2014
2014
-
[159]
The Memory Gap and the Future of High Performance Memories,
Maurice V Wilkes, “The Memory Gap and the Future of High Performance Memories, ” CAN, 2001
2001
-
[160]
Flipping Bits in Memory Without Accessing Them: An Experimental Study of DRAM Disturbance Errors,
Yoongu Kim, Ross Daly, Jeremie Kim, Chris Fallin, Ji Hye Lee, Donghyuk Lee, Chris Wilkerson, Konrad Lai, and Onur Mutlu, “Flipping Bits in Memory Without Accessing Them: An Experimental Study of DRAM Disturbance Errors, ” inISCA, 2014
2014
-
[161]
A Case for Exploiting Subarray-Level Parallelism (SALP) in DRAM,
Yoongu Kim, Vivek Seshadri, Donghyuk Lee, Jamie Liu, and Onur Mutlu, “A Case for Exploiting Subarray-Level Parallelism (SALP) in DRAM, ” inISCA, 2012. BIBLIOGRAPHY 210
2012
-
[162]
Architectural Techniques to Enhance DRAM Scaling,
Yoongu Kim, “Architectural Techniques to Enhance DRAM Scaling, ” Ph.D. dissertation, Carnegie Mellon University, 2015
2015
-
[163]
RAIDR: Retention-Aware In- telligent DRAM Refresh,
Jamie Liu, Ben Jaiyen, Richard Veras, and Onur Mutlu, “RAIDR: Retention-Aware In- telligent DRAM Refresh, ” inISCA, 2012
2012
-
[164]
The RowHammer Problem and Other Issues We May Face as Memory Becomes Denser,
Onur Mutlu, “The RowHammer Problem and Other Issues We May Face as Memory Becomes Denser, ” inDATE, 2017
2017
-
[165]
Decoupled Direct Memory Access: Isolating CPU and IO Traffic by Lever- aging a Dual-Data-Port DRAM,
Donghyuk Lee, Lavanya Subramanian, Rachata Ausavarungnirun, Jongmoo Choi, and Onur Mutlu, “Decoupled Direct Memory Access: Isolating CPU and IO Traffic by Lever- aging a Dual-Data-Port DRAM, ” inPACT, 2015
2015
-
[166]
Architecting Phase Change Memory as a Scalable DRAM Alternative,
Benjamin C. Lee, Engin Ipek, Onur Mutlu, and Doug Burger, “Architecting Phase Change Memory as a Scalable DRAM Alternative, ” inISCA, 2009
2009
-
[167]
Row Buffer Locality Aware Caching Policies for Hybrid Memories,
HanBin Yoon, Justin Meza, Rachata Ausavarungnirun, Rachael A Harding, and Onur Mutlu, “Row Buffer Locality Aware Caching Policies for Hybrid Memories, ” in ICCD, 2012
2012
-
[168]
Efficient Data Mapping and Buffering Techniques for Multilevel Cell Phase- Change Memories,
Hanbin Yoon, Justin Meza, Naveen Muralimanohar, Norman P. Jouppi, and Onur Mutlu, “Efficient Data Mapping and Buffering Techniques for Multilevel Cell Phase- Change Memories, ”ACM TACO, 2014
2014
-
[169]
Disaggregated Memory for Expansion and Sharing in Blade Servers,
Kevin Lim, Jichuan Chang, Trevor Mudge, Parthasarathy Ranganathan, Steven K. Rein- hardt, and Thomas F. Wenisch, “Disaggregated Memory for Expansion and Sharing in Blade Servers, ” inISCA, 2009
2009
-
[170]
Hitting the Memory Wall: Implications of the Obvi- ous,
Wm A Wulf and Sally A McKee, “Hitting the Memory Wall: Implications of the Obvi- ous, ”CAN, 1995
1995
-
[171]
Un- derstanding Latency Variation in Modern DRAM Chips: Experimental Characteriza- tion, Analysis, and Optimization,
Kevin K. Chang, Abhijith Kashyap, Hasan Hassan, Saugata Ghose, Kevin Hsieh, Donghyuk Lee, Tianshi Li, Gennady Pekhimenko, Samira Khan, and Onur Mutlu, “Un- derstanding Latency Variation in Modern DRAM Chips: Experimental Characteriza- tion, Analysis, and Optimization, ” inSIGM...
2016
-
[172]
Tiered-Latency DRAM: A Low Latency and Low Cost DRAM Architec- ture,
Donghyuk Lee, Yoongu Kim, Vivek Seshadri, Jamie Liu, Lavanya Subramanian, and Onur Mutlu, “Tiered-Latency DRAM: A Low Latency and Low Cost DRAM Architec- ture, ” inHPCA, 2013
2013
-
[173]
Adaptive-Latency DRAM: Optimizing DRAM Timing for the Common-Case,
Donghyuk Lee, Yoongu Kim, Gennady Pekhimenko, Samira Khan, Vivek Seshadri, Kevin Chang, and Onur Mutlu, “Adaptive-Latency DRAM: Optimizing DRAM Timing for the Common-Case, ” inHPCA, 2015
2015
-
[174]
Understanding Reduced-Voltage Operation in Modern DRAM Devices: Experimental Characterization, Analysis, and Mechanisms,
Kevin K Chang, A Giray Yağlıkçı, Saugata Ghose, Aditya Agrawal, Niladrish Chatterjee, Abhijith Kashyap, Donghyuk Lee, Mike O’Connor, Hasan Hassan, and Onur Mutlu, “Understanding Reduced-Voltage Operation in Modern DRAM Devices: Experimental Characterization, Analysis, and Mech...
2017
-
[175]
Design- Induced Latency Variation in Modern DRAM Chips: Characterization, Analysis, and Latency Reduction Mechanisms,
Donghyuk Lee, Samira Khan, Lavanya Subramanian, Saugata Ghose, Rachata Ausavarungnirun, Gennady Pekhimenko, Vivek Seshadri, and Onur Mutlu, “Design- Induced Latency Variation in Modern DRAM Chips: Characterization, Analysis, and Latency Reduction Mechanisms, ” inSIGMETRICS, 2017
2017
-
[176]
Character- izing Application Memory Error Vulnerability to Optimize Datacenter Cost via Heterogeneous-Reliability Memory,
Yixin Luo, Sriram Govindan, Bikash Sharma, Mark Santaniello, Justin Meza, Aman Kansal, Jie Liu, Badriddine Khessib, Kushagra Vaid, and Onur Mutlu, “Character- izing Application Memory Error Vulnerability to Optimize Datacenter Cost via Heterogeneous-Reliability Memory, ” inDSN, 2014
2014
-
[177]
Using ECC DRAM to Adaptively Increase Mem- ory Capacity,
Yixin Luo, Saugata Ghose, Tianshi Li, Sriram Govindan, Bikash Sharma, Bryan Kelly, Amirali Boroumand, and Onur Mutlu, “Using ECC DRAM to Adaptively Increase Mem- ory Capacity, ” arXiv:1706.08870 [cs:AR], 2017
2017 arXiv
-
[178]
SoftMC: A Flexible and Practical Open-Source Infrastructure for Enabling Experimental DRAM Studies,
Hasan Hassan, Nandita Vijaykumar, Samira Khan, Saugata Ghose, Kevin Chang, Gen- nady Pekhimenko, Donghyuk Lee, Oguz Ergin, and Onur Mutlu, “SoftMC: A Flexible and Practical Open-Source Infrastructure for Enabling Experimental DRAM Studies, ” in HPCA, 2017
2017
-
[179]
ChargeCache: Reducing DRAM Latency by Exploit- ing Row Access Locality,
Hasan Hassan, Gennady Pekhimenko, Nandita Vijaykumar, Vivek Seshadri, Donghyuk Lee, Oguz Ergin, and Onur Mutlu, “ChargeCache: Reducing DRAM Latency by Exploit- ing Row Access Locality, ” inHPCA, 2016
2016
-
[180]
The Reach Profiler (REAPER): Enabling the Mitigation of DRAM Retention Failures via Profiling at Aggressive Conditions,
Minesh Patel, Jeremie S Kim, and Onur Mutlu, “The Reach Profiler (REAPER): Enabling the Mitigation of DRAM Retention Failures via Profiling at Aggressive Conditions, ” in ISCA, 2017
2017
-
[181]
CROW: A Low-Cost Substrate for Improving DRAM Performance, Energy Efficiency, and Reliability,
Hasan Hassan, Minesh Patel, Jeremie S Kim, A Giray Yaglikci, Nandita Vijaykumar, Nika Mansouri Ghiasi, Saugata Ghose, and Onur Mutlu, “CROW: A Low-Cost Substrate for Improving DRAM Performance, Energy Efficiency, and Reliability, ” inISCA, 2019
2019
-
[182]
What Your DRAM Power Models Are Not Telling You: Lessons from a Detailed Experimental Study,
Saugata Ghose, Abdullah Giray Yaglikçi, Raghav Gupta, Donghyuk Lee, Kais Kudrolli, William X Liu, Hasan Hassan, Kevin K Chang, Niladrish Chatterjee, Aditya Agrawal et al., “What Your DRAM Power Models Are Not Telling You: Lessons from a Detailed Experimental Study, ” inSIGMETR...
2018
-
[183]
Solar-DRAM: Reducing DRAM Access Latency by Exploiting the Variation in Local Bitlines,
Jeremie Kim, Minesh Patel, Hasan Hassan, and Onur Mutlu, “Solar-DRAM: Reducing DRAM Access Latency by Exploiting the Variation in Local Bitlines, ” inICCD, 2018
2018
-
[184]
Revisiting RowHammer: An Experimental Analysis of Mod- ern DRAM Devices and Mitigation Techniques,
Jeremie S Kim, Minesh Patel, A Giray Yağlıkçı, Hasan Hassan, Roknoddin Azizi, Lois Orosa, and Onur Mutlu, “Revisiting RowHammer: An Experimental Analysis of Mod- ern DRAM Devices and Mitigation Techniques, ” inISCA, 2020
2020
-
[185]
RowHammer: A Retrospective,
Onur Mutlu and Jeremie S Kim, “RowHammer: A Retrospective, ” TCAD, 2019
2019
-
[186]
Reducing DRAM Latency via Charge-Level-Aware Look-Ahead Partial Restoration,
Yaohua Wang, Arash Tavakkol, Lois Orosa, Saugata Ghose, Nika Mansouri Ghiasi, Mi- nesh Patel, Jeremie S Kim, Hasan Hassan, Mohammad Sadrosadati, and Onur Mutlu, “Reducing DRAM Latency via Charge-Level-Aware Look-Ahead Partial Restoration, ” in MICRO, 2018. BIBLIOGRAPHY 212
2018
-
[187]
Recent Advances in DRAM and Flash Memory Architectures,
Onur Mutlu, Saugata Ghose, and Rachata Ausavarungnirun, “Recent Advances in DRAM and Flash Memory Architectures, ” Invited Journal Issue IPSI Transactions on Internet Research, 2018
2018
-
[188]
Main Memory Scaling: Challenges and Solution Directions,
Onur Mutlu, “Main Memory Scaling: Challenges and Solution Directions, ” inMore Than Moore Technologies for Next Generation Computer Design . Springer-Verlag, 2015
2015
-
[189]
Memory Technology Trend and Future Challenges,
Sungjoo Hong, “Memory Technology Trend and Future Challenges, ” inIEDM, 2010
2010
-
[190]
RowPress: Amplifying Read Disturbance in Modern DRAM Chips,
Haocong Luo, Ataberk Olgun, Abdullah Giray Yağlıkcı, Yahya Can Tuğrul, Steve Rhyner, Meryem Banu Cavlak, Joël Lindegger, Mohammad Sadrosadati, and Onur Mutlu, “RowPress: Amplifying Read Disturbance in Modern DRAM Chips, ” in ISCA, 2023
2023
-
[191]
Spatial Variation-Aware Read Dis- turbance Defenses: Experimental Analysis of Real DRAM Chips and Implications on Future Solutions,
A. Giray Yağlıkcı, Yahya Can Tuğrul, Geraldo F De Oliviera, Ismail Emir Yüksel, Ataberk Olgun, Haocong Luo, and Onur Mutlu, “Spatial Variation-Aware Read Dis- turbance Defenses: Experimental Analysis of Real DRAM Chips and Implications on Future Solutions, ” inHPCA, 2024
2024
-
[192]
ABACuS: All- Bank Activation Counters for Scalable and Low Overhead RowHammer Mitigation,
Ataberk Olgun, Yahya Can Tugrul, F. Nisa Bostancı, Ismail Emir Yüksel, Haocong Luo, Steve Rhyner, A. Giray Yaglıkçı, Geraldo F. Oliveira, and Onur Mutlu, “ABACuS: All- Bank Activation Counters for Scalable and Low Overhead RowHammer Mitigation, ” in USENIX Security, 2024
2024
-
[193]
CoMeT: Count-Min-Sketch-Based Row Tracking to Mitigate RowHammer at Low Cost,
F. Nisa Bostancı, Ismail Emir Yüksel, Ataberk Olgun, Konstantinos Kanellopoulos, Yahya Can Tugrul, A. Giray Yaglıkçı, Mohammad Sadrosadati, and Onur Mutlu, “CoMeT: Count-Min-Sketch-Based Row Tracking to Mitigate RowHammer at Low Cost, ” inHPCA, 2024
2024
-
[194]
HiRA: Hidden Row Activation for Reducing Re- fresh Latency of Off-the-Shelf DRAM Chips,
A Giray Yağlikçi, Ataberk Olgun, Minesh Patel, Haocong Luo, Hasan Hassan, Lois Orosa, Oğuz Ergin, and Onur Mutlu, “HiRA: Hidden Row Activation for Reducing Re- fresh Latency of Off-the-Shelf DRAM Chips, ” inMICRO, 2022
2022
-
[195]
Enabling Efficient and Scalable DRAM Read Disturbance Mitigation via New Experimental Insights into Modern DRAM Chips,
Abdullah Giray Yağlıkçı, “Enabling Efficient and Scalable DRAM Read Disturbance Mitigation via New Experimental Insights into Modern DRAM Chips, ” Ph.D. disser- tation, ETH Zürich, 2024
2024
-
[196]
Un- derstanding and Mitigating Side and Covert Channel Vulnerabilities Introduced by RowHammer Defenses,
F Bostancı, Oğuzhan Canpolat, Ataberk Olgun, İsmail Emir Yüksel, Konstantinos Kanellopoulos, Mohammad Sadrosadati, A Giray Yağlıkçı, and Onur Mutlu, “Un- derstanding and Mitigating Side and Covert Channel Vulnerabilities Introduced by RowHammer Defenses, ” arXiv:2503.17891 [cs...
2025
-
[197]
Revisiting DRAM Read Disturbance: Identifying Inconsistencies Between Experimen- tal Characterization and Device-Level Studies,
Haocong Luo, İsmail Emir Yüksel, Ataberk Olgun, A Giray Yağlıkçı, and Onur Mutlu, “Revisiting DRAM Read Disturbance: Identifying Inconsistencies Between Experimen- tal Characterization and Device-Level Studies, ” inVTS, 2025
2025
-
[198]
Chronus: Un- derstanding and Securing the Cutting-Edge Industry Solutions to DRAM Read Distur- bance,
Oğuzhan Canpolat, A. Giray Yağlıkçı, Geraldo F. Oliveira, Ataberk Olgun, F. Nisa Bostanci, Ismail E. Yüksel, Haocong Luo, Oğuz Ergin, and Onur Mutlu, “Chronus: Un- derstanding and Securing the Cutting-Edge Industry Solutions to DRAM Read Distur- bance, ” inHPCA, 2025. BIBLIOGRAPHY 213
2025
-
[199]
Under- standing RowHammer Under Reduced Refresh Latency: Experimental Analysis of Real DRAM Chips and Implications on Future Solutions,
Yahya Can Tuğrul, A Giray Yağlıkçı, İsmail Emir Yüksel, Ataberk Olgun, Oğuzhan Can- polat, Nisa Bostancı, Mohammad Sadrosadati, Oğuz Ergin, and Onur Mutlu, “Under- standing RowHammer Under Reduced Refresh Latency: Experimental Analysis of Real DRAM Chips and Implications on Fu...
2025
-
[200]
Flikker: Saving DRAM Refresh-Power through Critical Data Partitioning,
Song Liu, Karthik Pattabiraman, Thomas Moscibroda, and Benjamin G. Zorn, “Flikker: Saving DRAM Refresh-Power through Critical Data Partitioning, ” inASPLOS, 2011
2011
-
[201]
SE- CRET: Selective Error Correction for Refresh Energy Reduction in DRAMs,
Chung-Hsiang Lin, De-Yu Shen, Yi-Jung Chen, Chia-Lin Yang, and Michael Wang, “SE- CRET: Selective Error Correction for Refresh Energy Reduction in DRAMs, ” in ICCD, 2012
2012
-
[202]
Refresh Now and Then,
Seungjae Baek, Sangyeun Cho, and Rami Melhem, “Refresh Now and Then, ” IEEE TC, 2013
2013
-
[203]
Coordinated Refresh: Energy Efficient Techniques for DRAM Refresh Scheduling,
Ishwar Bhati, Zeshan Chishti, and Bruce Jacob, “Coordinated Refresh: Energy Efficient Techniques for DRAM Refresh Scheduling, ” inISPLED, 2013
2013
-
[204]
A Case for Refresh Paus- ing in DRAM Memory Systems,
Prashant Nair, Chia-Chen Chou, and Moinuddin K Qureshi, “A Case for Refresh Paus- ing in DRAM Memory Systems, ” inHPCA, 2013
2013
-
[205]
Refreshing Thoughts on DRAM: Power Saving vs. Data Integrity,
Amir Rahmati, Matthew Hicks, Daniel Holcomb, and Kevin Fu, “Refreshing Thoughts on DRAM: Power Saving vs. Data Integrity, ” inW ACAS, 2014
2014
-
[206]
Scalable and Energy Efficient DRAM Refresh Techniques,
Ishwar Singh Bhati, “Scalable and Energy Efficient DRAM Refresh Techniques, ” Ph.D. dissertation, 2014
2014
-
[207]
DTail: A Flexible Approach to DRAM Refresh Management,
Zehan Cui, Sally A McKee, Zhongbin Zha, Yungang Bao, and Mingyu Chen, “DTail: A Flexible Approach to DRAM Refresh Management, ” inSC, 2014
2014
-
[208]
Data-Aware DRAM Refresh to Squeeze the Margin of Retention Time in Hybrid Memory Cube,
Yinhe Han, Ying Wang, Huawei Li, and Xiaowei Li, “Data-Aware DRAM Refresh to Squeeze the Margin of Retention Time in Hybrid Memory Cube, ” inICCAD, 2014
2014
-
[209]
Optimized Active and Power-Down Mode Refresh Control in 3D-DRAMs,
Matthias Jung, Christian Weis, Norbert Wehn, Mohammadsadegh Sadri, and Luca Benini, “Optimized Active and Power-Down Mode Refresh Control in 3D-DRAMs, ” in VLSI-SoC, 2014
2014
-
[210]
Refresh Pausing in DRAM Memory Systems,
Prashant J Nair, Chia-Chen Chou, and Moinuddin K Qureshi, “Refresh Pausing in DRAM Memory Systems, ”TACO, 2014
2014
-
[211]
CREAM: A Concurrent-Refresh-Aware DRAM Memory Architecture,
Tao Zhang, Matt Poremba, Cong Xu, Guangyu Sun, and Yuan Xie, “CREAM: A Concurrent-Refresh-Aware DRAM Memory Architecture, ” inHPCA, 2014
2014
-
[212]
Flexible Auto-Refresh: Enabling Scalable and Energy-Efficient DRAM Refresh Reductions,
Ishwar Bhati, Zeshan Chishti, Shih-Lien Lu, and Bruce Jacob, “Flexible Auto-Refresh: Enabling Scalable and Energy-Efficient DRAM Refresh Reductions, ” inISCA, 2015
2015
-
[213]
Omitting Refresh: A Case Study for Commod- ity and Wide I/O DRAMs,
Matthias Jung, Éder Zulian, Deepak M. Mathew, Matthias Herrmann, Christian Brug- ger, Christian Weis, and Norbert Wehn, “Omitting Refresh: A Case Study for Commod- ity and Wide I/O DRAMs, ” inMEMSYS, 2015
2015
-
[214]
DRAM Refresh Mechanisms, Penalties, and Trade-Offs,
Ishwar Bhati, Mu-Tien Chang, Zeshan Chishti, Shih-Lien Lu, and Bruce Jacob, “DRAM Refresh Mechanisms, Penalties, and Trade-Offs, ” inTC, 2016. BIBLIOGRAPHY 214
2016
-
[215]
Hardware-Software Co-Design to Mitigate DRAM Refresh Overheads: A Case for Refresh-Aware Process Scheduling,
Jagadish B Kotra, Narges Shahidi, Zeshan A Chishti, and Mahmut T Kandemir, “Hardware-Software Co-Design to Mitigate DRAM Refresh Overheads: A Case for Refresh-Aware Process Scheduling, ”ASPLOS, 2017
2017
-
[216]
Nonblocking Memory Refresh,
Kate Nguyen, Kehan Lyu, Xianze Meng, Vilas Sridharan, and Xun Jian, “Nonblocking Memory Refresh, ” inISCA, 2018
2018
-
[217]
Content Aware Refresh: Exploiting the Asymmetry of DRAM Retention Errors to Reduce the Refresh Frequency of Less Vulnerable Data,
Shibo Wang, Mahdi Nazm Bojnordi, Xiaochen Guo, and Engin Ipek, “Content Aware Refresh: Exploiting the Asymmetry of DRAM Retention Errors to Reduce the Refresh Frequency of Less Vulnerable Data, ”IEEE TC, 2018
2018
-
[218]
Hiding DRAM Refresh Overhead in Real-Time Cyclic Executives,
Xing Pan and Frank Mueller, “Hiding DRAM Refresh Overhead in Real-Time Cyclic Executives, ” inRTSS, 2019
2019
-
[219]
The Colored Refresh Server for DRAM,
Xing Pan and Frank Mueller, “The Colored Refresh Server for DRAM, ” inISORC, 2019
2019
-
[220]
Reducing DRAM Refresh Power Consumption by Runtime Profiling of Retention Time and Dual-Row Activa- tion,
Haerang Choi, Dosun Hong, Jaesung Lee, and Sungjoo Yoo, “Reducing DRAM Refresh Power Consumption by Runtime Profiling of Retention Time and Dual-Row Activa- tion, ”MICPRO, 2020
2020
-
[221]
Charge-Aware DRAM Refresh Reduction with Value Transformation,
Seikwon Kim, Wonsang Kwak, Changdae Kim, Daehyeon Baek, and Jaehyuk Huh, “Charge-Aware DRAM Refresh Reduction with Value Transformation, ” inHPCA, 2020
2020
-
[222]
AVATAR: A Variable-Retention-Time (VRT) Aware Refresh for DRAM Systems,
Moinuddin K Qureshi, DaeHyun Kim, Samira Khan, Prashant J Nair, and Onur Mutlu, “AVATAR: A Variable-Retention-Time (VRT) Aware Refresh for DRAM Systems, ” in DSN, 2015
2015
-
[223]
Reducing Refresh Power in Mobile Devices with Morphable ECC,
Chiachen Chou, Prashant Nair, and Moinuddin K Qureshi, “Reducing Refresh Power in Mobile Devices with Morphable ECC, ” inDSN, 2015
2015
-
[224]
Improving DRAM Performance by Parallelizing Refreshes with Accesses,
Kevin Kai-Wei Chang, Donghyuk Lee, Zeshan Chishti, Alaa R Alameldeen, Chris Wilk- erson, Yoongu Kim, and Onur Mutlu, “Improving DRAM Performance by Parallelizing Refreshes with Accesses, ” inHPCA, 2014
2014
-
[225]
VRL-DRAM: Improving DRAM Perfor- mance via Variable Refresh Latency,
Anup Das, Hasan Hassan, and Onur Mutlu, “VRL-DRAM: Improving DRAM Perfor- mance via Variable Refresh Latency, ” inDAC, 2018
2018
-
[226]
Understanding and Mitigating Refresh Overheads in High-Density DDR4 DRAM Systems,
J. Mukundan, H. Hunter, K. H. Kim, J. Stuecheli, and J. F. Martínez, “Understanding and Mitigating Refresh Overheads in High-Density DDR4 DRAM Systems, ” inISCA, 2013
2013
-
[227]
Smart Refresh: An Enhanced Memory Con- troller Design for Reducing Energy in Conventional and 3D Die-Stacked DRAMs,
Mrinmoy Ghosh and Hsien-Hsin S Lee, “Smart Refresh: An Enhanced Memory Con- troller Design for Reducing Energy in Conventional and 3D Die-Stacked DRAMs, ” in MICRO, 2007
2007
-
[228]
Long-Retention-Time, High- Speed DRAM Array with 12-F 2 Twin Cell for Sub 1-V Operation,
Riichiro Takemura, Kiyoo Itoh, Tomonori Sekiguchi, Satoru Akiyama, Satoru Han- zawa, Kazuhiko Kajigaya, and Takayuki Kawahara, “Long-Retention-Time, High- Speed DRAM Array with 12-F 2 Twin Cell for Sub 1-V Operation, ”TOE, 2007
2007
-
[229]
Analysis of Retention Time Distribution of Embedded DRAM-A New Method to Characterize Across-Chip Threshold Voltage Variation,
Wei Kong, Paul C Parries, G Wang, and Subramanian S Iyer, “Analysis of Retention Time Distribution of Embedded DRAM-A New Method to Characterize Across-Chip Threshold Voltage Variation, ” inITC, 2008. BIBLIOGRAPHY 215
2008
-
[230]
Characterization of the Variable Retention Time in Dynamic Random Access Memory,
Heesang Kim, Byoungchan Oh, Younghwan Son, Kyungdo Kim, Seon-Yong Cha, Jae- Goan Jeong, Sung-Joo Hong, and Hyungcheol Shin, “Characterization of the Variable Retention Time in Dynamic Random Access Memory, ”TED, 2011
2011
-
[231]
Study of Trap Models Related to the Variable Retention Time Phenomenon in DRAM,
Heesang Kim, Byoungchan Oh, Younghwan Son, Kyungdo Kim, Seon-Yong Cha, Jae- Goan Jeong, Sung-Joo Hong, and Hyungcheol Shin, “Study of Trap Models Related to the Variable Retention Time Phenomenon in DRAM, ”TED, 2011
2011
-
[232]
Characterization of Data Retention Faults in DRAM Devices,
Angelo Bacchini, Marco Rovatti, Gianluca Furano, and Marco Ottavi, “Characterization of Data Retention Faults in DRAM Devices, ” inDFT, 2014
2014
-
[233]
Total Ionizing Dose Effects On DRAM Data Retention Time,
Angelo Bacchini, Gianluca Furano, Marco Rovatti, and Marco Ottavi, “Total Ionizing Dose Effects On DRAM Data Retention Time, ”IEEE Trans. Nucl. Sci. , 2014
2014
-
[234]
ProactiveDRAM: A DRAM-Initiated Reten- tion Management Scheme,
Jue Wang, Xiangyu Dong, and Yuan Xie, “ProactiveDRAM: A DRAM-Initiated Reten- tion Management Scheme, ” inICCD, 2014
2014
-
[235]
RADAR: A Case for Retention-Aware DRAM Assembly and Repair in Future FGR DRAM Memory,
Ying Wang, Yinhe Han, Cheng Wang, Huawei Li, and Xiaowei Li, “RADAR: A Case for Retention-Aware DRAM Assembly and Repair in Future FGR DRAM Memory, ” in DAC, 2015
2015
-
[236]
Retention Time Measurements and Modelling of Bit Error Rates of Wide I/O DRAM in MPSoCs,
Christian Weis, Matthias Jung, Peter Ehses, Cristiano Santos, Pascal Vivet, Sven Goossens, Martijn Koedam, and Norbert Wehn, “Retention Time Measurements and Modelling of Bit Error Rates of Wide I/O DRAM in MPSoCs, ” inDATE, 2015
2015
-
[237]
NDA: Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules,
Amin Farmahini-Farahani, Jung Ho Ahn, Katherine Morrow, and Nam Sung Kim, “NDA: Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules, ” inHPCA, 2015
2015
-
[238]
JAFAR: Near-Data Processing for Databases,
Oreoluwatomiwa O Babarinsa and Stratos Idreos, “JAFAR: Near-Data Processing for Databases, ” inSIGMOD, 2015
2015
-
[239]
The True Processing in Memory Accelerator,
Fabrice Devaux, “The True Processing in Memory Accelerator, ” inHot Chips, 2019
2019
-
[240]
GenStore: A High-Performance and Energy-Efficient In-Storage Computing System for Genome Sequence Analysis,
Nika Mansouri Ghiasi, Jisung Park, Harun Mustafa, Jeremie Kim, Ataberk Olgun, Arvid Gollwitzer, Damla Senol Cali, Can Firtina, Haiyu Mao, Nour Almadhoun Alserr et al., “GenStore: A High-Performance and Energy-Efficient In-Storage Computing System for Genome Sequence Analysis, ...
2022
-
[241]
Benchmarking Memory-Centric Computing Systems: Anal- ysis of Real Processing-in-Memory Hardware,
Juan Gómez-Luna, Izzat El Hajj, Ivan Fernandez, Christina Giannoula, Geraldo F Oliveira, and Onur Mutlu, “Benchmarking Memory-Centric Computing Systems: Anal- ysis of Real Processing-in-Memory Hardware, ” inCUT, 2021
2021
-
[242]
Benchmarking a New Paradigm: An Experimental Anal- ysis of a Real Processing-in-Memory Architecture,
Juan Gómez-Luna, Izzat El Hajj, Ivan Fernández, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu, “Benchmarking a New Paradigm: An Experimental Anal- ysis of a Real Processing-in-Memory Architecture, ” arXiv:2105.03814 [cs.AR], 2021
2021 arXiv
-
[243]
SynCron: Efficient Synchronization Support for Near-Data- Processing Architectures,
Christina Giannoula, Nandita Vijaykumar, Nikela Papadopoulou, Vasileios Karakostas, Ivan Fernandez, Juan Gómez-Luna, Lois Orosa, Nectarios Koziris, Georgios Goumas, and Onur Mutlu, “SynCron: Efficient Synchronization Support for Near-Data- Processing Architectures, ” inHPCA, 2...
2021
-
[244]
NERO: A Near High- Bandwidth Memory Stencil Accelerator for Weather Prediction Modeling,
Gagandeep Singh, Dionysios Diamantopoulos, Christoph Hagleitner, Juan Gomez- Luna, Sander Stuijk, Onur Mutlu, and Henk Corporaal, “NERO: A Near High- Bandwidth Memory Stencil Accelerator for Weather Prediction Modeling, ” inFPL, 2020
2020
-
[245]
A 1ynm 1.25V 8Gb, 16Gb/s/pin GDDR6-based Accelerator-in-Memory Supporting 1TFLOPS MAC Opera- tion and Various Activation Functions for Deep-Learning Applications,
S. Lee, K. Kim, S. Oh, J. Park, G. Hong, D. Ka, K. Hwang, J. Park, K. Kang, J. Kim, J. Jeon, N. Kim, Y. Kwon, K. Vladimir, W. Shin, J. Won, M. Lee, H. Jooet al., “A 1ynm 1.25V 8Gb, 16Gb/s/pin GDDR6-based Accelerator-in-Memory Supporting 1TFLOPS MAC Opera- tion and Various Acti...
2022
-
[246]
Near-Memory Processing in Action: Accelerating Personalized Recommendation with AxDIMM,
Liu Ke, Xuan Zhang, Jinin So, Jong-Geon Lee, Shin-Haeng Kang, Sukhan Lee, Songyi Han, Yeongon Cho, Jin Hyun Kim, Yongsuk Kwonet al., “Near-Memory Processing in Action: Accelerating Personalized Recommendation with AxDIMM, ”IEEE Micro, 2021
2021
-
[247]
SparseP: Towards Efficient Sparse Matrix Vector Multipli- cation on Real Processing-in-Memory Architectures,
Christina Giannoula, Ivan Fernandez, Juan Gómez Luna, Nectarios Koziris, Georgios Goumas, and Onur Mutlu, “SparseP: Towards Efficient Sparse Matrix Vector Multipli- cation on Real Processing-in-Memory Architectures, ” inSIGMETRICS, 2022
2022
-
[248]
McDRAM: Low Latency and Energy-Efficient Matrix Computations in DRAM,
Hyunsung Shin, Dongyoung Kim, Eunhyeok Park, Sungho Park, Yongsik Park, and Sungjoo Yoo, “McDRAM: Low Latency and Energy-Efficient Matrix Computations in DRAM, ”IEEE TCADICS, 2018
2018
-
[249]
McDRAM v2: In-Dynamic Random Access Memory Systolic Array Accelerator to Ad- dress the Large Model Problem in Deep Neural Networks on the Edge,
Seunghwan Cho, Haerang Choi, Eunhyeok Park, Hyunsung Shin, and Sungjoo Yoo, “McDRAM v2: In-Dynamic Random Access Memory Systolic Array Accelerator to Ad- dress the Large Model Problem in Deep Neural Networks on the Edge, ” IEEE Access, 2020
2020
-
[250]
Casper: Accelerating Stencil Computation using Near-Cache Processing,
Alain Denzler, Rahul Bera, Nastaran Hajinazar, Gagandeep Singh, Geraldo F Oliveira, Juan Gómez-Luna, and Onur Mutlu, “Casper: Accelerating Stencil Computation using Near-Cache Processing, ” arXiv:2112.14216 [cs.AR], 2021
2021 arXiv
-
[251]
Chameleon: Versatile and Practical Near-DRAM Acceleration Architecture for Large Memory Systems,
Hadi Asghari-Moghaddam, Young Hoon Son, Jung Ho Ahn, and Nam Sung Kim, “Chameleon: Versatile and Practical Near-DRAM Acceleration Architecture for Large Memory Systems, ” inMICRO, 2016
2016
-
[252]
A Case for Intelligent RAM,
D. Patterson, T. Anderson, N. Cardwell et al., “A Case for Intelligent RAM, ”IEEE Micro, 1997
1997
-
[253]
Computational RAM: Implementing Processors in Memory,
D. G. Elliott, M. Stumm, W. M. Snelgrove et al., “Computational RAM: Implementing Processors in Memory, ”Design and Test of Computers , 1999
1999
-
[254]
Saving Memory Movements Through Vector Processing in the DRAM,
M. A. Z. Alves, P. C. Santos, F. B. Moreira, and opthers, “Saving Memory Movements Through Vector Processing in the DRAM, ” inCASES, 2015
2015
-
[255]
Beyond the Wall: Near-Data Processing for Databases,
S. L. Xi, O. Babarinsa, M. Athanassoulis, and S. Idreos, “Beyond the Wall: Near-Data Processing for Databases, ” inDaMoN, 2015
2015
-
[256]
ABC-DIMM: Alleviat- ing the Bottleneck of Communication in DIMM-Based Near-Memory Processing with Inter-DIMM Broadcast,
Weiyi Sun, Zhaoshi Li, Shouyi Yin, Shaojun Wei, and Leibo Liu, “ABC-DIMM: Alleviat- ing the Bottleneck of Communication in DIMM-Based Near-Memory Processing with Inter-DIMM Broadcast, ” inISCA, 2021. BIBLIOGRAPHY 217
2021
-
[257]
GraphSSD: Graph Semantics Aware SSD,
Kiran Kumar Matam, Gunjae Koo, Haipeng Zha, Hung-Wei Tseng, and Murali An- navaram, “GraphSSD: Graph Semantics Aware SSD, ” inISCA, 2019
2019
-
[258]
Processing in Memory: The Terasys Mas- sively Parallel PIM Array,
Maya Gokhale, Bill Holmes, and Ken Iobst, “Processing in Memory: The Terasys Mas- sively Parallel PIM Array, ”Computer, 1995
1995
-
[259]
Mapping Irregular Applications to DIVA, a PIM-Based Data-Intensive Architecture,
Mary Hall, Peter Kogge, Jeff Koller, Pedro Diniz, Jacqueline Chame, Jeff Draper, Jeff LaCoss, John Granacki, Jay Brockman, Apoorv Srivastava et al. , “Mapping Irregular Applications to DIVA, a PIM-Based Data-Intensive Architecture, ” inSC, 1999
1999
-
[260]
Opportunities and Challenges of Performing Vector Operations Inside the DRAM,
Marco A. Z. Alves, Paulo C. Santos, Matthias Diener, and Luigi Carro, “Opportunities and Challenges of Performing Vector Operations Inside the DRAM, ” inMEMSYS, 2015
2015
-
[261]
Livia: Data-Centric Com- puting Throughout the Memory Hierarchy,
Elliot Lockerman, Axel Feldmann, Mohammad Bakhshalipour, Alexandru Stanescu, Shashwat Gupta, Daniel Sanchez, and Nathan Beckmann, “Livia: Data-Centric Com- puting Throughout the Memory Hierarchy, ” inASPLOS, 2020
2020
-
[262]
LazyPIM: An Efficient Cache Coherence Mech- anism for Processing-in-Memory,
Amirali Boroumand, Saugata Ghose, Brandon Lucia, Kevin Hsieh, Krishna Malladi, Hongzhong Zheng, and Onur Mutlu, “LazyPIM: An Efficient Cache Coherence Mech- anism for Processing-in-Memory, ”CAL, 2017
2017
-
[263]
TOP-PIM: Throughput-Oriented Programmable Process- ing in Memory,
Dongping Zhang, Nuwan Jayasena, Alexander Lyashevsky, Joseph L Greathouse, Lifan Xu, and Michael Ignatowski, “TOP-PIM: Throughput-Oriented Programmable Process- ing in Memory, ” inHPDC, 2014
2014
-
[264]
HRL: Efficient and Flexible Reconfigurable Logic for Near-Data Processing,
Mingyu Gao and Christos Kozyrakis, “HRL: Efficient and Flexible Reconfigurable Logic for Near-Data Processing, ” inHPCA, 2016
2016
-
[265]
The Mondrian Data Engine,
Mario Drumond, Alexandros Daglis, Nooshin Mirzadeh, Dmitrii Ustiugov, Javier Pi- corel, Babak Falsafi, Boris Grot, and Dionisios Pnevmatikatos, “The Mondrian Data Engine, ” inISCA, 2017
2017
-
[266]
Operand Size Reconfiguration for Big Data Processing in Memory,
P. C. Santos, G. F. Oliveira, D. G. Tomé, M. A. Z. Alves, E. C. Almeida, and L. Carro, “Operand Size Reconfiguration for Big Data Processing in Memory, ” inDATE, 2017
2017
-
[267]
NIM: An HMC- Based Machine for Neuron Computation,
Geraldo F Oliveira, Paulo C Santos, Marco AZ Alves, and Luigi Carro, “NIM: An HMC- Based Machine for Neuron Computation, ” inARC, 2017
2017
-
[268]
TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory,
Mingyu Gao, Jing Pu, Xuan Yang, Mark Horowitz, and Christos Kozyrakis, “TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory, ” inASPLOS, 2017
2017
-
[269]
Neurocube: A Programmable Digital Neuromorphic Architecture with High- Density 3D Memory,
Duckhwan Kim, Jaeha Kung, Sek Chai, Sudhakar Yalamanchili, and Saibal Mukhopad- hyay, “Neurocube: A Programmable Digital Neuromorphic Architecture with High- Density 3D Memory, ” inISCA, 2016
2016
-
[270]
Leveraging 3D Technologies for Hardware Security: Opportunities and Chal- lenges,
Peng Gu, Shuangchen Li, Dylan Stow, Russell Barnes, Liu Liu, Yuan Xie, and Eren Kursun, “Leveraging 3D Technologies for Hardware Security: Opportunities and Chal- lenges, ” inGLSVLSI, 2016. BIBLIOGRAPHY 218
2016
-
[271]
CoNDA: Efficient Cache Coherence Support for Near-Data Accelerators,
Amirali Boroumand, Saugata Ghose, Minesh Patel, Hasan Hassan, Brandon Lu- cia, Rachata Ausavarungnirun, Kevin Hsieh, Nastaran Hajinazar, Krishna T Malladi, Hongzhong Zheng et al., “CoNDA: Efficient Cache Coherence Support for Near-Data Accelerators, ” inISCA, 2019
2019
-
[272]
Transparent Offloading and Mapping (TOM) Enabling Programmer-Transparent Near-Data Processing in GPU Systems,
Kevin Hsieh, Eiman Ebrahimi, Gwangsun Kim, Niladrish Chatterjee, Mike O’Connor, Nandita Vijaykumar, Onur Mutlu, and Stephen W Keckler, “Transparent Offloading and Mapping (TOM) Enabling Programmer-Transparent Near-Data Processing in GPU Systems, ” inISCA, 2016
2016
-
[273]
NDC: Analyzing the Im- pact of 3D-Stacked Memory+Logic Devices on MapReduce Workloads,
S. H. Pugsley, J. Jestes, H. Zhang, R. Balasubramonian et al., “NDC: Analyzing the Im- pact of 3D-Stacked Memory+Logic Devices on MapReduce Workloads, ” inISPASS, 2014
2014
-
[274]
Scheduling Techniques for GPU Architec- tures with Processing-in-Memory Capabilities,
Ashutosh Pattnaik, Xulong Tang, Adwait Jog, Onur Kayiran, Asit K Mishra, Mahmut T Kandemir, Onur Mutlu, and Chita R Das, “Scheduling Techniques for GPU Architec- tures with Processing-in-Memory Capabilities, ” inPACT, 2016
2016
-
[275]
Data Reorganization in Memory Using 3D-Stacked DRAM,
Berkin Akin, Franz Franchetti, and James C Hoe, “Data Reorganization in Memory Using 3D-Stacked DRAM, ” inISCA, 2015
2015
-
[276]
BSSync: Processing Near Mem- ory for Machine Learning Workloads with Bounded Staleness Consistency Models,
Joo Hwan Lee, Jaewoong Sim, and Hyesoon Kim, “BSSync: Processing Near Mem- ory for Machine Learning Workloads with Bounded Staleness Consistency Models, ” in PACT, 2015
2015
-
[277]
Mitigating Edge Machine Learn- ing Inference Bottlenecks: An Empirical Study on Accelerating Google Edge Models,
Amirali Boroumand, Saugata Ghose, Berkin Akin, Ravi Narayanaswami, Geraldo F Oliveira, Xiaoyu Ma, Eric Shiu, and Onur Mutlu, “Mitigating Edge Machine Learn- ing Inference Bottlenecks: An Empirical Study on Accelerating Google Edge Models, ” arXiv:2103.00768 [cs.AR], 2021
2021 arXiv
-
[278]
Polyne- sia: Enabling High-Performance and Energy-Efficient Hybrid Transactional/Analytical Databases with Hardware/Software Co-Design,
Amirali Boroumand, Saugata Ghose, Geraldo F Oliveira, and Onur Mutlu, “Polyne- sia: Enabling High-Performance and Energy-Efficient Hybrid Transactional/Analytical Databases with Hardware/Software Co-Design, ” inICDE, 2022
2022
-
[279]
Polynesia: Enabling Effective Hybrid Transactional/Analytical Databases with Specialized Hard- ware/Software Co-Design,
Amirali Boroumand, Saugata Ghose, Geraldo F Oliveira, and Onur Mutlu, “Polynesia: Enabling Effective Hybrid Transactional/Analytical Databases with Specialized Hard- ware/Software Co-Design, ” arXiv:2103.00798 [cs.AR], 2021
2021 arXiv
-
[280]
SISA: Set-Centric Instruction Set Ar- chitecture for Graph Mining on Processing-in-Memory Systems,
Maciej Besta, Raghavendra Kanakagiri, Grzegorz Kwasniewski, Rachata Ausavarung- nirun, Jakub Beránek, Konstantinos Kanellopoulos, Kacper Janda, Zur Vonarburg- Shmaria, Lukas Gianinazzi, Ioana Stefan et al., “SISA: Set-Centric Instruction Set Ar- chitecture for Graph Mining on ...
2021
-
[281]
NATSA: A Near-Data Process- ing Accelerator for Time Series Analysis,
Ivan Fernandez, Ricardo Quislant, Eladio Gutiérrez, Oscar Plata, Christina Giannoula, Mohammed Alser, Juan Gómez-Luna, and Onur Mutlu, “NATSA: A Near-Data Process- ing Accelerator for Time Series Analysis, ” inICCD, 2020
2020
-
[282]
NAPEL: Near-Memory Computing Application Perfor- mance Prediction via Ensemble Learning,
Gagandeep Singh, Giovanni , Geraldo F Oliveira, Stefano Corda, Sander Stuijk, Onur Mutlu, and Henk Corporaal, “NAPEL: Near-Memory Computing Application Perfor- mance Prediction via Ensemble Learning, ” inDAC, 2019. BIBLIOGRAPHY 219
2019
-
[283]
A 20nm 6GB Function- in-Memory DRAM, Based on HBM2 with a 1.2 TFLOPS Programmable Computing Unit using Bank-Level Parallelism, for Machine Learning Applications,
Young-Cheon Kwon, Suk Han Lee, Jaehoon Lee, Sang-Hyuk Kwon, Je Min Ryu, Jong-Pil Son, O Seongil, Hak-Soo Yu, Haesuk Lee, Soo Young Kimet al., “A 20nm 6GB Function- in-Memory DRAM, Based on HBM2 with a 1.2 TFLOPS Programmable Computing Unit using Bank-Level Parallelism, for Mac...
2021
-
[284]
Hardware Ar- chitecture and Software Stack for PIM Based on Commercial DRAM Technology: In- dustrial Product,
Sukhan Lee, Shin-haeng Kang, Jaehoon Lee, Hyeonsu Kim, Eojin Lee, Seungwoo Seo, Hosang Yoon, Seungwon Lee, Kyounghwan Lim, Hyunsung Shinet al., “Hardware Ar- chitecture and Software Stack for PIM Based on Commercial DRAM Technology: In- dustrial Product, ” inISCA, 2021
2021
-
[285]
184QPS/W 64Mb/𝑚𝑚2 3D Logic-to-DRAM Hybrid Bonding with Process-Near-Memory Engine for Recommendation System,
Dimin Niu, Shuangchen Li, Yuhao Wang, Wei Han, Zhe Zhang, Yijin Guan, Tianchan Guan, Fei Sun, Fei Xue, Lide Duan et al., “184QPS/W 64Mb/𝑚𝑚2 3D Logic-to-DRAM Hybrid Bonding with Process-Near-Memory Engine for Recommendation System, ” in ISSCC, 2022
2022
-
[286]
Accelerating Sparse Matrix- Matrix Multiplication with 3D-Stacked Logic-in-Memory Hardware,
Q. Zhu, T. Graf, H. E. Sumbul, L. Pileggi, and F. Franchetti, “Accelerating Sparse Matrix- Matrix Multiplication with 3D-Stacked Logic-in-Memory Hardware, ” inHPEC, 2013
2013
-
[287]
Logic- Base Interconnect Design for Near Memory Computing in the Smart Memory Cube,
Erfan Azarkhish, Christoph Pfister, Davide Rossi, Igor Loi, and Luca Benini, “Logic- Base Interconnect Design for Near Memory Computing in the Smart Memory Cube, ” IEEE VLSI, 2016
2016
-
[288]
Neurostream: Scalable and Energy Efficient Deep Learning with Smart Memory Cubes,
Erfan Azarkhish, Davide Rossi, Igor Loi, and Luca Benini, “Neurostream: Scalable and Energy Efficient Deep Learning with Smart Memory Cubes, ”TPDS, 2018
2018
-
[289]
3D-Stacked Memory-Side Acceler- ation: Accelerator and System Design,
Qi Guo, Nikolaos Alachiotis, Berkin Akin, Fazle Sadi, Guanglin Xu, Tze Meng Low, Larry Pileggi, James C Hoe, and Franz Franchetti, “3D-Stacked Memory-Side Acceler- ation: Accelerator and System Design, ” inWoNDP, 2014
2014
-
[290]
Design Space Exploration for PIM Architectures in 3D-Stacked Mem- ories,
Jo textasciitilde ao Paulo C de Lima, Paulo Cesar Santos, Marco AZ Alves, Antonio Beck, and Luigi Carro, “Design Space Exploration for PIM Architectures in 3D-Stacked Mem- ories, ” inCF, 2018
2018
-
[291]
HAMLeT: Hardware Accelerated Memory Layout Transform within 3D-Stacked DRAM,
Berkin Akın, James C Hoe, and Franz Franchetti, “HAMLeT: Hardware Accelerated Memory Layout Transform within 3D-Stacked DRAM, ” inHPEC, 2014
2014
-
[292]
A Heterogeneous PIM Hardware-Software Co-Design for Energy-Efficient Graph Processing,
Yu Huang, Long Zheng, Pengcheng Yao, Jieshan Zhao, Xiaofei Liao, Hai Jin, and Jin- gling Xue, “A Heterogeneous PIM Hardware-Software Co-Design for Energy-Efficient Graph Processing, ” inIPDPS, 2020
2020
-
[293]
GraphH: A Processing-in-Memory Archi- tecture for Large-Scale Graph Processing,
Guohao Dai, Tianhao Huang, Yuze Chi, Jishen Zhao, Guangyu Sun, Yongpan Liu, Yu Wang, Yuan Xie, and Huazhong Yang, “GraphH: A Processing-in-Memory Archi- tecture for Large-Scale Graph Processing, ”TCAD, 2018
2018
-
[294]
Processing- in-Memory for Energy-Efficient Neural Network Training: A Heterogeneous Ap- proach,
Jiawen Liu, Hengyu Zhao, Matheus A Ogleari, Dong Li, and Jishen Zhao, “Processing- in-Memory for Energy-Efficient Neural Network Training: A Heterogeneous Ap- proach, ” inMICRO, 2018. BIBLIOGRAPHY 220
2018
-
[295]
iPIM: Programmable In-Memory Image Processing Accelerator using Near- Bank Architecture,
Peng Gu, Xinfeng Xie, Yufei Ding, Guoyang Chen, Weifeng Zhang, Dimin Niu, and Yuan Xie, “iPIM: Programmable In-Memory Image Processing Accelerator using Near- Bank Architecture, ” inISCA, 2020
2020
-
[296]
DRAMA: An Architec- ture for Accelerated Processing Near Memory,
A. Farmahini-Farahani, J. H. Ahn, K. Compton, and N. S. Kim, “DRAMA: An Architec- ture for Accelerated Processing Near Memory, ”Computer Architecture Letters, 2014
2014
-
[297]
Near-DRAM Ac- celeration with Single-ISA Heterogeneous Processing in Standard Memory Modules,
H. Asghari-Moghaddam, A. Farmahini-Farahani, K. Morrow et al., “Near-DRAM Ac- celeration with Single-ISA Heterogeneous Processing in Standard Memory Modules, ” IEEE Micro, 2016
2016
-
[298]
Active-Routing: Compute on the Way for Near-Data Processing,
Jiayi Huang, Ramprakash Reddy Puli, Pritam Majumder, Sungkeun Kim, Rahul Boy- apati, Ki Hwan Yum, and Eun Jung Kim, “Active-Routing: Compute on the Way for Near-Data Processing, ” inHPCA, 2019
2019
-
[299]
Lightweight SIMT Core Designs for Intelligent 3D Stacked DRAM,
Chad D Kersey, Hyesoon Kim, and Sudhakar Yalamanchili, “Lightweight SIMT Core Designs for Intelligent 3D Stacked DRAM, ” inMEMSYS, 2017
2017
-
[300]
PIMS: A Lightweight Processing-in-Memory Accelerator for Stencil Computations,
Jie Li, Xi Wang, Antonino Tumeo, Brody Williams, John D Leidel, and Yong Chen, “PIMS: A Lightweight Processing-in-Memory Accelerator for Stencil Computations, ” in MEMSYS, 2019
2019
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.