Pith. sign in

REVIEW 4 major objections 5 minor 111 references

Near-Memory Computing: Past, Present, and Future

T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Near-memory computing can cut data movement, but the hard part is software.

desk verdict A useful NMC survey with a solid taxonomy, but the forward-looking quantitative claim rests on an idealized, unvalidated analytic model that should be either fixed or explicitly softened before publication. read the letter →

arxiv 1908.02640 v1 pith:F7FTJXZK submitted 2019-08-07 cs.AR cs.DCcs.PF

classification cs.ARcs.DCcs.PF
keywords near-memorycomputingprocessing-in-memorydata-centric3D-stackedmemoryapplicationcharacterizationcompileroffloadinganalyticalperformancemodelwall
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Near-memory computing moves processing into the logic layer of 3D-stacked memory so that data-intensive workloads no longer pay the full cost of shuttling data to the CPU. This paper tries to establish that the decade-old NMC concept is now technologically viable, and that the remaining obstacles are mostly software. It supports that view with a taxonomy organizing the field by memory level, processing-unit type, tooling, interoperability, and application domain, and with a concrete software stack: platform-agnostic metrics for spotting NMC-friendly kernels, a compiler that transparently swaps those kernels for accelerator-library calls, and a first-order analytic model. The model predicts that as L1/L2 miss rates rise, a CPU-centric system falls further behind an NMC system, so the payoff concentrates in low-locality, memory-bound code. A sympathetic reader is left with the conclusion that the key barriers are coherence, virtual memory, programming models, and design-space tooling, not basic hardware feasibility.

What carries the argument

The load-bearing machinery is two-tier. The hardware enabler is the 3D-stacked memory cube, whose DRAM layers sit on a logic layer connected by through-silicon vias, with memory divided into vaults that provide high internal bandwidth. On the software side, the paper's framework rests on four microarchitecture-independent characterization metrics: memory entropy (randomness of the address stream), spatial locality (likelihood of nearby accesses), data-level parallelism per opcode, and basic-block-level parallelism, computed from instruction traces. These metrics let the compiler identify NMC candidates without depending on a specific microarchitecture. The compiler flow extracts kernels in a polyhedral intermediate representation, matches them against a library of accelerator routines, and swaps them in without programmer intervention. The classification table does the organizing work for the survey: every architecture is placed by memory level, processing-unit type and granularity, host, evaluation technique, and the presence or absence of coherence and virtual-memory support, which is what lets the paper locate recurring gaps.

What would settle it

Run the paper's analytic model and a cycle-accurate HMC-style simulator on the same low-locality kernels, replacing the perfect-vault-parallelism assumption with a realistic vault scheduler and row-buffer conflicts. If the simulated NMC delay and energy are not lower than the CPU-centric values whenever L1/L2 miss rates are high, the paper's illustrative claim that low-locality applications benefit from NMC would be overturned.

Watch

Extended reading notes

Core claim

The central claim is that processing 'at the home of data' can substantially reduce the data-movement bottleneck that dominates emerging scale-out workloads with limited data reuse. The paper defends this claim historically and architecturally: 3D stacking with through-silicon vias makes it practical to place programmable, fixed-function, or reconfigurable compute on the memory logic layer, and the survey of prior systems shows the idea has repeatedly delivered performance and energy gains in simulators and prototypes, yet failed to penetrate markets because of earlier technology limits and, now, because of system-software gaps. For the present, the paper contributes a classification scheme covering memory hierarchy, memory type, integration style, NMC unit, implementation granularity, host, evaluation technique, programming model, coherence, virtual memory, and workload. It also presents its own framework: microarchitecture-independent metrics (memory entropy, spatial locality, data-level parallelism, basic-block-level parallelism) to identify NMC-suitable kernels; a compiler that uses a polyhedral intermediate representation and pattern matching to replace detected kernels with accelerator runtime calls; and an analytic model comparing a multicore host with an HMC-like NMC system. The model's illustrative result is that increasing L1 and L2 miss rates degrades the multicore system more than the NMC system in both delay and energy, so low-locality applications benefit from NMC while high-locality applications are better off on the host.

Load-bearing premise

The analytic model assumes that in the NMC case all accesses are made to different vaults inside the HMC memory, that is, perfect vault-level parallelism with no contention, while the CPU-centric alternative is charged with specified L1/L2 miss penalties; if real NMC workloads suffer vault imbalance, row conflicts, or coherence traffic, the modeled advantage shrinks.

Editorial extensions

If this is right

  • If the framework works as argued, existing C/C++ code can gain NMC acceleration without rewriting, because kernel detection and offloading happen in the compiler.
  • As cache miss rates climb, the gap between CPU-centric and NMC systems widens in both delay and energy, so NMC's payoff grows exactly where today's data-intensive workloads are worst.
  • Low-locality kernels such as graph traversal, sparse linear algebra, and streaming analytics are the natural NMC candidates; high-locality kernels should stay on the host.
  • Widespread NMC adoption depends on solving coherence, virtual memory, data mapping, and programming-model support, since these recur across nearly all surveyed systems.
  • A standard benchmark suite and open-source simulators are prerequisites for comparing future NMC designs, which the paper identifies as an open gap.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An editorial extension: the four metrics could be turned into a predictive model, so that a kernel's entropy and spatial-locality scores become a cheap pre-silicon filter for offload decisions; the paper hints at this but does not establish the correlation.
  • The compiler's pattern-matching approach may generalize to heterogeneous NMC units and to emerging cache-coherent interconnects, potentially allowing NMC to be adopted incrementally in data centers without new programming languages.
  • The analytic model's assumption of perfect vault-level parallelism means real HMC-style systems with vault imbalance, row conflicts, and coherence traffic may realize a smaller fraction of the modeled gains; a sensitivity study would bound that fraction.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This survey paper organizes the near-memory computing (NMC) literature, proposing a taxonomy by memory hierarchy, processing unit, implementation type, tool support, interoperability, and application domain. It then outlines the authors' own NMC efforts: microarchitecture-independent application characterization metrics, an LLVM/polyhedral-based compiler flow for transparent offloading of NMC kernels, and a first-order analytic model comparing a CPU-centric multicore system with an NMC+host system. The paper argues that NMC can significantly reduce data movement for data-intensive, low-locality workloads, and identifies open challenges including coherence, virtual memory, programming models, and design-space exploration tools.

Significance. If the quantitative claims were fully supported, this paper would be a valuable systematization of the NMC field and a useful position piece on the software and modeling infrastructure needed for NMC adoption. The survey portion is genuinely useful: the taxonomy in Tables 1 and 2 and the discussion of challenges in Section VI synthesize a broad body of work and should serve as a convenient reference. The proposed platform-agnostic characterization metrics and compiler framework are promising ideas that align with the field's need for transparent offloading, though they are not evaluated in this manuscript. The analytic model in Section VIII-D, however, is the only quantitative support for the paper's central claim, and its assumptions are currently too idealized and opaque to sustain the conclusions drawn from it.

major comments (4)
  1. [Section VIII-D, Figure 8] The NMC-vs-CPU comparison is asymmetric and the NMC side is granted an idealized access pattern. The text states, in the 'Performance and Energy Exploration' paragraph, that for the near-memory system 'we consider the scenario when all accesses are done to the different vaults inside the HMC memory,' which effectively gives the NMC system perfect vault-level parallelism, zero bank conflicts, zero row-buffer misses, zero vault-controller contention, and no inter-vault data movement, coherence traffic, or kernel-offload cost. The CPU-centric side, by contrast, is charged with L1/L2 miss penalties through the varied miss rates in Figure 8. As a result, the observations that increasing miss rates degrade the multicore system relative to NMC, and that low-locality applications benefit from NMC, are partly built into the modeling assumptions rather than emerging from empirical or even well-specified quantitative analysis. Because this comparison is the only quantitative evidence for the paper's headline claim, this idealization is load-bearing.
  2. [Section VIII-D, model specification] The analytic model is not presented in a reproducible form. The text describes the performance calculation as 'the ratio between the number of memory accesses required to move the block of data to and from the I/O system and the available I/O bandwidth of each memory subsystem,' but no equations are given, no parameter table lists the access counts, bandwidths, energy values, or how vault count and link count enter the formulas, and no derivation is shown. The energy values (3.7 pJ/b for DRAM, 1.5 pJ/b for logic, 0.96 W static power) are cited to [9] and [37], but the reader cannot verify how these combine with miss rates to produce the normalized delay and energy surfaces in Figure 8. This lack of specification means the model cannot be checked, replicated, or extended by other researchers.
  3. [Section VIII-C, compiler framework] The claimed 'completely transparent' offloading of NMC kernels is not demonstrated. The compiler flow is described in words and in Figure 7, but no experimental results are presented: there are no benchmarks showing that the polyhedral pattern-matching framework detects kernels, that the generated runtime calls are correct, or what performance/energy overhead the offload incurs. Given that the compiler framework is one of the paper's stated contributions, the absence of any evaluation leaves the central claim about transparent NMC adoption unsupported, though it could be addressed by adding experiments or by clearly labeling this section as a research vision.
  4. [Section VIII-B and VIII-D, link between metrics and model] The proposed platform-agnostic metrics (memory entropy, spatial locality, data-level parallelism, basic-block-level parallelism) are not connected to the analytic model. The paper does not specify how a given characterization result, for example the Gramschmidt vs. Jacobi-1d comparison in Figure 6, translates into the miss rates or access-count parameters used in the Section VIII-D comparison. Without this mapping, the characterization metrics and the analytic model remain separate exercises, and the claim that the metrics can 'identify the kernels that can potentially benefit from NMC' is not quantitatively grounded.
minor comments (5)
  1. [Abstract and Section I] The phrase 'significantly diminish the data movement problem' is used as a definitive claim, but the supporting evidence is only the preliminary analytic model; consider softening the wording to 'may diminish' until more validation is available.
  2. [Table 2] The 'NMC Unit' column contains mixed entries (e.g., 'CPU', 'GPU', 'ACC', 'CGRA+FPGA') while the legend in Table 1 defines 'Processing Unit' types; consistency between the column headers and the legend would improve readability.
  3. [Section VII-B, Table 3] Ramulator-PIM [91] is listed with a URL instead of a formal citation, and the 'NMC capabilities' column uses subjective labels ('Yes', 'Limited', 'No') without a definition; a brief explanation of what constitutes 'Limited' would help.
  4. [Section II] In the sentence 'researches have proposed various NMC designs,' 'researches' should be 'researchers'; several other minor typos and grammatical issues (e.g., 'the above mentioned behaviour', 'offloading') appear throughout, so a careful copyedit is recommended.
  5. [Figure 2] Figure 2 is referenced in Section I but is not described in the text; adding a sentence explaining its content would make the figure self-contained.

Circularity Check

1 steps flagged · score 6.0 of 10

Section VIII-D's NMC-versus-CPU comparison assumes perfect vault-level parallelism and then reports the low-locality NMC advantage as an observation; the result is an artifact of the model's defining premises.

  1. self definitional [Section VIII-D, 'Performance and Energy Exploration' (Figure 8 discussion, pages 11-12)]
    "In case of a near-memory system, we consider the scenario when all accesses are done to the different vaults inside the HMC memory, which can exploit the inherent parallelism offered by the 3D memory. ... Based on our evaluation, we make the following three observations. First, ... if the miss rate (L1 and L2) increases the multi-core system performance degrades compared to the NMC systems."

    The NMC path is defined as servicing every access through different HMC vaults with no miss penalties, row conflicts, or offload overhead, while the CPU-centric path is charged with L1/L2 miss penalties. The paper's first and third 'observations' — higher miss rates hurt the multicore system relative to NMC, and low-locality applications favor NMC — are direct consequences of that asymmetry, not empirical findings. No equation or measurement is given to show an NMC workload incurring comparable miss or contention costs, so the NMC advantage is inserted as a premise ('all accesses ... to the different vaults') and then reported as a conclusion. The quantitative illustration therefore reduces by construction to the model's own inputs.

full rationale

The survey taxonomy in Sections II-VII is independent and self-contained: it organizes prior NMC architectures under memory-hierarchy, processing-unit, tool, interoperability, and application dimensions, and its classification table is a systematization rather than a prediction. The compiler flow in Section VIII-C is described concretely (LLVM-IR, Polly, Presburger sets, pattern matching) and depends on normal engineering citations rather than on a circular uniqueness claim. The main circularity concern is Section VIII-D: the analytic comparison defines the NMC case as 'all accesses are done to the different vaults inside the HMC memory' and then reports the resulting NMC advantage as an observation. Because the CPU side alone is charged with cache-miss penalties, the conclusion that low-locality, high-miss-rate kernels prefer NMC is true by construction. The paper's broader claim that NMC can diminish data-movement problems is supported by the surveyed literature, so the circularity is partial rather than total; however, the one quantitative demonstration offered by the authors is structurally built from its own favorable assumption.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central survey content does not depend on fitted parameters. The illustrative analytic model in Section VIII-D uses hand-chosen device parameters from prior HMC literature and sweeps cache miss rates, but it is not validated against real or simulated systems. The compiler framework assumes an accelerator library and pattern-matching capability that are not demonstrated.

free parameters (4)
  • HMC DRAM layer energy per bit = 3.7 pJ/b
    Taken from HMC references [9] and [37]; central to the energy comparison in Section VIII-D.
  • HMC logic layer energy per bit = 1.5 pJ/b
    Assumed for operations in the logic layer; used in the NMC energy model.
  • HMC logic layer static power = 0.96 W
    Assumed from Pugsley et al. [37] to model added logic; used in static energy estimates.
  • NMC core frequency = 1.2 GHz
    Chosen for the low-power cores in the logic layer (Table 4); affects execution time estimates.
assumptions (4)
  • domain assumption Data movement is a dominant bottleneck for data-intensive applications
    Motivates the whole paper (Introduction); likely true but not proven here.
  • domain assumption 3D-stacked memory with a logic layer provides sufficient bandwidth and energy efficiency for near-memory compute units
    Assumed from HMC/HBM literature, e.g., [9]-[12]; underlies the NMC promise.
  • domain assumption Near-memory accelerators expose a library of memory-bound kernels via an API that the compiler can match
    Stated in Section VIII-C; the compiler flow depends on this library existing.
  • domain assumption The HMC vault model with all accesses spread across vaults represents the NMC execution scenario
    Section VIII-D idealizes the NMC memory system; if real workloads cause vault contention, the performance/energy advantage is overstated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Near-Memory Computing: Past, Present, and Future." pith.science (2026). https://pith.science/paper/F7FTJXZK

@misc{pith2026190802640,
  author       = {Pith},
  title        = {Pith review of: Near-Memory Computing: Past, Present, and Future},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F7FTJXZK}},
  note         = {Machine review of arXiv:1908.02640}
}
read the original abstract

The conventional approach of moving data to the CPU for computation has become a significant performance bottleneck for emerging scale-out data-intensive applications due to their limited data reuse. At the same time, the advancement in 3D integration technologies has made the decade-old concept of coupling compute units close to the memory --- called near-memory computing (NMC) --- more viable. Processing right at the "home" of data can significantly diminish the data movement problem of data-intensive applications. In this paper, we survey the prior art on NMC across various dimensions (architecture, applications, tools, etc.) and identify the key challenges and open issues with future research directions. We also provide a glimpse of our approach to near-memory computing that includes i) NMC specific microarchitecture independent application characterization ii) a compiler framework to offload the NMC kernels on our target NMC platform and iii) an analytical model to evaluate the potential of NMC.

Figures

Figures reproduced from arXiv: 1908.02640 by the authors.

Figure 1
Figure 1. Classification of computing systems based on working set location, which is referred to as a working set [13]. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Processing options in the memory hierarchy high [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Micron’s Hybrid Memory Cube (HMC) [10] com [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Design space exploration highlighting application [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Overview of a system with NMC capability [83]. [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: Spatial Locality and Memory Entropy characteri [PITH_FULL_IMAGE:figures/full_fig_p011_6.png]
Figure 7
Figure 7. Figure 7: Proposed compilation flow for the system with NMC capabilities depicted in Figure 5. [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: Performance and energy comparison between multi-core and an NMC system, which is attached to a host [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

111 extracted references · 70 canonical work pages

  1. [9]

    Hybrid Memory Cube New DRAM Ar- chitecture Increases Density and Performance,

    J. Jeddeloh and B. Keeth, “Hybrid Memory Cube New DRAM Ar- chitecture Increases Density and Performance,” in 2012 Symposium on VLSI Technology (VLSIT), June 2012, pp. 87–88

  2. [37]

    NDC: Analyzing the Impact of 3D-Stacked Memory+Logic Devices on MapReduce Workloads,

    S. H. Pugsley, J. Jestes, H. Zhang, R. Balasubramonian, V . Srinivasan, A. Buyuktosunoglu, A. Davis, and F. Li, “NDC: Analyzing the Impact of 3D-Stacked Memory+Logic Devices on MapReduce Workloads,” in 2014 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) , March 2014, pp. 190–200

  3. [1]

    Hitting the Memory Wall: Implications of the Obvious,

    W. A. Wulf and S. A. McKee, “Hitting the Memory Wall: Implications of the Obvious,” SIGARCH Comput. Archit. News , vol. 23, no. 1, pp. 20–24, Mar. 1995

  4. [2]

    Design of Ion-implanted MOSFET’s with Very Small Physical Dimensions,

    R. H. Dennard, F. H. Gaensslen, V . L. Rideout, E. Bassous, and A. R. LeBlanc, “Design of Ion-implanted MOSFET’s with Very Small Physical Dimensions,” IEEE Journal of Solid-State Circuits , vol. 9, no. 5, pp. 256–268, Oct 1974

  5. [3]

    Dark Silicon and The End of Multicore Scaling,

    H. Esmaeilzadeh, E. Blem, R. S. Amant, K. Sankaralingam, and D. Burger, “Dark Silicon and The End of Multicore Scaling,” IEEE Micro, vol. 32, no. 3, pp. 122–134, 2012

  6. [4]

    Active Memory Cube: A Processing-in-Memory Architecture for Exascale Systems,

    R. Nair, S. F. Antao, C. Bertolli, P. Bose, J. R. Brunheroto, T. Chen, C. . Cher, C. H. A. Costa, J. Doi, C. Evangelinos, B. M. Fleischer, T. W. Fox, D. S. Gallo, L. Grinberg, J. A. Gunnels, A. C. Jacob, P. Jacob, H. M. Jacobson, T. Karkhanis, C. Kim, J. H. Moreno, J. K. O’Brien, M. Ohmacht, Y . Park, D. A. Prener, B. S. Rosenburg, K. D. Ryu, O. Sallenave...

  7. [5]

    An End- to-End Computing Model for the Square Kilometre Array,

    R. Jongerius, S. Wijnholds, R. Nijboer, and H. Corporaal, “An End- to-End Computing Model for the Square Kilometre Array,” Computer, vol. 47, no. 9, pp. 48–54, Sept 2014

  8. [6]

    Micro- Architectural Characterization of Apache Spark on Batch and Stream Processing Workloads,

    A. J. Awan, M. Brorsson, V . Vlassov, and E. Ayguade, “Micro- Architectural Characterization of Apache Spark on Batch and Stream Processing Workloads,” in Big Data and Cloud Computing (BD- Cloud), Social Computing and Networking (SocialCom), Sustainable Computing and Communications (SustainCom)(BDCloud-SocialCom- SustainCom), 2016 IEEE International Confe...

Show all 111 references
  1. [7]

    Performance Characterization of In-Memory Data Analytics on a Modern Cloud Server,

    ——, “Performance Characterization of In-Memory Data Analytics on a Modern Cloud Server,” in Big Data and Cloud Computing (BDCloud), 2015 IEEE Fifth International Conference on . IEEE, 2015, pp. 1–8

  2. [8]

    Node Architecture Implications for In-Memory Data Analytics on Scale-in Clusters,

    Awan, Ahsan Javed and Brorsson, Mats and Vlassov, Vladimir and Ayguade, Eduard, “Node Architecture Implications for In-Memory Data Analytics on Scale-in Clusters,” in Big Data Computing Applica- tions and Technologies (BDCAT), 2016 IEEE/ACM 3rd International Conference on. IEE...

  3. [10]

    Hybrid Memory Cube (HMC),

    J. T. Pawlowski, “Hybrid Memory Cube (HMC),” in 2011 IEEE Hot Chips 23 Symposium (HCS) , Aug 2011, pp. 1–24

  4. [11]

    25.2 A 1.2V 8Gb 8-Channel 128GB/s High-Bandwidth Memory (HBM) Stacked DRAM with Effective Microbump I/O Test Methods Using 29nm Process and TSV,

    D. U. Lee, K. W. Kim, K. W. Kim, H. Kim, J. Y . Kim, Y . J. Park, J. H. Kim, D. S. Kim, H. B. Park, J. W. Shin, J. H. Cho, K. H. Kwon, M. J. Kim, J. Lee, K. W. Park, B. Chung, and S. Hong, “25.2 A 1.2V 8Gb 8-Channel 128GB/s High-Bandwidth Memory (HBM) Stacked DRAM with Effecti...

  5. [12]

    A 1.2 V 12.8 GB/s 2 Gb Mobile Wide-I/O DRAM With 4×128 I/Os Using TSV Based Stacking,

    J. Kim, C. S. Oh, H. Lee, D. Lee, H. R. Hwang, S. Hwang, B. Na, J. Moon, J. Kim, H. Park, J. Ryu, K. Park, S. K. Kang, S. Kim, H. Kim, J. Bang, H. Cho, M. Jang, C. Han, J. LeeLee, J. S. Choi, and Y . Jun, “A 1.2 V 12.8 GB/s 2 Gb Mobile Wide-I/O DRAM With 4×128 I/Os Using TSV B...

  6. [13]

    Memristor Based Computation-in-memory Architecture for Data-Intensive Applications,

    S. Hamdioui, L. Xie, H. A. D. Nguyen, M. Taouil, K. Bertels, H. Corporaal, H. Jiao, F. Catthoor, D. Wouters, L. Eike, and J. van Lunteren, “Memristor Based Computation-in-memory Architecture for Data-Intensive Applications,” in 2015 Design, Automation Test in Europe Conference...

  7. [14]

    In-Memory Big Data Management and Processing: A Survey,

    H. Zhang, G. Chen, B. C. Ooi, K. L. Tan, and M. Zhang, “In-Memory Big Data Management and Processing: A Survey,” IEEE Transactions on Knowledge and Data Engineering , vol. 27, no. 7, pp. 1920–1948, July 2015

  8. [15]

    A Case for Intelligent Disks (IDISKs),

    K. Keeton, D. A. Patterson, and J. M. Hellerstein, “A Case for Intelligent Disks (IDISKs),” SIGMOD Rec., vol. 27, no. 3, pp. 42–52, Sep. 1998

  9. [16]

    Active Storage for Large- Scale Data Mining and Multimedia,

    E. Riedel, G. A. Gibson, and C. Faloutsos, “Active Storage for Large- Scale Data Mining and Multimedia,” in Proceedings of the 24rd International Conference on Very Large Data Bases , ser. VLDB ’98. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 1998, pp. 62–73

  10. [17]

    Evolution of Memory Architecture,

    R. Nair, “Evolution of Memory Architecture,” Proceedings of the IEEE, vol. 103, no. 8, pp. 1331–1345, Aug 2015

  11. [18]

    Data-Centric Computing Frontiers: A Survey On Processing-In-Memory,

    P. Siegl, R. Buchty, and M. Berekovic, “Data-Centric Computing Frontiers: A Survey On Processing-In-Memory,” in Proceedings of the Second International Symposium on Memory Systems . ACM, 2016, pp. 295–308

  12. [19]

    Enabling the Adoption of Processing-in-Memory: Challenges, Mech- anisms, Future Research Directions,

    S. Ghose, K. Hsieh, A. Boroumand, R. Ausavarungnirun, and O. Mutlu, “Enabling the Adoption of Processing-in-Memory: Challenges, Mech- anisms, Future Research Directions,” arXiv preprint arXiv:1802.00320, 2018

  13. [20]

    A Review of Near-Memory Computing Architectures: Opportunities and Challenges,

    G. Singh, L. Chelini, S. Corda, A. Javed Awan, S. Stuijk, R. Jor- dans, H. Corporaal, and A. Boonstra, “A Review of Near-Memory Computing Architectures: Opportunities and Challenges,” in 2018 21st Euromicro Conference on Digital System Design (DSD) , Aug 2018, pp. 608–617

  14. [21]

    A Logic-in-Memory Computer,

    H. S. Stone, “A Logic-in-Memory Computer,” IEEE Transactions on Computers, vol. C-19, no. 1, pp. 73–78, Jan 1970

  15. [22]

    Computational Ram: A Memory-SIMD Hybrid and its Application to DSP,

    D. G. Elliott, W. M. Snelgrove, and M. Stumm, “Computational Ram: A Memory-SIMD Hybrid and its Application to DSP,” in 1992 Proceedings of the IEEE Custom Integrated Circuits Conference , May 1992, pp. 30.6.1–30.6.4

  16. [23]

    EXECUBE-A New Architecture for Scaleable MPPs,

    P. M. Kogge, “EXECUBE-A New Architecture for Scaleable MPPs,” in Proceedings of the 1994 International Conference on Parallel Processing - Volume 01, ser. ICPP ’94. Washington, DC, USA: IEEE Computer Society, 1994, pp. 77–84

  17. [24]

    FBRAM: A New Form of Memory Optimized for 3D Graphics,

    M. F. Deering, S. A. Schlapp, and M. G. Lavelle, “FBRAM: A New Form of Memory Optimized for 3D Graphics,” in Proceedings of the 21st Annual Conference on Computer Graphics and Interactive Techniques, ser. SIGGRAPH ’94. New York, NY , USA: ACM, 1994, pp. 167–174

  18. [25]

    Processing in Memory: The Terasys Massively Parallel PIM Array,

    M. Gokhale, B. Holmes, and K. Iobst, “Processing in Memory: The Terasys Massively Parallel PIM Array,” Computer, vol. 28, no. 4, pp. 23–31, Apr 1995

  19. [26]

    A Case for Intelligent RAM,

    D. Patterson, T. Anderson, N. Cardwell, R. Fromm, K. Keeton, C. Kozyrakis, R. Thomas, and K. Yelick, “A Case for Intelligent RAM,” IEEE Micro, vol. 17, no. 2, pp. 34–44, Mar 1997

  20. [27]

    FlexRAM: Toward an Advanced Intelligent Memory System,

    Y . Kang, W. Huang, S.-M. Yoo, D. Keen, Z. Ge, V . Lam, P. Pat- tnaik, and J. Torrellas, “FlexRAM: Toward an Advanced Intelligent Memory System,” in Proceedings 1999 IEEE International Confer- ence on Computer Design: VLSI in Computers and Processors (Cat. No.99CB37040), 1999,...

  21. [28]

    Scalable Processors in the Billion- Transistor Era: IRAM,

    C. E. Kozyrakis, S. Perissakis, D. Patterson, T. Anderson, K. Asanovic, N. Cardwell, R. Fromm, J. Golbus, B. Gribstad, K. Keeton, R. Thomas, N. Treuhaft, and K. Yelick, “Scalable Processors in the Billion- Transistor Era: IRAM,” Computer, vol. 30, no. 9, pp. 75–78, Sep 1997

  22. [29]

    A Scalable Processing-in-Memory Accelerator for Parallel Graph Processing,

    J. Ahn, S. Hong, S. Yoo, O. Mutlu, and K. Choi, “A Scalable Processing-in-Memory Accelerator for Parallel Graph Processing,” in 2015 ACM/IEEE 42nd Annual International Symposium on Computer Architecture (ISCA), June 2015, pp. 105–117

  23. [30]

    Accelerating Pointer Chasing in 3D-Stacked Memory: Challenges, Mechanisms, Evaluation,

    K. Hsieh, S. Khan, N. Vijaykumar, K. K. Chang, A. Boroumand, S. Ghose, and O. Mutlu, “Accelerating Pointer Chasing in 3D-Stacked Memory: Challenges, Mechanisms, Evaluation,” in 2016 IEEE 34th International Conference on Computer Design (ICCD) . IEEE, 2016, pp. 25–32

  24. [31]

    Design and Evaluation of a Processing-in-Memory Architecture for the Smart Memory Cube,

    E. Azarkhish, D. Rossi, I. Loi, and L. Benini, “Design and Evaluation of a Processing-in-Memory Architecture for the Smart Memory Cube,” in Proceedings of the 29th International Conference on Architecture of Computing Systems – ARCS 2016 - Volume 9637 . New York, NY , USA: Spr...

  25. [32]

    Operand Size Reconfiguration for Big Data Processing in Memory,

    P. C. Santos, G. F. Oliveira, D. G. Tom ´e, M. A. Alves, E. C. Almeida, and L. Carro, “Operand Size Reconfiguration for Big Data Processing in Memory,” in Proceedings of the Conference on Design, Automation & Test in Europe . European Design and Automation Association, 2017, pp...

  26. [33]

    A Processing in Memory Taxonomy and a Case for Studying Fixed-Function PIM,

    G. Loh, N. Jayasena, M. Oskin, M. Nutter, D. Roberts, M. Meswani, D. Zhang, and M. Ignatowski, “A Processing in Memory Taxonomy and a Case for Studying Fixed-Function PIM,” in Workshop on Near- Data Processing (WoNDP), 2013

  27. [34]

    XSD: Accelerating MapReduce by Harnessing the GPU inside an SSD,

    B. Y . Cho, W. S. Jeong, D. Oh, and W. W. Ro, “XSD: Accelerating MapReduce by Harnessing the GPU inside an SSD,” in Proceedings of the 1st Workshop on Near-Data Processing , 2013

  28. [35]

    Enabling Cost-effective Data Processing with Smart SSD,

    Y . Kang, Y .-s. Kee, E. L. Miller, and C. Park, “Enabling Cost-effective Data Processing with Smart SSD,” in Mass Storage Systems and Technologies (MSST), 2013 IEEE 29th Symposium on . IEEE, 2013, pp. 1–12

  29. [36]

    Willow: A User-programmable SSD,

    S. Seshadri, M. Gahagan, S. Bhaskaran, T. Bunker, A. De, Y . Jin, Y . Liu, and S. Swanson, “Willow: A User-programmable SSD,” in Proceedings of the 11th USENIX Conference on Operating Systems Design and Implementation , ser. OSDI’14. Berkeley, CA, USA: USENIX Association, 2014...

  30. [38]

    TOP-PIM: Throughput-Oriented Programmable Processing in Memory,

    D. Zhang, N. Jayasena, A. Lyashevsky, J. L. Greathouse, L. Xu, and M. Ignatowski, “TOP-PIM: Throughput-Oriented Programmable Processing in Memory,” in Proceedings of the 23rd international symposium on High-performance parallel and distributed computing . ACM, 2014, pp. 85–98

  31. [39]

    Beyond the Wall: Near-Data Processing for Databases,

    S. L. Xi, O. Babarinsa, M. Athanassoulis, and S. Idreos, “Beyond the Wall: Near-Data Processing for Databases,” in Proceedings of the 11th International Workshop on Data Management on New Hardware . ACM, 2015, p. 2

  32. [40]

    Near Memory Data Structure Rearrangement,

    M. Gokhale, S. Lloyd, and C. Hajas, “Near Memory Data Structure Rearrangement,” in Proceedings of the 2015 International Symposium on Memory Systems . ACM, 2015, pp. 283–290

  33. [41]

    HRL: Efficient and Flexible Reconfigurable Logic for Near-Data Processing,

    M. Gao and C. Kozyrakis, “HRL: Efficient and Flexible Reconfigurable Logic for Near-Data Processing,” in 2016 IEEE International Sym- posium on High Performance Computer Architecture (HPCA) , March 2016, pp. 126–137

  34. [42]

    ProPRAM: Exploiting the Transparent Logic Resources in Non-V olatile Memory for Near Data Computing,

    Y . Wang, Y . Han, L. Zhang, H. Li, and X. Li, “ProPRAM: Exploiting the Transparent Logic Resources in Non-V olatile Memory for Near Data Computing,” in Proceedings of the 52nd Annual Design Automa- tion Conference. ACM, 2015, p. 47

  35. [43]

    BlueDBM: An Appliance for Big Data analytics,

    S. Jun, M. Liu, S. Lee, J. Hicks, J. Ankcorn, M. King, S. Xu, and Arvind, “BlueDBM: An Appliance for Big Data analytics,” in 2015 ACM/IEEE 42nd Annual International Symposium on Computer Architecture (ISCA), June 2015, pp. 1–13

  36. [44]

    NDA: Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules,

    A. Farmahini-Farahani, J. H. Ahn, K. Morrow, and N. S. Kim, “NDA: Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules,” in 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA), Feb 2015, pp. 283–295

  37. [45]

    PIM-Enabled Instructions: A Low-Overhead, Locality-Aware Processing-in-Memory Architecture,

    J. Ahn, S. Yoo, O. Mutlu, and K. Choi, “PIM-Enabled Instructions: A Low-Overhead, Locality-Aware Processing-in-Memory Architecture,” in Proceedings of the 42nd Annual International Symposium on Com- puter Architecture. ACM, 2015, pp. 336–348

  38. [46]

    Transparent Offloading and Mapping (TOM): Enabling Programmer-Transparent Near-Data Pro- cessing in GPU Systems,

    K. Hsieh, E. Ebrahimi, G. Kim, N. Chatterjee, M. O’Connor, N. Vi- jaykumar, O. Mutlu, and S. W. Keckler, “Transparent Offloading and Mapping (TOM): Enabling Programmer-Transparent Near-Data Pro- cessing in GPU Systems,” in ACM SIGARCH Computer Architecture News, vol. 44, no. 3....

  39. [47]

    Biscuit: A Framework for Near-data Processing of Big Data Workloads,

    B. Gu, A. S. Yoon, D.-H. Bae, I. Jo, J. Lee, J. Yoon, J.-U. Kang, M. Kwon, C. Yoon, S. Cho, J. Jeong, and D. Chang, “Biscuit: A Framework for Near-data Processing of Big Data Workloads,” SIGARCH Comput. Archit. News , vol. 44, no. 3, pp. 153–165, Jun. 2016. [Online]. Available...

  40. [48]

    Scheduling Techniques for GPU Architectures with Processing-in-Memory Capabilities,

    A. Pattnaik, X. Tang, A. Jog, O. Kayiran, A. K. Mishra, M. T. Kandemir, O. Mutlu, and C. R. Das, “Scheduling Techniques for GPU Architectures with Processing-in-Memory Capabilities,” in 2016 International Conference on Parallel Architecture and Compilation Techniques (PACT), S...

  41. [49]

    Caribou: Intelligent Distributed Storage,

    Z. Istv ´an, D. Sidler, and G. Alonso, “Caribou: Intelligent Distributed Storage,” Proceedings of the VLDB Endowment , vol. 10, no. 11, pp. 1202–1213, 2017

  42. [50]

    Sorting Big Data on Heterogeneous Near-data Processing Systems,

    E. Vermij, L. Fiorin, C. Hagleitner, and K. Bertels, “Sorting Big Data on Heterogeneous Near-data Processing Systems,” in Proceedings of the Computing Frontiers Conference , ser. CF’17. New York, NY , USA: ACM, 2017, pp. 349–354. [Online]. Available: http://doi.acm.org/10.1145...

  43. [51]

    Summarizer: Trading Communication with Computing Near Storage,

    G. Koo, K. K. Matam, T. I, H. V . K. G. Narra, J. Li, H.-W. Tseng, S. Swanson, and M. Annavaram, “Summarizer: Trading Communication with Computing Near Storage,” in Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture , ser. MICRO-50 ’17. New Yo...

  44. [52]

    The Mondrian Data Engine,

    D. L. De Oliveira, M. Paulo, A. Daglis, N. Mirzadeh, D. Ustiugov, J. Picorel Obando, B. Falsafi, B. Grot, and D. Pnevmatikatos, “The Mondrian Data Engine,” in Proceedings of the 44th International Symposium on Computer Architecture, no. EPFL-CONF-227947, 2017

  45. [53]

    Graph- PIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks,

    L. Nai, R. Hadidi, J. Sim, H. Kim, P. Kumar, and H. Kim, “Graph- PIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks,” in 2017 IEEE International Symposium on High Perfor- mance Computer Architecture (HPCA) , Feb 2017, pp. 457–468

  46. [54]

    Application-Transparent Near-Memory Processing Architecture with Memory Channel Network,

    M. Alian, S. W. Min, H. Asgharimoghaddam, A. Dhar, D. K. Wang, T. Roewer, A. McPadden, O. O’Halloran, D. Chen, J. Xiong, D. Kim, W. Hwu, and N. S. Kim, “Application-Transparent Near-Memory Processing Architecture with Memory Channel Network,” in 2018 51st Annual IEEE/ACM Inter...

  47. [55]

    Application Codesign of Near-Data Processing for Similarity Search,

    V . T. Lee, A. Mazumdar, C. C. del Mundo, A. Alaghi, L. Ceze, and M. Oskin, “Application Codesign of Near-Data Processing for Similarity Search,” in2018 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 2018, pp. 896–907

  48. [56]

    Processing- in-Memory for Energy-efficient Neural Network Training: A Hetero- geneous Approach,

    J. Liu, H. Zhao, M. A. Ogleari, D. Li, and J. Zhao, “Processing- in-Memory for Energy-efficient Neural Network Training: A Hetero- geneous Approach,” in 2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2018, pp. 655– 668

  49. [57]

    Google Workloads for Consumer Devices: Mitigating Data Movement Bottlenecks,

    A. Boroumand, S. Ghose, Y . Kim, R. Ausavarungnirun, E. Shiu, R. Thakur, D. Kim, A. Kuusela, A. Knies, P. Ranganathan, and O. Mutlu, “Google Workloads for Consumer Devices: Mitigating Data Movement Bottlenecks,” SIGPLAN Not., vol. 53, no. 2, pp. 316–331, Mar. 2018. [Online]. A...

  50. [58]

    Comp- Stor: An In-storage Computation Platform for Scalable Distributed Processing,

    M. Torabzadehkashi, S. Rezaei, V . Alves, and N. Bagherzadeh, “Comp- Stor: An In-storage Computation Platform for Scalable Distributed Processing,” in 2018 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). IEEE, 2018, pp. 1260– 1267

  51. [59]

    IBM POWER9 Opens up a New Era of Acceleration Enablement: OpenCAPI,

    J. Stuecheli, W. J. Starke, J. D. Irish, L. B. Arimilli, D. Dreps, B. Blaner, C. Wollbrink, and B. Allison, “IBM POWER9 Opens up a New Era of Acceleration Enablement: OpenCAPI,” vol. 62, no. 4/5. IBM, 2018, pp. 8–1

  52. [60]

    Scalable High Performance Main Memory System Using Phase-Change Memory Technology,

    M. K. Qureshi, V . Srinivasan, and J. A. Rivers, “Scalable High Performance Main Memory System Using Phase-Change Memory Technology,”SIGARCH Comput. Archit. News , vol. 37, no. 3, pp. 24– 33, Jun. 2009

  53. [61]

    A Novel Nonvolatile Memory with Spin Torque Transfer Magnetization Switching: Spin-Ram,

    M. Hosomi, H. Yamagishi, T. Yamamoto, K. Bessho, Y . Higo, K. Ya- mane, H. Yamada, M. Shoji, H. Hachino, C. Fukumoto, H. Nagao, and H. Kano, “A Novel Nonvolatile Memory with Spin Torque Transfer Magnetization Switching: Spin-Ram,” in IEEE InternationalElectron Devices Meeting,...

  54. [62]

    From Microprocessors to Nanostores: Rethinking Data-Centric Systems,

    P. Ranganathan, “From Microprocessors to Nanostores: Rethinking Data-Centric Systems,” Computer, vol. 44, no. 01, pp. 39–48, jan 2011

  55. [63]

    Self-sorting SSD: Producing Sorted Data Inside Active SSDs,

    L. C. Quero, Y .-S. Lee, and J.-S. Kim, “Self-sorting SSD: Producing Sorted Data Inside Active SSDs,” in Mass Storage Systems and Technologies (MSST), 2015 31st Symposium on . IEEE, 2015, pp. 1–7

  56. [64]

    Catalina: In-Storage Processing Ac- celeration for Scalable Big Data Analytics,

    M. Torabzadehkashi, S. Rezaei, A. Heydarigorji, H. Bobarshad, V . Alves, and N. Bagherzadeh, “Catalina: In-Storage Processing Ac- celeration for Scalable Big Data Analytics,” in 2019 27th Euromicro International Conference on Parallel, Distributed and Network-Based Processing ...

  57. [65]

    Query Processing on Smart SSDs,

    K. Park, Y .-S. Kee, J. M. Patel, J. Do, C. Park, and D. J. Dewitt, “Query Processing on Smart SSDs,” IEEE Data Eng. Bull. , vol. 37, no. 2, pp. 19–26, 2014

  58. [66]

    GraFBoost: Using Accelerated Flash Storage for External Graph Analytics,

    S. Jun, A. Wright, S. Zhang, S. Xu, and Arvind, “GraFBoost: Using Accelerated Flash Storage for External Graph Analytics,” in 2018 ACM/IEEE 45th Annual International Symposium on Computer Ar- chitecture (ISCA), June 2018, pp. 411–424

  59. [67]

    Performance characterization and optimization of in- memory data analytics on a scale-up server,

    A. J. Awan, “Performance characterization and optimization of in- memory data analytics on a scale-up server,” Ph.D. dissertation, KTH Royal Institute of Technology and Universitat Polit`ecnica de Catalunya, 2017

  60. [68]

    CAIRO: A Compiler- Assisted Technique for Enabling Instruction-Level Offloading of Processing-In-Memory,

    R. Hadidi, L. Nai, H. Kim, and H. Kim, “CAIRO: A Compiler- Assisted Technique for Enabling Instruction-Level Offloading of Processing-In-Memory,” ACM Trans. Archit. Code Optim. , vol. 14, no. 4, pp. 48:1–48:25, Dec. 2017. [Online]. Available: http: //doi.acm.org/10.1145/3155287

  61. [69]

    Identifying the Potential of Near Data Processing for Apache Spark,

    A. J. Awan et al. , “Identifying the Potential of Near Data Processing for Apache Spark,” in Proceedings of the International Symposium on Memory Systems, MEMSYS 2017, Alexandria, VA, USA, October 02 - 05, 2017, 2017, pp. 60–67

  62. [70]

    http://software.intel.com/en-us/node/ 544393

    Intel Vtune Amplifier XE 2013. http://software.intel.com/en-us/node/ 544393

  63. [71]

    Practical Near-Data Processing for In-Memory Analytics Frameworks,

    M. Gao, G. Ayers, and C. Kozyrakis, “Practical Near-Data Processing for In-Memory Analytics Frameworks,” in 2015 International Confer- ence on Parallel Architecture and Compilation (PACT) , Oct 2015, pp. 113–124

  64. [72]

    Data Access Optimization in a Processing-in-memory System,

    Z. Sura, A. Jacob, T. Chen, B. Rosenburg, O. Sallenave, C. Bertolli, S. Antao, J. Brunheroto, Y . Park, K. O’Brien, and R. Nair, “Data Access Optimization in a Processing-in-memory System,” in Proceedings of the 12th ACM International Conference on Computing Frontiers, ser. CF...

  65. [73]

    A Near-Memory Processor for Vector, Streaming and Bit manipulation Workloads,

    M. Wei, M. Snir, J. Torrellas, and R. B. Tremaine, “A Near-Memory Processor for Vector, Streaming and Bit manipulation Workloads,” in In The Second Watson Conference on Interaction between Architecture, Circuits, and Compilers , 2005

  66. [74]

    Near-Memory Address Trans- lation,

    J. Picorel, D. Jevdjic, and B. Falsafi, “Near-Memory Address Trans- lation,” in Parallel Architectures and Compilation Techniques (PACT), 2017 26th International Conference on . Ieee, 2017, pp. 303–317

  67. [75]

    Buffered Compares: Excavating the Hidden Parallelism Inside DRAM Architectures with Lightweight Logic,

    J. Lee, J. H. Ahn, and K. Choi, “Buffered Compares: Excavating the Hidden Parallelism Inside DRAM Architectures with Lightweight Logic,” in 2016 Design, Automation Test in Europe Conference Exhi- bition (DATE), March 2016, pp. 1243–1248

  68. [76]

    Exploring Specialized Near-memory Processing for Data Intensive Operations,

    Yitbarek, Salessawi Ferede and Yang, Tao and Das, Reetuparna and Austin, Todd, “Exploring Specialized Near-memory Processing for Data Intensive Operations,” in Proceedings of the 2016 Conference on Design, Automation & Test in Europe , ser. DATE ’16. San Jose, CA, USA: EDA Con...

  69. [77]

    Efficient Virtual Memory for Big Memory Servers,

    A. Basu, J. Gandhi, J. Chang, M. D. Hill, and M. M. Swift, “Efficient Virtual Memory for Big Memory Servers,” SIGARCH Comput. Archit. News, vol. 41, no. 3, pp. 237–248, Jun. 2013

  70. [78]

    CoNDA: Efficient Cache Coherence Support for Near-data Accelerators,

    A. Boroumand, S. Ghose, M. Patel, H. Hassan, B. Lucia, R. Ausavarungnirun, K. Hsieh, N. Hajinazar, K. T. Malladi, H. Zheng, and O. Mutlu, “CoNDA: Efficient Cache Coherence Support for Near-data Accelerators,” in Proceedings of the 46th International Symposium on Computer Archit...

  71. [79]

    Prometheus: Processing-in- Memory Heterogeneous Architecture Design from a Multi-Layer Net- work Theoretic Strategy,

    Y . Xiao, S. Nazarian, and P. Bogdan, “Prometheus: Processing-in- Memory Heterogeneous Architecture Design from a Multi-Layer Net- work Theoretic Strategy,” in2018 Design, Automation & Test in Europe Conference & Exhibition (DATE) . IEEE, 2018, pp. 1387–1392

  72. [80]

    ISA-Independent Workload Character- ization and its Implications for Specialized Architectures,

    Y . S. Shao and D. Brooks, “ISA-Independent Workload Character- ization and its Implications for Specialized Architectures,” in 2013 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), April 2013, pp. 245–255

  73. [81]

    Microarchitecture-Independent Workload Characterization,

    K. Hoste and L. Eeckhout, “Microarchitecture-Independent Workload Characterization,” IEEE Micro, vol. 27, no. 3, pp. 63–72, May 2007

  74. [82]

    Generic Pro- cessing in Memory Cycle Accurate Simulator under Hybrid Memory Cube Architecture,

    G. F. Oliveira, P. C. Santos, M. A. Alves, and L. Carro, “Generic Pro- cessing in Memory Cycle Accurate Simulator under Hybrid Memory Cube Architecture,” in 2017 International Conference on Embedded Computer Systems: Architectures, Modeling, and Simulation (SAMOS). IEEE, 2017,...

  75. [83]

    NAPEL: Near-Memory Computing Application Performance Prediction via Ensemble Learning,

    G. Singh, J. G ´omez-Luna, G. Mariani, G. F. Oliveira, S. Corda, S. Stuijk, O. Mutlu, and H. Corporaal, “NAPEL: Near-Memory Computing Application Performance Prediction via Ensemble Learning,” in Proceedings of the 56th Annual Design Automation Conference 2019 , ser. DAC ’19. ...

  76. [84]

    Analytic Multi-Core Processor Model for Fast Design- Space Exploration,

    R. Jongerius, A. Anghel, G. Dittmann, G. Mariani, E. Vermij, and H. Corporaal, “Analytic Multi-Core Processor Model for Fast Design- Space Exploration,” IEEE Transactions on Computers , vol. 67, no. 6, pp. 755–770, 2017

  77. [85]

    Sort vs. Hash Join Revisited for Near-Memory Execution,

    N. Mirzadeh, Y . O. Koc ¸berber, B. Falsafi, and B. Grot, “Sort vs. Hash Join Revisited for Near-Memory Execution,” in 5th Workshop on Architectures and Systems for Big Data (ASBD 2015) , no. EPFL- CONF-209121, 2015

  78. [86]

    Scaling Deep Learning on Multiple In-Memory Processors,

    L. Xu, D. P. Zhang, and N. Jayasena, “Scaling Deep Learning on Multiple In-Memory Processors,” in Proceedings of the 3rd Workshop on Near-Data Processing , 2015

  79. [87]

    Design Space Exploration for PIM Architectures in 3D-stacked Memories,

    J. a. P. C. de Lima, P. C. Santos, M. A. Z. Alves, A. C. S. Beck, and L. Carro, “Design Space Exploration for PIM Architectures in 3D-stacked Memories,” in Proceedings of the 15th ACM International Conference on Computing Frontiers , ser. CF ’18. New York, NY , USA: ACM, 2018,...

  80. [88]

    SiNUCA: A Validated Micro-Architecture Simulator,

    M. A. Z. Alves, C. Villavieja, M. Diener, F. B. Moreira, and P. O. A. Navaux, “SiNUCA: A Validated Micro-Architecture Simulator,” in 2015 IEEE 17th International Conference on High Performance Com- puting and Communications, 2015 IEEE 7th International Symposium on Cyberspace ...

  81. [89]

    HMC-Sim-2.0: A Simulation Platform for Exploring Custom Memory Cube Operations,

    J. D. Leidel and Y . Chen, “HMC-Sim-2.0: A Simulation Platform for Exploring Custom Memory Cube Operations,” in 2016 IEEE Inter- national Parallel and Distributed Processing Symposium Workshops (IPDPSW), May 2016, pp. 621–630

  82. [90]

    CasHMC: A Cycle-Accurate Simulator for Hybrid Memory Cube,

    D. I. Jeon and K. S. Chung, “CasHMC: A Cycle-Accurate Simulator for Hybrid Memory Cube,” IEEE Computer Architecture Letters , vol. 16, no. 1, pp. 10–13, Jan 2017

  83. [91]

    Ramulator for Processing-in-Memory,

    SAFARI Research Group, “Ramulator for Processing-in-Memory,” https://github.com/CMU-SAFARI/ramulator-pim/

  84. [92]

    Near-Data Processing for Machine Learning,

    H. Choe, S. Lee, H. Nam, S. Park, S. Kim, E.-Y . Chung, and S. Yoon, “Near-Data Processing for Machine Learning,” CoRR, vol. abs/1610.02273, 2016

  85. [93]

    Data Mining in Intelligent SSD: Simulation-Based Evaluation,

    Y .-Y . Jo, S.-W. Kim, M. Chung, and H. Oh, “Data Mining in Intelligent SSD: Simulation-Based Evaluation,” in 2016 International Conference on Big Data and Smart Computing (BigComp) . IEEE, 2016, pp. 123–128

  86. [94]

    The Gem5 Simulator,

    N. Binkert, B. Beckmann, G. Black, S. K. Reinhardt, A. Saidi, A. Basu, J. Hestness, D. R. Hower, T. Krishna, S. Sardashti, R. Sen, K. Sewell, M. Shoaib, N. Vaish, M. D. Hill, and D. A. Wood, “The Gem5 Simulator,” SIGARCH Comput. Archit. News, vol. 39, no. 2, pp. 1–7, Aug. 2011...

  87. [95]

    A Limits Study of Benefits from Nanostore-Based Future Data- Centric System Architectures,

    J. Chang, P. Ranganathan, T. Mudge, D. Roberts, M. A. Shah, and K. T. Lim, “A Limits Study of Benefits from Nanostore-Based Future Data- Centric System Architectures,” in Proceedings of the 9th conference on Computing Frontiers. ACM, 2012, pp. 33–42

  88. [96]

    COTSon: Infrastructure for Full System Simulation,

    E. Argollo, A. Falc ´on, P. Faraboschi, M. Monchiero, and D. Ortega, “COTSon: Infrastructure for Full System Simulation,” SIGOPS Oper. Syst. Rev., vol. 43, no. 1, pp. 52–61, 2009

  89. [97]

    Notary: Hardware techniques to enhance signatures,

    L. Yen, S. C. Draper, and M. D. Hill, “Notary: Hardware techniques to enhance signatures,” in Proceedings of the 41st annual IEEE/ACM International Symposium on Microarchitecture . IEEE Computer Society, 2008, pp. 234–245

  90. [98]

    A Component Model of Spatial Locality,

    X. Gu, I. Christopher, T. Bai, C. Zhang, and C. Ding, “A Component Model of Spatial Locality,” in Proceedings of the 2009 International Symposium on Memory Management , ser. ISMM ’09. New York, NY , USA: ACM, 2009, pp. 99–108. [Online]. Available: http://doi.acm.org/10.1145/15...

  91. [99]

    An Instrumentation Approach for Hardware-agnostic Software Char- acterization,

    A. Anghel, L. M. Vasilescu, G. Mariani, R. Jongerius, and G. Dittmann, “An Instrumentation Approach for Hardware-agnostic Software Char- acterization,” International Journal of Parallel Programming , vol. 44, no. 5, pp. 924–948, 2016

  92. [100]

    Polybench: The Polyhedral Benchmark Suite,

    L.-N. Pouchet, “Polybench: The Polyhedral Benchmark Suite,” URL: http://www. cs. ucla. edu/pouchet/software/polybench , 2012

  93. [101]

    Memory and Parallelism Analysis Using a Platform-Independent Approach,

    S. Corda, G. Singh, A. J. Awan, R. Jordans, and H. Corporaal, “Memory and Parallelism Analysis Using a Platform-Independent Approach,” in Proceedings of the 22nd International Workshop on Software and Compilers for Embedded Systems , ser. SCOPES ’19. New York, NY , USA: ACM, 2...

  94. [102]

    Platform Independent Software Analysis for Near Memory Computing,

    ——, “Platform Independent Software Analysis for Near Memory Computing,” in”Proceedings of 22nd Euromicro Conference on Digital System Design, DSD” , 2019

  95. [103]

    Declarative Transformations in the Polyhedral Model,

    O. Zinenko, L. Chelini, and T. Grosser, “Declarative Transformations in the Polyhedral Model,” Inria ; ENS Paris - Ecole Normale Sup ´erieure de Paris ; ETH Zurich ; TU Delft ; IBM Z ¨urich, Research Report RR- 9243, Dec. 2018. [Online]. Available: https://hal.inria.fr/hal-01965599

  96. [104]

    Polly-Performing Poly- hedral Optimizations on a Low-Level Intermediate Representation,

    T. Grosser, A. Groesslinger, and C. Lengauer, “Polly-Performing Poly- hedral Optimizations on a Low-Level Intermediate Representation,” Parallel Processing Letters, vol. 22, no. 04, p. 1250010, 2012

  97. [105]

    Data Reorganization in Mem- ory Using 3D-Stacked DRAM,

    B. Akin, F. Franchetti, and J. C. Hoe, “Data Reorganization in Mem- ory Using 3D-Stacked DRAM,” in Proceedings of the 42nd Annual International Symposium on Computer Architecture . ACM, 2015, pp. 131–143

  98. [106]

    Enabling Portable Energy Efficiency with Memory Accelerated Library,

    Q. Guo, T. Low, N. Alachiotis, B. Akin, L. Pileggi, J. C. Hoe, and F. Franchetti, “Enabling Portable Energy Efficiency with Memory Accelerated Library,” in 2015 48th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , Dec 2015, pp. 750–761

  99. [107]

    isl: An Integer Set Library for the Polyhedral Model,

    S. Verdoolaege, “isl: An Integer Set Library for the Polyhedral Model,” in Mathematical Software – ICMS 2010 , K. Fukuda, J. v. d. Hoeven, M. Joswig, and N. Takayama, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2010, pp. 299–302

  100. [108]

    A Study on Non-V olatile 3D Stacked Memory for Big Data Applications,

    C. Qian, L. Huang, P. Xie, N. Xiao, and Z. Wang, “A Study on Non-V olatile 3D Stacked Memory for Big Data Applications,” in International Conference on Algorithms and Architectures for Parallel Processing. Springer, 2015, pp. 103–118

  101. [109]

    Gen-Z DRAM and Persistent Memory Theory of Operation,

    M. Krause and M. Witkowski, “Gen-Z DRAM and Persistent Memory Theory of Operation,” Gen-Z Consortium White Paper , 2019. [Online]. Available: http://genzconsortium.org/wp-content/uploads/ 2019/03/Gen-Z-DRAM-PM-Theory-of-Operation-WP.pdf

  102. [110]

    Compute Express Link,

    D. D. Sharma, “Compute Express Link,” CXL Consortium White Paper. [Online]. Available: https://docs.wixstatic.com/ugd/0c1418 d9878707bbb7427786b70c3c91d5fbd1.pdf

  103. [111]

    An Introduction to CCIX White Paper,

    “An Introduction to CCIX White Paper,” CCIX Consortium Inc. [Online]. Available: https://docs.wixstatic.com/ugd/0c1418 c6d7ec2210ae47f99f58042df0006c3d.pdf

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.