Pith. sign in

REVIEW 3 major objections 4 minor 68 references

ATiM: Autotuning Tensor Programs for Processing-in-DRAM

T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper presents ATiM as the first fully automated, autotuning-driven tensor compiler for DRAM-PIM, reporting speedups over hand-tuned libraries of up to 6.18x on benchmark kernels and 8.21x on GPT-J layers.

desk verdict A solid, reproducible tensor compiler for UPMEM with a genuinely new joint host/kernel autotuning space; the headline speedups rest partly on author-written baselines and lack variance stats, so treat them as indicative rather than exact. read the letter →

arxiv 2412.19630 v3 pith:NUBJ4DSS submitted 2024-12-27 cs.AR

classification cs.AR
keywords processing-in-DRAMUPMEMtensorcompilerautotuningIRscheduleprimitivesboundarycheckeliminationGPT-J
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

ATiM is presented as the first tensor compiler whose code generation for a commercial processing-in-DRAM system is fully automated and driven by autotuning rather than hand-written libraries. The paper claims that high-level tensor programs—vector arithmetic, reductions, matrix–vector and tensor–vector products, and GPT-J's fully-connected and multi-head attention layers—can be lowered automatically into optimized UPMEM host and kernel code. The enabling mechanism is a joint search over host-side data distribution and per-DPU kernel loop structure, expressed with ordinary tensor-scheduling primitives and filtered by UPMEM-specific constraint checks, plus three tensor-IR passes that strip boundary-check branches from DPU kernels. On UPMEM hardware the compiled kernels are reported to run up to 6.18x faster than hand-tuned PrIM code for benchmark kernels and up to 8.21x faster for GPT-J layers, with the largest gains from tiling the reduction dimension to engage otherwise idle DPUs. If the claims hold, the dominant software obstacle to processing-in-DRAM—expert, per-workload hand tuning—would be removed.

What carries the argument

Two mechanisms carry the argument. First, the joint schedule-primitive search space: TVM primitives (split, reorder, bind, rfactor, cache_read/write, compute_at) are reused so a single evolutionary autotuner can vary host-to-DPU tiling factors and data order, choose whether to apply hierarchical reduction across DPUs, set tasklet-level kernel tiling and WRAM caching placement, and parallelize host post-reduction loops; an early code verifier filters candidates that violate UPMEM constraints such as the 64 KB WRAM limit, and balanced sampling plus an adaptive epsilon-greedy schedule keep non-rfactor candidates from being dropped prematurely. Second, the PIM-aware TIR transformation trio: DMA-aware boundary-check elimination deletes guards on WRAM load/store loops so they become vectorizable DMA transfers, loop-bound tightening intersects boundary inequalities with loop bounds to shrink loop extents, and invariant branch hoisting combines loop unswitching with partial dead code elimination to lift invariant branches and their DMA dependencies out of the loop nest. The passes are safe because they target lowered loop-based TIR with affine accesses and structural properties guaranteed by the lowering process itself.

What would settle it

Run the canonical PrIM implementations of GEVA, TTV, MMTV, and GEMV, taken from the official PrIM release, on the same 2048-DPU UPMEM server with the same grid-searched DPU counts as the paper, and time them against the ATiM-compiled binaries; if the canonical baselines erase most of the 6.18x gap, the autotuned-generator claim collapses. A complementary check: restrict ATiM's search to the PrIM+search parameters (DPU count, tasklet count, caching tile size) and confirm that performance falls back to roughly PrIM+search levels, which would validate that the joint search space is genuinely responsible for the gains.

Watch

Extended reading notes

Core claim

Working within the TVM tensor-compiler stack, ATiM repurposes schedule primitives so one autotuner controls both halves of a UPMEM program: host code that tiles tensors across DPUs, chooses the reduction strategy (partial on-DPU reduction via rfactor with a final host reduction), and post-processes results, and kernel code that binds loops to tasklets, applies multi-level tiling, and decides MRAM-to-WRAM caching location and size. The paper's central contention is that this joint space is what makes automated optimization work, because inter-DPU data distribution and intra-DPU kernel structure are tightly coupled; prior 1D spatial tiling leaves DPUs idle and misses the longer kernels that make 2D tiling with hierarchical reduction up to 8.84x faster. A second contribution is a set of PIM-aware passes—DMA-aware boundary-check elimination, loop-bound tightening, and invariant branch hoisting—that exploit the structural guarantees of lowered tensor IR to remove the branch penalties that stall UPMEM's in-order DPU cores. ATiM is claimed to be the first tensor compiler to provide fully automated, autotuning-integrated code generation for a DRAM-PIM system, with average speedups of 2.49x, 1.85x, and 2.86x over PrIM, PrIM with searched parameters, and SimplePIM, and up to 23.3x over autotuned CPU code for larger tensors.

Load-bearing premise

The reported speedups assume that the authors' self-written PrIM-style kernels for GEVA, TTV, MMTV, and GEMV perform as well as the canonical hand-tuned PrIM implementations they stand in for, since those rewrites are the baseline every gain is measured against.

Editorial extensions

If this is right

  • High-level tensor programs can be compiled into UPMEM host and kernel binaries without handwritten PIM code, replacing per-workload tuning with a one-time automated search.
  • Tiling the reduction dimension alongside the spatial dimension activates DPUs that one-dimensional distribution leaves idle and enables hierarchical reduction, cutting kernel time by up to 8.84x and host-to-DPU transfer time by up to 8.37x for matrix–vector kernels.
  • Boundary checks that barely matter on CPUs and GPUs become decisive on UPMEM's in-order DPUs, where ATiM's branch-removal passes deliver up to 20.5% speedups over hand-tuned code.
  • Compiled GPT-J fully-connected and multi-head attention layers outperform autotuned CPU code by up to 7.65x and 6.07x, with the advantage growing for larger tensors as data movement dominates.
  • Combining balanced sampling with adaptive epsilon-greedy search converges to results up to 21.2% better than the stock evolutionary search after 1,000 trials.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The boundary-check passes rely on structural properties of lowered tensor IR rather than UPMEM-only features, so they should transfer to other in-order processors with DMA engines (such as the RISC-V edge cores the paper names); that portability is plausible but not demonstrated by the experiments.
  • Because the GEVA, TTV, MMTV, and GEMV baselines were written by the authors 'PrIM-style' rather than taken from the canonical library, a head-to-head rerun against the official PrIM kernels would be the cleanest check of whether the reported speedups come from ATiM's search or from weaker baselines.
  • The search budget matters: ATiM's advantage over the simpler PrIM+search grid only emerges after the balanced exploration phase (roughly the first 400 trials), so users with very small autotuning budgets might not realize the reported gains.
  • If the joint-search thesis is correct, near-memory architectures with programmable per-bank units generally will need coupled data-mapping and kernel-loop autotuning; the paper's preliminary HBM-PIM simulator results suggest the machinery extends, but that extension is not part of the main evaluation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. ATiM extends the Apache TVM/TensorIR compiler stack to target UPMEM DRAM-PIM systems. It defines a joint autotuning search space that covers host-side data distribution across DPUs, kernel-side multi-level tiling and tasklet binding, WRAM caching decisions, and reduction strategies, and it adds TIR lowering passes that generate host and DPU code, including data-transfer intrinsics and hierarchical reduction. The paper also contributes three PIM-aware optimization passes (DMA-aware boundary-check elimination, loop-bound tightening, and invariant branch hoisting) and two search-control techniques (balanced sampling and adaptive epsilon-greedy) to address the expanded search space. Experimental results on a 2048-DPU UPMEM server report speedups of up to 6.18x over PrIM-style baselines for benchmark kernels and up to 8.21x for GPT-J MTV layers, with an artifact that includes source code, Docker image, and scripts for reproducing Figures 9, 10, and 12.

Significance. If the reported results hold, ATiM is a meaningful step toward a practical software stack for commercial DRAM-PIM: it demonstrates that a mainstream tensor compiler can be extended to generate optimized host and kernel code for UPMEM, reducing the need for hand-tuned libraries. The artifact is a genuine strength: the code is publicly available, the evaluation workflow is scripted, and pre-autotuned modules are provided, which lowers the barrier for independent verification. The claims are not circular: the speedups are measured against external baselines, and the cost model is trained on measured candidates rather than defining the outcome. The main weakness is that four of the seven benchmark-kernel baselines are author-written PrIM-style implementations, and the performance numbers are point estimates without repeated-run statistics, so the quantitative support for the headline speedups is weaker than the qualitative claims suggest.

major comments (3)
  1. [Section 6, Experimental setup] For GEVA, TTV, MMTV, and GEMV, the baseline is not the original PrIM library but kernels written by the authors 'based on PrIM's codes and programming methodology.' The manuscript does not demonstrate that these reimplementations match canonical PrIM in caching tile sizes, tasklet mapping, DMA patterns, or host-side reduction behavior. Since the benchmark-kernel speedup up to 6.18x is reported for GEMV (Section 7.1) and the GPT-J layer comparisons rely on MTV and MMTV, an under-tuned PrIM-style comparator could inflate the headline numbers. Please validate the baselines against the original PrIM implementations where available, document any code-level differences, or re-state the claims with explicit qualification.
  2. [Section 7.1 / Figures 9 and 10] All reported speedups are point estimates ('up to ...') without confidence intervals or repeated-run statistics, and Section 8 with Figure 15 explicitly acknowledges run-to-run variability on UPMEM. Because the strongest claims are maxima, the reported numbers may be favorable outliers. Please report median (and ideally min/max or confidence intervals) over multiple repetitions for at least the headline workloads and sizes, and state how many runs were used for each reported point.
  3. [Section 8 / Figure 15] The autotuning loop selects the best candidate from noisy hardware measurements, but the final evaluation protocol is not described. If the reported speedups come from the same measured trials used to select the best candidate, the results may overstate the expected performance of the selected schedule. Please specify how many repeated evaluations were performed per candidate, how the best candidate was chosen, and whether the final numbers were re-measured on fresh runs.
minor comments (4)
  1. [Section 7.4] The text says 'as detailed in Section 7.4' when referring to the balanced sampling and adaptive epsilon-greedy strategies; this should refer to Section 5.2.3.
  2. [References] References [12] and [13] are the same paper (Devaux, Hot Chips 2019) and should be merged into a single reference.
  3. [Section 3, Figure 4] The caption reports 'up to 23.7% speedup' while the text reports 'an average 20% runtime reduction'; please clarify whether 23.7% is the maximum over all configurations shown and state the corresponding workload explicitly.
  4. [Section 6, Table 3] For the PrIM and PrIM+search columns, DPU counts are grid-searched but tasklet and caching tile sizes are only listed for PrIM+search; adding the selected tasklet counts for PrIM and PrIM(E) would make the comparison easier to audit.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: ATiM's performance claims are measured on hardware against external or clearly labeled baselines, and the paper's only self-citation (PIMFlow) is used as related-work contrast, not as load-bearing evidence.

full rationale

The paper's central claim is a compiler construction plus a measured evaluation. The autotuner uses a cost model trained on measured candidates, but the reported speedups are obtained by executing the final compiled kernels on UPMEM hardware; no result is defined by the cost model or by a fitted parameter renamed as a prediction. The schedule-primitive lowering, joint host-kernel search, and boundary-check elimination passes are justified by the paper's own TIR structural arguments and by external prior work (TVM, TensorIR, DietCode), not by an unverified theorem imported from the authors' prior publications. The single self-citation, PIMFlow [55], appears only in a related-work comparison and is not used to justify any ATiM design decision or quantitative claim. The one legitimate concern is methodological rather than circular: in Section 6, the authors state that they 'wrote GEVA, TTV, MMTV, and GEMV based on PrIM's codes and programming methodology (PrIM-style)', so the reported speedups against 'PrIM' for those four kernels depend on how faithfully the author-written baselines reproduce the original PrIM implementations. This is a baseline-fidelity and external-validity concern, not a case where a prediction is equivalent to an input by construction. Section 8 also candidly discusses autotuning overheads and UPMEM timing variability, which further supports the interpretation that the evaluation is measured rather than self-referential. No equation, fitted value, or cited prior result is reduced to its own input; therefore the circularity score is 0.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

No hidden free parameters: the autotuned schedule variables (reported in Table 3 for each workload) are the product of measurement-based evolutionary search on the target UPMEM hardware, not ex-ante assumptions, hand-chosen fits, or inputs to a derivation. The paper postulates no new entities. The central claims rest on domain assumptions about TIR lowering guarantees, UPMEM hardware behavior, the self-written PrIM-style baselines, and the uPIMulator simulator, listed below.

free parameters (1)
  • ATiM autotuned schedule (DPU counts per dimension, tasklet counts, caching tile sizes, tiling factors) = Workload-dependent; examples in Table 3 (e.g., MTV 4MB: DPUs/dim (16,16), 16 tasklets, caching (32,32,4))
    These are the outputs of the measurement-based evolutionary search, not hand-chosen inputs or model-fit parameters; they are empirically selected per workload on UPMEM hardware. The reported speedups are measured after search, so no hidden degrees of freedom are used to manufacture the result.
assumptions (4)
  • domain assumption TIR semantics for compute_at and reverse_compute_at guarantee that all consumer operations of a loop stay under the loop's boundary condition.
    Invoked in Section 5.3 to justify aggressive boundary-check elimination and branch hoisting; the guarantee comes from TensorIR [16].
  • domain assumption Local MRAM padding in DPU memory makes overfetched DMA transfers safe (no corruption of meaningful data), and the remaining boundary checks in compute and host readout preserve correctness.
    Stated in Section 5.3.1 as the correctness argument for DMA-aware boundary check elimination.
  • domain assumption PrIM-style kernels written by the authors for GEVA, TTV, MMTV, and GEMV follow PrIM's programming methodology closely enough to represent expert-tuned baseline performance.
    Stated in Section 6; load-bearing for the claimed speedups over PrIM.
  • domain assumption uPIMulator simulation (commit 870d916) accurately models DPU instruction timing and stalls.
    Used in Section 7.3 for the instruction breakdown in Fig. 13.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ATiM: Autotuning Tensor Programs for Processing-in-DRAM." pith.science (2026). https://pith.science/paper/NUBJ4DSS

@misc{pith2026241219630,
  author       = {Pith},
  title        = {Pith review of: ATiM: Autotuning Tensor Programs for Processing-in-DRAM},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NUBJ4DSS}},
  note         = {Machine review of arXiv:2412.19630}
}
abstract

Processing-in-DRAM (DRAM-PIM) has emerged as a promising technology for accelerating memory-intensive operations in modern applications, such as Large Language Models (LLMs). Despite its potential, current software stacks for DRAM-PIM face significant challenges, including reliance on hand-tuned libraries that hinder programmability, limited support for high-level abstractions, and the lack of systematic optimization frameworks. To address these limitations, we present ATiM, a search-based optimizing tensor compiler for UPMEM. Key features of ATiM include: (1) automated searches of the joint search space for host and kernel tensor programs, (2) PIM-aware optimizations for efficiently handling boundary conditions, and (3) improved search algorithms for the expanded search space of UPMEM systems. Our experimental results on UPMEM hardware demonstrate performance gains of up to 6.18$\times$ for various UPMEM benchmark kernels and 8.21$\times$ for GPT-J layers. To the best of our knowledge, ATiM is the first tensor compiler to provide fully automated, autotuning-integrated code generation support for a DRAM-PIM system. By bridging the gap between high-level tensor computation abstractions and low-level hardware-specific requirements, ATiM establishes a foundation for advancing DRAM-PIM programmability and enabling streamlined optimization.

Figures

Figures reproduced from arXiv: 2412.19630 by the authors.

Figure 1
Figure 1. UPMEM architecture [12]. high-level semantic knowledge and UPMEM’s architectural features for significant performance gains. • We enhance the TVM autotuning framework to handle the expanded and more complex search space for UPMEM. ATiM refines the search mechanism to minimize sampling noises and biases early in the autotuning process, improving the final result quality. In the rest of the paper, we provide backgroun… view at source ↗
Figure 2
Figure 2. TIR transformations according to schedule primitives. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Performance impact of caching tile sizes, tiling schemes, and the [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Boundary checks’ impact on GEMV kernel (M [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: The overall system overview of ATiM. existing TIR lowering passes in TVM to translate schedule primitives to represent correct tensor programs for UPMEM host and kernel operations. ATiM implements address calculation logic for kernel memory accesses to tiled data per b…
Figure 6
Figure 6. Figure 6: Autotuning-driven code generation process in ATiM. [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: Data transfer code generation and optimization. [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]
Figure 8
Figure 8. Figure 8: Example of PIM-aware optimizations applied to a 7 [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 9
Figure 9. Figure 9: Tensor operation performance (normalized to PrIM). [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 10
Figure 10. Figure 10: Performance of FC (MTV) and MMTV operations of the MHA layers of GPT-J 6B and 30B (normalized to PrIM). [PITH_FULL_IMAGE:figures/full_fig_p011_10.png]
Figure 12
Figure 12. Figure 12: Kernel performance with PIM-aware optimizations in Section 5.3 applied. (a) and (b) are misaligned on one axis, (c) on both axes, and (d) misaligned VA. [PITH_FULL_IMAGE:figures/full_fig_p012_12.png]
Figure 14
Figure 14. Figure 14: Autotuning efficiency of balanced sampling and adaptive epsilon [PITH_FULL_IMAGE:figures/full_fig_p012_14.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

68 extracted references · 36 canonical work pages

  1. [1]

    2016.{TensorFlow}: a system for{Large-Scale} machine learning

    Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. 2016.{TensorFlow}: a system for{Large-Scale} machine learning. In12th USENIX symposium on operating systems design and implementation (OSDI 16) . 265–283

  2. [2]

    Krste Asanovic, Rimas Avizienis, Jonathan Bachrach, Scott Beamer, David Bian- colin, Christopher Celio, Henry Cook, Daniel Dabbelt, John Hauser, Adam Izraele- vitz, et al . 2016. The rocket chip generator. EECS Department, University of California, Berkeley, Tech. Rep. UCB/EECS-2016-17 4 (2016), 6–2

  3. [3]

    Junjie Bai, Fang Lu, Ke Zhang, et al. 2019. ONNX: Open Neural Network Exchange. https://github.com/onnx/onnx

  4. [4]

    Arthur Bernhardt, Andreas Koch, and Ilia Petrov. 2023. pimDB: From Main- Memory DBMS to Processing-In-Memory DBMS-Engines on Intelligent Memories. In Proceedings of the 19th International Workshop on Data Management on New Hardware (Seattle, WA, USA)(DaMoN ’23). Association for Computing Machinery, New York, NY, USA, 44–52. https://doi.org/10.1145/3592980.3595312

  5. [5]

    J. Chen, J. Gomez-Luna, I. El Hajj, Y. Guo, and O. Mutlu. 2023. SimplePIM: A Software Framework for Productive and Efficient Processing-in-Memory. In 2023 32nd International Conference on Parallel Architectures and Compilation Techniques (PACT). IEEE Computer Society, Los Alamitos, CA, USA, 99–111. https://doi.org/ 10.1109/PACT58117.2023.00017

  6. [6]

    Liang-Chi Chen, Chien-Chung Ho, and Yuan-Hao Chang. 2023. UpPipe: A Novel Pipeline Management on In-Memory Processors for RNA-seq Quantification. In 2023 60th ACM/IEEE Design Automation Conference (DAC) . 1–6. https://doi.org/ 10.1109/DAC56929.2023.10247915

  7. [7]

    Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). USENIX Association, Carlsb...

  8. [8]

    Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. Learning to Optimize Tensor Programs. In Advances in Neural Information Processing Systems , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.), Vol. 31. Curran Associates, Inc. https://proc...

Show all 68 references
  1. [9]

    Cho, Jeageun Jung, and Mattan Erez

    Benjamin Y. Cho, Jeageun Jung, and Mattan Erez. 2021. Accelerating bandwidth- bound deep learning inference with main-memory accelerators. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (St. Louis, Missouri) (SC ...

  2. [10]

    Lucian Codrescu, Willie Anderson, Suresh Venkumanhanti, Mao Zeng, Erich Plondke, Chris Koob, Ajay Ingle, Charles Tabony, and Rick Maule. 2014. Hexagon DSP: An architecture optimized for mobile multimedia and communications.IEEE Micro 34, 2 (2014), 34–43

  3. [11]

    Prangon Das, Purab Ranjan Sutradhar, Mark Indovina, Sai Manoj Pudukotai Dinakarrao, and Amlan Ganguly. 2022. Implementation and Evaluation of Deep Neural Networks in Commercially Available Processing in Memory Hardware. In 2022 IEEE 35th International System-on-Chip Conference...

  4. [13]

    Fabrice Devaux. 2019. The true Processing In Memory accelerator. In 2019 IEEE Hot Chips 31 Symposium (HCS) . 1–24. https://doi.org/10.1109/HOTCHIPS.2019. 8875680

  5. [14]

    Jeff Draper, Jacqueline Chame, Mary Hall, Craig Steele, Tim Barrett, Jeff LaCoss, John Granacki, Jaewook Shin, Chun Chen, Chang Woo Kang, Ihn Kim, and Gokhan Daglikoca. 2002. The architecture of the DIVA processing-in-memory chip. In Proceedings of the 16th International Confe...

  6. [15]

    Andi Drebes, Lorenzo Chelini, Oleksandr Zinenko, Albert Cohen, Henk Cor- poraal, Tobias Grosser, Kanishkan Vadivel, and Nicolas Vasilache. 2020. TC- CIM: Empowering Tensor Comprehensions for Computing-In-Memory. http: //impact.gforge.inria.fr/impact2020/ 10th International Wor...

  7. [16]

    Siyuan Feng, Bohan Hou, Hongyi Jin, Wuwei Lin, Junru Shao, Ruihang Lai, Zihao Ye, Lianmin Zheng, Cody Hao Yu, Yong Yu, and Tianqi Chen. 2023. TensorIR: An Abstraction for Automatic Tensorized Program Optimization. In Proceedings of the 28th ACM International Conference on Arch...

  8. [17]

    Christina Giannoula, Ivan Fernandez, Juan Gómez-Luna, Nectarios Koziris, Geor- gios Goumas, and Onur Mutlu. 2022. SparseP: Efficient Sparse Matrix Vector Multiplication on Real Processing-In-Memory Architectures. In 2022 IEEE Com- puter Society Annual Symposium on VLSI (ISVLSI...

  9. [18]

    Kailash Gogineni, Sai Santosh Dayapule, Juan Gómez-Luna, Karthikeya Gogi- neni, Peng Wei, Tian Lan, Mohammad Sadrosadati, Onur Mutlu, and Guru Venkataramani. 2024. SwiftRL: Towards Efficient Reinforcement Learning on Real Processing-In-Memory Systems. In 2024 IEEE Internationa...

  10. [19]

    Gokhale, B

    M. Gokhale, B. Holmes, and K. Iobst. 1995. Processing in memory: the Terasys massively parallel PIM array. Computer 28, 4 (1995), 23–31. https://doi.org/10. 1109/2.375174

  11. [20]

    Tobias Grosser, Armin Groesslinger, and Christian Lengauer. 2012. Polly—performing polyhedral optimizations on a low-level intermediate representation. Parallel Processing Letters 22, 04 (2012), 1250010

  12. [21]

    Peng Gu, Xinfeng Xie, Yufei Ding, Guoyang Chen, Weifeng Zhang, Dimin Niu, and Yuan Xie. 2020. iPIM: Programmable In-Memory Image Processing Accelera- tor Using Near-Bank Architecture. In 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA) . 804–81...

  13. [22]

    Oliveira, Gagandeep Singh, and Onur Mutlu

    Juan Gómez-Luna, Yuxin Guo, Sylvan Brocard, Julien Legriel, Remy Cimadomo, Geraldo F. Oliveira, Gagandeep Singh, and Onur Mutlu. 2023. Evaluating Machine LearningWorkloads on Memory-Centric Computing Systems. In 2023 IEEE Inter- national Symposium on Performance Analysis of Sy...

  14. [23]

    Oliveira, and Onur Mutlu

    Juan Gómez-Luna, Izzat El Hajj, Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu. 2022. Benchmarking a New Paradigm: Experimental Analysis and Characterization of a Real Processing-in-Memory System. IEEE Access 10 (2022), 52565–52608. https://doi.org/10...

  15. [24]

    Bongjoon Hyun, Taehun Kim, Dongjae Lee, and Minsoo Rhu. 2024. Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 263–279

  16. [25]

    Mohamed Assem Ibrahim, Mahzabeen Islam, and Shaizeen Aga. 2024. Bal- anced Data Placement for GEMV Acceleration with Processing-In-Memory. arXiv:2403.20297 [cs.AR]

  17. [26]

    Asif Ali Khan, Hamid Farzaneh, Karl F. A. Friebel, Clément Fournier, Lorenzo Chelini, and Jeronimo Castrillon. 2023. CINM (Cinnamon): A Compilation Infras- tructure for Heterogeneous Compute In-Memory and Compute Near-Memory Paradigms. arXiv:2301.07486 [cs.AR]

  18. [27]

    Jin Hyun Kim, Shin-Haeng Kang, Sukhan Lee, Hyeonsu Kim, Yuhwan Ro, Se- ungwon Lee, David Wang, Jihyun Choi, Jinin So, YeonGon Cho, JoonHo Song, Jeonghyeon Cho, Kyomin Sohn, and Nam Sung Kim. 2022. Aquabolt-XL HBM2- PIM, LPDDR5-PIM With In-Memory Processing, and AXDIMM With Acc...

  19. [28]

    Jin Hyun Kim, Shin-haeng Kang, Sukhan Lee, Hyeonsu Kim, Woongjae Song, Yuhwan Ro, Seungwon Lee, David Wang, Hyunsung Shin, Bengseng Phuah, Jihyun Choi, Jinin So, YeonGon Cho, JoonHo Song, Jangseok Choi, Jeonghyeon Cho, Kyomin Sohn, Youngsoo Sohn, Kwangil Park, and Nam Sung Kim...

  20. [29]

    Jens Knoop, Oliver Rüthing, and Bernhard Steffen. 1994. Partial dead code elimi- nation. ACM Sigplan Notices 29, 6 (1994), 147–158

  21. [30]

    Michael Kruse and Hal Finkel. 2018. User-Directed Loop-Transformations in Clang. In 2018 IEEE/ACM 5th Workshop on the LLVM Compiler Infrastructure in HPC (LLVM-HPC). 49–58. https://doi.org/10.1109/LLVM-HPC.2018.8639402

  22. [31]

    Yongkee Kwon, Kornijcuk Vladimir, Nahsung Kim, Woojae Shin, Jongsoon Won, Minkyu Lee, Hyunha Joo, Haerang Choi, Guhyun Kim, Byeongju An, Jeongbin Kim, Jaewook Lee, Ilkon Kim, Jaehan Park, Chanwook Park, Yosub Song, Byeongsu Yang, Hyungdeok Lee, Seho Kim, Daehan Kwon, Seongju L...

  23. [32]

    Ruihang Lai, Junru Shao, Siyuan Feng, Steven S Lyubomirsky, Bohan Hou, Wuwei Lin, Zihao Ye, Hongyi Jin, Yuchen Jin, Jiawei Liu, et al . 2023. Relax: Compos- able Abstractions for End-to-End Dynamic Machine Learning. arXiv preprint arXiv:2311.02103 (2023)

  24. [33]

    Lattner and V

    C. Lattner and V. Adve. 2004. LLVM: a compilation framework for lifelong program analysis & transformation. In International Symposium on Code Generation and Optimization, 2004. CGO 2004. 75–86. https://doi.org/10.1109/CGO.2004.1281665

  25. [34]

    Chris Lattner, Mehdi Amini, Uday Bondhugula, Albert Cohen, Andy Davis, Jacques Pienaar, River Riddle, Tatiana Shpeisman, Nicolas Vasilache, and Oleksandr Zi- nenko. 2021. MLIR: Scaling Compiler Infrastructure for Domain Specific Com- ATiM: Autotuning Tensor Programs for Proces...

  26. [35]

    Sukhan Lee, Shin-haeng Kang, Jaehoon Lee, Hyeonsu Kim, Eojin Lee, Seungwoo Seo, Hosang Yoon, Seungwon Lee, Kyounghwan Lim, Hyunsung Shin, Jinhyun Kim, O Seongil, Anand Iyer, David Wang, Kyomin Sohn, and Nam Sung Kim. 2021. Hardware Architecture and Software Stack for PIM Based...

  27. [36]

    Seongju Lee, Kyuyoung Kim, Sanghoon Oh, Joonhong Park, Gimoon Hong, Dongy- oon Ka, Kyudong Hwang, Jeongje Park, Kyeongpil Kang, Jungyeon Kim, Junyeol Jeon, Nahsung Kim, Yongkee Kwon, Kornijcuk Vladimir, Woojae Shin, Jongsoon Won, Minkyu Lee, Hyunha Joo, Haerang Choi, Jaewook L...

  28. [37]

    Chaemin Lim, Suhyun Lee, Jinwoo Choi, Jounghoo Lee, Seongyeon Park, Hanjun Kim, Jinho Lee, and Youngsok Kim. 2023. Design and Analysis of a Processing-in- DIMM Join Algorithm: A Case Study with UPMEM DIMMs. Proc. ACM Manag. Data 1, 2, Article 113 (June 2023), 27 pages. https:/...

  29. [38]

    Junfeng Lin, Huanyu Qu, Songchen Ma, Xinglong Ji, Hongyi Li, Xiaochuan Li, Chenhang Song, and Weihao Zhang. 2023. SongC: A Compiler for Hybrid Near- Memory and In-Memory Many-Core Architecture. IEEE Trans. Comput. (2023), 1–14. https://doi.org/10.1109/TC.2023.3311948

  30. [39]

    LLVM. [n. d.]. LLVM’s Analysis and Transform Passes. https://llvm.org/docs/ Passes.html

  31. [40]

    Martin Paul Lücke, Oleksandr Zinenko, William S Moses, Michel Steuwer, and Albert Cohen. 2024. The MLIR Transform Dialect. Your compiler is more powerful than you think. arXiv preprint arXiv:2409.03864 (2024)

  32. [41]

    Onur Mutlu, Saugata Ghose, Juan Gómez-Luna, and Rachata Ausavarungnirun

  33. [42]

    R. Nair, S. F. Antao, C. Bertolli, P. Bose, J. R. Brunheroto, T. Chen, C.-Y. Cher, C. H. A. Costa, J. Doi, C. Evangelinos, B. M. Fleischer, T. W. Fox, D. S. Gallo, L. Grinberg, J. A. Gunnels, A. C. Jacob, P. Jacob, H. M. Jacobson, T. Karkhanis, C. Kim, J. H. Moreno, J. K. O’Br...

  34. [43]

    John Nickolls, Ian Buck, Michael Garland, and Kevin Skadron. 2008. Scalable parallel programming with CUDA. In ACM SIGGRAPH 2008 Classes (Los Angeles, California) (SIGGRAPH ’08). Association for Computing Machinery, New York, NY, USA, Article 16, 14 pages. https://doi.org/10.1...

  35. [44]

    G. F. Oliveira, A. Olgun, A. Yaglikci, F. Bostanci, J. Gomez-Luna, S. Ghose, and O. Mutlu. 2024. MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple- Instruction Multiple-Data Computing. In 2024 IEEE Int...

  36. [45]

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fan...

  37. [46]

    Thomas Pawlowski

    J. Thomas Pawlowski. 2011. Hybrid memory cube (HMC). In 2011 IEEE Hot Chips 23 Symposium (HCS). 1–24. https://doi.org/10.1109/HOTCHIPS.2011.7477494

  38. [47]

    Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe. 2013. Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines. In Proceedings of the 34th ACM SIGPLAN Confere...

  39. [48]

    Jared Roesch, Steven Lyubomirsky, Logan Weber, Josh Pollock, Marisa Kirisame, Tianqi Chen, and Zachary Tatlock. 2018. Relay: a new IR for machine learning frameworks. In Proceedings of the 2nd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages (Ph...

  40. [49]

    Samsung Advanced Institute of Technology. 2022. OneMCC. https://github.com/ SAITPublic/OneMCC

  41. [50]

    Samsung Advanced Institute of Technology. 2022. PIMLibrary. https://github. com/SAITPublic/PIMLibrary

  42. [51]

    Samsung Advanced Institute of Technology. 2022. PIMSimulator. https://github. com/SAITPublic/PIMSimulator

  43. [52]

    Junru Shao, Xiyou Zhou, Siyuan Feng, Bohan Hou, Ruihang Lai, Hongyi Jin, Wuwei Lin, Masahiro Masuda, Cody Hao Yu, and Tianqi Chen. 2022. Tensor Program Optimization with Probabilistic Programs. In Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Ag...

  44. [53]

    ShareGPT Team. 2023. ShareGPT. https://sharegpt.com,2023

  45. [54]

    Haichen Shen, Jared Roesch, Zhi Chen, Wei Chen, Yong Wu, Mu Li, Vin Sharma, Zachary Tatlock, and Yida Wang. 2021. Nimble: Efficiently Com- piling Dynamic Neural Networks for Model Inference. In Proceedings of Machine Learning and Systems , A. Smola, A. Dimakis, and I. Stoica (...

  46. [55]

    Yongwon Shin, Juseong Park, Sungjun Cho, and Hyojin Sung. 2023. PIMFlow: Compiler and Runtime Support for CNN Models on Processing-in-Memory DRAM. In Proceedings of the 21st ACM/IEEE International Symposium on Code Generation and Optimization (<conf-loc>, <city>Montréal</city>...

  47. [56]

    SiFive. 2021. SiFive E21 Core Complex Manual

  48. [57]

    Stone, David Gohara, and Guochun Shi

    John E. Stone, David Gohara, and Guochun Shi. 2010. OpenCL: A Parallel Pro- gramming Standard for Heterogeneous Computing Systems.Computing in Science & Engineering 12, 3 (2010), 66–73. https://doi.org/10.1109/MCSE.2010.69

  49. [58]

    Hashimoto

    Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford Alpaca: An Instruction-following LLaMA model

  50. [59]

    Philippe Tillet, H. T. Kung, and David Cox. 2019. Triton: an intermediate language and compiler for tiled neural network computations. InProceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Lan- guages (Phoenix, AZ, USA) (MAPL 2019). As...

  51. [60]

    UPMEM. [n. d.]. llvm-project. https://github.com/upmem/llvm-project

  52. [61]

    Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen

    Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S. Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen. 2018. Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions. arXiv:1802.04730 [cs.PL]

  53. [62]

    Pirmin Vogel, Andrea Marongiu, and Luca Benini. 2018. Exploring shared virtual memory for FPGA accelerators with a configurable IOMMU. IEEE Trans. Comput. 68, 4 (2018), 510–525

  54. [63]

    Ben Wang. 2021. Mesh-Transformer-JAX: Model-Parallel Implementation of Trans- former Language Model with JAX

  55. [64]

    XILINX. 2020. Versal ACAP AI Engine Architecture Manual

  56. [65]

    Joseph Yiu. 2009. The definitive guide to the ARM Cortex-M3 . Newnes

  57. [66]

    Bojian Zheng, Ziheng Jiang, Cody Hao Yu, Haichen Shen, Josh Fromm, Yizhi Liu, Yida Wang, Luis Ceze, Tianqi Chen, and Gennady Pekhimenko. 2022. DietCode: Automatic optimization for dynamic tensor program. (2022)

  58. [67]

    Gonzalez, and Ion Stoica

    Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, and Ion Stoica. 2020. Ansor: Generating High-Performance Tensor Programs for Deep Learning. In 14th USENIX Symposium on Operating S...

  59. [2015]

    IBM Journal of Research and Development 59, 2/3 (2015), 17:1–17:14

    Active Memory Cube: A processing-in-memory architecture for exascale systems. IBM Journal of Research and Development 59, 2/3 (2015), 17:1–17:14. https://doi.org/10.1147/JRD.2015.2409732

  60. [2022]

    In Emerging Computing: From Devices to Systems: Looking Beyond Moore and Von Neumann

    A modern primer on processing in memory. In Emerging Computing: From Devices to Systems: Looking Beyond Moore and Von Neumann . Springer, 171–243

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.