Pith. sign in

REVIEW 4 major objections 4 minor 89 references

L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference

T0 review · 4 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read The paper claims that long-context LLM inference can escape the HBM capacity-bandwidth trade-off by offloading decoding attention and the whole KV cache to DIMM-PIM, with up to 6.1x speedup over HBM-PIM solutions.

desk verdict L3 is a thoughtful, well-scoped attempt to offload decode-attention to DIMM-PIM, but its flagship 6.1x speedup depends on an unvalidated SPD timing spoof that a real memory controller would likely reject; worth a serious referee, not a desk reject. read the letter →

arxiv 2504.17584 v1 pith:ENF2DUYO submitted 2025-04-24 cs.AR cs.LG

classification cs.ARcs.LG
keywords LLMinferenceprocessing-in-memoryDIMM-PIMKVcachelong-contextheterogeneousschedulingDDR4timingmulti-headattention
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

L3 claims that the capacity and bandwidth pressure of long-context LLM inference comes almost entirely from one operation, the decoding-phase multi-head attention (MHA) that must repeatedly read the per-request key-value (KV) cache. The paper therefore offloads that operation together with the entire KV cache to DIMM-PIM, ordinary host-memory modules augmented with processing units near the memory banks, while keeping the compute-heavy fully-connected layers on the GPU. It reports up to a 6.1x throughput gain over HBM-PIM accelerators, up to 5.0x over a GPU-only server, and up to 14.3x larger batch sizes on a simulated eight-GPU server with 2 TB of DIMM-PIM, without increasing the time between generated tokens. The broader point is that scaling memory capacity and bandwidth together matters more for long-context serving than either alone.

What carries the argument

The central mechanism is the zero-latency in-flight re-layout unit on the DIMM's rank buffer chip, backed by a double-buffering stage and deliberately altered timing values reported to the host memory controller so that burst writes start early and the unit has a window to swap bits without stalling the DDR bus. Around it sit two complementary KV-cache mappings, one that distributes the elements of a new K or V vector across chips and fixed banks for broadcast inner products, and one that scatters tokens across banks in burst-sized chunks for outer-product context computation, plus configurable bank-level and rank-level processing elements. The re-layout removes the bit-level mismatch between DRAM chip width and FP16 element width, the mappings remove the element-level mismatch between DDR layout and PIM's need for locality and regularity, and the spoofed timing is what makes the re-layout appear free.

What would settle it

Use a DDR4 memory controller that enforces the JEDEC timing parameters exactly as the DIMM's original SPD describes; if the controller rejects the spoofed values or the early burst data collides with the still-busy re-layout unit, causing corruption or bus contention, L3's zero-latency claim is falsified.

Watch

Extended reading notes

Core claim

On its own terms, the paper's discovery is that the memory bottleneck in LLM decode is not general but operation-specific: multi-head attention in the decoding phase needs both large storage for the KV cache and high bandwidth to stream it, and DIMM-PIM can supply both because its capacity and bandwidth scale with the number of plugged-in modules. L3 makes this work by fixing two data-placement mismatches: a re-layout unit on the rank buffer chip rearranges bits in flight during ordinary DDR burst writes so that each FP16 element resides in a single DRAM chip, and two KV-cache mapping schemes place co-processed K and V elements contiguously for score and context computations. Rank-level and bank-level processing units then fuse score, softmax, and context into a pipelined kernel, while a rankset-based communication scheme overlaps PCIe transfers with computation and an adaptive two-sub-batch scheduler overlaps prefilling with decoding to keep both devices busy. The reported result is a long-context inference system whose throughput and batch size scale with host memory rather than with HBM capacity.

Load-bearing premise

The load-bearing premise is that a real DDR4 memory controller will tolerate the memory module reporting deliberately altered timing values, so the module's re-layout unit can swap bits during a write burst without corrupting data or stalling the bus; the paper does not validate this against a real controller or a protocol-level simulation.

Editorial extensions

If this is right

  • A GPU server with DIMM-PIM host memory can serve long-context workloads at batch sizes that would run out of HBM, converting spare host capacity into higher GPU utilization.
  • Time-between-tokens need not rise when KV caches leave the GPU; the paper reports TBT comparable to GPU-only at the smallest configuration and 29-53% of it at 16 ranksets.
  • Scaling capacity alone or bandwidth alone gives little (1.1-1.6x at 8x scale-up), whereas scaling both gives 5.1x, so memory-system scaling must be coordinated.
  • Bank-level PUs on DIMM-PIM provide roughly 8x the bandwidth of rank-level DIMM-PIM, which is why the bank-level design is needed to keep up with server GPUs.
  • Prefill and decode can run on opposite sub-batches on the two devices, letting the scheduler hide most KV offload and projection/feed-forward behind each other.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the spoofed-SPD timing trick is the link most likely to break under a stricter memory controller; if protocol-level validation fails, L3's re-layout becomes an on-the-critical-path cost and the 6.1x speedup shrinks by the amount of that cost.
  • The same bit-level re-layout idea transfers to other PIM media (GDDR, HBM) and to mixed-precision KV caches, where rank or buffer logic could transpose elements during burst or refresh windows rather than with CPU copies.
  • Because L3's scheduler deliberately chunks at most one prefilling request per batch, the approach may combine cleanly with sparse-attention or KV-quantization schemes, which would cut the KV stream further and make DIMM-PIM bandwidth stretch even farther.
  • If DIMM-PIM evolves into pooled CXL-attached memory, the rankset load-balancing and two-sub-batch scheduler give a concrete recipe for using pooled capacity without exposing memory latency to the decode loop.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes L3, a hardware-software co-designed system that fully offloads decoding-phase multi-head attention and the entire KV cache to DIMM-based processing-in-memory in a GPU server. It claims three contributions: a zero-latency in-flight re-layout mechanism based on deliberately spoofed DDR4 SPD timing values, KV-cache mapping and cross-level processing units that enable bank-level PIM attention kernels, and a scheduler plus communication-overlap scheme that hides PCIe transfers and minimizes idle bubbles. Evaluation is carried out in a DRAMsim3/AttAcc-based simulator on three LLM models and four real-world traces, reporting up to 6.1x speedup over an HBM-PIM baseline and significantly improved batch sizes.

Significance. If the results hold, the paper targets a real and growing bottleneck: long-context decoding is simultaneously capacity- and bandwidth-bound, and DIMM-PIM offers a modular path to scale both dimensions. The authors correctly isolate decoding MHA as the critical operation, and the layered rank-PU/bank-PU design with two distinct KV mappings is a thoughtful response to the bit-level and element-level layout mismatches. The paper also ships a cycle-accurate simulator, synthesizes the PU logic for area and power estimates, evaluates on real traces, and includes an ablation study. However, the headline speedup rests on the unvalidated SPD-spoofing mechanism, the 'bubble-free' pipeline claim is asserted rather than proven, and the evaluation is closed-loop in the sense that the scheduler models are trained on the same simulator used for measurement. These issues are load-bearing and need to be resolved before the central claims can be considered reliable.

major comments (4)
  1. [§4.1 (Spoofed timing constraints)] The central enabler of the zero-latency in-flight re-layout is the claim that deliberately reporting a tWL of one cycle and enlarged post-write latencies through SPD is safe and entails 'no DDR protocol violations'. No protocol-level validation is provided. The simulator described in §6.1 is DRAMsim3 plus AttAcc; it does not model a host memory controller reacting to a DIMM whose programmed write latency differs from the physical value by roughly 15 cycles. DDR4 write leveling, mode-register training, and JEDEC minimum timing constraints exist precisely to reject this configuration. If a real controller retrains, stalls, or issues a subsequent command while the re-layout buffer is busy, the 'zero-latency' re-layout becomes a multi-cycle stall and the speedup collapses to the CPU-side transpose cost estimated in §2.5. Please provide a concrete protocol-level analysis or a cycle-accurate MC-plus-DIMM simulation demonstrating that a standards-compliant controller will tolerate the spoofed timing, or revise the claim to remove the zero-latency assumption.
  2. [§4.3 (Kernel fusion with bubble-free pipelining)] The text states that 'Quantitative analysis of the pipeline execution proves it is bubble-free', but no such analysis appears in the paper. The bubble-free property is load-bearing because it underlies the claim that kernel fusion adds no overhead and that the PIM-side pipeline is fully utilized. The proof must be supplied with explicit assumptions about chunk size, on-chip buffer capacity, softmax unit latency, and the synchronization points between bank PUs, rank PU, and the host. As written, the claim is unsupported.
  3. [§5.3 and §6.1 (Evaluation methodology)] The scheduler's latency models (Eqs. 5-6) are trained on profiling data collected from the same DRAMsim3-based simulator that is then used to measure end-to-end throughput. This creates a closed loop: the scheduler is fitted to the simulator, and the simulator is used to demonstrate the benefit of the scheduler. Additionally, all baselines are author re-implementations rather than the original published systems. At minimum, please report variance across multiple simulator runs or seeds, validate the prediction models on held-out configurations rather than held-out batches from the same distribution, and clarify which baseline parameters are taken directly from the original papers versus assumed by the authors.
  4. [§1 and §6.2 (Batch-size claim)] The abstract and introduction claim 'significantly improved batch sizes (up to 14.3x on DGX-A100)', but the evaluation never reports batch sizes or the 14.3x figure. §6.2 only gives qualitative statements such as 'L3 achieves much larger batch sizes compared to the HBM-based baselines'. Please add a quantitative batch-size comparison for the traced workloads and report the configuration that yields the 14.3x number.
minor comments (4)
  1. [§2.6 and Table 1] The citation for NEO is inconsistent: §2.6 text cites 'NEO [33]' while Table 1 lists 'NEO [40]'. The reference list entry [40] is the NEO paper; please correct the in-text citation.
  2. [§4.1 (Double buffering)] The relationship between the proposed double buffer and the conventional LRDIMM data-buffer path should be clarified with a timing diagram. The text says the double buffer 'replaces the conventional single-buffer approach' but does not specify what the conventional buffer is or how the re-layout unit interacts with the DDR4 burst timing.
  3. [§5.1 (Communication hiding)] The claim that 'the transfer of the prefilling KV cache can always be hidden' and that it is 'typically <16% of the Feed-forward latency' is not backed by any figure or table. Please provide supporting data or move this statement to the evaluation.
  4. [§6.4 (Table 5)] The bank PU area and power are synthesized in a logic process, and the footnote states that a DRAM process would incur 10x area overhead. Please report the resulting per-DIMM area and power overhead so the reader can assess the total hardware cost of the design.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: L3's architectural claims are self-contained proposals with simulation-based evaluation, not derivations that reduce to their inputs.

full rationale

The paper's central claims—offloading decoding MHA and KV cache to DIMM-PIM, the in-flight re-layout, the KV mapping methods, communication-computation overlap, and the adaptive scheduler—are presented as new designs and evaluated in a simulator built on DRAMsim3 and AttAcc. None of the claimed speedups is obtained by fitting a parameter to the target quantity and then 'predicting' it. The scheduler's latency models are trained on profiled data and validated on a held-out 20% test set, which is standard empirical practice and does not make the end-to-end result circular. The 'spoofed timing constraints' mechanism is an architectural assumption about memory-controller behavior, not a circular derivation; its validity is a correctness risk (unvalidated protocol-level behavior), not a logical reduction of the paper's conclusions to its premises. No load-bearing self-citation chain or imported uniqueness theorem is present. The evaluation is self-contained (same simulator for training and testing), which limits external confirmation but is not circularity under the stated criteria.

Assumptions & free parameters 5 free parameters · 5 assumptions · 3 invented entities

The contribution is a system design, so the ledger records the design assumptions and simulator constants that the speedup claim rests on. The largest items are the assumed bandwidth multipliers for DIMM-PIM and HBM-PIM, the unvalidated SPD timing spoof, the 256 KB rank-buffer design carried from Hermes, and the fitted scheduler latency models. There are no mathematical derivations that smuggle in the result.

free parameters (5)
  • HBM-PIM baseline aggregate bandwidth = 260.8 TB/s
    Table 2 sets HBM-PIM bandwidth to 16x the GPU HBM figure of 16.3 TB/s. This internal-bandwidth multiplier is an assumption for the baseline, not a measured value from NeuPIMs or AttAcc, and it shapes how strong the comparison baseline is.
  • DIMM-PIM aggregate bandwidth = 13.0 TB/s
    Table 2 derives 13.0 TB/s from 16 channels of DDR4-3200 at 406 GB/s nominal with a bank-level PIM multiplier. The quote of more than 30x from prior work is cited, but the exact 32x multiplier for this system is assumed.
  • Rank PU on-chip buffer size = 256 KB
    Section 4.4 sets the rank-PU buffer to 256 KB following Hermes, which supports about 128K tokens in one pass. Sequences beyond that require repeated fetching, so the buffer size bounds the latency behavior for very long contexts.
  • Scheduler latency model parameters = not disclosed
    Section 5.3 trains Random Forest Regression and linear models on runtime data to predict GPU and PIM latencies. The fitted coefficients and RFR hyperparameters are not given, yet the scheduler alignment and the reported speedup depend on these models.
  • tWL and tWR SPD spoof offsets = tWL reduced by one cycle; tWR increased
    Section 4.1 chooses the spoofed values by hand to create a re-layout window. The exact timing margins are not analyzed, so the zero-latency claim depends on these unquantified offsets.
assumptions (5)
  • domain assumption DIMM-PIM capacity and bandwidth scale linearly with the number of DIMMs and ranksets.
    Used in Sections 2.4 and 6.2 to justify the architectural match with decoding MHA. Based on cited PIM and DIMM work, but assumed for this system.
  • domain assumption Bank-level PUs can be integrated into commercial DRAM banks with acceptable overhead.
    Section 6.4 reports synthesis in a TSMC 28nm logic process and notes that a DRAM process would incur 10x area overhead. The ability to build the bank PUs inside DRAM is assumed, not demonstrated.
  • ad hoc to paper The host memory controller will behave correctly under deliberately spoofed SPD timing parameters.
    Section 4.1 relies on this to hide re-layout inside DDR burst transfers. No protocol-level or silicon validation is provided, and this is the weakest hardware premise.
  • domain assumption PCIe communication can overlap with DIMM-PIM computation without significant interference.
    Section 5.1 assumes one rankset per channel can receive data while other ranksets compute, preserving 75% of DIMM-PIM compute. The interference between PCIe DMA and PIM traffic is modeled in the simulator, not measured on hardware.
  • domain assumption The linear and Random Forest latency models generalize to previously unseen batches.
    Section 5.3 uses the models to align GPU and DIMM-PIM execution. Figure 10-b reports held-out prediction error from simulated data only, so generalization to real hardware is unverified.
invented entities (3)
  • Rank PU with re-layout, softmax, and adder units
    purpose: Performs in-flight bit-level re-layout, chunk softmax, and partial-sum aggregation on the DIMM rank's buffer chip.
    Described in Sections 4.1 to 4.3 and synthesized in SystemVerilog at TSMC 28nm. No working DIMM-PIM rank PU exists to independently validate the design.
  • Bank-level PU in each physical DRAM bank
    purpose: Executes the score and context multiply-accumulate operations using data from the bank row buffer and shared broadcast buffers.
    Logic-level synthesis only. DRAM-process integration and the 10x area overhead caveat in Table 5 mean the bank PU is a proposed entity, not an existing product.
  • Rankset-based communication mode hardware extension
    purpose: Allows one rank per channel to receive PCIe offload data while other ranks in the same channel continue PIM computation.
    Introduced in Section 5.1 as a new hardware extension to enable concurrent communication and computation. No implementation details or validation are provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference." pith.science (2026). https://pith.science/paper/ENF2DUYO

@misc{pith2026250417584,
  author       = {Pith},
  title        = {Pith review of: L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ENF2DUYO}},
  note         = {Machine review of arXiv:2504.17584}
}
abstract

Large Language Models (LLMs) increasingly require processing long text sequences, but GPU memory limitations force difficult trade-offs between memory capacity and bandwidth. While HBM-based acceleration offers high bandwidth, its capacity remains constrained. Offloading data to host-side DIMMs improves capacity but introduces costly data swapping overhead. We identify that the critical memory bottleneck lies in the decoding phase of multi-head attention (MHA) exclusively, which demands substantial capacity for storing KV caches and high bandwidth for attention computation. Our key insight reveals this operation uniquely aligns with modern DIMM-based processing-in-memory (PIM) architectures, which offers scalability of both capacity and bandwidth. Based on this observation and insight, we propose L3, a hardware-software co-designed system integrating DIMM-PIM and GPU devices. L3 introduces three innovations: First, hardware redesigns resolve data layout mismatches and computational element mismatches in DIMM-PIM, enhancing LLM inference utilization. Second, communication optimization enables hiding the data transfer overhead with the computation. Third, an adaptive scheduler coordinates GPU-DIMM-PIM operations to maximize parallelism between devices. Evaluations using real-world traces show L3 achieves up to 6.1$\times$ speedup over state-of-the-art HBM-PIM solutions while significantly improving batch sizes.

Figures

Figures reproduced from arXiv: 2504.17584 by the authors.

Figure 1
Figure 1. LLM inference process. The FC operations for each request can be batched and are compute-intensive, whereas MHA cannot be batched and is memory bandwidth-intensive. improved batch sizes (up to 14.3× on DGX-A100 [3]). L3 is the first host memory offloading system for higher throughput without sacrificing latencies (time-between-tokens) compared with server￾grade GPUs such as A100, and will be open source. 2 Motivatio… view at source ↗
Figure 2
Figure 2. GPU memory bottlenecks with Llama-7B on A100. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Challenges of DIMM-PIM integration. are exclusive to “decoding MHA”. On the contrary, for the FC oper￾ations in decoding, their computation efficiency is solely related to the batch size, and an increase in token length does not incur additional computational load or memory capacity requirements. This insight motivates our architectural proposal: decoupling and fully offloading “decoding MHA” along with the entire K… view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: L3 overview. Hardware-Software co-optimization. Prior research has inves￾tigated GPU-PIM integration for inference systems to capitalize on PIM’s bandwidth advantages. Systems such as NeuPIMs [33], AttAcc [72], and PAPI [32] enhance GPUs with HBM-PIM, achiev￾ing consid…
Figure 5
Figure 5. Figure 5: DIMM-PIM architecure for LLM attention. the continuous data flow. Consequently, the double buffer creates a processing window where re-layout can occur. Spoofed timing constraints. To address potential DDR timing violations caused by re-layout processing, we implement …
Figure 6
Figure 6. Figure 6: Communication-computation overlapping with [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Computation graph in L3. Data transfer between devices can be overlapped with computation. “Gen” denotes QKV Generation and “Proj, FF” denotes the projection and feed-forward operations. layer is assigned to different channels within a rankset. In this man￾ner, L3 ensu…
Figure 8
Figure 8. Figure 8: End-to-end inference throughput on real-world traces. [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: Scalability analysis. The results are evaluated using the OpenR1 trace. Scalability analysis. We further evaluate L3 performance with various host memory configurations, showing that performance scalability requires simultaneous scaling in memory capacity and bandwidth…
Figure 10
Figure 10. Figure 10: Ablation study and prediction errors. The throughput with the most naive implementation in (a) is normalized to 1. 0 1 2 3 GPU RS-2 RS-4 RS-8 RS-16 GPU RS-2 RS-4 RS-8 RS-16 GPU RS-2 RS-4 RS-8 RS-16 GPU RS-2 RS-4 RS-8 RS-16 Normalized TBT MHA QKVGen Proj,FF Batch Size=…
Figure 11
Figure 11. Figure 11: Latency analysis. RS-N denotes to L3 with N ranksets. The GPU baseline is assumed infinite memory capacity. communication-computation overlap further improves the through￾put by ∼1.7×. These results demonstrate that our system designs effectively enhance throughput in…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

89 extracted references · 14 canonical work pages

  1. [1]

    DDR4 SDRAM LRDIMM

    2015. DDR4 SDRAM LRDIMM. [Online]. Avalable: https://www.micron.com/ content/dam/micron/global/secure/spectek/data-sheet/dram/ddr4/spectek- 8gb-ddr4-sdram.pdf

  2. [2]

    JEDEC Standard: DDR4 SDRAM Load Reduced DIMM (LRDIMM) Design Specification

    2015. JEDEC Standard: DDR4 SDRAM Load Reduced DIMM (LRDIMM) Design Specification. [Online]. Avalable: https://www.jedec.org/standards-documents/ docs/module4_20_27

  3. [3]

    NVIDIA DGX A100

    2023. NVIDIA DGX A100. [Online]. Avalable: https://resources.nvidia.com/enus- dgx-systems/dgx-ai

  4. [4]

    cognitivecomputations/dolphin-r1 · Datasets at Hugging Face

    2025. cognitivecomputations/dolphin-r1 · Datasets at Hugging Face. https: //huggingface.co/datasets/cognitivecomputations/dolphin-r1. Referenced April 2025

  5. [5]

    open-r1/OpenR1-Math-220k · Datasets at Hugging Face

    2025. open-r1/OpenR1-Math-220k · Datasets at Hugging Face. https:// huggingface.co/datasets/open-r1/OpenR1-Math-220k. Referenced April 2025

  6. [6]

    open-r1/OpenThoughts-114k-math · Datasets at Hugging Face

    2025. open-r1/OpenThoughts-114k-math · Datasets at Hugging Face. https: //huggingface.co/datasets/open-r1/OpenThoughts-114k-math. Referenced April 2025

  7. [7]

    Gulavani, Alexey Tumanov, and Ramachandran Ramjee

    Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024. Taming throughput-latency tradeoff in LLM inference with sarathi-serve. In Proceedings of the 18th USENIX Conference on Operating Systems Design and Implementation (Santa Clara, CA, USA) (OSDI’24). USENIX Association...

  8. [8]

    Mohammad Alian, Seung Won Min, Hadi Asgharimoghaddam, Ashutosh Dhar, Dong Kai Wang, Thomas Roewer, Adam McPadden, Oliver O’Halloran, Deming Chen, Jinjun Xiong, Daehoon Kim, Wen-mei Hwu, and Nam Sung Kim. 2018. Application-Transparent Near-Memory Processing Architecture with Memory Channel Network. In 2018 51st Annual IEEE/ACM International Symposium on Mi...

Show all 89 references
  1. [9]

    Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H

    Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier- Hellstern, Gaurav Mishra, Erica Moreira, Mark Omer...

  2. [11]

    Hadi Asghari-Moghaddam, Young Hoon Son, Jung Ho Ahn, and Nam Sung Kim

  3. [12]

    Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2024. LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding. arXiv:2308.14508 [cs.CL] https://ar...

  4. [13]

    Malladi, Hongzhong Zheng, and Onur Mutlu

    Amirali Boroumand, Saugata Ghose, Minesh Patel, Hasan Hassan, Brandon Lu- cia, Rachata Ausavarungnirun, Kevin Hsieh, Nastaran Hajinazar, Krishna T. Malladi, Hongzhong Zheng, and Onur Mutlu. 2019. CoNDA: efficient cache coherence support for near-data accelerators. In Proceedin...

  5. [14]

    Dan Chen, Haiheng He, Hai Jin, Long Zheng, Yu Huang, Xinyang Shen, and Xiaofei Liao. 2023. MetaNMP: Leveraging Cartesian-Like Product to Acceler- ate HGNNs with Near-Memory Processing. In Proceedings of the 50th Annual International Symposium on Computer Architecture (Orlando,...

  6. [15]

    Quan Chen, Hailong Yang, Minyi Guo, Ram Srivatsa Kannan, Jason Mars, and Lingjia Tang. 2017. Prophet: Precise QoS Prediction on Non-Preemptive Acceler- ators to Improve Utilization in Warehouse-Scale Computers. SIGARCH Comput. Archit. News 45, 1 (April 2017), 17–32. https://do...

  7. [17]

    Weihao Cui, Han Zhao, Quan Chen, Ningxin Zheng, Jingwen Leng, Jieru Zhao, Zhuo Song, Tao Ma, Yong Yang, Chao Li, and Minyi Guo. 2021. Enable Simultane- ous DNN Services Based on Deterministic Operator Overlap and Precise Latency Prediction. In SC21: International Conference fo...

  8. [18]

    Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

    Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

  9. [19]

    DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei F...

  10. [20]

    Fabrice Devaux. 2019. The true Processing In Memory accelerator. In 2019 IEEE Hot Chips 31 Symposium (HCS) . 1–24. https://doi.org/10.1109/HOTCHIPS.2019. 8875680

  11. [22]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 [cs.CL] https://arxiv.org/abs/1810.04805

  12. [23]

    Siying Feng, Xin He, Kuan-Yu Chen, Liu Ke, Xuan Zhang, David Blaauw, Trevor Mudge, and Ronald Dreslinski. 2022. MeNDA: a near-memory multi-way merge solution for sparse transposition and dataflows. In Proceedings of the 49th Annual International Symposium on Computer Architect...

  13. [24]

    Elias Frantar and Dan Alistarh. 2023. SparseGPT: massive language models can be accurately pruned in one-shot. In Proceedings of the 40th International Conference on Machine Learning (Honolulu, Hawaii, USA) (ICML’23). JMLR.org, Article 414, 15 pages

  14. [25]

    Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, and Pengfei Zuo. 2024. Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention. In 2024 USENIX Annual Technical Conference (USENIX ATC 24)...

  15. [26]

    Yizhao Gao, Zhichen Zeng, Dayou Du, Shijie Cao, Peiyuan Zhou, Jiaxing Qi, Junjie Lai, Hayden Kwok-Hay So, Ting Cao, Fan Yang, and Mao Yang. 2025. SeerAtten- tion: Learning Intrinsic Sparse Attention in Your LLMs. arXiv:2410.13276 [cs.CL] https://arxiv.org/abs/2410.13276

  16. [27]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, A...

  17. [28]

    Yufeng Gu, Alireza Khadem, Sumanth Umesh, Ning Liang, Xavier Servot, Onur Mutlu, Ravi Iyer, and Reetuparna Das. 2025. PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference. In Proceedings of the 30th ACM International Conference on Architectural...

  18. [29]

    Oliveira, and Onur Mutlu

    Juan Gómez-Luna, Izzat El Hajj, Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu. 2022. Benchmarking a New Paradigm: Experimental Analysis and Characterization of a Real Processing-in-Memory System. IEEE Access 10 (2022), 52565–52608. https://doi.org/10...

  19. [30]

    Jiaao He and Jidong Zhai. 2024. FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines. arXiv:2403.11421 [cs.DC] https: //arxiv.org/abs/2403.11421

  20. [31]

    Mingxuan He, Choungki Song, Ilkon Kim, Chunseok Jeong, Seho Kim, Il Park, Mithuna Thottethodi, and T. N. Vijaykumar. 2020. Newton: A DRAM-maker’s Accelerator-in-Memory (AiM) Architecture for Machine Learning. In 2020 53rd Annual IEEE/ACM International Symposium on Microarchite...

  21. [32]

    Yintao He, Haiyu Mao, Christina Giannoula, Mohammad Sadrosadati, Juan Gómez-Luna, Huawei Li, Xiaowei Li, Ying Wang, and Onur Mutlu. 2025. PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing System. In Proceedings ...

  22. [34]

    Jie Huang and Kevin Chen-Chuan Chang. 2023. Towards Reasoning in Large Language Models: A Survey. arXiv:2212.10403 [cs.CL] https://arxiv.org/abs/ 2212.10403

  23. [35]

    P. K. Huang, C. Y. Lu, W. H. Wei, Christine Chiu, K. C. Ting, Clark Hu, C.H. Tsai, S. Y. Hou, W. C. Chiou, C. T. Wang, and Douglas Yu. 2021. Wafer Level System Integration of the Fifth Generation CoWoS®-S with High Performance Si Interposer at 2500 mm2. In 2021 IEEE 71st Elect...

  24. [36]

    Wenqin Huangfu, Xueqi Li, Shuangchen Li, Xing Hu, Peng Gu, and Yuan Xie

  25. [37]

    Malladi, Andrew Chang, and Yuan Xie

    Wenqin Huangfu, Krishna T. Malladi, Andrew Chang, and Yuan Xie. 2023. BEA- CON: Scalable Near-Data-Processing Accelerators for Genome Analysis near Memory Pool with the CXL Support. In Proceedings of the 55th Annual IEEE/ACM International Symposium on Microarchitecture (Chicag...

  26. [38]

    Son Hyojun, Jonatan Gilbert, Xiangyu Wu, Cho Haeyoon, Shivdikar Kaustubh, Abellán José L., Joshi Ajay, Kaeli David, and Kim John. 2025. PIMnet: A Domain- Specific Network for Efficient Collective Communication in Scalable PIM

  27. [40]

    Xuanlin Jiang, Yang Zhou, Shiyi Cao, Ion Stoica, and Minlan Yu. 2024. NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference. arXiv:2411.01142 [cs.DC] https://arxiv.org/abs/2411.01142

  28. [41]

    Hongshin Jun, Jinhee Cho, Kangseol Lee, Ho-Young Son, Kwiwook Kim, Hanho Jin, and Keith Kim. 2017. HBM (High Bandwidth Memory) DRAM Technology and Architecture. In 2017 IEEE International Memory Workshop (IMW). 1–4. https: //doi.org/10.1109/IMW.2017.7939084

  29. [42]

    Hongju Kal, Chanyoung Yoo, and Won Woo Ro. 2023. AESPA: Asynchronous Execution Scheme to Exploit Bank-Level Parallelism of Processing-in-Memory. In Proceedings of the 56th Annual IEEE/ACM International Symposium on Microar- chitecture (Toronto, ON, Canada) (MICRO ’23). Associa...

  30. [43]

    Katikapalli Subramanyam Kalyan. 2023. A Survey of GPT-3 Family Large Language Models Including ChatGPT and GPT-4. arXiv:2310.12321 [cs.CL] https://arxiv.org/abs/2310.12321

  31. [44]

    Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter, Ramachan- dran Ramjee, and Ashish Panwar

    Aditya K. Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter, Ramachan- dran Ramjee, and Ashish Panwar. 2025. POD-Attention: Unlocking Full Prefill- Decode Overlap for Faster LLM Inference. In Proceedings of the 30th ACM In- ternational Conference on Architectural Support for ...

  32. [45]

    Lee, Meng Li, Bert Maher, Dheevatsa Mudigere, Maxim Naumov, Martin Schatz, Mikhail Smelyanskiy, Xiaodong Wang, Brandon Reagen, Carole-Jean Wu, Mark Hempstead, and Xuan Zhang

    Liu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks, Vikas Chandra, Utku Diril, Amin Firoozshahian, Kim Hazelwood, Bill Jia, Hsien-Hsin S. Lee, Meng Li, Bert Maher, Dheevatsa Mudigere, Maxim Naumov, Martin Schatz, Mikhail Smelyanskiy, Xiaodong Wang, Brandon Reagen, Carole-...

  33. [47]

    Liu Ke, Xuan Zhang, Jinin So, Jong-Geon Lee, Shin-Haeng Kang, Sukhan Lee, Songyi Han, YeonGon Cho, Jin Hyun Kim, Yongsuk Kwon, KyungSoo Kim, Jin Jung, Ilkwon Yun, Sung Joo Park, Hyunsun Park, Joonho Song, Jeonghyeon Cho, Kyomin Sohn, Nam Sung Kim, and Hsien-Hsin S. Lee. 2022. ...

  34. [48]

    Guhyun Kim, Jinkwon Kim, Nahsung Kim, Woojae Shin, Jongsoon Won, Hyunha Joo, Haerang Choi, Byeongju An, Gyeongcheol Shin, Dayeon Yun, Jeongbin Kim, Changhyun Kim, Ilkon Kim, Jaehan Park, Yosub Song, Byeongsu Yang, Hyeongdeok Lee, Seungyeong Park, Wonjun Lee, Seonghun Kim, Yong...

  35. [49]

    Heesu Kim, Hanmin Park, Taehyun Kim, Kwanheum Cho, Eojin Lee, Soojung Ryu, Hyuk-Jae Lee, Kiyoung Choi, and Jinho Lee. 2021. GradPIM: A Practical Processing-in-DRAM Architecture for Gradient Descent. In 2021 IEEE Interna- tional Symposium on High-Performance Computer Architectu...

  36. [51]

    Jin Hyun Kim, Shin-haeng Kang, Sukhan Lee, Hyeonsu Kim, Woongjae Song, Yuhwan Ro, Seungwon Lee, David Wang, Hyunsung Shin, Bengseng Phuah, Jihyun Choi, Jinin So, YeonGon Cho, JoonHo Song, Jangseok Choi, Jeonghyeon Cho, Kyomin Sohn, Youngsoo Sohn, Kwangil Park, and Nam Sung Kim...

  37. [52]

    Kwiwook Kim and Myeong-jae Park. 2024. Present and Future, Challenges of High Bandwith Memory (HBM). In 2024 IEEE International Memory Workshop (IMW). 1–4. https://doi.org/10.1109/IMW59701.2024.10536972

  38. [53]

    Hyucksung Kwon, Kyungmo Koo, Janghyeon Kim, Woongkyu Lee, Minjae Lee, Hyungdeok Lee, Yousub Jung, Jaehan Park, Yosub Song, Byeongsu Yang, Haerang Choi, Guhyun Kim, Jongsoon Won, Woojae Shin, Changhyun Kim, Gyeongcheol Shin, Yongkee Kwon, Ilkon Kim, Euicheol Lim, John Kim, and ...

  39. [54]

    Se Jung Kwon, Jeonghoon Kim, Jeongin Bae, Kang Min Yoo, Jin-Hwa Kim, Bae- seong Park, Byeongwook Kim, Jung-Woo Ha, Nako Sung, and Dongsoo Lee. 2022. AlphaTuning: Quantization-Aware Parameter-Efficient Adaptation of Large-Scale 14 L3: DIMM-PIM Integrated Architecture and Coordi...

  40. [55]

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Princip...

  41. [56]

    Yongkee Kwon, Guhyun Kim, Nahsung Kim, Woojae Shin, Jongsoon Won, Hyunha Joo, Haerang Choi, Byeongju An, Gyeongcheol Shin, Dayeon Yun, Jeongbin Kim, Changhyun Kim, Ilkon Kim, Jaehan Park, Chanwook Park, Yosub Song, Byeongsu Yang, Hyeongdeok Lee, Seungyeong Park, Wonjun Lee, Se...

  42. [57]

    Youngeun Kwon, Yunjae Lee, and Minsoo Rhu. 2019. TensorDIMM: A Practical Near-Memory Processing Architecture for Embeddings and Tensor Operations in Deep Learning. In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture (Columbus, OH, USA) (MICR...

  43. [58]

    Yongkee Kwon, Kornijcuk Vladimir, Nahsung Kim, Woojae Shin, Jongsoon Won, Minkyu Lee, Hyunha Joo, Haerang Choi, Guhyun Kim, Byeongju An, Jeong- bin Kim, Jaewook Lee, Ilkon Kim, Jaehan Park, Chanwook Park, Yosub Song, Byeongsu Yang, Hyungdeok Lee, Seho Kim, Daehan Kwon, Seongju...

  44. [59]

    John H. Lau. 2022. Recent Advances and Trends in Multiple System and Heterogeneous Integration With TSV-Less Interposers. IEEE Transactions on Components, Packaging and Manufacturing Technology 12, 8 (2022), 1271–1281. https://doi.org/10.1109/TCPMT.2022.3194374

  45. [60]

    Dongjae Lee, Bongjoon Hyun, Taehun Kim, and Minsoo Rhu. 2024. PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems. In 2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). 627–642. https://doi.org/10.1109/MICRO61859.2024.00053

  46. [61]

    Donghun Lee, Jinin So, MINSEON AHN, Jong-Geon Lee, Jungmin Kim, Jeonghyeon Cho, Rebholz Oliver, Vishnu Charan Thummala, Ravi shankar JV, Sachin Suresh Upadhya, Mohammed Ibrahim Khan, and Jin Hyun Kim. 2022. Im- proving In-Memory Database Operations with Acceleration DIMM (AxDI...

  47. [62]

    Hyungdeok Lee, Guhyun Kim, Dayeon Yun, Ilkon Kim, Yongkee Kwon, and Euicheol Lim. 2024. Cost-Effective LLM Accelerator Using Processing in Memory Technology. In 2024 IEEE Symposium on VLSI Technology and Circuits (VLSI Tech- nology and Circuits). 1–2. https://doi.org/10.1109/V...

  48. [63]

    Wonbeom Lee, Jungi Lee, Junghwan Seo, and Jaewoong Sim. 2024. InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . USENIX Association, Santa Clar...

  49. [64]

    In 2022 IEEE Hot Chips 34 Symposium (HCS)

    System Architecture and Software Stack for GDDR6-AiM. In 2022 IEEE Hot Chips 34 Symposium (HCS). 1–25. https://doi.org/10.1109/HCS55958.2022.9895629

  50. [65]

    Jinhao Li, Jiaming Xu, Shan Huang, Yonghua Chen, Wen Li, Jun Liu, Yaoxiu Lian, Jiayi Pan, Li Ding, Hao Zhou, Yu Wang, and Guohao Dai. 2025. Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective. arXiv:2410.04466 [cs.AR] https://arxiv.org/abs/2410.04466

  51. [66]

    Shang Li, Zhiyuan Yang, Dhiraj Reddy, Ankur Srivastava, and Bruce Jacob. 2020. DRAMsim3: A Cycle-Accurate, Thermal-Capable DRAM Simulator.IEEE Comput. Archit. Lett. 19, 2 (July 2020), 106–109. https://doi.org/10.1109/LCA.2020.2973991

  52. [67]

    Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Guangxuan Xiao, and Song Han. 2025. AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration. GetMobile: Mobile Comp. and Comm. 28, 4 (Jan. 2025), 12–17. https://doi.org/10.1145/3714983.3714987

  53. [68]

    Haifeng Liu, Long Zheng, Yu Huang, Chaoqiang Liu, Xiangyu Ye, Jingrui Yuan, Xiaofei Liao, Hai Jin, and Jingling Xue. 2023. Accelerating Personalized Recom- mendation with Cross-level Near-Memory Processing. In Proceedings of the 50th Annual International Symposium on Computer ...

  54. [70]

    Cong Li, Zhe Zhou, Size Zheng, Jiaxi Zhang, Yun Liang, and Guangyu Sun

  55. [71]

    Anirban Nag and Rajeev Balasubramonian. 2021. OrderLight: Lightweight Memory-Ordering Primitive for Efficient Fine-Grained PIM Computations. In MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture (Virtual Event, Greece) (MICRO ’21). Association for Comp...

  56. [72]

    Jaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim, Yongsuk Kwon, Nam Sung Kim, and Jung Ho Ahn. 2024. AttAcc! Unleashing the Power of PIM for Batched Transformer-based Generative Model Inference. InProceedings of the 29th ACM International Conference on Architectura...

  57. [73]

    Jaehyun Park, Byeongho Kim, Sungmin Yun, Eojin Lee, Minsoo Rhu, and Jung Ho Ahn. 2021. TRiM: Enhancing Processor-Memory Interfaces with Scal- able Tensor Reduction in Memory. In MICRO-54: 54th Annual IEEE/ACM In- ternational Symposium on Microarchitecture (Virtual Event, Greec...

  58. [74]

    Sang-Soo Park, KyungSoo Kim, Jinin So, Jin Jung, Jonggeon Lee, Kyoungwan Woo, Nayeon Kim, Younghyun Lee, Hyungyo Kim, Yongsuk Kwon, Jinhyun Kim, Jieun Lee, YeonGon Cho, Yongmin Tai, Jeonghyeon Cho, Hoyoung Song, Jung Ho Ahn, and Nam Sung Kim. 2024. An LPDDR-based CXL-PNM Platf...

  59. [75]

    Aske Plaat, Annie Wong, Suzan Verberne, Joost Broekens, Niki van Stein, and Thomas Back. 2024. Reasoning with Large Language Models, a Survey. arXiv:2407.11511 [cs.AI] https://arxiv.org/abs/2407.11511

  60. [76]

    Minseok Seo, Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon, Guhyun Kim, Chanwook Park, Ilkon Kim, Jaehan Park, Jeongbin Kim, Woojae Shin, Jongsoon Won, Haerang Choi, Kyuyoung Kim, Daehan Kwon, Chunseok Jeong, Sangheon Lee, Yongseok Choi, Wooseok Byun, Seungcheol Baek, Hyu...

  61. [77]

    Lian Liu, Shixin Zhao, Bing Li, Haimeng Ren, Zhaohui Xu, Mengdi Wang, Xiaowei Li, Yinhe Han, and Ying Wang. 2025. Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM. https://arxiv.org/abs/2502.16963

  62. [78]

    Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023. FlexGen: high-throughput generative inference of large language models with a single GPU. In Proceedings of the 40th International Confer...

  63. [79]

    Weiyi Sun, Zhaoshi Li, Shouyi Yin, Shaojun Wei, and Leibo Liu. 2021. ABC-DIMM: alleviating the bottleneck of communication in DIMM-based near-memory pro- cessing with inter-DIMM broadcast. InProceedings of the 48th Annual International Symposium on Computer Architecture (Virtu...

  64. [81]

    Gomez, Łukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, ...

  65. [82]

    Ying Wei, Yi Chieh Huang, Haiming Tang, Nithya Sankaran, Ish Chadha, Dai Dai, Olakanmi Oluwole, Vishnu Balan, and Edward Lee. 2023. 9.3 NVLink-C2C: A Coherent Off Package Chip-to-Chip Interconnect with 40Gbps/pin Single- ended Signaling. In 2023 IEEE International Solid-State ...

  66. [83]

    Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. 2022. ZeroQuant: efficient and affordable post-training quantization for large-scale transformers. In Proceedings of the 36th International Conference on Neural Information Processing Sy...

  67. [84]

    Gon- zalez, and Ion Stoica

    Ying Sheng, Shiyi Cao, Dacheng Li, Coleman Hooper, Nicholas Lee, Shuo Yang, Christopher Chou, Banghua Zhu, Lianmin Zheng, Kurt Keutzer, Joseph E. Gon- zalez, and Ion Stoica. 2023. S-LoRA: Serving Thousands of Concurrent LoRA Adapters. arXiv preprint arXiv:2311.03285 (2023)

  68. [85]

    Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Y. X. Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, and Wangding Zeng. 2025. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Att...

  69. [86]

    Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, and Beidi Chen. 2023. H2O: heavy-hitter oracle for efficient generative inference of large language models. In Proceedings o...

  70. [87]

    Yilong Zhao, Mingyu Gao, Fangxin Liu, Yiwei Hu, Zongwu Wang, Han Lin, Ji Li, He Xian, Hanlin Dong, Tao Yang, Naifeng Jing, Xiaoyao Liang, and Li Jiang

  71. [88]

    Tianyang Zhong, Zhengliang Liu, Yi Pan, Yutong Zhang, Yifan Zhou, Shizhe Liang, Zihao Wu, Yanjun Lyu, Peng Shu, Xiaowei Yu, Chao Cao, Hanqi Jiang, Hanxu Chen, Yiwei Li, Junhao Chen, Huawen Hu, Yihen Liu, Huaqin Zhao, Shaochen Xu, Haixing Dai, Lin Zhao, Ruidong Zhang, Wei Zhao,...

  72. [89]

    Zhe Zhou, Cong Li, Xuechao Wei, Xiaoyang Wang, and Guangyu Sun. 2023. GNNear: Accelerating Full-Batch Training of Graph Neural Networks with near-Memory Processing. In Proceedings of the International Conference on Par- allel Architectures and Compilation Techniques (Chicago, ...

  73. [90]

    Zhe Zhou, Cong Li, Fan Yang, and Guangyu Sun. 2023. DIMM-Link: Enabling Efficient Inter-DIMM Communication for Near-Memory Processing. In 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . 302–316. https://doi.org/10.1109/HPCA56546.2023.10071005

  74. [91]

    Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung- Gon Chun. 2022. Orca: A Distributed Serving System for Transformer-Based Generative Models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). USENIX Association, Carlsbad, CA...

  75. [95]

    In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA)

    UM-PIM: DRAM-based PIM with Uniform & Shared Memory Space. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). 644–659. https://doi.org/10.1109/ISCA59077.2024.00053

  76. [99]

    Zixuan Zhou, Xuefei Ning, Ke Hong, Tianyu Fu, Jiaming Xu, Shiyao Li, Yuming Lou, Luning Wang, Zhihang Yuan, Xiuhong Li, Shengen Yan, Guohao Dai, Xiao- Ping Zhang, Yuhan Dong, and Yu Wang. 2024. A Survey on Efficient Inference for Large Language Models. arXiv:2404.14294 [cs.CL]...

  77. [385]

    https://doi.org/10.1109/MICRO50266.2020.00040

  78. [2016]

    InThe 49th Annual IEEE/ACM International Symposium on Microarchitecture (Taipei, Taiwan) (MICRO-49)

    Chameleon: versatile and practical near-DRAM acceleration architecture for large memory systems. InThe 49th Annual IEEE/ACM International Symposium on Microarchitecture (Taipei, Taiwan) (MICRO-49). IEEE Press, Article 50, 13 pages

  79. [2019]

    In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture (Columbus, OH, USA) (MICRO ’52)

    MEDAL: Scalable DIMM based Near Data Processing Accelerator for DNA Seeding Algorithm. In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture (Columbus, OH, USA) (MICRO ’52). Association for Computing Machinery, New York, NY, USA, 587–599. http...

  80. [2022]

    In Proceedings of the 36th International Conference on Neural Informa- tion Processing Systems (New Orleans, LA, USA) (NIPS ’22)

    FLASHATTENTION: fast and memory-efficient exact attention with IO- awareness. In Proceedings of the 36th International Conference on Neural Informa- tion Processing Systems (New Orleans, LA, USA) (NIPS ’22). Curran Associates Inc., Red Hook, NY, USA, Article 1189, 16 pages

  81. [2025]

    arXiv:2412.20166 [cs.AR] https://arxiv.org/abs/2412.20166

    LoL-PIM: Long-Context LLM Decoding with Scalable DRAM-PIM System. arXiv:2412.20166 [cs.AR] https://arxiv.org/abs/2412.20166

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.