REVIEW 4 major objections 4 minor 89 references
L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference
T0 review · 4 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The paper claims that long-context LLM inference can escape the HBM capacity-bandwidth trade-off by offloading decoding attention and the whole KV cache to DIMM-PIM, with up to 6.1x speedup over HBM-PIM solutions.
desk verdict L3 is a thoughtful, well-scoped attempt to offload decode-attention to DIMM-PIM, but its flagship 6.1x speedup depends on an unvalidated SPD timing spoof that a real memory controller would likely reject; worth a serious referee, not a desk reject. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the zero-latency in-flight re-layout unit on the DIMM's rank buffer chip, backed by a double-buffering stage and deliberately altered timing values reported to the host memory controller so that burst writes start early and the unit has a window to swap bits without stalling the DDR bus. Around it sit two complementary KV-cache mappings, one that distributes the elements of a new K or V vector across chips and fixed banks for broadcast inner products, and one that scatters tokens across banks in burst-sized chunks for outer-product context computation, plus configurable bank-level and rank-level processing elements. The re-layout removes the bit-level mismatch between DRAM chip width and FP16 element width, the mappings remove the element-level mismatch between DDR layout and PIM's need for locality and regularity, and the spoofed timing is what makes the re-layout appear free.
What would settle it
Use a DDR4 memory controller that enforces the JEDEC timing parameters exactly as the DIMM's original SPD describes; if the controller rejects the spoofed values or the early burst data collides with the still-busy re-layout unit, causing corruption or bus contention, L3's zero-latency claim is falsified.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that the memory bottleneck in LLM decode is not general but operation-specific: multi-head attention in the decoding phase needs both large storage for the KV cache and high bandwidth to stream it, and DIMM-PIM can supply both because its capacity and bandwidth scale with the number of plugged-in modules. L3 makes this work by fixing two data-placement mismatches: a re-layout unit on the rank buffer chip rearranges bits in flight during ordinary DDR burst writes so that each FP16 element resides in a single DRAM chip, and two KV-cache mapping schemes place co-processed K and V elements contiguously for score and context computations. Rank-level and bank-level processing units then fuse score, softmax, and context into a pipelined kernel, while a rankset-based communication scheme overlaps PCIe transfers with computation and an adaptive two-sub-batch scheduler overlaps prefilling with decoding to keep both devices busy. The reported result is a long-context inference system whose throughput and batch size scale with host memory rather than with HBM capacity.
Load-bearing premise
The load-bearing premise is that a real DDR4 memory controller will tolerate the memory module reporting deliberately altered timing values, so the module's re-layout unit can swap bits during a write burst without corrupting data or stalling the bus; the paper does not validate this against a real controller or a protocol-level simulation.
Editorial extensions
If this is right
- A GPU server with DIMM-PIM host memory can serve long-context workloads at batch sizes that would run out of HBM, converting spare host capacity into higher GPU utilization.
- Time-between-tokens need not rise when KV caches leave the GPU; the paper reports TBT comparable to GPU-only at the smallest configuration and 29-53% of it at 16 ranksets.
- Scaling capacity alone or bandwidth alone gives little (1.1-1.6x at 8x scale-up), whereas scaling both gives 5.1x, so memory-system scaling must be coordinated.
- Bank-level PUs on DIMM-PIM provide roughly 8x the bandwidth of rank-level DIMM-PIM, which is why the bank-level design is needed to keep up with server GPUs.
- Prefill and decode can run on opposite sub-batches on the two devices, letting the scheduler hide most KV offload and projection/feed-forward behind each other.
Reading between the lines
- Editorial extension: the spoofed-SPD timing trick is the link most likely to break under a stricter memory controller; if protocol-level validation fails, L3's re-layout becomes an on-the-critical-path cost and the 6.1x speedup shrinks by the amount of that cost.
- The same bit-level re-layout idea transfers to other PIM media (GDDR, HBM) and to mixed-precision KV caches, where rank or buffer logic could transpose elements during burst or refresh windows rather than with CPU copies.
- Because L3's scheduler deliberately chunks at most one prefilling request per batch, the approach may combine cleanly with sparse-attention or KV-quantization schemes, which would cut the KV stream further and make DIMM-PIM bandwidth stretch even farther.
- If DIMM-PIM evolves into pooled CXL-attached memory, the rankset load-balancing and two-sub-batch scheduler give a concrete recipe for using pooled capacity without exposing memory latency to the decode loop.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes L3, a hardware-software co-designed system that fully offloads decoding-phase multi-head attention and the entire KV cache to DIMM-based processing-in-memory in a GPU server. It claims three contributions: a zero-latency in-flight re-layout mechanism based on deliberately spoofed DDR4 SPD timing values, KV-cache mapping and cross-level processing units that enable bank-level PIM attention kernels, and a scheduler plus communication-overlap scheme that hides PCIe transfers and minimizes idle bubbles. Evaluation is carried out in a DRAMsim3/AttAcc-based simulator on three LLM models and four real-world traces, reporting up to 6.1x speedup over an HBM-PIM baseline and significantly improved batch sizes.
Significance. If the results hold, the paper targets a real and growing bottleneck: long-context decoding is simultaneously capacity- and bandwidth-bound, and DIMM-PIM offers a modular path to scale both dimensions. The authors correctly isolate decoding MHA as the critical operation, and the layered rank-PU/bank-PU design with two distinct KV mappings is a thoughtful response to the bit-level and element-level layout mismatches. The paper also ships a cycle-accurate simulator, synthesizes the PU logic for area and power estimates, evaluates on real traces, and includes an ablation study. However, the headline speedup rests on the unvalidated SPD-spoofing mechanism, the 'bubble-free' pipeline claim is asserted rather than proven, and the evaluation is closed-loop in the sense that the scheduler models are trained on the same simulator used for measurement. These issues are load-bearing and need to be resolved before the central claims can be considered reliable.
major comments (4)
- [§4.1 (Spoofed timing constraints)] The central enabler of the zero-latency in-flight re-layout is the claim that deliberately reporting a tWL of one cycle and enlarged post-write latencies through SPD is safe and entails 'no DDR protocol violations'. No protocol-level validation is provided. The simulator described in §6.1 is DRAMsim3 plus AttAcc; it does not model a host memory controller reacting to a DIMM whose programmed write latency differs from the physical value by roughly 15 cycles. DDR4 write leveling, mode-register training, and JEDEC minimum timing constraints exist precisely to reject this configuration. If a real controller retrains, stalls, or issues a subsequent command while the re-layout buffer is busy, the 'zero-latency' re-layout becomes a multi-cycle stall and the speedup collapses to the CPU-side transpose cost estimated in §2.5. Please provide a concrete protocol-level analysis or a cycle-accurate MC-plus-DIMM simulation demonstrating that a standards-compliant controller will tolerate the spoofed timing, or revise the claim to remove the zero-latency assumption.
- [§4.3 (Kernel fusion with bubble-free pipelining)] The text states that 'Quantitative analysis of the pipeline execution proves it is bubble-free', but no such analysis appears in the paper. The bubble-free property is load-bearing because it underlies the claim that kernel fusion adds no overhead and that the PIM-side pipeline is fully utilized. The proof must be supplied with explicit assumptions about chunk size, on-chip buffer capacity, softmax unit latency, and the synchronization points between bank PUs, rank PU, and the host. As written, the claim is unsupported.
- [§5.3 and §6.1 (Evaluation methodology)] The scheduler's latency models (Eqs. 5-6) are trained on profiling data collected from the same DRAMsim3-based simulator that is then used to measure end-to-end throughput. This creates a closed loop: the scheduler is fitted to the simulator, and the simulator is used to demonstrate the benefit of the scheduler. Additionally, all baselines are author re-implementations rather than the original published systems. At minimum, please report variance across multiple simulator runs or seeds, validate the prediction models on held-out configurations rather than held-out batches from the same distribution, and clarify which baseline parameters are taken directly from the original papers versus assumed by the authors.
- [§1 and §6.2 (Batch-size claim)] The abstract and introduction claim 'significantly improved batch sizes (up to 14.3x on DGX-A100)', but the evaluation never reports batch sizes or the 14.3x figure. §6.2 only gives qualitative statements such as 'L3 achieves much larger batch sizes compared to the HBM-based baselines'. Please add a quantitative batch-size comparison for the traced workloads and report the configuration that yields the 14.3x number.
minor comments (4)
- [§2.6 and Table 1] The citation for NEO is inconsistent: §2.6 text cites 'NEO [33]' while Table 1 lists 'NEO [40]'. The reference list entry [40] is the NEO paper; please correct the in-text citation.
- [§4.1 (Double buffering)] The relationship between the proposed double buffer and the conventional LRDIMM data-buffer path should be clarified with a timing diagram. The text says the double buffer 'replaces the conventional single-buffer approach' but does not specify what the conventional buffer is or how the re-layout unit interacts with the DDR4 burst timing.
- [§5.1 (Communication hiding)] The claim that 'the transfer of the prefilling KV cache can always be hidden' and that it is 'typically <16% of the Feed-forward latency' is not backed by any figure or table. Please provide supporting data or move this statement to the evaluation.
- [§6.4 (Table 5)] The bank PU area and power are synthesized in a logic process, and the footnote states that a DRAM process would incur 10x area overhead. Please report the resulting per-DIMM area and power overhead so the reader can assess the total hardware cost of the design.
Circularity Check
No significant circularity: L3's architectural claims are self-contained proposals with simulation-based evaluation, not derivations that reduce to their inputs.
full rationale
The paper's central claims—offloading decoding MHA and KV cache to DIMM-PIM, the in-flight re-layout, the KV mapping methods, communication-computation overlap, and the adaptive scheduler—are presented as new designs and evaluated in a simulator built on DRAMsim3 and AttAcc. None of the claimed speedups is obtained by fitting a parameter to the target quantity and then 'predicting' it. The scheduler's latency models are trained on profiled data and validated on a held-out 20% test set, which is standard empirical practice and does not make the end-to-end result circular. The 'spoofed timing constraints' mechanism is an architectural assumption about memory-controller behavior, not a circular derivation; its validity is a correctness risk (unvalidated protocol-level behavior), not a logical reduction of the paper's conclusions to its premises. No load-bearing self-citation chain or imported uniqueness theorem is present. The evaluation is self-contained (same simulator for training and testing), which limits external confirmation but is not circularity under the stated criteria.
Assumptions & free parameters
free parameters (5)
- HBM-PIM baseline aggregate bandwidth =
260.8 TB/s
- DIMM-PIM aggregate bandwidth =
13.0 TB/s
- Rank PU on-chip buffer size =
256 KB
- Scheduler latency model parameters =
not disclosed
- tWL and tWR SPD spoof offsets =
tWL reduced by one cycle; tWR increased
assumptions (5)
- domain assumption DIMM-PIM capacity and bandwidth scale linearly with the number of DIMMs and ranksets.
- domain assumption Bank-level PUs can be integrated into commercial DRAM banks with acceptable overhead.
- ad hoc to paper The host memory controller will behave correctly under deliberately spoofed SPD timing parameters.
- domain assumption PCIe communication can overlap with DIMM-PIM computation without significant interference.
- domain assumption The linear and Random Forest latency models generalize to previously unseen batches.
invented entities (3)
-
Rank PU with re-layout, softmax, and adder units
-
Bank-level PU in each physical DRAM bank
-
Rankset-based communication mode hardware extension
Cite this review
Pith. "Pith review of L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference." pith.science (2026). https://pith.science/paper/ENF2DUYO
@misc{pith2026250417584,
author = {Pith},
title = {Pith review of: L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference},
year = {2026},
howpublished = {\url{https://pith.science/paper/ENF2DUYO}},
note = {Machine review of arXiv:2504.17584}
}
abstract
Large Language Models (LLMs) increasingly require processing long text sequences, but GPU memory limitations force difficult trade-offs between memory capacity and bandwidth. While HBM-based acceleration offers high bandwidth, its capacity remains constrained. Offloading data to host-side DIMMs improves capacity but introduces costly data swapping overhead. We identify that the critical memory bottleneck lies in the decoding phase of multi-head attention (MHA) exclusively, which demands substantial capacity for storing KV caches and high bandwidth for attention computation. Our key insight reveals this operation uniquely aligns with modern DIMM-based processing-in-memory (PIM) architectures, which offers scalability of both capacity and bandwidth. Based on this observation and insight, we propose L3, a hardware-software co-designed system integrating DIMM-PIM and GPU devices. L3 introduces three innovations: First, hardware redesigns resolve data layout mismatches and computational element mismatches in DIMM-PIM, enhancing LLM inference utilization. Second, communication optimization enables hiding the data transfer overhead with the computation. Third, an adaptive scheduler coordinates GPU-DIMM-PIM operations to maximize parallelism between devices. Evaluations using real-world traces show L3 achieves up to 6.1$\times$ speedup over state-of-the-art HBM-PIM solutions while significantly improving batch sizes.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
DDR4 SDRAM LRDIMM
2015. DDR4 SDRAM LRDIMM. [Online]. Avalable: https://www.micron.com/ content/dam/micron/global/secure/spectek/data-sheet/dram/ddr4/spectek- 8gb-ddr4-sdram.pdf
2015
-
[2]
JEDEC Standard: DDR4 SDRAM Load Reduced DIMM (LRDIMM) Design Specification
2015. JEDEC Standard: DDR4 SDRAM Load Reduced DIMM (LRDIMM) Design Specification. [Online]. Avalable: https://www.jedec.org/standards-documents/ docs/module4_20_27
2015
-
[3]
NVIDIA DGX A100
2023. NVIDIA DGX A100. [Online]. Avalable: https://resources.nvidia.com/enus- dgx-systems/dgx-ai
2023
-
[4]
cognitivecomputations/dolphin-r1 · Datasets at Hugging Face
2025. cognitivecomputations/dolphin-r1 · Datasets at Hugging Face. https: //huggingface.co/datasets/cognitivecomputations/dolphin-r1. Referenced April 2025
2025
-
[5]
open-r1/OpenR1-Math-220k · Datasets at Hugging Face
2025. open-r1/OpenR1-Math-220k · Datasets at Hugging Face. https:// huggingface.co/datasets/open-r1/OpenR1-Math-220k. Referenced April 2025
2025
-
[6]
open-r1/OpenThoughts-114k-math · Datasets at Hugging Face
2025. open-r1/OpenThoughts-114k-math · Datasets at Hugging Face. https: //huggingface.co/datasets/open-r1/OpenThoughts-114k-math. Referenced April 2025
2025
-
[7]
Gulavani, Alexey Tumanov, and Ramachandran Ramjee
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024. Taming throughput-latency tradeoff in LLM inference with sarathi-serve. In Proceedings of the 18th USENIX Conference on Operating Systems Design and Implementation (Santa Clara, CA, USA) (OSDI’24). USENIX Association...
2024
-
[8]
Mohammad Alian, Seung Won Min, Hadi Asgharimoghaddam, Ashutosh Dhar, Dong Kai Wang, Thomas Roewer, Adam McPadden, Oliver O’Halloran, Deming Chen, Jinjun Xiong, Daehoon Kim, Wen-mei Hwu, and Nam Sung Kim. 2018. Application-Transparent Near-Memory Processing Architecture with Memory Channel Network. In 2018 51st Annual IEEE/ACM International Symposium on Mi...
arXiv 2018
Show all 89 references
-
[9]
Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H
Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier- Hellstern, Gaurav Mishra, Erica Moreira, Mark Omer...
2023 arXiv
-
[11]
Hadi Asghari-Moghaddam, Young Hoon Son, Jung Ho Ahn, and Nam Sung Kim
-
[12]
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2024. LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding. arXiv:2308.14508 [cs.CL] https://ar...
2024 arXiv
-
[13]
Malladi, Hongzhong Zheng, and Onur Mutlu
Amirali Boroumand, Saugata Ghose, Minesh Patel, Hasan Hassan, Brandon Lu- cia, Rachata Ausavarungnirun, Kevin Hsieh, Nastaran Hajinazar, Krishna T. Malladi, Hongzhong Zheng, and Onur Mutlu. 2019. CoNDA: efficient cache coherence support for near-data accelerators. In Proceedin...
2019
-
[14]
Dan Chen, Haiheng He, Hai Jin, Long Zheng, Yu Huang, Xinyang Shen, and Xiaofei Liao. 2023. MetaNMP: Leveraging Cartesian-Like Product to Acceler- ate HGNNs with Near-Memory Processing. In Proceedings of the 50th Annual International Symposium on Computer Architecture (Orlando,...
2023
-
[15]
Quan Chen, Hailong Yang, Minyi Guo, Ram Srivatsa Kannan, Jason Mars, and Lingjia Tang. 2017. Prophet: Precise QoS Prediction on Non-Preemptive Acceler- ators to Improve Utilization in Warehouse-Scale Computers. SIGARCH Comput. Archit. News 45, 1 (April 2017), 17–32. https://do...
2017
-
[17]
Weihao Cui, Han Zhao, Quan Chen, Ningxin Zheng, Jingwen Leng, Jieru Zhao, Zhuo Song, Tao Ma, Yong Yang, Chao Li, and Minyi Guo. 2021. Enable Simultane- ous DNN Services Based on Deterministic Operator Overlap and Precise Latency Prediction. In SC21: International Conference fo...
2021
-
[18]
Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
-
[19]
DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei F...
2017 arXiv
-
[20]
Fabrice Devaux. 2019. The true Processing In Memory accelerator. In 2019 IEEE Hot Chips 31 Symposium (HCS) . 1–24. https://doi.org/10.1109/HOTCHIPS.2019. 8875680
2019 doi
-
[22]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 [cs.CL] https://arxiv.org/abs/1810.04805
2019 arXiv
-
[23]
Siying Feng, Xin He, Kuan-Yu Chen, Liu Ke, Xuan Zhang, David Blaauw, Trevor Mudge, and Ronald Dreslinski. 2022. MeNDA: a near-memory multi-way merge solution for sparse transposition and dataflows. In Proceedings of the 49th Annual International Symposium on Computer Architect...
2022
-
[24]
Elias Frantar and Dan Alistarh. 2023. SparseGPT: massive language models can be accurately pruned in one-shot. In Proceedings of the 40th International Conference on Machine Learning (Honolulu, Hawaii, USA) (ICML’23). JMLR.org, Article 414, 15 pages
2023
-
[25]
Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, and Pengfei Zuo. 2024. Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention. In 2024 USENIX Annual Technical Conference (USENIX ATC 24)...
2024
-
[26]
Yizhao Gao, Zhichen Zeng, Dayou Du, Shijie Cao, Peiyuan Zhou, Jiaxing Qi, Junjie Lai, Hayden Kwok-Hay So, Ting Cao, Fan Yang, and Mao Yang. 2025. SeerAtten- tion: Learning Intrinsic Sparse Attention in Your LLMs. arXiv:2410.13276 [cs.CL] https://arxiv.org/abs/2410.13276
2025 arXiv
-
[27]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, A...
2017 arXiv
-
[28]
Yufeng Gu, Alireza Khadem, Sumanth Umesh, Ning Liang, Xavier Servot, Onur Mutlu, Ravi Iyer, and Reetuparna Das. 2025. PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference. In Proceedings of the 30th ACM International Conference on Architectural...
2025
-
[29]
Oliveira, and Onur Mutlu
Juan Gómez-Luna, Izzat El Hajj, Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu. 2022. Benchmarking a New Paradigm: Experimental Analysis and Characterization of a Real Processing-in-Memory System. IEEE Access 10 (2022), 52565–52608. https://doi.org/10...
2022
-
[30]
Jiaao He and Jidong Zhai. 2024. FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines. arXiv:2403.11421 [cs.DC] https: //arxiv.org/abs/2403.11421
2024 arXiv
-
[31]
Mingxuan He, Choungki Song, Ilkon Kim, Chunseok Jeong, Seho Kim, Il Park, Mithuna Thottethodi, and T. N. Vijaykumar. 2020. Newton: A DRAM-maker’s Accelerator-in-Memory (AiM) Architecture for Machine Learning. In 2020 53rd Annual IEEE/ACM International Symposium on Microarchite...
2020
-
[32]
Yintao He, Haiyu Mao, Christina Giannoula, Mohammad Sadrosadati, Juan Gómez-Luna, Huawei Li, Xiaowei Li, Ying Wang, and Onur Mutlu. 2025. PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing System. In Proceedings ...
2025
-
[34]
Jie Huang and Kevin Chen-Chuan Chang. 2023. Towards Reasoning in Large Language Models: A Survey. arXiv:2212.10403 [cs.CL] https://arxiv.org/abs/ 2212.10403
2023 arXiv
-
[35]
P. K. Huang, C. Y. Lu, W. H. Wei, Christine Chiu, K. C. Ting, Clark Hu, C.H. Tsai, S. Y. Hou, W. C. Chiou, C. T. Wang, and Douglas Yu. 2021. Wafer Level System Integration of the Fifth Generation CoWoS®-S with High Performance Si Interposer at 2500 mm2. In 2021 IEEE 71st Elect...
2021
-
[36]
Wenqin Huangfu, Xueqi Li, Shuangchen Li, Xing Hu, Peng Gu, and Yuan Xie
-
[37]
Malladi, Andrew Chang, and Yuan Xie
Wenqin Huangfu, Krishna T. Malladi, Andrew Chang, and Yuan Xie. 2023. BEA- CON: Scalable Near-Data-Processing Accelerators for Genome Analysis near Memory Pool with the CXL Support. In Proceedings of the 55th Annual IEEE/ACM International Symposium on Microarchitecture (Chicag...
2023
-
[38]
Son Hyojun, Jonatan Gilbert, Xiangyu Wu, Cho Haeyoon, Shivdikar Kaustubh, Abellán José L., Joshi Ajay, Kaeli David, and Kim John. 2025. PIMnet: A Domain- Specific Network for Efficient Collective Communication in Scalable PIM
2025
-
[40]
Xuanlin Jiang, Yang Zhou, Shiyi Cao, Ion Stoica, and Minlan Yu. 2024. NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference. arXiv:2411.01142 [cs.DC] https://arxiv.org/abs/2411.01142
2024 arXiv
-
[41]
Hongshin Jun, Jinhee Cho, Kangseol Lee, Ho-Young Son, Kwiwook Kim, Hanho Jin, and Keith Kim. 2017. HBM (High Bandwidth Memory) DRAM Technology and Architecture. In 2017 IEEE International Memory Workshop (IMW). 1–4. https: //doi.org/10.1109/IMW.2017.7939084
2017
-
[42]
Hongju Kal, Chanyoung Yoo, and Won Woo Ro. 2023. AESPA: Asynchronous Execution Scheme to Exploit Bank-Level Parallelism of Processing-in-Memory. In Proceedings of the 56th Annual IEEE/ACM International Symposium on Microar- chitecture (Toronto, ON, Canada) (MICRO ’23). Associa...
2023
-
[43]
Katikapalli Subramanyam Kalyan. 2023. A Survey of GPT-3 Family Large Language Models Including ChatGPT and GPT-4. arXiv:2310.12321 [cs.CL] https://arxiv.org/abs/2310.12321
2023 arXiv
-
[44]
Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter, Ramachan- dran Ramjee, and Ashish Panwar
Aditya K. Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter, Ramachan- dran Ramjee, and Ashish Panwar. 2025. POD-Attention: Unlocking Full Prefill- Decode Overlap for Faster LLM Inference. In Proceedings of the 30th ACM In- ternational Conference on Architectural Support for ...
2025
-
[45]
Lee, Meng Li, Bert Maher, Dheevatsa Mudigere, Maxim Naumov, Martin Schatz, Mikhail Smelyanskiy, Xiaodong Wang, Brandon Reagen, Carole-Jean Wu, Mark Hempstead, and Xuan Zhang
Liu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks, Vikas Chandra, Utku Diril, Amin Firoozshahian, Kim Hazelwood, Bill Jia, Hsien-Hsin S. Lee, Meng Li, Bert Maher, Dheevatsa Mudigere, Maxim Naumov, Martin Schatz, Mikhail Smelyanskiy, Xiaodong Wang, Brandon Reagen, Carole-...
2020
-
[47]
Liu Ke, Xuan Zhang, Jinin So, Jong-Geon Lee, Shin-Haeng Kang, Sukhan Lee, Songyi Han, YeonGon Cho, Jin Hyun Kim, Yongsuk Kwon, KyungSoo Kim, Jin Jung, Ilkwon Yun, Sung Joo Park, Hyunsun Park, Joonho Song, Jeonghyeon Cho, Kyomin Sohn, Nam Sung Kim, and Hsien-Hsin S. Lee. 2022. ...
2022
-
[48]
Guhyun Kim, Jinkwon Kim, Nahsung Kim, Woojae Shin, Jongsoon Won, Hyunha Joo, Haerang Choi, Byeongju An, Gyeongcheol Shin, Dayeon Yun, Jeongbin Kim, Changhyun Kim, Ilkon Kim, Jaehan Park, Yosub Song, Byeongsu Yang, Hyeongdeok Lee, Seungyeong Park, Wonjun Lee, Seonghun Kim, Yong...
2024
-
[49]
Heesu Kim, Hanmin Park, Taehyun Kim, Kwanheum Cho, Eojin Lee, Soojung Ryu, Hyuk-Jae Lee, Kiyoung Choi, and Jinho Lee. 2021. GradPIM: A Practical Processing-in-DRAM Architecture for Gradient Descent. In 2021 IEEE Interna- tional Symposium on High-Performance Computer Architectu...
2021
-
[51]
Jin Hyun Kim, Shin-haeng Kang, Sukhan Lee, Hyeonsu Kim, Woongjae Song, Yuhwan Ro, Seungwon Lee, David Wang, Hyunsung Shin, Bengseng Phuah, Jihyun Choi, Jinin So, YeonGon Cho, JoonHo Song, Jangseok Choi, Jeonghyeon Cho, Kyomin Sohn, Youngsoo Sohn, Kwangil Park, and Nam Sung Kim...
2021
-
[52]
Kwiwook Kim and Myeong-jae Park. 2024. Present and Future, Challenges of High Bandwith Memory (HBM). In 2024 IEEE International Memory Workshop (IMW). 1–4. https://doi.org/10.1109/IMW59701.2024.10536972
2024
-
[53]
Hyucksung Kwon, Kyungmo Koo, Janghyeon Kim, Woongkyu Lee, Minjae Lee, Hyungdeok Lee, Yousub Jung, Jaehan Park, Yosub Song, Byeongsu Yang, Haerang Choi, Guhyun Kim, Jongsoon Won, Woojae Shin, Changhyun Kim, Gyeongcheol Shin, Yongkee Kwon, Ilkon Kim, Euicheol Lim, John Kim, and ...
-
[54]
Se Jung Kwon, Jeonghoon Kim, Jeongin Bae, Kang Min Yoo, Jin-Hwa Kim, Bae- seong Park, Byeongwook Kim, Jung-Woo Ha, Nako Sung, and Dongsoo Lee. 2022. AlphaTuning: Quantization-Aware Parameter-Efficient Adaptation of Large-Scale 14 L3: DIMM-PIM Integrated Architecture and Coordi...
2022
-
[55]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Princip...
2023
-
[56]
Yongkee Kwon, Guhyun Kim, Nahsung Kim, Woojae Shin, Jongsoon Won, Hyunha Joo, Haerang Choi, Byeongju An, Gyeongcheol Shin, Dayeon Yun, Jeongbin Kim, Changhyun Kim, Ilkon Kim, Jaehan Park, Chanwook Park, Yosub Song, Byeongsu Yang, Hyeongdeok Lee, Seungyeong Park, Wonjun Lee, Se...
2023
-
[57]
Youngeun Kwon, Yunjae Lee, and Minsoo Rhu. 2019. TensorDIMM: A Practical Near-Memory Processing Architecture for Embeddings and Tensor Operations in Deep Learning. In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture (Columbus, OH, USA) (MICR...
2019
-
[58]
Yongkee Kwon, Kornijcuk Vladimir, Nahsung Kim, Woojae Shin, Jongsoon Won, Minkyu Lee, Hyunha Joo, Haerang Choi, Guhyun Kim, Byeongju An, Jeong- bin Kim, Jaewook Lee, Ilkon Kim, Jaehan Park, Chanwook Park, Yosub Song, Byeongsu Yang, Hyungdeok Lee, Seho Kim, Daehan Kwon, Seongju...
-
[59]
John H. Lau. 2022. Recent Advances and Trends in Multiple System and Heterogeneous Integration With TSV-Less Interposers. IEEE Transactions on Components, Packaging and Manufacturing Technology 12, 8 (2022), 1271–1281. https://doi.org/10.1109/TCPMT.2022.3194374
2022
-
[60]
Dongjae Lee, Bongjoon Hyun, Taehun Kim, and Minsoo Rhu. 2024. PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems. In 2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). 627–642. https://doi.org/10.1109/MICRO61859.2024.00053
2024
-
[61]
Donghun Lee, Jinin So, MINSEON AHN, Jong-Geon Lee, Jungmin Kim, Jeonghyeon Cho, Rebholz Oliver, Vishnu Charan Thummala, Ravi shankar JV, Sachin Suresh Upadhya, Mohammed Ibrahim Khan, and Jin Hyun Kim. 2022. Im- proving In-Memory Database Operations with Acceleration DIMM (AxDI...
2022
-
[62]
Hyungdeok Lee, Guhyun Kim, Dayeon Yun, Ilkon Kim, Yongkee Kwon, and Euicheol Lim. 2024. Cost-Effective LLM Accelerator Using Processing in Memory Technology. In 2024 IEEE Symposium on VLSI Technology and Circuits (VLSI Tech- nology and Circuits). 1–2. https://doi.org/10.1109/V...
2024
-
[63]
Wonbeom Lee, Jungi Lee, Junghwan Seo, and Jaewoong Sim. 2024. InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . USENIX Association, Santa Clar...
2024
-
[64]
In 2022 IEEE Hot Chips 34 Symposium (HCS)
System Architecture and Software Stack for GDDR6-AiM. In 2022 IEEE Hot Chips 34 Symposium (HCS). 1–25. https://doi.org/10.1109/HCS55958.2022.9895629
2022
-
[65]
Jinhao Li, Jiaming Xu, Shan Huang, Yonghua Chen, Wen Li, Jun Liu, Yaoxiu Lian, Jiayi Pan, Li Ding, Hao Zhou, Yu Wang, and Guohao Dai. 2025. Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective. arXiv:2410.04466 [cs.AR] https://arxiv.org/abs/2410.04466
2025 arXiv
-
[66]
Shang Li, Zhiyuan Yang, Dhiraj Reddy, Ankur Srivastava, and Bruce Jacob. 2020. DRAMsim3: A Cycle-Accurate, Thermal-Capable DRAM Simulator.IEEE Comput. Archit. Lett. 19, 2 (July 2020), 106–109. https://doi.org/10.1109/LCA.2020.2973991
2020
-
[67]
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Guangxuan Xiao, and Song Han. 2025. AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration. GetMobile: Mobile Comp. and Comm. 28, 4 (Jan. 2025), 12–17. https://doi.org/10.1145/3714983.3714987
2025
-
[68]
Haifeng Liu, Long Zheng, Yu Huang, Chaoqiang Liu, Xiangyu Ye, Jingrui Yuan, Xiaofei Liao, Hai Jin, and Jingling Xue. 2023. Accelerating Personalized Recom- mendation with Cross-level Near-Memory Processing. In Proceedings of the 50th Annual International Symposium on Computer ...
2023
-
[70]
Cong Li, Zhe Zhou, Size Zheng, Jiaxi Zhang, Yun Liang, and Guangyu Sun
-
[71]
Anirban Nag and Rajeev Balasubramonian. 2021. OrderLight: Lightweight Memory-Ordering Primitive for Efficient Fine-Grained PIM Computations. In MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture (Virtual Event, Greece) (MICRO ’21). Association for Comp...
2021
-
[72]
Jaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim, Yongsuk Kwon, Nam Sung Kim, and Jung Ho Ahn. 2024. AttAcc! Unleashing the Power of PIM for Batched Transformer-based Generative Model Inference. InProceedings of the 29th ACM International Conference on Architectura...
2024
-
[73]
Jaehyun Park, Byeongho Kim, Sungmin Yun, Eojin Lee, Minsoo Rhu, and Jung Ho Ahn. 2021. TRiM: Enhancing Processor-Memory Interfaces with Scal- able Tensor Reduction in Memory. In MICRO-54: 54th Annual IEEE/ACM In- ternational Symposium on Microarchitecture (Virtual Event, Greec...
2021
-
[74]
Sang-Soo Park, KyungSoo Kim, Jinin So, Jin Jung, Jonggeon Lee, Kyoungwan Woo, Nayeon Kim, Younghyun Lee, Hyungyo Kim, Yongsuk Kwon, Jinhyun Kim, Jieun Lee, YeonGon Cho, Yongmin Tai, Jeonghyeon Cho, Hoyoung Song, Jung Ho Ahn, and Nam Sung Kim. 2024. An LPDDR-based CXL-PNM Platf...
2024
-
[75]
Aske Plaat, Annie Wong, Suzan Verberne, Joost Broekens, Niki van Stein, and Thomas Back. 2024. Reasoning with Large Language Models, a Survey. arXiv:2407.11511 [cs.AI] https://arxiv.org/abs/2407.11511
2024
-
[76]
Minseok Seo, Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon, Guhyun Kim, Chanwook Park, Ilkon Kim, Jaehan Park, Jeongbin Kim, Woojae Shin, Jongsoon Won, Haerang Choi, Kyuyoung Kim, Daehan Kwon, Chunseok Jeong, Sangheon Lee, Yongseok Choi, Wooseok Byun, Seungcheol Baek, Hyu...
2024
-
[77]
Lian Liu, Shixin Zhao, Bing Li, Haimeng Ren, Zhaohui Xu, Mengdi Wang, Xiaowei Li, Yinhe Han, and Ying Wang. 2025. Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM. https://arxiv.org/abs/2502.16963
2025 arXiv
-
[78]
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023. FlexGen: high-throughput generative inference of large language models with a single GPU. In Proceedings of the 40th International Confer...
2023
-
[79]
Weiyi Sun, Zhaoshi Li, Shouyi Yin, Shaojun Wei, and Leibo Liu. 2021. ABC-DIMM: alleviating the bottleneck of communication in DIMM-based near-memory pro- cessing with inter-DIMM broadcast. InProceedings of the 48th Annual International Symposium on Computer Architecture (Virtu...
2021
-
[81]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, ...
2017
-
[82]
Ying Wei, Yi Chieh Huang, Haiming Tang, Nithya Sankaran, Ish Chadha, Dai Dai, Olakanmi Oluwole, Vishnu Balan, and Edward Lee. 2023. 9.3 NVLink-C2C: A Coherent Off Package Chip-to-Chip Interconnect with 40Gbps/pin Single- ended Signaling. In 2023 IEEE International Solid-State ...
2023
-
[83]
Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. 2022. ZeroQuant: efficient and affordable post-training quantization for large-scale transformers. In Proceedings of the 36th International Conference on Neural Information Processing Sy...
2022
-
[84]
Gon- zalez, and Ion Stoica
Ying Sheng, Shiyi Cao, Dacheng Li, Coleman Hooper, Nicholas Lee, Shuo Yang, Christopher Chou, Banghua Zhu, Lianmin Zheng, Kurt Keutzer, Joseph E. Gon- zalez, and Ion Stoica. 2023. S-LoRA: Serving Thousands of Concurrent LoRA Adapters. arXiv preprint arXiv:2311.03285 (2023)
2023 arXiv
-
[85]
Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Y. X. Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, and Wangding Zeng. 2025. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Att...
2025 arXiv
-
[86]
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, and Beidi Chen. 2023. H2O: heavy-hitter oracle for efficient generative inference of large language models. In Proceedings o...
2023
-
[87]
Yilong Zhao, Mingyu Gao, Fangxin Liu, Yiwei Hu, Zongwu Wang, Han Lin, Ji Li, He Xian, Hanlin Dong, Tao Yang, Naifeng Jing, Xiaoyao Liang, and Li Jiang
-
[88]
Tianyang Zhong, Zhengliang Liu, Yi Pan, Yutong Zhang, Yifan Zhou, Shizhe Liang, Zihao Wu, Yanjun Lyu, Peng Shu, Xiaowei Yu, Chao Cao, Hanqi Jiang, Hanxu Chen, Yiwei Li, Junhao Chen, Huawen Hu, Yihen Liu, Huaqin Zhao, Shaochen Xu, Haixing Dai, Lin Zhao, Ruidong Zhang, Wei Zhao,...
2024
-
[89]
Zhe Zhou, Cong Li, Xuechao Wei, Xiaoyang Wang, and Guangyu Sun. 2023. GNNear: Accelerating Full-Batch Training of Graph Neural Networks with near-Memory Processing. In Proceedings of the International Conference on Par- allel Architectures and Compilation Techniques (Chicago, ...
2023
-
[90]
Zhe Zhou, Cong Li, Fan Yang, and Guangyu Sun. 2023. DIMM-Link: Enabling Efficient Inter-DIMM Communication for Near-Memory Processing. In 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . 302–316. https://doi.org/10.1109/HPCA56546.2023.10071005
2023
-
[91]
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung- Gon Chun. 2022. Orca: A Distributed Serving System for Transformer-Based Generative Models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). USENIX Association, Carlsbad, CA...
2022
-
[95]
In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA)
UM-PIM: DRAM-based PIM with Uniform & Shared Memory Space. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). 644–659. https://doi.org/10.1109/ISCA59077.2024.00053
2024
-
[99]
Zixuan Zhou, Xuefei Ning, Ke Hong, Tianyu Fu, Jiaming Xu, Shiyao Li, Yuming Lou, Luning Wang, Zhihang Yuan, Xiuhong Li, Shengen Yan, Guohao Dai, Xiao- Ping Zhang, Yuhan Dong, and Yu Wang. 2024. A Survey on Efficient Inference for Large Language Models. arXiv:2404.14294 [cs.CL]...
2024 arXiv
-
[385]
https://doi.org/10.1109/MICRO50266.2020.00040
2020
-
[2016]
InThe 49th Annual IEEE/ACM International Symposium on Microarchitecture (Taipei, Taiwan) (MICRO-49)
Chameleon: versatile and practical near-DRAM acceleration architecture for large memory systems. InThe 49th Annual IEEE/ACM International Symposium on Microarchitecture (Taipei, Taiwan) (MICRO-49). IEEE Press, Article 50, 13 pages
-
[2019]
In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture (Columbus, OH, USA) (MICRO ’52)
MEDAL: Scalable DIMM based Near Data Processing Accelerator for DNA Seeding Algorithm. In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture (Columbus, OH, USA) (MICRO ’52). Association for Computing Machinery, New York, NY, USA, 587–599. http...
-
[2022]
In Proceedings of the 36th International Conference on Neural Informa- tion Processing Systems (New Orleans, LA, USA) (NIPS ’22)
FLASHATTENTION: fast and memory-efficient exact attention with IO- awareness. In Proceedings of the 36th International Conference on Neural Informa- tion Processing Systems (New Orleans, LA, USA) (NIPS ’22). Curran Associates Inc., Red Hook, NY, USA, Article 1189, 16 pages
-
[2025]
arXiv:2412.20166 [cs.AR] https://arxiv.org/abs/2412.20166
LoL-PIM: Long-Context LLM Decoding with Scalable DRAM-PIM System. arXiv:2412.20166 [cs.AR] https://arxiv.org/abs/2412.20166
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.