REVIEW 4 major objections 5 minor 121 references
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference
T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read CENT, a GPU-free CXL memory system, claims 2.3x higher LLM inference throughput than A100 GPUs while using 2.9x less energy.
desk verdict Solid systems work with one load-bearing gap: no accuracy validation for BF16/Taylor softmax, plus an asymmetric GPU comparison; deserves peer review with major-revision expectations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the hierarchical PIM-PNM memory device. Each CXL device holds 16 GDDR6-PIM chips; each chip has two channels, and every bank in a channel has a near-bank processing unit with a 16-MAC reduction tree that reads operands directly from the bank at 1 GHz, giving the device 512 TB/s of internal bandwidth. Because near-bank units only do MACs, a second level, PNM units with accumulators, reduction trees, exponent accelerators, and RISC-V cores, finishes the non-MAC operations through a 64 KB shared buffer. Arithmetic uses BF16 values, and exponentiation inside the PNM units uses a 10-order Taylor-series approximation. This hierarchy, plus the CXL communication primitives, is what lets a transformer block run end-to-end inside memory without a GPU or host compute.
What would settle it
Run the paper's provided CENT simulator on Llama2-70B with the published traces, measure end-to-end perplexity on a held-out text set, and compare against FP16 A100 inference at batch 128; if perplexity diverges by more than a preset tolerance, the accuracy premise fails. Separately, reproduce the 2.3x throughput and 2.9x energy numbers with the supplied scripts; if they do not reproduce, the headline comparison fails.
Extended reading notes
Core claim
The paper's central claim is that a GPU-free server built from CXL-connected GDDR6-PIM devices can outperform a four-GPU A100 server for LLM inference by moving arithmetic to where the weights live. CENT maps the arithmetic of a transformer block onto two levels: near-bank processing units inside each DRAM bank perform the multiply-accumulate operations that make up over 99% of the arithmetic, while processing-near-memory accelerators and RISC-V cores next to the memory handle softmax, normalization, square roots, and divisions. A CXL 3.0 switch network with broadcast, multicast, gather, send, and receive primitives lets CENT distribute a model across devices using pipeline, tensor, or hybrid parallelism. On Llama2 7B/13B/70B with 4K contexts and maximum supported batch sizes, the paper reports 2.3x geomean throughput, 2.9x lower energy, and 5.2x more tokens per dollar than the GPU baseline; with 32K contexts the decode-throughput advantage grows to 3.3x.
Load-bearing premise
The premise that BF16 near-bank MACs and a 10-order Taylor-series exponential produce transformer output indistinguishable from FP16 GPU inference; the functional simulator checks instruction-level behavior but the paper reports no perplexity, accuracy, or downstream-task comparison, so if output quality degrades, CENT cannot replace GPUs regardless of speed or cost.
Editorial extensions
If this is right
- If CENT's numbers hold, LLM serving can shift from compute-optimized GPUs to memory-optimized CXL-attached PIM devices, with the host CPU only orchestrating and sampling.
- The advantage grows with context length: at 32K contexts the paper measures 3.3x decoding throughput over GPUs, so long-context and reasoning workloads benefit most.
- Prefill stays compute-bound and GPUs remain roughly 2.5x faster there, but because prefill is only about 2% of end-to-end GPU time, the decode-dominated total still favors CENT.
- Scaling is feasible up to about 64 CXL devices per server with one switch and 128 devices with two-level switching, giving a path from small models to 70B-class serving without GPUs.
- TCO gains are substantial: CENT's owned and rental costs are about 2.5x lower per hour, which is why tokens per dollar reach 5.2x.
Reading between the lines
- If the unverified BF16/Taylor arithmetic preserves model quality, a natural next step is aggressive weight quantization, because the same memory-bound argument suggests CENT would tolerate it at least as well as GPUs.
- The cost comparison is a single-server figure: in a multi-tenant fleet, sharing the host CPU and CXL switch across racks could improve or erode the 5.2x tokens-per-dollar number, so the headline TCO claim is not automatically a fleet-level bound.
- The manuscript reports 2.9x lower energy in the abstract and results but 2.3x lower energy in the conclusion; this discrepancy is not flagged in the paper and should be resolved before relying on either number.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. CENT is a proposed GPU-free server for LLM inference built from CXL-attached devices containing GDDR6-PIM channels and near-memory PNM units. The paper describes the hierarchical PIM-PNM microarchitecture, a CXL 3.0-based network with send/receive/broadcast/multicast primitives, mappings for pipeline, tensor, and hybrid parallelism, and an evaluation of Llama2-7B/13B/70B using a modified Ramulator2 simulator plus power and cost models. Against a vLLM baseline on A100 GPUs, the authors report 2.3x higher end-to-end throughput, 2.9x lower energy, and 5.2x more tokens per dollar. The artifact is open source and includes trace generation, simulation, power modeling, and figure-generation scripts.
Significance. If the reported results hold, CENT would be a significant data point showing that memory-centric hardware can serve memory-bound LLM decoding more efficiently than GPUs. The paper's strengths include the open-source artifact, the use of a timing-aware DRAM simulator, the grounding in real PIM prototypes (AiM, UPMEM), the inclusion of CXL-PNM and heterogeneous GPU-PIM baselines, and the scalability exploration to 128 devices. The central limitation is that output fidelity is never validated, so the throughput and TCO claims currently refer to a system whose tokens are not shown to be as useful as GPU-generated tokens.
major comments (4)
- [§4.2, §6, Appendix A.3.3] The numerical fidelity of the proposed datapath is load-bearing and is not evaluated. Section 4.2 specifies BF16 near-bank MACs, BF16 reduction trees that reduce 16 BF16 inputs to one BF16, BF16 accumulators, and 10th-order Taylor-series exponent accelerators in the PNM units. For Llama2-70B, the GEMV dimension is 8K and attention scores reach 4K elements, so repeated BF16 rounding and reduction-tree accumulation can accumulate substantial error. Section 6 and Appendix A.3.3 state that the functional simulator verifies instruction-level correctness only, that model weights are not required, and that the performance/power simulator does not model real values. No perplexity, logit-level, or downstream-task comparison against FP16/FP32 GPU inference is reported. If these approximations degrade model quality, the 2.3x throughput and 5.2x tokens-per-dollar results describe a device that cannot replace GPUs. Please add accuracy validation (e.g., WikiText perplexity, common QA benchmarks, and logit similarity for representative prompts across all three model sizes), or revise the numerical pipeline to use higher-precision accumulation and show that the approximation error is negligible.
- [§5.1, §7.1, Figure 13] The main throughput comparison uses asymmetric concurrency. The GPU runs vLLM with batch size 128, while CENT uses pipeline parallelism with 32, 40, and 80 concurrent prompts for the 7B, 13B, and 70B models, respectively (Section 5.1 states that CENT does not batch within a pipeline stage; the 'batch' is the pipeline depth). Because reported throughput is aggregated over all concurrent queries, a larger batch directly increases tokens/s. The paper does not show CENT throughput at batch 128 or at matched concurrency levels, nor does it report the maximum feasible batch for the 7B and 13B GPU configurations, which use 1 and 2 GPUs rather than 4. Please report throughput at matched batch sizes and/or provide an SLA-constrained comparison; otherwise the 2.3x geomean does not isolate memory-system efficiency.
- [§6, Table 6, Figure 12] The TCO and tokens-per-dollar results depend on unvalidated cost assumptions. PIM module cost is set to 10x standard DRAM, the CXL controller production volume is set to 3 million units, and the A100 price is set to $10,000 (Section 6, Table 6, Figure 12). These are plausible inputs but they are point estimates, and the 5.2x tokens-per-dollar figure has no sensitivity analysis. For example, if the PIM cost multiplier is actually 20x or the production volume is 300K, the TCO advantage will shrink substantially. Please provide a one-way (or ideally multi-way) sensitivity analysis over these parameters and report the range of tokens-per-dollar, so readers can see how robust the central TCO claim is.
- [§4.1, §6] The CXL network evaluation rests on a protocol-level assumption that is not validated. Section 4.1 repurposes a reserved header code in the PBR flit to implement broadcast/multicast and requires write acknowledgements from all destinations. Section 6 models the multicast-capable switch by halving bandwidth and doubling latency of a baseline switch. No CXL-protocol simulation, credit/ordering analysis, or sensitivity study supports this model. Since the TP and hybrid mappings (Figures 9 and 14) rely on frequent broadcast and gather transactions, and the latency speedups in Figure 13(a) depend on this cost model, the authors should either validate the multicast model with a cycle-level CXL switch simulation or show that the results are insensitive to the factor-of-two bandwidth/latency penalty.
minor comments (5)
- [§9] The conclusion states that CENT 'consumes 2.3× less energy,' while the abstract and Section 7.2 report 2.9×; please reconcile the numbers.
- [§5.1, Figure 13] The term 'batch' is used for pipeline depth in the CENT configuration; consider calling it 'concurrent prompts' or 'pipeline width' to avoid confusion with GPU batching.
- [Figure 14(a)] The 16K and 32K context results use a 1TB CENT configuration with 16Gb GDDR6-PIM modules, but Table 4 lists 512GB; this configuration change should be stated in the main text and not only in the caption.
- [Appendix A.3.3] The statement that 'model weights and parameters are not required for this appendix' should be reconciled with the claim that functional correctness is verified; as written, it reinforces the missing accuracy validation.
- [Throughout] Minor typographical issues include 'hierachical' in Section 4, 'Transfmer' in reference [51], and inconsistent use of 'Llama2' versus 'Llama 2'.
Circularity Check
No circularity found: CENT's throughput, energy, and TCO results are simulated from stated models with external GPU baselines, not equivalent by construction to their inputs.
full rationale
I walked the claimed derivation chain for the three headline results. The 2.3x throughput figure comes from a Ramulator2-based performance simulation of CENT (Section 6, Table 4) against vLLM/A100 GPU baselines measured at batch=128; it is a modeled outcome, not a fitted parameter. The 2.9x energy figure is computed from the simulated power model (Section 7.2) with per-component power estimates, and the 5.2x tokens/dollar figure follows from the throughput and an explicitly stated TCO model (Section 6, Tables 4-6, Figure 12) using assumed hardware costs and a 3M-unit volume estimate. None of these models encode the target ratios as inputs; changing the DRAM timing, power, or cost assumptions would change the outputs, so the results are not forced by definition. The only author-overlapping citation is Ramulator2 [67], used as open-source simulation infrastructure; it is code-reproduced and carries no load-bearing theoretical premise. The paper's own artifact note (Appendix A.3.3) that the performance/power simulator 'does not model real values' is a numerical-fidelity limitation (BF16 MACs, BF16 reduction trees, 10th-order Taylor exp are unvalidated against FP16 GPU output quality), which is a correctness/external-validity risk rather than a circularity: it does not make any derived quantity equal its input. No self-definitional, fitted-input, uniqueness-imported, or renaming circularity was found.
Assumptions & free parameters
free parameters (3)
- PIM module cost multiplier =
10x standard DRAM
- CENT device production volume =
3 million units
- A100 GPU price =
$10,000
assumptions (5)
- domain assumption BF16 multipliers and a 10-order Taylor approximation for exponentiation preserve model accuracy.
- ad hoc to paper CXL 3.0 reserved header codes can be repurposed to implement broadcast and multicast without violating the protocol.
- domain assumption Near-bank PUs in a DRAM process can operate at 1 GHz with 16 MAC units and all-bank activation (ACTab) without exceeding power or area budgets.
- ad hoc to paper A CXL switch supporting multicast can be modeled by halving bandwidth and doubling latency of a baseline switch.
- domain assumption DRAM power and timing parameters of GDDR6-PIM match Samsung GDDR6 and AiM current specifications.
Cite this review
Pith. "Pith review of PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference." pith.science (2026). https://pith.science/paper/55LXYXN7
@misc{pith2026250207578,
author = {Pith},
title = {Pith review of: PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference},
year = {2026},
howpublished = {\url{https://pith.science/paper/55LXYXN7}},
note = {Machine review of arXiv:2502.07578}
}
abstract
Large Language Model (LLM) inference uses an autoregressive manner to generate one token at a time, which exhibits notably lower operational intensity compared to earlier Machine Learning (ML) models such as encoder-only transformers and Convolutional Neural Networks. At the same time, LLMs possess large parameter sizes and use key-value caches to store context information. Modern LLMs support context windows with up to 1 million tokens to generate versatile text, audio, and video content. A large key-value cache unique to each prompt requires a large memory capacity, limiting the inference batch size. Both low operational intensity and limited batch size necessitate a high memory bandwidth. However, contemporary hardware systems for ML model deployment, such as GPUs and TPUs, are primarily optimized for compute throughput. This mismatch challenges the efficient deployment of advanced LLMs and makes users pay for expensive compute resources that are poorly utilized for the memory-bound LLM inference tasks. We propose CENT, a CXL-ENabled GPU-Free sysTem for LLM inference, which harnesses CXL memory expansion capabilities to accommodate substantial LLM sizes, and utilizes near-bank processing units to deliver high memory bandwidth, eliminating the need for expensive GPUs. CENT exploits a scalable CXL network to support peer-to-peer and collective communication primitives across CXL devices. We implement various parallelism strategies to distribute LLMs across these devices. Compared to GPU baselines with maximum supported batch sizes and similar average power, CENT achieves 2.3$\times$ higher throughput and consumes 2.9$\times$ less energy. CENT enhances the Total Cost of Ownership (TCO), generating 5.2$\times$ more tokens per dollar than GPUs.
Figures
Figures from the paper (15 more)
Reference graph
Works this paper leans on
-
[1]
URL: https://azure.microsoft.com/en-us/ pricing/calculator/
Azure pricing calculator. URL: https://azure.microsoft.com/en-us/ pricing/calculator/
-
[2]
Enabling cxl memory expansion for in-memory database management systems
Minseon Ahn, Andrew Chang, Donghun Lee, Jongmin Gim, Jungmin Kim, Jaemin Jung, Oliver Rebholz, Vincent Pham, Krishna Malladi, and Yang Seok Ki. Enabling cxl memory expansion for in-memory database management systems. InProceedings of the 18th International Workshop on Data Management on New Hardware , pages 1–5, 2022
2022
-
[3]
Gqa: Training generalized multi- query transformer models from multi-head checkpoints, 2023
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. Gqa: Training generalized multi- query transformer models from multi-head checkpoints, 2023. URL: https://arxiv.org/abs/2305.13245, arXiv:2305.13245
arXiv 2023
-
[4]
Introducing the next generation of claude
Anthropic. Introducing the next generation of claude. URL: https: //www.anthropic.com/news/claude-3-family
-
[5]
Exploiting CXL-based memory for distributed deep learn- ing
Moiz Arif, Kevin Assogba, M Mustafa Rafique, and Sudharshan Vazhkudai. Exploiting CXL-based memory for distributed deep learn- ing. In Proceedings of the 51st International Conference on Parallel Processing, pages 1–11, 2022
2022
-
[6]
Llama 2 70b: An mlperf inference benchmark for large language models
Thomas Atta-fosu. Llama 2 70b: An mlperf inference benchmark for large language models. URL: https://mlcommons.org/2024/03/mlperf- llama2-70b/
2024
-
[7]
Improving image generation with better captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, Wesam Manassra, Prafulla Dhariwal, Casey Chu, Yunxin Jiao, and Aditya Ramesh. Improving image generation with better captions. Computer Science., 2(3):8, 2023. URL: https://cdn.openai.com/papers/dall-e- 3.pdf
2023
-
[8]
144-lane, 72-port, pci express gen 5.0 pex89144 express- fabric platform
Broadcom. 144-lane, 72-port, pci express gen 5.0 pex89144 express- fabric platform. URL: https://www.broadcom.com/products/pcie- switches-bridges/expressfabric/gen5/pex89144
Show all 121 references
-
[9]
The berkeley out-of-order machine (boom): An open-source industry- competitive, synthesizable, parameterized risc-v processor
Christopher Celio, Krste Asanovic, and David Patterson. The berkeley out-of-order machine (boom): An open-source industry- competitive, synthesizable, parameterized risc-v processor. URL: https://riscv.org/wp-content/uploads/2016/01/Wed1345-RISCV- Workshop-3-BOOM.pdf
2016
-
[10]
Patterson, and Krste Asanović
Christopher Celio, Pi-Feng Chiu, Borivoje Nikolic, David A. Patterson, and Krste Asanović. BOOM v2: an open-source out-of-order RISC- V core. Technical Report UCB/EECS-2017-157, EECS Department, University of California, Berkeley, Sep 2017. URL: http://www2.eecs. berkeley.edu/...
2017
-
[11]
Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks
Yu-Hsin Chen, Tushar Krishna, Joel S Emer, and Vivienne Sze. Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks. IEEE journal of solid-state circuits , 52(1):127–138, 2016
2016
-
[12]
Longlora: Efficient fine-tuning of long-context large language models
Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia. Longlora: Efficient fine-tuning of long-context large language models. arXiv preprint arXiv:2309.12307, 2023
2023 arXiv
-
[13]
The true processing in memory accelerator
Fabrice Devaux. The true processing in memory accelerator. In 2019 IEEE Hot Chips 31 Symposium (HCS) , pages 1–24. IEEE Computer Society, 2019
2019
-
[14]
To pim or not for emerging general purpose processing in ddr memory systems
Alexandar Devic, Siddhartha Balakrishna Rai, Anand Sivasubrama- niam, Ameen Akel, Sean Eilert, and Justin Eno. To pim or not for emerging general purpose processing in ddr memory systems. In Proceedings of the 49th Annual International Symposium on Computer Architecture, pages...
2022
-
[16]
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. CoRR, abs/1810.04805, 2018. URL: http://arxiv.org/ abs/1810.04805, arXiv:1810.04805
2018 arXiv
-
[17]
Dram spot price
dramexchange. Dram spot price. URL: https://www.dramexchange. com/
-
[18]
The llama 3 herd of models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Ka- dian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024
2024 arXiv
-
[19]
The inference cost of search disruption – large language model cost analysis
Afzal Ahmad Dylan Patel. The inference cost of search disruption – large language model cost analysis. URL: https://www.semianalysis. com/p/the-inference-cost-of-search-disruption
-
[20]
Nvidia tesla a100 80gb gpu sxm4 deep learning com- puting graphics card oem
ebay. Nvidia tesla a100 80gb gpu sxm4 deep learning com- puting graphics card oem. URL: https://www.ebay.com/itm/ 126596600113?chn=ps&mkevt=1&mkcid=28&srsltid=AfmBOop8- DCL9WiHC15MU05ZikXFveIxl95uEuqd55d5LHBrMjRXNMiwSTg
-
[21]
Pci interface ic
Mouser Electronics. Pci interface ic. URL: https://www.mouser.com/ c/semiconductors/interface-ics/pci-interface-ic/
-
[22]
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017
Stefan Elfwing, Eiji Uchibe, and Kenji Doya. Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017. URL: https://arxiv.org/abs/1702.03118, arXiv:1702. 03118
2017 arXiv
-
[23]
Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design
Hongxiang Fan, Thomas Chau, Stylianos I Venieris, Royson Lee, Alexandros Kouris, Wayne Luk, Nicholas D Lane, and Mohamed S Abdelfattah. Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design. In 2022 55th IEEE/ACM International Symposium o...
2022
-
[24]
NDA: Near-DRAM acceleration architecture lever- aging commodity DRAM devices and standard memory modules
Amin Farmahini-Farahani, Jung Ho Ahn, Katherine Morrow, and Nam Sung Kim. NDA: Near-DRAM acceleration architecture lever- aging commodity DRAM devices and standard memory modules. In 2015 IEEE 21st International Symposium on High Performance Com- puter Architecture (HPCA), pag...
2015
-
[25]
Reinhardt, Adrian M
Jeremy Fowers, Kalin Ovtcharov, Michael Papamichael, Todd Mas- sengill, Ming Liu, Daniel Lo, Shlomi Alkalay, Michael Haselman, Logan Adams, Mahdi Ghandi, Stephen Heil, Prerak Patel, Adam Sapek, Gabriel Weisz, Lisa Woods, Sitaram Lanka, Steven K. Reinhardt, Adrian M. Caulfield,...
2018
-
[26]
Sparsep: Towards ef- ficient sparse matrix vector multiplication on real processing-in- memory architectures
Christina Giannoula, Ivan Fernandez, Juan Gómez Luna, Nectarios Koziris, Georgios Goumas, and Onur Mutlu. Sparsep: Towards ef- ficient sparse matrix vector multiplication on real processing-in- memory architectures. Proceedings of the ACM on Measurement and Analysis of Computi...
2022
-
[27]
Benchmarking a new paradigm: Experimental analysis and characterization of a real processing-in- memory system
Juan Gómez-Luna, Izzat El Hajj, Ivan Fernandez, Christina Giannoula, Geraldo F Oliveira, and Onur Mutlu. Benchmarking a new paradigm: Experimental analysis and characterization of a real processing-in- memory system. IEEE Access, 10:52565–52608, 2022
2022
-
[28]
Evaluating machine learningworkloads on memory-centric comput- ing systems
Juan Gómez-Luna, Yuxin Guo, Sylvan Brocard, Julien Legriel, Remy Cimadomo, Geraldo F Oliveira, Gagandeep Singh, and Onur Mutlu. Evaluating machine learningworkloads on memory-centric comput- ing systems. In 2023 IEEE International Symposium on Performance Analysis of Systems a...
2023
-
[29]
Our next-generation model: Gemini 1.5
Google. Our next-generation model: Gemini 1.5. URL: https://blog.google/technology/ai/google-gemini-next-generation- model-february-2024/
2024
-
[30]
Memory pooling with cxl.IEEE Micro, 43(2):48– 57, 2023
Donghyun Gouk, Miryeong Kwon, Hanyeoreum Bae, Sangwon Lee, and Myoungsoo Jung. Memory pooling with cxl.IEEE Micro, 43(2):48– 57, 2023
2023
-
[31]
Direct access, High-Performance memory disaggregation with DirectCXL
Donghyun Gouk, Sangwon Lee, Miryeong Kwon, and Myoungsoo Jung. Direct access, High-Performance memory disaggregation with DirectCXL. In 2022 USENIX Annual Technical Conference (USENIX ATC 22), pages 287–294, 2022
2022
-
[32]
OliVe: Acceler- ating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
Cong Guo, Jiaming Tang, Weiming Hu, Jingwen Leng, Chen Zhang, Fan Yang, Yunxin Liu, Minyi Guo, and Yuhao Zhu. OliVe: Acceler- ating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization. In Proceedings of the 50th Annual International PIM Is All You Need...
2025
-
[33]
Deepseek-r1: Incentivizing reasoning capability in llms via reinforce- ment learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforce- ment learning. arXiv preprint arXiv:2501.12948, 2025
2025 arXiv
-
[34]
ELSA: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks
Tae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim, Hyunji Choi, Sung Jun Jung, and Jae W Lee. ELSA: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Architectu...
2021
-
[35]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 770–778, 2016
2016
-
[36]
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016
2016 arXiv
-
[37]
Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing
Guseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi, Sanghyeon Lee, Hyungkyu Ham, Gwangsun Kim, Divya Mahajan, and Jongse Park. Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing. In Proceedings of the 29th ACM International Conference on Architectural Sup...
-
[38]
Le, Yonghui Wu, and Zhifeng Chen
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen. GPipe: efficient training of giant neural networks using pipeline parallelism . Curran Associates Inc., Red Hook, NY, USA, 2019
2019
-
[39]
BEACON: Scalable Near-Data-Processing Accelerators for Genome Analysis near Memory Pool with the CXL Support
Wenqin Huangfu, Krishna T Malladi, Andrew Chang, and Yuan Xie. BEACON: Scalable Near-Data-Processing Accelerators for Genome Analysis near Memory Pool with the CXL Support. In 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) , pages 727–743. IEEE, 2022
2022
-
[40]
Floatpim: In-memory acceleration of deep neural network training with high precision
Mohsen Imani, Saransh Gupta, Yeseong Kim, and Tajana Rosing. Floatpim: In-memory acceleration of deep neural network training with high precision. In Proceedings of the 46th International Sympo- sium on Computer Architecture, pages 802–815, 2019
2019
-
[41]
Ultra-efficient processing in-memory for data intensive applications
Mohsen Imani, Saransh Gupta, and Tajana Rosing. Ultra-efficient processing in-memory for data intensive applications. In Proceedings of the 54th Annual Design Automation Conference 2017 , pages 1–6, 2017
2017
-
[42]
Intel xeon gold 6430 processor, 60m cache, 2.10 ghz
Intel. Intel xeon gold 6430 processor, 60m cache, 2.10 ghz. URL: https://www.intel.com/content/www/us/en/products/ sku/231737/intel-xeon-gold-6430-processor-60m-cache-2-10- ghz/specifications.html
-
[43]
Intel ® xeon® gold 6430 processor
Intel. Intel ® xeon® gold 6430 processor. URL: https://www.intel. com/content/www/us/en/products/sku/231737/intel-xeon-gold- 6430-processor-60m-cache-2-10-ghz/specifications.html
-
[44]
CXL-ANNS: Software- Hardware Collaborative Memory Disaggregation and Computation for Billion-Scale Approximate Nearest Neighbor Search
Junhyeok Jang, Hanjin Choi, Hanyeoreum Bae, Seungjun Lee, Miryeong Kwon, and Myoungsoo Jung. CXL-ANNS: Software- Hardware Collaborative Memory Disaggregation and Computation for Billion-Scale Approximate Nearest Neighbor Search. In 2023 USENIX Annual Technical Conference (USEN...
2023
-
[45]
Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings
Norm Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, Clifford Young, Xiang Zhou, Zongwei Zhou, and David A Patterson. Tpu v4: An optically reconfigurable supercomputer for machine learning wi...
2023
-
[46]
Noise-resilient DNN: Tolerating noise in PCM- based AI accelerators via noise-aware training
Sanjay Kariyappa, Hsinyu Tsai, Katie Spoon, Stefano Ambrogio, Pri- tish Narayanan, Charles Mackin, An Chen, Moinuddin Qureshi, and Geoffrey W Burr. Noise-resilient DNN: Tolerating noise in PCM- based AI accelerators via noise-aware training. IEEE Transactions on Electron Devic...
2021
-
[47]
ChatGPT for good? On opportunities and challenges of large language models for education
Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Jürgen Pfeffer, Oleksandra Poquet, Michael Saile...
2023
-
[48]
Rec- nmp: Accelerating personalized recommendation with near-memory processing
Liu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks, Vikas Chandra, Utku Diril, Amin Firoozshahian, Kim Hazelwood, Bill Jia, Hsien-Hsin S Lee, Meng Li, Bert Maher, Dheevatsa Mudigere, Maxim Naumov, Martin Schatz, Mikhail Smelyanskiy, Xiaodong Wang, Bran- don Reagen, Carole...
2020
-
[49]
Moonwalk: Nre optimization in asic clouds.ACM SIGARCH Computer Architecture News, 45(1):511–526, 2017
Moein Khazraee, Lu Zhang, Luis Vega, and Michael Bedford Taylor. Moonwalk: Nre optimization in asic clouds.ACM SIGARCH Computer Architecture News, 45(1):511–526, 2017
2017
-
[50]
Aquabolt-XL HBM2-PIM, LPDDR5-PIM with in-memory processing, and AXDIMM with accel- eration buffer
Jin Hyun Kim, Shin-Haeng Kang, Sukhan Lee, Hyeonsu Kim, Yuh- wan Ro, Seungwon Lee, David Wang, Jihyun Choi, Jinin So, YeonGon Cho, Kyomin Sohn, and Nam Sung Kim. Aquabolt-XL HBM2-PIM, LPDDR5-PIM with in-memory processing, and AXDIMM with accel- eration buffer. IEEE Micro, 42(3...
2022
-
[51]
Samsung PIM/PNM for Transfmer Based AI: Energy Efficiency on PIM/PNM Cluster
Jin Hyun Kim, Yuhwan Ro, Jinin So, Sukhan Lee, Shin-haeng Kang, YeonGon Cho, Hyeonsu Kim, Byeongho Kim, Kyungsoo Kim, Sangsoo Park, Jin-Seong Kim, Sanghoon Cha, Won-Jo Lee, Jin Jung, Jong- Geon Lee, Jieun Lee, JoonHo Song, Seungwon Lee, Jeonghyeon Cho, Jaehoon Yu, and Kyomin S...
2023
-
[52]
A 1ynm 1.25v 8gb 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep learning application
Daehan Kwon, Seongju Lee, Kyuyoung Kim, Sanghoon Oh, Joonhong Park, Gi-Moon Hong, Dongyoon Ka, Kyudong Hwang, Jeongje Park, Kyeongpil Kang, Jungyeon Kim, Junyeol Jeon, Nahsung Kim, Yongkee Kwon, Vladimir Kornijcuk, Woojae Shin, Jongsoon Won, Minkyu Lee, Hyunha Joo, Haerang Cho...
2023
-
[53]
Maeri: Enabling flexible dataflow mapping over dnn accelerators via recon- figurable interconnects
Hyoukjun Kwon, Ananda Samajdar, and Tushar Krishna. Maeri: Enabling flexible dataflow mapping over dnn accelerators via recon- figurable interconnects. ACM SIGPLAN Notices, 53(2):461–475, 2018
2018
-
[54]
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles, p...
2023
-
[55]
Memory-centric computing with sk hynix’s domain-specific memory
Yongkee Kwon, Guhyun Kim, Nahsung Kim, Woojae Shin, Jongsoon Won, Hyunha Joo, Haerang Choi, Byeongju An, Gyeongcheol Shin, Dayeon Yun, Jeongbin Kim, Changhyun Kim, Ilkon Kim, Jaehan Park, Chanwook Park, Yosub Song, Byeongsu Yang, Hyeongdeok Lee, Seungyeong Park, Wonjun Lee, Se...
2023
-
[56]
System architecture and software stack for GDDR6-AiM
Yongkee Kwon, Kornijcuk Vladimir, Nahsung Kim, Woojae Shin, Jong- soon Won, Minkyu Lee, Hyunha Joo, Haerang Choi, Guhyun Kim, Byeongju An, Jeongbin Kim, Jaewook Lee, Ilkon Kim, Jaehan Park, Chanwook Park, Yosub Song, Byeongsu Yang, Hyungdeok Lee, Seho Kim, Daehan Kwon, Seongju...
2022
-
[57]
25.4 a 20nm 6gb function-in-memory DRAM, based on HBM2 with a 1.2 tflops pro- grammable computing unit using bank-level parallelism, for machine learning applications
Young-Cheon Kwon, Suk Han Lee, Jaehoon Lee, Sang-Hyuk Kwon, Je Min Ryu, Jong-Pil Son, O Seongil, Hak-Soo Yu, Haesuk Lee, Soo Young Kim, Youngmin Cho, Jin Guk Kim, Jongyoon Choi, Hyun- Sung Shin, Jin Kim, BengSeng Phuah, HyoungMin Kim, Myeong Jun Song, Ahn Choi, Daeho Kim, SooY...
2021
-
[58]
Improving in-memory database operations with acceleration DIMM (AxDIMM)
Donghun Lee, Jinin So, Minseon Ahn, Jong-Geon Lee, Jungmin Kim, Jeonghyeon Cho, Rebholz Oliver, Vishnu Charan Thummala, Ravi shankar JV, Sachin Suresh Upadhya, Donghun Lee, Jinin So, Min- seon Ahn, Jong-Geon Lee, Jungmin Kim, Jeonghyeon Cho, Rebholz Oliver, Vishnu Charan Thumm...
2022
-
[59]
Using machine learning to increase yield and lower packaging costs
Melvin Lee. Using machine learning to increase yield and lower packaging costs. URL: https://semiengineering.com/using-machine- learning-to-increase-yield-and-lower-packaging-costs/
-
[60]
A 1ynm 1.25 v 8gb, 16gb/s/pin gddr6-based accelerator-in-memory supporting 1tflops mac operation and various activation functions for deep-learning applications
Seongju Lee, Kyuyoung Kim, Sanghoon Oh, Joonhong Park, Gimoon Hong, Dongyoon Ka, Kyudong Hwang, Jeongje Park, Kyeongpil Kang, Jungyeon Kim, Junyeol Jeon, Nahsung Kim, Yongkee Kwon, Kornijcuk Vladimir, Woojae Shin, Jongsoon Won, Minkyu Lee, Hyunha Joo, Haerang Choi, Jaewook Lee...
2022
-
[61]
Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Ra- jadnya, Scott Lee, Ishwar Agarwal, Mark D
Huaicheng Li, Daniel S Berger, Lisa Hsu, Daniel Ernst, Pantea Zar- doshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Huaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Ra- jadnya, Scott...
2023
-
[62]
Accelerating distributed reinforcement learning with in-switch computing
Youjie Li, Iou-Jen Liu, Yifan Yuan, Deming Chen, Alexander Schwing, and Jian Huang. Accelerating distributed reinforcement learning with in-switch computing. In Proceedings of the 46th International Symposium on Computer Architecture, pages 279–291, 2019
2019
-
[63]
Specification
Compute Express Link ™. Specification. URL: https://www. computeexpresslink.org/
-
[64]
Deepseek-v3 technical report
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024
2024 arXiv
-
[65]
Enmc: Extreme near-memory classification via approximate screening
Liu Liu, Jilan Lin, Zheng Qu, Yufei Ding, and Yuan Xie. Enmc: Extreme near-memory classification via approximate screening. In MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture, pages 1309–1322, 2021
2021
-
[66]
Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture
Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li, Tao Wang, and Yun Liang. Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture. InMICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture , pages 977– 991, 2021
2021
-
[67]
Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator
Haocong Luo, Yahya Can Tuğrul, F Bostancı, Ataberk Olgun, A Giray Yağlıkçı, and Onur Mutlu. Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator. arXiv preprint arXiv:2308.11030, 2023
2023 arXiv
-
[68]
A Binary-activation, Multi-level Weight RNN and Training Algorithm for ADC-/DAC- free and Noise-resilient Processing-in-memory Inference with eNVM
Siming Ma, David Brooks, and Gu-Yeon Wei. A Binary-activation, Multi-level Weight RNN and Training Algorithm for ADC-/DAC- free and Noise-resilient Processing-in-memory Inference with eNVM. IEEE Transactions on Emerging Topics in Computing , 2023
2023
-
[69]
Dram power calculator
Micron. Dram power calculator. URL: https://www.micron.com/ support/tools-and-utilities/power-calc
-
[70]
Nvidia shipped 3.76m data center gpus in 2023
Jowi Morales. Nvidia shipped 3.76m data center gpus in 2023. URL: https://www.tomshardware.com/tech-industry/nvidia-shipped- 376m-data-center-gpus-in-2023-dominates-business-with-98- revenue-share
2023
-
[71]
Supply chain aware computer architecture
August Ning, Georgios Tziantzioulis, and David Wentzlaff. Supply chain aware computer architecture. In Proceedings of the 50th Annual International Symposium on Computer Architecture , pages 1–15, 2023
2023
-
[72]
184QPS/W 64Mb/mm 2 3D logic-to-DRAM hybrid bonding with process-near-memory engine for recommendation system
Dimin Niu, Shuangchen Li, Yuhao Wang, Wei Han, Zhe Zhang, Yijin Guan, Tianchan Guan, Fei Sun, Fei Xue, Lide Duan, Yuanwei Fang, Hongzhong Zheng, Xiping Jiang, Song Wang, Fengguo Zuo, Yubing Wang, Bing Yu, Qiwei Ren, and Yuan Xie. 184QPS/W 64Mb/mm 2 3D logic-to-DRAM hybrid bond...
2022
-
[73]
Introduction to the nvidia dgx a100 system
NVIDIA. Introduction to the nvidia dgx a100 system. URL: https://docs.nvidia.com/dgx/dgxa100-user-guide/introduction-to- dgxa100.html#power-specifications
-
[74]
Nvidia a100 tensor core gpu
NVIDIA. Nvidia a100 tensor core gpu. URL: https://www.nvidia. com/content/dam/en-zz/Solutions/Data-Center/a100/pdf/nvidia- a100-datasheet-us-nvidia-1758950-r4-web.pdf
-
[75]
Nvidia hgx a100, the most powerful end-to-end ai supercom- puting platform
Nvidia. Nvidia hgx a100, the most powerful end-to-end ai supercom- puting platform. URL: https://www.nvidia.com/content/dam/en- zz/Solutions/Data-Center/HGX/a100-80gb-hgx-a100-datasheet- us-nvidia-1485640-r6-web.pdf
-
[76]
Nvlink and nvlink switch
Nvidia. Nvlink and nvlink switch. URL: https://www.nvidia.com/en- us/data-center/nvlink/
-
[77]
NVIDIA. Ptx isa. URL: https://docs.nvidia.com/cuda/pdf/ptx_isa_8.5. pdf
-
[78]
Fine- grained dram: Energy-efficient dram for extreme bandwidth systems
Mike O’Connor, Niladrish Chatterjee, Donghyuk Lee, John Wilson, Aditya Agrawal, Stephen W Keckler, and William J Dally. Fine- grained dram: Energy-efficient dram for extreme bandwidth systems. In Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitect...
2017
-
[79]
Bureau of Labor Statistics
U.S. Bureau of Labor Statistics. Average energy prices for the united states, regions, census divisions, and selected metro- politan areas. URL: https://www.bls.gov/regions/midwest/data/ averageenergyprices_selectedareas_table.htm
-
[80]
Accelerating neural network inference with processing-in-dram: From the edge to the cloud
Geraldo F Oliveira, Juan Gómez-Luna, Saugata Ghose, Amirali Boroumand, and Onur Mutlu. Accelerating neural network inference with processing-in-dram: From the edge to the cloud. IEEE Micro, 42(6):25–38, 2022
2022
-
[81]
Gpt-4 turbo and gpt-4
OpenAI. Gpt-4 turbo and gpt-4. URL: https://platform.openai.com/ docs/models/gpt-4-turbo-and-gpt-4
-
[82]
Learning to reason with llms
OpenAI. Learning to reason with llms. URL: https://openai.com/ index/learning-to-reason-with-llms/ . PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference ASPLOS ’25, March 30-April 3, 2025, Rotterdam, Netherlands
2025
-
[83]
Video generation models as world simulators
OpenAI. Video generation models as world simulators. URL: https://openai.com/research/video-generation-models-as-world- simulators
- [84]
-
[85]
Cost and yield analysis of multi-die packaging using 2.5 d technology compared to fan-out wafer level packaging
Chet Palesko, Amy Palesko, and E Jan Vardaman. Cost and yield analysis of multi-die packaging using 2.5 d technology compared to fan-out wafer level packaging. In Proceedings of the 5th Electronics System-integration Technology Conference (ESTC) , pages 1–5. IEEE, 2014
2014
-
[86]
Attacc! unleashing the power of pim for batched transformer-based generative model inference
Jaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim, Yongsuk Kwon, Nam Sung Kim, and Jung Ho Ahn. Attacc! unleashing the power of pim for batched transformer-based generative model inference. In Proceedings of the 29th ACM International Conference on Architectural Sup...
-
[87]
Trim: Enhancing processor-memory interfaces with scalable tensor reduction in memory
Jaehyun Park, Byeongho Kim, Sungmin Yun, Eojin Lee, Minsoo Rhu, and Jung Ho Ahn. Trim: Enhancing processor-memory interfaces with scalable tensor reduction in memory. In MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture , pages 268– 281, 2021
2021
-
[88]
An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language Models
Sang-Soo Park, KyungSoo Kim, Jinin So, Jin Jung, Jonggeon Lee, Ky- oungwan Woo, Nayeon Kim, Younghyun Lee, Hyungyo Kim, Yong- suk Kwon, Jinhyun Kim, Jieun Lee, YeonGon Cho, Yongmin Tai, Jeonghyeon Cho, Hoyoung Song, Jung Ho Ahn, and Nam Sung Kim. An LPDDR-based CXL-PNM Platfor...
2024
-
[89]
Association for Computing Machinery.doi:10.1145/3620665. 3640422
-
[90]
Splitwise: Efficient gen- erative llm inference using phase splitting
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. Splitwise: Efficient gen- erative llm inference using phase splitting. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) , pages 118–...
2024
-
[91]
Movie gen: A cast of media foundation models
Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, et al. Movie gen: A cast of media foundation models. arXiv preprint arXiv:2410.13720, 2024
2024 arXiv
-
[92]
Nvidia ada lovelace leaked specifications, die sizes, architecture, cost, and performance analysis
Dylan Patel. Nvidia ada lovelace leaked specifications, die sizes, architecture, cost, and performance analysis. URL: https://www. semianalysis.com/p/nvidia-ada-lovelace-leaked-specifications
-
[93]
FACT: FFN- Attention Co-optimized Transformer Architecture with Eager Cor- relation Prediction
Yubin Qin, Yang Wang, Dazheng Deng, Zhiren Zhao, Xiaolong Yang, Leibo Liu, Shaojun Wei, Yang Hu, and Shouyi Yin. FACT: FFN- Attention Co-optimized Transformer Architecture with Eager Cor- relation Prediction. In Proceedings of the 50th Annual International Symposium on Compute...
2023
-
[94]
Dota: detect and omit weak attentions for scalable transformer acceleration
Zheng Qu, Liu Liu, Fengbin Tu, Zhaodong Chen, Yufei Ding, and Yuan Xie. Dota: detect and omit weak attentions for scalable transformer acceleration. In Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems...
2022
-
[95]
Train short, test long: Attention with linear biases enables input length extrapolation.arXiv preprint arXiv:2108.12409, 2021
Ofir Press, Noah A Smith, and Mike Lewis. Train short, test long: Attention with linear biases enables input length extrapolation.arXiv preprint arXiv:2108.12409, 2021
2021 arXiv
-
[96]
Impala: Algorithm/architecture co-design for in- memory multi-stride pattern matching
Elaheh Sadredini, Reza Rahimi, Marzieh Lenjani, Mircea Stan, and Kevin Skadron. Impala: Algorithm/architecture co-design for in- memory multi-stride pattern matching. In 2020 IEEE international symposium on high performance computer architecture (HPCA) , pages 86–98. IEEE, 2020
2020
-
[97]
8gb gddr6 sgram c-die
Samsung. 8gb gddr6 sgram c-die. URL: https://datasheet.lcsc.com/ lcsc/2204251615_Samsung-K4Z80325BC-HC14_C2920181.pdf
-
[98]
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V Le. Searching for activation functions. arXiv preprint arXiv:1710.05941, 2017
2017 arXiv
-
[99]
Glu variants improve transformer
Noam Shazeer. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020
2002 arXiv
-
[100]
Megatron-lm: Training multi- billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi- billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019
1909 arXiv
-
[101]
Generative ai winds in memory semiconduc- tors, total demand for server drams is declining
Kiwoom Securities. Generative ai winds in memory semiconduc- tors, total demand for server drams is declining. URL: https://www. businesspost.co.kr/BP?command=article_view&num=316574
-
[102]
Ro- Former: Enhanced Transformer with Rotary Position Embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu. Ro- Former: Enhanced Transformer with Rotary Position Embedding. arXiv preprint arXiv:2104.09864, 2021
2021 arXiv
-
[103]
Demystifying CXL Memory with Genuine CXL- Ready Systems and Devices
Yan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper, Ipoom Jeong, Ren Wang, and Nam Sung Kim. Demystifying CXL Memory with Genuine CXL- Ready Systems and Devices. arXiv preprint arXiv:2303.15375, 2023
2023 arXiv
-
[104]
Scaling equations for the accu- rate prediction of cmos device performance from 180nm to 7nm
Aaron Stillmaker and Bevan Baas. Scaling equations for the accu- rate prediction of cmos device performance from 180nm to 7nm. Integration, 58:74–81, 2017. URL: https://www.sciencedirect.com/ science/article/pii/S0167926017300755, doi:10.1016/j.vlsi.2017. 02.002
2017 doi
-
[105]
Sharegpt
ShareGPT Team. Sharegpt. URL: https://sharegpt.com/
-
[106]
Llama 2: Open foun- dation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Alma- hairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu,...
2023 arXiv
-
[107]
Design compiler
Synopsys. Design compiler. concurrent timing, area, power, and test optimization. URL: https://www.synopsys.com/implementation-and- signoff/rtl-synthesis-test/dc-ultra.html
-
[108]
Memory systems and interconnects for scale-out servers
Stavros Volos. Memory systems and interconnects for scale-out servers. Technical report, EPFL, 2015
2015
-
[109]
Spatten: Efficient sparse attention architecture with cascade token and head pruning
Hanrui Wang, Zhekai Zhang, and Song Han. Spatten: Efficient sparse attention architecture with cascade token and head pruning. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), pages 97–110. IEEE, 2021
2021
-
[110]
Accelerating compute by cramming it into dram mem- ory
UPMEM. Accelerating compute by cramming it into dram mem- ory. URL: https://www.upmem.com/nextplatform-com-2019-10-03- accelerating-compute-by-cramming-it-into-dram/
2019
-
[111]
Open release of grok-1
XAI. Open release of grok-1. URL: https://x.ai/blog/grok-os
-
[112]
Sparse attention acceleration with synergistic in-memory prun- ing and on-chip recomputation
Amir Yazdanbakhsh, Ashkan Moradifirouzabadi, Zheng Li, and Mingu Kang. Sparse attention acceleration with synergistic in-memory prun- ing and on-chip recomputation. In 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) , pages 744–762. IEEE, 2022
2022
-
[113]
Die shot of the tu104 gpu used in rtx 2080 cards
Wikipedia. Die shot of the tu104 gpu used in rtx 2080 cards. URL: https://en.wikipedia.org/wiki/Turing_(microarchitecture) #/media/File:Nvidia@12nm@Turing@TU104@GeForce_ RTX_2080@S_TAIWAN_1841A1_PKYN44.000_TU104-400- A1_DSCx7_poly@5xExt.jpg
-
[114]
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Vic- toria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer....
2022 arXiv
-
[115]
Xing, Joseph E
Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P. Xing, Joseph E. Gonzalez, and Ion Stoica. Alpa: Automat- ing inter- and Intra-Operator parallelism for distributed deep learn- ing. In 16th USENIX ...
-
[116]
Root mean square layer normalization
Biao Zhang and Rico Sennrich. Root mean square layer normalization. Advances in Neural Information Processing Systems , 32, 2019. ASPLOS ’25, March 30-April 3, 2025, Rotterdam, Netherlands Yufeng Gu and Alireza Khadem et al
2019
-
[117]
Transpim: A memory-based acceleration via software-hardware co- design for transformer
Minxuan Zhou, Weihong Xu, Jaeyoung Kang, and Tajana Rosing. Transpim: A memory-based acceleration via software-hardware co- design for transformer. In 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA), pages 1071–1085. IEEE, 2022
2022
-
[118]
Sky-Sorter: A Processing-in-Memory Architecture for Large-Scale Sorting
Farzaneh Zokaee, Fan Chen, Guangyu Sun, and Lei Jiang. Sky-Sorter: A Processing-in-Memory Architecture for Large-Scale Sorting. IEEE Transactions on Computers, 72(2):480–493, 2022. Received 24 June 2024; revised 2 October 2024; accepted 27 January 2025
2022
-
[120]
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xu- anzhe Liu, Xin Jin, and Hao Zhang. Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving. arXiv preprint arXiv:2401.09670, 2024
2024 arXiv
-
[2018]
URL: https://doi.org/10.1109/isca
-
[2022]
URL: https://www.usenix.org/conference/ osdi22/presentation/zheng-lianmin
USENIX Association. URL: https://www.usenix.org/conference/ osdi22/presentation/zheng-lianmin
-
[2024]
Association for Computing Machinery.doi:10.1145/3620666. 3651380
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.