Pith. sign in

REVIEW 3 major objections 5 minor 80 references

FlashAccel integrates high-bandwidth flash into GPUs so that capacity, not HBM size, sets LLM decode throughput, delivering 2.54× tokens per GPU under a 100 ms latency budget.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Six HBF stacks plus FlashAccel co-design deliver 2.54× decode throughput and 1.93× energy efficiency per GPU versus HBM-only under a 100 ms latency constraint.

T0 review reviewed 2026-07-14 challenge →

load-bearing objection Practical HBF-GPU co-design that fixes the three real blockers for LLM serving; 2.5× is sim-only but the mechanisms and ablations are concrete and useful. the 3 major comments →

arxiv 2607.10186 v1 pith:EHHJYOH6 submitted 2026-07-11 cs.AR

FlashAccel: Leveraging High-Bandwidth Flash for High-Throughput LLM Inference

classification cs.AR
keywords High Bandwidth FlashLLM InferenceKV CacheHeterogeneous MemoryGPU ArchitectureData LayoutPrefetch
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

LLM decode is memory-capacity limited: model weights and KV caches outgrow HBM, which caps batch size, forces early KV eviction, and pushes systems toward multi-GPU setups with communication overhead. High-bandwidth flash (HBF) offers roughly 8× denser capacity at HBM-class bandwidth, yet its multi-microsecond read latency, need for thousands of planes to hit peak bandwidth, and lack of heterogeneous management make it unusable out of the box. FlashAccel co-designs the stack—SRAM prefetch buffers on every plane and base die, execution-order weight layout and hyper-page-aware KV write/offload/read policies, plus an FTL-free storage layer and programming model that expose HBF, HBM and SRAM through one virtual address space. Simulation of CSI and CLI configurations shows that six HBF stacks raise per-GPU decode throughput 2.54× and energy efficiency 1.93× versus an HBM-only GPU under a 100 ms SLO, while also raising multi-turn KV hit rates and cutting recomputation. A sympathetic reader cares because the same capacity bottleneck now throttles every large-model serving stack; if the co-design works, denser, cheaper, more energy-efficient inference becomes practical without proportionally more GPUs.

Core claim

By integrating six HBF stacks into an HBM-based GPU and applying latency-hiding SRAM prefetch, specialized layouts that keep plane load balanced for both static weights and dynamic KV cache, and an HBF-aware programming model, FlashAccel removes the HBM capacity ceiling on batch size. Under a 100 ms decode latency constraint the resulting system delivers average 2.54× higher throughput per GPU and 1.93× higher tokens-per-joule than the pure-HBM baseline.

What carries the argument

The hyper-page abstraction (one page from every plane treated as a single access unit) together with the GroupMmap/GroupArrange/SramPrefetch interfaces. They force all weight and KV traffic to activate the full plane array, hide the 4 µs tR behind computation, and keep plane load balanced even as the active request set changes every step.

Load-bearing premise

The simulated HBF stack (96 planes per die, 4 µs read latency, 768 GB/s per stack, and a 10× endurance gain from relaxed retention) plus the event-driven LLMCompass-based simulator accurately capture real silicon timing, power, and software overheads.

What would settle it

Build or cycle-accurate model a real HBF stack with the stated plane count and latencies; measure end-to-end decode throughput and energy of a Qwen3-235B or LLaMA-405B workload under a 100 ms SLO. If the measured speedup versus an H200 falls well below the reported 2.5×, the central claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • A single GPU can hold far larger models or far more concurrent sessions without multi-GPU communication, cutting both hardware cost and failure domains.
  • Multi-turn agent and long-context workloads retain nearly all prior-turn KV caches, eliminating most recomputation that currently dominates prefill energy.
  • Decode throughput becomes limited by the latency SLO rather than by HBM capacity, so operators can trade latency budget for batch size and tokens-per-second more freely.
  • Energy efficiency (tokens/J) improves even though flash read energy is higher than HBM, because the larger batches amortize fixed costs.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If HBF stacks can be produced at HBM-comparable area and cost, the economic optimum for inference clusters may shift from many small HBM GPUs to fewer high-capacity HBF GPUs, altering interconnect and rack design.
  • The same hyper-page and GroupArrange techniques could be applied to other capacity-bound, read-mostly structures such as embedding tables or retrieval indices, not only transformer KV caches.
  • Relaxed-retention flash for short-lived KV data may become a standard tier in heterogeneous memory hierarchies once the programming model is stabilized.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. FlashAccel proposes a hardware–software co-design that integrates High-Bandwidth Flash (HBF) stacks with HBM-based GPUs for capacity-constrained LLM inference. The system addresses three obstacles—high Flash access latency, low plane-level bandwidth utilization, and heterogeneous resource management—via distributed SRAM caches and a SramPrefetch interface, specialized hyper-page layouts for weights and KV cache (including GroupArrange offloading), and an FTL-free HBF-aware storage layer plus programming model (NandMmap, GroupMmap, GroupWrite, GroupArrange). Evaluated with an event-driven simulator extended from LLMCompass on four models (Qwen3-235B/480B, LLaMA3.1-405B, DeepSeek-V3) under 50 ms and 100 ms SLOs, the paper reports that six HBF stacks (CSI) yield average 2.54× throughput per GPU and 1.93× energy efficiency versus an 8×H200 baseline under a 100 ms latency constraint, with ablations attributing gains primarily to prefetching and secondarily to layout optimizations.

Significance. If the reported gains hold under realistic silicon and software overheads, the work would be a substantial contribution to LLM serving architecture: it shows how to convert Flash’s density advantage into higher decode batch sizes and better multi-turn KV reuse without multi-GPU scaling costs, while remaining compatible with modern GQA/MLA/MoE models. Strengths include explicit endurance and write-bandwidth calculations grounded in published DeepSeek token volumes and P/E-cycle data, multi-model/SLO coverage, and ablations that isolate prefetch, weight layout, and KV layout. The programming model and hyper-page abstractions are concrete and potentially reusable. The central limitation is that all quantitative claims rest on an unvalidated device model and simulator; the result is therefore best read as a carefully argued design study rather than a measured system result.

major comments (3)
  1. §7.1–7.2 and Table 2: The headline 2.54× throughput / 1.93× energy claims are produced entirely by an event-driven simulator (LLMCompass + custom NAND model) whose HBF parameters (96 planes/die, tR = 4 µs and tProg = 75 µs retained after 4× plane-capacity reduction, 768 GB/s per stack, SRAM sizing) are never validated against silicon, RTL, or a public artifact. Because the largest feasible batch sizes under the 100 ms SLO are what drive the reported gains, even moderate optimism in latency, bandwidth utilization, or software overhead would shrink those batch sizes and collapse the headline. The manuscript needs either (a) a sensitivity study that shows the 2.54× remains under plausible 20–30 % degradations of tR, effective bandwidth, and prefetch overlap, or (b) a clear statement that the numbers are upper-bound projections pending silicon validation.
  2. §7.4 and the endurance argument: The claim that KV-cache writes (988 MB/s/GPU) fit a 5-year TBW budget relies on a 10× endurance boost from relaxed retention (from 100K to 1M P/E cycles) plus the assumption that append-only writes and isolated-block allocation eliminate FTL overhead without correctness or wear-leveling cost. The 10× factor is presented as “conservative” relative to literature that claims up to 50×, but no retention-time target, error-rate model, or refresh policy is specified for multi-turn sessions that may last longer than “3 days.” A load-bearing claim of the paper is that Flash is a practical medium for KV cache; this needs a more precise retention/endurance model or an explicit sensitivity bound.
  3. §5.2 and §6.2.2 (GroupArrange / hyper-page packing): The KV-cache layout and offload policy are central to claiming near-peak bandwidth under dynamic active-request sets. The evaluation reports only aggregate latency breakdowns and throughput (Fig. 15); it does not quantify residual plane-load imbalance, offload volume to HBM, or the frequency of plane conflicts after GroupArrange. Without these intermediate metrics it is hard to judge whether the 15 % throughput loss attributed to “disabling KV layout” fully captures the mechanism, or whether HBM pressure under CLI (explicitly noted for 512 KB blocks) reintroduces capacity limits that the abstract claims to remove.
minor comments (5)
  1. Fig. 1 and the model-size trend discussion would benefit from explicit year labels and a clearer distinction between dense and MoE parameter counts, since MoE models dominate the later points.
  2. §3.2: the balls-into-bins 52 % imbalance figure for 100 GB KV cache is useful; stating the exact page size and number of blocks used would make the calculation reproducible.
  3. Table 1: “DP @ 188 KB/token” etc. is dense; a short footnote defining the per-token KV footprint formula would help readers unfamiliar with GQA/MLA sizing.
  4. §7.6 energy: the 8 pJ/bit Flash read energy is taken from a hybrid-bonded prototype; a one-sentence comparison to the HBM3e figure’s measurement conditions would strengthen the tokens/J claim.
  5. Typos / consistency: “t 𝑃𝑟𝑜𝑔” spacing in Table 2; occasional “FlashAccel” vs “FlashAccel” capitalization; arXiv ID year (2607) is future-dated relative to the 2026 citations—worth a consistency check.

Circularity Check

0 steps flagged

No circularity: throughput/energy claims are discrete-event simulation outcomes of concrete layouts and schedules, not algebraic identities or fitted parameters renamed as predictions.

full rationale

FlashAccel is a hardware–software co-design paper whose central quantitative claims (2.54× throughput/GPU and 1.93× energy efficiency under a 100 ms SLO with six HBF stacks) are produced by an event-driven simulator extended from LLMCompass plus a plane-granularity NAND model (Section 7.1, Table 2). The simulator takes as inputs independently stated device parameters (tR = 4 µs, tProg = 75 µs from XL-Flash, 96 planes/die, 768 GB/s per stack, SLC endurance figures) and the authors’ proposed data layouts, GroupArrange offload, and SramPrefetch schedule; it then reports measured latency and throughput under those assumptions. Nothing in the derivation chain reduces by construction to a free parameter or to a self-citation of an unverified uniqueness theorem. Endurance arguments combine published P/E-cycle data with measured token volumes from DeepSeek reports; they are not tautological. Self-citations are absent from the load-bearing path. The paper is therefore self-contained against its own simulation methodology; any remaining concerns are about external validity of the device model, not circularity.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 2 invented entities

The central performance claims rest on a set of device and workload parameters taken from prior flash literature or chosen by the authors, plus standard assumptions about LLM decode being memory-bound and about plane-level parallelism. No new physical entities are postulated beyond the engineered HBF stack itself.

free parameters (4)
  • KV block size = 256 KB
    Chosen as 256 KB to balance plane-load skew against HBM pressure (Section 5.2); directly affects reported bandwidth utilization and HBM footprint.
  • planes per Flash die = 96
    Set to 96 following Lincoln and SanDisk references; determines hyper-page size and required concurrency for peak bandwidth.
  • endurance multiplier from relaxed retention = 10×
    Authors conservatively assume 10× (literature claims up to 50×) to claim 1 M P/E cycles; load-bearing for the claim that KV-cache writes are sustainable.
  • SRAM capacity per stack = 32 MB
    Sized > 2 × peak_bandwidth × read_latency for double buffering; chosen to make the prefetch pipeline work.
axioms (4)
  • domain assumption HBF can be realized with HBM3e-comparable bandwidth (≈4.8 TB/s aggregate) by scaling planes and TSVs while retaining Flash density and non-volatility.
    Stated in Sections 2.3 and 4; underpins the entire capacity-bandwidth premise.
  • domain assumption LLM decode is memory-bound and benefits monotonically from larger batch size up to the latency SLO.
    Used throughout Sections 2 and 7; standard Roofline argument.
  • ad hoc to paper Append-only write patterns of weights and KV cache allow elimination of a full FTL without correctness loss.
    Section 6.1; enables the lightweight storage layer.
  • domain assumption tR = 4 µs and tPROG = 75 µs remain valid even after plane capacity is reduced 4× relative to the XL-Flash reference.
    Table 2 and Section 7.1; conservative but unvalidated for the proposed geometry.
invented entities (2)
  • hyper page no independent evidence
    purpose: Logical access unit that aggregates one page from every plane to expose full inter-plane parallelism.
    Defined in Section 3.2 and used for all weight and KV layouts; engineering abstraction rather than a physical discovery.
  • FlashAccel programming model (NandMmap, GroupMmap, GroupArrange, SramPrefetch) no independent evidence
    purpose: Expose heterogeneous HBM/HBF/SRAM through a unified virtual address space and hide Flash latency.
    Section 6.2; new software interface required by the hardware.

reviewed 2026-07-14 · how reviews work

0 comments
Cite this review

Pith. "Pith review of FlashAccel: Leveraging High-Bandwidth Flash for High-Throughput LLM Inference." pith.science (2026). https://pith.science/paper/EHHJYOH6

@misc{pith2026260710186,
  author       = {Pith},
  title        = {Pith review of: FlashAccel: Leveraging High-Bandwidth Flash for High-Throughput LLM Inference},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EHHJYOH6}},
  note         = {Machine review of arXiv:2607.10186}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large language model (LLM) inference is increasingly limited by the capacity of High-Bandwidth Memory (HBM) in GPUs, as model weights and KV cache grow rapidly. High-Bandwidth Flash (HBF) provides higher capacity than HBM while retaining comparable bandwidth, making it a promising substrate for capacity-constrained LLM inference. However, its inherently high access latency, low bandwidth utilization, and lack of support for heterogeneous resource management make it difficult to integrate HBF into GPUs for LLM inference. We present FlashAccel, a co-designed system that enables efficient LLM inference using HBF. FlashAccel integrates HBF into HBM-based GPUs, providing architectural support to mitigate access latency. It improves bandwidth utilization through specialized data layouts for both model weights and KV cache, and introduces an HBF-aware storage management layer together with a programming model to organize persistent data in HBF and coordinate heterogeneous memory resources at the system level. Experimental results demonstrate that integrating six HBF stacks into the GPU enables FlashAccel to deliver an average improvement of 2.54$\times$ and 1.93$\times$ in throughput per GPU and energy efficiency over the HBM-only GPU under 100ms latency constraint, respectively.

Figures

Figures reproduced from arXiv: 2607.10186 by Chunmeng Dou, Xiaoming Chen, Xiaotian Sun, Xiaoyu Zhang, Xinyu Wang, Xueqi Li, Yalong Xue.

Figure 1
Figure 1. Figure 1: Trends in HBM memory capacity and model size. cache, preventing reuse opportunities in multi-turn inter￾actions and increasing recomputation overhead [33]. Third, due to the limited capacity of a single GPU, storing large models requires multi-GPU systems, which incur significant hardware cost and introduce additional complexity in inter￾GPU communication, task scheduling, and fault tolerance. Prior works … view at source ↗
Figure 2
Figure 2. Figure 2: (a) Benefits of batching. All results are obtained using an 8-layer configuration of Qwen3-32B due to capacity limits (For the full model, H100 only supports batch size less than 8). (b) Benefits of KV cache reuse. cache. For an 8K context length, each request requires about 2GB KV cache, meaning the GPU can support a batch size of at most 8. According to the Roofline model [60], larger batch sizes increas… view at source ↗
Figure 4
Figure 4. Figure 4: Low bandwidth utilization of accessing weights in HBF. (a) Accessing one weight. (b) Accessing multiple weights. Plane-0 Req-1 #1 Plane-1 Plane-2 Plane-3 Req-2 #2 Req-3 #1 Req-1 #2 Req-2 #3 Req-4 #1 Req-1 #3 Req-2 #4 Req-4 #2 Req-2 #1 Req-2 #5 Req-4 #3 Loading #1 of Req-1 and Req-3 encounters plane conflict Req-1 #1 The 1st KV block of request-1 For clarity, only 4 planes are shown Next Inference Step [PI… view at source ↗
Figure 5
Figure 5. Figure 5: Imbalance of accessing KV cache in HBF. The problem is even more complicated for KV cache be￾cause its data layout changes at runtime. At each inference step, the system accesses a unified KV cache constructed from the KV caches of all active requests. Since active re￾quests change over steps, even if the unified KV cache is evenly distributed across planes at one step, the balance does not persist. As req… view at source ↗
Figure 6
Figure 6. Figure 6: Architecture of FlashAccel. (a) Two options for integrating HBF. (b) Architecture of an HBF stack. (c) Design of array die, circuit die, and base die inside HBF. in the same plane, leading to contention analogous to bank conflicts, which further degrades bandwidth utilization. 3.3 System Resource Management The third challenge is system-level support for HBF. Unlike HBM, HBF is non-volatile, so the system … view at source ↗
Figure 7
Figure 7. Figure 7: Data layout for weights in HBF. ❷ Offloading Plane Plane Plane Plane #1 #2 #3 #4 #5 Plane Plane Plane Plane HBM #6 #7 #8 #9 #10 #1 #2 #3 #4 #5 #6 #8 #9 #6 #5 #8 #9 GPU #1 #6 #7 #10 #7 #10 Nonresident KV Blocks (Outside HBF) Hyper Page ❹ Write KV Cache to HBM #7 #10 #11 #13 #12 #14 #1 #2 #3 #4 Hyper Page ❸Load KV Cache to GPU ❺ Write Back to HBF #11 HBF Reused KV Blocks in HBF New generated KV Blocks KV Blo… view at source ↗
Figure 8
Figure 8. Figure 8: Data layout of KV cache in HBF and optimizations for achieving peak bandwidth. Combined with the prefetch mechanism described in Sec￾tion 6.2.3, which loads these weights sequentially, this data layout enables high HBF bandwidth utilization. 5.2 KV Cache in HBF The layout of the KV cache is more complicated than that of static weights because the active request set changes at every inference step. FlashAcc… view at source ↗
Figure 10
Figure 10. Figure 10: Page table mapping (a) after NandMmap, and (b) further after SramPrefetch. Plane Plane Plane Plane Object-1 #1 HBF Addr - #1 HBF Addr - #4 HBF Addr - #5 HBF Addr - #2 1 1 1 1 V. Physical Address Object-2 1 HBF Addr - #3 #2 #3 #4 #5 #1 #3 HBM Addr - #2 #2 Hyper Page #4 #5 Plane Distribution GroupArrange 0 - #2 maps to an overloaded plane->offload to HBM GroupMmap Hyper page in contiguous logical address sp… view at source ↗
Figure 11
Figure 11. Figure 11: Page table mapping after GroupMmap and GroupArrange. 6.2 Programming Model FlashAccel exposes HBF through a programming model com￾prising an interface set and a runtime. After the program invokes these interfaces, the runtime maintains page table mappings from logical addresses to the physical addresses of HBM, HBF, or SRAM. This design exposes heterogeneous memory through a unified virtual address space.… view at source ↗
Figure 12
Figure 12. Figure 12: Proposed prefetch mechanism. (a) Modification to original compute kernels after introducing prefetch. (b) Ker￾nel launch order and prefetch queue. (c) Overlap of prefetch and computation. to data in HBF and a prefetch size as input, and returns an out￾put logical address referring to prefetched data in SRAM for subsequent accesses. FlashAccel also provides SramRelease to release the allocated SRAM space a… view at source ↗
Figure 13
Figure 13. Figure 13: Throughput per GPU comparison among the baseline, CLI, and CSI. Normalized Throughput / GPU [PITH_FULL_IMAGE:figures/full_fig_p010_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Throughput per GPU comparison after extending to 16 GPUs. batch sizes. 8×CLI shows lower throughput than 8×CSI, as the CLI architecture has only 5 HBF stacks, resulting in 16.7% lower bandwidth compared to CSI. Under the strict 50ms SLO, this bandwidth limitation causes it to slightly underperform the baseline on DeepSeek-V3 model and Qwen3-235B model. Under a 100ms SLO, CLI can load more KV cache per ste… view at source ↗
Figure 16
Figure 16. Figure 16: KV cache hit rate and normalized tokens requir￾ing computation (normalized to 8×H200). 7.4 Write Issues Flash suffers from write-side limitations in both bandwidth and endurance [PITH_FULL_IMAGE:figures/full_fig_p011_16.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

80 extracted references · 13 linked inside Pith

  1. [1]

    Nitin Agrawal, Vijayan Prabhakaran, Ted Wobber, John D Davis, Mark Manasse, and Rina Panigrahy. 2008. Design tradeoffs for {SSD} per- formance. In2008 USENIX Annual Technical Conference (USENIX ATC 08)

  2. [2]

    Joshua Ainslie, James Lee-Thorp, Michiel De Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. 2023. Gqa: Training general- ized multi-query transformer models from multi-head checkpoints. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 4895–4901

  3. [3]

    Keivan Alizadeh, Seyed Iman Mirzadeh, Dmitry Belenko, S Khatam- ifard, Minsik Cho, Carlo C Del Mundo, Mohammad Rastegari, and Mehrdad Farajtabar. 2024. Llm in a flash: Efficient large language model inference with limited memory. InProceedings of the 62nd An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 12562–12584

  4. [4]

    Oscar Antepara, Zhengji Zhao, Brian Austin, Nan Ding, Leonid Oliker, Nicholas J Wright, and Samuel Williams. 2025. Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Su- percomputing. InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 888–904

  5. [5]

    Malika Aubakirova, Alex Atallah, Chris Clark, Justin Summerville, and Anjney Midha. 2026. State of AI: An Empirical 100 Trillion Token Study with OpenRouter.arXiv preprint arXiv:2601.10088(2026)

  6. [6]

    badlogic. 2026. pi-mono.https://github.com/badlogic/pi-mono. Ver- sion: main branch, accessed 2026-04-15

  7. [7]

    Yushi Bai, Xin Lv, Jiajie Zhang, Hong Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2024. Longbench: A bilingual, multitask bench- mark for long context understanding. InProceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers). 3119–3137

  8. [8]

    Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020. Longformer: The long-document transformer.arXiv preprint arXiv:2004.05150 (2020)

  9. [9]

    Adithya Bhaskar, Alexander Wettig, Tianyu Gao, Yihe Dong, and Danqi Chen. 2025. Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?arXiv preprint arXiv:2506.17121 (2025)

  10. [10]

    Jalil Boukhobza, Pierre Olivier, Wen Sheng Lim, Liang-Chi Chen, Yun- Shan Hsieh, Shin-Ting Wu, Chien-Chung Ho, Po-Chun Huang, and Yuan-Hao Chang. 2025. A survey on flash-memory storage systems: A host-side perspective.ACM Transactions on Storage21, 3 (2025), 1–59

  11. [11]

    Yu Cai, Saugata Ghose, Erich F Haratsch, Yixin Luo, and Onur Mutlu

  12. [12]

    IEEE105, 9 (2017), 1666–1704

    Error characterization, mitigation, and recovery in flash- memory-based solid-state drives.Proc. IEEE105, 9 (2017), 1666–1704

  13. [13]

    Yu Cai, Gulay Yalcin, Onur Mutlu, Erich F Haratsch, Adrian Cristal, Os- man S Unsal, and Ken Mai. 2012. Flash correct-and-refresh: Retention- aware error management for increased flash memory lifetime. In2012 IEEE 30th International Conference on Computer Design (ICCD). IEEE, 94–101

  14. [14]

    Wanik Cho, Chanhui Jeong, Jongwoo Kim, Jongseok Jung, Keunseon Ahn, Jayoon Goo, Sangkyu Lee, Kayoung Cho, Tei Cho, Dauni Kim, Gwan Park, Yushin Ahn, Sooyeol Chai, Gwihan Ko, Sunyoung Jung, Eunwoo Jo, Taehun Park, Jinhyun Ban, Cheoljoong Park, Jae Hyun Park, Sanghoon Oh, Sojin Jeong, Youngjun Kwak, Kyungsoo Jeong, Jinyeop Kim, Minchol Shin, Eunho Yang, Tai...

  15. [15]

    Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

  16. [16]

    Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in neural information processing systems35 (2022), 16344–16359

  17. [17]

    DeepSeek-AI. 2025.Day 6: One More Thing, DeepSeek-V3/R1 Inference System Overview.https://github.com/deepseek-ai/open-infra- index/blob/main/202502OpenSourceWeek/day_6_one_more_thing_ deepseekV3R1_inference_system_overview.mdGitHub repository

  18. [18]

    Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

    DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huaj...

  19. [19]

    Kaustubh Deshpande, Ved Sirdeshmukh, Johannes Baptist Mols, Lifeng Jin, Ed-Yeremai Hernandez-Cardona, Dean Lee, Jeremy Kritz, Willow E Primack, Summer Yue, and Chen Xing. 2025. Multichallenge: A re- alistic multi-turn conversation evaluation benchmark challenging to frontier llms. InFindings of the Association for Computational Linguis- tics: ACL 2025. 18...

  20. [20]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Ka- dian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schel- ten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony S. Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aur’elien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Rozière,...

  21. [21]

    Zehao Fan, Garrett Gagnon, Zhenyu Liu, and Liu Liu. 2025. Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM. arXiv:2505.05772 [cs.CL]https://arxiv.org/abs/2505.05772

  22. [22]

    Raja Gond, Nipun Kwatra, and Ramachandran Ramjee. 2025. Token- Weave: Efficient Compute-Communication Overlap for Distributed LLM Inference.https://arxiv.org/abs/2505.11329

  23. [23]

    Sylvain Gugger, Lysandre Debut, Thomas Wolf, Philipp Schmid, Zachary Mueller, Sourab Mangrulkar, Marc Sun, and Benjamin Bossan

  24. [24]

    Accelerate: Training and inference at scale made simple, efficient and adaptable.https://github.com/huggingface/accelerate

  25. [25]

    Aayush Gupta, Youngjae Kim, and Bhuvan Urgaonkar. 2009. DFTL: a flash translation layer employing demand-based selective caching of page-level address mappings.Acm Sigplan Notices44, 3 (2009), 229–240

  26. [26]

    Minho Ha, Euiseok Kim, and Hoshik Kim. 2026. H 3: Hybrid Archi- tecture Using High Bandwidth Memory and High Bandwidth Flash for Cost-Efficient LLM Inference.IEEE Computer Architecture Letters (2026)

  27. [27]

    Guseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi, Sanghyeon Lee, Hyungkyu Ham, Gwangsun Kim, Divya Mahajan, and Jongse Park. 2024. Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing. InProceedings of the 29th ACM International Confer- ence on Architectural Support for Programming Languages and Operat- ing Systems, Volume 3. 722–737

  28. [28]

    Po-Kai Hsu, Weihong Xu, Qunyou Liu, Tajana Rosing, and Shimeng Yu. 2026. HAVEN: High-Bandwidth Flash Augmented Vector Engine for Large-Scale Approximate Nearest-Neighbor Search Acceleration. arXiv preprint arXiv:2603.01175(2026)

  29. [29]

    Yitao Hu, Xiulong Liu, Guotao Yang, Linxuan Li, Kai Zeng, Zhixin Zhao, Sheng Chen, Laiping Zhao, Wenxin Li, and Keqiu Li. 2025. TightLLM: Maximizing throughput for LLM inference via adaptive offloading policy.IEEE Trans. Comput.(2025)

  30. [30]

    Hongjing Huang, Zeke Wang, Jie Zhang, Zhenhao He, Chao Wu, Jun Xiao, and Gustavo Alonso. 2021. Shuhai: A tool for benchmarking high bandwidth memory on FPGAs.IEEE Trans. Comput.71, 5 (2021), 1133–1144

  31. [31]

    Hongshin Jun, Jinhee Cho, Kangseol Lee, Ho-Young Son, Kwiwook Kim, Hanho Jin, and Keith Kim. 2017. Hbm (high bandwidth memory) dram technology and architecture. In2017 IEEE International Memory Workshop (IMW). IEEE, 1–4

  32. [32]

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Ben- jamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361(2020)

  33. [33]

    KIOXIA Corporation. 2022.New Storage Class Memory Solution Accelerates Non-Relational Database Performance: Aerospike® NoSQL Databases Using KIOXIA FL6 Series Enterprise NVMe® SCM SSDs Deliver Heightened Performance Gains versus TLC SSDs. Application Brief Rev. 1.0. KIOXIA Corporation. https://americas.kioxia.com/content/dam/kioxia/en-us/business/ ssd/ent...

  34. [34]

    Apostolos Kokolis, Michael Kuchnik, John Hoffman, Adithya Kumar, Parth Malani, Faye Ma, Zachary DeVito, Shubho Sengupta, Kalyan Saladi, and Carole-Jean Wu. 2025. Revisiting reliability in large-scale machine learning research clusters. In2025 IEEE International Sym- posium on High Performance Computer Architecture (HPCA). IEEE, 1259–1274

  35. [35]

    Toshiyuki Kouchi, Mami Kakoi, Noriyasu Kumazaki, Akio Sugahara, Akihiro Imamoto, Yasufumi Kajiyama, Yuri Terada, Bushnaq Sanad, Naoaki Kanagawa, Takuyo Kodama, Ryo Fukuda, Hiromitsu Ko- mai, Norichika Asaoka, Hidekazu Ohnishi, Ryosuke Isomura, Takaya Handa, Kensuke Yamamoto, Yuki Ishizaki, Yoko Deguchi, Atsushi Okuyama, Junichi Sato, Hiroki Yabe, Hua-Ling...

  36. [36]

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Yu, Joey Gonzalez, Hao Zhang, and Ion Stoica. 2023. vllm: Easy, fast, and cheap llm serving with pagedattention.See https://vllm. ai/(accessed 9 August 2023)(2023)

  37. [37]

    Chae Yeon Lee, Chae Ho Won, Seyeon Jung, Eun Su Jung, Tae Min Choi, Hwa Rim Lee, JinUk Yoo, Songhun Yoon, and Sung Gyu Pyo

  38. [38]

    3D integrated process and hybrid bonding of high bandwidth memory (HBM).Electronic Materials Letters21, 3 (2025), 395–419

  39. [39]

    Jaeyong Lee, Hyeunjoo Kim, Sanghun Oh, Myoungjun Chun, Myung- suk Kim, and Jihong Kim. 2025. Aif: Accelerating on-device llm in- ference using in-flash processing. InProceedings of the 52nd Annual International Symposium on Computer Architecture. 529–543

  40. [40]

    Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the middle: How language models use long contexts.Transactions of the association for computational linguistics12 (2024), 157–173

  41. [41]

    Yixin Luo, Yu Cai, Saugata Ghose, Jongmoo Choi, and Onur Mutlu

  42. [42]

    In2015 31st Symposium on Mass Storage Systems and Technologies (MSST)

    WARM: Improving NAND flash memory lifetime with write- hotness aware retention management. In2015 31st Symposium on Mass Storage Systems and Technologies (MSST). IEEE, 1–14

  43. [43]

    Micron Technology. 2024.HBM3E: Powering the future of AI with high-bandwidth memory.https://www.micron.com/about/blog/ applications/ai/microns-hbm3e-powering-the-future-of-ai-with- high-bandwidth-memoryAccessed: 2026-04-16

  44. [44]

    Vidyabhushan Mohan, Taniya Siddiqua, Sudhanva Gurumurthi, and Mircea R Stan. 2010. How I learned to stop worrying and love flash endurance. In2nd Workshop on Hot Topics in Storage and File Systems (HotStorage 10)

  45. [45]

    NVIDIA Corporation. 2022. NVIDIA H100 GPU.https://resources. nvidia.com/en-us-gpu-resources/h100-datasheet-24306. Accessed: 2026-04-14

  46. [46]

    NVIDIA Corporation. 2024. NVIDIA DGX H200 System Architec- ture.https://www.nvidia.com/en-us/data-center/dgx-h200/. Includes ConnectX-7 InfiniBand networking, implying RDMA capability

  47. [47]

    Jaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim, Yongsuk Kwon, Nam Sung Kim, and Jung Ho Ahn. 2024. Attacc! unleashing the power of pim for batched transformer-based gener- ative model inference. InProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. 103–119

  48. [48]

    Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. InProceedings of the 36th annual acm symposium on user interface software and technology. 1–22

  49. [49]

    Balls into bins

    Martin Raab and Angelika Steger. 1998. “Balls into bins”—A simple and tight analysis. InInternational Workshop on Randomization and Approximation Techniques in Computer Science. Springer, 159–170

  50. [50]

    Ananda Samajdar, Yuhao Zhu, Paul Whatmough, Matthew Mattina, and Tushar Krishna. 2018. Scale-sim: Systolic cnn accelerator simula- tor.arXiv preprint arXiv:1811.02883(2018)

  51. [51]

    SanDisk. 2025.Memory-Centric AI: Sandisk’s High Bandwidth Flash Will Redefine AI Infrastructure.https://www.sandisk.com/ company/newsroom/blogs/2025/memory-centric-ai-sandisks-high- bandwidth-flash-will-redefine-ai-infrastructure[Online]

  52. [52]

    Minseok Seo, Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon, Guhyun Kim, Chanwook Park, Ilkon Kim, Jaehan Park, Jeong- bin Kim, Woojae Shin, Jongsoon Won, Haerang Choi, Kyuyoung Kim, Daehan Kwon, Chunseok Jeong, Sangheon Lee, Yongseok Choi, Wooseok Byun, Seungcheol Baek, Hyuk-Jae Lee, and John Kim. 2024. Ianus: Integrated accelerator based on npu-pim ...

  53. [53]

    Zhihong Shao, Damai Dai, Daya Guo, Bo Liu (Benjamin Liu), Zihan Wang, and Huajian Xin. 2024. Deepseek-v2: A strong, economi- cal, and efficient mixture-of-experts language model.arXiv preprint arXiv:2405.04434(2024)

  54. [54]

    Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.arXiv preprint arXiv:1701.06538(2017)

  55. [55]

    Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023. Flexgen: High-throughput generative inference of large language models with a single gpu. InInternational Conference on Machine Learning. PMLR, 31094–31116

  56. [56]

    Ramesh Sitaraman. 2001. The power of two random choices: A survey of techniques and results. (2001)

  57. [57]

    Weiyi Sun, Mingyu Gao, Zhaoshi Li, Aoyang Zhang, Iris Ying Chou, Jianfeng Zhu, Shaojun Wei, and Leibo Liu. 2025. Lincoln: Real-Time 50˜ 100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash Memory. In2025 IEEE International Sym- posium on High Performance Computer Architecture (HPCA). IEEE, 1734–1750

  58. [58]

    Mojumder, Shi Dong, Rafael Ubal, Xiang Gong, Shane Treadway, Yuhui Bao, Vincent Zhao, José L

    Yifan Sun, Trinayan Baruah, Saiful A. Mojumder, Shi Dong, Rafael Ubal, Xiang Gong, Shane Treadway, Yuhui Bao, Vincent Zhao, José L. Abellán, John Kim, Ajay Joshi, and David Kaeli. 2018. MGSim + MGMark: A Framework for Multi-GPU System Research. arXiv:1811.02884 [cs.DC]https://arxiv.org/abs/1811.02884

  59. [59]

    Jayanth M. Thimmaiah, Ryuji Yamashita, In-Soo Yoon, Jason Li, Cynthia Hsu, Takuya Ariki, Naoki Ookuma, Yosuke Kato, Koichiro Hayashi, Kazuki Yamauchi, Indra K V, Masahiro Kano, Sirisha Bhamidi- pati, Sneha Bhatia, Seema Malhotra, Naoki Ojima, Ella Wu, Zhiyong Yang, Frank W. Tsai, Mathias Bayle, Naoyuki Minami, Yasuyuki Fuji- hara, Kei Kitamura, Tomofumi K...

  60. [60]

    In2026 IEEE International Solid-State Circuits Conference (ISSCC), Vol

    A 2Tb 4b/Cell 6-Plane 3D-Flash Memory with 37.6 Gb/mm 2 Bit Density and> 85MB/s Write Throughput. In2026 IEEE International Solid-State Circuits Conference (ISSCC), Vol. 69. IEEE, 254–256

  61. [61]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need.Advances in neural information processing systems30 (2017)

  62. [62]

    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023. Voyager: An open-ended embodied agent with large language models.arXiv preprint arXiv:2305.16291(2023)

  63. [63]

    Jiahao Wang, Jinbo Han, Xingda Wei, Sijie Shen, Dingyan Zhang, Chenguang Fang, Rong Chen, Wenyuan Yu, and Haibo Chen. 2025. {KVCache} Cache in the Wild: Characterizing and Optimizing {KVCache} Cache at a Large Cloud Provider. In2025 USENIX An- nual Technical Conference (USENIX ATC 25). 465–482

  64. [64]

    Qian Wang, Zhenheng Tang, Zichen Jiang, Nuo Chen, Tianyu Wang, and Bingsheng He. 2025. Agenttaxo: Dissecting and benchmarking token distribution of llm multi-agent systems. InICLR 2025 Workshop on Foundation Models in the Wild

  65. [65]

    Wei Wang, Wen Pan, Tao Xie, and Deng Zhou. 2016. How many MLCs should impersonate SLCs to optimize SSD performance?. In Proceedings of the Second International Symposium on Memory Systems. 238–247

  66. [66]

    Samuel Williams, Andrew Waterman, and David Patterson. 2009. Roofline: an insightful visual performance model for multicore ar- chitectures.Commun. ACM52, 4 (2009), 65–76

  67. [67]

    Kosuke Yanagidaira, Mario Sako, Yasuhiro Hirashima, Yumi Higashi, Yutaka Shimizu, Takeshi Nakano, Yusuke Ochi, Hiroaki Yamada, Nobushi Matsuura, Akihiro Imamoto, Kazuaki Kawaguchi, Koji Tabata, Hiroaki Hoshino, Takeshi Hioka, Shigehito Saigusa, Hiroki Date, Masaki Unno, Jumpei Sato, You Kamata, Takahiro Shimizu, Akio Sug- ahara, Taira Shibuya, Atsushi Oku...

  68. [68]

    Kosuke Yanagidaira, Mario Sako, Yasuhiro Hirashima, Junya Matsuno, Yumi Higashi, Yutaka Shimizu, Akihiro Imamoto, Kazuaki Kawaguchi, Koji Tabata, Takeshi Nakano, Yusuke Ochi, Hiroaki Hoshino, Takeshi , Xinyu Wang, Yalong Xue, Xiaotian Sun, Xiaoyu Zhang, Chunmeng Dou, Xueqi Li, and Xiaoming Chen Hioka, Shigehito Saigusa, Hiroki Date, Masaki Unno, Jumpei Sa...

  69. [69]

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxin Yang, Jingren Zhou, Jingren Zhou, Junyan Lin, Kai Dang, Keqin Bao, Ke-Pei Ya...

  70. [70]

    Qwen3 technical report.arXiv preprint arXiv:2505.09388(2025)

  71. [71]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. InThe eleventh international conference on learning representations

  72. [72]

    Xi Ye, Fangcong Yin, Yinghui He, Joie Zhang, Howard Yen, Tianyu Gao, Greg Durrett, and Danqi Chen. 2025. LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation. In Second Conference on Language Modeling.https://openreview.net/ forum?id=ruWC5LIMSo

  73. [73]

    Zihao Yi, Jiarui Ouyang, Zhe Xu, Yuwen Liu, Tianhao Liao, Haohao Luo, and Ying Shen. 2025. A survey on recent advances in llm-based multi-turn dialogue systems.Comput. Surveys58, 6 (2025), 1–38

  74. [74]

    Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022. Orca: A distributed serving system for {Transformer-Based} generative models. In16th USENIX symposium on operating systems design and implementation (OSDI 22). 521–538

  75. [75]

    Zhongkai Yu, Shengwen Liang, Tianyun Ma, Yunke Cai, Ziyuan Nan, Di Huang, Xinkai Song, Yifan Hao, Jie Zhang, Tian Zhi, Yongwei Zhao, Zidong Du, Xing Hu, Qi Guo, and Tianshi Chen. 2024. Cambricon-llm: A chiplet-based hybrid architecture for on-device inference of 70b llm. In2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1474–1488

  76. [76]

    Hengrui Zhang, August Ning, Rohan Baskar Prabhakar, and David Wentzlaff. 2024. Llmcompass: Enabling efficient hardware design for large language model inference. In2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). IEEE, 1080– 1096

  77. [77]

    Xing, Haotong Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhang- hao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Haotong Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judg- ing llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems36 (2023), 46595–46623

  78. [78]

    Gonzalez, Clark W

    Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark W. Barrett, and Ying Sheng. 2024. Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems37 (2024), 62557–62583

  79. [79]

    Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xu- anzhe Liu, Xin Jin, and Hao Zhang. 2024. {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). 193–210

  80. [80]

    Anqi Zhou, Yu Zhang, Fei Ding, Ziqi Lian, Renxi Jin, Yudong Yang, Qidong Wang, and Liqiang Cao. 2024. Research progress of hybrid bonding technology for three-dimensional integration.Microelectron- ics Reliability155 (2024), 115372

This paper was first reviewed by grok-4.5 on July 14, 2026.