REVIEW 3 major objections 5 minor 80 references
FlashAccel integrates high-bandwidth flash into GPUs so that capacity, not HBM size, sets LLM decode throughput, delivering 2.54× tokens per GPU under a 100 ms latency budget.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 13:39 UTC pith:EHHJYOH6
load-bearing objection Practical HBF-GPU co-design that fixes the three real blockers for LLM serving; 2.5× is sim-only but the mechanisms and ablations are concrete and useful. the 3 major comments →
FlashAccel: Leveraging High-Bandwidth Flash for High-Throughput LLM Inference
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
By integrating six HBF stacks into an HBM-based GPU and applying latency-hiding SRAM prefetch, specialized layouts that keep plane load balanced for both static weights and dynamic KV cache, and an HBF-aware programming model, FlashAccel removes the HBM capacity ceiling on batch size. Under a 100 ms decode latency constraint the resulting system delivers average 2.54× higher throughput per GPU and 1.93× higher tokens-per-joule than the pure-HBM baseline.
What carries the argument
The hyper-page abstraction (one page from every plane treated as a single access unit) together with the GroupMmap/GroupArrange/SramPrefetch interfaces. They force all weight and KV traffic to activate the full plane array, hide the 4 µs tR behind computation, and keep plane load balanced even as the active request set changes every step.
Load-bearing premise
The simulated HBF stack (96 planes per die, 4 µs read latency, 768 GB/s per stack, and a 10× endurance gain from relaxed retention) plus the event-driven LLMCompass-based simulator accurately capture real silicon timing, power, and software overheads.
What would settle it
Build or cycle-accurate model a real HBF stack with the stated plane count and latencies; measure end-to-end decode throughput and energy of a Qwen3-235B or LLaMA-405B workload under a 100 ms SLO. If the measured speedup versus an H200 falls well below the reported 2.5×, the central claim fails.
If this is right
- A single GPU can hold far larger models or far more concurrent sessions without multi-GPU communication, cutting both hardware cost and failure domains.
- Multi-turn agent and long-context workloads retain nearly all prior-turn KV caches, eliminating most recomputation that currently dominates prefill energy.
- Decode throughput becomes limited by the latency SLO rather than by HBM capacity, so operators can trade latency budget for batch size and tokens-per-second more freely.
- Energy efficiency (tokens/J) improves even though flash read energy is higher than HBM, because the larger batches amortize fixed costs.
Where Pith is reading between the lines
- If HBF stacks can be produced at HBM-comparable area and cost, the economic optimum for inference clusters may shift from many small HBM GPUs to fewer high-capacity HBF GPUs, altering interconnect and rack design.
- The same hyper-page and GroupArrange techniques could be applied to other capacity-bound, read-mostly structures such as embedding tables or retrieval indices, not only transformer KV caches.
- Relaxed-retention flash for short-lived KV data may become a standard tier in heterogeneous memory hierarchies once the programming model is stabilized.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. FlashAccel proposes a hardware–software co-design that integrates High-Bandwidth Flash (HBF) stacks with HBM-based GPUs for capacity-constrained LLM inference. The system addresses three obstacles—high Flash access latency, low plane-level bandwidth utilization, and heterogeneous resource management—via distributed SRAM caches and a SramPrefetch interface, specialized hyper-page layouts for weights and KV cache (including GroupArrange offloading), and an FTL-free HBF-aware storage layer plus programming model (NandMmap, GroupMmap, GroupWrite, GroupArrange). Evaluated with an event-driven simulator extended from LLMCompass on four models (Qwen3-235B/480B, LLaMA3.1-405B, DeepSeek-V3) under 50 ms and 100 ms SLOs, the paper reports that six HBF stacks (CSI) yield average 2.54× throughput per GPU and 1.93× energy efficiency versus an 8×H200 baseline under a 100 ms latency constraint, with ablations attributing gains primarily to prefetching and secondarily to layout optimizations.
Significance. If the reported gains hold under realistic silicon and software overheads, the work would be a substantial contribution to LLM serving architecture: it shows how to convert Flash’s density advantage into higher decode batch sizes and better multi-turn KV reuse without multi-GPU scaling costs, while remaining compatible with modern GQA/MLA/MoE models. Strengths include explicit endurance and write-bandwidth calculations grounded in published DeepSeek token volumes and P/E-cycle data, multi-model/SLO coverage, and ablations that isolate prefetch, weight layout, and KV layout. The programming model and hyper-page abstractions are concrete and potentially reusable. The central limitation is that all quantitative claims rest on an unvalidated device model and simulator; the result is therefore best read as a carefully argued design study rather than a measured system result.
major comments (3)
- §7.1–7.2 and Table 2: The headline 2.54× throughput / 1.93× energy claims are produced entirely by an event-driven simulator (LLMCompass + custom NAND model) whose HBF parameters (96 planes/die, tR = 4 µs and tProg = 75 µs retained after 4× plane-capacity reduction, 768 GB/s per stack, SRAM sizing) are never validated against silicon, RTL, or a public artifact. Because the largest feasible batch sizes under the 100 ms SLO are what drive the reported gains, even moderate optimism in latency, bandwidth utilization, or software overhead would shrink those batch sizes and collapse the headline. The manuscript needs either (a) a sensitivity study that shows the 2.54× remains under plausible 20–30 % degradations of tR, effective bandwidth, and prefetch overlap, or (b) a clear statement that the numbers are upper-bound projections pending silicon validation.
- §7.4 and the endurance argument: The claim that KV-cache writes (988 MB/s/GPU) fit a 5-year TBW budget relies on a 10× endurance boost from relaxed retention (from 100K to 1M P/E cycles) plus the assumption that append-only writes and isolated-block allocation eliminate FTL overhead without correctness or wear-leveling cost. The 10× factor is presented as “conservative” relative to literature that claims up to 50×, but no retention-time target, error-rate model, or refresh policy is specified for multi-turn sessions that may last longer than “3 days.” A load-bearing claim of the paper is that Flash is a practical medium for KV cache; this needs a more precise retention/endurance model or an explicit sensitivity bound.
- §5.2 and §6.2.2 (GroupArrange / hyper-page packing): The KV-cache layout and offload policy are central to claiming near-peak bandwidth under dynamic active-request sets. The evaluation reports only aggregate latency breakdowns and throughput (Fig. 15); it does not quantify residual plane-load imbalance, offload volume to HBM, or the frequency of plane conflicts after GroupArrange. Without these intermediate metrics it is hard to judge whether the 15 % throughput loss attributed to “disabling KV layout” fully captures the mechanism, or whether HBM pressure under CLI (explicitly noted for 512 KB blocks) reintroduces capacity limits that the abstract claims to remove.
minor comments (5)
- Fig. 1 and the model-size trend discussion would benefit from explicit year labels and a clearer distinction between dense and MoE parameter counts, since MoE models dominate the later points.
- §3.2: the balls-into-bins 52 % imbalance figure for 100 GB KV cache is useful; stating the exact page size and number of blocks used would make the calculation reproducible.
- Table 1: “DP @ 188 KB/token” etc. is dense; a short footnote defining the per-token KV footprint formula would help readers unfamiliar with GQA/MLA sizing.
- §7.6 energy: the 8 pJ/bit Flash read energy is taken from a hybrid-bonded prototype; a one-sentence comparison to the HBM3e figure’s measurement conditions would strengthen the tokens/J claim.
- Typos / consistency: “t 𝑃𝑟𝑜𝑔” spacing in Table 2; occasional “FlashAccel” vs “FlashAccel” capitalization; arXiv ID year (2607) is future-dated relative to the 2026 citations—worth a consistency check.
Circularity Check
No circularity: throughput/energy claims are discrete-event simulation outcomes of concrete layouts and schedules, not algebraic identities or fitted parameters renamed as predictions.
full rationale
FlashAccel is a hardware–software co-design paper whose central quantitative claims (2.54× throughput/GPU and 1.93× energy efficiency under a 100 ms SLO with six HBF stacks) are produced by an event-driven simulator extended from LLMCompass plus a plane-granularity NAND model (Section 7.1, Table 2). The simulator takes as inputs independently stated device parameters (tR = 4 µs, tProg = 75 µs from XL-Flash, 96 planes/die, 768 GB/s per stack, SLC endurance figures) and the authors’ proposed data layouts, GroupArrange offload, and SramPrefetch schedule; it then reports measured latency and throughput under those assumptions. Nothing in the derivation chain reduces by construction to a free parameter or to a self-citation of an unverified uniqueness theorem. Endurance arguments combine published P/E-cycle data with measured token volumes from DeepSeek reports; they are not tautological. Self-citations are absent from the load-bearing path. The paper is therefore self-contained against its own simulation methodology; any remaining concerns are about external validity of the device model, not circularity.
Axiom & Free-Parameter Ledger
free parameters (4)
- KV block size =
256 KB
- planes per Flash die =
96
- endurance multiplier from relaxed retention =
10×
- SRAM capacity per stack =
32 MB
axioms (4)
- domain assumption HBF can be realized with HBM3e-comparable bandwidth (≈4.8 TB/s aggregate) by scaling planes and TSVs while retaining Flash density and non-volatility.
- domain assumption LLM decode is memory-bound and benefits monotonically from larger batch size up to the latency SLO.
- ad hoc to paper Append-only write patterns of weights and KV cache allow elimination of a full FTL without correctness loss.
- domain assumption tR = 4 µs and tPROG = 75 µs remain valid even after plane capacity is reduced 4× relative to the XL-Flash reference.
invented entities (2)
-
hyper page
no independent evidence
-
FlashAccel programming model (NandMmap, GroupMmap, GroupArrange, SramPrefetch)
no independent evidence
Cite this review
Pith. "Pith review of FlashAccel: Leveraging High-Bandwidth Flash for High-Throughput LLM Inference." pith.science (2026). https://pith.science/paper/EHHJYOH6
@misc{pith2026260710186,
author = {Pith},
title = {Pith review of: FlashAccel: Leveraging High-Bandwidth Flash for High-Throughput LLM Inference},
year = {2026},
howpublished = {\url{https://pith.science/paper/EHHJYOH6}},
note = {Machine review of arXiv:2607.10186}
}
read the original abstract
Large language model (LLM) inference is increasingly limited by the capacity of High-Bandwidth Memory (HBM) in GPUs, as model weights and KV cache grow rapidly. High-Bandwidth Flash (HBF) provides higher capacity than HBM while retaining comparable bandwidth, making it a promising substrate for capacity-constrained LLM inference. However, its inherently high access latency, low bandwidth utilization, and lack of support for heterogeneous resource management make it difficult to integrate HBF into GPUs for LLM inference. We present FlashAccel, a co-designed system that enables efficient LLM inference using HBF. FlashAccel integrates HBF into HBM-based GPUs, providing architectural support to mitigate access latency. It improves bandwidth utilization through specialized data layouts for both model weights and KV cache, and introduces an HBF-aware storage management layer together with a programming model to organize persistent data in HBF and coordinate heterogeneous memory resources at the system level. Experimental results demonstrate that integrating six HBF stacks into the GPU enables FlashAccel to deliver an average improvement of 2.54$\times$ and 1.93$\times$ in throughput per GPU and energy efficiency over the HBM-only GPU under 100ms latency constraint, respectively.
Figures
Reference graph
Works this paper leans on
-
[1]
Nitin Agrawal, Vijayan Prabhakaran, Ted Wobber, John D Davis, Mark Manasse, and Rina Panigrahy. 2008. Design tradeoffs for {SSD} per- formance. In2008 USENIX Annual Technical Conference (USENIX ATC 08)
2008
-
[2]
Joshua Ainslie, James Lee-Thorp, Michiel De Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. 2023. Gqa: Training general- ized multi-query transformer models from multi-head checkpoints. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 4895–4901
2023
-
[3]
Keivan Alizadeh, Seyed Iman Mirzadeh, Dmitry Belenko, S Khatam- ifard, Minsik Cho, Carlo C Del Mundo, Mohammad Rastegari, and Mehrdad Farajtabar. 2024. Llm in a flash: Efficient large language model inference with limited memory. InProceedings of the 62nd An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 12562–12584
2024
-
[4]
Oscar Antepara, Zhengji Zhao, Brian Austin, Nan Ding, Leonid Oliker, Nicholas J Wright, and Samuel Williams. 2025. Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated Su- percomputing. InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 888–904
2025
-
[5]
Malika Aubakirova, Alex Atallah, Chris Clark, Justin Summerville, and Anjney Midha. 2026. State of AI: An Empirical 100 Trillion Token Study with OpenRouter.arXiv preprint arXiv:2601.10088(2026)
arXiv 2026
-
[6]
badlogic. 2026. pi-mono.https://github.com/badlogic/pi-mono. Ver- sion: main branch, accessed 2026-04-15
2026
-
[7]
Yushi Bai, Xin Lv, Jiajie Zhang, Hong Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2024. Longbench: A bilingual, multitask bench- mark for long context understanding. InProceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers). 3119–3137
2024
-
[8]
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020. Longformer: The long-document transformer.arXiv preprint arXiv:2004.05150 (2020)
Pith/arXiv arXiv 2020
-
[9]
Adithya Bhaskar, Alexander Wettig, Tianyu Gao, Yihe Dong, and Danqi Chen. 2025. Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?arXiv preprint arXiv:2506.17121 (2025)
Pith/arXiv arXiv 2025
-
[10]
Jalil Boukhobza, Pierre Olivier, Wen Sheng Lim, Liang-Chi Chen, Yun- Shan Hsieh, Shin-Ting Wu, Chien-Chung Ho, Po-Chun Huang, and Yuan-Hao Chang. 2025. A survey on flash-memory storage systems: A host-side perspective.ACM Transactions on Storage21, 3 (2025), 1–59
2025
-
[11]
Yu Cai, Saugata Ghose, Erich F Haratsch, Yixin Luo, and Onur Mutlu
-
[12]
IEEE105, 9 (2017), 1666–1704
Error characterization, mitigation, and recovery in flash- memory-based solid-state drives.Proc. IEEE105, 9 (2017), 1666–1704
2017
-
[13]
Yu Cai, Gulay Yalcin, Onur Mutlu, Erich F Haratsch, Adrian Cristal, Os- man S Unsal, and Ken Mai. 2012. Flash correct-and-refresh: Retention- aware error management for increased flash memory lifetime. In2012 IEEE 30th International Conference on Computer Design (ICCD). IEEE, 94–101
2012
-
[14]
Wanik Cho, Chanhui Jeong, Jongwoo Kim, Jongseok Jung, Keunseon Ahn, Jayoon Goo, Sangkyu Lee, Kayoung Cho, Tei Cho, Dauni Kim, Gwan Park, Yushin Ahn, Sooyeol Chai, Gwihan Ko, Sunyoung Jung, Eunwoo Jo, Taehun Park, Jinhyun Ban, Cheoljoong Park, Jae Hyun Park, Sanghoon Oh, Sojin Jeong, Youngjun Kwak, Kyungsoo Jeong, Jinyeop Kim, Minchol Shin, Eunho Yang, Tai...
2025
-
[15]
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
-
[16]
Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in neural information processing systems35 (2022), 16344–16359
2022
-
[17]
DeepSeek-AI. 2025.Day 6: One More Thing, DeepSeek-V3/R1 Inference System Overview.https://github.com/deepseek-ai/open-infra- index/blob/main/202502OpenSourceWeek/day_6_one_more_thing_ deepseekV3R1_inference_system_overview.mdGitHub repository
2025
-
[18]
DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huaj...
Pith/arXiv arXiv 2025
-
[19]
Kaustubh Deshpande, Ved Sirdeshmukh, Johannes Baptist Mols, Lifeng Jin, Ed-Yeremai Hernandez-Cardona, Dean Lee, Jeremy Kritz, Willow E Primack, Summer Yue, and Chen Xing. 2025. Multichallenge: A re- alistic multi-turn conversation evaluation benchmark challenging to frontier llms. InFindings of the Association for Computational Linguis- tics: ACL 2025. 18...
2025
-
[20]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Ka- dian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schel- ten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony S. Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aur’elien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Rozière,...
Pith/arXiv arXiv 2024
-
[21]
Zehao Fan, Garrett Gagnon, Zhenyu Liu, and Liu Liu. 2025. Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM. arXiv:2505.05772 [cs.CL]https://arxiv.org/abs/2505.05772
Pith/arXiv arXiv 2025
-
[22]
Raja Gond, Nipun Kwatra, and Ramachandran Ramjee. 2025. Token- Weave: Efficient Compute-Communication Overlap for Distributed LLM Inference.https://arxiv.org/abs/2505.11329
Pith/arXiv arXiv 2025
-
[23]
Sylvain Gugger, Lysandre Debut, Thomas Wolf, Philipp Schmid, Zachary Mueller, Sourab Mangrulkar, Marc Sun, and Benjamin Bossan
-
[24]
Accelerate: Training and inference at scale made simple, efficient and adaptable.https://github.com/huggingface/accelerate
-
[25]
Aayush Gupta, Youngjae Kim, and Bhuvan Urgaonkar. 2009. DFTL: a flash translation layer employing demand-based selective caching of page-level address mappings.Acm Sigplan Notices44, 3 (2009), 229–240
2009
-
[26]
Minho Ha, Euiseok Kim, and Hoshik Kim. 2026. H 3: Hybrid Archi- tecture Using High Bandwidth Memory and High Bandwidth Flash for Cost-Efficient LLM Inference.IEEE Computer Architecture Letters (2026)
2026
-
[27]
Guseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi, Sanghyeon Lee, Hyungkyu Ham, Gwangsun Kim, Divya Mahajan, and Jongse Park. 2024. Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing. InProceedings of the 29th ACM International Confer- ence on Architectural Support for Programming Languages and Operat- ing Systems, Volume 3. 722–737
2024
-
[28]
Po-Kai Hsu, Weihong Xu, Qunyou Liu, Tajana Rosing, and Shimeng Yu. 2026. HAVEN: High-Bandwidth Flash Augmented Vector Engine for Large-Scale Approximate Nearest-Neighbor Search Acceleration. arXiv preprint arXiv:2603.01175(2026)
arXiv 2026
-
[29]
Yitao Hu, Xiulong Liu, Guotao Yang, Linxuan Li, Kai Zeng, Zhixin Zhao, Sheng Chen, Laiping Zhao, Wenxin Li, and Keqiu Li. 2025. TightLLM: Maximizing throughput for LLM inference via adaptive offloading policy.IEEE Trans. Comput.(2025)
2025
-
[30]
Hongjing Huang, Zeke Wang, Jie Zhang, Zhenhao He, Chao Wu, Jun Xiao, and Gustavo Alonso. 2021. Shuhai: A tool for benchmarking high bandwidth memory on FPGAs.IEEE Trans. Comput.71, 5 (2021), 1133–1144
2021
-
[31]
Hongshin Jun, Jinhee Cho, Kangseol Lee, Ho-Young Son, Kwiwook Kim, Hanho Jin, and Keith Kim. 2017. Hbm (high bandwidth memory) dram technology and architecture. In2017 IEEE International Memory Workshop (IMW). IEEE, 1–4
2017
-
[32]
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Ben- jamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361(2020)
Pith/arXiv arXiv 2020
-
[33]
KIOXIA Corporation. 2022.New Storage Class Memory Solution Accelerates Non-Relational Database Performance: Aerospike® NoSQL Databases Using KIOXIA FL6 Series Enterprise NVMe® SCM SSDs Deliver Heightened Performance Gains versus TLC SSDs. Application Brief Rev. 1.0. KIOXIA Corporation. https://americas.kioxia.com/content/dam/kioxia/en-us/business/ ssd/ent...
2022
-
[34]
Apostolos Kokolis, Michael Kuchnik, John Hoffman, Adithya Kumar, Parth Malani, Faye Ma, Zachary DeVito, Shubho Sengupta, Kalyan Saladi, and Carole-Jean Wu. 2025. Revisiting reliability in large-scale machine learning research clusters. In2025 IEEE International Sym- posium on High Performance Computer Architecture (HPCA). IEEE, 1259–1274
2025
-
[35]
Toshiyuki Kouchi, Mami Kakoi, Noriyasu Kumazaki, Akio Sugahara, Akihiro Imamoto, Yasufumi Kajiyama, Yuri Terada, Bushnaq Sanad, Naoaki Kanagawa, Takuyo Kodama, Ryo Fukuda, Hiromitsu Ko- mai, Norichika Asaoka, Hidekazu Ohnishi, Ryosuke Isomura, Takaya Handa, Kensuke Yamamoto, Yuki Ishizaki, Yoko Deguchi, Atsushi Okuyama, Junichi Sato, Hiroki Yabe, Hua-Ling...
2020
-
[36]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Yu, Joey Gonzalez, Hao Zhang, and Ion Stoica. 2023. vllm: Easy, fast, and cheap llm serving with pagedattention.See https://vllm. ai/(accessed 9 August 2023)(2023)
2023
-
[37]
Chae Yeon Lee, Chae Ho Won, Seyeon Jung, Eun Su Jung, Tae Min Choi, Hwa Rim Lee, JinUk Yoo, Songhun Yoon, and Sung Gyu Pyo
-
[38]
3D integrated process and hybrid bonding of high bandwidth memory (HBM).Electronic Materials Letters21, 3 (2025), 395–419
2025
-
[39]
Jaeyong Lee, Hyeunjoo Kim, Sanghun Oh, Myoungjun Chun, Myung- suk Kim, and Jihong Kim. 2025. Aif: Accelerating on-device llm in- ference using in-flash processing. InProceedings of the 52nd Annual International Symposium on Computer Architecture. 529–543
2025
-
[40]
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the middle: How language models use long contexts.Transactions of the association for computational linguistics12 (2024), 157–173
2024
-
[41]
Yixin Luo, Yu Cai, Saugata Ghose, Jongmoo Choi, and Onur Mutlu
-
[42]
In2015 31st Symposium on Mass Storage Systems and Technologies (MSST)
WARM: Improving NAND flash memory lifetime with write- hotness aware retention management. In2015 31st Symposium on Mass Storage Systems and Technologies (MSST). IEEE, 1–14
-
[43]
Micron Technology. 2024.HBM3E: Powering the future of AI with high-bandwidth memory.https://www.micron.com/about/blog/ applications/ai/microns-hbm3e-powering-the-future-of-ai-with- high-bandwidth-memoryAccessed: 2026-04-16
2024
-
[44]
Vidyabhushan Mohan, Taniya Siddiqua, Sudhanva Gurumurthi, and Mircea R Stan. 2010. How I learned to stop worrying and love flash endurance. In2nd Workshop on Hot Topics in Storage and File Systems (HotStorage 10)
2010
-
[45]
NVIDIA Corporation. 2022. NVIDIA H100 GPU.https://resources. nvidia.com/en-us-gpu-resources/h100-datasheet-24306. Accessed: 2026-04-14
2022
-
[46]
NVIDIA Corporation. 2024. NVIDIA DGX H200 System Architec- ture.https://www.nvidia.com/en-us/data-center/dgx-h200/. Includes ConnectX-7 InfiniBand networking, implying RDMA capability
2024
-
[47]
Jaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim, Yongsuk Kwon, Nam Sung Kim, and Jung Ho Ahn. 2024. Attacc! unleashing the power of pim for batched transformer-based gener- ative model inference. InProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. 103–119
2024
-
[48]
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. InProceedings of the 36th annual acm symposium on user interface software and technology. 1–22
2023
-
[49]
Balls into bins
Martin Raab and Angelika Steger. 1998. “Balls into bins”—A simple and tight analysis. InInternational Workshop on Randomization and Approximation Techniques in Computer Science. Springer, 159–170
1998
-
[50]
Ananda Samajdar, Yuhao Zhu, Paul Whatmough, Matthew Mattina, and Tushar Krishna. 2018. Scale-sim: Systolic cnn accelerator simula- tor.arXiv preprint arXiv:1811.02883(2018)
Pith/arXiv arXiv 2018
-
[51]
SanDisk. 2025.Memory-Centric AI: Sandisk’s High Bandwidth Flash Will Redefine AI Infrastructure.https://www.sandisk.com/ company/newsroom/blogs/2025/memory-centric-ai-sandisks-high- bandwidth-flash-will-redefine-ai-infrastructure[Online]
2025
-
[52]
Minseok Seo, Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon, Guhyun Kim, Chanwook Park, Ilkon Kim, Jaehan Park, Jeong- bin Kim, Woojae Shin, Jongsoon Won, Haerang Choi, Kyuyoung Kim, Daehan Kwon, Chunseok Jeong, Sangheon Lee, Yongseok Choi, Wooseok Byun, Seungcheol Baek, Hyuk-Jae Lee, and John Kim. 2024. Ianus: Integrated accelerator based on npu-pim ...
2024
-
[53]
Zhihong Shao, Damai Dai, Daya Guo, Bo Liu (Benjamin Liu), Zihan Wang, and Huajian Xin. 2024. Deepseek-v2: A strong, economi- cal, and efficient mixture-of-experts language model.arXiv preprint arXiv:2405.04434(2024)
Pith/arXiv arXiv 2024
-
[54]
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.arXiv preprint arXiv:1701.06538(2017)
Pith/arXiv arXiv 2017
-
[55]
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023. Flexgen: High-throughput generative inference of large language models with a single gpu. InInternational Conference on Machine Learning. PMLR, 31094–31116
2023
-
[56]
Ramesh Sitaraman. 2001. The power of two random choices: A survey of techniques and results. (2001)
2001
-
[57]
Weiyi Sun, Mingyu Gao, Zhaoshi Li, Aoyang Zhang, Iris Ying Chou, Jianfeng Zhu, Shaojun Wei, and Leibo Liu. 2025. Lincoln: Real-Time 50˜ 100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash Memory. In2025 IEEE International Sym- posium on High Performance Computer Architecture (HPCA). IEEE, 1734–1750
2025
-
[58]
Mojumder, Shi Dong, Rafael Ubal, Xiang Gong, Shane Treadway, Yuhui Bao, Vincent Zhao, José L
Yifan Sun, Trinayan Baruah, Saiful A. Mojumder, Shi Dong, Rafael Ubal, Xiang Gong, Shane Treadway, Yuhui Bao, Vincent Zhao, José L. Abellán, John Kim, Ajay Joshi, and David Kaeli. 2018. MGSim + MGMark: A Framework for Multi-GPU System Research. arXiv:1811.02884 [cs.DC]https://arxiv.org/abs/1811.02884
Pith/arXiv arXiv 2018
-
[59]
Jayanth M. Thimmaiah, Ryuji Yamashita, In-Soo Yoon, Jason Li, Cynthia Hsu, Takuya Ariki, Naoki Ookuma, Yosuke Kato, Koichiro Hayashi, Kazuki Yamauchi, Indra K V, Masahiro Kano, Sirisha Bhamidi- pati, Sneha Bhatia, Seema Malhotra, Naoki Ojima, Ella Wu, Zhiyong Yang, Frank W. Tsai, Mathias Bayle, Naoyuki Minami, Yasuyuki Fuji- hara, Kei Kitamura, Tomofumi K...
-
[60]
In2026 IEEE International Solid-State Circuits Conference (ISSCC), Vol
A 2Tb 4b/Cell 6-Plane 3D-Flash Memory with 37.6 Gb/mm 2 Bit Density and> 85MB/s Write Throughput. In2026 IEEE International Solid-State Circuits Conference (ISSCC), Vol. 69. IEEE, 254–256
-
[61]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need.Advances in neural information processing systems30 (2017)
2017
-
[62]
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023. Voyager: An open-ended embodied agent with large language models.arXiv preprint arXiv:2305.16291(2023)
Pith/arXiv arXiv 2023
-
[63]
Jiahao Wang, Jinbo Han, Xingda Wei, Sijie Shen, Dingyan Zhang, Chenguang Fang, Rong Chen, Wenyuan Yu, and Haibo Chen. 2025. {KVCache} Cache in the Wild: Characterizing and Optimizing {KVCache} Cache at a Large Cloud Provider. In2025 USENIX An- nual Technical Conference (USENIX ATC 25). 465–482
2025
-
[64]
Qian Wang, Zhenheng Tang, Zichen Jiang, Nuo Chen, Tianyu Wang, and Bingsheng He. 2025. Agenttaxo: Dissecting and benchmarking token distribution of llm multi-agent systems. InICLR 2025 Workshop on Foundation Models in the Wild
2025
-
[65]
Wei Wang, Wen Pan, Tao Xie, and Deng Zhou. 2016. How many MLCs should impersonate SLCs to optimize SSD performance?. In Proceedings of the Second International Symposium on Memory Systems. 238–247
2016
-
[66]
Samuel Williams, Andrew Waterman, and David Patterson. 2009. Roofline: an insightful visual performance model for multicore ar- chitectures.Commun. ACM52, 4 (2009), 65–76
2009
-
[67]
Kosuke Yanagidaira, Mario Sako, Yasuhiro Hirashima, Yumi Higashi, Yutaka Shimizu, Takeshi Nakano, Yusuke Ochi, Hiroaki Yamada, Nobushi Matsuura, Akihiro Imamoto, Kazuaki Kawaguchi, Koji Tabata, Hiroaki Hoshino, Takeshi Hioka, Shigehito Saigusa, Hiroki Date, Masaki Unno, Jumpei Sato, You Kamata, Takahiro Shimizu, Akio Sug- ahara, Taira Shibuya, Atsushi Oku...
2025
-
[68]
Kosuke Yanagidaira, Mario Sako, Yasuhiro Hirashima, Junya Matsuno, Yumi Higashi, Yutaka Shimizu, Akihiro Imamoto, Kazuaki Kawaguchi, Koji Tabata, Takeshi Nakano, Yusuke Ochi, Hiroaki Hoshino, Takeshi , Xinyu Wang, Yalong Xue, Xiaotian Sun, Xiaoyu Zhang, Chunmeng Dou, Xueqi Li, and Xiaoming Chen Hioka, Shigehito Saigusa, Hiroki Date, Masaki Unno, Jumpei Sa...
2025
-
[69]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxin Yang, Jingren Zhou, Jingren Zhou, Junyan Lin, Kai Dang, Keqin Bao, Ke-Pei Ya...
-
[70]
Qwen3 technical report.arXiv preprint arXiv:2505.09388(2025)
Pith/arXiv arXiv 2025
-
[71]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. InThe eleventh international conference on learning representations
2022
-
[72]
Xi Ye, Fangcong Yin, Yinghui He, Joie Zhang, Howard Yen, Tianyu Gao, Greg Durrett, and Danqi Chen. 2025. LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation. In Second Conference on Language Modeling.https://openreview.net/ forum?id=ruWC5LIMSo
2025
-
[73]
Zihao Yi, Jiarui Ouyang, Zhe Xu, Yuwen Liu, Tianhao Liao, Haohao Luo, and Ying Shen. 2025. A survey on recent advances in llm-based multi-turn dialogue systems.Comput. Surveys58, 6 (2025), 1–38
2025
-
[74]
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022. Orca: A distributed serving system for {Transformer-Based} generative models. In16th USENIX symposium on operating systems design and implementation (OSDI 22). 521–538
2022
-
[75]
Zhongkai Yu, Shengwen Liang, Tianyun Ma, Yunke Cai, Ziyuan Nan, Di Huang, Xinkai Song, Yifan Hao, Jie Zhang, Tian Zhi, Yongwei Zhao, Zidong Du, Xing Hu, Qi Guo, and Tianshi Chen. 2024. Cambricon-llm: A chiplet-based hybrid architecture for on-device inference of 70b llm. In2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1474–1488
2024
-
[76]
Hengrui Zhang, August Ning, Rohan Baskar Prabhakar, and David Wentzlaff. 2024. Llmcompass: Enabling efficient hardware design for large language model inference. In2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). IEEE, 1080– 1096
2024
-
[77]
Xing, Haotong Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhang- hao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Haotong Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judg- ing llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems36 (2023), 46595–46623
2023
-
[78]
Gonzalez, Clark W
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark W. Barrett, and Ying Sheng. 2024. Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems37 (2024), 62557–62583
2024
-
[79]
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xu- anzhe Liu, Xin Jin, and Hao Zhang. 2024. {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). 193–210
2024
-
[80]
Anqi Zhou, Yu Zhang, Fei Ding, Ziqi Lian, Renxi Jin, Yudong Yang, Qidong Wang, and Liqiang Cao. 2024. Research progress of hybrid bonding technology for three-dimensional integration.Microelectron- ics Reliability155 (2024), 115372
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.