REVIEW 3 major objections 6 minor 49 references
OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read OasisKV claims that decode-time attention sparsity can be predicted one step ahead from speculative-draft queries, letting KV caches live off-GPU without losing accuracy.
desk verdict Solid systems paper with a genuinely new draft-token lookahead idea, but the headline accuracy bound is contradicted by its own Table 1 and the load-bearing prediction claim rests on a figure that doesn't validate the production operating point. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is look-ahead attention: one shared attention kernel processes the normal token and the draft token over the same resident sparse KV set, then the draft query scans compressed key summaries (coordinate-wise min/max per block, in the style of Quest) kept in HBM at about 1/16 of the full KV-cache size. The scan produces a head-wise top-$K$ ranking of all logical blocks without restoring any keys. A coordinator-driven asynchronous pipeline runs three background CUDA streams per layer—top-$K$ prediction, KV selection, and KV transfer—with layer-local synchronization, and a capped-eviction policy pairs each admitted nonresident block with a least-recently-selected resident block so per-step PCIe traffic is bounded. A head-wise logical-to-logical mapping layer above the unchanged page tables lets different KV heads hold different sparse block sets without breaking the batched attention path.
What would settle it
Run the Fig. 6 per-layer top-$K$ agreement profile on a different model (for example Llama-3.1-8B-Instruct) and on a multi-turn agentic trace with tool observations; if average per-layer agreement falls below the reported 98.74%, prefetch misses will grow and OasisKV will either drop attention context or stall on on-demand fetches, degrading the reported accuracy and throughput.
Extended reading notes
Core claim
The central claim is that the next token's attention pattern is predictable from the current step: if the draft query from speculative decoding is propagated forward through the current sparse resident KV set, the top-$K$ block set it ranks agrees with the true next-token query's exact top-$K$ set at least 98.2% in every layer, 98.74% on average, as measured on Qwen3-8B with GSM8K. OasisKV builds a serving system around that signal. The full KV cache lives in CPU or remote memory; HBM holds a bounded head-wise working set plus per-block min/max key summaries. Each layer's draft query scans those summaries, ranks all logical blocks, selects missing top-$K$ blocks, and prefetches them over PCIe or the network in the background, with per-step admission capped so transfers stay within the bandwidth a decode step can hide. The result is that sparsity pays off as throughput: larger batches fit in HBM, and per-request memory footprint no longer grows with context length.
Load-bearing premise
The design depends on the draft query, computed from the GPU's currently resident sparse KV set, selecting the same most-relevant KV blocks as the true next-token query at least 98.2% of the time in every layer, a number measured on one model and one dataset.
Editorial extensions
If this is right
- HBM capacity no longer caps decode batch size: with a 2,048-token KV budget OasisKV sustains 90-95 concurrent requests at 16K context where dense attention supports about 22.
- Decode throughput can exceed a dense engine while holding accuracy near full attention: 1.69x on AIME24 at 0.1 points of accuracy loss, and up to 1.89x if the top-$K$ budget is shrunk at a cost of about 4.4 points.
- Prefill-decode disaggregation no longer requires full KV transfer: admission traffic drops 6.5-9.7x and decode-node host DRAM drops 2.2-2.6x, while throughput remains about 2x dense.
- Per-step PCIe traffic, not attention compute, is the binding constraint: capping the fetch ratio at 0.05 keeps accuracy within 0.1 point of dense while more than doubling throughput over an uncapped fetch in the ablation.
- The same lookahead signal works for both dense 8B-class models and a 235B MoE model under tensor parallelism, so the mechanism is not tied to a single architecture.
Reading between the lines
- If the 98.2% per-layer agreement floor generalizes beyond the one measured model and dataset, the lookahead mechanism could also prefetch model weights or activations, since it only needs a future-query signal and rankable summaries.
- Because the paper's prefix-caching TTFT analysis is analytic rather than implemented, the promised 2.0-2.2x TTFT reduction at 90% hit rate remains untested; wiring remote partial fetching into a prefix-caching engine would settle it.
- The fetch-cap ablation suggests the right operating point is set by interconnect bandwidth rather than HBM capacity, so faster interconnects should shift the optimal cap upward and extend the accuracy-throughput frontier.
- Re-running the Fig. 6 agreement profile on the paper's own Llama-3.1-8B and Qwen3-235B workloads would show whether the lookahead signal is robust across model families or specific to the Qwen3-8B/GSM8K pairing.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. OasisKV is a vLLM-based sparse KV prefetching system for LLM decode. It keeps the full KV cache in CPU DRAM or remote memory, retains only a bounded, head-wise working set in GPU HBM (default B=16, K=128 blocks, i.e., 2,048 tokens), and uses draft tokens from an EAGLE-3-style speculative-decoding module as a lookahead signal. The draft query is propagated through each layer, scans compressed per-block key summaries, predicts the next step's top-K blocks, and a fully asynchronous pipeline prefetches missing blocks over PCIe or the network under a capped-eviction policy. The paper evaluates accuracy on LongBench v2, AIME24/25, and GPQA-Diamond; throughput on Qwen3-8B, Qwen3-235B, and Llama-3.1-8B in single-GPU, TP8, and PD-disaggregated settings; and reports up to 1.69x–2.1x dense throughput within 0.7 points of full-attention accuracy.
Significance. The core idea is attractive and the systems work is substantial: reusing the speculative-decoding draft as a training-free lookahead signal, overlapping prefetch behind per-layer forward compute, bounding per-step traffic with a capped-eviction policy, and extending the same signal to remote partial fetching. The paper is careful to compare accuracy only against the full-attention anchor of the same stack, and Table 2 cleanly demonstrates that decode throughput is PCIe-bandwidth-bound rather than attention-compute-bound. If the prediction-agreement claim holds at the production operating point on long contexts, the throughput and memory-savings numbers would be a useful contribution. The main gaps are that the load-bearing agreement evidence is not measured at K=128 on long contexts, the headline accuracy bound is contradicted by several Table 1 entries, and the default fetch ratio is selected and evaluated on the same benchmark. These concerns are addressable with additional experiments and revised claims; I do not see an internal inconsistency that would require rejection.
major comments (3)
- [§4.2.1, Fig. 6] The entire prefetch pipeline rests on the claim that the propagated draft query predicts the next step's exact top-K block set with ≥98.2% per-layer agreement. The only evidence, Fig. 6, is measured on Qwen3-8B with GSM8K and does not state K; the caption also does not state whether the draft query was computed with the bounded sparse resident set or with the full KV context. With the production configuration (B=16, K=128, 2,048 tokens) and typical GSM8K prompts shorter than 2,048 tokens, the top-128 set would be the entire context and agreement would be trivially 100%, whereas the reported mean of 98.74% suggests a much smaller K was used, possibly the K=20 of Fig. 4. No experiment isolates prediction agreement from the eviction/cap policy on the long-context workloads where the throughput claims are made. Because §3.1 correctly states that a miss is not free, the prefetch-miss behavior at the real operating point is load-bearing and currently unmeasured. Please report per-layer agreement at K=128 on AIME24, LongBench v2, or an agentic trace, with context lengths, resident-set size, and K explicitly stated.
- [Abstract, §1 contribution bullet, §5.4, Table 1] The claim that accuracy stays 'within 0.7 points' of full attention is not supported by Table 1. The table shows per-split and per-benchmark deltas of -2.78 on Qwen3-8B LongBench v2 long, -2.50 on AIME24 avg@8, -1.99 on GPQA-Diamond pass@4, and -1.67 on Llama-3.1-8B LongBench short, all under the same 2,048-token KV budget. The 0.7 figure only describes selected overall averages such as LongBench overall (-0.40) and the long-output pass@8 overall (-0.66). The contribution bullet that says 'all within 0.7 points' is therefore too strong, and §5.4's 'within about a point of its anchor' is also inconsistent with the table. Please report the worst-case delta or the full distribution, and revise the abstract and §5.4 accordingly.
- [§5.5.1, Table 2; §5.3.3; Fig. 12] The default operating point (fetch ratio 0.05) and the headline '1.69x at 0.1 points accuracy loss' are selected and evaluated on the same benchmark, AIME24. Table 2 sweeps the fetch cap on AIME24 and then Fig. 13 reports the AIME24 result at the best row; the paper describes no held-out validation or tuning/validation split. The same 0.05 cap is then used in the PD-disaggregation experiments of Fig. 12, so the selection-on-test concern propagates to the 2.1–2.3x PD claims. Please validate the chosen fetch ratio on held-out benchmarks such as AIME25, GPQA, and LongBench, report accuracy variability, or clearly designate Table 2 as a tuning study and the main claims as evaluated on separate data.
minor comments (6)
- [Fig. 6 caption] The caption should state K, the context length, and whether the draft query was evaluated with the bounded sparse resident set or with full KV; without these parameters the measurement cannot be reproduced or compared with the production configuration.
- [§4.2.3] There is a typo: 'measures how how limiting' should read 'measures how limiting'.
- [Abstract] The phrase '2.2-2.6 less decode-node host memory' appears to be missing the multiplication sign; it should read '2.2–2.6× less'.
- [§5.4, Table 1] The statement that the method stays 'within about a point of its anchor' is inconsistent with several Table 1 entries (e.g., -2.78, -2.50, -1.99); please align the prose with the reported numbers.
- [Fig. 12 vs Fig. 14(a)] The relation between Fig. 12's per-request host-memory occupancy (1.54 and 1.73 GiB) and Fig. 14(a)'s total transferred bytes per request (about 2.1–3.0 GiB at 32K) should be clarified: is the occupancy a final value or an average over time, and how does it relate to the monotonic growth of the resident set described in §4.4.2?
- [§5.3.1] The sentence 'the unplotted points are the cases that are unreachable in their frameworks' is vague; please list the concrete configurations that OOM and the configurations that are unsupported for each baseline.
Circularity Check
No significant circularity: the lookahead signal comes from an external EAGLE-3-style draft model and prediction agreement is measured against exact top-K sets; the narrow Fig. 6 validation is a correctness risk, not a circular step.
full rationale
The paper's central derivation chain is self-contained rather than circular. The lookahead signal in §4.2.1 is supplied by an external EAGLE-3-style draft model ([16]), not by OasisKV's own KV-selection mechanism, and the load-bearing accuracy claim (Fig. 6) is an empirical profile comparing the propagated draft query's top-K set with the exact top-K set of the true next-token query; agreement is measured, not defined into existence. The bandwidth-pipeline constraints (Eqs. 3–4) are roofline inequalities that do not encode the conclusion. Accuracy comparisons (Table 1) are read as deltas against each stack's own full-attention anchor, which is a fair protocol rather than a self-referential reduction. The only author self-citation in the paper is [43], used alongside the external trace study [18] for the contextual claim about agentic workloads; it is incidental and non-load-bearing. The operating hyperparameters (K=128, fetch ratio 0.05) are chosen by ablation on the evaluation benchmark before reporting the headline results; that is an overfitting/validity concern, not circularity, because the reported throughput and accuracy still depend on genuinely measured behavior of the external draft model, PCIe staging, and network prefetch. The narrow validation in Fig. 6 (one model, one dataset, unspecified K) is a real correctness risk for the load-bearing prediction premise, but the premise is an empirical measurement rather than a definitional or self-citation-based reduction.
Assumptions & free parameters
free parameters (5)
- KV block size B =
16 tokens
- Top-K blocks per head K =
128 blocks (2,048 tokens)
- Per-step fetch ratio cap =
0.05
- Sparsification threshold (token budget) =
not reported
- Compressed-key summary overhead =
1/16 of full KV size
assumptions (6)
- domain assumption Decode-time attention is sparse enough that a fixed 2,048-token per-head budget preserves accuracy.
- domain assumption Adjacent decode steps have strong temporal locality in block importance, so blocks selected for the current token contain the next step's top-K set.
- domain assumption Draft tokens from EAGLE-3-like modules provide an accurate lookahead signal without an extra full forward pass.
- domain assumption Block-level min/max key summaries rank the global top-K blocks correctly.
- domain assumption PCIe/network prefetch traffic can be hidden within the decode-step overlap window C_token.
- domain assumption The underlying hardware (pinned memory, UVA, CUDA streams, RoCE) behaves as assumed.
Cite this review
Pith. "Pith review of OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching." pith.science (2026). https://pith.science/paper/RCWXODBN
@misc{pith2026260808097,
author = {Pith},
title = {Pith review of: OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching},
year = {2026},
howpublished = {\url{https://pith.science/paper/RCWXODBN}},
note = {Machine review of arXiv:2608.08097}
}
abstract
Large language model (LLM) inference serving is increasingly constrained by memory rather than compute. As long-context and long-form reasoning workloads become more prevalent, the key-value (KV) cache dominates both memory footprint and memory traffic during LLM token generation, i.e., decode. In particular, HBM capacity has become a scarce and costly resource that heavily limits inference batch size and system throughput. This paper presents OasisKV, a memory-centric LLM inference system design that alleviates HBM capacity pressure by decoupling full KV-cache storage from HBM during LLM decoding. Because decode-time attention is naturally sparse, OasisKV keeps only the KV entries of the most relevant tokens in HBMs for attention computation. We observe that future important tokens can be predicted accurately in advance using lookahead tokens drafted by speculative decoding (SD). OasisKV employs an efficient attention background pipeline to identify important KV blocks. They are then prefetched from higher-capacity memory tiers (e.g., host or remote memory) and staged in HBMs before being used in the next decode step. We implement OasisKV based on vLLM. The lookahead prediction is accurate enough to keep accuracy within 0.7 points of full attention under a 2,048-token KV budget. This lets OasisKV turn sparsity into throughput gain: $1.69\times$ over dense vLLM on the reasoning workload at 0.1 points of accuracy loss, and up to $2.1\times$ on multi-GPU long-context serving. Under prefill--decode disaggregation, OasisKV reaches about $2\times$ dense throughput while admitting each request with $6.5$--$9.7\times$ less KV and holding $2.2$-$2.6$ less decode-node host memory than full KV transfer.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, and Sumit Sanghai. 2023. GQA: Training General- ized Multi-Query Transformer Models from Multi-Head Checkpoints. arXiv preprint arXiv:2305.13245(2023)
arXiv 2023
-
[2]
Renze Chen, Zhuofeng Wang, Beiquan Cao, Tong Wu, Size Zheng, Xiuhong Li, Xuechao Wei, Shengen Yan, Meng Li, and Yun Liang. 2024. ArkVale: Efficient Generative LLM Inference with Recallable Key- Value Eviction. InAdvances in Neural Information Processing Systems 37
work page 2024
-
[3]
Yaoqi Chen, Jinkai Zhang, Baotong Lu, Qianxi Zhang, Chengruidong Zhang, Jing Liu, Jingjia Luo, Di Liu, Huiqiang Jiang, Qi Chen, Bailu Ding, Xiao Yan, Jiawei Jiang, Chen Chen, Mingxing Zhang, Cheng Li, Yuqing Yang, Fan Yang, and Mao Yang. 2026. RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference.Proceedings of the VLDB Endowment19, ...
-
[4]
DeepSeek-AI. 2024. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.arXiv preprint arXiv:2405.04434 (2024)
arXiv 2024
-
[5]
et.al. DeepSeek-AI. 2025. DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models. arXiv:2512.02556 [cs.CL] https: //arxiv.org/abs/2512.02556
arXiv 2025
-
[6]
Hendryx, Zifan Wang, Chen Bo Calvin Zhang, Noah Jacobson, Bing Liu, and Brad Kenstler
Xiang Deng, Jeff Da, Edwin Pan, Yan He, Charles Ide, Kanak Garg, Niklas Lauffer, Andrew Park, Nitin Pasari, Chetan Rane, Karmini Sampath, Maya Krishnan, Srivatsa Kundurthy, Sean M. Hendryx, Zifan Wang, Chen Bo Calvin Zhang, Noah Jacobson, Bing Liu, and Brad Kenstler. 2025. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? https:/...
work page 2025
-
[7]
Yaosheng Fu, Guangxuan Xiao, Xin Dong, Song Han, and Oreste Villa
-
[8]
Yizhao Gao, Zhichen Zeng, Dayou Du, Shijie Cao, Peiyuan Zhou, Jiaxing Qi, Junjie Lai, Hayden Kwok-Hay So, Ting Cao, Fan Yang, and Mao Yang. 2025. SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs. arXiv:2410.13276 [cs.CL] https://arxiv.org/abs/ 2410.13276
arXiv 2025
Show all 49 references
-
[9]
In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khan- delwal, and Lin Zhong. 2024. Prompt Cache: Modular Attention Reuse for Low-Latency Inference. InProceedings of Machine Learning and Systems, P. Gibbons, G. Pekhimenko, and C. De Sa (Eds.), Vol. 6. 325–338. https:/...
2024
-
[10]
Shiyu Ji, Yixuan Wang, Yijun Liu, Qingfu Zhu, and Wanxiang Che
-
[11]
Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu
Huiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu, Xu- fang Luo, Surin Ahn, Zhenhua Han, Amir H. Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2024. MInference 1.0: Acceler- ating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention. arXiv:240...
2024 arXiv
-
[12]
arXiv:2603.22910 [cs.CL] https://arxiv.org/abs/ 2603.22910
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction. arXiv:2603.22910 [cs.CL] https://arxiv.org/abs/ 2603.22910
-
[13]
Xunhao Lai, Jianqiao Lu, Yao Luo, Yiyuan Ma, and Xun Zhou. 2025. FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference. arXiv:2502.20766 [cs.LG] https://arxiv. org/abs/2502.20766 OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookah...
2025 arXiv
-
[14]
Shibo Jie, Yehui Tang, Kai Han, Zhi-Hong Deng, and Jing Han. 2025. SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs. InProceedings of the 42nd International Conference on Machine Learning, Vol. 267. 27917–27928. https://mlanthology.org/icml /2025/jie202...
2025
-
[15]
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast Inference from Transformers via Speculative Decoding. InProceedings of the 40th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 202), Andreas Krause, Emma Brunskill, Kyungh...
2023
-
[16]
Wonbeom Lee, Jungi Lee, Junghwan Seo, and Jaewoong Sim. 2024. InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management. arXiv:2406.19707 [cs.LG] https://arxiv.org/abs/2406.19707
2024 arXiv
-
[17]
Jian Lin, Jiazhi Mi, Zicong Hong, Haodong Wang, Qianli Liu, Haodyue Zhang, Peng Li, and Song Guo. 2026. KVDrive: A Holistic Multi- Tier KV Cache Management System for Long-Context LLM Inference. arXiv:2605.18071 [cs.CL]https://arxiv.org/abs/2605.18071
2026 arXiv
-
[18]
Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. 2025. EAGLE-3: Scaling up Inference Acceleration of Large Language Mod- els via Training-Time Test. InAdvances in Neural Information Process- ing Systems, D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi,...
2025
-
[19]
Guangda Liu, Wenhao Chen, Chengwei Li, Zhenyu Ning, Jing Lin, Yiwu Yao, Quan Chen, Shixuan Sun, Jieru Zhao, and Minyi Guo. 2026. ECHO: Efficient KV Cache Offloading with Lossless Prefetching for Serving Native Sparse Attention LLMs. In20th USENIX Symposium on Operating Systems...
2026
-
[20]
Banruo Liu, Haoran Qiu, Íñigo Goiri, Rodrigo Fonseca, Ricardo Bianchini, and Esha Choukse. 2026. Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale. arXiv:2608.00101 [cs.AI]https://arxiv.org/abs/2608.00101
2026 arXiv
-
[21]
Hong Liu, Rui Cen, Junhan Shi, Guangshuo Qin, Jiebin Zhang, Tianyu Liu, Runzhi Fan, Guoliang Zhao, Ruobing Xie, Kai Zhang, Song Liu, Guanghua Yu, and Jianchen Zhu. 2026. AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding. arXiv:2607.25852 [cs.CL...
2026 arXiv
-
[22]
Guangda Liu, Chengwei Li, Zhenyu Ning, Jing Lin, Yiwu Yao, Danning Ke, Minyi Guo, and Jieru Zhao. 2026. FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference. arXiv:2505.13109 [cs.LG]https: //arxiv.org/abs/2505.13109
2026
-
[23]
NVIDIA Corporation. 2026. NIXL: NVIDIA Inference Xfer Library. https://github.com/ai-dynamo/nixl. Accessed August 2026
2026
-
[24]
NVIDIA Corporation. [n. d.]. Understanding Memory — CUDA C++ Programming Guide. https://docs.nvidia.com/cuda/cuda- programming-guide/02-basics/understanding-memory.html . Accessed July 2026
2026
-
[25]
Shishir G Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E Gonzalez. 2025. The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models. InForty-second International Conference on Ma...
2025
-
[26]
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024. Splitwise: Efficient Generative LLM Inference Using Phase Splitting. In2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). 118–132....
2024
-
[27]
Qwen Team. 2025. Qwen3 Technical Report. arXiv:2505.09388 [cs.CL] https://arxiv.org/abs/2505.09388
2025 arXiv
-
[28]
Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu. 2025. Mooncake: Trading More Storage for Less Computation — A KVCache-centric Architecture for Serving LLM Chatbot. In23rd USENIX Conference on File and Storage Tec...
2025
-
[29]
Graham Lopez, Matthew B
Pavel Shamis, Manjunath Gorentla Venkata, M. Graham Lopez, Matthew B. Baker, Oscar R. Hernandez, Yossi Itigin, Mike Dubman, Gilad Shainer, Richard L. Graham, Liran Liss, Yiftah Shahar, Sreeram Potluri, Davide Rossetti, Donald Becker, Duncan Poole, Christopher Lamb, Sameer Kuma...
2015
-
[30]
SGLang Project. 2026. Hierarchical KV Caching (HiCache). https: //docs.sglang.io/advanced_features/hicache.html. Accessed 2026-06-23
2026
-
[31]
Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, and Song Han. 2024. Quest: Query-Aware Sparsity for Efficient Long- Context LLM Inference. arXiv:2406.10774 [cs.CL] https://arxiv. org/abs/2406.10774
2024 arXiv
-
[32]
Hanshi Sun, Li-Wen Chang, Wenlei Bao, Size Zheng, Ningxin Zheng, Xin Liu, Harry Dong, Yuejie Chi, and Beidi Chen. 2025. ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference. arXiv:2410.21465 [cs.LG]https://arxiv.org/abs/2410.21465
2025 arXiv
-
[33]
vLLM Project. 2026. Speculative Decoding. https://docs.vll m.ai/en/latest/features/speculative_decoding/ . vLLM Documentation, accessed July 2026
2026
-
[34]
vLLM Project. 2026. Automatic Prefix Caching. https://docs.vll m.ai/en/latest/features/automatic_prefix_caching.html . Accessed 2026-06-23
2026
-
[35]
et.al. Wei An. 2024. Fire-Flyer AI-HPC: A Cost-Effective Software- Hardware Co-Design for Deep Learning. arXiv:2408.14158 [cs.DC] https://arxiv.org/abs/2408.14158
2024 arXiv
-
[36]
Yan Wang, Qifan Zhang, Jiachen Yu, Tian Liang, Dongyang Ma, Xiang Hu, Zibo Lin, Chunyang Li, Zhichao Wang, Miao Peng, Nuo Chen, Jia Li, Yujiu Yang, Haitao Mi, and Dong Yu. 2026. FlashMemory- DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention. arXiv:...
2026 arXiv
-
[37]
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024. Efficient Streaming Language Models with Attention Sinks. arXiv:2309.17453 [cs.CL] https://arxiv.org/abs/2309.1 7453
2024 arXiv
-
[38]
Yongtong Wu, Shaoyuan Chen, Yinmin Zhong, Rilin Huang, Yixuan Tan, Wentao Zhang, Liyue Zhang, Shangyan Zhou, Yuxuan Liu, Shun- feng Zhou, Mingxing Zhang, Xin Jin, and Panpan Huang. 2026. Du- alPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference. arXiv:2602...
2026
-
[39]
Jiaming Xu, Jiayi Pan, Hanzhen Wang, Yongkang Zhou, Jiancai Ye, Yu Wang, and Guohao Dai. 2026. SpeContext: Enabling Efficient Long- context Reasoning with Speculative Context Sparsity in LLMs. In Proceedings of the 31st ACM International Conference on Architectural Support for...
2026
-
[40]
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. 2024. OSWorld: Benchmarking Multi- modal Agent...
2024 arXiv
-
[41]
Jiawei Yi, Ping Gong, Youhui Bai, Zewen Jin, Shengnan Wang, Jiaqi Ruan, Jia He, Jiaan Zhu, Pengcheng Wang, Haibo Wang, Weiguang Wang, Xia Zhu, and Cheng Li. 2026. LiteCache: A Query Similarity- Driven, GPU-Centric KVCache Subsystem for Efficient LLM Inference. arXiv:2511.14510...
2026
-
[42]
Can Xiao, Sukmin Cho, Junbong We, Zhixiong Niu, Jianyi Cheng, Yiren Zhao, Youngjin Kwon, Yongqiang Xiong, Rui Ma, and Junyi Liu
Qingyue Yang, Jie Wang, Xing Li, Zhihai Wang, Chen Chen, Lei Chen, Xianzhi Yu, Wulong Liu, Jianye Hao, Mingxuan Yuan, and Bin Li. Can Xiao, Sukmin Cho, Junbong We, Zhixiong Niu, Jianyi Cheng, Yiren Zhao, Youngjin Kwon, Yongqiang Xiong, Rui Ma, and Junyi Liu
-
[43]
Aaron Zhao and Junyi Liu. 2026. Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference .Computer59, 04 (April 2026), 165–171. doi:10.1109/MC.2026.3659288
2026
-
[44]
Zihan Zhao, Baotong Lu, Shengjie Lin, Yizou Chen, Jing Liu, Yanqi Zhang, Ziming Miao, Ming-Chang Yang, Haiying Shen, Qi Chen, and Fan Yang. 2026. Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving. arXiv:2604.26837 [cs.LG] https://arxiv.or...
2026 arXiv
-
[45]
Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Yuxing Wei, Lean Wang, Zhiping Xiao, et al. 2025. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.arXiv preprint arXiv:2502.11089(2025)
2025 arXiv
-
[46]
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xu- anzhe Liu, Xin Jin, and Hao Zhang. 2024. DistServe: disaggregating prefill and decoding for goodput-optimized large language model serv- ing. InProceedings of the 18th USENIX Conference on Operating Systems Design...
2024
-
[48]
2026.HiSparse: Turbocharging Sparse Attention with Hierarchical Memory
Tingwei Huang Zhiqiang Xie, Zhangheng Huang. 2026.HiSparse: Turbocharging Sparse Attention with Hierarchical Memory. https: //www.lmsys.org/blog/2026-04-10-sglang-hisparse/
2026
-
[2025]
InAdvances in Neural Information Processing Systems, Vol
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression. InAdvances in Neural Information Processing Systems, Vol. 38. Curran Associates, Inc
-
[2026]
arXiv:2606.04511 [cs.CL] https://arxiv.org/abs/ 2606.04511
SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference. arXiv:2606.04511 [cs.CL] https://arxiv.org/abs/ 2606.04511
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.