REVIEW 3 major objections 6 minor 1 cited by
Hypic is the first system that makes position-independent caching work for hybrid-attention LLMs, cutting time-to-first-token by 3.25× on average while nearly preserving task quality.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 16:48 UTC pith:BFXK4QWP
load-bearing objection Clean systems paper that makes PIC work on hybrid-attention models via an exact linear-state composition law, with solid multi-model numbers and only minor pragmatic soft spots. the 3 major comments →
HYPIC: Accelerating Hybrid-Attention LLM Serving with Position-Independent Caching
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Hybrid-attention LLMs can support true position-independent caching once linear-attention layers cache a segment-cumulative transition operator alongside each segment’s zero-start end-state, residual full-attention layers recompute a small boundary seam window, and cold segments are prefilled in parallel across instances. The result is large, quality-preserving reductions in prefill latency and higher sustainable throughput on production hybrid models.
What carries the argument
The segment-cumulative transition operator: the product of the per-token transition matrices over a segment. Cached with the segment’s zero-start end-state, it lets any prefix recurrent state be left-multiplied and composed with the segment in constant time, removing the structural error of naive state addition.
Load-bearing premise
After independent segment prefill, the largest full-attention deviations sit at the start of each reused segment, so recomputing a fixed small window of tokens there is enough to restore cross-segment attention.
What would settle it
On a held-out hybrid model and multi-segment RAG or agent workload, compare task accuracy of the default eight-token seam against much wider seams and against full recompute of all full-attention tokens; a large accuracy jump when the seam is widened would show that boundary concentration is not sufficient.
If this is right
- Production hybrid-attention models can reuse non-contiguous RAG and agent segments the way pure full-attention stacks already can.
- Long cold requests stop being an unavoidable tail-latency tax and become parallelizable across serving instances.
- Prefill cost on multi-document and multi-turn agent prompts can fall several-fold at nearly the same task quality.
- Linear-layer cache storage stays small relative to full-attention KV because only fixed-size states and compact transitions are retained.
- Prefix caching and position-independent segment caching can run as complementary reuse paths rather than rivals.
Where Pith is reading between the lines
- As more frontier models adopt hybrid linear/full stacks, position-independent caching becomes a first-class systems primitive rather than a full-attention-only technique.
- The same transition-operator composition may transfer to pure linear or state-space sequence models outside language serving.
- Seam width may need to become adaptive if future architectures weaken or strengthen segment-boundary attention sinks.
- Segment self-containment suggests a general scatter-combine pattern for any cache unit whose prefill depends only on its own tokens.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Hypic is presented as the first serving system that enables position-independent caching (PIC) on hybrid-attention LLMs, where most layers use linear attention and a minority retain full attention. The paper identifies three obstacles that prevent existing PIC methods from transferring: (i) linear layers expose only a per-request recurrent state rather than per-token KV, so splice/correction primitives fail; (ii) full-attention correction in a hybrid stack lacks the per-token hidden states that linear layers suppress; and (iii) cold long prefills remain monolithic under prior PIC. The proposed remedies are a cached segment-cumulative transition operator T_C together with the zero-start end-state for constant-time linear-state composition (Eq. 6), a small boundary seam window (default w=8) that propagates hidden states to repair full-attention cross-segment attention, and inter-instance segment parallelism (scatter/combine with LPT) for cold misses. Implemented on SGLang and evaluated on four hybrid models and five workloads, Hypic reports 3.25× average TTFT reduction and 1.66× QPS improvement over Prefix Cache, with a 1.71-point average quality gap from Full Recompute, plus a 5.7× cold-prefill speedup at 8 workers.
Significance. The work addresses a timely and practically important collision between two production trends: hybrid-attention architectures (Qwen3.5, Ring, MiniMax-M1, Kimi-Linear, etc.) and PIC for RAG/agentic segment assembly. The linear-attention contribution is algebraically clean: T_C is derived directly from the published recurrence families in Table 1, is independent of the prefix, and yields layer-exact composition under matching hidden states, with measured near-FP16 fidelity at layer 0 and bounded deep-layer drift (Table 4). Segment parallelism is a genuine systems lever that prior PIC work left unused. Strengths include a unified treatment across scalar/diagonal/dense transition families (Table 3), an end-to-end implementation with public/private cache pools, and multi-model multi-workload evaluation against the production baseline (Prefix Cache) and Full Recompute. If the results hold under broader deployment, the paper is a solid systems contribution for hybrid LLM serving.
major comments (3)
- §3.2 Insight 2 and Fig. 3: The seam-window design rests on the claim that full-attention deviations concentrate at segment beginnings. Fig. 3 shows this for a single layer/head (layer 7, head 0) on one prompt. While §6.4’s w-sweep and the 1.71-point quality envelope support practical sufficiency of w=8, the locality claim would be much stronger with aggregate statistics (e.g., fraction of attention-weighted KV deviation mass in the first w tokens, across layers, heads, and prompts). Without that, it remains unclear how often a fixed small window is enough versus cases where deviation is more diffuse.
- §6.2 Prod-RAG (Fig. 11): The TTFT–QPS and throughput gains are reported after a 10-minute warm-up with residual cold misses from churn/eviction, but segment hit rates, miss rates, and the fraction of TTFT spent on cold vs. hit paths are not stated. Because Hypic’s gains come from both PIC hits (composition + seam) and segment parallelism on misses, the attribution of the 1.66× QPS / 3.25× TTFT numbers is hard to interpret without those rates. Please report hit/miss statistics (and ideally a breakdown of TTFT components under load) so readers can judge how much each contribution drives the headline results.
- §4.2 / Table 3 and §6.3: For dense-transition models (Qwen3.5), T_C is a full d_k×d_k matrix (~32 KB per head per layer). Construction overhead is shown as 5.2–6.7% of prefill (Fig. 13), but end-to-end public-pool HBM footprint versus Prefix Cache / Full Recompute under realistic segment cardinality and cache capacity is not quantified. Given that the public pool is shared and LRU-managed, a short capacity/pressure experiment (or a closed-form estimate at the evaluated segment counts) would confirm that the dense-family storage cost does not erode the claimed serving gains under memory pressure.
minor comments (6)
- Abstract and §1: “near-exact” composition is accurate at layer 0 under matching hidden states, but deep-layer drift is ~9% relative L2 (Table 4). A brief clarifying phrase in the abstract (e.g., “algebraically exact under matching hidden states; bounded residual drift in deeper layers”) would avoid over-reading.
- Fig. 10: Accuracy–TTFT scatter is dense; a small table of absolute F1/ROUGE-L and p50 TTFT per cell would make the 1.71-point average gap easier to verify without visual estimation.
- §4.2 Causal convolution warm-up and RoPE re-rotation: these are important for Qwen3.5 and Ring respectively; a short note on which evaluated models use which path would help readers map Table 3 families to the four models.
- §5 Implementation: “14k lines of Python and Triton” is useful; if any of the FLA double-invocation or NIXL transfer path is planned for release, stating that would strengthen reproducibility.
- Notation: T_C is introduced as both the product of per-token T_t and the cached object; a single consistent definition near Eq. (4)–(6) would reduce momentary ambiguity when reading Table 3’s closed forms.
- Related work: CacheBlend, EPIC, MEPIC, and recent hybrid models are cited appropriately; a one-sentence explicit statement that prior PIC systems cannot run on hybrid stacks (hence no direct numerical comparison) would pre-empt a common reviewer question.
Circularity Check
No significant circularity: linear-state composition is algebraic from the published recurrence; seam width and speedups are measured, not forced by construction.
full rationale
The load-bearing linear-attention claim is derived by unrolling the standard advanced linear-attention recurrence S_i = T_i S_{i-1} + u_i (Eq. 2, Tab. 1) to obtain the exact composition S_{C1 C2|0} = T_{C2} S_{C1|0} + S_{C2|0} (Eq. 4) and the multi-segment law (Eq. 6). Caching the segment-cumulative transition T_C with the zero-start end-state is the natural algebraic dual of that unrolling; layer-0 fidelity is reported against full recompute (relative norm ~6e-5), and residual deep-layer drift is measured separately (Tab. 4), not assumed away. Seam windows rest on an empirical concentration observation (Fig. 3) and a fixed default w=8 chosen by an accuracy–TTFT sweep (§6.4), not on a fitted parameter that is then re-labeled as a prediction. Segment parallelism follows from PIC self-containment and is evaluated by cold-miss scaling. Headline TTFT/QPS/quality numbers are end-to-end measurements against Full Recompute and Prefix Cache on public and production workloads. Self-citations (e.g., EPIC, CacheSlide) supply background on full-attention PIC and are not used as uniqueness theorems that force the hybrid design. No step reduces a claimed prediction to its own inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (1)
- seam window width w =
8
axioms (4)
- domain assumption Advanced linear-attention layers obey the unified recurrence S_i = T_i S_{i-1} + u_i (Eq. 2) for the scalar, diagonal and dense families listed in Table 1.
- domain assumption After independent segment prefill, the largest full-attention deviations concentrate at the beginning of each reused segment.
- domain assumption Application-provided segment boundaries correctly identify semantically independent reusable units.
- standard math Standard floating-point arithmetic and RoPE re-rotation preserve the algebraic identities used for composition and key re-positioning.
invented entities (3)
-
segment-cumulative transition operator T_C
independent evidence
-
seam window
no independent evidence
-
segment parallelism (scatter/combine workers + LPT)
independent evidence
read the original abstract
In retrieval-augmented generation and agentic LLM serving, prompts are assembled from independent segments into long contexts, making the prefill stage dominate per-request cost. Two directions have emerged to reduce this cost: position-independent caching (PIC) admits KV reuse for non-contiguous segments shared across requests, while hybrid-attention models cut computation by replacing most full-attention layers with linear attention. However, they cannot coexist: applying existing PIC methods to hybrid-attention models breaks down because per-token KV-cache reuse primitives do not transfer to the per-request recurrent state. We present Hypic, the first system to accelerate hybrid-attention LLM serving with position-independent caching. For linear-attention layers, we identify the segment-cumulative transition operator as the missing algebraic primitive and cache it alongside each segment's zero-start end-state, enabling near-exact and constant-time composition of independently cached segments. For the remaining full-attention layers, existing PIC methods also fail because linear layers do not expose the per-token hidden states needed for selective recomputation. We show that the largest deviations concentrate at segment beginnings and construct a small seam window that propagates hidden states through the hybrid-attention stack to repair cross-segment attention. Finally, Hypic introduces segment parallelism, which exploits PIC's segment-level self-containment to parallelize cache-miss prefill across instances, turning long cold requests into an accelerable workload. Evaluated across four hybrid-attention models and five workloads, Hypic reduces time-to-first-token by $3.25\times$ on average and improves QPS by $1.66\times$ over Prefix Cache, while preserving task quality with a 1.71-point gap from Full Recompute.
Figures
Forward citations
Cited by 1 Pith paper
-
Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV
Source-omitted event KV rows can carry compact upstream state (semantic materialization); deliberate answer-free carriers raise recovery from 6% to 51% on Qwen3-8B, while passive natural mentions do not.
Reference graph
Works this paper leans on
-
[1]
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ram- jee. 2024. Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’24). USENIX Association, 117–134
2024
-
[2]
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhid- ian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al
-
[3]
InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics
LongBench: A Bilingual, Multitask Benchmark for Long Con- text Understanding. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 3119–3137
-
[4]
Ziyi Cao, Qingsi Si, Jingbin Zhang, and Bingquan Liu. 2026. Sparse Attention Across Multiple-Context KV Cache. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 30165–30173. doi:10. 1609/aaai.v40i36.40266
2026
-
[5]
Aili Chen, Aonian Li, Bangwei Gong, Binyang Jiang, Bo Fei, Bo Yang, Boji Shan, Changqing Yu, Chao Wang, Cheng Zhu, et al. 2025. MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.arXiv preprint arXiv:2506.13585(2025)
Pith/arXiv arXiv 2025
-
[6]
Chuangtao Chen, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Bing Li, and Ulf Schlichtmann. 2026. KV Packet: Recomputation-Free Context- Independent KV Caching for LLMs. arXiv:2604.13226 [cs.CL]
Pith/arXiv arXiv 2026
-
[7]
Tri Dao and Albert Gu. 2024. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Du- ality. InProceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235). PMLR, 10041–10071
2024
-
[8]
Alexander Richard Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir Radev. 2019. Multi-News: A Large-Scale Multi-Document Summariza- tion Dataset and Abstractive Hierarchical Model. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 1074–1084
2019
-
[9]
In Gim, Guojun Chen, Seung-Seob Lee, Nikhil Sarda, Anurag Khandel- wal, and Lin Zhong. 2024. Prompt Cache: Modular Attention Reuse for Low-Latency Inference. InProceedings of Machine Learning and Systems, Vol. 6
2024
-
[10]
Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer
-
[11]
InProceedings of the 2nd Workshop on New Frontiers in Summarization (EMNLP-IJCNLP 2019 Workshop)
SAMSum Corpus: A Human-Annotated Dialogue Dataset for Abstractive Summarization. InProceedings of the 2nd Workshop on New Frontiers in Summarization (EMNLP-IJCNLP 2019 Workshop). 70–79
2019
-
[12]
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing a Multi-hop QA Dataset for Compre- hensive Evaluation of Reasoning Steps. InProceedings of the 28th International Conference on Computational Linguistics. 6609–6625
2020
-
[13]
Junhao Hu, Wenrui Huang, Weidong Wang, Haoyi Wang, Tiancheng Hu, Qin Zhang, Hao Feng, Xusheng Chen, Yizhou Shan, and Tao Xie
-
[14]
InProceedings of the 42nd International Confer- ence on Machine Learning (Proceedings of Machine Learning Research, Vol
EPIC: Efficient Position-Independent Caching for Serving Large Language Models. InProceedings of the 42nd International Confer- ence on Machine Learning (Proceedings of Machine Learning Research, Vol. 267). PMLR, 24391–24402
-
[15]
Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, and Lu Wang. 2021. Efficient Attentions for Long Document Summarization. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics. 1419–1436
2021
-
[16]
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. Swe-bench: Can language models resolve real-world github issues?. InInternational Conference on Learning Representations, Vol. 2024. 54107–54157
2024
-
[17]
Weld, and Luke Zettlemoyer
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer
-
[18]
InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Vol
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Vol. 1. 1601–1611
-
[19]
Fu, Christo- pher Ré, and Azalia Mirhoseini
Jordan Juravsky, Bradley Brown, Ryan Ehrlich, Daniel Y. Fu, Christo- pher Ré, and Azalia Mirhoseini. 2024. Hydragen: High-Throughput LLM Inference with Shared Prefixes. InProceedings of the 41st Interna- tional Conference on Machine Learning (ICML ’24)
2024
-
[20]
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020. Transformers are RNNs: Fast Autoregressive Trans- formers with Linear Attention. InProceedings of the 37th International Conference on Machine Learning (ICML)
2020
-
[21]
Kimi Team. 2025. Kimi Linear: An Expressive, Efficient Attention Architecture. arXiv:2510.26692 [cs.CL]https://arxiv.org/abs/2510. 26692
Pith/arXiv arXiv 2025
-
[22]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica
-
[23]
InProceedings of the 29th ACM Symposium on Operating Systems Principles (SOSP ’23)
Efficient Memory Management for Large Language Model Serv- ing with PagedAttention. InProceedings of the 29th ACM Symposium on Operating Systems Principles (SOSP ’23). ACM, 611–626
-
[24]
Shenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li, and Yang You. 2023. Sequence Parallelism: Long Sequence Training from System Perspective. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL). 2391–2404
2023
-
[25]
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, et al . 2024. Jamba: A hybrid transformer-mamba language model.arXiv preprint arXiv:2403.19887(2024)
Pith/arXiv arXiv 2024
-
[26]
Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. InText Summarization Branches Out. 74–81
2004
-
[27]
Yang Liu, Yunfei Gu, Liqiang Zhang, Chentao Wu, Guangtao Xue, Jie Li, Minyi Guo, Junhao Hu, and Jie Meng. 2026. CacheSlide: Unlocking Cross Position-Aware KV Cache Reuse for Accelerating LLM Serv- ing. InProceedings of the 24th USENIX Conference on File and Storage Technologies (FAST ’26). USENIX Association, 83–99
2026
-
[28]
Songshuo Lu, Hua Wang, Yutian Rong, Zhi Chen, and Yaohua Tang
-
[29]
InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 6588–6601. doi:10.18653/v1/2025.emnlp-main.334
-
[30]
Dongyang Ma, Yan Wang, and Tian Lan. 2025. Block-Attention for Effi- cient Prefilling. InThe Thirteenth International Conference on Learning Representations
2025
-
[31]
NVIDIA. 2026. NVIDIA Inference Xfer Library (NIXL). Accessed: 2026-07-08.https://github.com/ai-dynamo/nixl
2026
-
[32]
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024. Splitwise: Efficient Generative LLM Inference Using Phase Splitting. InProceedings of the 51st Annual International Symposium on Computer Architecture (ISCA ’24). IEEE, 118–132
2024
-
[33]
Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Minxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu. 2025. Mooncake: Trading More Storage for Less Computation – A KVCache-Centric Architecture for Serving LLM Chatbot. In23rd USENIX Conference on File and Storage Technologies (FAST ’25). USENIX Association. 13 Liu et al
2025
-
[34]
Zhen Qin, Weigao Sun, Dong Li, Xuyang Shen, Weixuan Sun, and Yiran Zhong. 2024. Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models. arXiv:2401.04658 [cs.CL]
Pith/arXiv arXiv 2024
-
[35]
Qwen Team. 2026. Qwen3.5: Towards Native Multimodal Agents. Qwen Technical Blog.https://qwen.ai/blog?id=qwen3.5
2026
-
[36]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang
-
[37]
InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing
SQuAD: 100,000+ Questions for Machine Comprehension of Text. InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2383–2392
2016
-
[38]
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. arXiv preprint arXiv:1909.08053(2019)
Pith/arXiv arXiv 2019
-
[39]
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2024. RoFormer: Enhanced Transformer with Rotary Position Embedding.Neurocomputing568 (2024), 127063
2024
-
[40]
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. 2023. Retentive Net- work: A Successor to Transformer for Large Language Models. arXiv:2307.08621 [cs.CL]
Pith/arXiv arXiv 2023
-
[41]
Ling Team, Bin Han, Caizhi Tang, Chen Liang, Donghao Zhang, Fan Yuan, Feng Zhu, Jie Gao, Jingyu Hu, Longfei Li, et al . 2025. Every attention matters: An efficient hybrid architecture for long-context reasoning.arXiv preprint arXiv:2510.19338(2025)
arXiv 2025
-
[42]
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. MuSiQue: Multihop Questions via Single-hop Ques- tion Composition.Transactions of the Association for Computational Linguistics10 (2022), 539–554
2022
-
[43]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need. InAdvances in Neural Information Processing Systems, Vol. 30
2017
-
[44]
Jiahao Wang, Weiyu Xie, Mingxing Zhang, Boxing Zhang, Jianwei Dong, Yuening Zhu, Chen Lin, Jinqi Tang, Yaochen Han, Zhiyuan Ai, Xianglin Chen, Yongwei Wu, and Congfeng Jiang. 2026. From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval- Augmented Generation.Proceedings of the ACM on Management of Data4, 1 (2026). doi:10.1145/3786655
doi:10.1145/3786655 2026
-
[45]
Qian Wang, Zahra Yousefijamarani, Morgan Lindsay Heisler, Rongzhi Gu, Xiaolong Bai, Yizhou Shan, Wei Zhang, Lan Wang, Ying Xiong, Yong Zhang, and Zhenan Fan. 2025. MEPIC: Memory Efficient Position Independent Caching for LLM Serving. arXiv:2512.16822 [cs.LG]
arXiv 2025
-
[46]
Shihao Wang, Jiahao Chen, Yanqi Pan, Hao Huang, Yichen Hao, Xi- angyu Zou, Wen Xia, Wentao Zhang, Chongyang Qiu, and Pengfei Wang. 2026. ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation. arXiv:2602.02579 [cs.AI]
arXiv 2026
-
[47]
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024. Efficient Streaming Language Models with Attention Sinks. InThe Twelfth International Conference on Learning Representa- tions
2024
-
[48]
Bin Yang, Qiuyu Leng, Jun Zeng, and Zhenhua Wu. 2025. CacheClip: Accelerating RAG with Effective KV Cache Reuse. arXiv:2510.10129 [cs.LG]
Pith/arXiv arXiv 2025
-
[49]
Huan Yang, Renji Zhang, Mingzhe Huang, Weijun Wang, Yin Tang, Yuanchun Li, Yunxin Liu, and Deyu Zhang. 2025. KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse. arXiv:2503.16525 [cs.LG]
Pith/arXiv arXiv 2025
-
[50]
Jingbo Yang, Bairu Hou, Wei Wei, Yujia Bao, and Shiyu Chang. 2025. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse. InAdvances in Neural Information Processing Systems, Vol. 38
2025
-
[51]
Songlin Yang, Jan Kautz, and Ali Hatamizadeh. 2025. Gated Delta Networks: Improving Mamba2 with Delta Rule. InThe Thirteenth International Conference on Learning Representations
2025
-
[52]
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. 2024. Gated Linear Attention Transformers with Hardware- Efficient Training. InProceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235). PMLR, 56501–56523
2024
-
[53]
Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim
-
[54]
InAdvances in Neural Information Processing Systems, Vol
Parallelizing Linear Transformers with the Delta Rule over Se- quence Length. InAdvances in Neural Information Processing Systems, Vol. 37
-
[55]
Songlin Yang and Yu Zhang. 2024. FLA: A Triton-Based Library for Hardware-Efficient Implementations of Linear Attention Mechanism. https://github.com/fla-org/flash-linear-attention
2024
-
[56]
Cohen, Ruslan Salakhutdinov, and Christopher D
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2369–2380
2018
-
[57]
Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, and Junchen Jiang. 2025. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowl- edge Fusion. InProceedings of the Twentieth European Conference on Computer Systems (EuroSys ’25). ACM. doi:10.1145/3689031.3696098
-
[58]
Hancheng Ye, Zhengqi Gao, Mingyuan Ma, Qinsi Wang, Yuzhe Fu, Ming-Yu Chung, Yueqian Lin, Zhijian Liu, Jianyi Zhang, Danyang Zhuo, and Yiran Chen. 2025. KVCOMM: Online Cross-context KV Cache Communication for Efficient LLM-based Multi-agent Systems. InAdvances in Neural Information Processing Systems
2025
-
[59]
Lu Ye, Ze Tao, Yong Huang, and Yang Li. 2024. ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 11608–11620
2024
-
[60]
Qingfei Zhao, Ruobing Wang, Yukuo Cen, Daren Zha, Shicheng Tan, Yuxiao Dong, and Jie Tang. 2024. LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Ques- tion Answering. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 22600–22632
2024
-
[61]
Shiju Zhao, Junhao Hu, Jiaqi Zheng, and Guihai Chen. 2026. You Need an Encoder for Native Position-Independent Caching. arXiv:2602.01519 [cs.CL]https://arxiv.org/abs/2602.01519
arXiv 2026
-
[62]
Gonzalez, Clark Barrett, and Ying Sheng
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024. SGLang: Efficient Execution of Structured Language Model Programs. InAdvances in Neural Information Processing Systems
2024
-
[63]
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xu- anzhe Liu, Xin Jin, and Hao Zhang. 2024. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’24). USENIX Association, 193–210
2024
-
[64]
Yuechi Zhou, Yi Su, Jianxin Zhang, Juntao Li, Qingrong Xia, Zhefeng Wang, Xinyu Duan, and Baoxing Huai. 2025. A3: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving. arXiv:2511.17560 [cs.CL] 14
arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.