Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Hypic is the first system that makes position-independent caching work for hybrid-attention LLMs, cutting time-to-first-token by 3.25× on average while nearly preserving task quality.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 16:48 UTC pith:BFXK4QWP

load-bearing objection Clean systems paper that makes PIC work on hybrid-attention models via an exact linear-state composition law, with solid multi-model numbers and only minor pragmatic soft spots. the 3 major comments →

arxiv 2607.01299 v2 pith:BFXK4QWP submitted 2026-07-01 cs.DC

HYPIC: Accelerating Hybrid-Attention LLM Serving with Position-Independent Caching

classification cs.DC
keywords position-independent cachinghybrid attentionlinear attentionKV cacheLLM servingRAGsegment parallelismtime-to-first-token
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

In retrieval-augmented generation and agent workloads, long prompts are assembled from independent segments, so prefill dominates serving cost. Position-independent caching reuses those segments behind arbitrary prefixes, and hybrid models replace most full-attention layers with linear attention; until now the two could not be combined, because linear layers expose only a fixed-size recurrent state with no per-token KV handle for splice-and-correct methods. Hypic supplies the missing algebraic object—the segment-cumulative transition operator—caches it with each segment’s zero-start end-state, and composes independently prefilled segments in constant time. For residual full-attention layers it recomputes a small fixed window at each segment start so hidden states can propagate through the hybrid stack and repair cross-segment attention. It also parallelizes cold-segment prefill across instances. Across four hybrid models and five workloads, Hypic reduces TTFT by 3.25× and raises sustainable QPS by 1.66× over prefix caching, with only a 1.71-point gap from full recompute.

Core claim

Hybrid-attention LLMs can support true position-independent caching once linear-attention layers cache a segment-cumulative transition operator alongside each segment’s zero-start end-state, residual full-attention layers recompute a small boundary seam window, and cold segments are prefilled in parallel across instances. The result is large, quality-preserving reductions in prefill latency and higher sustainable throughput on production hybrid models.

What carries the argument

The segment-cumulative transition operator: the product of the per-token transition matrices over a segment. Cached with the segment’s zero-start end-state, it lets any prefix recurrent state be left-multiplied and composed with the segment in constant time, removing the structural error of naive state addition.

Load-bearing premise

After independent segment prefill, the largest full-attention deviations sit at the start of each reused segment, so recomputing a fixed small window of tokens there is enough to restore cross-segment attention.

What would settle it

On a held-out hybrid model and multi-segment RAG or agent workload, compare task accuracy of the default eight-token seam against much wider seams and against full recompute of all full-attention tokens; a large accuracy jump when the seam is widened would show that boundary concentration is not sufficient.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Production hybrid-attention models can reuse non-contiguous RAG and agent segments the way pure full-attention stacks already can.
  • Long cold requests stop being an unavoidable tail-latency tax and become parallelizable across serving instances.
  • Prefill cost on multi-document and multi-turn agent prompts can fall several-fold at nearly the same task quality.
  • Linear-layer cache storage stays small relative to full-attention KV because only fixed-size states and compact transitions are retained.
  • Prefix caching and position-independent segment caching can run as complementary reuse paths rather than rivals.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • As more frontier models adopt hybrid linear/full stacks, position-independent caching becomes a first-class systems primitive rather than a full-attention-only technique.
  • The same transition-operator composition may transfer to pure linear or state-space sequence models outside language serving.
  • Seam width may need to become adaptive if future architectures weaken or strengthen segment-boundary attention sinks.
  • Segment self-containment suggests a general scatter-combine pattern for any cache unit whose prefill depends only on its own tokens.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. Hypic is presented as the first serving system that enables position-independent caching (PIC) on hybrid-attention LLMs, where most layers use linear attention and a minority retain full attention. The paper identifies three obstacles that prevent existing PIC methods from transferring: (i) linear layers expose only a per-request recurrent state rather than per-token KV, so splice/correction primitives fail; (ii) full-attention correction in a hybrid stack lacks the per-token hidden states that linear layers suppress; and (iii) cold long prefills remain monolithic under prior PIC. The proposed remedies are a cached segment-cumulative transition operator T_C together with the zero-start end-state for constant-time linear-state composition (Eq. 6), a small boundary seam window (default w=8) that propagates hidden states to repair full-attention cross-segment attention, and inter-instance segment parallelism (scatter/combine with LPT) for cold misses. Implemented on SGLang and evaluated on four hybrid models and five workloads, Hypic reports 3.25× average TTFT reduction and 1.66× QPS improvement over Prefix Cache, with a 1.71-point average quality gap from Full Recompute, plus a 5.7× cold-prefill speedup at 8 workers.

Significance. The work addresses a timely and practically important collision between two production trends: hybrid-attention architectures (Qwen3.5, Ring, MiniMax-M1, Kimi-Linear, etc.) and PIC for RAG/agentic segment assembly. The linear-attention contribution is algebraically clean: T_C is derived directly from the published recurrence families in Table 1, is independent of the prefix, and yields layer-exact composition under matching hidden states, with measured near-FP16 fidelity at layer 0 and bounded deep-layer drift (Table 4). Segment parallelism is a genuine systems lever that prior PIC work left unused. Strengths include a unified treatment across scalar/diagonal/dense transition families (Table 3), an end-to-end implementation with public/private cache pools, and multi-model multi-workload evaluation against the production baseline (Prefix Cache) and Full Recompute. If the results hold under broader deployment, the paper is a solid systems contribution for hybrid LLM serving.

major comments (3)
  1. §3.2 Insight 2 and Fig. 3: The seam-window design rests on the claim that full-attention deviations concentrate at segment beginnings. Fig. 3 shows this for a single layer/head (layer 7, head 0) on one prompt. While §6.4’s w-sweep and the 1.71-point quality envelope support practical sufficiency of w=8, the locality claim would be much stronger with aggregate statistics (e.g., fraction of attention-weighted KV deviation mass in the first w tokens, across layers, heads, and prompts). Without that, it remains unclear how often a fixed small window is enough versus cases where deviation is more diffuse.
  2. §6.2 Prod-RAG (Fig. 11): The TTFT–QPS and throughput gains are reported after a 10-minute warm-up with residual cold misses from churn/eviction, but segment hit rates, miss rates, and the fraction of TTFT spent on cold vs. hit paths are not stated. Because Hypic’s gains come from both PIC hits (composition + seam) and segment parallelism on misses, the attribution of the 1.66× QPS / 3.25× TTFT numbers is hard to interpret without those rates. Please report hit/miss statistics (and ideally a breakdown of TTFT components under load) so readers can judge how much each contribution drives the headline results.
  3. §4.2 / Table 3 and §6.3: For dense-transition models (Qwen3.5), T_C is a full d_k×d_k matrix (~32 KB per head per layer). Construction overhead is shown as 5.2–6.7% of prefill (Fig. 13), but end-to-end public-pool HBM footprint versus Prefix Cache / Full Recompute under realistic segment cardinality and cache capacity is not quantified. Given that the public pool is shared and LRU-managed, a short capacity/pressure experiment (or a closed-form estimate at the evaluated segment counts) would confirm that the dense-family storage cost does not erode the claimed serving gains under memory pressure.
minor comments (6)
  1. Abstract and §1: “near-exact” composition is accurate at layer 0 under matching hidden states, but deep-layer drift is ~9% relative L2 (Table 4). A brief clarifying phrase in the abstract (e.g., “algebraically exact under matching hidden states; bounded residual drift in deeper layers”) would avoid over-reading.
  2. Fig. 10: Accuracy–TTFT scatter is dense; a small table of absolute F1/ROUGE-L and p50 TTFT per cell would make the 1.71-point average gap easier to verify without visual estimation.
  3. §4.2 Causal convolution warm-up and RoPE re-rotation: these are important for Qwen3.5 and Ring respectively; a short note on which evaluated models use which path would help readers map Table 3 families to the four models.
  4. §5 Implementation: “14k lines of Python and Triton” is useful; if any of the FLA double-invocation or NIXL transfer path is planned for release, stating that would strengthen reproducibility.
  5. Notation: T_C is introduced as both the product of per-token T_t and the cached object; a single consistent definition near Eq. (4)–(6) would reduce momentary ambiguity when reading Table 3’s closed forms.
  6. Related work: CacheBlend, EPIC, MEPIC, and recent hybrid models are cited appropriately; a one-sentence explicit statement that prior PIC systems cannot run on hybrid stacks (hence no direct numerical comparison) would pre-empt a common reviewer question.

Circularity Check

0 steps flagged

No significant circularity: linear-state composition is algebraic from the published recurrence; seam width and speedups are measured, not forced by construction.

full rationale

The load-bearing linear-attention claim is derived by unrolling the standard advanced linear-attention recurrence S_i = T_i S_{i-1} + u_i (Eq. 2, Tab. 1) to obtain the exact composition S_{C1 C2|0} = T_{C2} S_{C1|0} + S_{C2|0} (Eq. 4) and the multi-segment law (Eq. 6). Caching the segment-cumulative transition T_C with the zero-start end-state is the natural algebraic dual of that unrolling; layer-0 fidelity is reported against full recompute (relative norm ~6e-5), and residual deep-layer drift is measured separately (Tab. 4), not assumed away. Seam windows rest on an empirical concentration observation (Fig. 3) and a fixed default w=8 chosen by an accuracy–TTFT sweep (§6.4), not on a fitted parameter that is then re-labeled as a prediction. Segment parallelism follows from PIC self-containment and is evaluated by cold-miss scaling. Headline TTFT/QPS/quality numbers are end-to-end measurements against Full Recompute and Prefix Cache on public and production workloads. Self-citations (e.g., EPIC, CacheSlide) supply background on full-attention PIC and are not used as uniqueness theorems that force the hybrid design. No step reduces a claimed prediction to its own inputs by construction.

Axiom & Free-Parameter Ledger

1 free parameters · 4 axioms · 3 invented entities

The central claims rest on the standard linear-attention recurrence already published for the evaluated model families, the empirical observation that attention/KV deviation concentrates at segment starts, and a small hand-chosen seam width. No new physical entities are postulated; the transition operator is algebraically derived rather than invented. Free parameters are limited to the seam width and ordinary systems knobs (cache sizes, LPT).

free parameters (1)
  • seam window width w = 8
    Default w=8 tokens chosen after a sensitivity sweep on two model/dataset pairs; larger values raise TTFT without meaningful accuracy gain. Directly controls the quality–latency trade-off of full-attention repair.
axioms (4)
  • domain assumption Advanced linear-attention layers obey the unified recurrence S_i = T_i S_{i-1} + u_i (Eq. 2) for the scalar, diagonal and dense families listed in Table 1.
    Taken from the cited model papers (RetNet, GLA, DeltaNet, GDN, KDA, Qwen3.5, Ring, etc.); the composition law is derived from it.
  • domain assumption After independent segment prefill, the largest full-attention deviations concentrate at the beginning of each reused segment.
    Empirical observation (Fig. 3, Insight 2) used to justify a constant-size seam rather than arbitrary-token recompute.
  • domain assumption Application-provided segment boundaries correctly identify semantically independent reusable units.
    Assumed throughout; PIC_SEPARATOR markers and RAG/agent templates supply the boundaries.
  • standard math Standard floating-point arithmetic and RoPE re-rotation preserve the algebraic identities used for composition and key re-positioning.
    Used for fidelity claims and RoPE adjustment (Eqs. 7–8).
invented entities (3)
  • segment-cumulative transition operator T_C independent evidence
    purpose: Algebraic primitive that lets independently cached linear-attention segments be composed in constant time without structural error.
    Derived by unrolling the existing recurrence; not a free postulate. Cached alongside the zero-start end-state.
  • seam window no independent evidence
    purpose: Small fixed-length recomputation region at segment starts that restores cross-segment full attention without per-token recurrent-state storage.
    Engineering construct justified by the observed deviation locality; width is a free parameter.
  • segment parallelism (scatter/combine workers + LPT) independent evidence
    purpose: Inter-instance scheme that parallelizes cold-segment prefill by exploiting PIC self-containment.
    Systems mechanism; correctness follows from segment independence already granted by PIC.

pith-pipeline@v1.1.0-grok45 · 25964 in / 2911 out tokens · 31483 ms · 2026-07-14T16:48:07.927118+00:00 · methodology

0 comments
read the original abstract

In retrieval-augmented generation and agentic LLM serving, prompts are assembled from independent segments into long contexts, making the prefill stage dominate per-request cost. Two directions have emerged to reduce this cost: position-independent caching (PIC) admits KV reuse for non-contiguous segments shared across requests, while hybrid-attention models cut computation by replacing most full-attention layers with linear attention. However, they cannot coexist: applying existing PIC methods to hybrid-attention models breaks down because per-token KV-cache reuse primitives do not transfer to the per-request recurrent state. We present Hypic, the first system to accelerate hybrid-attention LLM serving with position-independent caching. For linear-attention layers, we identify the segment-cumulative transition operator as the missing algebraic primitive and cache it alongside each segment's zero-start end-state, enabling near-exact and constant-time composition of independently cached segments. For the remaining full-attention layers, existing PIC methods also fail because linear layers do not expose the per-token hidden states needed for selective recomputation. We show that the largest deviations concentrate at segment beginnings and construct a small seam window that propagates hidden states through the hybrid-attention stack to repair cross-segment attention. Finally, Hypic introduces segment parallelism, which exploits PIC's segment-level self-containment to parallelize cache-miss prefill across instances, turning long cold requests into an accelerable workload. Evaluated across four hybrid-attention models and five workloads, Hypic reduces time-to-first-token by $3.25\times$ on average and improves QPS by $1.66\times$ over Prefix Cache, while preserving task quality with a 1.71-point gap from Full Recompute.

Figures

Figures reproduced from arXiv: 2607.01299 by Junhao Hu, Juntong Wu, Minghao Li, Weihang Chen, Xiaoxu Chen, Yang Liu, Yifei Liu.

Figure 1
Figure 1. Figure 1: Existing PIC methods reuse per-token KV cache in full-attention models via splice and correction (left); on hybrid stacks, both primitives fail because linear-attention layers expose only a per-request recurrent state, with no per-token handle (right). 1 Introduction Large language model (LLM) serving is shifting from single￾turn chat toward retrieval-augmented question answer￾ing [11, 15, 34, 47], multi-d… view at source ↗
Figure 2
Figure 2. Figure 2: Memory-access footprint of correction. (a) Full￾attention stack: every token’s prefix state is in the KV cache, so correction can read it directly. (b) Hybrid stack: linear layers retain only the per-request recurrent state, leaving non-final tokens’ prefix states uncached. C1: Thenum District -> Chrysan Company. C2: Derek lives in Thenum District. C3: answer with company name only. Query: Which company do… view at source ↗
Figure 3
Figure 3. Figure 3 [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: (a) Cache-miss prefill under PDC and existing PIC vs. (b) Parallel execution enabled by segment self￾containment. However, this migration assumes a prerequisite that does not hold in a hybrid stack. As shown in [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 3
Figure 3. Figure 3: Deviations between Full Recompute and Naive Splice for Qwen3.5-35B-A3B layer 7, head 0. 𝑆 (ℓ) 𝑖−1 requires 𝑖 − 1 recurrence steps from 𝑆0, so recomputing a single token degrades to re-running all preceding tokens, eliminating the caching benefit entirely. Neither option is acceptable. We therefore ask: is the ability to recompute an arbitrary token truly necessary? As a diagnostic, we run Qwen3.5- 35B-A3B … view at source ↗
Figure 6
Figure 6. Figure 6: Linear-attention state composition with cached transitions. Each segment caches the tuple (𝑇𝐶, 𝑆𝐶|0) at first prefill; at reuse time Hypic composes the prefix end-state and the cached tuples via Equation (6). segment-cumulative transition operator—a quantity computed as a transient intermediate at every recurrence step yet never persisted by current serving systems. To address this, Hypic caches not only t… view at source ↗
Figure 7
Figure 7. Figure 7: Seam window across adjacent segments (𝐶1,𝐶2): the last 𝑤 tokens of 𝐶1 and the first 𝑤 tokens of 𝐶2 are excluded from each segment’s cached state and recomputed jointly at splice time. At cache time, Hypic stores the zero-start end-state 𝑆𝐶|0 in the public pool. At reuse time, Hypic replaces each 𝑆𝐶𝑖 |0 in Equation (6) with 𝑅(𝑝𝑖) 𝑆𝐶𝑖 |0 before the prefix𝑇 -products act, where 𝑝𝑖 is segment 𝐶𝑖 ’s global star… view at source ↗
Figure 7
Figure 7. Figure 7: Seam-window recomputation. Hypic excludes the first 𝑤 tokens of each interior segment from the cached KV and recomputes them under the assembled prefix at reuse time. The two boundary segments of a request are handled specially. The leading segment is typically the system prompt, which has no left neighbor and always anchors at position 0, so no cross-segment deviation needs repair and Hypic caches it in f… view at source ↗
Figure 8
Figure 8. Figure 8: Seam-window handling at linear-attention layers. Each segment caches (𝑇𝐶, 𝑆𝐶|0) over interior tokens only; at splice time Hypic recomputes the seam window’s own 𝑇 and 𝑆 on the fly and inserts them into the composition law, jointly advancing the running state and forwarding per￾token outputs to the layer above. key from start 𝑎 to start 𝑏 reduces to one left-multiplication: 𝐾𝑏 = 𝑅(𝑏 − 𝑎) 𝐾𝑎. (8) These inter… view at source ↗
Figure 9
Figure 9. Figure 9: Accelerate long cold requests with segment paral￾lelism. The Hypic Router probes hit status for each segment (Seg 1, 3 hit; Seg 2, 4, 5 miss), LPT-dispatches the miss seg￾ments across the worker pool (Seg 2 and 4 to Worker 2; Seg 5 to Worker 3), and designates Worker 1 as the combine node, which pulls cache from peers and assembles the running state. to a prefill worker pool—each worker prefills its segmen… view at source ↗
Figure 10
Figure 10. Figure 10: Accuracy–TTFT tradeoff across four models and four datasets. 6 Evaluation 6.1 Setup Hardware. We run all experiments on a node with 8×NVIDIA H20-3e GPUs, each with 141 GB HBM and fully connected by 18-link NVLink, dual-socket Intel Xeon 6759P￾C totaling 120 physical cores, 2 TB DDR5 DRAM, and six Mel￾lanox ConnectX RDMA NICs at 200 Gbps HDR and 400 Gbps NDR. Models. We evaluate four production hybrid-atte… view at source ↗
Figure 10
Figure 10. Figure 10: Accuracy–TTFT Pareto across four models and four datasets. 0.0 0.5 1.0 1.5 TTFT p50 (s) Ring-mini (TP=1) 0.0 0.5 1.0 1.5 Ring-flash (TP=4) 0.0 0.5 1.0 1.5 Qwen3.5-35B (TP=2) 0.0 0.5 1.0 1.5 Qwen3.5-122B (TP=4) 5 10 15 20 Request rate (req/s) 50k 100k 150k Throughput (tokens/s/GPU) 2 4 6 8 10 Request rate (req/s) 5k 10k 15k 2 4 6 8 10 Request rate (req/s) 10k 20k 30k 1 2 3 4 5 Request rate (req/s) 2.5k 5k … view at source ↗
Figure 11
Figure 11. Figure 11: P50 TTFT and per-GPU token throughput at various QPS on the Prod-RAG trace. 1000 2000 3000 4000 (a) Segment length 0.2 0.4 0.6 TTFT p50 (s) 5 10 15 (b) Number of Segments Full Recompute HYPIC [PITH_FULL_IMAGE:figures/full_fig_p011_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Linear-attention composition scaling: accuracy and TTFT against (a) per-segment length at a fixed segment count of 4, and (b) segment count at a fixed per-segment length of 1k tokens. Recompute grows from 0.141 s to 0.624 s as the prompt be￾comes 4× longer, while Hypic grows only from 0.103 s to 0.127 s—a speedup that rises from 1.37× at 4k tokens to 4.91× at 16k tokens. We next vary the number of retriev… view at source ↗
Figure 15
Figure 15. Figure 15: Segment parallelism TTFT breakdown. (a) Scaling with the number of prefill workers. (b) Round-robin vs. LPT load balancing at four workers. in both cases (0.1418 and 0.1164), while larger seam windows increase TTFT without significant accuracy gains. 𝑤=8 is thus a comfortable default. 6.5 Segment parallelism for cache-miss prefill Finally, we stress the cold-miss path—no retrieved segment is cached—to see… view at source ↗
Figure 14
Figure 14. Figure 14: Segment parallelism TTFT breakdown into dis￾patch forward (parallel per-segment prefill), comm (cross￾node KV pull), and combine forward (state composition and seam recompute) as we sweep prefill worker count 𝑛. from 8 to 32 raises TTFT by 76 ms while ROUGE-L varies within 0.15 points. Thus, 𝑤=8 suffices as the default. 6.5 Segment parallelism for cache-miss prefill Segment parallelism TTFT breakdown. Her… view at source ↗
Figure 13
Figure 13. Figure 13: Task accuracy and TTFT against window width 𝑤 per segment boundary. compute per-segment end-states and transitions indepen￾dently, compose the running state via Equation (6), and com￾pare it against Full Recompute [PITH_FULL_IMAGE:figures/full_fig_p012_13.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV

    cs.AI 2026-07 conditional novelty 6.5

    Source-omitted event KV rows can carry compact upstream state (semantic materialization); deliberate answer-free carriers raise recovery from 6% to 51% on Qwen3-8B, while passive natural mentions do not.

Reference graph

Works this paper leans on

64 extracted references · 9 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ram- jee. 2024. Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’24). USENIX Association, 117–134

  2. [2]

    Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhid- ian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al

  3. [3]

    InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics

    LongBench: A Bilingual, Multitask Benchmark for Long Con- text Understanding. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 3119–3137

  4. [4]

    Ziyi Cao, Qingsi Si, Jingbin Zhang, and Bingquan Liu. 2026. Sparse Attention Across Multiple-Context KV Cache. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 30165–30173. doi:10. 1609/aaai.v40i36.40266

  5. [5]

    Aili Chen, Aonian Li, Bangwei Gong, Binyang Jiang, Bo Fei, Bo Yang, Boji Shan, Changqing Yu, Chao Wang, Cheng Zhu, et al. 2025. MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.arXiv preprint arXiv:2506.13585(2025)

  6. [6]

    Chuangtao Chen, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Bing Li, and Ulf Schlichtmann. 2026. KV Packet: Recomputation-Free Context- Independent KV Caching for LLMs. arXiv:2604.13226 [cs.CL]

  7. [7]

    Tri Dao and Albert Gu. 2024. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Du- ality. InProceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235). PMLR, 10041–10071

  8. [8]

    Alexander Richard Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir Radev. 2019. Multi-News: A Large-Scale Multi-Document Summariza- tion Dataset and Abstractive Hierarchical Model. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 1074–1084

  9. [9]

    In Gim, Guojun Chen, Seung-Seob Lee, Nikhil Sarda, Anurag Khandel- wal, and Lin Zhong. 2024. Prompt Cache: Modular Attention Reuse for Low-Latency Inference. InProceedings of Machine Learning and Systems, Vol. 6

  10. [10]

    Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer

  11. [11]

    InProceedings of the 2nd Workshop on New Frontiers in Summarization (EMNLP-IJCNLP 2019 Workshop)

    SAMSum Corpus: A Human-Annotated Dialogue Dataset for Abstractive Summarization. InProceedings of the 2nd Workshop on New Frontiers in Summarization (EMNLP-IJCNLP 2019 Workshop). 70–79

  12. [12]

    Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing a Multi-hop QA Dataset for Compre- hensive Evaluation of Reasoning Steps. InProceedings of the 28th International Conference on Computational Linguistics. 6609–6625

  13. [13]

    Junhao Hu, Wenrui Huang, Weidong Wang, Haoyi Wang, Tiancheng Hu, Qin Zhang, Hao Feng, Xusheng Chen, Yizhou Shan, and Tao Xie

  14. [14]

    InProceedings of the 42nd International Confer- ence on Machine Learning (Proceedings of Machine Learning Research, Vol

    EPIC: Efficient Position-Independent Caching for Serving Large Language Models. InProceedings of the 42nd International Confer- ence on Machine Learning (Proceedings of Machine Learning Research, Vol. 267). PMLR, 24391–24402

  15. [15]

    Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, and Lu Wang. 2021. Efficient Attentions for Long Document Summarization. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics. 1419–1436

  16. [16]

    Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. Swe-bench: Can language models resolve real-world github issues?. InInternational Conference on Learning Representations, Vol. 2024. 54107–54157

  17. [17]

    Weld, and Luke Zettlemoyer

    Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer

  18. [18]

    InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Vol

    TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Vol. 1. 1601–1611

  19. [19]

    Fu, Christo- pher Ré, and Azalia Mirhoseini

    Jordan Juravsky, Bradley Brown, Ryan Ehrlich, Daniel Y. Fu, Christo- pher Ré, and Azalia Mirhoseini. 2024. Hydragen: High-Throughput LLM Inference with Shared Prefixes. InProceedings of the 41st Interna- tional Conference on Machine Learning (ICML ’24)

  20. [20]

    Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020. Transformers are RNNs: Fast Autoregressive Trans- formers with Linear Attention. InProceedings of the 37th International Conference on Machine Learning (ICML)

  21. [21]

    Kimi Team. 2025. Kimi Linear: An Expressive, Efficient Attention Architecture. arXiv:2510.26692 [cs.CL]https://arxiv.org/abs/2510. 26692

  22. [22]

    Gonzalez, Hao Zhang, and Ion Stoica

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica

  23. [23]

    InProceedings of the 29th ACM Symposium on Operating Systems Principles (SOSP ’23)

    Efficient Memory Management for Large Language Model Serv- ing with PagedAttention. InProceedings of the 29th ACM Symposium on Operating Systems Principles (SOSP ’23). ACM, 611–626

  24. [24]

    Shenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li, and Yang You. 2023. Sequence Parallelism: Long Sequence Training from System Perspective. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL). 2391–2404

  25. [25]

    Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, et al . 2024. Jamba: A hybrid transformer-mamba language model.arXiv preprint arXiv:2403.19887(2024)

  26. [26]

    Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. InText Summarization Branches Out. 74–81

  27. [27]

    Yang Liu, Yunfei Gu, Liqiang Zhang, Chentao Wu, Guangtao Xue, Jie Li, Minyi Guo, Junhao Hu, and Jie Meng. 2026. CacheSlide: Unlocking Cross Position-Aware KV Cache Reuse for Accelerating LLM Serv- ing. InProceedings of the 24th USENIX Conference on File and Storage Technologies (FAST ’26). USENIX Association, 83–99

  28. [28]

    Songshuo Lu, Hua Wang, Yutian Rong, Zhi Chen, and Yaohua Tang

  29. [29]

    InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

    TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 6588–6601. doi:10.18653/v1/2025.emnlp-main.334

  30. [30]

    Dongyang Ma, Yan Wang, and Tian Lan. 2025. Block-Attention for Effi- cient Prefilling. InThe Thirteenth International Conference on Learning Representations

  31. [31]

    NVIDIA. 2026. NVIDIA Inference Xfer Library (NIXL). Accessed: 2026-07-08.https://github.com/ai-dynamo/nixl

  32. [32]

    Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024. Splitwise: Efficient Generative LLM Inference Using Phase Splitting. InProceedings of the 51st Annual International Symposium on Computer Architecture (ISCA ’24). IEEE, 118–132

  33. [33]

    Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Minxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu. 2025. Mooncake: Trading More Storage for Less Computation – A KVCache-Centric Architecture for Serving LLM Chatbot. In23rd USENIX Conference on File and Storage Technologies (FAST ’25). USENIX Association. 13 Liu et al

  34. [34]

    Zhen Qin, Weigao Sun, Dong Li, Xuyang Shen, Weixuan Sun, and Yiran Zhong. 2024. Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models. arXiv:2401.04658 [cs.CL]

  35. [35]

    Qwen Team. 2026. Qwen3.5: Towards Native Multimodal Agents. Qwen Technical Blog.https://qwen.ai/blog?id=qwen3.5

  36. [36]

    Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang

  37. [37]

    InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing

    SQuAD: 100,000+ Questions for Machine Comprehension of Text. InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2383–2392

  38. [38]

    Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. arXiv preprint arXiv:1909.08053(2019)

  39. [39]

    Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2024. RoFormer: Enhanced Transformer with Rotary Position Embedding.Neurocomputing568 (2024), 127063

  40. [40]

    Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. 2023. Retentive Net- work: A Successor to Transformer for Large Language Models. arXiv:2307.08621 [cs.CL]

  41. [41]

    Ling Team, Bin Han, Caizhi Tang, Chen Liang, Donghao Zhang, Fan Yuan, Feng Zhu, Jie Gao, Jingyu Hu, Longfei Li, et al . 2025. Every attention matters: An efficient hybrid architecture for long-context reasoning.arXiv preprint arXiv:2510.19338(2025)

  42. [42]

    Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. MuSiQue: Multihop Questions via Single-hop Ques- tion Composition.Transactions of the Association for Computational Linguistics10 (2022), 539–554

  43. [43]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need. InAdvances in Neural Information Processing Systems, Vol. 30

  44. [44]

    Jiahao Wang, Weiyu Xie, Mingxing Zhang, Boxing Zhang, Jianwei Dong, Yuening Zhu, Chen Lin, Jinqi Tang, Yaochen Han, Zhiyuan Ai, Xianglin Chen, Yongwei Wu, and Congfeng Jiang. 2026. From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval- Augmented Generation.Proceedings of the ACM on Management of Data4, 1 (2026). doi:10.1145/3786655

  45. [45]

    Qian Wang, Zahra Yousefijamarani, Morgan Lindsay Heisler, Rongzhi Gu, Xiaolong Bai, Yizhou Shan, Wei Zhang, Lan Wang, Ying Xiong, Yong Zhang, and Zhenan Fan. 2025. MEPIC: Memory Efficient Position Independent Caching for LLM Serving. arXiv:2512.16822 [cs.LG]

  46. [46]

    Shihao Wang, Jiahao Chen, Yanqi Pan, Hao Huang, Yichen Hao, Xi- angyu Zou, Wen Xia, Wentao Zhang, Chongyang Qiu, and Pengfei Wang. 2026. ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation. arXiv:2602.02579 [cs.AI]

  47. [47]

    Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024. Efficient Streaming Language Models with Attention Sinks. InThe Twelfth International Conference on Learning Representa- tions

  48. [48]

    Bin Yang, Qiuyu Leng, Jun Zeng, and Zhenhua Wu. 2025. CacheClip: Accelerating RAG with Effective KV Cache Reuse. arXiv:2510.10129 [cs.LG]

  49. [49]

    Huan Yang, Renji Zhang, Mingzhe Huang, Weijun Wang, Yin Tang, Yuanchun Li, Yunxin Liu, and Deyu Zhang. 2025. KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse. arXiv:2503.16525 [cs.LG]

  50. [50]

    Jingbo Yang, Bairu Hou, Wei Wei, Yujia Bao, and Shiyu Chang. 2025. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse. InAdvances in Neural Information Processing Systems, Vol. 38

  51. [51]

    Songlin Yang, Jan Kautz, and Ali Hatamizadeh. 2025. Gated Delta Networks: Improving Mamba2 with Delta Rule. InThe Thirteenth International Conference on Learning Representations

  52. [52]

    Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. 2024. Gated Linear Attention Transformers with Hardware- Efficient Training. InProceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235). PMLR, 56501–56523

  53. [53]

    Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim

  54. [54]

    InAdvances in Neural Information Processing Systems, Vol

    Parallelizing Linear Transformers with the Delta Rule over Se- quence Length. InAdvances in Neural Information Processing Systems, Vol. 37

  55. [55]

    Songlin Yang and Yu Zhang. 2024. FLA: A Triton-Based Library for Hardware-Efficient Implementations of Linear Attention Mechanism. https://github.com/fla-org/flash-linear-attention

  56. [56]

    Cohen, Ruslan Salakhutdinov, and Christopher D

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2369–2380

  57. [57]

    Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, and Junchen Jiang. 2025. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowl- edge Fusion. InProceedings of the Twentieth European Conference on Computer Systems (EuroSys ’25). ACM. doi:10.1145/3689031.3696098

  58. [58]

    Hancheng Ye, Zhengqi Gao, Mingyuan Ma, Qinsi Wang, Yuzhe Fu, Ming-Yu Chung, Yueqian Lin, Zhijian Liu, Jianyi Zhang, Danyang Zhuo, and Yiran Chen. 2025. KVCOMM: Online Cross-context KV Cache Communication for Efficient LLM-based Multi-agent Systems. InAdvances in Neural Information Processing Systems

  59. [59]

    Lu Ye, Ze Tao, Yong Huang, and Yang Li. 2024. ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 11608–11620

  60. [60]

    Qingfei Zhao, Ruobing Wang, Yukuo Cen, Daren Zha, Shicheng Tan, Yuxiao Dong, and Jie Tang. 2024. LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Ques- tion Answering. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 22600–22632

  61. [61]

    Shiju Zhao, Junhao Hu, Jiaqi Zheng, and Guihai Chen. 2026. You Need an Encoder for Native Position-Independent Caching. arXiv:2602.01519 [cs.CL]https://arxiv.org/abs/2602.01519

  62. [62]

    Gonzalez, Clark Barrett, and Ying Sheng

    Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024. SGLang: Efficient Execution of Structured Language Model Programs. InAdvances in Neural Information Processing Systems

  63. [63]

    Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xu- anzhe Liu, Xin Jin, and Hao Zhang. 2024. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’24). USENIX Association, 193–210

  64. [64]

    Yuechi Zhou, Yi Su, Jianxin Zhang, Juntao Li, Qingrong Xia, Zhefeng Wang, Xinyu Duan, and Baoxing Huai. 2025. A3: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving. arXiv:2511.17560 [cs.CL] 14