Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T11:16:31.904921Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 58 inbound Pith citation observations for arXiv:2407.11550.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T11:16:31.904921Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:57:15.694151Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
69 of 69 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 6e7faeaa-a576-47ec-8d7f-75305605368d · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 144dd309-7e56-4854-920c-fcd77533bc31 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Summedits: measuring llm ability at factual reasoning through the lens of summarization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 52cb7d50-e276-468f-a91b-0c1948b9822d · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Llm-based code generation method for golang compiler testing
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f24c49b7-2128-49ca-80cd-7b77651fd41d · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fb2f066c-506f-4508-a6d4-3377504b0a21 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference The claude 3 model family: Opus, sonnet, haiku, March 2024
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 159bcee7-e37a-4900-8939-363686e5e9d8 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e247733d-70e0-440e-96fa-6098d8520233 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 29cf75f6-909c-41aa-b680-225ae8061826 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference H2o: Heavy-hitter oracle for efficient generative inference of large language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 91cbc77b-ae31-4941-953e-bab9d84f440a · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference PyramidInfer: Pyramid KV cache compression for high-throughput LLM inference
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5db3f52e-6649-45e8-9603-4ddd49a703be · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c2b4aaac-0dc2-4647-a4ea-ffa1b5dbc548 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference SnapKV: LLM knows what you are looking for before generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 485cf98b-5527-4526-9341-82de8ea46220 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Llm kv cache compression made easy
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fd5c8696-81ca-4b25-b752-90921024fa1a · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Catalyst: Optimizing cache management for large in-memory key-value systems
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2dd34f3a-72dc-40a8-a5ba-5093857f671e · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Longformer: The Long-Document Transformer
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eff59542-4dd8-41a3-92d5-4eb3ccd9cee3 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Lm-infinite: Zero-shot extreme length generalization for large language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ba508c6-0cc0-4e36-9495-180237eefdc4 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Efficient Streaming Language Models with Attention Sinks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 40d20c6a-c160-4d4a-8f16-849b124557e6 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Scissorhands: Exploiting the persistence of impor- tance hypothesis for llm kv cache compression at test time
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f0367217-df7e-48ff-b663-070b8cdaadd4 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 56663dad-0c07-4e19-93c0-13bffa7734f0 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 54676c93-14c6-49be-b5a7-f9b12295e6ef · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0e5c3793-50d4-462e-903e-5c8340166dfe · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a804081c-3403-4281-befb-1f02ba20aa63 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Arkvale: Efficient generative llm inference with recallable key-value eviction
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 42a2eb19-cc67-414d-aaf5-48f46890d3aa · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Pqcache: Product quantization-based kvcache for long context llm inference
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9933c708-0a8d-4273-8971-2369de29cb48 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 00ee77e9-1e50-4904-9ec6-9c218efbadae · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Deja vu: Contextual sparsity for efficient llms at inference time
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 29963ae4-5f82-4d45-96a6-b7d89aafd617 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in Neural Information Processing Systems, 35:16344–16359
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3cf3797b-d1cc-4f4c-900e-dbd69cf3ffec · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 33c67fdc-d48f-4b52-b0f5-0a5565ad6ad0 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Efficient memory management for large language model serving with pagedattention
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f684ebbd-af63-4be8-85fc-84561b063492 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference The llama 3 herd of models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f1802368-12fb-469b-ae7e-285d5a4f5f58 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Mistral 7B
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6d0feab6-a4fe-4240-87cd-d1856d040a83 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ebba2fa9-a7f8-4386-8b23-c4a196d7a320 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 881c827d-22a1-41b4-853d-275e2f81b98c · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 888260fa-eedc-4e31-81ed-9504421253f7 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Prompt cache: Modular attention reuse for low-latency inference
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c34b8ce4-76bf-4cc0-b6fe-4eca1d791962 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference SGLang: Efficient Execution of Structured Language Model Programs
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7b29fa69-6d24-4661-b68a-814c11724ce0 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Kvzip: Query-agnostic kv cache compression with context reconstruction
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9ff5bc83-ac8b-42c2-85f1-80dc7ed31a8c · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Expected attention: KV cache compression by estimating attention from future queries distribution
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3066825f-7ce1-44da-9c44-7079e82dfe8f · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Draft-based approximate inference for llms
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 878c43ec-0a9c-4c42-88a5-870334e7df6c · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation baa7e92d-55ac-44be-8258-189d42b518f1 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Kevin Zhou, and Xike Xie
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f75913c8-e146-41e6-ae50-9eda9dffd818 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8002f619-d4c6-4141-b7ba-177b8900a38d · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Not all heads matter: A head-level KV cache compression method with integrated retrieval and reasoning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9afbe2d5-bf62-4dca-8beb-82a8acdd06ec · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 67793422-d0da-48fe-b7ee-cb62b56261b4 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31a68326-fa4a-4a09-a223-3158675baa1e · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference QAQ: Quality Adaptive Quantization for LLM KV Cache
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae37c6f8-7817-4566-a999-0629ccfaa6a5 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 332355d8-0f65-4d5c-aafd-94ea9845a0d8 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Longspec: Long-context lossless speculative decoding with efficient drafting and verification
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dd3f8a06-a3df-43a0-9bdc-cbffc120524d · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 05ba65c9-975f-4cd9-8ab3-0f27e64dfd22 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Needle In A Haystack - pressure testing LLMs
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ff9240fd-688c-422e-b8ea-04691a7129f6 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Zoology: Measuring and improving recall in efficient language models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a0a9affe-2bae-48ed-a176-7ef6e0cec775 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference The narrativeqa reading comprehension challenge
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 045b2a96-9a3a-4eb3-831c-b90c51bcdeb6 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9e10dcc3-d9a1-4269-a067-56d40627b938 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9b2ad6bf-8a6f-45c1-b259-956c9962ce22 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9727104d-89be-42bc-b647-bb5fab16d6a4 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Musique: Multihop questions via single-hop question composition
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 58304fc4-6592-4cf3-916f-6b12d957eeb9 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Efficient Attentions for Long Document Summarization
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0fda1443-aa59-46ae-a7fb-01ad7f27c369 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df2560f3-8962-41cc-9726-3b1cd63b4999 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1e9e909c-87ee-4b7d-96cf-aa963c6151ad · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Weld, and Luke Zettlemoyer
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 64227a81-109f-45a8-a76b-75afadec75d6 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 12e5b9d5-93e2-4051-b2a4-54f83b952702 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Learning question classifiers
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b9fd830c-5b19-4062-8ad8-3b2b5e0e0134 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Longcoder: A long-range pre-trained language model for code completion
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 64b8c0e7-4a09-4653-af0e-2f894b8b12a4 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Repobench: Benchmarking repository-level code auto-completion systems
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5fa0aa7-9481-43c1-ae89-56182d47fa58 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Lost in the Middle: How Language Models Use Long Contexts
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 570d521e-c74f-4c58-b934-a37305a94e27 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b936adde-37fc-45f8-8b68-46ff94f0b7e9 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2c18b871-6722-4361-a7a5-97b9ec96ced0 · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference LongCoder: A Long-Range Pre-trained Language Model for Code Completion
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 15a10e3a-e594-451f-855f-1642d46a6b6a · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a53ed5a-b99e-4e70-babd-5adcda5a929c · outbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference unanswerable
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f17e24f-6c54-4e27-8f67-b00442291c14 · inbound
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f827319a-be95-42c8-9a27-420ba7675998 · inbound
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6acdf6fb-83ba-4780-b293-712fea310160 · inbound
From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 96d90404-72c7-4a92-a26e-07c29d737a99 · inbound
CaliDrop: KV Cache Compression with Calibration Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6079f28-5534-4db1-87df-d7bc15b06115 · inbound
Adaptive KV-Cache Compression without Manually Setting Budget Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa3065d9-cbd7-4529-8aa1-7c2233fcc46f · inbound
PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd52760f-6744-4b3d-8d2c-f8b730e6a790 · inbound
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f4409d38-b709-431f-8cd0-4cf258b23b19 · inbound
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7b153f9-48c0-4353-b88e-c577f109e60f · inbound
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f0141be-a84a-42f3-9332-6ac9eafa64cc · inbound
Predicting Future Utility: Global Combinatorial Optimization for Task-Agnostic KV Cache Eviction Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6235209a-5567-4d26-a903-76b19b66df2c · inbound
Predicting Future Utility: Global Combinatorial Optimization for Task-Agnostic KV Cache Eviction Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd059498-2f91-45cc-bfdf-a30890ff5533 · inbound
CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2cb08b7e-855e-47bc-a2bc-429babc17a2b · inbound
CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb52036e-e225-4445-bf97-74ac0da5d12f · inbound
RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ba045b40-e115-43b9-85aa-7592be0e4e31 · inbound
RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e6346723-9918-45b5-b9d7-69a8d3376839 · inbound
RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66b1bc7e-180a-459c-9973-3782ac42978f · inbound
OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c9e8d310-d388-41fa-8342-51ded957d3d1 · inbound
Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb783ee4-13e0-4dc3-971f-9a5d8292bbe9 · inbound
cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b856209d-dcf8-4559-b657-c00aedf98241 · inbound
CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 406d98a9-c4c0-4f47-88f3-5e6dff90c453 · inbound
AudioKV: KV Cache Eviction in Efficient Large Audio Language Models Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 35fc7be4-c413-4c23-b3ff-fa4fc69a1a89 · inbound
StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation acd2012f-995e-488c-92a1-4ac0447d9b1b · inbound
HieraSparse: Hierarchical Semi-Structured Sparse KV Attention Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e1d3ab66-5641-4e35-8342-dfea167095bb · inbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae260121-b8da-4d6e-8047-4988258c5c16 · inbound
Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5d0e01fd-cc17-4c20-9610-ee008ee8743f · inbound
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f4cb9951-4ede-4629-884f-d0b59830d73b · inbound
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 19d8247b-a1cb-4c75-ab1b-26617f8683ac · inbound
ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 54352432-fa9f-43b7-a2ed-e73125c78673 · inbound
Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3b0a3732-f212-49b2-aadc-6c042827756e · inbound
Compute Where it Counts: Self Optimizing Language Models Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0b95bdb8-d046-4b8a-af92-370d300ee2cf · inbound
Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0e656eb6-73be-486d-8506-36094e2c6f71 · inbound
Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d6937745-42fd-4af9-80f6-16e1f681b9a6 · inbound
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 57c2cafe-3be5-496e-baaf-029560cfef9a · inbound
Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f4035e90-3c4b-48ad-bef5-617fffdac953 · inbound
Head-Aware Key-Value Compression for Efficient Autoregressive Image Generation Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 664cd1d1-fd5e-4a19-8f98-ccfa86236adf · inbound
Runtime-Certified Bounded-Error Quantized Attention Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 618041a0-879a-4b29-910b-00e4247ace06 · inbound
ArborKV: Structure-Aware KV Cache Management for Scaling Tree-based LLM Reasoning Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f6ef1f4f-8a40-4a43-8f63-a5272cdff75d · inbound
Adaptive Mass-Segmented KV Compression for Long-Context Reasoning Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 840d6d94-0f12-4b8e-a316-855aa691d53e · inbound
Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ceb41bf3-04b2-4d55-beb5-38cb1984c52c · inbound
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ab6e242f-0bd2-44c7-8c31-21f407d802c6 · inbound
AURA: Action-Gated Memory for Robot Policies at Constant VRAM Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1878f6e1-7c5b-4f4d-a45f-ff80f8ec31e8 · inbound
TGV-KV: Text-Grounded KV Eviction for Vision-Language Models Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f696f251-de49-4933-aeb5-6970a3fa6c40 · inbound
QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3bcfccae-87c0-489e-bada-5fc3e17d7144 · inbound
Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f020fc52-9aa1-4c1d-82f1-3cf5d6ad275f · inbound
Recency/Frequency Adaptive KV Caching for Large Language Model Serving Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 43d52b19-1db8-4de6-a152-0055d561e04e · inbound
RoPE-Aware Bit Allocation for KV-Cache Quantization Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 34819430-05d8-4ab5-9f0b-71969161c1c3 · inbound
Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cdcca5e3-b201-4475-84b2-8babd76b69a8 · inbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e14af459-92f8-4ac3-9f53-934c91a98c70 · inbound
HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b1c927ed-d880-4e1b-9773-c91e39e52da1 · inbound
Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cac1ad83-7a9e-4a88-a6e2-df8080ee2990 · inbound
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fd1634c4-70f1-4929-ba94-e707d266f298 · inbound
MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5f68943-dc4d-448c-b85b-5c1cd7d85bf5 · inbound
PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73850121-342a-4844-aad9-807694ad86cd · inbound
Error Certificates for KV-Cache Eviction via Randomized Design Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e0d93f9-2ca8-4c7a-8d0c-75e38cda5bc0 · inbound
MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4301edce-d036-4326-a5ac-7512a5da7da9 · inbound
LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f250c808-b7b1-416a-b174-215876bb254b · inbound
GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c1f21c3-1568-423c-95da-973cf561f3e2 · inbound
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.