Pith. sign in

Paper Citation Record · LEDGER

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

As of 6 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 58 inbound Pith citation observations for arXiv:2407.11550.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.11550 v5

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T11:16:31.904921Z

measured 127 of 127 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 58 of 58 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:57:15.694151Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact32
  • verified fuzzy34
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6e7faeaa-a576-47ec-8d7f-75305605368d · outbound

This paper cites A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.120808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:5e2493cc14dbfee66557fc7e47b36bcd50ef4b13f3c67825b41fd9eaf0858a86

Observation 144dd309-7e56-4854-920c-fcd77533bc31 · outbound

This paper cites Summedits: measuring llm ability at factual reasoning through the lens of summarization.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Summedits: measuring llm ability at factual reasoning through the lens of summarization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.133814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:912278e1ffe55c870d61b31edc4e5d5cbc0fa87e1885d3e0c08e1e01e09e78c8

Observation 52cb7d50-e276-468f-a91b-0c1948b9822d · outbound

This paper cites Llm-based code generation method for golang compiler testing.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Llm-based code generation method for golang compiler testing

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.137587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:5286b832e640ac4693eed0929cc810baa60bf8bfb24e98290d86deafe6561905

Observation f24c49b7-2128-49ca-80cd-7b77651fd41d · outbound

This paper cites GPT-4 Technical Report.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference GPT-4 Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:32.087204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:a2510b33791def2bd877713a26c63769601ad1521dd251602eea3c8aaaaf24c7

Observation fb2f066c-506f-4508-a6d4-3377504b0a21 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku, March 2024.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference The claude 3 model family: Opus, sonnet, haiku, March 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.142781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:2b1f5d75c81b70bf30200f226a65a81ad0adb468f5126b7f8ff3531d7e001116

Observation 159bcee7-e37a-4900-8939-363686e5e9d8 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:31.989850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:121147f579cb822e77063e8dfbe245222f753396ee2f9383671bd9f37b788d4b

Observation e247733d-70e0-440e-96fa-6098d8520233 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:32.055445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:06cc228da9a540c3017c4951b227be4fa062a1c52a4c0b0fd0d18efc6459c43d

Observation 29cf75f6-909c-41aa-b680-225ae8061826 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.147044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:5a1380e33c0773997053dbd005dcaa222394bebadb567cf808c39ac13bfc0a9d

Observation 91cbc77b-ae31-4941-953e-bab9d84f440a · outbound

This paper cites PyramidInfer: Pyramid KV cache compression for high-throughput LLM inference.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference PyramidInfer: Pyramid KV cache compression for high-throughput LLM inference

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.151003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:a35eecc5d94f6cd3d13e95c334508fb050c859e12907d823852e4f2e0db6add7

Observation 5db3f52e-6649-45e8-9603-4ddd49a703be · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:32.027501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:8f17d6024c5464ed6a84fb12889321571f6783490c07202c7c14ec752d756945

Observation c2b4aaac-0dc2-4647-a4ea-ffa1b5dbc548 · outbound

This paper cites SnapKV: LLM knows what you are looking for before generation.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference SnapKV: LLM knows what you are looking for before generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.157063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:7fd0a49215a39e12f28b31bdf47ec5607701e40b5983208ec036927fa69835d5

Observation 485cf98b-5527-4526-9341-82de8ea46220 · outbound

This paper cites Llm kv cache compression made easy.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Llm kv cache compression made easy

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.160934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:a7d474071730246934f193eca870c6fed93f3f240957bf191c4a3c8f4f55f4d1

Observation fd5c8696-81ca-4b25-b752-90921024fa1a · outbound

This paper cites Catalyst: Optimizing cache management for large in-memory key-value systems.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Catalyst: Optimizing cache management for large in-memory key-value systems

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.165596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:6c2cedddf25ada249f89f638979b609a0b56ae0108075f66ae65410e91d5ed93

Observation 2dd34f3a-72dc-40a8-a5ba-5093857f671e · outbound

This paper cites Longformer: The Long-Document Transformer.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Longformer: The Long-Document Transformer

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:32.009823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:4a59089794cacaace527bd3d3cd5f4f552a04baf4513a2c2c6b346b98357b6da

Observation eff59542-4dd8-41a3-92d5-4eb3ccd9cee3 · outbound

This paper cites Lm-infinite: Zero-shot extreme length generalization for large language models.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Lm-infinite: Zero-shot extreme length generalization for large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.169379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:15d827f570cbf2852f4d8f4c99b02092d6ae2dd27f2567f5568a58e54b899fe2

Observation 4ba508c6-0cc0-4e36-9495-180237eefdc4 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Efficient Streaming Language Models with Attention Sinks

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:32.032488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:b1cf6c77a1f4f86d8c60de174ffc0f1550a9924f5d43073b56cb3ca1a6735bd1

Observation 40d20c6a-c160-4d4a-8f16-849b124557e6 · outbound

This paper cites Scissorhands: Exploiting the persistence of impor- tance hypothesis for llm kv cache compression at test time.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Scissorhands: Exploiting the persistence of impor- tance hypothesis for llm kv cache compression at test time

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.173735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:09999bf51857720b866c9517caa18348608650c0eac77d34d5c00a3578ed48a4

Observation f0367217-df7e-48ff-b663-070b8cdaadd4 · outbound

This paper cites On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.074216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:e5995c23989ad0d3c199143dbe6c54c1ee58db27a7e64f4e67b0fe82b26e972a

Observation 56663dad-0c07-4e19-93c0-13bffa7734f0 · outbound

This paper cites Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.177378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:20ae4a3f6aa3423a30f64c455d8319d2ced3da12fa22aac2532709b0dd31d203

Observation 54676c93-14c6-49be-b5a7-f9b12295e6ef · outbound

This paper cites MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T11:16:32.106754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:f1a9e872d0115f76f214bf5b9ffbede69e72fdfcb18789f9bea025aa362275ca

Observation 0e5c3793-50d4-462e-903e-5c8340166dfe · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:95b8e0b85acf993205df680f81c67afa6506b5e5a07a0eec803e31a0233337df

Observation a804081c-3403-4281-befb-1f02ba20aa63 · outbound

This paper cites Arkvale: Efficient generative llm inference with recallable key-value eviction.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Arkvale: Efficient generative llm inference with recallable key-value eviction

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.180782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:3fc39868460753661fefd5f7fd42a8cf61dcaf1202716e61aa24d2884b7b7121

Observation 42a2eb19-cc67-414d-aaf5-48f46890d3aa · outbound

This paper cites Pqcache: Product quantization-based kvcache for long context llm inference.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Pqcache: Product quantization-based kvcache for long context llm inference

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.184132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:0c589dc75f5a10e954ef833bd3ca094818e39faf03588c13d7815003df282fc1

Observation 9933c708-0a8d-4273-8971-2369de29cb48 · outbound

This paper cites Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.022030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:e2a23d22e5d5cbb57232f1240d1273cc548e4815815c69e44423812dfeef5952

Observation 00ee77e9-1e50-4904-9ec6-9c218efbadae · outbound

This paper cites Deja vu: Contextual sparsity for efficient llms at inference time.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Deja vu: Contextual sparsity for efficient llms at inference time

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.187182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:0b20906e7c8bcff467e0fbe34a0643c11500a6bf86870c864f91333326cbb64f

Observation 29963ae4-5f82-4d45-96a6-b7d89aafd617 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in Neural Information Processing Systems, 35:16344–16359.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in Neural Information Processing Systems, 35:16344–16359

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.190958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:dc5dae165f44d9371215cfe9edfb3900d56934ee3282b7481c6e5eaaf157e382

Observation 3cf3797b-d1cc-4f4c-900e-dbd69cf3ffec · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:32.049856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:b40291a9a17f8b27bed74613b48f153820688b3ce0458b2706c020b9edd9159c

Observation 33c67fdc-d48f-4b52-b0f5-0a5565ad6ad0 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Efficient memory management for large language model serving with pagedattention

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.194755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:a6b3b842160b957e0e196021178dd1520de0db1d27fa43862db650f0a868271b

Observation f684ebbd-af63-4be8-85fc-84561b063492 · outbound

This paper cites The llama 3 herd of models.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference The llama 3 herd of models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.197601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:7c57d628b6fb080b4b403c668b8d996ba78799b4cbee8ca9fe2b431a6c36510e

Observation f1802368-12fb-469b-ae7e-285d5a4f5f58 · outbound

This paper cites Mistral 7B.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Mistral 7B

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:32.082945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:5f44a57ade44a9c53de1ad857c331f13a23de8c4182df2d774e6fce8d6c05df7

Observation 6d0feab6-a4fe-4240-87cd-d1856d040a83 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.200995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:d688e5c86059ab265e9aec35ebcaff0feed766d0a40e11805cd320a13f69d750

Observation ebba2fa9-a7f8-4386-8b23-c4a196d7a320 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:32.092177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:6f4613190ce470e6874de53d525f5d73c85d7fef25867d32ade738bea936967b

Observation 881c827d-22a1-41b4-853d-275e2f81b98c · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:32.101465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:e05875da4d066e18ed57372088084434405ec3125e19c374b24da65c4c0f893a

Observation 888260fa-eedc-4e31-81ed-9504421253f7 · outbound

This paper cites Prompt cache: Modular attention reuse for low-latency inference.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Prompt cache: Modular attention reuse for low-latency inference

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.204467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:d901c1f867c865847726da7897c1c6a182b960b691c361d1071e8f8a896e409d

Observation c34b8ce4-76bf-4cc0-b6fe-4eca1d791962 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference SGLang: Efficient Execution of Structured Language Model Programs

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:32.126000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:4830b6af2025ac4edd0818eee6c6489eeeb9d9b4f0f08e6bc1c369e0523b972f

Observation 7b29fa69-6d24-4661-b68a-814c11724ce0 · outbound

This paper cites Kvzip: Query-agnostic kv cache compression with context reconstruction.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Kvzip: Query-agnostic kv cache compression with context reconstruction

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:31.953933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:b0525f99bfe78d0709dd18ca38b0533eb5f6b34dc3a0b9b47f52ae8c99ef56c4

Observation 9ff5bc83-ac8b-42c2-85f1-80dc7ed31a8c · outbound

This paper cites Expected attention: KV cache compression by estimating attention from future queries distribution.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Expected attention: KV cache compression by estimating attention from future queries distribution

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:31.960407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:7dfb9705f68574f4135caf73138a8726964d19726572e405a6eb464fa3ac24f5

Observation 3066825f-7ce1-44da-9c44-7079e82dfe8f · outbound

This paper cites Draft-based approximate inference for llms.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Draft-based approximate inference for llms

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:31.965616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:b3ee91952012fb8291df71e546521bcb37e3215f7ad6e18e3f812b9801926f57

Observation 878c43ec-0a9c-4c42-88a5-870334e7df6c · outbound

This paper cites CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:52.226407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:daf136ef785c2ce86b566651154766a802e48313a7a5d32c8ce60cfda7e042a5

Observation baa7e92d-55ac-44be-8258-189d42b518f1 · outbound

This paper cites Kevin Zhou, and Xike Xie.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Kevin Zhou, and Xike Xie

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.208025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:9a8409b3d1e2d985980bb160154f5d6f1bd7fc08055312184a3ca3247ca30e26

Observation f75913c8-e146-41e6-ae50-9eda9dffd818 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:49:16.947404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:ecf9addfe28d9605d9fe6b16fa7b30e75320c60c82d2a71fd89db5c678d2e68f

Observation 8002f619-d4c6-4141-b7ba-177b8900a38d · outbound

This paper cites Not all heads matter: A head-level KV cache compression method with integrated retrieval and reasoning.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Not all heads matter: A head-level KV cache compression method with integrated retrieval and reasoning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:31.985382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:afec1c474dc20d55fef5c858d7d84a2706d79c5494f21ee43fd1a0652a5871f8

Observation 9afbe2d5-bf62-4dca-8beb-82a8acdd06ec · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.212024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:65dd759117e9bd5cc6bbab3b4e39be837c9b450bb25a6b85b79ac8dce002ddf8

Observation 67793422-d0da-48fe-b7ee-cb62b56261b4 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:31.994648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:6f083d97667678b254a0844d28a609ea35a9b72c5378820406f7d66950a33cdf

Observation 31a68326-fa4a-4a09-a223-3158675baa1e · outbound

This paper cites QAQ: Quality Adaptive Quantization for LLM KV Cache.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference QAQ: Quality Adaptive Quantization for LLM KV Cache

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T11:16:32.000177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:2622386d8b17f97a499f97264b0fee7686e6a8dfaefce7005d6ed275a5af4903

Observation ae37c6f8-7817-4566-a999-0629ccfaa6a5 · outbound

This paper cites TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.005322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:536504b0236dc7f99266477edae302c13266d3920a7ee495ef6211f69576bc50

Observation 332355d8-0f65-4d5c-aafd-94ea9845a0d8 · outbound

This paper cites Longspec: Long-context lossless speculative decoding with efficient drafting and verification.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Longspec: Long-context lossless speculative decoding with efficient drafting and verification

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.215607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:732018b370a300d1e5d546dfff46158a98f34f65b6c86d862afab8b31747d681

Observation dd3f8a06-a3df-43a0-9bdc-cbffc120524d · outbound

This paper cites SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.015576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:fa52ad29136b8611f9928dd23ad57cb6a22fb30b54b7331c64b2202e63ee49a2

Observation 05ba65c9-975f-4cd9-8ab3-0f27e64dfd22 · outbound

This paper cites Needle In A Haystack - pressure testing LLMs.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Needle In A Haystack - pressure testing LLMs

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.218999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:97c539ab10b72a29e7148584816ba4794ca464bb01e7fa14300773f29d2702e5

Observation ff9240fd-688c-422e-b8ea-04691a7129f6 · outbound

This paper cites Zoology: Measuring and improving recall in efficient language models.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Zoology: Measuring and improving recall in efficient language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.222444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:b9deb367c9be95f2e6b700599b333cb9aaaeb1dd7a3cab2a754724ea44f9e094

Observation a0a9affe-2bae-48ed-a176-7ef6e0cec775 · outbound

This paper cites The narrativeqa reading comprehension challenge.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference The narrativeqa reading comprehension challenge

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.225564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:b7745e6626d6d61080269dd616200f81600dff63e1ffe38e5ca39e02cb6da89c

Observation 045b2a96-9a3a-4eb3-831c-b90c51bcdeb6 · outbound

This paper cites A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.038376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:39295b047a5f2c9fc313e662605a2bf96a6544a09c7d6c0aed17c7a5b6ff50a1

Observation 9e10dcc3-d9a1-4269-a067-56d40627b938 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:32.043803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:ad681cfc5badff6b1cbf7c54b1ae33bd53be4253014d36a67da1cc4db7e1df5f

Observation 9b2ad6bf-8a6f-45c1-b259-956c9962ce22 · outbound

This paper cites Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.228381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:318a236d0b74ecf7bda85eec4d271bce06d6d44446905a19ac0664da65d4a115

Observation 9727104d-89be-42bc-b647-bb5fab16d6a4 · outbound

This paper cites Musique: Multihop questions via single-hop question composition.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Musique: Multihop questions via single-hop question composition

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.231534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:a87298629aeca83a58b1718d04eb6f7949602c04c98b1613eae46d887b222b21

Observation 58304fc4-6592-4cf3-916f-6b12d957eeb9 · outbound

This paper cites Efficient Attentions for Long Document Summarization.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Efficient Attentions for Long Document Summarization

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.060641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:a619bd91bad155b5db01c466fe3cb9e388a27316e6d4dbaaf1f0bde21175d609

Observation 0fda1443-aa59-46ae-a7fb-01ad7f27c369 · outbound

This paper cites QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.065330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:3a6847e66f2df5acf217e316f5cc69146b7f2a9c0933712503dc6f0cfab39446

Observation df2560f3-8962-41cc-9726-3b1cd63b4999 · outbound

This paper cites Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:32.069748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:45066c0570993c3cedd5053b511bd4099f43c660e415ced2ec6ee04bea618550

Observation 1e9e909c-87ee-4b7d-96cf-aa963c6151ad · outbound

This paper cites Weld, and Luke Zettlemoyer.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Weld, and Luke Zettlemoyer

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.234499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:492094a62e73c87858e1a209208ab25187c232327ae2ab42219467af55de2b95

Observation 64227a81-109f-45a8-a76b-75afadec75d6 · outbound

This paper cites SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.079227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:82b11db56a5d28f75bd3713630492f0e1ffe43f5bed2c4fd794644b316e79191

Observation 12e5b9d5-93e2-4051-b2a4-54f83b952702 · outbound

This paper cites Learning question classifiers.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Learning question classifiers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.237749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:bc434f0c542ac8ff46304f37befea7f8e61bc8896c614ad899882004c9c976bf

Observation b9fd830c-5b19-4062-8ad8-3b2b5e0e0134 · outbound

This paper cites Longcoder: A long-range pre-trained language model for code completion.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Longcoder: A long-range pre-trained language model for code completion

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.240745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:206d7fe426cb8633d6bfbfc202c3712dfc85693bd3a1b3b48b171b50c170289a

Observation 64b8c0e7-4a09-4653-af0e-2f894b8b12a4 · outbound

This paper cites Repobench: Benchmarking repository-level code auto-completion systems.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Repobench: Benchmarking repository-level code auto-completion systems

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.243644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:f5be7c026a3d1928f42574ee9fde429101f3a692321ec5baf8bc668c0724bdc4

Observation a5fa0aa7-9481-43c1-ae89-56182d47fa58 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Lost in the Middle: How Language Models Use Long Contexts

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:16:32.097026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:97215e31c9e73a446c18084f612f1a3f6fe69178ebaf33d8cc25586ce59fc915

Observation 570d521e-c74f-4c58-b934-a37305a94e27 · outbound

This paper cites Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.246861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:2187ad41286a767291bda1a9ada8dc1651d67ceaf9fdd2da130274ad339fec89

Observation b936adde-37fc-45f8-8b68-46ff94f0b7e9 · outbound

This paper cites Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.249528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:d82931b804bd611759ebd5ca1ccf5a8663c2e277ae82c0fb3f94fa2c7c3835e5

Observation 2c18b871-6722-4361-a7a5-97b9ec96ced0 · outbound

This paper cites LongCoder: A Long-Range Pre-trained Language Model for Code Completion.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference LongCoder: A Long-Range Pre-trained Language Model for Code Completion

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.111686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:f54289b2181d90a5397e6ac218ea33e2f31157fbb968dafabb046efdf1417961

Observation 15a10e3a-e594-451f-855f-1642d46a6b6a · outbound

This paper cites RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems

Reference 68

Resolution
malformed identifier
local_arxiv, observed 2026-05-17T11:16:32.115902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:c305a3e1b1e437372ddd7e4b6e64baa27258475d7d8582205cb1239cc9e3d2a3

Observation 8a53ed5a-b99e-4e70-babd-5adcda5a929c · outbound

This paper cites unanswerable.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference unanswerable

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:16:32.129916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:26ab5727e9a30c36f40563848f5e995872fd4c4bf90e7fc5950bca09bd5cfd4a

Pith citing papers

Observation 0f17e24f-6c54-4e27-8f67-b00442291c14 · inbound

FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration cites this paper.

FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:02:30.377552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T03:59:58.634512Z digest=sha256:90f6c406b838eb950333388e0ba6df3c585edc12cf1459963026fbbafa217570

Observation f827319a-be95-42c8-9a27-420ba7675998 · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:17:31.067489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:3ab8aa8b45a9d677f903577dbb3acebeb5c23197fee46fd47acf0c7b5b9c90c3

Observation 6acdf6fb-83ba-4780-b293-712fea310160 · inbound

From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs cites this paper.

From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:05:09.588491Z digest=sha256:9ccd717d233c1ac80d9577409e7ba4cbd85c07f84f8a8547a141d4b6695f039d

Observation 96d90404-72c7-4a92-a26e-07c29d737a99 · inbound

CaliDrop: KV Cache Compression with Calibration cites this paper.

CaliDrop: KV Cache Compression with Calibration Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.694151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.694151Z digest=sha256:c69be830adc8c55c75f541a97eba9e9f6f16bc600e4e84117a68b724dea49674

Observation f6079f28-5534-4db1-87df-d7bc15b06115 · inbound

Adaptive KV-Cache Compression without Manually Setting Budget cites this paper.

Adaptive KV-Cache Compression without Manually Setting Budget Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T11:10:56.741169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:10:56.741169Z digest=sha256:72db7309de92ea12b1b57f34e3e4d32bb4821a4bae91df57f59722e4babee4a8

Observation aa3065d9-cbd7-4529-8aa1-7c2233fcc46f · inbound

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference cites this paper.

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:13.056228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:13.056228Z digest=sha256:a46cd6bc5b01ce759f072bff3d57d35e8df30b88d8044565425708e94c763b1c

Observation bd52760f-6744-4b3d-8d2c-f8b730e6a790 · inbound

EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments cites this paper.

EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:54:22.753517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T21:54:05.442769Z digest=sha256:5186b8712d1fe9d59d2a49ccefa800935363a1c3d790fc929094a4557e0c2575

Observation f4409d38-b709-431f-8cd0-4cf258b23b19 · inbound

Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection cites this paper.

Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T05:09:28.823705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:09:28.823705Z digest=sha256:d8b275ffcc830372ff6ae5cb72bbe28a88439a3c7715860b2e06324cbe75f419

Observation e7b153f9-48c0-4353-b88e-c577f109e60f · inbound

ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs cites this paper.

ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:37:50.106345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:37:50.106345Z digest=sha256:fba84d3eb892d3fbd0fc8ea6baa83983cb9ab9127a031d4e1a953ee47b1f4bd6

Observation 3f0141be-a84a-42f3-9332-6ac9eafa64cc · inbound

Predicting Future Utility: Global Combinatorial Optimization for Task-Agnostic KV Cache Eviction cites this paper.

Predicting Future Utility: Global Combinatorial Optimization for Task-Agnostic KV Cache Eviction Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T03:17:41.901032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:17:41.901032Z digest=sha256:c80d1456b0ff8528cd2139b524a0ec5df2f84370c0cb932b10b8364d006f949a

Observation 6235209a-5567-4d26-a903-76b19b66df2c · inbound

Predicting Future Utility: Global Combinatorial Optimization for Task-Agnostic KV Cache Eviction cites this paper.

Predicting Future Utility: Global Combinatorial Optimization for Task-Agnostic KV Cache Eviction Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T03:17:41.954664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:17:41.954664Z digest=sha256:fa721f6c9fab4860985d1322a27a55cabafd77b1391499fa2a06b2281beb7380

Observation bd059498-2f91-45cc-bfdf-a30890ff5533 · inbound

CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation cites this paper.

CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:34:11.386035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T13:32:15.598047Z digest=sha256:40f8b07496a5fbb3a8ea74ae4ad0de4aac8bc5e24ca332aad6479c86dc9abb77

Observation 2cb08b7e-855e-47bc-a2bc-429babc17a2b · inbound

CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation cites this paper.

CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:16:19.869117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:16:19.869117Z digest=sha256:80f00fc9eb60bc545da1fc06c36066b8af57c745ab6193a44568df4d8dd40882

Observation bb52036e-e225-4445-bf97-74ac0da5d12f · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:59:33.902420Z digest=sha256:96622812aeac3c2a4c22a1edf9a20c0b5c24e71eccb11074ff42d0eed9736aeb

Observation ba045b40-e115-43b9-85aa-7592be0e4e31 · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T12:50:09.472921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T12:45:27.150368Z digest=sha256:e0aa2438ff61427627622848bd85ab495c32d588b8f791bb784287bd53009063

Observation e6346723-9918-45b5-b9d7-69a8d3376839 · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T22:05:42.996217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:05:42.996217Z digest=sha256:d0934978f9222250890156ee9a7400127a8eeb4a2dcd654e5e28c0cd9e49bda5

Observation 66b1bc7e-180a-459c-9973-3782ac42978f · inbound

OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer cites this paper.

OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:39:55.826057Z digest=sha256:2109bc62281e4c9f044612d9a294d13d08ddfd9581c3cacf40a2baac0bfc6186

Observation c9e8d310-d388-41fa-8342-51ded957d3d1 · inbound

Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs cites this paper.

Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:31:17.420639Z digest=sha256:7c2a41988b961aad0603e08cfe82bb530cdc5e0129fbd1a367683f96c3f65b84

Observation cb783ee4-13e0-4dc3-971f-9a5d8292bbe9 · inbound

cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations cites this paper.

cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T09:09:32.932051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:09:32.932051Z digest=sha256:0c3790098cd325a668da50fa69f1fe61aba6e0bead4f1709380f8ad2df3b9bb1

Observation b856209d-dcf8-4559-b657-c00aedf98241 · inbound

CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference cites this paper.

CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:45:26.555395Z digest=sha256:0153fb15c5a6bec093276e1ab6e4d4d34083a0f5709025406f940fa309b46850

Observation 406d98a9-c4c0-4f47-88f3-5e6dff90c453 · inbound

AudioKV: KV Cache Eviction in Efficient Large Audio Language Models cites this paper.

AudioKV: KV Cache Eviction in Efficient Large Audio Language Models Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:28:49.768210Z digest=sha256:42bf6cde3d7afcefe25abce9f3ad3f3033d07ef008d525014a7846ac018fafae

Observation 35fc7be4-c413-4c23-b3ff-fa4fc69a1a89 · inbound

StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression cites this paper.

StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T11:16:13.269662Z digest=sha256:1d6969710579f7daa820f5e723b4cb1a378e3b3a8cf50412dde1acf9325f116a

Observation acd2012f-995e-488c-92a1-4ac0447d9b1b · inbound

HieraSparse: Hierarchical Semi-Structured Sparse KV Attention cites this paper.

HieraSparse: Hierarchical Semi-Structured Sparse KV Attention Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T07:15:19.184970Z digest=sha256:6cb5c5d4edcc654c6fa00a4a8e05e7dfbe4a0650ad03a7aa1f1c5c393c118a65

Observation e1d3ab66-5641-4e35-8342-dfea167095bb · inbound

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing cites this paper.

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:48:18.330339Z digest=sha256:8206733a34bb85338416803dcb38ecdd1add183755bc62bcd178202217697125

Observation ae260121-b8da-4d6e-8047-4988258c5c16 · inbound

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache cites this paper.

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:15:35.871863Z digest=sha256:81812c252d869a557d9cbaab266114e92b83d1bb7a2e440805948d212ef819a9

Observation 5d0e01fd-cc17-4c20-9610-ee008ee8743f · inbound

Reformulating KV Cache Eviction Problem for Long-Context LLM Inference cites this paper.

Reformulating KV Cache Eviction Problem for Long-Context LLM Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T02:37:52.545943Z digest=sha256:a38c3c6077d0e4afed4f305201ad4c1ef46cc8d54ec41af91f68616830d5c78f

Observation f4cb9951-4ede-4629-884f-d0b59830d73b · inbound

RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache cites this paper.

RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:16:15.869198Z digest=sha256:94bd6cf7071c683f76a5eba6b3f939143d8283b2cd69e377b4730c7c06d76007

Observation 19d8247b-a1cb-4c75-ab1b-26617f8683ac · inbound

ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing cites this paper.

ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T02:52:33.076123Z digest=sha256:314d4fcb87213fa7943d2fefc9e29821602fd3e76be2b3ff1bb7e8f405e796c6

Observation 54352432-fa9f-43b7-a2ed-e73125c78673 · inbound

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction cites this paper.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:9425484fd21851df58ba75be935158b15d93420bf8546fc70731073369b09067

Observation 3b0a3732-f212-49b2-aadc-6c042827756e · inbound

Compute Where it Counts: Self Optimizing Language Models cites this paper.

Compute Where it Counts: Self Optimizing Language Models Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:57:58.649358Z digest=sha256:d5d8f9e11a02eef6ba66c36c5624d72cd6d46262cbd339074a654d20e167fa2a

Observation 0b95bdb8-d046-4b8a-af92-370d300ee2cf · inbound

Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation cites this paper.

Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T19:49:04.707675Z digest=sha256:e514db6e250d91cec770b434265fc961c507adf8529d5da29726a77f55feb378

Observation 0e656eb6-73be-486d-8506-36094e2c6f71 · inbound

Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility cites this paper.

Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T05:35:09.705532Z digest=sha256:4fe03b2784689a35b4517261b9a97c43704130c5f74110454d6acd46474e1538

Observation d6937745-42fd-4af9-80f6-16e1f681b9a6 · inbound

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference cites this paper.

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T22:12:51.079539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T22:09:15.782104Z digest=sha256:b3a416a530d9a96cc65d77561cb714fc1490493498c4f2b10ba5dae75037b352

Observation 57c2cafe-3be5-496e-baaf-029560cfef9a · inbound

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction cites this paper.

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:13:15.950561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T12:13:14.295659Z digest=sha256:1339aa8e1ce96e66448fd58786585dcc27daec1403742db6d888c4e343776e86

Observation f4035e90-3c4b-48ad-bef5-617fffdac953 · inbound

Head-Aware Key-Value Compression for Efficient Autoregressive Image Generation cites this paper.

Head-Aware Key-Value Compression for Efficient Autoregressive Image Generation Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:09:41.660092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T06:05:23.148656Z digest=sha256:f5c7e2f4e629e6725f6878139de162db345a4360483afb7b1af3d3bc93dfe0e8

Observation 664cd1d1-fd5e-4a19-8f98-ccfa86236adf · inbound

Runtime-Certified Bounded-Error Quantized Attention cites this paper.

Runtime-Certified Bounded-Error Quantized Attention Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:49:41.134082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T05:45:42.295527Z digest=sha256:4d5ca0e423a07d534b70aa3a8e30fa7ff16eed2c52b4ad7d86744c3095b6ecc9

Observation 618041a0-879a-4b29-910b-00e4247ace06 · inbound

ArborKV: Structure-Aware KV Cache Management for Scaling Tree-based LLM Reasoning cites this paper.

ArborKV: Structure-Aware KV Cache Management for Scaling Tree-based LLM Reasoning Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:51:08.742724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T05:48:30.118982Z digest=sha256:1f836493e48afccaecf30a3ea15bc5f0609c35ee225c76890be7c6df139a2d09

Observation f6ef1f4f-8a40-4a43-8f63-a5272cdff75d · inbound

Adaptive Mass-Segmented KV Compression for Long-Context Reasoning cites this paper.

Adaptive Mass-Segmented KV Compression for Long-Context Reasoning Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:25:22.824928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T05:22:47.198185Z digest=sha256:2052780e903cda9016b438c3075f0275376170668a8172b9b161e0b07c6744ec

Observation 840d6d94-0f12-4b8e-a316-855aa691d53e · inbound

Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression cites this paper.

Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-30T00:04:07.005277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-29T23:46:44.498959Z digest=sha256:9f0d864a700b3bc3722b7eab621f4ba6ddc666eda47667e5dc4729d0d6462da7

Observation ceb41bf3-04b2-4d55-beb5-38cb1984c52c · inbound

MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference cites this paper.

MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:56:15.429922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T16:07:22.196982Z digest=sha256:4a770e1b93013777e0f1ee7f98525a784306502c5e7eefa7355689190fa51368

Observation ab6e242f-0bd2-44c7-8c31-21f407d802c6 · inbound

AURA: Action-Gated Memory for Robot Policies at Constant VRAM cites this paper.

AURA: Action-Gated Memory for Robot Policies at Constant VRAM Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:16:24.341478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T14:30:49.006075Z digest=sha256:9cea8179f88d9809c1b2cc755d31b6934c8c632caa3f0fa0aaa53360dc56f94e

Observation 1878f6e1-7c5b-4f4d-a45f-ff80f8ec31e8 · inbound

TGV-KV: Text-Grounded KV Eviction for Vision-Language Models cites this paper.

TGV-KV: Text-Grounded KV Eviction for Vision-Language Models Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:16:26.712945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T11:04:54.171291Z digest=sha256:ddc5517321beefe71e611904ea5ca8087b5dafa0fff00d38d6b499b6597a86ee

Observation f696f251-de49-4933-aeb5-6970a3fa6c40 · inbound

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving cites this paper.

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:46:57.470290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T01:47:43.240850Z digest=sha256:be0ce6ba9d2ff9704a1cd80909b41ffa410d0d1b37388819ffc4fbec9991edc0

Observation 3bcfccae-87c0-489e-bada-5fc3e17d7144 · inbound

Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving cites this paper.

Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.142261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T02:25:46.234576Z digest=sha256:55d80299a26cc60940fa020439392e58101620cf1b638bdc77936226cb9735c0

Observation f020fc52-9aa1-4c1d-82f1-3cf5d6ad275f · inbound

Recency/Frequency Adaptive KV Caching for Large Language Model Serving cites this paper.

Recency/Frequency Adaptive KV Caching for Large Language Model Serving Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T07:29:39.057707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T13:24:44.218871Z digest=sha256:ac18017fbf705b6b90fdcf4c6af1af7cbd014f94950e496427ca08531983caba

Observation 43d52b19-1db8-4de6-a152-0055d561e04e · inbound

RoPE-Aware Bit Allocation for KV-Cache Quantization cites this paper.

RoPE-Aware Bit Allocation for KV-Cache Quantization Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:59:57.339429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T01:05:13.790381Z digest=sha256:73e4cff83340a77517fb2995ba29bc0db6bc36aed56a6b08c2665d9c9d604e64

Observation 34819430-05d8-4ab5-9f0b-71969161c1c3 · inbound

Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling cites this paper.

Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:09:57.153630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T00:53:43.629684Z digest=sha256:c18899bd0223f7d6b1fe897fa6be3f2013fede2e6f9abdc67225ed9e90f0bdad

Observation cdcca5e3-b201-4475-84b2-8babd76b69a8 · inbound

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference cites this paper.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:49:58.039510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:2d2992bac15e31967993c0d709dc7131e56e7e07ecfe6b670474527a3d83c750

Observation e14af459-92f8-4ac3-9f53-934c91a98c70 · inbound

HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression cites this paper.

HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:44:37.454177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T10:27:30.344430Z digest=sha256:0923a40d6253f5285e590041f730bda0cc8c4e18f7c535f18b6dd4c7f4a83e0a

Observation b1c927ed-d880-4e1b-9773-c91e39e52da1 · inbound

Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM cites this paper.

Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:24:21.317305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-30T07:19:48.530272Z digest=sha256:d09d55f856a190f1268cf846a0c26f9e8ed98c1d60af906c150f01ca1aa3fefe

Observation cac1ad83-7a9e-4a88-a6e2-df8080ee2990 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.281921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:eb64ed1e5d96a9cefda57b63bc8e5174c4db61ea7389a3d5de48a8c8b280b7ae

Observation fd1634c4-70f1-4929-ba94-e707d266f298 · inbound

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference cites this paper.

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T10:41:08.967261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:41:08.967261Z digest=sha256:92d570f4f7f463bdcb564d9d1eb47fe8c926d886d180e187fd57993b0750912c

Observation e5f68943-dc4d-448c-b85b-5c1cd7d85bf5 · inbound

PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents cites this paper.

PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-01T14:12:01.001000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:12:01.001000Z digest=sha256:25c1d2a7a1796fe90b60c1391e3ac754bf98af4b4afc77af240cc09f57925df3

Observation 73850121-342a-4844-aad9-807694ad86cd · inbound

Error Certificates for KV-Cache Eviction via Randomized Design cites this paper.

Error Certificates for KV-Cache Eviction via Randomized Design Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T07:23:37.013625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:23:37.013625Z digest=sha256:270c51c7063ca0e5ab1d826101f706a1bb33ed5f0cfaa242424e1cd0835e9741

Observation 0e0d93f9-2ca8-4c7a-8d0c-75e38cda5bc0 · inbound

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models cites this paper.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:30.056509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:30.056509Z digest=sha256:87c5c5cba084ed527779bf12804600bd878b6dc60e1d9e01de43f8f1edee77ec

Observation 4301edce-d036-4326-a5ac-7512a5da7da9 · inbound

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding cites this paper.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.643900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.643900Z digest=sha256:6fa5640f2c915929ce2924171f2dd6f8616f0d18125bc608a7e233639adcfae7

Observation f250c808-b7b1-416a-b174-215876bb254b · inbound

GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference cites this paper.

GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T09:52:42.397695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:52:42.397695Z digest=sha256:9dcbfecd0976b4751ad526a0af96ea5041979daa1dffc3af4187a68811635320

Observation 8c1f21c3-1568-423c-95da-973cf561f3e2 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:57.862186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:43:57.862186Z digest=sha256:6565cb5955b2038cfd992952373f5c3daaaad32a65a4d492a1ec14702fb8df25