Pith. sign in

Paper Citation Record · LEDGER

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction

As of 3 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2605.09649.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.09649 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T05:02:25.513351Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T02:23:36.089893Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact28
  • verified fuzzy2
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 013b6995-1260-4408-ae8b-3a7b2f547ad1 · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:38:00.425257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:9fcb1cb39af4ed81679195d13f6049f5f3c08cb46995bbd6482b7e7667b369a8

Observation f56a4737-ccb5-4148-8a48-05c14368f787 · outbound

This paper cites arXiv preprint arXiv:2512.03324 , year=.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction arXiv preprint arXiv:2512.03324 , year=

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.502225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:031659afa99c34f334c0e9faced19f6737d9bd6879b9a3ac2b262d5640a5471c

Observation 416e77be-388e-4f64-a49b-e09f86040c65 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:58:29.304691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:92707d587bcb5265494ef5b3054b166ff5cd1e93fff0816a7d0ee2d6b8f4651f

Observation 38f4cab9-de76-4f6c-85c8-ad27a820454d · outbound

This paper cites R-KV: Redundancy-aware KV cache compression for reasoning models.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction R-KV: Redundancy-aware KV cache compression for reasoning models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.451571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:8a18808e14be523b8ee20dbbea2e231893a72341d1939d04fbd8466038e515c2

Observation bd2546f2-e360-45c6-b995-80c4c599d8bf · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T05:41:26.485359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:b0c313f277ee65e7a0b6b151e1eff65f8a18387cc2373a44ab3d9f569cf7e725

Observation f6d53288-ed26-40d2-aec1-2619bf91110f · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:41:26.412685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:690fcb82b004515cd828df252c237ca3c4889056aaea33af8d58db620a2722c8

Observation 5ccde352-3862-424b-949e-1f850ec11e1f · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:41:26.675257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:98b8c187b6d56d45ce4ea1f656e67bdc4d681f3cdbb68f8e64d6e6147e304894

Observation 7bd1e6bb-a2e3-4e1f-aade-e22b7ce1ed6c · outbound

This paper cites How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.694598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:22a5fd6bf5059b412855071b312b93fbe39c6b4f443c6b55ed9bc92591440a4a

Observation 54352432-fa9f-43b7-a2ed-e73125c78673 · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.250514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:271c6b00cd4d4fbf269f74f03213f51cd8d7156ee668bb4c6febf016fc6653a9

Observation 4e375f8d-a619-4f26-8b6a-d8c8893f666e · outbound

This paper cites SeerAttention-R: Sparse Attention Adaptation for Long Reasoning.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.703884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:a385b7b87136e12048b9e6271f4244055406a2838b02f2805957e3a56bc96350

Observation b238eeaf-a668-4850-a0a1-a91ca84d10dc · outbound

This paper cites Dialogue Without Limits: Constant-Sized KV Caches for Extended Responses in LLMs.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Dialogue Without Limits: Constant-Sized KV Caches for Extended Responses in LLMs

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.595658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:54ddfff73d605ce0bb83eecb531dd9d947fa099762603faba3f4e4e4671c8412

Observation 5b879935-f2aa-4a79-a91e-c62fa0d794b6 · outbound

This paper cites LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.495354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:fb3fbf8d5f5e719bef2038abe0fc795fb50bbde14d147881cf983cf708719fc8

Observation 786afdd9-8099-401c-a79d-856af36a9362 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Measuring Mathematical Problem Solving With the MATH Dataset

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:41:26.606539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:ccf0131f2b2e9c5364fed5810f09e9cc7905ba5f86089f2b16270b691d34e794

Observation 3bbf2dbb-27cc-47c6-be66-704261458a7b · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:722f29335d05770c2b0d60c1d5d6473d269d7b7086845a2d1cd78d8a9cbfcf4f

Observation 1707df9e-d9c0-4a08-b85d-db4c98a22616 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:41:26.514243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:947aea9d173e67a9cd77e7c0511ebae4d2fe887878c3256e4177ec06cfa7de0b

Observation c04ca568-05a7-4eac-b93e-69b5335d3677 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:41:26.566711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:61b0d9abe1b8a6545efe966052a062cba5ec45c4d14afbd59a74a466780ab039

Observation 85741ffd-ef11-49db-9dd8-571b56e02b91 · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:3443e69eb8c89b3bfd83c9039eae3a1bda0dce18abadb466128be466ef85c028

Observation 546364dd-325b-4ad1-a1f1-8c5b47a9126e · outbound

This paper cites Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:26.589479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:fadb7ed2ea31a9fcb072dd44b8131f80384882700d56e2a8cb9d69ac66d6b027

Observation e09839fd-0a65-413a-8973-1ca43760afae · outbound

This paper cites Kediff: Key similarity-based KV cache eviction for long-context LLM inference in resource-constrained environments.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Kediff: Key similarity-based KV cache eviction for long-context LLM inference in resource-constrained environments

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.519981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:64f6ea827ac39902134a8108e856ce526f4879ea3106446a9d992f745f2cbe74

Observation 6328dd7b-6d31-48cb-9a7e-0d14b2507d5f · outbound

This paper cites Daniela Rus.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Daniela Rus

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:26.654056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:7b6c56521de5683c7f4f2720776aa0c81e150ffd58e36a826df6cf1f1751d2aa

Observation 5df17a2f-7838-4dcb-8173-216b94145f06 · outbound

This paper cites VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.670126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:421cff006bbb54fb02a3c5e2c3c3fa5c66c602d38f6e6d74d123cbe65086cbe6

Observation 74f78f2a-04be-427f-a45c-f635d39eb2d6 · outbound

This paper cites Vision-language- action models: Concepts, progress, applications and chal- lenges.arXiv preprint arXiv:2505.04769.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Vision-language- action models: Concepts, progress, applications and chal- lenges.arXiv preprint arXiv:2505.04769

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.445807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:43902a7ac791aa79664082eff95b41d6f6ff5fc4aea9ba606a9772cdf40dda0b

Observation 56215eab-bc41-47e8-9332-38d08d6b5f28 · outbound

This paper cites ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.478146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:87d76290f6c3745b13992d9fb50cdb52ab0f0457e9ebb0651d19e9adabca38e3

Observation 037032b0-6783-450e-b744-9a1f0f70bc3a · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:12:21.110645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:e3c0f802866da854c6e48d225947daa9ab7a0f7a4f0ffa6ecb2f4566b889e122

Observation abc86ec3-430d-4170-84e6-229078555ee7 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Gemini: A Family of Highly Capable Multimodal Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:41:26.571668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:5bd0109ff10951eb25c7a5234174da9cfd33a5765d33904ca2793cf854b92163

Observation 2542ebda-0081-4fd9-b658-cd54ead99e4f · outbound

This paper cites VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.551067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:990124474dc04f0f6d65fa3fe97f15d4ce4d09a1bbe1c9cb6d8c43334dbe1647

Observation 8d4df5aa-921c-45b6-8a7e-2d4d6e7d8870 · outbound

This paper cites MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.648505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:ccaad9930ed83d3aad70dd251740785194cb58d2a5ee4ded208634db24f3b995

Observation bb7f7fc5-0671-4d8e-9eff-b6e8f15902a7 · outbound

This paper cites LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:26.626805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:e1c23880f33b79bd91a15ddb3f142d4fd05f0596e0da5260e3e00a9d85aa621a

Observation d3afb04a-5cda-497d-ab69-06d54b324724 · outbound

This paper cites Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.460656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:c0ad298c9b2c3c1415ac53444d4a76ef04722b78fd6b657879bace9e614bef0b

Observation dcc96a6a-2556-4e52-b608-a0959b909e93 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Efficient Streaming Language Models with Attention Sinks

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:41:26.617659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:648e0a8fc73b679983cd67339c8289dc96189600e512ac32480996da64a66028

Observation a89fbc46-e19b-4389-b915-245443a00efc · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T00:19:20.780377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:b6348abe020e122ee0b7dbe26a5d61d96ab976742190fc6b5923b8c5eb48b63e

Observation e3ebbf7c-c88d-474c-9976-b1d2d5ffe071 · outbound

This paper cites doi: 10.18653/v1/2025.acl-long.736.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction doi: 10.18653/v1/2025.acl-long.736

Reference 32

Resolution
metadata mismatch
doi, observed 2026-05-12T05:06:21.144068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:3ef36c30bc5ec3133efcd86ae1414a82d7d8d73d2f90db6017d4f4f5f955b866

Observation b425d45b-25d9-4bef-9270-e7c6144f21a2 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:b8884ec41082731b4deed1170da909c59acdb9d435d9b288e8cf79669c6b05fa

Observation 14a8b709-b2b2-40bd-a491-717c9768bcae · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:58:32.742607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:e65c676331d0f595d62f0130a841f8cdc767ee8a857ea2f6438db9a334a6e99a

Observation 27aa9edd-9a3a-44da-b882-a9fd4188f70e · outbound

This paper cites Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:36:34.186386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:10171e8b5eb902af0716907141ae4ed8e5a27478d29437c65e7d1a6809df0da3

Observation 407be0f2-ccca-4ef8-b15e-b62afb385155 · outbound

This paper cites Visual Perception.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Visual Perception

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.420665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:2834321c062e3e8bc508240a96dd589e18018be6787a834793d74cfd68fe15a5

Observation 7820a812-dab0-40ed-acee-c8b76e865bcb · outbound

This paper cites an unresolved cited work.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-12T12:36:34.194631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:8e7c650af8d00ac8e39a65e6757ce38d4e45e4ad7cb3a50d3f404c273cf5ca76

Observation acd79c82-aa15-4f2f-9358-90700de6edb8 · outbound

This paper cites Results are averaged over 5 random seeds.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Results are averaged over 5 random seeds

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:36:34.190559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:e587eb69e97d766b89927898c542c40fa616063135992f6bd69afab428f29cea

Pith citing papers

Observation caa0a936-0ccb-42f8-95e6-f35a53116123 · inbound

Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns cites this paper.

Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T02:23:36.089893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:23:36.089893Z digest=sha256:d6a68f65d13985f56aca7dd6f030dec5efe6b9284256d562e42ca44108416a67