Pith. sign in

Paper Citation Record · LEDGER

Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?

As of 21 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 5 inbound Pith citation observations for arXiv:2506.17121.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17121 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:18:45.004345Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:50:08.621959Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T01:06:24.193651Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved8
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87f1e4ad-c051-43af-ad7b-be0e00ae38c1 · outbound

This paper cites an unresolved cited work.

Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs? Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:18:45.294863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:18:44.965845Z digest=sha256:5b42a6e24da3e8d0dc9a47511f2ff74cd73507047080117808ad9a8aca298ba3

Observation 30882127-aac2-4e4a-b152-62297aa62abf · outbound

This paper cites # Input” and “# Output.

Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs? # Input” and “# Output

Reference 2

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T19:18:45.308361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:18:44.962439Z digest=sha256:5f35567028a7ea6f404758dca27e61f95e7c2400d6763dbd93127a0feb8ddeae

Observation 27a5540a-dd0b-4fd9-8e73-4c31b2da8332 · outbound

This paper cites The combination of P + C is usually the best-performing variant in both the PyramidKV and SnapKV groups.

Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs? The combination of P + C is usually the best-performing variant in both the PyramidKV and SnapKV groups

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:18:45.271467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:18:44.972789Z digest=sha256:6db4f23c6a094fbb4bb8f619beeb056af42f316a0df559c98dfd5f7ddea67df3

Observation 003e20fb-272f-40be-a0cc-5474c04899e0 · outbound

This paper cites The Llama 3 Herd of Models.

Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs? The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:18:44.941847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:18:44.941847Z digest=sha256:7746bb4c7cac072c8a9bd02eb445e3e0747963ef048dae4d55d01cad36218bb7

Observation dacb5859-ecd7-4e94-8b30-b6a171695694 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs? DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:18:44.946385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:18:44.946385Z digest=sha256:3af77c08cc2067a99db8f5382c6c9240e5030985945839866fc6346f37e01d46

Observation 3fa5406b-da39-451d-b5a2-32a8095fc2a0 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs? MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:18:44.950535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:18:44.950535Z digest=sha256:9f9124b4d4073edbb60cb941f14c2660ad5f2f9ceb1916f18af7ed1580703c92

Observation da326e20-4b3f-4f2d-84fb-c3a94a0e7fef · outbound

This paper cites Howard Yen, Tianyu Gao, Minmin Hou, Ke Ding, Daniel Fleischer, Peter Izsak, Moshe Wasserblat, and Danqi Chen.

Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs? Howard Yen, Tianyu Gao, Minmin Hou, Ke Ding, Daniel Fleischer, Peter Izsak, Moshe Wasserblat, and Danqi Chen

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:18:44.958963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:18:44.958963Z digest=sha256:f71a3af2da4a0c22268dee10fd5ddaffad7d8b2429685dc7b41746a9e69129df

Observation 679de3fc-8f3d-4173-9c35-11b7481d0372 · outbound

This paper cites It does not usually affect performance.

Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs? It does not usually affect performance

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:18:45.282960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:18:44.969105Z digest=sha256:c3c016da6200043a2c2ea7220e6ce64e0c8641c014fde9e468b55ec5179e45f1

Observation 57d1b2c6-e657-4956-af76-6f42dd2ff6cb · outbound

This paper cites On the other hand, the precise values of the real metrics are noisy and show some variation across different runs.

Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs? On the other hand, the precise values of the real metrics are noisy and show some variation across different runs

Reference 13

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T19:18:45.258407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:18:45.004345Z digest=sha256:660ccf3a4f70a1228d4aedf37702c3b88fd1deb169658286b192cd7f58a1968d

Observation f66b5a4a-fff3-4a7c-95ee-cf1617ef0c7a · outbound

This paper cites SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation.

Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs? SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T19:18:44.954886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:18:44.954886Z digest=sha256:d0b9f248b97ea232b5f12abc9fcbd7d50f306bc05091615f8973749ee5dd6cf4

Observation 679c8929-54e9-4d5a-a57a-c61b4bf08153 · outbound

This paper cites TokenButler: Token Importance is Predictable.

Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs? TokenButler: Token Importance is Predictable

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T19:18:44.892668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:18:44.892668Z digest=sha256:27c9a8209553a1cf5c62fdbd9047b43d11b7157ef8b2b6c67b2b2f6c214b31dc

Observation f56a7f2e-fda1-48bb-b682-3f2d8a74509a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T19:18:44.937775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:18:44.937775Z digest=sha256:b793afa5ab71f97e9d1b9cd58cec0125901e374eb5cf5842dc75de175a759b78

Pith citing papers

Observation afb554fe-504c-4e41-ad74-79dab74009f2 · inbound

MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining cites this paper.

MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:11:42.728494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T18:09:18.157131Z digest=sha256:12d9414cc81f3b77222192faf668cc7250391c4beebc49aa978dc75596855887

Observation d1862ac4-1f89-4a40-8b1e-226026b7054a · inbound

Which Heads Matter for Reasoning? RL-Guided KV Cache Compression cites this paper.

Which Heads Matter for Reasoning? RL-Guided KV Cache Compression Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T10:50:08.621959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:50:08.621959Z digest=sha256:e96b8f9d5638c95defcca04af9230fe2700ebf98520fd18c926cfe9177e1aa0b

Observation 38aa2481-8770-4e99-ac9c-b5d3ea4f18aa · inbound

Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference cites this paper.

Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:25:58.321365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:39:18.387881Z digest=sha256:f699bf73d6cd0a96f91f6203258f03951142c95f4601449b95cdc7389328fe52

Observation 50c8fb55-e2fc-49ea-8e8d-74e3e7e4ef46 · inbound

Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads cites this paper.

Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:06:24.195403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T12:48:02.761437Z digest=sha256:b2ef659ffddece7af437a43dc63bc3bf17475963e5dc3431d4c86e5fe128ea57

Observation 45026465-d3c1-4c1e-b725-a61f43213aa9 · inbound

FlashAccel: Leveraging High-Bandwidth Flash for High-Throughput LLM Inference cites this paper.

FlashAccel: Leveraging High-Bandwidth Flash for High-Throughput LLM Inference Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T13:40:32.154860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:40:32.154860Z digest=sha256:f44bf9fbcdd9bb7fa40b38cf279a401e213ed239bd2d1ddadddc2f940d452f79