Pith. sign in

Paper Citation Record · LEDGER

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches

As of 23 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2605.18825.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.18825 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T22:08:45.016196Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T10:41:08.967261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact5
  • verified fuzzy12
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 25c6b07b-0a5e-4d25-9cdd-d47df59e1cf8 · outbound

This paper cites Gonzalez and Clark W.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Gonzalez and Clark W

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.787084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:8e6a41142eb79fedde1b5887d4f0d3642729b727be895463a8bf20f52e12e65e

Observation 9a54bd5d-6c5c-443a-baa2-bd1c3e134544 · outbound

This paper cites Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention , booktitle =.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention , booktitle =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.750572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:be932ba9e540b646fbf2c3f317c4f2cf97d9032a5f0489b91dd022a4986d875f

Observation 7e884ae1-2dbf-4ec7-b2c4-fa408dea7869 · outbound

This paper cites The Thirteenth International Conference on Learning Representations.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches The Thirteenth International Conference on Learning Representations

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.777996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:d2e3e9be171d13285e734dff8e9d0ac3e05bd5df9b6f6c8fdc6a8eecd0a36e44

Observation 287ad4b2-6007-4523-95f6-ee90fac8d837 · outbound

This paper cites 2025 , eprint=.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches 2025 , eprint=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.754299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:ff7bbdbea544d6858afb93982ff2bbc496ebcb358e8c0b9827105be05ceb6a7b

Observation 30a22117-dc70-4a1c-9726-12149b80bd92 · outbound

This paper cites 2025 , eprint=.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches 2025 , eprint=

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.759517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:d6c1eb3df60951895fa1191017f0bc1156faf321256e1df702e1f87f9962707c

Observation 3f6588c1-3427-483e-b8f6-71085a2ec512 · outbound

This paper cites Learned Prefix Caching for Efficient.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Learned Prefix Caching for Efficient

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.765623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:2c5ba7fc13d3130c004d10e368d54657bca99f2864fe1c52b37eb322d6a9d4f9

Observation 946af7ae-3d2b-44d2-b01b-9dc3e34ecdb6 · outbound

This paper cites QWEN Bailian usage traces (anon.).

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches QWEN Bailian usage traces (anon.)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.773394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:90bef399df697472504305d5216abefdaae627f5e1c5e8de5dc04734e7712517

Observation 176759fe-8140-4c63-b50d-ff26d65463aa · outbound

This paper cites Unicache: A unified batch-level learning-based content caching.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Unicache: A unified batch-level learning-based content caching

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:06.583928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:de152a787fd69692942a5693e370c72f6b37fa69388c628af76528eff392ae1f

Observation 3e535bcb-a957-4ae8-b71c-de68cddc909f · outbound

This paper cites Token prediction as implicit classification to identify LLM -generated text.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Token prediction as implicit classification to identify LLM -generated text

Reference 15

Resolution
verified exact
doi, observed 2026-05-20T22:09:06.591135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:fa24609de0c2133b6f9b38cb6e077afa01e8c657becf438af45ddbbc647deb33

Observation 13527145-d746-4e92-b1a2-107121e6b2fa · outbound

This paper cites Cost-efficient large language model serving for multi-turn conversations with cachedattention.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Cost-efficient large language model serving for multi-turn conversations with cachedattention

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.783225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:6e6ed5c5377b42da86f830f7c44a5cbdf2f76445c426e87824f9a2f87598087a

Observation 0fdbb4bb-1968-4bcf-a4d0-4431f548ebbd · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:09:06.601169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:e40f5cb2c5ba6442519b195e817544d9363da94d574b4938ad039cf78b966d26

Observation 7f108d38-d26a-4f55-9ae2-4b854f539519 · outbound

This paper cites Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:09:06.568706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:3ea601cad375ba570d74c14449378640e92434834659915440bd48896a0c524e

Observation e176b16c-08e1-4438-9d7e-68302af0c6c8 · outbound

This paper cites Marconi: Prefix Caching for the Era of Hybrid LLMs.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Marconi: Prefix Caching for the Era of Hybrid LLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:07.011959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:d7df26a5bd5001a051f0697e6b8ecf69c637fa12cc971c48360bb314a4ebfabc

Observation d7ec5bef-2697-4fe9-be9a-fa4b5f8e50a3 · outbound

This paper cites KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:07.016867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:25d0b5881982063fa01afcabe51ca77fc1a27bdc84d5790079ee0f2b55f8f02d

Observation 3b68995f-6bd6-435a-85eb-9440e5820ee4 · outbound

This paper cites Preble: Efficient distributed prompt scheduling for LLM serving.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Preble: Efficient distributed prompt scheduling for LLM serving

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.743158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:65b5d38e780c2c5d2b94fb8c6e3553ac76779b282f065299830a22450e45962f

Observation 7cea0dd1-b8c9-4162-b42d-0c0a7161c3be · outbound

This paper cites Learned prefix caching for efficient LLM inference.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Learned prefix caching for efficient LLM inference

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.735076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:40d4236b62c564b83b935bf4aed5f9794b5d8b2fab0ec457f9475d393d67ae59

Observation 88f6127d-6148-4e40-8d08-48b7bd619163 · outbound

This paper cites CC-Bench trajectories: Agentic coding task trajectories.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches CC-Bench trajectories: Agentic coding task trajectories

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.739223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:500786af7a3ace0a19005af78bd243b7314ccdd1e3eb652df3959ea1cb9b6350

Observation 70da8ef4-c32d-42a0-a79e-9146900d9507 · outbound

This paper cites Jenga: Effective memory management for serving LLM with heterogeneity.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Jenga: Effective memory management for serving LLM with heterogeneity

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:09:06.607602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:5751e5bf2d5765554e9814d36312042ecbdda4f60273d2bd295d821c0e922820

Observation 9dceac49-8284-4ee1-8c5b-62768e437d6c · outbound

This paper cites Gonzalez, Clark W.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Gonzalez, Clark W

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.746684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:59c8b7b308049ed232f1ccabe2bff14192098b6d6ecec85236dbd550235529ef

Pith citing papers

Observation cf72e7fd-b383-499b-a431-a20a9ebec4c2 · inbound

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference cites this paper.

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T10:41:08.967261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:41:08.967261Z digest=sha256:83915a468eb780413725cb100fb41c99a00862e4b4d612c7fc45314c5d0ea5bb