Pith. sign in

Paper Citation Record · LEDGER

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches

As of 23 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2605.18825.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.18825 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T22:08:45.016196Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T10:41:08.967261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact5
  • verified fuzzy12
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 25c6b07b-0a5e-4d25-9cdd-d47df59e1cf8 · outbound

This paper cites Gonzalez and Clark W.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Gonzalez and Clark W

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.787084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:9116a799b4c3df816d120345022a5ea1e7b8a3d8c9dd325e2d120eaf54d191e3

Observation 9a54bd5d-6c5c-443a-baa2-bd1c3e134544 · outbound

This paper cites Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention , booktitle =.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention , booktitle =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.750572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:9b64758c8c0c5a14a79148726d897c6cb7b6b5aae68340b9c0e2a8aa376dcfbd

Observation 7e884ae1-2dbf-4ec7-b2c4-fa408dea7869 · outbound

This paper cites The Thirteenth International Conference on Learning Representations.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches The Thirteenth International Conference on Learning Representations

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.777996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:c1a512792c09abca6cb0b31eb13f2533f2fd18314ab27c07dfb6624b1bf7dec8

Observation 287ad4b2-6007-4523-95f6-ee90fac8d837 · outbound

This paper cites 2025 , eprint=.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches 2025 , eprint=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.754299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:2c2caae02f0878ff30b9af577d00ea4423e74eee4236b743d93720be44cc7cb4

Observation 30a22117-dc70-4a1c-9726-12149b80bd92 · outbound

This paper cites 2025 , eprint=.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches 2025 , eprint=

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.759517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:93cac670c636c17f5002ca39b87535015d351601747d55d5e486b9d1c084b9f8

Observation 3f6588c1-3427-483e-b8f6-71085a2ec512 · outbound

This paper cites Learned Prefix Caching for Efficient.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Learned Prefix Caching for Efficient

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.765623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:45d5c6ed640a134af4e79b25a3ed47f701495e1afed4eb41d6bf66632c8d8498

Observation 946af7ae-3d2b-44d2-b01b-9dc3e34ecdb6 · outbound

This paper cites QWEN Bailian usage traces (anon.).

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches QWEN Bailian usage traces (anon.)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.773394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:4f8e725409cfedb4d8230381453cd515dfdd289d2192c57785f71ca74a987591

Observation 176759fe-8140-4c63-b50d-ff26d65463aa · outbound

This paper cites Unicache: A unified batch-level learning-based content caching.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Unicache: A unified batch-level learning-based content caching

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:06.583928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:068a16608952f0ce57541a40898b78a43925de9f3ef9302768f9c356829e0e22

Observation 3e535bcb-a957-4ae8-b71c-de68cddc909f · outbound

This paper cites Token prediction as implicit classification to identify LLM -generated text.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Token prediction as implicit classification to identify LLM -generated text

Reference 15

Resolution
verified exact
doi, observed 2026-05-20T22:09:06.591135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:15a5233cc9cc00ec4c841fcd9502e9aac829140e6abcaa0a195ff02daa72bc35

Observation 13527145-d746-4e92-b1a2-107121e6b2fa · outbound

This paper cites Cost-efficient large language model serving for multi-turn conversations with cachedattention.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Cost-efficient large language model serving for multi-turn conversations with cachedattention

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.783225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:9821b85cb43776f4e55a69c3aea1be4c84b470c80417ca962ce509d1c73499e2

Observation 0fdbb4bb-1968-4bcf-a4d0-4431f548ebbd · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:09:06.601169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:399dfb49a006ea46125cac2e12bee0308b3445d60f73b71eaee5697ba2232089

Observation 7f108d38-d26a-4f55-9ae2-4b854f539519 · outbound

This paper cites Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:09:06.568706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:0234acb0f2fe6025e41143217684a4b950d718733a3e05222096173abe008e1a

Observation e176b16c-08e1-4438-9d7e-68302af0c6c8 · outbound

This paper cites Marconi: Prefix Caching for the Era of Hybrid LLMs.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Marconi: Prefix Caching for the Era of Hybrid LLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:07.011959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:9fff4005da8bc3aab5d91ed51ddb56818a4c83ed6c02c2481190f624951a954d

Observation d7ec5bef-2697-4fe9-be9a-fa4b5f8e50a3 · outbound

This paper cites KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:07.016867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:b8ad7135a0ce1cdde9fd1a40f00167564a2d8e7eaf820a560946b7aed341c591

Observation 3b68995f-6bd6-435a-85eb-9440e5820ee4 · outbound

This paper cites Preble: Efficient distributed prompt scheduling for LLM serving.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Preble: Efficient distributed prompt scheduling for LLM serving

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.743158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:b065cb700ecf0e7b71e2576e3db2a7c00c8dad450b9aefc8c81e0cbcfc71710e

Observation 7cea0dd1-b8c9-4162-b42d-0c0a7161c3be · outbound

This paper cites Learned prefix caching for efficient LLM inference.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Learned prefix caching for efficient LLM inference

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.735076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:74f17f675e9915c26a35ce9c41b106894dc6bf47cbba7a8abb9f510cd9aeb729

Observation 88f6127d-6148-4e40-8d08-48b7bd619163 · outbound

This paper cites CC-Bench trajectories: Agentic coding task trajectories.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches CC-Bench trajectories: Agentic coding task trajectories

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.739223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:38f4e921226b59dfc2a4de1bddf6024e7b7d4a94857141cd18915c1da51ac4bb

Observation 70da8ef4-c32d-42a0-a79e-9146900d9507 · outbound

This paper cites Jenga: Effective memory management for serving LLM with heterogeneity.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Jenga: Effective memory management for serving LLM with heterogeneity

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:09:06.607602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:34a48be2ce1b4c620ad28ae4a7c8f38ed8c8975130e839abce41fc523def7f28

Observation 9dceac49-8284-4ee1-8c5b-62768e437d6c · outbound

This paper cites Gonzalez, Clark W.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches Gonzalez, Clark W

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:09:07.746684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:4c79ed7c763ebeccdbfa5aed6a2e2f1876a91c64921842414faa8e987701bfbd

Pith citing papers

Observation cf72e7fd-b383-499b-a431-a20a9ebec4c2 · inbound

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference cites this paper.

MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T10:41:08.967261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:41:08.967261Z digest=sha256:83915a468eb780413725cb100fb41c99a00862e4b4d612c7fc45314c5d0ea5bb