Pith. sign in

Paper Citation Record · LEDGER

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference

As of 3 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2604.19769.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.19769 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-14T23:46:59.863921Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:09:07.252227Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29ddb2a5-6661-4574-8292-f5ea4a6da519 · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference LongBench: A bilingual, multitask benchmark for long context understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.438410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:ad69cb6ae6cffaaaf0305ed3253506d78d9a74ffb248f3c83cd664e2a3cdd6ff

Observation 72188809-63a1-469f-ae2f-4cbd91f52452 · outbound

This paper cites an unresolved cited work.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-14T23:48:19.434347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:bf11c1597d7eca03915000c189944622b343a65f1aca2678313c3c1e61e81780

Observation 031e9997-33ae-4377-88c0-6fac96767e37 · outbound

This paper cites Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.430208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:0f31aa16a8eef24972bec3095fc58a327974dd429e18c2a2d509319021d30147

Observation c4b586df-b9e7-405c-8e79-9edf7d377d8a · outbound

This paper cites Ruler: What’s the real context size of your long-context language models?.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Ruler: What’s the real context size of your long-context language models?

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.425546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:6ca5310576b532940d48c18e949e9198f2dccf1cebfc13c99c3d75bc9815a33b

Observation bb6d5432-39d0-445a-a660-77f9b0b38cf7 · outbound

This paper cites Qwen2.5-coder technical report.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Qwen2.5-coder technical report

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.498603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:1bab1a8ff8691d35ca304c21afa29c321cb4793fb1deb02785497d2a86afd768

Observation 665954c3-1d4a-4ad8-b0ec-0231bf90ee5f · outbound

This paper cites an unresolved cited work.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-14T23:48:19.460331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:7297caab1d813903bad5546ec41fc61f32baa4f138953d38b37dfd21e4474f45

Observation bbfb79de-7fd1-4f14-a545-8bd5ab549c36 · outbound

This paper cites KVPR: Efficient LLM in- ference with I/O-aware KV cache partial recomputation.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference KVPR: Efficient LLM in- ference with I/O-aware KV cache partial recomputation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.464413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:dacb144ff4331b15127f3708406ed6da4cd045d858afc29990e0d48fc094dc0c

Observation 78f74fd8-0de9-47c0-9e61-78802fb8e95d · outbound

This paper cites [Jianget al., 2026 ] Bo Jiang, Taolue Yang, Youyuan Liu, Xubin He, Sheng Di, and Sian Jin.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference [Jianget al., 2026 ] Bo Jiang, Taolue Yang, Youyuan Liu, Xubin He, Sheng Di, and Sian Jin

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.522699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:ab2287f6c09027003fc0c1ec92f769a96224fd38e1ca7e4f74a0246c3f09df39

Observation 8aec22ed-6b29-4f7e-af93-2730afc50783 · outbound

This paper cites Reformer: The efficient transformer.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Reformer: The efficient transformer

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.476854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:6a0738d97ea49fc223c697221457f155a24cc38730f497854ac918bf899e37e4

Observation 4a9bde02-e4e0-4b1b-8458-99c1e830c54e · outbound

This paper cites Cachegen: Kv cache compression and streaming for fast large language model serving.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Cachegen: Kv cache compression and streaming for fast large language model serving

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.491644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:04f68c5c0894c14ecf1c85c4540b4cfae8479f9c10fcf9ab55873908aa094d49

Observation 32f5ccc4-85cf-425f-86b8-afd6adfd0a48 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T23:48:19.056976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:17ac0d9fc48dcffb44c99961ad52533c82eb8f557e41406e85b9f2058eff8693

Observation 3fb4c6b4-fb6f-490e-addd-e96cb520112a · outbound

This paper cites Freekv: Boosting kv cache retrieval for efficient llm inference.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Freekv: Boosting kv cache retrieval for efficient llm inference

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.451759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:004e037378306c770ebefe47d00fc162f83477755757f4bdfccae9bdec0eba51

Observation 80c4a8ef-c933-43cd-91e3-e8a9cb5aeeeb · outbound

This paper cites MiniKV: Pushing the limits of 2-bit KV cache via compression and system co-design for efficient long context inference.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference MiniKV: Pushing the limits of 2-bit KV cache via compression and system co-design for efficient long context inference

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.502500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:a1cf0781c3e4b70b272878d1f588e58c5f624f764c29893718a4a4e861bdb348

Observation 8a5ebbf6-a19d-47cb-8ff1-76c3f7f96e58 · outbound

This paper cites [Shenget al., 2023 ] Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Re, Ion Stoica, and Ce Zhang.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference [Shenget al., 2023 ] Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Re, Ion Stoica, and Ce Zhang

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.495342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:749076396620d58fba8a0093aee85acbb5e8f71bebfac842df1a8dfd4f87ae10

Observation 155289ca-e674-4f90-aa5b-72ed8b94d694 · outbound

This paper cites Shadowkv: Kv cache in shadows for high-throughput long-context llm inference.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Shadowkv: Kv cache in shadows for high-throughput long-context llm inference

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.481161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:995c6c9ee17fae196213a6448a6e4a40867d7d912cc394fd8100bccf04787d98

Observation 998898b5-b396-4997-9aee-6206a0d8b1f3 · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Llama 2: Open foundation and fine-tuned chat models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.485461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:789fcef0f75bf62cdbcc22af7cbe574d6e7beb04959f34fc434fddf33c668ec0

Observation 1525ca07-aedf-4f8e-b6e6-15973a7367fa · outbound

This paper cites Leave no document behind: Benchmarking long-context LLMs with extended multi-doc QA.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Leave no document behind: Benchmarking long-context LLMs with extended multi-doc QA

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.549878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:498966069969954683ed82c69ef44f133d84d3a3434a737089bd18b0035d7bd1

Observation 811fba38-13a6-4b3f-8364-a0fdeb32bd69 · outbound

This paper cites [Wanget al., 2025 ] Dongwei Wang, Zijie Liu, Song Wang, Yuxin Ren, Jianing Deng, Jingtong Hu, Tianlong Chen, and Huanrui Yang.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference [Wanget al., 2025 ] Dongwei Wang, Zijie Liu, Song Wang, Yuxin Ren, Jianing Deng, Jingtong Hu, Tianlong Chen, and Huanrui Yang

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.557757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:f5dacbe9de2672d8d7fc717b55eb58ab65b966a910fd3ed7bd75fe1a3ede9420

Observation 7ff87bee-b022-4eca-a3a7-7a6a9ef61a1a · outbound

This paper cites Efficient streaming language models with attention sinks.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Efficient streaming language models with attention sinks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.447436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:4e352df7cbf703d8b0c349c753dc13eeb5ba160a9c3291508b2d048048930060

Observation 311491fb-a1f3-4324-a233-f63d5b7a9d7c · outbound

This paper cites H 2o: Heavy-hitter oracle for efficient generative inference of large language models.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference H 2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.472189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:7f5fd9c12924267430ba61e23381ae04551cd00a76b0d1bdc1e5f31ffface3f9

Observation 8632b857-300a-4f2e-b393-00db69115fa9 · outbound

This paper cites an unresolved cited work.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-14T23:48:19.442909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:ff6af7d49946c60003e1520f1a686708f1b8674ead671d3981a5e08bfc4040a2

Observation 9f456916-6588-4b93-bcd4-7ba6904fd3d4 · outbound

This paper cites ProcessBench: Identify- ing process errors in mathematical reasoning.

TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference ProcessBench: Identify- ing process errors in mathematical reasoning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T23:48:19.456353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T23:46:59.863921Z digest=sha256:e417dab89b02c7e6a8ef970f1a08eed2c17ae620b0678d429a7601c21dc78792

Pith citing papers

Observation c758b92d-0a3e-4a3c-9ab5-6252a2e02f4a · inbound

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization cites this paper.

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:07.252227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:07.252227Z digest=sha256:bce72dbe28cdc5868f31b459c4fcba2eb4aa2b377198b0ac055b8996cedfa5dc