Pith. sign in

Paper Citation Record · LEDGER

Prompt Cache: Modular Attention Reuse for Low-Latency Inference

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2311.04934.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.04934 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:27:45.534694Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dd159a32-1bc2-4141-8132-23950eee6907 · inbound

SGLang: Efficient Execution of Structured Language Model Programs cites this paper.

SGLang: Efficient Execution of Structured Language Model Programs Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:20:01.060678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:4d8fc74e43be1cb593d140b752f1ed32ccff69fc5a5f91966996c5b679771e9a

Observation faa06f4f-055a-4740-b990-f77bbe3509f4 · inbound

Accelerating Retrieval-Augmented Generation cites this paper.

Accelerating Retrieval-Augmented Generation Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:45:23.028887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:45:23.028887Z digest=sha256:31b1d00b97b6ced7a4df9e9f777428cabd1bf47dc1084d67f73b832ee739bce6

Observation f5c29039-d30e-42c7-8632-b2361c11eb78 · inbound

Offline Learning for Combinatorial Multi-armed Bandits cites this paper.

Offline Learning for Combinatorial Multi-armed Bandits Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-09T20:45:51.269548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:45:51.269548Z digest=sha256:04b9fa959d84fc0368f07c13e98ba9bc02755b6a19c3e600373ab6e53aaec3d0

Observation e53a4bae-80f1-46c9-b5d3-d5a35ebd824d · inbound

The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware) cites this paper.

The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware) Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:19.864940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:19.864940Z digest=sha256:8251c9837b19b84a444fcb31cb030d98c561bceb23eebace1c83e51e75459c16

Observation 16ed6f4d-bfae-4ca9-baa0-85949ef8450c · inbound

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models cites this paper.

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:47.904448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:47.904448Z digest=sha256:057b4d170807ed2884d3d14b61c50756c50debd8470116439ba5f1fd9e873135

Observation 0963909e-d08c-4d8d-a699-ec5caa12d701 · inbound

Efficient Remote KV Cache Reuse with GPU-native Video Codec cites this paper.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.751793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:b85d653cf9a780cd64f1685cbc2cb19db3973ac0b0e909bbfc28f1ce4319a743

Observation 92fa110f-13f9-48be-b398-b59bb357c2cb · inbound

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems cites this paper.

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:15:06.759644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T12:14:09.509302Z digest=sha256:40abca2bfb69d9b65d5449e5241d56d4e2e978cbb095b453b1490dccd4e3f6b9

Observation 2a2e0a21-0aea-4f25-ba68-87e6aa32af2d · inbound

HieraSparse: Hierarchical Semi-Structured Sparse KV Attention cites this paper.

HieraSparse: Hierarchical Semi-Structured Sparse KV Attention Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:16:54.115010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T07:15:19.184970Z digest=sha256:62de53e477c4c5a99e7f8028e28d0c0e1504e067796142c4178676a458ffaa97

Observation bd4add9d-645e-4228-a9c8-d058c812b0d9 · inbound

Continuous Semantic Caching for Low-Cost LLM Serving cites this paper.

Continuous Semantic Caching for Low-Cost LLM Serving Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:01:03.329620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T02:32:28.923367Z digest=sha256:86b6c30aeee23365b6970de0c627a17842878615276ec518d025c365ab3bba52

Observation b3ac640f-1592-41ff-baf2-0a54928892e1 · inbound

Rethinking LLMOps for Fraud and AML: Building a Compliance-Grade LLM Serving Stack cites this paper.

Rethinking LLMOps for Fraud and AML: Building a Compliance-Grade LLM Serving Stack Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:12:07.117712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T02:11:18.259161Z digest=sha256:e0502f671d7c058b01e246f8e8a69e65ea71063f37bb065a3007af51004349af

Observation c60f87ed-cae7-406e-aecf-459bebd1366a · inbound

CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs cites this paper.

CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:23:08.727622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T06:22:28.511930Z digest=sha256:70de437f4de028a9e03d93707d382a7e4368c95e874b08f3cc1199bae303c092

Observation bfaf9d39-e4d1-4421-be0b-24ae52176da5 · inbound

Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics cites this paper.

Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:56:15.724672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T16:02:53.294235Z digest=sha256:a8c8f83e67139dd8d5d5f30e35cc736c687cafe0643966e9c1fc929f20aefe49

Observation 85dadf70-ad99-4a12-bbfd-c82eb014cff8 · inbound

SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance cites this paper.

SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:31.262604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T16:24:31.109508Z digest=sha256:2f6d0fe55544a8361207169080cc9b12a46998abc2e8b60039db071efdc1d3b3

Observation 37c68cb6-0a91-4c8a-87bd-31d350c349f0 · inbound

MiniPIC: Flexible Position-Independent Caching in <100LOC cites this paper.

MiniPIC: Flexible Position-Independent Caching in <100LOC Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-27T07:20:42.194167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T07:12:45.449658Z digest=sha256:c4bda033002b1d7565d2079b1ff59d0647b18dd7ae8e5d82ea4c0b8128a93d38

Observation a4ce376b-a86c-4d18-8d68-2843cdf4baaa · inbound

Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving cites this paper.

Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:29.942545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T17:39:47.477338Z digest=sha256:51903dd20607fa22907484766831eff38efd0c637159dc64463367da55bcfdba

Observation 26caed26-3309-46d9-89ef-878ff58d3e8a · inbound

CRAwLeR -- Cross-Reference Aware Legal Retrieval cites this paper.

CRAwLeR -- Cross-Reference Aware Legal Retrieval Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:49:39.642868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T12:37:33.877749Z digest=sha256:58d574379e0e8d6068dea7d299fe4eaefdf71b6c0fe9b5c1222a42c3c571b7a6

Observation 624a7fda-bf59-4d3b-ac12-10ace44a4d26 · inbound

A Deterministic Control Plane for LLM Coding Agents cites this paper.

A Deterministic Control Plane for LLM Coding Agents Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-26T03:58:57.012444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T03:58:28.523545Z digest=sha256:37e525e30c909e57350ab099d38705dc258bcd18cfcb201403551db98e93873f

Observation 93d08a1b-9ecd-4804-b9fc-8f6a1bac1d30 · inbound

KernelSight-LM: A Kernel-Level LLM Inference Simulator cites this paper.

KernelSight-LM: A Kernel-Level LLM Inference Simulator Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:54:06.243600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T00:48:19.207465Z digest=sha256:7b4ee59ac2bf2af5603e12e0c78312d562669e0872b2f3aa7bb454fec1d1d976

Observation 32144030-bfe6-4ae6-9006-c298a7780836 · inbound

KernelSight-LM: A Kernel-Level LLM Inference Simulator cites this paper.

KernelSight-LM: A Kernel-Level LLM Inference Simulator Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.214792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T23:09:38.092583Z digest=sha256:e3c583889ee43c00c8879f5e467f025b0a08798a3182b56416c1f41d38479c56

Observation a13ebabe-ba70-4929-aba6-63c486ffc31a · inbound

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel cites this paper.

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T02:10:49.354541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:10:49.354541Z digest=sha256:63424dec249873b418fa33671ac3b5193a2bfece0c8226ec98c38e0a1348ba78

Observation de1d7fba-23a8-4471-8cb9-b49677e9a0b0 · inbound

Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images cites this paper.

Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T08:35:41.852997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:35:41.852997Z digest=sha256:71ab50e3c4a5b14de83682eccb4430396dcd008e402d33ef3d3a7f994cf03582

Observation 0be0343d-266a-43a5-9403-52a7359ec686 · inbound

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation cites this paper.

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-30T23:47:57.858093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T23:47:57.858093Z digest=sha256:b8377ed704d814955b381b29186ca769577a2be9b428a0be4c4b4622bc006a39

Observation 5490f1a0-917e-4e28-9f39-e4d82775af4e · inbound

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving cites this paper.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.662090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.662090Z digest=sha256:363a8ec6d7d1453bea98fc8548af2440b43e30816fd38bbbe1311fd3f1e87ee3

Observation 406282b1-80b3-4997-bf03-855d288ec60d · inbound

Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems cites this paper.

Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:45.534694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:45.534694Z digest=sha256:5ac15e09f8481b2ce5a1688eb3f73267f25b789e4b81e46039db16943c17f2fe