Pith. sign in

Paper Citation Record · LEDGER

Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2403.19708.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.19708 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:22:42.522976Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 120312c5-8335-48a5-8248-6cb87de67a0f · inbound

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts cites this paper.

Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:42.522976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:42.522976Z digest=sha256:85f13eaf02ffdbc9dc611f2d096ce2a9d44f16d4ad7ff54e5fb7a805a3667f59

Observation 6da3e595-ad2f-45d1-9bb1-b5038c8a11c0 · inbound

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live cites this paper.

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:00:39.752838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T01:58:23.234348Z digest=sha256:88f65b585cac518446c317941975fb518e81171252ad0c2f8db03587ed49d93d

Observation b8e1e9e6-cf7e-48cc-af69-590c16514dfa · inbound

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live cites this paper.

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T00:19:17.443187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:19:17.443187Z digest=sha256:86b44d294a1919be1528a6cff67978ab5bec9cffd24762134949a030bbf74d68

Observation c62d904d-4177-4a83-87e2-f7e3a47616ac · inbound

Low-Scaling Many-Body Green's Function Calculations for Molecular Systems via Interacting-Bath Dynamical Embedding Theory cites this paper.

Low-Scaling Many-Body Green's Function Calculations for Molecular Systems via Interacting-Bath Dynamical Embedding Theory Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T13:32:41.086180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:32:41.086180Z digest=sha256:d5a8777a90a011c17e2d7014c69e7c5606f35111788fcc31ae0825c0925a9c53

Observation 2ded0a11-a249-4292-b383-e932afa77729 · inbound

TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing cites this paper.

TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:03:04.268884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T18:02:46.266111Z digest=sha256:e451174b6038870f6a95808535169e3a451722655271723057f1a2b0a10a3757

Observation ef6f2ec0-955f-444e-9a95-a4319a2c7b74 · inbound

Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling cites this paper.

Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:36:36.708261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T06:31:36.776819Z digest=sha256:95cb1499d8d5f41535275acef2afa80767aef54aed0f86916b546ad13b5f6f6c

Observation ed68c6de-3bca-45cd-aa64-cbabb88a2c73 · inbound

Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving cites this paper.

Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:31:10.503096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T16:19:33.613685Z digest=sha256:a993436e30fb96f18a2f7572313ee520a4eb254250e0be88a5eed0664861ffec

Observation a004ef01-6537-40fa-b2a6-994d5438e679 · inbound

Rethinking LLMOps for Fraud and AML: Building a Compliance-Grade LLM Serving Stack cites this paper.

Rethinking LLMOps for Fraud and AML: Building a Compliance-Grade LLM Serving Stack Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:12:07.134044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T02:11:18.259161Z digest=sha256:0b30a75d90ef245d88ab1a9bd17d56aab07b587003c25fd8cc5e7c7ce90f6065

Observation a3af4b67-9b5d-4ee1-b048-6b7d7a7b58e8 · inbound

Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving cites this paper.

Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:36:26.252091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T11:38:05.435505Z digest=sha256:8b0e5adde837f60ecd3c7aa271a6b447422597359d262b32ca9cd5628435c484

Observation 595e1dc6-daa4-413f-851e-8f9ee8550ae6 · inbound

ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories cites this paper.

ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:13.432502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T08:11:57.929452Z digest=sha256:f2813914e61655758eb523d87afa687f66a2cf2264aad9abe6ea27ccffa22064

Observation 1ff9b9ca-560b-4466-99b2-d36e0b56ea61 · inbound

KernelSight-LM: A Kernel-Level LLM Inference Simulator cites this paper.

KernelSight-LM: A Kernel-Level LLM Inference Simulator Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:54:06.292563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T00:48:19.207465Z digest=sha256:64f2d731343b9d7f44335eec197759e9284c22f303c9c32aefe4d05f19d01554

Observation fbf19b2e-d612-4291-95d5-fd9aae4c0e39 · inbound

KernelSight-LM: A Kernel-Level LLM Inference Simulator cites this paper.

KernelSight-LM: A Kernel-Level LLM Inference Simulator Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.272040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:09:38.092583Z digest=sha256:a01740a8116678a128265460b61c13d232210ff03d60a24736d84f0a4ed9cfb0

Observation 49cb944e-0b29-42de-8f4d-934d278c33a2 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.007406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:0a27ff0f2438f99f1cefcec88cb51f262cdd52159b8392a2cfdab8acba660e98

Observation 9f592382-7ce0-4cf9-965a-1f805bb5f982 · inbound

HyMCache: A KV Cache Framework for Multi-Turn LLM Serving with CXL-Hybrid Memory cites this paper.

HyMCache: A KV Cache Framework for Multi-Turn LLM Serving with CXL-Hybrid Memory Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T15:58:25.475059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:58:25.475059Z digest=sha256:b1b0796eddc76263edf6cfcf73476a4d6108d84239c1880156cdb7b8d47fcc05