Pith. sign in

Paper Citation Record · LEDGER

P/D-Serve: Serving Disaggregated Large Language Model at Scale

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2408.08147.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.08147 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:49:09.890421Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:37:36.387611Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0f66a650-246b-4c2b-aa0d-2e1ff2812c67 · inbound

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching cites this paper.

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:58:11.931791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-23T16:57:46.645061Z digest=sha256:cac2eb796e296dabc37527ef67c4751a7779991d7ea218a5f5d0478c7f5559d4

Observation 73cd2cb3-f961-43ff-bf15-056782726c6c · inbound

A System for Microserving of LLMs cites this paper.

A System for Microserving of LLMs P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:01.802775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:01.802775Z digest=sha256:7cb4478d67318bd3f81b3a6c0c9daea6f9188f303aa129d0520900abc5f39722

Observation 9857202e-31fb-4f66-baee-4a8f4072702d · inbound

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels cites this paper.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.508469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.508469Z digest=sha256:642698d116cffc7686388cf2c2f86427d6cd90acd5e907a9671948b6c84aa928

Observation a0117dca-f92c-49d7-9e15-5433c53e1873 · inbound

Efficiently Serving Large Multimodal Models Using EPD Disaggregation cites this paper.

Efficiently Serving Large Multimodal Models Using EPD Disaggregation P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:29:34.722701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:29:34.722701Z digest=sha256:8f75c03f00250f56bd73245a8784b2dd77b64d83d24486be547e5eb3eef4cda8

Observation 9669077d-a821-4144-b871-35e595bbefee · inbound

HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment cites this paper.

HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T11:32:21.519527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:32:21.519527Z digest=sha256:7d109bcf0d01caf8135a9c80570c4bf513e07ed79e8c5c70bc59dcd2fc943b83

Observation 98edf3ec-67d3-4150-bca6-d89a37b78ddb · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.915120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.915120Z digest=sha256:7d501de185a65fc74a93964cf04e172cfdc4ad7aa1dc53941f3dc8a472328d85

Observation 47c8d098-b2b8-41c6-b079-79c7f3870010 · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:09.890421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:09.890421Z digest=sha256:45e482d8e04dc54949947b3f532e70a608860dd8324e87ab343db13f7b028bb8

Observation 927d0b0a-ff32-4d29-b92a-b6fe002c767d · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:44:58.145493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:fef314c9dabc7fa02c0ecf02789d2a7f16a2b162d3682dad7473febe6eeead7a

Observation 8e2c26b6-936b-41dc-b964-713400ed1a7d · inbound

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation cites this paper.

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:32.597438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:32.597438Z digest=sha256:6ddbb9f62f3634b3580fcf4767c523fdab70f9dfaf80b10172af96fa86b85852

Observation 310aaf6a-68ed-41e1-a343-8c5a5695b27f · inbound

3DLS: A 3D Logic-Stacked Architecture for Disaggregated LLM Serving cites this paper.

3DLS: A 3D Logic-Stacked Architecture for Disaggregated LLM Serving P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:36.389208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-03T04:29:56.401898Z digest=sha256:630a74f23de511d00ab627d3ecdbd05ecd1e73ad17afa570a9fa72092b88fa8a

Observation 08698169-f112-45b2-a260-1d555cd15ccd · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:464ca6bf515685b4cb593dc13f2ed33a163296bc056a5578162763f8702bfd99

Observation 4f3ce9b1-5086-48fd-93b2-002d9156e69f · inbound

HorizonServe: Coordinating Request Scheduling with GPU Sharing for Omni-Model Serving cites this paper.

HorizonServe: Coordinating Request Scheduling with GPU Sharing for Omni-Model Serving P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:29.272650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:29.272650Z digest=sha256:9d9a1d536167e5ab72d24b8e4be63b0b98285b0faf09b9e2612e3e1d344ebe83