Pith. sign in

Paper Citation Record · LEDGER

P/D-Serve: Serving Disaggregated Large Language Model at Scale

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2408.08147.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.08147 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:49:09.890421Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:37:36.387611Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0f66a650-246b-4c2b-aa0d-2e1ff2812c67 · inbound

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching cites this paper.

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:58:11.931791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T16:57:46.645061Z digest=sha256:b2a50e06e197c6f68d6c2728f8784af54d38c05b4b54693e949a802aed90dce8

Observation 73cd2cb3-f961-43ff-bf15-056782726c6c · inbound

A System for Microserving of LLMs cites this paper.

A System for Microserving of LLMs P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:01.802775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:01.802775Z digest=sha256:7cb4478d67318bd3f81b3a6c0c9daea6f9188f303aa129d0520900abc5f39722

Observation 9857202e-31fb-4f66-baee-4a8f4072702d · inbound

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels cites this paper.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.508469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.508469Z digest=sha256:8d7aff84773b66a1fa3635526b68ee969591b883d22138b388c6fe84d6c5951b

Observation a0117dca-f92c-49d7-9e15-5433c53e1873 · inbound

Efficiently Serving Large Multimodal Models Using EPD Disaggregation cites this paper.

Efficiently Serving Large Multimodal Models Using EPD Disaggregation P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:29:34.722701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:29:34.722701Z digest=sha256:bcd81e709ea86164b6c7b565dac23021fc9a034b59bce10fe4abba628080fd9c

Observation 9669077d-a821-4144-b871-35e595bbefee · inbound

HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment cites this paper.

HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T11:32:21.519527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:32:21.519527Z digest=sha256:7d109bcf0d01caf8135a9c80570c4bf513e07ed79e8c5c70bc59dcd2fc943b83

Observation 98edf3ec-67d3-4150-bca6-d89a37b78ddb · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:22.915120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:22.915120Z digest=sha256:32b36563e5110cd3c07f1958740a2589a3ca874b01480562f5a334fa685d4955

Observation 47c8d098-b2b8-41c6-b079-79c7f3870010 · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:09.890421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:09.890421Z digest=sha256:5afe2ef1cc242378ceea9683ce0858fdfb6626980b824b1ab7a6079ee207bf04

Observation 927d0b0a-ff32-4d29-b92a-b6fe002c767d · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:44:58.145493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:62eb386ce25b16cf71a9c7f1ae4d2bcc39c9fca63c5fdf8c4eb14671604afc9c

Observation 8e2c26b6-936b-41dc-b964-713400ed1a7d · inbound

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation cites this paper.

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:32.597438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:32.597438Z digest=sha256:d83b28a84c3e8ac20ee1d4872a1caa4eb292f3ce22d23b36df82ac057aff3d1b

Observation 310aaf6a-68ed-41e1-a343-8c5a5695b27f · inbound

3DLS: A 3D Logic-Stacked Architecture for Disaggregated LLM Serving cites this paper.

3DLS: A 3D Logic-Stacked Architecture for Disaggregated LLM Serving P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:36.389208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-03T04:29:56.401898Z digest=sha256:f059f0bba37a5fe6257f2e359753971e3f41a2f8bc5bda2a46ca38b515db1351

Observation 08698169-f112-45b2-a260-1d555cd15ccd · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:464ca6bf515685b4cb593dc13f2ed33a163296bc056a5578162763f8702bfd99

Observation 4f3ce9b1-5086-48fd-93b2-002d9156e69f · inbound

HorizonServe: Coordinating Request Scheduling with GPU Sharing for Omni-Model Serving cites this paper.

HorizonServe: Coordinating Request Scheduling with GPU Sharing for Omni-Model Serving P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:29.272650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:29.272650Z digest=sha256:9d9a1d536167e5ab72d24b8e4be63b0b98285b0faf09b9e2612e3e1d344ebe83