Pith. sign in

Paper Citation Record · LEDGER

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2504.19867.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19867 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:03:52.902518Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:57:38.712067Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fd7d91d0-44e7-4401-95ab-fb53c88276eb · inbound

Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving cites this paper.

Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:03:52.902518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:03:52.902518Z digest=sha256:fbf216f3acdb2d44a49e106b8e79fadee4103354f12b64d384ed62016da69175

Observation 633d21dc-2e09-4afa-a53e-f7dfe3ee0580 · inbound

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing cites this paper.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:06.362657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:06.362657Z digest=sha256:c79e0dfceae1470cd8a8eaea47c36f2f79329b40e4dba765a824f9940effcf88

Observation e2ba28d7-9b2c-4bf3-9bb3-e753026a5fc8 · inbound

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities cites this paper.

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.262035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T09:45:57.201837Z digest=sha256:c2dcbd4de5edcc58e5911670ccc9567e1a826c25aac39c53c869eaadb4b2fe5a

Observation e3cf6b4c-9744-43e2-b0f6-bc0407270db5 · inbound

KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving cites this paper.

KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:42:31.180388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T17:42:06.870534Z digest=sha256:53642c97d6847a583b60ea3d10a92642ed7b1bc586b3a9b2d9806ff67d6553a0

Observation 249612ad-faab-49d3-b2e2-9277affae70a · inbound

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location cites this paper.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T10:46:52.308744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:2858690509ae7c2a95faab7a8a602f3bd27ada886522855c35adbb3c1f54ba45

Observation 478a67f2-2dab-4a49-9434-6844a04c21d7 · inbound

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models cites this paper.

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:57:38.713663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T13:30:19.620692Z digest=sha256:5c0f53096105f4070aab34c25c2aa23b37cbf1876801e90b966ee63add764980