Pith. sign in

Paper Citation Record · LEDGER

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing

As of 4 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2511.04791.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.04791 v2

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T23:39:08.036454Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T04:57:08.746551Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T10:46:52.298443Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation faee7694-f5fa-4c23-aa26-734461bdc49e · outbound

This paper cites S., Ramjee, R., and Tumanov, A.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing S., Ramjee, R., and Tumanov, A

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:05.344401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:05.344401Z digest=sha256:5fb44d0693fe138fc93b0434343a0d06658682f85bebd5ef0b2e4de65d695256

Observation d83e1e03-2924-46ef-9fb5-3cef716d1b72 · outbound

This paper cites ISBN 9798400712616.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing ISBN 9798400712616

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:05.963627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:05.963627Z digest=sha256:5703cd412882d97e866bc4e597c33fb60885a65c0fec514c3c0b5b6396174f45

Observation 24997dbb-dafd-4851-83f9-d7a0965b0dc5 · outbound

This paper cites The Llama 3 Herd of Models.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:06.217757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:06.217757Z digest=sha256:d13feb6b720c5a2b3eecf9348010f639308edb968398dbc0dea4e13d14cada3b

Observation ae877024-34f8-43a3-9549-b818ea3d9c92 · outbound

This paper cites MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:06.591452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:06.591452Z digest=sha256:a4638f7459498b2618163c9496e90acde420bc47282de14ee56c0faeea682be4

Observation 3a1e4f4d-6d0f-4b72-b764-30d9947e5e7d · outbound

This paper cites Bul- let: Boosting gpu utilization for llm serving via dy- namic spatial-temporal orchestration.arXiv preprint arXiv:2504.19516,.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Bul- let: Boosting gpu utilization for llm serving via dy- namic spatial-temporal orchestration.arXiv preprint arXiv:2504.19516,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:06.693807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:06.693807Z digest=sha256:654ead95b5901a9c60a94a6e65793078e43084fa46c79482f229edee1f972dfb

Observation 6b00b24c-47c9-44c4-a1ca-fceddc7379c9 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:06.769325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:06.769325Z digest=sha256:cef327d1b16804735f17dc68aa6126848991f3aacf9865c738e78effe9649600

Observation 89f5d42f-215a-4e1d-9c0e-7b0559827ac3 · outbound

This paper cites lmcache.ai/2025-04-29-pdbench/.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing lmcache.ai/2025-04-29-pdbench/

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:06.882234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:06.882234Z digest=sha256:ab24ab221ac8b3dbeee3ff062ac3c263fb4d150662b95cb15a6a6888266a4dfb

Observation 34f30212-c926-4155-8db8-ba996fb96953 · outbound

This paper cites Azure llm inference trace 2023,.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Azure llm inference trace 2023,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:07.099305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:07.099305Z digest=sha256:a8814bdb0231dbc5a1733892bc9b0b990663b71e55244ac627b523e0f994c826

Observation 089141c0-e46f-48fa-a5c0-9d4bbdda9e17 · outbound

This paper cites GPT-4 Technical Report.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing GPT-4 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:07.249012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:07.249012Z digest=sha256:ceac8078370d85d05f49f813bcbec998f895dbb67fecdfd76b82ea99306797ab

Observation eed99aaf-864d-4c80-a05c-649a38a3c4f9 · outbound

This paper cites an unresolved cited work.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:07.384783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:07.384783Z digest=sha256:920fc87f077625aa158c46f9277cd59820b8454e6bccd3eb9cfc7e431859ee6b

Observation 9455ab2d-4938-4969-8633-9e880bf8a1ca · outbound

This paper cites Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:07.565606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:07.565606Z digest=sha256:18d91f2317266c393cd0cb3b9b7535b753df7a5115b32baeead076c4f1b57716

Observation 4d4edeec-e363-4910-8dfb-0053f4fd3678 · outbound

This paper cites Cuda graphs compatibility of attention back- ends, 2025a.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Cuda graphs compatibility of attention back- ends, 2025a

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:07.694976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:07.694976Z digest=sha256:1ba65eb329e1b91fc76b11951e86b5e168eeca9dde78e94c437a8fbab774f27f

Observation 795a3d79-9edd-4103-9fbb-8ff49e53f522 · outbound

This paper cites Qwen3 Technical Report.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Qwen3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:07.864886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:07.864886Z digest=sha256:8f2211aaad91aa3d4b99f665aa928b59c98f170ecdb471197360a29f51d3249c

Observation badb81ca-3851-4295-8134-fad2b5dd08df · outbound

This paper cites Nanoflow: To- wards optimal large language model serving throughput.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Nanoflow: To- wards optimal large language model serving throughput

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:08.036454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:08.036454Z digest=sha256:b2ff88d82286a2b3b85d330756a716bbf84d55aec46d433d55647d91dc245bb0

Observation 3558dd2f-fd35-4251-b89e-46968595aba2 · outbound

This paper cites ISBN 9781450381376.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing ISBN 9781450381376

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:05.802487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:05.802487Z digest=sha256:fa9cf3588ce77a0b7ec7c217489ecec42afca2d13a9c832f387d1b2a7871b702

Observation 633d21dc-2e09-4afa-a53e-f7dfe3ee0580 · outbound

This paper cites semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:06.362657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:06.362657Z digest=sha256:e4676a3b03d0475337a957a46ea027a75e675371e7728309db609bae92186dcf

Observation 9d4d9977-53b8-4aba-a0ed-77fb97d7427d · outbound

This paper cites Optimizing slo- oriented llm serving with pd-multiplexing.arXiv preprint arXiv:2504.14489,.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Optimizing slo- oriented llm serving with pd-multiplexing.arXiv preprint arXiv:2504.14489,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:05.625797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:05.625797Z digest=sha256:1311b72ac4746ccd0767beba6d24ef796826eee69747585deeb533f1dc7a48c4

Observation 990953a8-7b1d-4b14-a681-dbdbfbac60a4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Gemini: A Family of Highly Capable Multimodal Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:06.136972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:06.136972Z digest=sha256:08582a4edd4bce0d182a4997fa5c5e169febdd684f13462ee0e184f0ffb1c25e

Observation f1e71ccb-e045-4960-9a4d-e425c2b5e784 · outbound

This paper cites Chow, M., Jahanshahi, A., and Wong, D.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Chow, M., Jahanshahi, A., and Wong, D

Reference 2025

Resolution
verified exact
doi, observed 2026-08-03T23:44:06.428629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-08-03T23:39:05.462331Z digest=sha256:37d3c66aac6b323ffd206363c4d05410613bbcb4d1105a848b0adfea505162e1

Pith citing papers

Observation 8c40914f-5fa0-4995-8bc9-b0d978ebdeeb · inbound

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location cites this paper.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-02T10:46:52.299777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:c8ecba851d203b5e6c5693bc6bf79ad0060e4bbecd12e8814d1749fa13d38f36