Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T23:39:08.036454Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2511.04791.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T23:39:08.036454Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T04:57:08.746551Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T10:46:52.298443Z
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation faee7694-f5fa-4c23-aa26-734461bdc49e · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing S., Ramjee, R., and Tumanov, A
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d83e1e03-2924-46ef-9fb5-3cef716d1b72 · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing ISBN 9798400712616
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24997dbb-dafd-4851-83f9-d7a0965b0dc5 · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing The Llama 3 Herd of Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae877024-34f8-43a3-9549-b818ea3d9c92 · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a1e4f4d-6d0f-4b72-b764-30d9947e5e7d · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Bul- let: Boosting gpu utilization for llm serving via dy- namic spatial-temporal orchestration.arXiv preprint arXiv:2504.19516,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b00b24c-47c9-44c4-a1ca-fceddc7379c9 · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89f5d42f-215a-4e1d-9c0e-7b0559827ac3 · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing lmcache.ai/2025-04-29-pdbench/
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34f30212-c926-4155-8db8-ba996fb96953 · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Azure llm inference trace 2023,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 089141c0-e46f-48fa-a5c0-9d4bbdda9e17 · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing GPT-4 Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eed99aaf-864d-4c80-a05c-649a38a3c4f9 · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9455ab2d-4938-4969-8633-9e880bf8a1ca · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d4edeec-e363-4910-8dfb-0053f4fd3678 · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Cuda graphs compatibility of attention back- ends, 2025a
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 795a3d79-9edd-4103-9fbb-8ff49e53f522 · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Qwen3 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation badb81ca-3851-4295-8134-fad2b5dd08df · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Nanoflow: To- wards optimal large language model serving throughput
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3558dd2f-fd35-4251-b89e-46968595aba2 · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing ISBN 9781450381376
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 633d21dc-2e09-4afa-a53e-f7dfe3ee0580 · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d4d9977-53b8-4aba-a0ed-77fb97d7427d · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Optimizing slo- oriented llm serving with pd-multiplexing.arXiv preprint arXiv:2504.14489,
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 990953a8-7b1d-4b14-a681-dbdbfbac60a4 · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Gemini: A Family of Highly Capable Multimodal Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1e71ccb-e045-4960-9a4d-e425c2b5e784 · outbound
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing Chow, M., Jahanshahi, A., and Wong, D
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8c40914f-5fa0-4995-8bc9-b0d978ebdeeb · inbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.