Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T04:57:08.746551Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2606.04415.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T04:57:08.746551Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1f85bfb3-65f8-4962-a13c-64c5c8aef64b · outbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Does FlexNPU introduce noticeable overhead compared with direct NPU passthrough? (2) Dynamic PD co-location vs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 249612ad-faab-49d3-b2e2-9277affae70a · outbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 07b84d3b-e19d-4f09-9e95-7ed35415c4fe · outbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Nixie: Efficient, Transparent Temporal Multiplexing for Consumer GPUs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6364194d-4a2f-40d1-ba21-a39fa04f04bc · outbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location StreamBox: A Lightweight GPU SandBox for Serverless Inference Workflow
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c27a46e5-9ec3-4e1d-a457-603bc1358061 · outbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Singularity: Planet-Scale, Preemptive and Elastic Scheduling of AI Workloads
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2946dcdf-102c-415b-9d4d-8d55f1e2bfdf · outbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3e9ddfb4-444e-444a-8c46-5a5f34b82c67 · outbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Tetris: Memory-Efficient Serverless Inference Through Tensor Sharing
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d18f28ca-ef34-4cac-987b-0e6cf618207b · outbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Pre-warming Is Not Enough: Accelerating Serverless Inference With Opportunistic Pre-loading
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9417f6ac-efe6-462a-a6a7-4678f10da5a9 · outbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d22f5b6-1caf-4750-9ef7-bb22ad742fba · outbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Splitwise: Efficient Generative LLM Inference Using Phase Splitting
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 907e3f9b-a739-4dc9-854e-e64753a0d43d · outbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5d615ec3-c1c4-4a6d-bf8a-55b0ab92bae6 · outbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8c40914f-5fa0-4995-8bc9-b0d978ebdeeb · outbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bcc3c02e-c147-4c78-b25a-ed843b950361 · outbound
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
No inbound Pith citation observations are available.