Pith. sign in

Paper Citation Record · LEDGER

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location

As of 11 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2606.04415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.04415 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T04:57:08.746551Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1f85bfb3-65f8-4962-a13c-64c5c8aef64b · outbound

This paper cites Does FlexNPU introduce noticeable overhead compared with direct NPU passthrough? (2) Dynamic PD co-location vs.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Does FlexNPU introduce noticeable overhead compared with direct NPU passthrough? (2) Dynamic PD co-location vs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T04:57:08.746551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:2f76de7f190b5db0fe31f75de0bcc1898c7ce0813a4986167f5a50171b25a0e6

Observation 249612ad-faab-49d3-b2e2-9277affae70a · outbound

This paper cites semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T10:46:52.308744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:39d0318b5e235b0994fea9ef8d8dcc6ec114fb45fa9600f6e496ae618a20f270

Observation 07b84d3b-e19d-4f09-9e95-7ed35415c4fe · outbound

This paper cites Nixie: Efficient, Transparent Temporal Multiplexing for Consumer GPUs.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Nixie: Efficient, Transparent Temporal Multiplexing for Consumer GPUs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:46:52.304403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:cd0426017f21bcc7c30f04fd158d49e8fe7b190a5e622bacab2d4ca0a65ee9c0

Observation 6364194d-4a2f-40d1-ba21-a39fa04f04bc · outbound

This paper cites StreamBox: A Lightweight GPU SandBox for Serverless Inference Workflow.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location StreamBox: A Lightweight GPU SandBox for Serverless Inference Workflow

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T04:57:08.746551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:c7ef3144b2b25cefb2756d5d2bf0e56bad895766d2c58ffde78b14d7341aff5d

Observation c27a46e5-9ec3-4e1d-a457-603bc1358061 · outbound

This paper cites Singularity: Planet-Scale, Preemptive and Elastic Scheduling of AI Workloads.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Singularity: Planet-Scale, Preemptive and Elastic Scheduling of AI Workloads

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:46:52.303080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:fc6501868b7f722f68778ecaa5d46ee7ea6b179715fde3f91efb07a25c835d3a

Observation 2946dcdf-102c-415b-9d4d-8d55f1e2bfdf · outbound

This paper cites Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T10:46:52.314989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:66134fa90f2af5e30c43ea0b4d2532423857b30a5a82d407ccadc4d4956cea69

Observation 3e9ddfb4-444e-444a-8c46-5a5f34b82c67 · outbound

This paper cites Tetris: Memory-Efficient Serverless Inference Through Tensor Sharing.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Tetris: Memory-Efficient Serverless Inference Through Tensor Sharing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T04:57:08.746551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:37214a7d8f5f67b1ad7e1be7fc3805659e46d199ba8079e8d5d5618726c4ba93

Observation d18f28ca-ef34-4cac-987b-0e6cf618207b · outbound

This paper cites Pre-warming Is Not Enough: Accelerating Serverless Inference With Opportunistic Pre-loading.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Pre-warming Is Not Enough: Accelerating Serverless Inference With Opportunistic Pre-loading

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T04:57:08.746551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:1f3f0f294427d2197231ec8a19d890934ca485f6f15ee1159af0e7c69dab2f57

Observation 9417f6ac-efe6-462a-a6a7-4678f10da5a9 · outbound

This paper cites Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T04:57:08.746551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:8eabf8cda076c387e7a92285ab63f2868557d261ce30592013138f4b46f562c1

Observation 3d22f5b6-1caf-4750-9ef7-bb22ad742fba · outbound

This paper cites Splitwise: Efficient Generative LLM Inference Using Phase Splitting.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Splitwise: Efficient Generative LLM Inference Using Phase Splitting

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T04:57:08.746551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:bc41c5335a1a582f9a5e04c0b0606c3a71b58de345ac0be9a48f882ca19f281a

Observation 907e3f9b-a739-4dc9-854e-e64753a0d43d · outbound

This paper cites MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T10:46:52.317722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:1f37d661839e57b27a57dea278a78d93ccc519116c76c516e31d6724e2c16e4a

Observation 5d615ec3-c1c4-4a6d-bf8a-55b0ab92bae6 · outbound

This paper cites Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:46:52.305907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:5591a013be3129605d16152710b8dd82ac1a811c7808b51fe7dd701bee95102b

Observation 8c40914f-5fa0-4995-8bc9-b0d978ebdeeb · outbound

This paper cites DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-02T10:46:52.299777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:2b451b198e4ba645d0b0b678a8f037ca29666ab4234455ad4692ebf306c3fb83

Observation bcc3c02e-c147-4c78-b25a-ed843b950361 · outbound

This paper cites RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T10:46:52.312389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:7676f2075f00fe99b3a35ec0ecc00c31598b23ad013d2c39eb473048f7e12943

Pith citing papers

No inbound Pith citation observations are available.