Pith. sign in

Paper Citation Record · LEDGER

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

As of 19 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 7 inbound Pith citation observations for arXiv:2504.19867.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19867 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:45:34.461813Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T05:50:05.675450Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:57:38.712067Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved24
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47776202-1c19-4d62-8f11-68028e9f4b96 · outbound

This paper cites Language models are few-shot learners.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Language models are few-shot learners

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.169537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.266574Z digest=sha256:28544015e6ce50a97ca971b56e46dc4e7c5c233ed20087f94c55b7fce6983ffb

Observation 90f40132-4b7b-4c0b-b0b0-a0ae3f23260a · outbound

This paper cites Towards a Human-like Open-Domain Chatbot.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Towards a Human-like Open-Domain Chatbot

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.270931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.270931Z digest=sha256:2c62be5990de8d8e32d126fd5ae0d35e587dbb085866a581cca3ab855d6b945b

Observation f200d716-a72a-42de-82ab-c576906a338f · outbound

This paper cites Recipes for building an open-domain chatbot.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Recipes for building an open-domain chatbot

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.275409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.275409Z digest=sha256:71ef54f11b5fa8e84319a449326d44105fcef8803d6652cd93d8dd5822e1e00e

Observation 1e52a746-0e93-483f-9159-863bfba6b574 · outbound

This paper cites CodeBERT: A Pre-Trained Model for Programming and Natural Languages.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage CodeBERT: A Pre-Trained Model for Programming and Natural Languages

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.279420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.279420Z digest=sha256:ae93d0b043d19df01fdb0519d2ba2d16a7a1d51fb2ea43a3bcac7394a6ecbdcb

Observation 9d959efb-140c-4ec9-84e2-620e70a15be0 · outbound

This paper cites IntelliCode Compose: Code Generation Using Transformer.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage IntelliCode Compose: Code Generation Using Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.283941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.283941Z digest=sha256:4819bbafcc056693e85e6ff03e2309f820757924913931cefeebe143f4e49021

Observation 5609e358-bb64-4558-bd09-671d4a44dbd5 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.288668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.288668Z digest=sha256:aab87695c899b1a2a7452f47aa271170c0c589f8cb7d8187e5ae45975ff8a09c

Observation b2190ff8-a254-4a92-bd22-e6e080cf52f3 · outbound

This paper cites Training language models to follow instructions with human feedback.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Training language models to follow instructions with human feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.297218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.297218Z digest=sha256:fbd3757dcf02890a3e14f8b1235a72c8ca02463bffa54415fd130d7adf7b5546

Observation 8a4363b6-5b7c-41bc-9644-f367071331d5 · outbound

This paper cites GPT-4 Technical Report.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage GPT-4 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.301227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.301227Z digest=sha256:3cd4ec62c37ba698a76dab152dc06cd570a8971310e30b780eb897c4f370a0f7

Observation 4cf210e3-3eee-4bd9-b266-266bc5c4b840 · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage A Survey on Efficient Inference for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.305534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.305534Z digest=sha256:5b678ea906551d9ec319d4c81fd289f9f2d3be37b0b95ef056de99abb5ddb043

Observation fc028040-1dc2-42bc-b63a-fc1a8eee89de · outbound

This paper cites Orca: A distributed serving system for transformer-based generative models.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Orca: A distributed serving system for transformer-based generative models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.158878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.310388Z digest=sha256:bba6eba06dd173380167f42e741da66e274e573db760811bb0d238d462c61df1

Observation 1ab96d66-c8f0-46d6-aaf3-ab7c9acf3998 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Efficient memory management for large language model serving with pagedattention

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.148157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.314009Z digest=sha256:508c2f586bdd3dc52e803c6c5b961e51f38e871f2ba49e42753d9c0f98d60d47

Observation fe41a937-246d-4b2e-9c57-3ab601119849 · outbound

This paper cites Fastertransformer: About transformer related optimization, including bert, gpt.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Fastertransformer: About transformer related optimization, including bert, gpt

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.137285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.317680Z digest=sha256:d2f565424bfa0813c22cbe9ea41adb35ae2c029904ffe8f0efc5e435df91df6b

Observation 8153ace8-1a71-4ff4-8aeb-3f593da7c1b1 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.321242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.321242Z digest=sha256:a0e02812832d2841f0735bfdc3ba03d5a0b7b2a963f69b3be799477e95328647

Observation 2f0a985d-a116-40b9-8d50-f3648ef4cc2f · outbound

This paper cites DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.325191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.325191Z digest=sha256:15f47c0e107988fe9a93e254de90e2440710200e723c2fba1cd7678edd65d18f

Observation e8fa9981-97ac-4d03-9c8a-c10d2a3bf1ce · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.328992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.328992Z digest=sha256:6470320d2deda62e4f0bca9dd94085fa582f19d06ed739e165c6d514f2e0a60f

Observation 5aceb0f4-8fe5-42df-aa2d-d5328c90a820 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Splitwise: Efficient generative llm inference using phase splitting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.124854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.332824Z digest=sha256:e134df9115f62be53d5b481e4302e553fa76e7592d73a479f944600fa565f0a9

Observation 53417fe7-6d49-414e-b39d-1ecb77c9aad1 · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.112281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.336304Z digest=sha256:5da2dc8e216584f421ecbf3baecda7bec20370cd8f74340d867c796f46553906

Observation 605ab4f6-e493-4c55-8681-8fe3fc915257 · outbound

This paper cites Exegpt: Constraint-aware resource scheduling for llm inference.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Exegpt: Constraint-aware resource scheduling for llm inference

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.099556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.339718Z digest=sha256:e2172f2e0d66035820a49fce03d04a56411ae162df39959ca4d22b293fc92353

Observation 6d18cfe1-2ed3-468e-a366-805219e82cad · outbound

This paper cites Mooncake: A kvcache-centric disaggregated architecture for llm serving.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Mooncake: A kvcache-centric disaggregated architecture for llm serving

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.085362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.342931Z digest=sha256:6bb7335d80b08a7f55453208124e4d0392eb0f56452ee7e15329185cffd94afe

Observation 2f0d0fd9-ae99-4432-9916-015faae932ce · outbound

This paper cites Nvlink and nvlink switch, 2024.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Nvlink and nvlink switch, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.072104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.346272Z digest=sha256:fa5713dc9570ddc016fc22ad93718fc01a90443e0dec1c47cea6c3f85bad2b39

Observation 19a7e791-1d27-4c19-b003-5b9ce4d540e3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage LLaMA: Open and Efficient Foundation Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.349353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.349353Z digest=sha256:7f7c64f61b5545b04a2a9d1570d8b33c29e6b71bb90e53167a504b3e5d4e3eb1

Observation 61a5b929-9972-443c-a059-0060fc856a34 · outbound

This paper cites Attention is all you need.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Attention is all you need

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.058629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.352750Z digest=sha256:ad09af81fca34b982d1f4335d4fe6578ed8406fb0e131c70dc2bd36cdec1a19f

Observation a8f64235-01cc-41b8-83f2-45ac5d5ac413 · outbound

This paper cites an unresolved cited work.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.356125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.356125Z digest=sha256:bb3de04754c38a1ecaf4c6016daa1db318f219d4f01ea90372477e190e5802d3

Observation 025766aa-b84a-4062-961a-d9611f355298 · outbound

This paper cites Gulavani, Alexey Tumanov, , and Ramachandran Ramjee.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Gulavani, Alexey Tumanov, , and Ramachandran Ramjee

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.037896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.359353Z digest=sha256:23171c8ab0540ced84d3308539c4eaa5966a62423f76515455714e7edf1a0eed

Observation 787b0509-0962-4de6-a495-60f5a6655c8e · outbound

This paper cites Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.025999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.362826Z digest=sha256:1cb21a5af9dbbbb2d56ce4d05d53662313de18324a5e556425855aeba1c11bc5

Observation 74289d5d-3bcd-448a-a81d-db02c871eaec · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage SGLang: Efficient Execution of Structured Language Model Programs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.366590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.366590Z digest=sha256:0c18d1bade8755e9e8151b91ec99a9d05f1442d9c6e341d3ffe06bd9015a5d6b

Observation 82e0b917-d1d3-449f-b5cc-63c673e43468 · outbound

This paper cites Optimizing inference on large language models with nvidia tensorrt-llm, now publicly available.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Optimizing inference on large language models with nvidia tensorrt-llm, now publicly available

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:35.014350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.370833Z digest=sha256:31abe70b3cedf8c28b93a935eb1e39f9e40777d86c1a0eb988944e329658dcb2

Observation 49ce9b54-90b2-4672-9260-5257d1ed35d5 · outbound

This paper cites Lmdeploy is a toolkit for compressing, deploying, and serving llms.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Lmdeploy is a toolkit for compressing, deploying, and serving llms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.891645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.374025Z digest=sha256:a713cfd0a61678a2b71fbdb6bcf6dc8727092b98e2e7f42cc7a0f521d19d2eb9

Observation af0b6ff2-100a-484a-b98e-e1f283babc91 · outbound

This paper cites MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.377394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.377394Z digest=sha256:c0cf9213eb6b0386b443d80ec07fc37d3e0ba2378f2187e4e667c406083b0690

Observation 6c86c13c-95ee-470e-9758-15b75b97e875 · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Fast Distributed Inference Serving for Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.381122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.381122Z digest=sha256:84e29e6a0d97f135d370b148eb86cac08fd97ff7ae6240b29728e3d410906b3c

Observation 3703adc7-a210-4eb9-9430-454b655b1f06 · outbound

This paper cites Gonzalez, and Ion Stoica.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Gonzalez, and Ion Stoica

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.880190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.384669Z digest=sha256:559c1054eb924680fc234450dfbdaee4c67b48a520be643f6c0864b5c96734fd

Observation 12fe33a9-2ade-4f83-ae34-b7a892e9bb2b · outbound

This paper cites Efficient LLM Scheduling by Learning to Rank.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Efficient LLM Scheduling by Learning to Rank

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.388256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.388256Z digest=sha256:3ec9efe255da4b921f96bec17ae8379ffc6abf43a2a1a4560add2df567185619

Observation 9adcf7f8-bb58-409d-857e-a029ef0afc5d · outbound

This paper cites NanoFlow: Towards Optimal Large Language Model Serving Throughput.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage NanoFlow: Towards Optimal Large Language Model Serving Throughput

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.392125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.392125Z digest=sha256:a94a76a953b02b91ee4cdb364ef349fa1f6aa229bf1e3b0cbd231847ca2652aa

Observation fb5fe41f-b203-4c5d-8b3b-4ca9ff7ddfeb · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.868461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.396280Z digest=sha256:b300ae2ba3a24a484a09cfb101b7b3d990b006954466af49c6c233b1b55f2743

Observation 0bbe9b5e-f846-4403-8b22-2c96df9ca148 · outbound

This paper cites Flashdecoding++: Faster large language model inference on gpus.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Flashdecoding++: Faster large language model inference on gpus

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.856980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.400476Z digest=sha256:541ae8759af0da2b950449c0f2466321b421d77073bd83e431f810602012826e

Observation 71241fdc-d701-47f7-b3b0-1a1919dc4262 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.846885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.405100Z digest=sha256:36b5384794148d0f067b41807ddf47b31b72efd712cf2134456dc7d774be1293

Observation 0e18ebe6-3235-46da-9089-1787c765d9d3 · outbound

This paper cites Nvidia mps, 2022.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Nvidia mps, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.832646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.408628Z digest=sha256:70a27a9bdf7c8e0cf19bbd91d0b675aef8cfc4003190f821447568d43f8dba19

Observation 1b218a1e-ec8d-4f71-9fac-6b5116b08159 · outbound

This paper cites Inter-process communication.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Inter-process communication

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.819752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.412068Z digest=sha256:aeb7582664952719438eab38785cca473d0978f75b5d8af25462aa15f57e0025

Observation a1e3dcc0-d14c-4ac3-8596-67cceaa20d7b · outbound

This paper cites Fundamentals of queueing theory, volume.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Fundamentals of queueing theory, volume

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.807656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.415688Z digest=sha256:d5149c5f5db66185b6b40ad48a6b93eed0c257d9ac818b24cc923d01103bac43

Observation 32c97de5-8db4-4ec2-aaf5-ab40b1ca800f · outbound

This paper cites Sharegpt.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Sharegpt

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.784756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.423215Z digest=sha256:c74c62e11cbd15496b3f4063a3084eb9e886cf7af9572a6fd162861671af6d10

Observation 5fff933f-4759-4a93-b4ec-aca5dcb6429d · outbound

This paper cites Nvidia dynamo: A datacenter scale distributed inference serving framework.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Nvidia dynamo: A datacenter scale distributed inference serving framework

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.774210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.426768Z digest=sha256:65e65fd06efdc3ee840319d5cee598c9b32cd2e630c815ea096c9c25fc411db7

Observation 3b079a6a-c177-4b00-a078-dcecb026a0dc · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.430466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.430466Z digest=sha256:d1ab750ee0bba6dc992a666a44a5f0734d3671f54edb5f0f7fbe208fecf7e7c4

Observation 50bb15ee-baaf-4d75-a44c-bc24ac1745a2 · outbound

This paper cites The Llama 3 Herd of Models.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage The Llama 3 Herd of Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.433792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.433792Z digest=sha256:62177dcf440389903b612e9868214e84e035043ca5092aa502a9a34e8b0171d4

Observation ce818738-dcb8-41e4-8fe0-d87c25ae686b · outbound

This paper cites Cuda toolkit.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Cuda toolkit

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.763472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.437398Z digest=sha256:54c221fbbf2c0936f874cb25afebfcb1b1af3e5bf67e34ad43312fc7e040345d

Observation 15cd71d6-93a4-42bc-ab0a-8f4e04ee8476 · outbound

This paper cites Efficient large message broadcast using nccl and cuda-aware mpi for deep learning.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Efficient large message broadcast using nccl and cuda-aware mpi for deep learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.753873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.440731Z digest=sha256:f4a80a39b8d07b7fb34adf246c976abd4573a3cf8429d8f4f3882f61952e50f3

Observation 9b40d1aa-b6cb-4fd6-a00a-c2d648a2434c · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Pytorch: An imperative style, high-performance deep learning library

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.443972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.443972Z digest=sha256:3103dc32125d911f283c00086b6a82e650e09d7cc7353d8430f531344ada18ad

Observation 77bb8015-bfae-4dc5-9d35-b7e2466b5ae5 · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date, April 2024.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Introducing meta llama 3: The most capable openly available llm to date, April 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.737330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.447264Z digest=sha256:df286ec0e0e51d42bd4c967001e4d2ffd3de64137312dd3242acc333e9533ff2

Observation cecaecdf-2c08-498a-a75b-57767c79f289 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.450930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.450930Z digest=sha256:e54bd3fa91f5eed3083b4b22dba0f675276141164e5c530c392f1c6b7e3d20e6

Observation c630acf1-ec20-4f47-8cb6-5a26ad0f6a86 · outbound

This paper cites DeepSeek-V3 Technical Report.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage DeepSeek-V3 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.454577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.454577Z digest=sha256:ad7cda8b7b4157cb0a9285d60c43f662078cbf64e013722bde0062de6c63d222

Observation dea17199-d56d-4780-a656-f998876d9994 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T05:45:34.458370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:45:34.458370Z digest=sha256:e60fe13d872b413167f01f2a35c55785dda2ab26ca1886ecac3edb28f89cd7be

Observation 321e6f93-1aec-497f-9af9-4b890bdf8f8c · outbound

This paper cites Math-500.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Math-500

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:45:34.725530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.461813Z digest=sha256:a10cd3922c1ebedb4903396cb46fd655e11ac1d6d35d416bdb7c4c6963db9933

Observation 59b81587-048a-473c-a81b-0722d232172f · outbound

This paper cites an unresolved cited work.

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage Unresolved cited work

Reference 399

Resolution
parse uncertain
raw_fallback, observed 2026-08-16T05:45:34.795365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:45:34.419634Z digest=sha256:6b838fb19223a10dea962b8365dd96de6f4de61bda4e1d3c44109e7665f8b6f2

Pith citing papers

Observation fd7d91d0-44e7-4401-95ab-fb53c88276eb · inbound

Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving cites this paper.

Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:03:52.902518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:03:52.902518Z digest=sha256:9a145e799e44b5679d0b6d9e06f3b8235d739fb9d25b9391c6f590a3a7ce5eca

Observation 633d21dc-2e09-4afa-a53e-f7dfe3ee0580 · inbound

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing cites this paper.

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T23:39:06.362657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:06.362657Z digest=sha256:db843df442787bb6da971dfda9e7a35162ad9605eb9ada542d33e68612daaecf

Observation e2ba28d7-9b2c-4bf3-9bb3-e753026a5fc8 · inbound

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities cites this paper.

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.262035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T09:45:57.201837Z digest=sha256:993f7c3ff5e11535134ab3684d2144fcec31ca0a39ddacc45a202a40497c335f

Observation e3cf6b4c-9744-43e2-b0f6-bc0407270db5 · inbound

KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving cites this paper.

KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:42:31.180388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T17:42:06.870534Z digest=sha256:4aa310af132f820cd574ec5502f4f4c4bcf9c29a85e4e9c9fc35c37d1bbeb106

Observation 249612ad-faab-49d3-b2e2-9277affae70a · inbound

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location cites this paper.

FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T10:46:52.308744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T04:57:08.746551Z digest=sha256:8d94741adf350bdd2bbd30ab1f1537270e208af2089b31e1e3b4d7a49511be86

Observation 478a67f2-2dab-4a49-9434-6844a04c21d7 · inbound

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models cites this paper.

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:57:38.713663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T13:30:19.620692Z digest=sha256:9295fb6baa6c2ad2419a105683a5062d7dffa8054773fde3a0dd035a4f88e639

Observation 3c53029e-3559-4473-98af-3781038d5753 · inbound

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving cites this paper.

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T05:50:05.675450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:50:05.675450Z digest=sha256:9ba5d4290598a1a5e8b0ec803f33b2036c4df400fd24c4763342f384dae57cdd