Pith. sign in

Paper Citation Record · LEDGER

Efficient Multi-round LLM Inference over Disaggregated Serving

As of 14 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 6 inbound Pith citation observations for arXiv:2602.14516.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.14516 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T23:16:46.889804Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T12:56:16.768455Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 555b6bfa-4922-4902-9880-e8402f78d693 · outbound

This paper cites an unresolved cited work.

Efficient Multi-round LLM Inference over Disaggregated Serving Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.025728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.025728Z digest=sha256:c5204d005c2772a924dbb196dfbdb15a65aa31762db1974192256895814b8d11

Observation d8783764-882b-4d78-b67e-3005a4749d0a · outbound

This paper cites Hydrainfer: Hybrid disaggregated scheduling for multimodal large language model serving.arXiv preprint arXiv:2505.12658,.

Efficient Multi-round LLM Inference over Disaggregated Serving Hydrainfer: Hybrid disaggregated scheduling for multimodal large language model serving.arXiv preprint arXiv:2505.12658,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.082617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.082617Z digest=sha256:3dc6afa57f8a9f2031e237b15d82f0275f35e7dae9a4599be42293eac69d91a3

Observation 17d9d2fd-8810-4899-9066-72b2c9535641 · outbound

This paper cites StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models.

Efficient Multi-round LLM Inference over Disaggregated Serving StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.239465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.239465Z digest=sha256:d616ff56242e7d87f3c34063d05568eafab65c0ec6f29dbd2fc5a72bd6ccf9a8

Observation f49d73c5-7804-4ea5-9e81-6b070a96660b · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Efficient Multi-round LLM Inference over Disaggregated Serving Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.511270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.511270Z digest=sha256:343608540528dfd2be2c6e12cb2dfbb54f895000dde068f41cfd5055289747f1

Observation d08313a1-b71c-4a19-8226-7c359e030f00 · outbound

This paper cites Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live.

Efficient Multi-round LLM Inference over Disaggregated Serving Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.857124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.857124Z digest=sha256:2f7128ac42769c82b6f95fb06a960c99f4cd88ba31b1818f6acfdf4543be37bb

Observation a4f233bb-a4d3-4a61-8f4a-cb42ccd4d76c · outbound

This paper cites DeepSeek-V3 Technical Report.

Efficient Multi-round LLM Inference over Disaggregated Serving DeepSeek-V3 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.971359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.971359Z digest=sha256:1d551ce6767d4924c1c72287af3360677d92c066ac0d459a896578535f03575e

Observation 3286a50a-54a8-4563-87bc-10b768a09779 · outbound

This paper cites Augmented Language Models: a Survey.

Efficient Multi-round LLM Inference over Disaggregated Serving Augmented Language Models: a Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:45.064125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:45.064125Z digest=sha256:8f4de337f6ba25e5dc8cbd701317e415076bbc48ab2a098318dc6d8ebb19438e

Observation 9beb1252-5373-4c29-a425-ef6ab4d18e37 · outbound

This paper cites Gaia: a benchmark for general ai assistants.

Efficient Multi-round LLM Inference over Disaggregated Serving Gaia: a benchmark for general ai assistants

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:45.124383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:45.124383Z digest=sha256:2829a064d5db7f51e6365738eb575dca772b0dce900c5c8a3cdea47e0fa2aff4

Observation 3b11e7e9-3526-48e5-a686-5cf746fc7b81 · outbound

This paper cites Nvidia dynamo documentation: Kv router.

Efficient Multi-round LLM Inference over Disaggregated Serving Nvidia dynamo documentation: Kv router

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:45.256552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:45.256552Z digest=sha256:711d30a9e3da66279ca3041e2f2580bf4aff25b3a06cde756f651c385927fc81

Observation 8d67bad8-cc56-4a4c-bbf2-d3352db4b33e · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting.

Efficient Multi-round LLM Inference over Disaggregated Serving Splitwise: Efficient generative llm inference using phase splitting

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:45.361663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:45.361663Z digest=sha256:64044b9fd094e060ebe1e7777fae387b89d2a4d61497861d114258dbdef81019

Observation 35fcc53c-6442-42f1-a182-be498968b2f2 · outbound

This paper cites Fast inference for augmented large lan- guage models.

Efficient Multi-round LLM Inference over Disaggregated Serving Fast inference for augmented large lan- guage models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:45.486415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:45.486415Z digest=sha256:c3a67febf3aa6c0b7e7d984815e5c837a8e148bc6d1eb2f6002a2edee7f91a74

Observation 2e335f2c-d89c-4958-8848-4d519209050a · outbound

This paper cites Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy.

Efficient Multi-round LLM Inference over Disaggregated Serving Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:45.704349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:45.704349Z digest=sha256:45582d0964bc2ca89440fd0bf5b827057c0df3d8015734b3906f14de44300cea

Observation 48212ac5-9595-48bf-8a43-e0470ab94bfe · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Efficient Multi-round LLM Inference over Disaggregated Serving Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:45.823877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:45.823877Z digest=sha256:c217778990f65bea54ffffcc7aa61a22fc9cc2bf21d64fa16632be9079518a41

Observation 7922afca-e345-4fbe-9aa3-90b613a355d7 · outbound

This paper cites N., Kaiser, Ł., and Polosukhin, I.

Efficient Multi-round LLM Inference over Disaggregated Serving N., Kaiser, Ł., and Polosukhin, I

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:46.045201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:46.045201Z digest=sha256:74cb6ac8fe5f9277b8d8e341c97f5ae39acaacfb46979a7933ef8761dde5aafc

Observation cfc5ffd0-586f-437b-aba6-1d3e925c3830 · outbound

This paper cites Qwen3 Technical Report.

Efficient Multi-round LLM Inference over Disaggregated Serving Qwen3 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:46.327419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:46.327419Z digest=sha256:9e2601a136a90a4c1d6a110f72dfe16c0c689a33476775777dd44674a6f9345d

Observation 3014c50a-7f7f-40ba-80ae-9ce98c722302 · outbound

This paper cites an unresolved cited work.

Efficient Multi-round LLM Inference over Disaggregated Serving Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:46.447356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:46.447356Z digest=sha256:9503b204f130711ac7a338cdcc2ce33314836a84ae105e1037a4723808fd8906

Observation 12a3dd87-6f0e-4bf7-b355-460bd5f095a2 · outbound

This paper cites R., and Cao, Y.

Efficient Multi-round LLM Inference over Disaggregated Serving R., and Cao, Y

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:46.561275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:46.561275Z digest=sha256:d63268b32641738e3968d28cd78b065e976fb2b968b96876de66976e2bfcb69d

Observation cef83b54-dfa1-4400-9015-6fffa96b8af5 · outbound

This paper cites LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind.

Efficient Multi-round LLM Inference over Disaggregated Serving LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:46.765490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:46.765490Z digest=sha256:cd5fe2ca997119fef5e0c0d5c7f444ebc5fd3b2cb36dfce89ee3acf3cc4e477d

Observation 4a758c01-6db5-4171-942d-213606ed8ca3 · outbound

This paper cites More Details about Offline Planning A.1.

Efficient Multi-round LLM Inference over Disaggregated Serving More Details about Offline Planning A.1

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:46.842880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:46.842880Z digest=sha256:f04856c53be939d93f9e2881bc0ec33a4293d0c4358a85401b8959c82cc53d3f

Observation f34115de-a571-441f-a92f-d07ae8ed1b5d · outbound

This paper cites More experimental results.

Efficient Multi-round LLM Inference over Disaggregated Serving More experimental results

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:46.889804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:46.889804Z digest=sha256:115076c9886d480e60774a262d26715920def49348feae4d92345759ae34ce9d

Observation 58273f0e-8fda-42c8-9cf4-085507b6335d · outbound

This paper cites H., Gonzalez, J., Zhang, H., and Stoica, I.

Efficient Multi-round LLM Inference over Disaggregated Serving H., Gonzalez, J., Zhang, H., and Stoica, I

Reference 1991

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.631418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.631418Z digest=sha256:a657929f8772264d9faa88f3b2403175ac27f8e0556a15ee01f06c73456fb3ad

Observation d20fdeaf-85f2-4752-af64-7f07217c58e6 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

Efficient Multi-round LLM Inference over Disaggregated Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.283021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.283021Z digest=sha256:7db0ddd87667a52ffb8c13acbb4c9043c041f1ca9dee9ea7aa893fdf6f0ee035

Observation d157a2be-c120-415b-95a0-a02d4a444cc1 · outbound

This paper cites AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving.

Efficient Multi-round LLM Inference over Disaggregated Serving AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:46.161659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:46.161659Z digest=sha256:5eec24bba588952fc3611394ce4c6cd2cc285a7b896c5a3ecbde942effa0b544

Observation c936e470-ab9d-4087-be4b-45c65b411738 · outbound

This paper cites Mixtral of Experts.

Efficient Multi-round LLM Inference over Disaggregated Serving Mixtral of Experts

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.393382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.393382Z digest=sha256:470d0a866d78d63050585782c0727e2b159fb5ab4a106d5e2c95947a6ca1624a

Observation 3b891953-6f2e-4a5a-a8da-ad5283f6a874 · outbound

This paper cites an unresolved cited work.

Efficient Multi-round LLM Inference over Disaggregated Serving Unresolved cited work

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:45.952855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:45.952855Z digest=sha256:e5e66f1b7931e74eaea743332e08688f12145d4191be410954fc6718af21cb2c

Observation 019801e1-5b70-41dd-87ac-34b7899ae808 · outbound

This paper cites Inference Scaling for Long-Context Retrieval Augmented Generation.

Efficient Multi-round LLM Inference over Disaggregated Serving Inference Scaling for Long-Context Retrieval Augmented Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:46.676879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:46.676879Z digest=sha256:c2ce21724e04dfba63447fd43709e958c738166ff2a7ba262e437c2f16b60162

Observation fb4cc8f0-0779-45b1-af60-9f0a249f39a3 · outbound

This paper cites Tokenscale: Timely and accurate autoscaling for disaggregated llm serving with token velocity.arXiv preprint arXiv:2512.03416,.

Efficient Multi-round LLM Inference over Disaggregated Serving Tokenscale: Timely and accurate autoscaling for disaggregated llm serving with token velocity.arXiv preprint arXiv:2512.03416,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.747173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.747173Z digest=sha256:c03ec8161ed4ec7aded26cb7ab35e7b97d182338a8af96458f7fa15b31a1b701

Observation 3a04a370-a51b-4fbc-8faf-556e64023072 · outbound

This paper cites The SCIP Optimization Suite 9.0.

Efficient Multi-round LLM Inference over Disaggregated Serving The SCIP Optimization Suite 9.0

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:43.933864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:43.933864Z digest=sha256:4784d45b8b3ea0ccbdf717a90e45b4b7c199d3be4a31009d5627dde56856f8b1

Observation b5a6dc06-2d54-42b9-9b24-8a4969c84d56 · outbound

This paper cites The Llama 3 Herd of Models.

Efficient Multi-round LLM Inference over Disaggregated Serving The Llama 3 Herd of Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.157699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.157699Z digest=sha256:1c99a880e5d17ba23b94fe1269a43ddc7fb26adf3741eceed04c99b58568ced9

Pith citing papers

Observation ff8791b6-7290-4405-9c08-178de7b687af · inbound

Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics cites this paper.

Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics Efficient Multi-round LLM Inference over Disaggregated Serving

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-22T03:22:04.971065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T17:30:40.021376Z digest=sha256:2100548ae55ee42aa7c8b353ab193d185e368a647a263d8482a504c9f4048bd2

Observation 8dc64203-2e53-4cf6-8e55-7c48ed101993 · inbound

KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving cites this paper.

KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving Efficient Multi-round LLM Inference over Disaggregated Serving

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-22T03:22:04.971065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T06:58:36.525442Z digest=sha256:a54420c9e4135cebb7dbafb104f09e79dea514471a901dddfae28afa84f291ce

Observation ef758152-80f3-4f82-9375-0cd1f82e4ba4 · inbound

DelAC: A Multi-agent Reinforcement Learning of Team-Symmetric Stochastic Games cites this paper.

DelAC: A Multi-agent Reinforcement Learning of Team-Symmetric Stochastic Games Efficient Multi-round LLM Inference over Disaggregated Serving

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-22T03:22:04.971065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T21:31:06.105461Z digest=sha256:659fc537bbc752793c3c2b46230d49fac2c2b2dff7e928db051d286d653f85b8

Observation 1511e2c0-23f1-439f-a641-fe8537a205ea · inbound

HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling cites this paper.

HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling Efficient Multi-round LLM Inference over Disaggregated Serving

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-22T03:22:04.971065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T20:58:51.630574Z digest=sha256:1641c7e777ec85df1cb0741f282c58dd17cedef63d611a35138736efd99a1b8c

Observation 8fe073dd-c74d-4740-bbf3-975126a37bb9 · inbound

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving cites this paper.

Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving Efficient Multi-round LLM Inference over Disaggregated Serving

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-22T03:22:04.971065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T12:56:16.768455Z digest=sha256:d3abd475f5519b1e17d6b1a29e1cd22da7fe7875a021c6e576a5fce4c9d3ad45

Observation a7b292d0-99e6-45bb-8ab1-78b1769ba6d2 · inbound

TurboServe: Serving Streaming Video Generation Efficiently and Economically cites this paper.

TurboServe: Serving Streaming Video Generation Efficiently and Economically Efficient Multi-round LLM Inference over Disaggregated Serving

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-22T03:22:04.971065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T19:23:10.356881Z digest=sha256:a2c597a2bf4691ae12f5da7e4f99834052aa8c3d961802b2cd6a14710d055e2f