Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T23:16:46.889804Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 6 inbound Pith citation observations for arXiv:2602.14516.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T23:16:46.889804Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T12:56:16.768455Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
29 of 29 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 555b6bfa-4922-4902-9880-e8402f78d693 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8783764-882b-4d78-b67e-3005a4749d0a · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Hydrainfer: Hybrid disaggregated scheduling for multimodal large language model serving.arXiv preprint arXiv:2505.12658,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17d9d2fd-8810-4899-9066-72b2c9535641 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f49d73c5-7804-4ea5-9e81-6b070a96660b · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d08313a1-b71c-4a19-8226-7c359e030f00 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4f233bb-a4d3-4a61-8f4a-cb42ccd4d76c · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving DeepSeek-V3 Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3286a50a-54a8-4563-87bc-10b768a09779 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Augmented Language Models: a Survey
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9beb1252-5373-4c29-a425-ef6ab4d18e37 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Gaia: a benchmark for general ai assistants
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b11e7e9-3526-48e5-a686-5cf746fc7b81 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Nvidia dynamo documentation: Kv router
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d67bad8-cc56-4a4c-bbf2-d3352db4b33e · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Splitwise: Efficient generative llm inference using phase splitting
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35fcc53c-6442-42f1-a182-be498968b2f2 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Fast inference for augmented large lan- guage models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e335f2c-d89c-4958-8848-4d519209050a · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48212ac5-9595-48bf-8a43-e0470ab94bfe · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7922afca-e345-4fbe-9aa3-90b613a355d7 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving N., Kaiser, Ł., and Polosukhin, I
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfc5ffd0-586f-437b-aba6-1d3e925c3830 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Qwen3 Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3014c50a-7f7f-40ba-80ae-9ce98c722302 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12a3dd87-6f0e-4bf7-b355-460bd5f095a2 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving R., and Cao, Y
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cef83b54-dfa1-4400-9015-6fffa96b8af5 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a758c01-6db5-4171-942d-213606ed8ca3 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving More Details about Offline Planning A.1
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f34115de-a571-441f-a92f-d07ae8ed1b5d · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving More experimental results
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58273f0e-8fda-42c8-9cf4-085507b6335d · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving H., Gonzalez, J., Zhang, H., and Stoica, I
Reference 1991
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d20fdeaf-85f2-4752-af64-7f07217c58e6 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Reference 1994
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d157a2be-c120-415b-95a0-a02d4a444cc1 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c936e470-ab9d-4087-be4b-45c65b411738 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Mixtral of Experts
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b891953-6f2e-4a5a-a8da-ad5283f6a874 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Unresolved cited work
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 019801e1-5b70-41dd-87ac-34b7899ae808 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Inference Scaling for Long-Context Retrieval Augmented Generation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb4cc8f0-0779-45b1-af60-9f0a249f39a3 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving Tokenscale: Timely and accurate autoscaling for disaggregated llm serving with token velocity.arXiv preprint arXiv:2512.03416,
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a04a370-a51b-4fbc-8faf-556e64023072 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving The SCIP Optimization Suite 9.0
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5a6dc06-2d54-42b9-9b24-8a4969c84d56 · outbound
Efficient Multi-round LLM Inference over Disaggregated Serving The Llama 3 Herd of Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff8791b6-7290-4405-9c08-178de7b687af · inbound
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics Efficient Multi-round LLM Inference over Disaggregated Serving
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8dc64203-2e53-4cf6-8e55-7c48ed101993 · inbound
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving Efficient Multi-round LLM Inference over Disaggregated Serving
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ef758152-80f3-4f82-9375-0cd1f82e4ba4 · inbound
DelAC: A Multi-agent Reinforcement Learning of Team-Symmetric Stochastic Games Efficient Multi-round LLM Inference over Disaggregated Serving
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1511e2c0-23f1-439f-a641-fe8537a205ea · inbound
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling Efficient Multi-round LLM Inference over Disaggregated Serving
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8fe073dd-c74d-4740-bbf3-975126a37bb9 · inbound
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving Efficient Multi-round LLM Inference over Disaggregated Serving
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a7b292d0-99e6-45bb-8ab1-78b1769ba6d2 · inbound
TurboServe: Serving Streaming Video Generation Efficiently and Economically Efficient Multi-round LLM Inference over Disaggregated Serving
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.