Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:45:41.645133Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 2 inbound Pith citation observations for arXiv:2508.19559.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:45:41.645133Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T19:23:10.356881Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T02:39:25.590182Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5296e3ec-05ea-4bcd-aa64-18062ea8bf1e · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Accessed 2025-7-24
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e9a84438-7d06-4387-9d3f-3117c7e02ccb · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference https:// developer.nvidia.com/docs,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b1aa73c1-0094-43ae-acc1-44eaa0ef514c · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2be6c643-712a-4b5b-b3c2-5aaacc822daa · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b78d49da-ca7c-4f1d-93d6-27b0c2b84725 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 005f0de8-1e5a-4f39-9539-e06fb5d312ae · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Gpu utilization is a misleading metric, 2025
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 37415a6e-20ca-453a-88e9-92b47eab4bc4 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference KVDirect: Distributed Disaggregated LLM Inference
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 917bc262-3221-4a86-ba8c-68677aa20ce7 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Leveraging endpoint flexibility in data- intensive clusters
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 434850f0-8903-417f-b7a6-ffd26b851236 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Efficient coflow scheduling with varys
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 409d0484-5e32-410c-a5e0-104957c72a60 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Resource central: Understanding and predict- ing workloads for improved resource management in large cloud platforms
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e875fca2-10f5-4384-852b-1a3a117d8998 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference A Complete Survey on LLM-based AI Chatbots
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 475ceb66-0e2a-402e-9b23-87108b8374f7 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Paragon: Qos-aware scheduling for heterogeneous datacenters
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6a2343fd-7f22-4243-9541-5ab22adf0c5c · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Deploy DeepSeek-V3/R1 671b on 8 × H100 and throughput bench- marks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5bbc3669-7426-4e42-b71e-95fbe657eed6 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Autoscale: Dynamic, robust capacity management for multi-tier data centers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation be13370d-eb17-418d-ba53-bd1ecfdce013 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Firmament: Fast, centralized cluster scheduling at scale
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6a74bdda-0d75-4dff-865e-b99caede7018 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aeedfa4-c07d-4dfe-8f40-33f896f9b853 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Tiresias: A {GPU} cluster manager for distributed deep learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 39312223-c144-4967-ba9d-d17370826359 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b62d94d-a129-4d69-9b26-5a56356e42ca · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f33f8277-8b6a-4991-b6f4-b43908aa318b · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45d01c18-09d5-4c7f-ba03-31439aed0250 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Quincy: fair scheduling for distributed computing clusters
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4440eb5b-1899-4b27-825d-5f28883f0231 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Amant, Chetan Bansal, Victor Rühle, Anoop Kulkarni, Steve Kofsky, and Saravan Rajmohan
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8e7c6e4-e32f-40d6-b289-aa0fb5da0410 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59ce4631-00ea-4533-a7ba-7c080344900d · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Netcache: Balancing key-value stores with fast in-network caching
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 617bfea6-471a-425a-bf6a-44dcb80fa309 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference P/d- serve: Serving disaggregated large language model at scale, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 21b1a28a-61fe-4592-9c44-72ab951f3b18 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Morpheus: Towards automated {SLOs} for en- terprise clusters
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d890f1f7-c6ea-4e78-9a31-aa1b0b2ff789 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Pod-attention: Unlocking full prefill-decode overlap for faster llm inference
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cbeeca16-12c9-4573-a9b3-c7050b81c92a · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Keda: Kubernetes event-driven autoscaling
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 322f0d49-79c5-449f-a8c1-166f093df45f · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Kubernetes horizontal pod au- toscaler
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation efaff735-3754-4881-bab5-24bbd0938ddb · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Kubernetes vertical pod autoscaler
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 61e57b32-a9c4-4e53-b13d-a709948684c1 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Kubernetes: Production-grade container orchestration
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c4311eac-23f7-4449-b276-c31f83235da3 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34e5fb31-0986-4b91-8ebf-86c0b4eda6a6 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 24b6c890-27de-4af7-8f90-d3dffe6ba77a · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Themis: Fair and efficient {GPU} cluster scheduling
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d9a25650-0111-4be7-94b2-9e8419d86b50 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Alpaserve: Statistical multiplexing with model parallelism for deep learning serving
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 008c09e5-d21c-48c1-b5c1-7e835ca18425 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Mastering llm techniques: Inference optimization, 2023
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 946a2472-09c5-4e0b-8c07-93415e59af99 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference {Heterogeneity-Aware} cluster scheduling policies for deep learning workloads
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e26c1d2a-85b5-4204-88f6-abbf66139cad · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Introducing chatgpt
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7dd67305-7e24-437b-b0e2-24a865b5d666 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Tensorrt-llm: A deep learning compiler for large language models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 06e8cfb6-5051-45c5-8a5b-5b8ad9c2bdcf · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Splitwise: Efficient generative LLM inference using phase splitting
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10af1f55-f75b-4194-9f3e-e5ebe6d56cb1 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Splitwise: Efficient generative llm inference using phase splitting
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation add94ac0-d2e8-450f-beb4-56798e2c52e7 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92be0001-d2d4-4b27-b8bb-b84232124710 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9321c0-4d56-464b-bfc7-346b861a505f · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cde3a46f-15b0-4b33-8203-5c39b29d5f5d · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Dynamollm: Designing llm inference clusters for performance and energy ef- ficiency, 2024
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 63815573-bf62-4253-a78e-11a2f9f2bee6 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Burstgpt: A real-world workload dataset to optimize llm serving systems, 2025
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fd14e2a9-b054-4988-883c-e1bdcdeae13d · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Attention is all you need
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e5c97bd-0af5-460d-b280-c1736cfdbb3d · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Fast Distributed Inference Serving for Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b5f8af-d27f-4efc-a55b-3e0e49cba8af · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Deepscaling: mi- croservices autoscaling for stable cpu utilization in large scale cloud systems
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 628fb8ca-daed-443a-9c09-98c934c15340 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Gandiva: Introspec- tive cluster scheduling for deep learning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8029a4fa-940d-4270-8a33-34a1f2fbeb3c · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Skylb: A locality-aware cross-region load balancer for llm inference
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ca7c4f-3319-459b-b0ad-2f5dae35a1e0 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference When Search Engine Services meet Large Language Models: Visions and Challenges
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a26dc84-0566-465a-b261-d159407ac033 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference {AntMan}: Dynamic scaling on {GPU} clusters for deep learning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9c1753c7-4196-4eef-8c5d-9605fafc83cf · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Orca: A distributed serving system for transformer-based generative models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e003e517-1b64-4008-83bc-c1ef61bfa38c · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Qwen3 Technical Report
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2226b0d-482c-40e2-9ff5-95ac47053a60 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Sglang: Efficient execution of structured language model programs
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 52c49da3-783b-4771-a9c6-28672a21f558 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Day zero benchmarks for qwen 3 with sglang on baseten
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e999b7ab-399e-41ac-b2a2-49922ad45b24 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05f34593-b95c-4dbd-bc11-7220b48a6446 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7d6a797b-c5be-45c7-936a-a5d037fa6939 · outbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Unresolved cited work
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cf870d12-d482-4314-8454-4d57614fcc5a · inbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 63695de3-9746-4158-9525-4fb243846207 · inbound
TurboServe: Serving Streaming Video Generation Efficiently and Economically Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.