Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:37:20.782815Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.06557.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:37:20.782815Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f7c086e7-da91-4cf2-9e73-416473ca3692 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Qwen2 technical report,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6df7e7c6-661c-4448-829f-664448a2ff39 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Vidur: A large-scale simulation framework for llm inference,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e40999ce-2470-470b-a2f0-1178eb087af6 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Taming throughput-latency tradeoff in llm inference with sarathi-serve,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e0c590-f737-4ce3-8d93-0fcd5148538c · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving No request left behind: Tackling heterogeneity in long-context llm inference with medha,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2c8554d3-2816-43c6-90f4-9db367884daa · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Llama 3 model card,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d35df27b-d3f1-4666-867e-a1851c6c7424 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving GQA: Training generalized multi-query transformer models from multi-head checkpoints,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09985aab-a5fe-4d0a-b00c-c0fca85c91eb · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Qwen-bailian anonymous dataset,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0b689cdd-67dd-48e1-8f04-25bec86b029b · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Evaluating Large Language Models Trained on Code
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c15315-24f1-4f0c-8d2d-322151282dd5 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving SLOs-Serve: Optimized Serving of Multi-SLO LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5a122f2-266f-42e2-9cab-25cc238628ab · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving ATP: Adaptive Tensor Parallelism for Foundation Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 293345e1-d02f-42ba-b9db-730dc9fa4821 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Jockey: guaranteed job latency in data parallel clusters,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c648caa0-8e40-47cf-afd1-8f33e914deb2 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Prompt cache: Modular attention reuse for low-latency inference,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4c25151-548c-47f8-8a7e-89bac272047a · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Qoserve: Breaking the silos of llm inference serving,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84208f07-9b5c-4c1c-9016-7f977104d077 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving [Online]
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c7baa3a3-8134-4d35-9bb9-30ab31b8392e · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Kvquant: towards 10 million context length llm inference with kv cache quantization,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6c523305-35dd-42e4-a6ee-e89a4984e8a5 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c541652-eada-4457-8dda-348ec2c2091c · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving A quantitative measure of fairness and discrimination for resource allocation in shared computer systems,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 666e0fd2-14f4-44d4-9676-d687734bee90 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Sageserve: Optimizing llm serving on cloud data centers with forecast aware auto-scaling,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 638bcc91-5ef2-475b-a47f-6f2599c2e29d · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Learned Best-Effort LLM Serving
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8870396d-c2fa-4414-909d-9ff6483b0ca2 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Efficient memory management for large language model serving with pagedattention,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c75f11ff-066f-45bc-97cd-339a9e211cd0 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Tokenscale: Timely and accurate autoscaling for disaggregated llm serving with token velocity,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0075d4b1-e848-48d3-8f0e-b3bc98dd3a63 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Revisiting disaggregated large language model serving for performance and energy implications,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 35e1ccb2-db4c-417a-b66f-869f1f0a0f41 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving AlpaServe: Statistical multiplexing with model parallelism for deep learning serving,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 09a2ad14-8c11-4154-931d-55f0e279e7c7 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Scheduling algorithms for multiprogramming in a hard-real-time environment,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c7146dc-68e0-4829-9115-ddf59a1a2310 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Lmcache: An efficient kv cache layer for enterprise-scale llm inference,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bac5c43-334a-4fb1-9e22-81af19b95c07 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Ai-dynamo,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8427720a-7675-417a-85e9-870ec2923103 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Nvidia gb200 nvl partition,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 61c4a06c-d981-41cb-8d49-e3032ab53b05 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Nvidia gb200 nvl72 delivers trillion-parameter llm training and real-time inference,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 503f3e8e-a236-4824-8677-3a68d0798b69 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Splitwise: Efficient generative llm inference using phase splitting,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 536b8c62-c3b9-4e5d-81ac-a0e129befb8b · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Conserve: Fine-grained gpu harvesting for llm online and offline co-serving,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4541a13c-fa37-471c-9027-06615b4e1c41 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Mooncake: trading more storage for less computation — a kvcache-centric architecture for serving llm chatbot,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 39b8a3ff-44f9-43c0-b4e8-e3347737ddb0 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 651a8d8a-658b-4efc-a648-8015e9df3fe2 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Timecard: controlling user-perceived delays in server-based mobile applications,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a41a1345-3f21-4cee-b901-5143ae3a48ea · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df6804cb-9c70-46ad-9ab9-a7373a2c6000 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b3e2c9db-a907-49e0-80fc-9c80050748cc · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Kvcache cache in the wild: Characterizing and optimizing kvcache cache at a large cloud provider,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3b7e6a22-325d-42ac-8e81-ba10fa374696 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Better never than late: meeting deadlines in datacenter networks,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c690888-fb43-4994-b527-8f98516dfc4c · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Preble: Efficient Distributed Prompt Scheduling for LLM Serving
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fd0007d-fff8-4e4c-b1d6-3162977ac553 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Aegaeon: Effective gpu pooling for concurrent llm serving on the market,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 613b04c9-77a8-40b0-98c2-98966bceb4b5 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Orca: A distributed serving system for transformer-based generative models,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8c723fcd-c9af-4d71-9627-8c8a942f1dd0 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Superinfer: Slo-aware rotary scheduling and memory management for llm inference on superchips,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9c7f6120-152e-4850-ad47-5545febf8105 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving FastServe: Iteration-Level preemptive scheduling for large language model inference,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3485f82c-ef46-496b-8e9d-8a9e3430313f · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Sglang: efficient execution of structured language model programs,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d1cdf254-95b2-4983-8adf-992d781c83ce · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1d212a09-726c-41d7-bfa2-d0fa4512fd9f · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving PolyServe: Efficient Multi-SLO Serving at Scale
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db9f465e-5014-4f7e-b25d-6a60a126d96f · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Jitserve: Slo-aware llm serving with imprecise request information,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a4fb85d8-adb6-42b8-aaff-e03ec7b03973 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Available: https://arxiv.org/abs/2504.20068
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 950d1db5-909a-46ba-b81d-2e1bfa2e58ac · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving A Quantitative Measure Of Fairness And Discrimination For Resource Allocation In Shared Computer Systems
Reference 1998
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5490f1a0-917e-4e28-9f39-e4d82775af4e · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75f0e0c6-d426-4116-b53f-a4db4bc05a15 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Available: https://arxiv.org/abs/2409.17264
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84c5990a-d01b-4f37-a842-c87d29bd2cc4 · outbound
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.