Pith. sign in

Paper Citation Record · LEDGER

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving

As of 17 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.06557.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06557 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:37:20.782815Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact3
  • verified fuzzy20
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f7c086e7-da91-4cf2-9e73-416473ca3692 · outbound

This paper cites Qwen2 technical report,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Qwen2 technical report,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.614826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.614826Z digest=sha256:5a3341174a1340325215571e6a05902a558ff89dcdc7644b8208098c14eb3374

Observation 6df7e7c6-661c-4448-829f-664448a2ff39 · outbound

This paper cites Vidur: A large-scale simulation framework for llm inference,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Vidur: A large-scale simulation framework for llm inference,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.996527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.619844Z digest=sha256:d557af1e572bcf1470557b7e3c0a82d643b833218b88e940f4342964bdc509f0

Observation e40999ce-2470-470b-a2f0-1178eb087af6 · outbound

This paper cites Taming throughput-latency tradeoff in llm inference with sarathi-serve,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Taming throughput-latency tradeoff in llm inference with sarathi-serve,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.623715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.623715Z digest=sha256:b22631557f18092863086a006fc2dea59b6aff1ca0e48a49ecd41ba1783a0918

Observation 31e0c590-f737-4ce3-8d93-0fcd5148538c · outbound

This paper cites No request left behind: Tackling heterogeneity in long-context llm inference with medha,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving No request left behind: Tackling heterogeneity in long-context llm inference with medha,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.981307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.627580Z digest=sha256:be11307123d7a4b733f041c0a5bc07d10d152b8c866f7b1e687fb857d869974e

Observation 2c8554d3-2816-43c6-90f4-9db367884daa · outbound

This paper cites Llama 3 model card,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Llama 3 model card,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.969707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.634509Z digest=sha256:c3a0481187d035775537056ed38e4b374f7d53884ddfc6ddb6397f01ac7a464b

Observation d35df27b-d3f1-4666-867e-a1851c6c7424 · outbound

This paper cites GQA: Training generalized multi-query transformer models from multi-head checkpoints,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving GQA: Training generalized multi-query transformer models from multi-head checkpoints,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.638243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.638243Z digest=sha256:256768bc2137bec2b5535713acaa3bab038301e18cb153ce9df42de1f513d554

Observation 09985aab-a5fe-4d0a-b00c-c0fca85c91eb · outbound

This paper cites Qwen-bailian anonymous dataset,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Qwen-bailian anonymous dataset,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.954386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.641506Z digest=sha256:63fc3e387d4ba252265ed11f6459bf8cb1eaae731e808b22e6b00951ba496b24

Observation 0b689cdd-67dd-48e1-8f04-25bec86b029b · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.644989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.644989Z digest=sha256:1d605b506f0b0523c6c81de6b5be263974ef87b0fd3e0dfc02f911f259dfca4f

Observation 10c15315-24f1-4f0c-8d2d-322151282dd5 · outbound

This paper cites SLOs-Serve: Optimized Serving of Multi-SLO LLMs.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving SLOs-Serve: Optimized Serving of Multi-SLO LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.648215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.648215Z digest=sha256:195e61ed94834b00f84428deca324c11aade122e580919b8fe07db1684831b45

Observation b5a122f2-266f-42e2-9cab-25cc238628ab · outbound

This paper cites ATP: Adaptive Tensor Parallelism for Foundation Models.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving ATP: Adaptive Tensor Parallelism for Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.651825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.651825Z digest=sha256:87daee70b6f8249444665057b84001f24b8b40f06161b32c9272fb4cbb93ef08

Observation 293345e1-d02f-42ba-b9db-730dc9fa4821 · outbound

This paper cites Jockey: guaranteed job latency in data parallel clusters,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Jockey: guaranteed job latency in data parallel clusters,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.655496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.655496Z digest=sha256:8c5fb33e8962849d55d3d586951ae7f15728eb80c996ff6ed5f2cb11b8d1337c

Observation c648caa0-8e40-47cf-afd1-8f33e914deb2 · outbound

This paper cites Prompt cache: Modular attention reuse for low-latency inference,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Prompt cache: Modular attention reuse for low-latency inference,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.658651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.658651Z digest=sha256:9f82aecec548902dd96593cb71400a3381ee3470d6f1f611dd65d6c6e8c1865e

Observation a4c25151-548c-47f8-8a7e-89bac272047a · outbound

This paper cites Qoserve: Breaking the silos of llm inference serving,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Qoserve: Breaking the silos of llm inference serving,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.665753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.665753Z digest=sha256:dae36de8c4ded848234ef59ec058b5c2af8c96c4c34e4711a14bea64fc269a91

Observation 84208f07-9b5c-4c1c-9016-7f977104d077 · outbound

This paper cites [Online].

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving [Online]

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.939258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.669126Z digest=sha256:39441c2dfdf35af6414fe8e68e028f147d0254d838be7eaece1f130424f0f850

Observation c7baa3a3-8134-4d35-9bb9-30ab31b8392e · outbound

This paper cites Kvquant: towards 10 million context length llm inference with kv cache quantization,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Kvquant: towards 10 million context length llm inference with kv cache quantization,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.930056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.672231Z digest=sha256:85bf35a512e53ed0119d658bf61ba93fdaad1d6a751ff64c9ae619393c955689

Observation 6c523305-35dd-42e4-a6ee-e89a4984e8a5 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.675426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.675426Z digest=sha256:23c5cb6cdedb4083750a228a1e90025533cb4c0b27dc60e31a2012fdd23c431c

Observation 7c541652-eada-4457-8dda-348ec2c2091c · outbound

This paper cites A quantitative measure of fairness and discrimination for resource allocation in shared computer systems,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving A quantitative measure of fairness and discrimination for resource allocation in shared computer systems,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.921044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.678834Z digest=sha256:bb934586ad71482e9a4f4a3cdcf99d73a63f0189eee22ef69d1ef89041cc3f9f

Observation 666e0fd2-14f4-44d4-9676-d687734bee90 · outbound

This paper cites Sageserve: Optimizing llm serving on cloud data centers with forecast aware auto-scaling,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Sageserve: Optimizing llm serving on cloud data centers with forecast aware auto-scaling,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.686402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.686402Z digest=sha256:777bd6fd755ae92a559fba0122f12543a132752cc5521197347a89c10be97f06

Observation 638bcc91-5ef2-475b-a47f-6f2599c2e29d · outbound

This paper cites Learned Best-Effort LLM Serving.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Learned Best-Effort LLM Serving

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:37:21.557527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.689429Z digest=sha256:ecf7a131b178654855e62ec2c36df1269d5ede868265e7b303a5eee96fbfc321

Observation 8870396d-c2fa-4414-909d-9ff6483b0ca2 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Efficient memory management for large language model serving with pagedattention,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.692788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.692788Z digest=sha256:a4952df0ca1e380ac8bde424ffe0d30fcbeab31b7457eff0aee3156547918f47

Observation c75f11ff-066f-45bc-97cd-339a9e211cd0 · outbound

This paper cites Tokenscale: Timely and accurate autoscaling for disaggregated llm serving with token velocity,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Tokenscale: Timely and accurate autoscaling for disaggregated llm serving with token velocity,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.696067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.696067Z digest=sha256:a9783b0bdbc3c75adf2777ef3f0c0b67d357f4c3cebf33b1468527eb24a944c3

Observation 0075d4b1-e848-48d3-8f0e-b3bc98dd3a63 · outbound

This paper cites Revisiting disaggregated large language model serving for performance and energy implications,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Revisiting disaggregated large language model serving for performance and energy implications,

Reference 22

Resolution
metadata mismatch
raw_fallback, observed 2026-08-15T14:37:21.434037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.699193Z digest=sha256:76d2168b53bc6b6b5a11ffc7b07c3b18bfafcfc5bcf7af1372f2a83d6cdb0c78

Observation 35e1ccb2-db4c-417a-b66f-869f1f0a0f41 · outbound

This paper cites AlpaServe: Statistical multiplexing with model parallelism for deep learning serving,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving AlpaServe: Statistical multiplexing with model parallelism for deep learning serving,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.911478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.702290Z digest=sha256:43bb8e87591760f82ef2d8661832ffe229e4a304b09538c94ab159a88e13471e

Observation 09a2ad14-8c11-4154-931d-55f0e279e7c7 · outbound

This paper cites Scheduling algorithms for multiprogramming in a hard-real-time environment,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Scheduling algorithms for multiprogramming in a hard-real-time environment,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.705254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.705254Z digest=sha256:52fad03b043326a5d896e5a1ad368331d9ba4cb7eb1e681ffbb3b7d3a352bf63

Observation 2c7146dc-68e0-4829-9115-ddf59a1a2310 · outbound

This paper cites Lmcache: An efficient kv cache layer for enterprise-scale llm inference,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Lmcache: An efficient kv cache layer for enterprise-scale llm inference,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.708627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.708627Z digest=sha256:c95339e30c972afb7bf0253f7df2f05f862e21a4989b96860e373e499d424cb0

Observation 9bac5c43-334a-4fb1-9e22-81af19b95c07 · outbound

This paper cites Ai-dynamo,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Ai-dynamo,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.902193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.711621Z digest=sha256:44746cc6fca9b0fb4bf2ac808d01de4d9adb2e02b01d6c0302b410947510bcf2

Observation 8427720a-7675-417a-85e9-870ec2923103 · outbound

This paper cites Nvidia gb200 nvl partition,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Nvidia gb200 nvl partition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.892252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.714811Z digest=sha256:3d78b81a511462ef7a987d70c2165b2d1dd3a415b9641d7bc9b21af0391f014f

Observation 61c4a06c-d981-41cb-8d49-e3032ab53b05 · outbound

This paper cites Nvidia gb200 nvl72 delivers trillion-parameter llm training and real-time inference,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Nvidia gb200 nvl72 delivers trillion-parameter llm training and real-time inference,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.882768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.718007Z digest=sha256:abe545ebc6c56938e0ccbb038ab7e421d35eddede8fdb21275db511d707f14e8

Observation 503f3e8e-a236-4824-8677-3a68d0798b69 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Splitwise: Efficient generative llm inference using phase splitting,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.721320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.721320Z digest=sha256:ac7aa715b6185a67a4a9488ce24139328cf505a70890e5b18e2a85684b35ec92

Observation 536b8c62-c3b9-4e5d-81ac-a0e129befb8b · outbound

This paper cites Conserve: Fine-grained gpu harvesting for llm online and offline co-serving,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Conserve: Fine-grained gpu harvesting for llm online and offline co-serving,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.872872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.724311Z digest=sha256:0f5583b208aff3d622a7164403afe1313832bfb9299c9079504b53cf151c266d

Observation 4541a13c-fa37-471c-9027-06615b4e1c41 · outbound

This paper cites Mooncake: trading more storage for less computation — a kvcache-centric architecture for serving llm chatbot,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Mooncake: trading more storage for less computation — a kvcache-centric architecture for serving llm chatbot,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.862931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.730929Z digest=sha256:bd4c60a37f39c9d563458967dcbdaf777ba95e55dc0b0ee921ded30c3f733545

Observation 39b8a3ff-44f9-43c0-b4e8-e3347737ddb0 · outbound

This paper cites Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:37:21.206726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.733967Z digest=sha256:5597e63042aaff2d570f7767bdeb38efecbf657d2acb38b9eef9d787b2cf6c83

Observation 651a8d8a-658b-4efc-a648-8015e9df3fe2 · outbound

This paper cites Timecard: controlling user-perceived delays in server-based mobile applications,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Timecard: controlling user-perceived delays in server-based mobile applications,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.737451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.737451Z digest=sha256:4866b7765b4942fd93a27677140a869f40dec6cd8457b7ed6b8cf8a636c77e2b

Observation a41a1345-3f21-4cee-b901-5143ae3a48ea · outbound

This paper cites ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.727436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.727436Z digest=sha256:24dd3c632c70842e9e09fe95d0ea5fd8e1f3095833d3dc4bd22eca08c9ec02e1

Observation df6804cb-9c70-46ad-9ab9-a7373a2c6000 · outbound

This paper cites Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:37:21.116238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.743928Z digest=sha256:5f3489f8888aa1c07c1419a8233fac69536296f518f3072072539640cbbb0cbe

Observation b3e2c9db-a907-49e0-80fc-9c80050748cc · outbound

This paper cites Kvcache cache in the wild: Characterizing and optimizing kvcache cache at a large cloud provider,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Kvcache cache in the wild: Characterizing and optimizing kvcache cache at a large cloud provider,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.852825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.747222Z digest=sha256:61a349b66a8f156fea414bf643c4a54430cc9cd41e1b67e9afa89fb6427a5233

Observation 3b7e6a22-325d-42ac-8e81-ba10fa374696 · outbound

This paper cites Better never than late: meeting deadlines in datacenter networks,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Better never than late: meeting deadlines in datacenter networks,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.750430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.750430Z digest=sha256:28f50cc06d06fdfddd0cef558a02abff80d78490a93d9eb753b447c731a4ca1e

Observation 9c690888-fb43-4994-b527-8f98516dfc4c · outbound

This paper cites Preble: Efficient Distributed Prompt Scheduling for LLM Serving.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Preble: Efficient Distributed Prompt Scheduling for LLM Serving

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.740627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.740627Z digest=sha256:1572a6ecc087e9feb9c1ed167a2582aff452d792a1fbbe18d9f97c4758ea7f7f

Observation 6fd0007d-fff8-4e4c-b1d6-3162977ac553 · outbound

This paper cites Aegaeon: Effective gpu pooling for concurrent llm serving on the market,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Aegaeon: Effective gpu pooling for concurrent llm serving on the market,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.757321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.757321Z digest=sha256:d318357f84d980004741352b9089b08a736778db3a7efd3ecbb430bdd6a60296

Observation 613b04c9-77a8-40b0-98c2-98966bceb4b5 · outbound

This paper cites Orca: A distributed serving system for transformer-based generative models,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Orca: A distributed serving system for transformer-based generative models,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.831594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.760766Z digest=sha256:c3b85af2f588695898d40da815341853c235f51f2ce11ac3fc0188eaf3b7fdd4

Observation 8c723fcd-c9af-4d71-9627-8c8a942f1dd0 · outbound

This paper cites Superinfer: Slo-aware rotary scheduling and memory management for llm inference on superchips,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Superinfer: Slo-aware rotary scheduling and memory management for llm inference on superchips,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.821546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.763697Z digest=sha256:42ef7380e4782751399c6066bcda1d65bf0a93772b6f35ed238dbc32b2fa2462

Observation 9c7f6120-152e-4850-ad47-5545febf8105 · outbound

This paper cites FastServe: Iteration-Level preemptive scheduling for large language model inference,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving FastServe: Iteration-Level preemptive scheduling for large language model inference,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.842741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.754259Z digest=sha256:6b7c35287e46fa67cac7ade99138d9e2a906df9994dbfc0723c82269c41d93cd

Observation 3485f82c-ef46-496b-8e9d-8a9e3430313f · outbound

This paper cites Sglang: efficient execution of structured language model programs,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Sglang: efficient execution of structured language model programs,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.801320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.776514Z digest=sha256:fbe6f11276ccda14951d4a94cb04e49c35e256718d799a7565ca6ffbf0d2f657

Observation d1cdf254-95b2-4983-8adf-992d781c83ce · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.791256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.779425Z digest=sha256:e6358b35302f0e594cf7b1f6f9a0f0d7b5ece6b0f82d39d94f7487620f9d19d6

Observation 1d212a09-726c-41d7-bfa2-d0fa4512fd9f · outbound

This paper cites PolyServe: Efficient Multi-SLO Serving at Scale.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving PolyServe: Efficient Multi-SLO Serving at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.782815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.782815Z digest=sha256:803b664161dc42606be6980c63a04aec3f68525e598e6e2f84c6e044aec7a45a

Observation db9f465e-5014-4f7e-b25d-6a60a126d96f · outbound

This paper cites Jitserve: Slo-aware llm serving with imprecise request information,.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Jitserve: Slo-aware llm serving with imprecise request information,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:37:21.811217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T14:37:20.770173Z digest=sha256:199213cc1652b888e8d64a5b04e4e5c3e43253cbfb229126efd12104e8c1a9c7

Observation a4fb85d8-adb6-42b8-aaff-e03ec7b03973 · outbound

This paper cites Available: https://arxiv.org/abs/2504.20068.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Available: https://arxiv.org/abs/2504.20068

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.773331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.773331Z digest=sha256:c0c498d9684ce821219b201109e6cddf614660c5932c8d2fd0cb6b480b894997

Observation 950d1db5-909a-46ba-b81d-2e1bfa2e58ac · outbound

This paper cites A Quantitative Measure Of Fairness And Discrimination For Resource Allocation In Shared Computer Systems.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving A Quantitative Measure Of Fairness And Discrimination For Resource Allocation In Shared Computer Systems

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.682201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.682201Z digest=sha256:6c52b8c69f001a4fbdc4f60c70895d44207a4c6e72a68d0ecd46a19b9313461d

Observation 5490f1a0-917e-4e28-9f39-e4d82775af4e · outbound

This paper cites Prompt Cache: Modular Attention Reuse for Low-Latency Inference.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.662090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.662090Z digest=sha256:3e4ed9cab705a739fce1e2124f0053cd338f08bbb52ca61421f87a7cee2c4ca5

Observation 75f0e0c6-d426-4116-b53f-a4db4bc05a15 · outbound

This paper cites Available: https://arxiv.org/abs/2409.17264.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving Available: https://arxiv.org/abs/2409.17264

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.631206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.631206Z digest=sha256:3621aab7ac171454cde9de2f8d1329d780ee57ab3d734d60f4064bc5ec87f1a5

Observation 84c5990a-d01b-4f37-a842-c87d29bd2cc4 · outbound

This paper cites SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips.

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:20.766968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:20.766968Z digest=sha256:1dae02c0b966207d1171b361abda2915c9491505ca5e7a335e5c1a2d6e2649bb

Pith citing papers

No inbound Pith citation observations are available.