Pith. sign in

Paper Citation Record · LEDGER

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees

As of 17 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2506.19677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19677 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:34:39.461915Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7b85b78-d2bc-4062-9493-ee49a6281da0 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Code Llama: Open Foundation Models for Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.135939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.135939Z digest=sha256:9f2586089758d9f88f521f0c5129c2380466a86b69c7c385d7f5d1440a2a79df

Observation 5ae8e1bb-65c9-45b7-9679-9d03881705af · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.142344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.142344Z digest=sha256:86decd96c8c951bf6d5d3c24b9dbf189f0e1993fc3d1a713d1e5e9bedcd52649

Observation fa6e4015-90ca-4285-8d57-536cdf6b889d · outbound

This paper cites Qwen2.5-Coder Technical Report.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Qwen2.5-Coder Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.148481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.148481Z digest=sha256:e6799114a6574c3a4cc06ce5baa36d9d20b356f9817212fb2791f07dd301605c

Observation 2487f045-311f-4609-b73a-42cec3188f6a · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees StarCoder 2 and The Stack v2: The Next Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.154863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.154863Z digest=sha256:588b4b9e3bb14b306f786822dc610bf5f19daeb15498815664e30636de864ffd

Observation 0419d7b9-9f88-44a2-bc25-02f8c6df7a6b · outbound

This paper cites Fine tuning large language model for secure code generation,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Fine tuning large language model for secure code generation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.836208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.160386Z digest=sha256:5547dc6c93e771337787cf1c5fb54112524e5a15a4db6268dc9d6635e8b7803a

Observation 975034c9-d087-43a5-9f57-3da2366bb72c · outbound

This paper cites RepoHyper: Search-Expand-Refine on Semantic Graphs for Repository-Level Code Completion.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees RepoHyper: Search-Expand-Refine on Semantic Graphs for Repository-Level Code Completion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.164978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.164978Z digest=sha256:129793e1bead68f0f1a3e650c59839f4ee64c9cf87383d9419024fbec20340d4

Observation d2e1e620-1d0c-4cde-94ba-99ea4809c32d · outbound

This paper cites A Survey on Large Language Models for Code Generation.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees A Survey on Large Language Models for Code Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.171573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.171573Z digest=sha256:6518480788aaaa63e6860093de86f13a790a6ee5940b1c0c3751f8457ed45401

Observation 81fda2bd-3062-49c6-b4bd-e9f3b4f75e8f · outbound

This paper cites Language models for code completion: A practical evaluation,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Language models for code completion: A practical evaluation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.808922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.177142Z digest=sha256:bc081320a790ce75cd53f1b933e4996817e62a6f22374f35c29afb31454c63b7

Observation 984a41e3-9824-4d89-9e67-147a8cc2a164 · outbound

This paper cites AI-assisted Code Authoring at Scale: Fine-tuning, deploying, and mixed methods evaluation.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees AI-assisted Code Authoring at Scale: Fine-tuning, deploying, and mixed methods evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.182174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.182174Z digest=sha256:9d79b5c24a7c8229bb6f8d3df67c17741b08006287c7ff9d5d058593560e863f

Observation de5cda9c-beef-486e-8e3b-f51dd0b660cf · outbound

This paper cites A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.188655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.188655Z digest=sha256:da6245372ad99758982caae00e74a1b155e2353739843b8e36b395f02c9b5396

Observation fb1b5d8a-fc0d-4f20-baaa-3956624fb948 · outbound

This paper cites Security and privacy challenges of large language models: A survey,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Security and privacy challenges of large language models: A survey,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.194272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.194272Z digest=sha256:40f76cee5d3b994411fc6621bae6dcd472113210f8d7c98947b18780fb27b253

Observation e4245da9-307c-43c7-9d71-8c95c8245d53 · outbound

This paper cites Hierarchical Repository-Level Code Summarization for Business Applications Using Local LLMs.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Hierarchical Repository-Level Code Summarization for Business Applications Using Local LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.201458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.201458Z digest=sha256:71ac019ec074bca99216d8b2e42823cfff3bf1a5d0a06bd280ae7934d4a2f199

Observation a0de50b7-d19b-447a-82fa-61a647a202e0 · outbound

This paper cites Using ollama,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Using ollama,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.206736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.206736Z digest=sha256:ce257a28666aa38f191df6883e8426f2eaaccbd07f60b0127532b57c00945d30

Observation a86f1378-8681-47e0-a439-8386e407f3f4 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Efficient memory management for large language model serving with pagedattention,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.212463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.212463Z digest=sha256:a9af41e492d85a861548b6bd9d8962899687d741846249c33e1ad6bc6367f382

Observation 6a3e9bb0-2cca-41c1-acce-1740743c9e2c · outbound

This paper cites Efficiently programming large language models using sglang.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Efficiently programming large language models using sglang

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.218200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.218200Z digest=sha256:fc1f16ea7919a02ba7f7ccf8a8af5081c39e8d4fe160a0f563e8a8b371425ae3

Observation 88e3acd8-986c-4367-ae0a-968afae4c9ba · outbound

This paper cites an unresolved cited work.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:34:40.707301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.223034Z digest=sha256:c1806f0f623871f32bcae586d1e0c106bbb8268d0efc41411d337f3d38c8f8ce

Observation f9dfc8c3-0af7-4947-9eb9-76716a57586a · outbound

This paper cites Orca: A distributed serving system for{Transformer-Based}generative models,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Orca: A distributed serving system for{Transformer-Based}generative models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.682085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.228640Z digest=sha256:67f7e6516cb49fd4372d17b99735c23d3a090c03d5cb43562032a67bdb5c2b77

Observation 6b06a26a-ab7e-4c13-8a75-f654254649d1 · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Llumnix: Dynamic scheduling for large language model serving,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.234092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.234092Z digest=sha256:714a4b47c5b838226fcc4a30b9dbe54199ae10631514b59e7c630ce901fd007b

Observation 14004c0c-c6cd-4aca-b686-b31b093e8020 · outbound

This paper cites Batch: Machine learn- ing inference serving on serverless platforms with adaptive batching,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Batch: Machine learn- ing inference serving on serverless platforms with adaptive batching,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.650655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.240589Z digest=sha256:edf51fdf6429035f0914f5729d4502a096fcb1919a74be9fb7bd1cef314a9c65

Observation 4b467c65-bc88-42d3-942d-1f11c637c742 · outbound

This paper cites an unresolved cited work.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:34:40.621627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.246143Z digest=sha256:fda591ae8538e8064ea402b6fa2b89d0cede1a0fb658313347ff541e830dbfe3

Observation 92ebd5ed-4418-4c73-aae6-d5b413ab1033 · outbound

This paper cites Efficient deep neural network serving: Fast and furious,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Efficient deep neural network serving: Fast and furious,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.593282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.251838Z digest=sha256:9a8a2aba74cd4b34af6d904205c3f00f061686a26a661bb84ede43b75ef5bdf7

Observation 7b077c26-7f02-4c97-bc65-36badbd8e441 · outbound

This paper cites A comprehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees A comprehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.257810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.257810Z digest=sha256:3fa6dcffc0dabc8d402ba367f65c87b55d27b115c4a6c887169a2e18d39d9dd1

Observation 89693262-08b4-4aca-9dad-bd170f4f25b9 · outbound

This paper cites Survey of uncertainty estimation in large language models-sources, methods, applications, and challenge,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Survey of uncertainty estimation in large language models-sources, methods, applications, and challenge,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.544746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.263884Z digest=sha256:995723be656924d91b4c4cf2905da2b15a0038e53b8dd9d1eb00804f94e455da

Observation c44c90cf-5c58-4362-ae34-376fab391135 · outbound

This paper cites Enabling efficient batch serving for lmaas via generation length prediction,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Enabling efficient batch serving for lmaas via generation length prediction,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.511175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.270108Z digest=sha256:5a169e2f505ce5f773e1bc09c1fe9622af7c94bd7153e630f622e5736fbdc00e

Observation 6a2720c6-3fe5-4df3-8cab-5ac2e9704ee3 · outbound

This paper cites [performance]: [v1] increasing the request batch size causes a significant drop in performance,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: [v1] increasing the request batch size causes a significant drop in performance,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.490366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.276247Z digest=sha256:0e2d37780383f6453d8e2bd074fe3f2c2d5593fc3d8fb8e96336fa48269ba46b

Observation 7c6a5c16-1a97-460e-8382-232caa403ddf · outbound

This paper cites [performance]: Added request take too much time, and the model will not run untill all the request are added into the cache,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: Added request take too much time, and the model will not run untill all the request are added into the cache,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.467521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.282691Z digest=sha256:ecd2fef5348c66ae54ff600c1eb7c5f2a51c824d64bf342ed1fc004d7925d627

Observation d333ecd6-60c1-43c3-a91c-7f704869250c · outbound

This paper cites [performance]: Why does the tpot increase with the request rate increase?.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: Why does the tpot increase with the request rate increase?

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.445506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.289069Z digest=sha256:70995e19f926a3ee7ddef932808654af8def4ad60073fa6e326ecf404cc517d3

Observation 1e11c2fc-df77-4507-a78f-0a568eedddaa · outbound

This paper cites [performance]: Ttft spikes when qps increases during deepseek- r1 testing with tp8 and pp2,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: Ttft spikes when qps increases during deepseek- r1 testing with tp8 and pp2,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.420072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.295612Z digest=sha256:5e5dbcaa5cd9708bb3f1b1aea0c5b908194c5ae994b23a783a0872c0ae9f730b

Observation 63039161-2ca4-4e2c-a2b2-1f7096b417f9 · outbound

This paper cites [performance]: poor performance in pipeline parallesm when batch-size is large,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: poor performance in pipeline parallesm when batch-size is large,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.389296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.302539Z digest=sha256:bd3ba2692c41e1076de31f0f392005f0f2e7480c4ff7149d155f5d7b6a729cfb

Observation 6a219196-1625-497f-9766-d0269f70d488 · outbound

This paper cites [performance]: How to improve performance under concurrency,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees [performance]: How to improve performance under concurrency,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.354487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.313230Z digest=sha256:ff3752ed1df6315f9074b9ff467e50dbcfff479490bf8c9e13b914331d9dce9b

Observation dac3f033-f474-435b-9908-ae705c93400c · outbound

This paper cites {SHEPHERD}: Serv- ing{DNNs}in the wild,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees {SHEPHERD}: Serv- ing{DNNs}in the wild,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.331059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.318307Z digest=sha256:5fac32ba3029fedc441e26fa5291909dd62b11dfa738942a8277d20dc9bd2954

Observation fc85e42c-445e-4c5f-9ca3-c736a7f1623f · outbound

This paper cites Serving{DNNs}like clockwork: Performance predictability from the bottom up,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Serving{DNNs}like clockwork: Performance predictability from the bottom up,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.299244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.324673Z digest=sha256:48917b4a01fc7eca1247f01da5a8f8edbc2469431017fa7d223d223f6b94ba47

Observation 7f1229fe-691c-4f25-841c-a679c64f1c80 · outbound

This paper cites {INFaaS}: Automated model-less inference serving,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees {INFaaS}: Automated model-less inference serving,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.270319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.332023Z digest=sha256:2f04e3490c67e335c145c8ba910f97d26ea2ae364f09e18e3ce097579d182fd5

Observation f50f9745-b263-4575-b556-0e150cb23a47 · outbound

This paper cites Llama: A heterogeneous & serverless framework for auto-tuning video analytics pipelines,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Llama: A heterogeneous & serverless framework for auto-tuning video analytics pipelines,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.239989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.337702Z digest=sha256:d1e041a37da60be02d4b054537aa32cfd5454986da9113bc65a95edf65fb9ba4

Observation 0abfbf75-a5bd-4d1e-ae6f-e57c0c0b8f48 · outbound

This paper cites Llumnix: Dynamic Scheduling for Large Language Model Serving.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.343001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.343001Z digest=sha256:ed19d3a4581f3226afb48d57d7d4d6d5e2e7ac06486cb278bea6099ad585c9bf

Observation 6d0980ed-da06-4666-a1f9-6df49d0f4ce8 · outbound

This paper cites {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.348915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.348915Z digest=sha256:450714758eaca25aa531cbe479710f232b64a3711fa7c8e8789f8a862ec42498

Observation 21f521eb-5287-4847-bfeb-ea744a2059ac · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.354289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.354289Z digest=sha256:1b7b43a0ad2b52b1607ea6dff387265d0623b1367bd8b7cfa5e3d827c2afde17

Observation 6d72c8f2-8d06-49c2-bfb3-2007f5d3bef8 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.360289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.360289Z digest=sha256:13acbaf87ec151385a231e4516f7392e044cc73e11752b0b87975507835dd23b

Observation 6b687260-c7e0-4390-a200-4ad5d118d0ff · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Splitwise: Efficient generative llm inference using phase splitting,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.366807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.366807Z digest=sha256:218ffd72f40e3b74916a009508a3d5259e448ecd75eecc8a6abced61988e7c23

Observation f4dfdf25-cd56-47a8-abc6-fd64155fa57c · outbound

This paper cites ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:34:39.600636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.373993Z digest=sha256:dc1087a205f104fc45a28c0d5a979d5760c3d3d117641f9c8a79158d04a7823f

Observation 94df0066-4de7-4408-a830-7d0e450f6759 · outbound

This paper cites Niyama : Breaking the Silos of LLM Inference Serving.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Niyama : Breaking the Silos of LLM Inference Serving

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.381247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.381247Z digest=sha256:056ecb3eb65e1eb398c4983ecf0e0b1b297ff0006a1eae6b6637bca9baa13931

Observation 558ef808-996f-4c05-8fbd-566875a77a00 · outbound

This paper cites Hl-codellama-chat-response dataset,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Hl-codellama-chat-response dataset,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.174778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.386759Z digest=sha256:f91614794cc33759a1287bd939038d1c96b2211ccac282cef5b82ec840c9f9f9

Observation 472ce62f-cb3a-439c-bf13-eacd8a120661 · outbound

This paper cites Synthetic code generations dataset,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Synthetic code generations dataset,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.143508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.392953Z digest=sha256:ca3630be525577dbb82ae6a47f74d6c9fc1d2fbb0bb8522ca87311639513fe42

Observation 4d48ba10-b845-4ee8-b8c9-133f8d9f1a82 · outbound

This paper cites Code summary java dataset,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Code summary java dataset,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.111151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.398191Z digest=sha256:8e4b94dfa7688ff8ca689c13eedfef33c5c5a5bb2f38be9582691d5c7917d08d

Observation 3e23a315-b542-4f1f-bce9-46d62a2dfad2 · outbound

This paper cites Code translation dataset,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Code translation dataset,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.075428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.403669Z digest=sha256:0d5e61c79433e9a39e9c3d92419e3b64ceaff6d673148280fe78df65563c85bf

Observation 651f9bd1-fc9f-4958-902d-92d799ad9696 · outbound

This paper cites Learned Best-Effort LLM Serving.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Learned Best-Effort LLM Serving

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.408803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.408803Z digest=sha256:9df6647a76ad529f2745df54b115e33af1e29a1263b104121e0fc39e03131aff

Observation be3f75a8-9c4f-4c8b-9263-3e8fb2d439ec · outbound

This paper cites Preble: Efficient Distributed Prompt Scheduling for LLM Serving.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Preble: Efficient Distributed Prompt Scheduling for LLM Serving

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.414400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.414400Z digest=sha256:088b1c291ec89696dc4699e24458dc70eee028b99926755f0ec6d9988ce3d72c

Observation 20cc696d-a5b2-448a-9bff-b0a632792fa4 · outbound

This paper cites Past-future scheduler for llm serving under sla guarantees,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Past-future scheduler for llm serving under sla guarantees,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.049337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.419712Z digest=sha256:c47c5ef7d30da9e7dc222c61bff561be8d0b28c5f8aed07db739965ae8a4fc32

Observation 44a109d9-1384-44a7-9802-babe4e010d8b · outbound

This paper cites ShareGPT Dataset: A Collection of ChatGPT Conversations,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees ShareGPT Dataset: A Collection of ChatGPT Conversations,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:40.019737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.425745Z digest=sha256:ed5797f92b08ca37477cfd684f8c4a45ca7bd536b345ad693370fe966baf5591

Observation 269a077a-b8b5-40d8-83d7-fd363b68b917 · outbound

This paper cites Integrating concurrency control in n-tier application scaling management in the cloud,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Integrating concurrency control in n-tier application scaling management in the cloud,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.992542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.431313Z digest=sha256:7f16a22e015db55bd32927f3538a4a388a9d0084fc7a064dacb61d6ce8c7b9b5

Observation 09d6d1a8-8d6c-4637-b690-604c3accdab1 · outbound

This paper cites An r-square coefficient based on final prediction error,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees An r-square coefficient based on final prediction error,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.966406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.436591Z digest=sha256:cf341edc66ed983a5b92c64cbfb811e688ebad6c4ba2432e7dadca6330ec7088

Observation b2a8c6f6-3f6f-4d29-8a61-959cdc3073f1 · outbound

This paper cites Scipy 1.0: fundamental algorithms for scientific computing in python,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Scipy 1.0: fundamental algorithms for scientific computing in python,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.442683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.442683Z digest=sha256:2cc3736913985f15e897efe210920490b442ec78a241a02dcfdc426fbe6f19e3

Observation 328358dd-936f-4e48-9342-cc8dc09198ec · outbound

This paper cites Multi-dimensional sla-based resource allocation for multi-tier cloud computing systems,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Multi-dimensional sla-based resource allocation for multi-tier cloud computing systems,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.919225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.448659Z digest=sha256:b01552c8e333a13eba673016389bdb5fd9c7a12b2aa4bafc0e65b11cf7f87473

Observation 780e5dfe-4108-4577-84f5-14c7ff256e7c · outbound

This paper cites Decision model for cloud comput- ing under sla constraints,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Decision model for cloud comput- ing under sla constraints,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.898075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.455244Z digest=sha256:7c3c3004b45d69967fdc95d2e298293107e6f8b00ceefe7b0f22c11118d20118

Observation 1018c585-1bd8-4c21-88f1-880752fcb863 · outbound

This paper cites When average is not average: large response time fluctuations in n-tier systems,.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees When average is not average: large response time fluctuations in n-tier systems,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:34:39.878885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:34:39.461915Z digest=sha256:77dfc54c2778be10d30dae9d55e4cf233897fddf1c67f9288b23d58ee0cd37a8

Pith citing papers

No inbound Pith citation observations are available.