Pith. sign in

Paper Citation Record · LEDGER

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling

As of 9 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2608.02244.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02244 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:54:03.289004Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e7973ad6-073e-4949-9cf5-cf86c4b8eaec · outbound

This paper cites an unresolved cited work.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:00.908542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:00.908542Z digest=sha256:2b273bef1d244de5ae51915b18446da21d9c3fcab313cc025e3f9760024b6939

Observation 2c6cbed0-7a3d-47ba-ac04-8671d85aed5f · outbound

This paper cites Bo Chen, Chris N Potts, and Gerhard J Woeginger.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Bo Chen, Chris N Potts, and Gerhard J Woeginger

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:01.018364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:01.018364Z digest=sha256:38f1276982fe300bc51d0f7e97abd7bb9e1f3b6e863634aa6c1167f2305ca5b8

Observation df53f8bc-474a-4790-9a86-10abbd444944 · outbound

This paper cites Adaptively Robust LLM Inference Optimization under Prediction Uncertainty.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Adaptively Robust LLM Inference Optimization under Prediction Uncertainty

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:01.285290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:01.285290Z digest=sha256:8e8ffe23bf7a89d956afb0a8802bea1f3f3be83e94ab333af9ccfb97e94b4fc8

Observation 5f6f5cf8-53dd-4cff-98ea-f88c73790568 · outbound

This paper cites 40 DeepSeek.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling 40 DeepSeek

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:01.350872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:01.350872Z digest=sha256:b65e2d22fbdcd966bb41504719253a665a3f25e5c832a9d6c106c1b02076217c

Observation cf9fde2a-7409-466c-9bd2-f4b44b203437 · outbound

This paper cites Accessed: 2026-06-10.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Accessed: 2026-06-10

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:01.383420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:01.383420Z digest=sha256:16da7123c4620ae4b9ec92f06f3bb59080b960451936097465e5ce635ab98211

Observation 2af17c6a-1691-456d-8e74-c68d5ec88bb5 · outbound

This paper cites Online Resource Allocation with Convex-set Machine-Learned Advice.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Online Resource Allocation with Convex-set Machine-Learned Advice

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:01.521832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:01.521832Z digest=sha256:fe9cd2f5a6e3a6fad360f8c934573fbf999735f4a0194771c2c4c9b0d5c78339

Observation 2571490d-1d02-41c2-8867-b5b9a920d62f · outbound

This paper cites Ac- cessed: 2026-06-10.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Ac- cessed: 2026-06-10

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:01.629578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:01.629578Z digest=sha256:70da3c71ee8645d180ba5bbd6fb22c69583a1e6826f8b335a343288c80b59876

Observation 5ad9925b-cd20-4058-9865-074128677b22 · outbound

This paper cites Patrick Jaillet, Jiashuo Jiang, Konstantina Mellou, Marco Molinaro, Chara Podimata, and Zijie Zhou.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Patrick Jaillet, Jiashuo Jiang, Konstantina Mellou, Marco Molinaro, Chara Podimata, and Zijie Zhou

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:01.788171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:01.788171Z digest=sha256:6f309a9ba744d51a7a8ef2c33c48cc5daaddf9dd44acf2e5537b4cb47d4992dc

Observation 4f84c343-0ef8-4ec0-86cf-166c78aa7bb0 · outbound

This paper cites Redwan Ibne Seraj Khan, Kunal Jain, Haiying Shen, Ankur Mallick, Anjaly Parayil, Anoop Kulkarni, Steve Kofsky, Pankhuri Choudhary, Renee St Amant, Rujia Wang, et al.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Redwan Ibne Seraj Khan, Kunal Jain, Haiying Shen, Ankur Mallick, Anjaly Parayil, Anoop Kulkarni, Steve Kofsky, Pankhuri Choudhary, Renee St Amant, Rujia Wang, et al

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:01.847860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:01.847860Z digest=sha256:7510c7c7a015dfb30a9d8fa0d0f33ebdad1243d9443ee7325d9c5ffdfe1a3cb6

Observation 5f8dbb02-5c22-47fd-b8be-a10e48552223 · outbound

This paper cites Ensuring Fair LLM Serving Amid Diverse Applications.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Ensuring Fair LLM Serving Amid Diverse Applications

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:01.930191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:01.930191Z digest=sha256:e7da5d6522b00dd0454ae7f1ac561633d42ccfc0fc4ea7b281e5572678682b45

Observation f2e77c7e-5fe4-49f3-a623-5e6b83fc75fd · outbound

This paper cites an unresolved cited work.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:02.214600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:02.214600Z digest=sha256:75c8a05d0a1b6b2ad29b19eff1446285faaa67fd48926296e9b6d59a6a86d962

Observation 668ffd06-84ce-4678-919d-6f8fe449c161 · outbound

This paper cites an unresolved cited work.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:02.456372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:02.456372Z digest=sha256:6a1886e73964237948ae20e3779b8e12b08e272e7293d0b88201dbe8f2ebaab6

Observation 6a3abbdd-bbda-463c-9b30-0398b321c20a · outbound

This paper cites GPT-4 Technical Report.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling GPT-4 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:02.588106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:02.588106Z digest=sha256:1eb5f9e38e6fcebae92d4f20773f1d56617fbd3820b97c828ac071a8d498b14d

Observation 7e181c7d-b3da-4e49-9fe9-7dc5e00ecd21 · outbound

This paper cites Accessed: 2026-06-10.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Accessed: 2026-06-10

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:02.719597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:02.719597Z digest=sha256:4341c82d3168411efbc5768f5b70f972bf8eece1b0e0f5aa810cd8b9247e4f72

Observation 641ec49e-2bf3-476b-bfca-41d7dcb912e9 · outbound

This paper cites Julien Robert and Nicolas Schabanel.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Julien Robert and Nicolas Schabanel

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:02.813457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:02.813457Z digest=sha256:2c74681fd3ec0a06e1269619327f12feb4e394bdb199e3913839f0499004e1ac

Observation e187760f-dbe4-42ce-bd71-3f81186a6a64 · outbound

This paper cites In Proceedings of the Nineteenth Annual ACM-SIAM Symposium on Discrete Algorithms, SODA 2008, San Francisco, California, USA, January 20-22, 2008, Shang-Hua Teng (Ed.).

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling In Proceedings of the Nineteenth Annual ACM-SIAM Symposium on Discrete Algorithms, SODA 2008, San Francisco, California, USA, January 20-22, 2008, Shang-Hua Teng (Ed.)

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:02.883870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:02.883870Z digest=sha256:ebf387fa7c23d2ae74054add10a36a2dd9f565fee9d1f8704db423a0e1ca3356

Observation 70da5781-0ec5-496e-bbc3-17de77757ce7 · outbound

This paper cites Rana Shahout, Eran Malach, Chunwei Liu, Weifan Jiang, Minlan Yu, and Michael Mitzenmacher.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Rana Shahout, Eran Malach, Chunwei Liu, Weifan Jiang, Minlan Yu, and Michael Mitzenmacher

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:02.973735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:02.973735Z digest=sha256:3a9b0590f2ac8c131a1ffb7f328b00865dd86e530db687e626af3e561ae00307

Observation 7a132beb-3140-4865-aba1-858a953f1282 · outbound

This paper cites Don't Stop Me Now: Embedding Based Scheduling for LLMs.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Don't Stop Me Now: Embedding Based Scheduling for LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:03.018347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:03.018347Z digest=sha256:bbf8e096bb0cb92a341bc0bbda522f28ca77ae8d43d57f1070d69afb81cb5afc

Observation a0baf015-af7d-44f8-a8b5-2be04de56ed9 · outbound

This paper cites LLM Serving Optimization with Variable Prefill and Decode Lengths.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling LLM Serving Optimization with Variable Prefill and Decode Lengths

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:03.079660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:03.079660Z digest=sha256:094f0cc4697d4433ef0fdbb8905d5fabe8c4dca181643dd6e08e0c3f3ffda90f

Observation db97d505-d2a9-403b-98fa-7ec2732478f8 · outbound

This paper cites Equinox: Holistic Fair Scheduling in Serving Large Language Models.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Equinox: Holistic Fair Scheduling in Serving Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:03.143633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:03.143633Z digest=sha256:e515f5d689394724491c3636afe382e995a5ba5e20b1ffea097d0364b50af90f

Observation cdef2cf6-f521-4abe-b227-83ae8c4a340d · outbound

This paper cites LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:03.201603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:03.201603Z digest=sha256:2ea53815158881664df85907ed5f3f7f0e782d49a7ce150c0e2ba404ddb1986a

Observation 1fc05e27-74b7-4cb7-ac38-3710fe703192 · outbound

This paper cites DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:03.289004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:03.289004Z digest=sha256:391b2f4e937da053e902021f7b5ab95a2fd210d08305abe8b8b1ae2ee0d9514d

Observation 087a142a-1a32-4d3d-885f-4d9a4fb9f36b · outbound

This paper cites 2023.Dissecting Batching Effects in GPT Inference.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling 2023.Dissecting Batching Effects in GPT Inference

Reference 1641

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:01.225727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:01.225727Z digest=sha256:1d29bd05d3753997b716793faf58e94475bd0072dde4f411f3b0adca33ed1c38

Observation ddcb80a6-0fc5-4b22-b1e3-ba91dabb83d5 · outbound

This paper cites an unresolved cited work.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Unresolved cited work

Reference 1969

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:01.699637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:01.699637Z digest=sha256:e581eb62473eb9d3abe831daac0b44e4922d92121080525510ccb24520cf9f69

Observation 4b9c7015-8c18-420d-bd5b-f1a797240a87 · outbound

This paper cites an unresolved cited work.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Unresolved cited work

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:01.138037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:01.138037Z digest=sha256:0b2fdcae5a42851b9195019c7aa7a234ac46e37c92c193361f7f16811714b6e9

Observation dff2a208-8646-43ab-8ac6-2911380ffef2 · outbound

This paper cites Wenhua Li, Libo Wang, Xing Chai, and Hang Yuan.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Wenhua Li, Libo Wang, Xing Chai, and Hang Yuan

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:02.116767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:02.116767Z digest=sha256:4d7bbac1b1c4806b8f779554803eb8e2504b606a336f64751608e24060fddd8a

Observation ead3a966-1d85-46d4-9779-90fa2866b0b7 · outbound

This paper cites doi:10.1016/S0166-218X(01)00272-4Special Issue devoted to Foundation of Heuristics in Combinatoria l Optimization.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling doi:10.1016/S0166-218X(01)00272-4Special Issue devoted to Foundation of Heuristics in Combinatoria l Optimization

Reference 2002

Resolution
verified exact
doi, observed 2026-08-04T10:58:25.697257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T10:54:00.678116Z digest=sha256:015524ae3f732b676234c66bdd87e8d937bf7b2a6e604e42cdf25eade696a636

Observation 9afc8553-44a8-4c1c-ac54-a1d4fb79d120 · outbound

This paper cites an unresolved cited work.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Unresolved cited work

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:00.361206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:00.361206Z digest=sha256:2a6afcceb7829bf9f7c10b20401e54d183e93ea392fad07d5030f840d82f72d8

Observation 0b1737ba-9f9c-4cb7-a2bc-edcccf6749d6 · outbound

This paper cites Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:02.024754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:02.024754Z digest=sha256:d702e85ae14684edb1822dfbf2cabd847382c8d9ff7beedcf7cb63d43c92de7b

Observation c9c7c267-c5d7-411c-817a-c603298c68e3 · outbound

This paper cites Thodoris Lykouris and Sergei Vassilvitskii.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Thodoris Lykouris and Sergei Vassilvitskii

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:02.323278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:02.323278Z digest=sha256:906714a70ec0124725a769220edc6528e1936cc1d3378adab4745f5c9d3e2fea

Observation 34dedbdf-e758-49cf-8dad-677d2e5335b7 · outbound

This paper cites In46th International Colloquium on Automata, Languages, and Programming, ICALP 2019, July 9-12, 2019, Patras, Greece (LIPIcs, Vol.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling In46th International Colloquium on Automata, Languages, and Programming, ICALP 2019, July 9-12, 2019, Patras, Greece (LIPIcs, Vol

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:01.447992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:01.447992Z digest=sha256:abf12e205bfa4834598f0041766eade2cc1f2cd9765eb22339cb928a31ba9afe

Observation 1f639b82-11bb-4afa-83d6-0ffab48c67e9 · outbound

This paper cites Marco Cascella, Jonathan Montomoli, Valentina Bellini, and Elena Bignami.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Marco Cascella, Jonathan Montomoli, Valentina Bellini, and Elena Bignami

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:00.786896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:00.786896Z digest=sha256:bb0f8d73a0812ec1a301e97344f5051f66a16b03a4b54ccda8e1eed951862b35

Observation dbf90813-f4f7-461c-9da8-eaafb5296015 · outbound

This paper cites Ho-Yin Mak, Ying Rong, and Jiawei Zhang.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Ho-Yin Mak, Ying Rong, and Jiawei Zhang

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:02.396636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:02.396636Z digest=sha256:da748c741ee5b977e0290ad01884d03dee9f72cf8484f1235be327f013820b91

Observation 33809743-9fe6-4fe0-afd9-442265d89329 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:00.298222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:00.298222Z digest=sha256:5cc2caf3ce00bd089b53fce299963a46d5f869056b0bfdedc22a2fecbe3b8b16

Observation 46da00a1-c551-488c-8393-79b665384216 · outbound

This paper cites Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:00.179459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:00.179459Z digest=sha256:5f5ad51451c99f7efa34bb9ed29a58e7786e0eb54f44c90602ea6300743809ab

Observation 17797206-444c-43c2-b75e-c2251d36ac58 · outbound

This paper cites Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:00.567728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:00.567728Z digest=sha256:49673c67cfdefe373f4a2991712ad763dbb166e718b9e71db922d4f6f578076a

Observation fdbdf55b-6bee-4ffd-97a8-c5705fbf8826 · outbound

This paper cites Accessed: 2026-06-10.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Accessed: 2026-06-10

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:00.460788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:00.460788Z digest=sha256:a8b825fa25aee0e7fd12979c9307e9f15bb427c64c24ba53ae16aa081db2bebc

Pith citing papers

No inbound Pith citation observations are available.