Pith. sign in

Paper Citation Record · LEDGER

Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2403.02310.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.02310 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:53:30.772845Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

15
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8f3425e2-02d2-4560-9573-4846b10184c5 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 283

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:39:33.509031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:d4d70e2392cde696ba5351a4252910e99a1c24f82d65a3397fa2b58c2bc59416

Observation 1e99bc29-4825-41dc-bf6e-482fac912f0b · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:44:58.114111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:895abb6d566c593308036de65f4a8e385912a7e352fe9477f96e9566f1e85489

Observation b80e34f7-c743-48f1-a56a-f98d807df92c · inbound

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism cites this paper.

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:30.772845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:53:30.772845Z digest=sha256:22f2015c208946bd10a852bc5b94d16bc2f05a79972e7757017177226a652264

Observation 1a28d3b4-f021-4bb3-b926-0ec0c9ee9ddc · inbound

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators cites this paper.

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:58:42.965092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T23:54:08.052482Z digest=sha256:ffb055bf7f7e7cd712e4bb0cfde73afa4f97e85df40bb38b31fa62f78490894e

Observation 97bcbd00-63ac-4df6-83be-2e57677b482c · inbound

PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction cites this paper.

PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:47:54.091286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T12:45:27.028757Z digest=sha256:822980c19aa2ae570a513197f7fb0da9be366757f047e18626c1594ba643f1a6

Observation c0e65acb-5eca-404e-823f-06ae489253f3 · inbound

Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling cites this paper.

Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:36:36.672643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:31:36.776819Z digest=sha256:9263a89fb94c91713d1b78b72df59ecb0abdb9c6b808d79360b27f3323945505

Observation 9aa1a6b3-df44-4281-9401-d6426f28a503 · inbound

Agentic Witnessing: Pragmatic and Scalable TEE-Enabled Privacy-Preserving Auditing cites this paper.

Agentic Witnessing: Pragmatic and Scalable TEE-Enabled Privacy-Preserving Auditing Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-09T00:34:30.279799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T02:57:11.715370Z digest=sha256:a0aac6ed7a79abf083c3b82b6ee748a9a06808b8d55695884a385d75303df7d4

Observation 43579636-675c-44d7-b5d5-786a088a8e5c · inbound

Agentic AI Systems Should Be Designed as Marginal Token Allocators cites this paper.

Agentic AI Systems Should Be Designed as Marginal Token Allocators Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:05.704482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T15:11:41.570592Z digest=sha256:e8ef6c6a6841bd0fc2e69692998f58e38f16f30895a5970ffeb7afda57a5ec54

Observation 01a9a38f-abc6-4197-a51c-c117f799eff8 · inbound

Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation cites this paper.

Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:45:58.613756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:42:32.739352Z digest=sha256:3d1c3b46cbd7f2e3bc674d57d4a17a57a4893b4a7997cdce3fa530e91640b68f

Observation 74d58261-838b-4fd0-808d-b7d5c2fa05f4 · inbound

Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation cites this paper.

Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:23.908790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T10:20:45.375375Z digest=sha256:8358ab054283ea67b4fb5c988919aff0291424cecb55240052038bbae34b2e6e

Observation dab488d6-f358-4543-9ec7-46d02a4c250b · inbound

Computational Challenges in Token Economics: Bridging Economic Theory and AI System Design cites this paper.

Computational Challenges in Token Economics: Bridging Economic Theory and AI System Design Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T13:03:17.881842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T13:00:56.755353Z digest=sha256:da861452f0e447748c5f8c3aa4bba479086ecc42320782520fe528a19e961367

Observation 9fce4adf-ae6f-4a6d-a3a4-4ef3950b03fc · inbound

HarnessAPI: A Skill-First Framework for Unified Streaming APIs and MCP Tools cites this paper.

HarnessAPI: A Skill-First Framework for Unified Streaming APIs and MCP Tools Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:24:38.270384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T05:21:46.992441Z digest=sha256:53ebe7365c4c55e565677416937252f6844faeb48cd70ffbc26de84dd2c6bc64

Observation bd969b09-0574-4b31-9ffa-01e9be155afa · inbound

DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving cites this paper.

DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:53:57.624787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T20:46:51.896596Z digest=sha256:70eb5dc83dedb25a85cdef1db64298f7f97efa424bc754b9b0a8eff4d2b66381

Observation 3748cd8c-fe39-488d-b958-23b9b146be6f · inbound

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving cites this paper.

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:03:47.840958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T18:03:19.798168Z digest=sha256:446c44f7da0e146466983d9e766a13fe8abad530c696f1401758112bfc307498

Observation ea36a92f-e078-4765-92ad-07ed17376370 · inbound

ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving cites this paper.

ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:13.223881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T18:07:52.038640Z digest=sha256:c4dccf63e1011797ef69f7eb8dba10f915714cb1999fced0040d7aaa86a39ce6

Observation f3070805-629f-4d65-ac3a-4b88cb3866ef · inbound

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents cites this paper.

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:36:59.461051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T01:07:14.691347Z digest=sha256:1b9acaf90d890bbfb38ff17f05bc1d1e958d87d8ce4630d859a5115b45fefe6a

Observation f11ace76-a222-43e0-a171-c74632f0b438 · inbound

Beyond Prediction: Tail-Aware Scheduling for LLM Inference cites this paper.

Beyond Prediction: Tail-Aware Scheduling for LLM Inference Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:58:57.825283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T01:01:34.655458Z digest=sha256:4ee2b9619d8bd2ddd0f3a425932c8313fe7a3fab923d31f61c58552d03b5d38f

Observation 2ca9d91d-efa0-46a4-826a-af33269b5ace · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 183

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T05:09:36.881011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:f6c52bdefe63923718a69107974bece200a79673737786e14554b0628883118e

Observation 67cb0682-5fcb-49f9-9ace-d0fb8fe0be3c · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 183

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:20.252837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:20.252837Z digest=sha256:48890a710c0df07b683cbc41d9491c7abbfa52782b42e7f14ef2021a6d4092b5

Observation bd03649a-e348-4420-b995-3264601b0d54 · inbound

Load Testing for Machine Learning Model Serving Systems at Scale cites this paper.

Load Testing for Machine Learning Model Serving Systems at Scale Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:19:44.171292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T11:54:48.056285Z digest=sha256:8bd753bb52d082de12da21b641779f03184aba248f526a40a14993061b6d4629

Observation c478561c-37b0-40dd-88e9-9e9473efcb13 · inbound

Speculative Decoding at Temperature Zero: A Scoped Safety-Invariance Screen with a 48,072-Sample Expansion cites this paper.

Speculative Decoding at Temperature Zero: A Scoped Safety-Invariance Screen with a 48,072-Sample Expansion Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:59:58.321693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T00:05:06.647295Z digest=sha256:aa1568b5fce276a5864c78fb8a3c3ce50b678346d07a4cac0c5f4bbea1d7cb78

Observation 0970c22c-a640-4754-a322-26d9829e9c6c · inbound

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs cites this paper.

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:23.420190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T21:04:43.874528Z digest=sha256:6c5ba80e418c1da159c0f9b65d193c6cdf1f6e933ea465a5c62fdf027770494a

Observation cb35a35e-61eb-43a1-8743-61017d6e8f5e · inbound

KernelSight-LM: A Kernel-Level LLM Inference Simulator cites this paper.

KernelSight-LM: A Kernel-Level LLM Inference Simulator Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:54:06.298161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T00:48:19.207465Z digest=sha256:906f173ba8bf13b36697ecacf2dce11b1165097e5cd96396a5a6307d68d51a85

Observation c25b08c1-89d8-4736-af91-0e1d640d5bc4 · inbound

KernelSight-LM: A Kernel-Level LLM Inference Simulator cites this paper.

KernelSight-LM: A Kernel-Level LLM Inference Simulator Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.267496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T23:09:38.092583Z digest=sha256:ff90a4eab081cd43f742c95911d940adf52ab23b0aac81ad0e2d82864cc13a00

Observation d2306179-367e-4b31-a14f-2eea20d99c93 · inbound

Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference cites this paper.

Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:35:43.651571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-01T04:13:53.722933Z digest=sha256:2f06477e279ded828f7e388c54b0a60f418511696da8353c496400c8973a626e

Observation 4c37c6f3-b581-4bfd-aa59-fab3fe947a51 · inbound

OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters cites this paper.

OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T06:37:42.164177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T06:30:30.308713Z digest=sha256:1cbbade2598c52b00f409cee40ca6224e7b865998fa62c5deb505632003afe1f

Observation fbfecaf3-58cc-4976-a941-5373c19bff38 · inbound

Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS Processes cites this paper.

Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS Processes Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T15:30:19.741437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T15:30:19.741437Z digest=sha256:b2d94b129f922b529006056d56dd680504816f63457b2603f5c8c24ba03d08cb

Observation 37ddf9c1-8738-4d6a-a367-f9c2f8226ff3 · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:45:40.096209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-08T22:38:12.637901Z digest=sha256:a38e041021e74d676d5d3443ce52c961ee5fb765c1abe9797c9515b02d5096b1

Observation bfd610e5-6faf-4584-9844-83ca056397f3 · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:51.380223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T01:55:09.658053Z digest=sha256:733ffb41083570cf62d55a6bb1a98eaadd5f02faf0ed2ee4d77e2c6bbec14c6d

Observation b7863588-a8d2-43b4-8caf-9c84ee16a394 · inbound

Persistent Computational State: A Session-Centric Runtime for Generative World Models cites this paper.

Persistent Computational State: A Session-Centric Runtime for Generative World Models Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T07:46:26.617631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:46:26.617631Z digest=sha256:cbb141aa3c4c829c4827cd624cf2ab2497cba7f38ab2882f5be4acbcd174c18c

Observation 46da00a1-c551-488c-8393-79b665384216 · inbound

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling cites this paper.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:00.179459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:00.179459Z digest=sha256:194f10c99e601b78a24851113d78e3a12885748ec36a0d06583846cf87d05d0c