Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2403.02310.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T20:53:30.772845Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
15
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 8f3425e2-02d2-4560-9573-4846b10184c5 · inbound
A Survey on Efficient Inference for Large Language Models Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 283
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1e99bc29-4825-41dc-bf6e-482fac912f0b · inbound
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b80e34f7-c743-48f1-a56a-f98d807df92c · inbound
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a28d3b4-f021-4bb3-b926-0ec0c9ee9ddc · inbound
ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 97bcbd00-63ac-4df6-83be-2e57677b482c · inbound
PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c0e65acb-5eca-404e-823f-06ae489253f3 · inbound
Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9aa1a6b3-df44-4281-9401-d6426f28a503 · inbound
Agentic Witnessing: Pragmatic and Scalable TEE-Enabled Privacy-Preserving Auditing Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 43579636-675c-44d7-b5d5-786a088a8e5c · inbound
Agentic AI Systems Should Be Designed as Marginal Token Allocators Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 01a9a38f-abc6-4197-a51c-c117f799eff8 · inbound
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 74d58261-838b-4fd0-808d-b7d5c2fa05f4 · inbound
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dab488d6-f358-4543-9ec7-46d02a4c250b · inbound
Computational Challenges in Token Economics: Bridging Economic Theory and AI System Design Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9fce4adf-ae6f-4a6d-a3a4-4ef3950b03fc · inbound
HarnessAPI: A Skill-First Framework for Unified Streaming APIs and MCP Tools Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bd969b09-0574-4b31-9ffa-01e9be155afa · inbound
DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3748cd8c-fe39-488d-b958-23b9b146be6f · inbound
A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ea36a92f-e078-4765-92ad-07ed17376370 · inbound
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3070805-629f-4d65-ac3a-4b88cb3866ef · inbound
Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f11ace76-a222-43e0-a171-c74632f0b438 · inbound
Beyond Prediction: Tail-Aware Scheduling for LLM Inference Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2ca9d91d-efa0-46a4-826a-af33269b5ace · inbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 183
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 67cb0682-5fcb-49f9-9ace-d0fb8fe0be3c · inbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 183
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd03649a-e348-4420-b995-3264601b0d54 · inbound
Load Testing for Machine Learning Model Serving Systems at Scale Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c478561c-37b0-40dd-88e9-9e9473efcb13 · inbound
Speculative Decoding at Temperature Zero: A Scoped Safety-Invariance Screen with a 48,072-Sample Expansion Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0970c22c-a640-4754-a322-26d9829e9c6c · inbound
PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb35a35e-61eb-43a1-8743-61017d6e8f5e · inbound
KernelSight-LM: A Kernel-Level LLM Inference Simulator Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c25b08c1-89d8-4736-af91-0e1d640d5bc4 · inbound
KernelSight-LM: A Kernel-Level LLM Inference Simulator Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d2306179-367e-4b31-a14f-2eea20d99c93 · inbound
Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c37c6f3-b581-4bfd-aa59-fab3fe947a51 · inbound
OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fbfecaf3-58cc-4976-a941-5373c19bff38 · inbound
Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS Processes Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ddf9c1-8738-4d6a-a367-f9c2f8226ff3 · inbound
Think Before You Grid-Search: Floor-First Triage for LLM Serving Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bfd610e5-6faf-4584-9844-83ca056397f3 · inbound
Think Before You Grid-Search: Floor-First Triage for LLM Serving Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b7863588-a8d2-43b4-8caf-9c84ee16a394 · inbound
Persistent Computational State: A Session-Centric Runtime for Generative World Models Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46da00a1-c551-488c-8393-79b665384216 · inbound
Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.