Pith. sign in

Paper Citation Record · LEDGER

Llumnix: Dynamic Scheduling for Large Language Model Serving

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2406.03243.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.03243 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:39:53.844797Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T05:09:36.948141Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a8662879-5233-4297-891e-10034a42c6f2 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 288

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.533295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:fadec63ad8492f08b465fc0b487c943a3abd7506577830bc7af959fa36419c8d

Observation 4fedd64d-faa8-4223-a09b-a1d9e33d15df · inbound

FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving cites this paper.

FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T11:17:01.075475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:17:01.075475Z digest=sha256:52b435c0c389c0fa9b74f202a6f7e78d1776d07d2779da17ba487d7a6599aa27

Observation 4ad4c246-0ca0-4201-9a5f-5d6b97938b77 · inbound

SYMPHONY: Improving Memory Management for LLM Inference Workloads cites this paper.

SYMPHONY: Improving Memory Management for LLM Inference Workloads Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T10:39:11.677104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:39:11.677104Z digest=sha256:5502c4b1acaeb9bb9cfd00c966106871f6b555dc8b18e077137588c150470657

Observation 61a8e9f1-e8e6-413d-97f5-cec861972ef5 · inbound

Hierarchical Autoscaling for Large Language Model Serving with Chiron cites this paper.

Hierarchical Autoscaling for Large Language Model Serving with Chiron Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.144196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.144196Z digest=sha256:ee3b600d0058d0f40a44c94c5fa73c0f1896057b953235bfbb2f51bbe1188385

Observation a2107f8d-b8b7-4e61-ab44-c7932de99b71 · inbound

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization cites this paper.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.262738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.262738Z digest=sha256:7f97456685aee4bcafc55cee2b72256e2ad82693e500cb9d98d9527535086e11

Observation c194e30f-14df-4134-9cfc-1ba29cc268e7 · inbound

SLO-Aware Scheduling for Large Language Model Inferences cites this paper.

SLO-Aware Scheduling for Large Language Model Inferences Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:39:53.844797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:39:53.844797Z digest=sha256:c63d0e2d888181e7d3759949b13a02e969066e2755604172e254ac4b2196ca14

Observation 8c624cac-0437-4c13-811c-aed11907cc3a · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:10.166931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:10.166931Z digest=sha256:6b064012e0efaf5544eefcac05c8a39466a36fc7a48ebaf9ec8b4b3c1acf5443

Observation 0abfbf75-a5bd-4d1e-ae6f-e57c0c0b8f48 · inbound

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees cites this paper.

Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T18:34:39.343001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:34:39.343001Z digest=sha256:ed19d3a4581f3226afb48d57d7d4d6d5e2e7ac06486cb278bea6099ad585c9bf

Observation d580ae2e-c08a-4398-9254-a1a722f9e353 · inbound

EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices cites this paper.

EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T20:59:01.514175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:59:01.514175Z digest=sha256:a1802e853e0cdd9b1f713236400bea4ea9f81e8c99aad614107fe0bf081048ac

Observation ec890e84-bc22-4e75-a72c-e1508795c22d · inbound

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators cites this paper.

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:51.976790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T19:04:17.725111Z digest=sha256:255377b09805fee6f4ada8f74fc0601fce0537f41680331385f4209337649676

Observation 9b8ecfe1-4204-47c1-86a8-ff631983be04 · inbound

PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR cites this paper.

PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:29:25.459494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T02:24:48.872065Z digest=sha256:9931594e39af69643985fb95967bafc0d5582fef50af21c1a07b0c023f1c9562

Observation 69c54834-2df4-4c20-99a8-fd16db5453c1 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 209

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:09:36.950269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:f937e5de46fb1dbbc18fdd013553f1ef18fa5c4e00776580ba5312e3cfd52fb6

Observation 12f4dc19-0ea6-498b-9e11-7c6542185cf7 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 209

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:23.421625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:23.421625Z digest=sha256:6336034b0666078a8407d3d904e8a918f6ef18608d6a9061e8cae4678cae2be2

Observation 69adefae-59f4-4d84-b581-8b282b50b1bd · inbound

OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters cites this paper.

OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:37:42.178086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-03T06:30:30.308713Z digest=sha256:860a6863aa145adf63d9b5d9d4a047c4308d46c4814e6ba5b2f190f057dbf068

Observation 606ba1af-719c-4a2e-98f9-0ee96a7ab5f8 · inbound

A Workflow-Aware Serving Layer for Agentic Applications cites this paper.

A Workflow-Aware Serving Layer for Agentic Applications Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T05:56:18.797766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:56:18.797766Z digest=sha256:fcfd28466b160d1780444a0a011c410903d1753ca5ee78fb13ba78b8b4b81a91

Observation d9952405-2363-41ae-9c04-8732e4777c9a · inbound

Kalypso: Relational LLM Serving cites this paper.

Kalypso: Relational LLM Serving Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-30T11:40:16.380604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:40:16.380604Z digest=sha256:96964f0baf107ca48e879a0bc757873094e39757c3de2088b275ddd813e4cd6b

Observation dee33baa-89dd-4005-8575-75202281c9fe · inbound

Spatial Prefix Caching for Wireless Edge LLM Inference: A Stochastic-Geometry and Queueing Framework cites this paper.

Spatial Prefix Caching for Wireless Edge LLM Inference: A Stochastic-Geometry and Queueing Framework Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:33:27.295482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:33:27.295482Z digest=sha256:f40e102aff436fa081a55e84f0eaa647910c1666412fafb4d68b9feaf22b26d7