Pith. sign in

Paper Citation Record · LEDGER

Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2406.13511.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.13511 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:42:31.698219Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T16:17:09.083875Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bc527404-68ee-4cf3-843e-bd8d3726b516 · inbound

Multi-Bin Batching for Increasing LLM Inference Throughput cites this paper.

Multi-Bin Batching for Increasing LLM Inference Throughput Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T23:56:23.156010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:56:23.156010Z digest=sha256:d837ab929fc6425c0f314738c2fe6b6f7e3f42a8ef4b3ae52e8452cbd162be5f

Observation 84af620b-0654-4e1f-82ba-fa65bf115b9f · inbound

DeServe: Towards Affordable Offline LLM Inference via Decentralization cites this paper.

DeServe: Towards Affordable Offline LLM Inference via Decentralization Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:16:53.025399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:16:53.025399Z digest=sha256:8f20102e2abcdc0b8f9ab7a5a25c3846fcfda9e51fe7d215b567f8f9e2f909b2

Observation f7d0968c-0f9a-4d02-9a7f-2f646c6de948 · inbound

WindVE: Collaborative CPU-NPU Vector Embedding cites this paper.

WindVE: Collaborative CPU-NPU Vector Embedding Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.698219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:42:31.698219Z digest=sha256:e1a9cdfb82a80730f9afdb55738c7c03ff0ab638f8faf7b03f08f9c953baa4cf

Observation 02d3b520-2a3d-4bf2-92f6-2f222987f600 · inbound

SLO-Aware Scheduling for Large Language Model Inferences cites this paper.

SLO-Aware Scheduling for Large Language Model Inferences Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:39:53.708221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:39:53.708221Z digest=sha256:e76890f76c8eae401d94e47c8c728ab29bed5dfe52c2508e06c0ac6661869c26

Observation dd405bdf-6ffd-4768-9eab-b96b043a70c5 · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:09.712328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:09.712328Z digest=sha256:b0097add391badc39ebe1a8537bcf4733f0c00aadaf3758c6e6015a4b3b9b6a4

Observation 2d3a9bbc-708f-4f10-a4c0-9c73b723dee9 · inbound

SCALE: Scalable Cross-Attention Learning with Extrapolation for Agentic Workflow Scheduling cites this paper.

SCALE: Scalable Cross-Attention Learning with Extrapolation for Agentic Workflow Scheduling Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:09.085114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T22:52:37.127993Z digest=sha256:a72874d6031d12ac1459b3359fafea5b0512e91608a7b33404245183c48e2e6d

Observation 23fc5ea6-32fc-4583-b2ba-e2690a3d4469 · inbound

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving cites this paper.

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:10.303506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:54:10.303506Z digest=sha256:4de55ecb400c7d8749a6f91235a20410c0c4b32ad146498db8c3a948f35e16e4