Pith. sign in

Paper Citation Record · LEDGER

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization

As of 21 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2502.15763.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.15763 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:55:04.355426Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec4585e5-84c2-4920-865d-64c0a125d9be · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Efficient memory management for large language model serving with pagedattention,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.247328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.247328Z digest=sha256:e083e10b474241c9634625ed775f91c47914dd7397887d6b413f6f4eddbd39cf

Observation 63f79b81-13e9-4886-9dc2-16b15356ebf1 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Orca: A distributed serving system for {Transformer-Based} generative models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.251658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.251658Z digest=sha256:7d54e66bd614984322050d0f8dead6e1738824a0b6e3f371506819907bf1ecd9

Observation 5d34f8df-52dc-4982-b129-85a6998b33a8 · outbound

This paper cites NVIDIA Announces Financial Results for Fourth Quarter and Fiscal 2024,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization NVIDIA Announces Financial Results for Fourth Quarter and Fiscal 2024,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:55:04.630193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T18:55:04.255137Z digest=sha256:c79fec5852fb4a300020bf36dffe1efa2ef5110306b0046ea0f010df2fda2e55

Observation 4e8fa402-4777-446f-996c-6d17079cd198 · outbound

This paper cites Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.258838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.258838Z digest=sha256:e6d7bd66367cbe0ac5df0ac8169729416a2396e5163b3183b8ae628d4fb000b8

Observation a2107f8d-b8b7-4e61-ab44-c7932de99b71 · outbound

This paper cites Llumnix: Dynamic Scheduling for Large Language Model Serving.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Llumnix: Dynamic Scheduling for Large Language Model Serving

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.262738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.262738Z digest=sha256:ca5a0e024b7318fd832210198d62983924ea014e97bf186e85e604ddfccb3e81

Observation 26fdc164-e8ec-4580-8b0b-10d9aacad476 · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache manage- ment,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache manage- ment,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.266415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.266415Z digest=sha256:54d39c84a3c9d37145ca1c674756463d4d2d3c856d251907c79ce3ecac794574

Observation ebd6cff6-6b2b-4c88-8ba0-cc7ff836ad7d · outbound

This paper cites {dLoRA}: Dynamically orchestrating requests and adapters for {LoRA}{LLM} serving,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization {dLoRA}: Dynamically orchestrating requests and adapters for {LoRA}{LLM} serving,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.269646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.269646Z digest=sha256:145d8c583acb57e826944cae024f749d707304c09f1a8502e4100730afb8dac4

Observation 217c071e-2182-432b-b3ae-446d1a85b001 · outbound

This paper cites Fairness in serving large language models,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Fairness in serving large language models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.272430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.272430Z digest=sha256:a86e8c41c444baf0cd2d7a53fa6e1d7b3b20d3f1f198812c177ab9c066da44b2

Observation a3bcd99a-78b5-4f9d-833d-58e845c4ca66 · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Fast Distributed Inference Serving for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.275192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.275192Z digest=sha256:b0f6b2c3cfeeae609c34aeb0f9fd40e5735b7d7dda051ba90502e580d225e83c

Observation 7b6b5e0a-91e0-4a1d-b752-c6ca16e455ee · outbound

This paper cites {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:55:04.605687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T18:55:04.278374Z digest=sha256:765a204c72ba220a21966e66adb89f0357d7a6ab34dd42dc1bbe9d8f0b2b1a27

Observation ed1c43cf-2700-4225-b701-1a38c196612e · outbound

This paper cites Quantization and training of neural networks for efficient integer-arithmetic-only inference,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Quantization and training of neural networks for efficient integer-arithmetic-only inference,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.281167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.281167Z digest=sha256:b47334ee5c6f8842db3b70c7e8a097bae539817df2a600f1342b126f4b0b5d29

Observation 823abd1b-c07b-4383-92d5-c8851d0c2391 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Generating Long Sequences with Sparse Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.284157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.284157Z digest=sha256:d42e9dfe8c047d7d701b922061fce406c4844d6e84b6d7b409c0c21595bd6967

Observation f65b6bef-05a6-489f-9a6a-62826fc87103 · outbound

This paper cites Sparsellm: Towards global pruning of pre-trained language models,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Sparsellm: Towards global pruning of pre-trained language models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:55:04.591216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T18:55:04.287314Z digest=sha256:a1baad9dcebe815eac15b32b2a99482ca0cbd489540efdac945fa83797ea2685

Observation 7f8e762f-27f7-4f02-b404-e7c7af33c6c9 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Distilling the Knowledge in a Neural Network

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.290035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.290035Z digest=sha256:d6078ac283f2793c123783bad9383678bbdbb91345addfb94b280428cd1ffce8

Observation 107e4a54-c0db-4ab5-af3d-448ba831ddb7 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Fast Transformer Decoding: One Write-Head is All You Need

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.293244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.293244Z digest=sha256:2c8a1cf9fe32b14f950bde45bec6a0a6edcafe9a451d20c307b25e7828203dc1

Observation 24055fa7-e84d-4ca6-890e-34633648d845 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.296963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.296963Z digest=sha256:a70d7c3185fe0bdaced08e7be2d9d88055494ce5d50daf5de406f785c9700d83

Observation c5c2018f-f620-41fc-933b-7130c305012f · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Efficient large-scale language model training on gpu clusters using megatron-lm,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.301223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.301223Z digest=sha256:4ff2f4506000e9360d2951d72a8c49a4cd8df28b7747a944ebe52f95aca68675

Observation 8d97fafc-e9a8-493e-b97a-1ba3dfd2933c · outbound

This paper cites Gpipe: Easy scaling with micro-batch pipeline parallelism,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Gpipe: Easy scaling with micro-batch pipeline parallelism,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:55:04.576513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T18:55:04.304314Z digest=sha256:c7128cef018c489a09341b785823f337b825e8af077132c4dd393de8994879c2

Observation 4d92bd38-203e-4592-8366-78d03d20d628 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.307566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.307566Z digest=sha256:f6c8a274775edd419069730d196ca32b9cac7f9391583cfdd970d0ab41f31267

Observation 205fc4b0-30e8-4e9f-b287-5ade9ebdc7fc · outbound

This paper cites Reducing activation recomputation in large transformer models,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Reducing activation recomputation in large transformer models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:55:04.567209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T18:55:04.311288Z digest=sha256:9d86db364ae1f71a69ca4bb7a30f2c7bf18715427e5c6f6a1912e075f984cd9d

Observation 7dfe8a2d-538a-4e54-b789-7bdbb4f56be1 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.314448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.314448Z digest=sha256:f4b75d10ce4dbc4d2bbe217df1bbbc30ef3beed842ae6f629b5d25a89f08cafa

Observation b6e1049c-ed18-4882-80fb-70cdc048af05 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.317745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.317745Z digest=sha256:e4bac32ca2bc06df3aa2991ef072721fe38ce7891cb57ac41cedc17b795a93cb

Observation 32072092-3d38-486d-81a2-bf588d04d3d4 · outbound

This paper cites Fast inference from transform- ers via speculative decoding,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Fast inference from transform- ers via speculative decoding,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.320856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.320856Z digest=sha256:e4c9164dc7fb6469296942173c2b4ec29dd778b455db0bebe607de5216ca849e

Observation e40bfe89-6d39-4759-af7e-92a886af9f78 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Accelerating Large Language Model Decoding with Speculative Sampling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.323833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.323833Z digest=sha256:c2e863e8a3f977b1afbd6a76272988fe086e329489ddc45ab97de6a2589a5a16

Observation 44d01514-e713-48a2-b989-d6811b0121fa · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Splitwise: Efficient generative llm inference using phase splitting,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.327290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.327290Z digest=sha256:95ad3971c3d835bd3e1d6202b34d5733d7a1776025791477596aa76750861011

Observation bf3a698b-6d10-4bad-bedf-03331e1c72c7 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.330334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.330334Z digest=sha256:535b45dcf7f4c0b293811e1668110ff61d8b8798041d38b2049a46a4610dd1c3

Observation 91da8edb-25d0-4524-bc2e-f0c338051bc0 · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization A Survey on Efficient Inference for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.333568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.333568Z digest=sha256:fdfa2d7ddefa8a0184bd4079d17025bb2163c62fd4c0d337f4ad06da6ff2eb0d

Observation fd3f04f4-9b2c-4b92-9c97-d860bc8f9eb7 · outbound

This paper cites Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.337027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.337027Z digest=sha256:c009929cdb48dc317e5616e37a7466df13cb92bd79a2ff266551fbca5407efd7

Observation 2ea1a0e1-949a-495b-a9cf-decd88342df5 · outbound

This paper cites Print surface thermal modeling and layer time control for large-scale additive manu- facturing,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Print surface thermal modeling and layer time control for large-scale additive manu- facturing,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:55:04.541190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T18:55:04.340598Z digest=sha256:8a94f35f750aca34ad50f7e1855ab55d57e4bbde9f35e219b23086f09fac5f41

Observation 511422f8-4566-4ff9-8e02-3fa4f0181b00 · outbound

This paper cites Surgery scheduling under case cancellation and surgery duration uncertainty,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Surgery scheduling under case cancellation and surgery duration uncertainty,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:55:04.531937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T18:55:04.343558Z digest=sha256:21f12594fd11ddb6ec92e826c7ea35ea1d0a7fb68357e07d9fdac704b8f3f389

Observation be03ed37-ee6e-469a-af1d-1cc1bfa93f56 · outbound

This paper cites Online appointment sequencing and scheduling,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Online appointment sequencing and scheduling,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:55:04.522557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T18:55:04.346744Z digest=sha256:8dc52bee6497a033fb9a497e49222613cd1aad91ff2d5e3f0bb2a89b37e08463

Observation 59ab1e9d-8c59-4396-9b79-99006beebba4 · outbound

This paper cites A dynamic sequential decision- making model on mri real-time scheduling with simulation-based opti- mization,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization A dynamic sequential decision- making model on mri real-time scheduling with simulation-based opti- mization,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:55:04.513317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T18:55:04.349497Z digest=sha256:a3e737fd73888ed3d9b3a65510374ceb23c479147ffe24d1f4a8d7d3b47574ca

Observation 9df9f3c7-785b-43cc-bb9e-9b3aca5dabf1 · outbound

This paper cites Online scheduling of ordered flow shops,.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Online scheduling of ordered flow shops,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:55:04.503561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T18:55:04.352473Z digest=sha256:488a120477828ac5fb16b4bf7df75ac2573485fca8bd4790b0e6ea37d31450c8

Observation fcff8344-4f85-46d3-a7b6-1cb35329bc14 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Training Verifiers to Solve Math Word Problems

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:55:04.355426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:55:04.355426Z digest=sha256:8b206a35a81b8f280193a3aaf3c9750dedb34e8f3e2c61da770c99ff43c19be5

Pith citing papers

No inbound Pith citation observations are available.