Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T18:55:04.355426Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2502.15763.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T18:55:04.355426Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ec4585e5-84c2-4920-865d-64c0a125d9be · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Efficient memory management for large language model serving with pagedattention,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63f79b81-13e9-4886-9dc2-16b15356ebf1 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Orca: A distributed serving system for {Transformer-Based} generative models,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d34f8df-52dc-4982-b129-85a6998b33a8 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization NVIDIA Announces Financial Results for Fourth Quarter and Fiscal 2024,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4e8fa402-4777-446f-996c-6d17079cd198 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2107f8d-b8b7-4e61-ab44-c7932de99b71 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Llumnix: Dynamic Scheduling for Large Language Model Serving
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26fdc164-e8ec-4580-8b0b-10d9aacad476 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache manage- ment,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebd6cff6-6b2b-4c88-8ba0-cc7ff836ad7d · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization {dLoRA}: Dynamically orchestrating requests and adapters for {LoRA}{LLM} serving,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 217c071e-2182-432b-b3ae-446d1a85b001 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Fairness in serving large language models,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3bcd99a-78b5-4f9d-833d-58e845c4ca66 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Fast Distributed Inference Serving for Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b6b5e0a-91e0-4a1d-b752-c6ca16e455ee · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ed1c43cf-2700-4225-b701-1a38c196612e · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Quantization and training of neural networks for efficient integer-arithmetic-only inference,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 823abd1b-c07b-4383-92d5-c8851d0c2391 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Generating Long Sequences with Sparse Transformers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f65b6bef-05a6-489f-9a6a-62826fc87103 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Sparsellm: Towards global pruning of pre-trained language models,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7f8e762f-27f7-4f02-b404-e7c7af33c6c9 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Distilling the Knowledge in a Neural Network
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 107e4a54-c0db-4ab5-af3d-448ba831ddb7 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Fast Transformer Decoding: One Write-Head is All You Need
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24055fa7-e84d-4ca6-890e-34633648d845 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5c2018f-f620-41fc-933b-7130c305012f · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Efficient large-scale language model training on gpu clusters using megatron-lm,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d97fafc-e9a8-493e-b97a-1ba3dfd2933c · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Gpipe: Easy scaling with micro-batch pipeline parallelism,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4d92bd38-203e-4592-8366-78d03d20d628 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 205fc4b0-30e8-4e9f-b287-5ade9ebdc7fc · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Reducing activation recomputation in large transformer models,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7dfe8a2d-538a-4e54-b789-7bdbb4f56be1 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Ring Attention with Blockwise Transformers for Near-Infinite Context
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6e1049c-ed18-4882-80fb-70cdc048af05 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Flashattention: Fast and memory-efficient exact attention with io-awareness,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32072092-3d38-486d-81a2-bf588d04d3d4 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Fast inference from transform- ers via speculative decoding,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e40bfe89-6d39-4759-af7e-92a886af9f78 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Accelerating Large Language Model Decoding with Speculative Sampling
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44d01514-e713-48a2-b989-d6811b0121fa · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Splitwise: Efficient generative llm inference using phase splitting,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf3a698b-6d10-4bad-bedf-03331e1c72c7 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91da8edb-25d0-4524-bc2e-f0c338051bc0 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization A Survey on Efficient Inference for Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd3f04f4-9b2c-4b92-9c97-d860bc8f9eb7 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea1a0e1-949a-495b-a9cf-decd88342df5 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Print surface thermal modeling and layer time control for large-scale additive manu- facturing,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 511422f8-4566-4ff9-8e02-3fa4f0181b00 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Surgery scheduling under case cancellation and surgery duration uncertainty,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation be03ed37-ee6e-469a-af1d-1cc1bfa93f56 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Online appointment sequencing and scheduling,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 59ab1e9d-8c59-4396-9b79-99006beebba4 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization A dynamic sequential decision- making model on mri real-time scheduling with simulation-based opti- mization,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9df9f3c7-785b-43cc-bb9e-9b3aca5dabf1 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Online scheduling of ordered flow shops,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fcff8344-4f85-46d3-a7b6-1cb35329bc14 · outbound
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization Training Verifiers to Solve Math Word Problems
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.