Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T22:53:58.352008Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2606.17787.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T22:53:58.352008Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a50ef0c8-068d-4fa3-9305-9fbaff6a797e · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving https:// docs.vllm.ai/en/stable/deployment/k8s/, 2026
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df5d1cf0-a1c7-4c6d-b9b1-43563d9ff6c8 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving https: //github.com/flashinfer-ai/flashinfer, 2026
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eec8535-1c49-4233-8c49-e7445cde2f5c · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acf3725c-eac4-47f5-94b1-c5d9e7be79c4 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving https: //huggingface.co/docs/text-generation- inference, 2026
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 701ca8b4-f9b7-4b55-8585-2de165840038 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving https://github.com/kserve/kserve, 2026
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23f45e03-73f7-413e-a239-e7085c7510c8 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving https: //github.com/triton-inference-server/server, 2026
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4695858-47cc-434c-a8e0-c99a4a748883 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving https://pytorch.org, 2026
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afd8ff5c-6475-49f8-9b2f-9e36f1cd873e · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving https://zeromq.org, 2026
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b713c946-798f-4d39-aad2-cdf79a4cd3a1 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Vidur: A large-scale simulation framework for LLM inference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf4fbf55-0dfd-4729-8e6a-f573a0dda2e6 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Taming throughput- latency tradeoff in LLM inference with Sarathi-Serve
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31f6ed29-fa9a-4e33-98c0-5ede73efd04a · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98d1a7b1-350b-44ed-9c07-be3785c4ef8b · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Accelerating Large Language Model Decoding with Speculative Sampling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8d98aca0-3450-4bb5-8e64-ca52f26e3ecd · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Recycle: Resilient training of large DNNs using pipeline adaptation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94ea7503-a455-4be8-9f72-d01228cb1099 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Cost-efficient large language model serving for multi-turn conversations with CachedAtten- tion
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a747b482-a093-44e9-b235-fc9c34e81dfb · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Characterization of large language model devel- opment in the datacenter
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d1b946-fa90-4ad8-ad8f-bfb6e3f6d005 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Oobleck: Resilient distributed training of large models using pipeline templates
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 515ecadb-84bc-437c-8c44-7a9e3072bdeb · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving GhostServe: A lightweight checkpointing system in the shadow for fault-tolerant LLM serving
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95689ca0-8f85-440a-8936-2ef05fadd90c · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving MegaScale: Scaling large language model training to more than 10,000 GPUs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed395e94-2832-470a-ad0d-f0f5027ebac4 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Revisiting reliability in large-scale machine learning research clusters
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bc8901c-819e-41fb-ba2a-8d082e8e96a7 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Efficient memory man- agement for large language model serving with Page- dAttention
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5175a2b5-3a68-4f8b-8a1d-cfa9fdadb63e · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Fast inference from transformers via speculative decoding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation added2ae-d8c7-463b-9ed1-25ef75b9edc5 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving PEARL: Parallel speculative decoding with adaptive draft length
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb39cfe3-cd5f-4adb-87eb-ef22e76cc1ee · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving CacheGen: KV cache compression and streaming for fast large language model serving
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bac4c3d-c534-47ac-855a-ef8d24467246 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving The Llama 3 Herd of Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dba359e2-87a3-45fd-a6a3-4205127ea8ab · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving AMUSD: Asynchronous multi- device speculative decoding for LLM acceleration
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f31f6144-a051-47f3-a1be-f18954711ad6 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Splitwise: Efficient generative LLM inference using phase splitting
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19009165-96f7-46d6-9df1-04dc08ef38c7 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving ECCheck: Enhancing in-memory check- point with erasure coding in distributed DNN training
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ad37d1b-49c6-43b4-a423-a497bfc98a1b · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Towards resiliency in large language model serving with KevlarFlow
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 153437a6-363a-4fe3-88a0-09f5421cd3e1 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Mooncake: Trading more storage for less computation – a KVCache-centric architecture for serving LLM chat- bot
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 360ecb2f-7a53-46a8-aebe-021cd4b73805 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving ShareGPT conversation dataset
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fc473c8-2533-4ddb-9c19-f87ba7c755f4 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving FlexGen: High- throughput generative inference of large language mod- els with a single GPU
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0befe7a2-ac5b-4a55-a9cc-07357fb3f436 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving DéjàVu: KV-cache streaming for fast, fault-tolerant generative LLM serv- ing
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94d765e1-aebf-4d92-b496-5c9e849f8c99 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Llumnix: Dy- namic scheduling for large language model serving
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb7d7f18-bf5b-48ad-b556-301b453006ff · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Bamboo: Making preemptible in- stances resilient for affordable training of large DNNs
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e352b913-5f7a-4493-8a1f-644d3e34d3d3 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 60346195-0523-4335-bf39-8b4459dc9add · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b86e601-9fd8-484d-879f-9d759b9a2487 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving ByteCheckpoint: A unified checkpointing system for large foundation model development
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 767c681d-bcb1-4ac1-aa04-ad2c3c248818 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d92a333d-db58-44b8-a9e6-6bed58a0a42a · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Fast distributed inference serving for large language models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d92c734-89a3-4b37-8d4f-b355a1466328 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving FailSafe: High-performance resilient serv- ing
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 132556ef-0f08-4109-b294-bf595c950621 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Qwen3 Technical Report
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 147a22fe-75dd-4747-8ed3-2c5113223e8e · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Orca: A distributed serving system for transformer-based generative models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c57243e-6044-475c-a09e-84826308776b · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Gonzalez, Clark Bar- rett, and Ying Sheng
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2349f40-b36b-4963-bab5-575d3f837a11 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Dist- Serve: Disaggregating prefill and decoding for goodput- optimized large language model serving
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 650d2676-44dd-4bc8-8308-af8833cd27e8 · outbound
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving Resiliency at scale: Managing Google’s TPUv4 machine learning supercomputer
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.