Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:24:42.887971Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 2 inbound Pith citation observations for arXiv:2604.06664.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:24:42.887971Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T13:52:43.089348Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-20T02:12:58.334659Z
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 519cfe44-bf3c-467b-b820-f4e77c6ad26f · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8d5fb11d-d504-4237-b058-3cec193d1766 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 63b49a5d-b279-46c3-accc-2954e933d687 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start 2025.Let Tensors Fly — Accelerating Large Model Weight Loading with R-Fork.https: //lmsys.org/blog/2025-12-10-rfork/
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 186aa887-d2fc-4238-8922-a969df56c673 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7f56689b-cfb4-4201-8f38-a0d3cdb13556 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c3c566b1-61cc-4bae-8ce6-3bf6b0e9c3f9 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start In2023 USENIX Annual Technical Conference (USENIX ATC 23)
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c89709f7-58e0-430d-9c5b-fe6491c37bb5 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2afeb18a-cc60-449d-9053-557c4904f7b2 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9abf8720-7ade-45d1-8233-be53d6b0f4a2 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 20b3a2fa-b8ef-41d5-9838-23cb30cee85e · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7602c92c-29fb-4875-ab13-c97d76be6915 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7c6379b2-11ae-4985-81bc-0708e7c90252 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9c01bece-c97f-476c-9a22-5099491c69f0 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7fd5788c-2869-4622-8af6-196ddf78a43a · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 990cc7c3-ee5c-49a5-af6e-b7e6187f83b2 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 705d941b-b6cb-49b2-a6dd-7c255def34c0 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e7c2437e-ea26-4aa6-9d8e-30c785b7d041 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Flying serving: On-the-fly parallelism switching for large language model serving.arXiv preprint arXiv:2602.22593,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7ccd7787-6ea0-44e0-9b8a-5bed3c491658 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d9f5e39f-d63c-4cf2-9da8-c8c61cdfffd0 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start In2018 IEEE International Conference on Cluster Computing (CLUSTER)
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 92431362-0309-4381-bdb7-5a890f3f554f · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start The Llama 3 Herd of Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4689a1a8-540e-454a-b67f-5594342f0897 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7b309813-33be-4dcf-a8d1-deaab7673ca4 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 68e6cee9-7bae-4b04-89a8-2c71685a1b7d · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9006a53b-4ca8-4a2f-a9e0-a34da324770f · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Gemma 3 Technical Report
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1e425e3c-d96a-4636-be7a-791bc5522ade · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ce5e41f3-48ff-4017-9b8a-fe54742cd04b · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Gonzalez, Hao Zhang, and Ion Stoica
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 38f53fd8-803c-48d5-bf93-ec36b1a46d89 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP ’23)
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1c09efa7-efda-4364-87c7-bc22f61d0033 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b9cfe369-cf68-4f08-8169-c6797cc53325 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start 2025.Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8690f2af-ff41-4dd8-af9a-0d0249c9f330 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4b76ed9b-46e7-4196-bd54-4b0b51d2cef8 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7a1ca8c5-1ac7-4c92-809d-de4ec042e2d9 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0e79af0d-79e8-4159-a2e7-28e61f98c845 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start NVIDIA.https://docs .nvidia.com/cuda/cuda-driver- api/group__CUDA__GRAPH.htmlAccessed March 5
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f3befc6f-a929-4e44-a59c-bbc3eae0895e · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start 2026.NCCL: Optimized primitives for inter-GPU communica- tion.https://github.com/NVIDIA/nccl
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1474ef13-4aca-4613-ae0e-b445a0c1a69f · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start 2026.NVSHMEM: The NVIDIA SHMEM library for GPU clusters.https://github.com/NVIDIA/nvshmem
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c809b55c-65f9-4a96-8edc-bf690e0605c4 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f02537f4-b941-40d2-9884-60ba5cbf0fe1 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start 2025.Weight Transfer for RL Post-Training in under 2 seconds.https://research .perplexity.ai/articles/weight-transfer-for- rl-post-training-in-under-2-seconds 13
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 846ebf8b-b70c-4221-848d-b15cd88445c1 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3e5ec513-48a4-42b8-8eb0-94724751e168 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8023bd15-7971-4431-81cc-67daf1548ef8 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5746d400-1a22-4aee-9149-f9f9cb639847 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 73c1365c-f44f-49ba-b087-b1f4766ba024 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c7d3d83b-52ad-499c-ac98-aec12e62e2ae · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8ef62e13-4beb-40ae-933d-0062e009a5a3 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1b84e5cf-4996-45c1-85db-43c1edc0481e · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d4ec762d-a45b-4be9-8994-4857cfb3603f · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start 2026.Elastic EP in SGLang: Achieving Partial Failure Tolerance for DeepSeek MoE Deploy- ments.https://www .lmsys.org/blog/2026-03-25-eep-partial-failure- tolerance/
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 676db65c-613c-4a2e-8780-be8b4f148866 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 641a95e8-8605-4104-b286-61800c8a5e65 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7b005a35-8f36-439f-8299-860dd4056bba · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Aegaeon: Effective gpu pooling for concurrent llm serving on the market
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 30824795-e557-48e9-b1c5-b104a954eaf1 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6bbe1455-bc65-43fb-98c7-76990269a634 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Qwen3 Technical Report
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4ce7ae01-0e8e-4f79-9dc1-1d343ee9bce2 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Lambdas- cale: Enabling fast scaling for serverless large language model inference
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a346645f-8653-4794-a46d-c1e1bab47eb4 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 56c8730f-71af-4b94-bb1d-fc5bbf8bf3ef · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation becdce75-5260-41e4-8e72-9db465b26e60 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4dc64771-106b-40be-b9b6-15cfe5311fb7 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e6290f43-6bdc-4bd1-8821-8821b73447f7 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 81b70ee4-afb9-4d47-9c30-afa191e090ef · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b36572a8-67f5-488c-8bf0-c51eda5febf9 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c279c4d6-7aa4-431f-a765-a64b3403e2f2 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ec0b2086-aab8-445c-8071-d1c8385d7ad9 · outbound
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start id": 7
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d036bfed-c54f-4977-9299-1608fca2e7af · inbound
C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bf594ba4-f0cf-473f-9b6f-0b0cd97ea08d · inbound
InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.