Pith. sign in

Paper Citation Record · LEDGER

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start

As of 4 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 2 inbound Pith citation observations for arXiv:2604.06664.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.06664 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:24:42.887971Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:52:43.089348Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-20T02:12:58.334659Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact13
  • verified fuzzy10
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 519cfe44-bf3c-467b-b820-f4e77c6ad26f · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.300025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:e5c720daa11ac8322936fb69e6e3d02db2c4bcd34eda6607aca82aa7d6c15dfd

Observation 8d5fb11d-d504-4237-b058-3cec193d1766 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.322432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:2205eb948ea324cb8e310d0aaabe05c8208e5eb9ce76f8a575fd4d2d8da6bf8b

Observation 63b49a5d-b279-46c3-accc-2954e933d687 · outbound

This paper cites 2025.Let Tensors Fly — Accelerating Large Model Weight Loading with R-Fork.https: //lmsys.org/blog/2025-12-10-rfork/.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start 2025.Let Tensors Fly — Accelerating Large Model Weight Loading with R-Fork.https: //lmsys.org/blog/2025-12-10-rfork/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:54:00.286925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:5e11d5572650704acdc6f18030ff33a328d9b325a416faf1739988af99a5b1b6

Observation 186aa887-d2fc-4238-8922-a969df56c673 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.290459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:55c2a3d82061390852bb1bffde4290bbaae503fd96bd681651a31aff22bf2a1b

Observation 7f56689b-cfb4-4201-8f38-a0d3cdb13556 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.296869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:88fa48a854f5d72fbf9dc4edf446711464ca310b56b880dbab92df8cde1d476c

Observation c3c566b1-61cc-4bae-8ce6-3bf6b0e9c3f9 · outbound

This paper cites In2023 USENIX Annual Technical Conference (USENIX ATC 23).

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start In2023 USENIX Annual Technical Conference (USENIX ATC 23)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:54:00.306914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:b0722a7f82c8bd605410db10a353ca1d32acd334d1b201a7834bc5ca800885be

Observation c89709f7-58e0-430d-9c5b-fe6491c37bb5 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.318770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:5a4f0b78c2c10c2d2015f75505abd4ac345b99edd75fe8c6e1deab9e9483c7c4

Observation 2afeb18a-cc60-449d-9053-557c4904f7b2 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.303338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:f268fdd072a1a2fab19aa1084fbcb61dc7f00b14727a6c1ff61a37f4806584e9

Observation 9abf8720-7ade-45d1-8233-be53d6b0f4a2 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.311028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:db7a51267ea4856c84c9a6403c163e50bef121574d74d9869e74811f32775197

Observation 20b3a2fa-b8ef-41d5-9838-23cb30cee85e · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.293569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:f072318ffc6e5443e719b60457fb6bb4538f5f67482f6dc8042183b525ae6817

Observation 7602c92c-29fb-4875-ab13-c97d76be6915 · outbound

This paper cites Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:41:01.512723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:99a1668e6a7d501acd3367fad5852401e3143de065cd5d08c43850b17878f939

Observation 7c6379b2-11ae-4985-81bc-0708e7c90252 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.213564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:d3bb3b7f7f7761a7b08573856301663e04589f299b33613f557069d6866bbcd5

Observation 9c01bece-c97f-476c-9a22-5099491c69f0 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.177952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:5fe1c36b0eecd1fd5564a1a03b23f4842185267d17e350c6fcb6933e206e9912

Observation 7fd5788c-2869-4622-8af6-196ddf78a43a · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.203811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:efd31e1d4802a5a44fbc1e19bc6b64de67dadd6ae7fffc6e8478a0a7c697d542

Observation 990cc7c3-ee5c-49a5-af6e-b7e6187f83b2 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.245902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:100809cc0aeea037235257aba6fbbf205efa67bedceb912dd98c6e2e18211ae2

Observation 705d941b-b6cb-49b2-a6dd-7c255def34c0 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.218145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:64d48fb94283704d2e7462517555ba95afdc4fea5914648aeffdc67b505d0860

Observation e7c2437e-ea26-4aa6-9d8e-30c785b7d041 · outbound

This paper cites Flying serving: On-the-fly parallelism switching for large language model serving.arXiv preprint arXiv:2602.22593,.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Flying serving: On-the-fly parallelism switching for large language model serving.arXiv preprint arXiv:2602.22593,

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:01.387653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:3472ab147a4cb6299669085846acb02380d3f1efa2f3e9beec50c31a581ba9a0

Observation 7ccd7787-6ea0-44e0-9b8a-5bed3c491658 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.239430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:ba3777947e8f2402986f511d17494628b8c2c395b1108cd3b6faecaa35d19f35

Observation d9f5e39f-d63c-4cf2-9da8-c8c61cdfffd0 · outbound

This paper cites In2018 IEEE International Conference on Cluster Computing (CLUSTER).

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start In2018 IEEE International Conference on Cluster Computing (CLUSTER)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:54:00.282990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:651ee334704f4ea665c9c8ac7218e3df697d856d2ee00acbb12f7865668b912b

Observation 92431362-0309-4381-bdb7-5a890f3f554f · outbound

This paper cites The Llama 3 Herd of Models.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start The Llama 3 Herd of Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:41:01.572787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:c734e0ed38178f6a8ea9fcf0d1a6a30c555b56a720633ceecf59beee34e792c5

Observation 4689a1a8-540e-454a-b67f-5594342f0897 · outbound

This paper cites Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:42:01.775082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:74000e39a676bd91b7f83d17585ffd7bb0e72eb11f2d24ee0768a2d8310ee2cd

Observation 7b309813-33be-4dcf-a8d1-deaab7673ca4 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.158879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:f2122a92b3ce93b5710d36e0764595d42a2f8fbc9f9cc0cb6e37efcc08d6d241

Observation 68e6cee9-7bae-4b04-89a8-2c71685a1b7d · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.242653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:71cc3d1cf99a79dbd098ceeae9e56f6c97358f46f7f74b34eb4f9716dc27e07c

Observation 9006a53b-4ca8-4a2f-a9e0-a34da324770f · outbound

This paper cites Gemma 3 Technical Report.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Gemma 3 Technical Report

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:41:01.567245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:c82f70bc001812d5f1a75b5ba0ebcb83a734863155f9b39ce5db492942a204bf

Observation 1e425e3c-d96a-4636-be7a-791bc5522ade · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.188201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:3d101fcac04aebc53978b8e0528ca9dbb97cf696e949972114553612784d5116

Observation ce5e41f3-48ff-4017-9b8a-fe54742cd04b · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Gonzalez, Hao Zhang, and Ion Stoica

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:54:00.200222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:cde87d12fb44bc1382227e3def7bbef3af1058fffd66ac6116da57380a9602c4

Observation 38f53fd8-803c-48d5-bf93-ec36b1a46d89 · outbound

This paper cites In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP ’23).

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP ’23)

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:41:01.769468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:57cd152874379c9419b7e0cd53d293faef7ae90845f836301152687760fd2494

Observation 1c09efa7-efda-4364-87c7-bc22f61d0033 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.257482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:6708153b59d01f98ecfce936158802ccfb5c6038acbb13f3b29f54c1204f3b96

Observation b9cfe369-cf68-4f08-8169-c6797cc53325 · outbound

This paper cites 2025.Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start 2025.Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:01.620019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:10e89d7aae69472f8e78a300b054a67ef63c469587909c162b00deb8ebeb692e

Observation 8690f2af-ff41-4dd8-af9a-0d0249c9f330 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.232207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:22bb1d679bebceaf291727d89108a5622b9582d2b4946db1253c1dfe7015c901

Observation 4b76ed9b-46e7-4196-bd54-4b0b51d2cef8 · outbound

This paper cites WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:52.823910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:2bc1ceb6ceba1750de7a407ad39d32ab857b482e6d666284a8550c5edd7d3db1

Observation 7a1ca8c5-1ac7-4c92-809d-de4ec042e2d9 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.264987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:7394163889e41874a46fc1bead289b7619d2528955f8071bf83a2d1b473f1a8f

Observation 0e79af0d-79e8-4159-a2e7-28e61f98c845 · outbound

This paper cites NVIDIA.https://docs .nvidia.com/cuda/cuda-driver- api/group__CUDA__GRAPH.htmlAccessed March 5.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start NVIDIA.https://docs .nvidia.com/cuda/cuda-driver- api/group__CUDA__GRAPH.htmlAccessed March 5

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:54:00.272827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:c8282f5b87773f36a20272ee4e691889c9b5f0c3291e9e671812a5a609dc57b9

Observation f3befc6f-a929-4e44-a59c-bbc3eae0895e · outbound

This paper cites 2026.NCCL: Optimized primitives for inter-GPU communica- tion.https://github.com/NVIDIA/nccl.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start 2026.NCCL: Optimized primitives for inter-GPU communica- tion.https://github.com/NVIDIA/nccl

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:54:00.196143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:5f9cade6818974aeb131f8295c2b1a2b54717328d9c1cd26a6bec63c7b4fadfe

Observation 1474ef13-4aca-4613-ae0e-b445a0c1a69f · outbound

This paper cites 2026.NVSHMEM: The NVIDIA SHMEM library for GPU clusters.https://github.com/NVIDIA/nvshmem.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start 2026.NVSHMEM: The NVIDIA SHMEM library for GPU clusters.https://github.com/NVIDIA/nvshmem

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:54:00.209149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:85c93ffa7ebb408bb78880da7991ac0d16eae33eb99ff357f80f0b39f7290e60

Observation c809b55c-65f9-4a96-8edc-bf690e0605c4 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.181278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:acca63e29dfc08a267041db0a54e8297eb1877143134a96ec0db4569c60a9f60

Observation f02537f4-b941-40d2-9884-60ba5cbf0fe1 · outbound

This paper cites 2025.Weight Transfer for RL Post-Training in under 2 seconds.https://research .perplexity.ai/articles/weight-transfer-for- rl-post-training-in-under-2-seconds 13.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start 2025.Weight Transfer for RL Post-Training in under 2 seconds.https://research .perplexity.ai/articles/weight-transfer-for- rl-post-training-in-under-2-seconds 13

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:54:00.224870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:e8ed5cba8222178e99e3f77469434ba1d81c897c593e745b0b2aff7a63945f43

Observation 846ebf8b-b70c-4221-848d-b15cd88445c1 · outbound

This paper cites vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention,.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention,

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:25:42.224988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:562f5755b144987130b64e58869452240fa0175ff226ef321953de195fd1aa35

Observation 3e5ec513-48a4-42b8-8eb0-94724751e168 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.235508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:6297785c13b7f6e59caaddc48180835d0cec49826dde4cd1de3f8fce7dd11e7a

Observation 8023bd15-7971-4431-81cc-67daf1548ef8 · outbound

This paper cites ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:01.482759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:9eca7c621205087d6a5fa308d7e20b186718d0b028592ff9d08a1d74a0465f58

Observation 5746d400-1a22-4aee-9149-f9f9cb639847 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.221305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:8109b1563321625cd356ae8c1459d91676f327cf75eb2753a9133d5049cbcaa6

Observation 73c1365c-f44f-49ba-b087-b1f4766ba024 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.162495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:b724fdfb4fad55cf3e6fcc1149c662d7d783ffb339eb50fdeb34b7ebef6ef3a5

Observation c7d3d83b-52ad-499c-ac98-aec12e62e2ae · outbound

This paper cites CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:01.675221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:9865689b5c3858a1fa6339c7ec0f749a1df4357f94e0ae43a2b70632ac61238c

Observation 8ef62e13-4beb-40ae-933d-0062e009a5a3 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.228620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:0f8304de8bea50fba1bd96e5bbf5213c61cdf49609344421d40c56ebabc1b044

Observation 1b84e5cf-4996-45c1-85db-43c1edc0481e · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.174512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:7efd984c4377929cb6a28075ff33c66cf35e2ac7b52ed4b354ef01c4c6d4f00b

Observation d4ec762d-a45b-4be9-8994-4857cfb3603f · outbound

This paper cites 2026.Elastic EP in SGLang: Achieving Partial Failure Tolerance for DeepSeek MoE Deploy- ments.https://www .lmsys.org/blog/2026-03-25-eep-partial-failure- tolerance/.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start 2026.Elastic EP in SGLang: Achieving Partial Failure Tolerance for DeepSeek MoE Deploy- ments.https://www .lmsys.org/blog/2026-03-25-eep-partial-failure- tolerance/

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:54:00.269134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:6bcedb9fd1ddd10cfa100692373da00b9e6f34df70921f6029d8c37503c8f308

Observation 676db65c-613c-4a2e-8780-be8b4f148866 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.276313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:db49bec0b06a28817abaf78735da8e734f26f972cad04c339128bde3334c0127

Observation 641a95e8-8605-4104-b286-61800c8a5e65 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.279265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:6c6d54ca7127d237a7d1d8f75815cfb5f6c8b8120d422ffac642af97cb1ddabf

Observation 7b005a35-8f36-439f-8299-860dd4056bba · outbound

This paper cites Aegaeon: Effective gpu pooling for concurrent llm serving on the market.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Aegaeon: Effective gpu pooling for concurrent llm serving on the market

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:25:42.220672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:1a7a7fbf91797a23685864e408b82b046f4e25feb99816b833e62818e95516e1

Observation 30824795-e557-48e9-b1c5-b104a954eaf1 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.261540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:5ac6a087c7cffb0add9bf27341953c289d68a46d4c6a53f8500f6b09f292187d

Observation 6bbe1455-bc65-43fb-98c7-76990269a634 · outbound

This paper cites Qwen3 Technical Report.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Qwen3 Technical Report

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:41:01.502719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:565fe3d247613e39f107555008dfbc1eeb214446f193dbdc2dde4acd51e0aa9c

Observation 4ce7ae01-0e8e-4f79-9dc1-1d343ee9bce2 · outbound

This paper cites Lambdas- cale: Enabling fast scaling for serverless large language model inference.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Lambdas- cale: Enabling fast scaling for serverless large language model inference

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:01.236616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:d658001aa4b77dbc519c52cf390d6698d10abeae171687421cf5ae069f798f5b

Observation a346645f-8653-4794-a46d-c1e1bab47eb4 · outbound

This paper cites Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-12T02:08:18.534775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:240af41dfa611c4e41bbc17c6fa6eb3daa14ef9577ffc9743c52a82b06e3f5a1

Observation 56c8730f-71af-4b94-bb1d-fc5bbf8bf3ef · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.250538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:d2930a0e21a115da9b26efb86de6c1ee447d1a80e43769db0b08520b12912c62

Observation becdce75-5260-41e4-8e72-9db465b26e60 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:25:42.216438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:882d0ca3cd0f16a320e8e2756e0b218ce24b47f521cea06d9c5a4dcf34969816

Observation 4dc64771-106b-40be-b9b6-15cfe5311fb7 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.191775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:0de789f367a5a2b001d7baeb5173e2639d3de7252c7ed662bed841b5e172610f

Observation e6290f43-6bdc-4bd1-8821-8821b73447f7 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.166212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:3d68bf080e7b99c9d4e7910ce9dafbc360a03adb1fdb28ec669a03a1a72b438d

Observation 81b70ee4-afb9-4d47-9c30-afa191e090ef · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.169841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:32d290abe7853c7c57e4ebbedf4fdf3de5f7d10959afbf3cef18c71a7f0a6f21

Observation b36572a8-67f5-488c-8bf0-c51eda5febf9 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:54:00.184683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:0b9259842807f3edeae9a70331036a7933d521e2107f9a575b430a3cb1a97848

Observation c279c4d6-7aa4-431f-a765-a64b3403e2f2 · outbound

This paper cites an unresolved cited work.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start Unresolved cited work

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:01.494180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:9e619588017e8c8606150d66a8b5e13334b770f452a907e66f476f7bba232c62

Observation ec0b2086-aab8-445c-8071-d1c8385d7ad9 · outbound

This paper cites id": 7.

Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start id": 7

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:54:00.253918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:42.887971Z digest=sha256:4eddf47681e932615bde61a38e4aaa5456b436a0486dc7bb6ef2e8a86f9040f8

Pith citing papers

Observation d036bfed-c54f-4977-9299-1608fca2e7af · inbound

C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG cites this paper.

C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-20T02:12:58.337827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T02:10:57.582345Z digest=sha256:5288f818e9c50506667c4b8a5501e78b789fda36b01bc5047a9fa76311e00c77

Observation bf594ba4-f0cf-473f-9b6f-0b0cd97ea08d · inbound

InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata cites this paper.

InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T13:52:43.089348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:52:43.089348Z digest=sha256:86a4df32d9906f55e61414d35fd886d186d9f343a126cbaf3a5d05e46adeede4