Pith. sign in

Paper Citation Record · LEDGER

Speeding up Model Loading with fastsafetensors

As of 9 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 4 inbound Pith citation observations for arXiv:2505.23072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23072 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:07.427165Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:52:43.151100Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1eb807ef-31d1-4c1d-9bdf-4ecf60c58ed7 · outbound

This paper cites Introducing ChatGPT,.

Speeding up Model Loading with fastsafetensors Introducing ChatGPT,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.656524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:03.832875Z digest=sha256:11be00fc763022891c6bdfd62ded169739d2dd25c103a095856ed39ba1fad3de

Observation f5b86204-bff4-4e3f-ac2b-9f89a0b3f10e · outbound

This paper cites Introducing Gemini: our largest and most capable AI model,.

Speeding up Model Loading with fastsafetensors Introducing Gemini: our largest and most capable AI model,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.492922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:03.902873Z digest=sha256:9addd5e494be24087dafce07f01e8d958567660789df79bfc9b232faa55a9008

Observation 9a68980f-964e-44a0-9507-846aeffe203f · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence,.

Speeding up Model Loading with fastsafetensors Granite Code Models: A Family of Open Foundation Models for Code Intelligence,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.366624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:03.973943Z digest=sha256:5f0af30a0191567d81faa75fc5ae1327d8ef757ef9a85d1d26f41492d1978f75

Observation 25defce5-5c40-4fd9-a40e-73806d7154ac · outbound

This paper cites The Shift from Models to Compound AI Systems,.

Speeding up Model Loading with fastsafetensors The Shift from Models to Compound AI Systems,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.176639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:04.077312Z digest=sha256:cad3ae4195e50949fb7c47ad8f8edc8e7b5055e7121212690adca2e989485a16

Observation 87dc08d5-980d-4bd8-b7d9-83a130d7e7d3 · outbound

This paper cites FlashAttention: Fast and memory-efficient exact attention with IO-awareness,.

Speeding up Model Loading with fastsafetensors FlashAttention: Fast and memory-efficient exact attention with IO-awareness,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.059280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:04.145099Z digest=sha256:e99df1a325bf1f58f9d58a9ac242da75cf609ef5514336c39f185d11c4968703

Observation fd73d497-5dc8-42c4-9eb6-1b2a8d87bdc7 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning,.

Speeding up Model Loading with fastsafetensors FlashAttention-2: Faster attention with better parallelism and work partitioning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.895787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:04.214691Z digest=sha256:ba6d53493a0d050cc72919e886b7c1409bffd8a9d299d2415ee6c1b62b2dc607

Observation 12cc850f-9c15-4005-87b1-9ae7c7d25ae2 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention,.

Speeding up Model Loading with fastsafetensors Efficient Memory Management for Large Language Model Serving with PagedAttention,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.758834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:04.304079Z digest=sha256:94aac1e83c0b47573703bdc332af99897adeb8aa07906e9b6913f5976aca8be9

Observation ae9808a8-4a8a-421d-9b8b-1cd773232cd5 · outbound

This paper cites Orca: A Distributed Serving System for Transformer-Based Generative Models,.

Speeding up Model Loading with fastsafetensors Orca: A Distributed Serving System for Transformer-Based Generative Models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.611414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:04.350083Z digest=sha256:e7105d40d15e76a2c5f0d39d2dea83848499d3ab3a1fe123b4095f47c0182ebd

Observation e30d29f1-ab5e-4cdb-9cc2-d96cb0e427d8 · outbound

This paper cites Accelerating Production LLMs with Combined Token/Embedding Speculators.

Speeding up Model Loading with fastsafetensors Accelerating Production LLMs with Combined Token/Embedding Speculators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.425862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.425862Z digest=sha256:e22179489db67106caf3e77c00eceafa2ab32c89240ab459c3f5a6ef998d329e

Observation e64cecdc-7e37-422b-9e32-862ec995d3f4 · outbound

This paper cites DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving.

Speeding up Model Loading with fastsafetensors DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.496281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.496281Z digest=sha256:73e10bd3601dbe190f67d72aad85529c973a40dc3d3ebe466c862e7108a83361

Observation c2718510-a082-46bb-aa53-00fac5a49f0d · outbound

This paper cites Taming throughput-latency tradeoff in llm inference with sarathi-serve,.

Speeding up Model Loading with fastsafetensors Taming throughput-latency tradeoff in llm inference with sarathi-serve,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.472887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:04.561087Z digest=sha256:f21a5fd89e1e9c06ed9b0025319a1322d549c6884ffc654d7af2dc3e50ca81e4

Observation 910f4ac6-e5ae-4023-9856-08bdf6a695fc · outbound

This paper cites Decrease PyTorch Model Load Times with CoreWeave’s Tensorizer,.

Speeding up Model Loading with fastsafetensors Decrease PyTorch Model Load Times with CoreWeave’s Tensorizer,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.309831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:04.639769Z digest=sha256:11691ed02a83fe9e762ec7414cf950a76f01e5c2d9b94c913887ba2373de2b84

Observation cda45ea4-2e7e-4ab7-952f-e4009f79e25f · outbound

This paper cites ServerlessLLM: Low-Latency Serverless Inference for Large Language Models.

Speeding up Model Loading with fastsafetensors ServerlessLLM: Low-Latency Serverless Inference for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.729125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.729125Z digest=sha256:6603e71c5ec91c70a60f9017765f0dd01a14ed08a73be18211385850022f309b

Observation 2b1adefe-1e6a-4e8c-8fe1-fb2f309c00e3 · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

Speeding up Model Loading with fastsafetensors Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.791744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.791744Z digest=sha256:7181a6bdc9255b9e13d2569c41c0a5130b9066b976c80e83cf0174b2ab7616cf

Observation 89ee66e9-bbba-4b12-ab18-81eab9f57679 · outbound

This paper cites Safetensors,.

Speeding up Model Loading with fastsafetensors Safetensors,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.115727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:04.876060Z digest=sha256:c2b03dd13b7d498f408a130b11361e3a76cc4126ab20d72bbd14c0dcd8e46615

Observation bc3dc6c4-e620-4184-8408-f68174e86e7b · outbound

This paper cites Models – Hugging Face,.

Speeding up Model Loading with fastsafetensors Models – Hugging Face,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.934865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:04.943400Z digest=sha256:b62b34a6e9f037a3d4801411782be20d84bba78c554ac5ed5349691c33575a2d

Observation 8718fe95-bf6e-4e3d-8bfd-c7ac573d3059 · outbound

This paper cites pickle — Python object Serialization,.

Speeding up Model Loading with fastsafetensors pickle — Python object Serialization,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.697345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:05.012673Z digest=sha256:1f6e09d14b239fbd6562c514effe0e8321d091ff9a2ea0377d530fcb6daab52b

Observation a4160a11-7b76-4e06-ac30-f8ef8dab91ce · outbound

This paper cites Pytorch,.

Speeding up Model Loading with fastsafetensors Pytorch,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.287909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:05.160118Z digest=sha256:baaf30877c29c02d02f8ec33520dfb29885a635b5a7eaf470bb0757deb8b90ae

Observation d5df08d8-f0c5-4082-a97e-39e272121d0e · outbound

This paper cites Available: https://docs .python.org/3/library/pickle.html.

Speeding up Model Loading with fastsafetensors Available: https://docs .python.org/3/library/pickle.html

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.472708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:05.086172Z digest=sha256:63a788ce23ae9b7e8af709fdcff5ee4187e22b4e47785fd75023fdecf93cdf34

Observation eac066d1-55ad-4914-9134-18f059fbd4c8 · outbound

This paper cites ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning.

Speeding up Model Loading with fastsafetensors ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.322950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.322950Z digest=sha256:e6eba1df168b20aa55109fd9630aefd79e11639ed1cef8ec1ab7edba8065b3d8

Observation 1c0e8d69-88b5-4a1d-9c5e-c55ee495787a · outbound

This paper cites TensorFlow: A System for Large-Scale Machine Learning,.

Speeding up Model Loading with fastsafetensors TensorFlow: A System for Large-Scale Machine Learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.162396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:05.256431Z digest=sha256:29fdf23f887e00bb4ff16644113f057847b8f759de8332db38dd38f9238c92c8

Observation f87c9321-eb6f-4360-a83c-21e94e87b89e · outbound

This paper cites ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development,.

Speeding up Model Loading with fastsafetensors ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.974543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:05.583944Z digest=sha256:6a6daf36d559441a07377b87ca699e578811e37ef8ff2d1191c9ee033693b8a7

Observation da9d9610-51ee-4e8f-8d8d-fd5c57aa3058 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel,.

Speeding up Model Loading with fastsafetensors PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.415353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.415353Z digest=sha256:24af3b54b02f5c0edb8de782d5ae0a10a556f767d8c2631ceb0f0438c69af398

Observation f5fdf0db-d9a9-45f0-a3a6-9d4f12f43a1b · outbound

This paper cites NVIDIA Magnum IO GPUDirect Storage Design Guide,.

Speeding up Model Loading with fastsafetensors NVIDIA Magnum IO GPUDirect Storage Design Guide,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.563051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:05.795124Z digest=sha256:bebc046d61e4f0dad4d352a22b8fd1f8a149693bdd85260ad69c0285c56eb95d

Observation ea280cc0-7b24-4930-8870-e6faefdd52ae · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Speeding up Model Loading with fastsafetensors Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.994728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.994728Z digest=sha256:8f4bbd0bd83cedb078aff5183de77d5f040edb30b6603f7c19db4be53812c7fa

Observation 17d7f8ea-1c41-42aa-b6e3-5e2dfc00f514 · outbound

This paper cites ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development.

Speeding up Model Loading with fastsafetensors ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.667602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.667602Z digest=sha256:4749fcd7ea9c6f9a8cb8c10c6b67090f84c20f52b4a05ea2bf1895296c33ec6f

Observation bcd5b2ae-657d-4b94-bf1a-6bd50878fd60 · outbound

This paper cites Welcome to DLPack’s documentation! — DLPack 0.6.0 documentation,.

Speeding up Model Loading with fastsafetensors Welcome to DLPack’s documentation! — DLPack 0.6.0 documentation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.782963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:05.733164Z digest=sha256:af40c512750f33967f7a76544b72deaf95e2215352fa85583664b65776659782

Observation 44d43244-a860-49cb-a83a-d032e69233e5 · outbound

This paper cites huggingface/text-generation-inference: Large Language Model Text Generation Inference,.

Speeding up Model Loading with fastsafetensors huggingface/text-generation-inference: Large Language Model Text Generation Inference,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.177797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:06.274780Z digest=sha256:25c47f956a736de157576b24e5c8ad8b2cf315aef7871920d563bd13af35313b

Observation 6356c0c6-767c-49d9-842a-31481fc4d10e · outbound

This paper cites vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs,.

Speeding up Model Loading with fastsafetensors vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.966023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:06.378960Z digest=sha256:bd45363b836d7dfa04cc437614e5f1b8b0acc0fd7335dd9a512b9ee173eecb51

Observation 1145e70a-359e-48e6-99fa-8aa855cee725 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs,.

Speeding up Model Loading with fastsafetensors SGLang: Efficient Execution of Structured Language Model Programs,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.449698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.449698Z digest=sha256:0ca8fbc0f22ccc89287ea9a427d06d595d8d6d3106432b012ed15c7d2c3cd664

Observation 11caba6e-a8c6-410f-a6e8-0f52e01635ec · outbound

This paper cites The Falcon Series of Open Language Models.

Speeding up Model Loading with fastsafetensors The Falcon Series of Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.100629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.100629Z digest=sha256:4127b57fcf8e2ed6521744ba54b001a14d0e06eff5975228e9edc0d818130d53

Observation 911f8b35-c10d-40d5-bf03-ef139478ad66 · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Speeding up Model Loading with fastsafetensors BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.183400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.183400Z digest=sha256:2164a166d46ce4b2854b5b8d40bbcae7e85f15e4db474c22cf02e09c5d1a0475

Observation 0471cae3-d345-484f-a461-0339ac493343 · outbound

This paper cites PEP 703 — Making the Global Interpreter Lock Optional in CPython,.

Speeding up Model Loading with fastsafetensors PEP 703 — Making the Global Interpreter Lock Optional in CPython,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.485679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:06.692824Z digest=sha256:4c3a0489943dd65274a11fa6b5b2aad1fdc84ff7b1315f113eed435de30b949c

Observation beddeef5-3e13-41b1-9354-49ff7190c4d2 · outbound

This paper cites SPIN: Seam- less Operating System Integration of Peer-to-Peer DMA Between SSDs and GPUs,.

Speeding up Model Loading with fastsafetensors SPIN: Seam- less Operating System Integration of Peer-to-Peer DMA Between SSDs and GPUs,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.306983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:06.791585Z digest=sha256:43410be59254afb71f0f5cf3bd9a6642e4c56837485169e0918fd3c853ec7be3

Observation 55590a15-344c-4e65-857b-a27540e8f8c1 · outbound

This paper cites How beneficial is peer-to-peer DMA?.

Speeding up Model Loading with fastsafetensors How beneficial is peer-to-peer DMA?

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.178224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:06.863437Z digest=sha256:ef620fabbeef515e5992dda2057271cb925f2f295d54904c0ac59afa78e429a2

Observation 3ee878dd-2212-47c6-ba76-acdff335049c · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Speeding up Model Loading with fastsafetensors SGLang: Efficient Execution of Structured Language Model Programs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.513769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.513769Z digest=sha256:84ce692a90c289d088f1350b0f6c9700649d323415e758a25e62bf359244781a

Observation bee2d220-7aeb-4748-819b-4dc82321b264 · outbound

This paper cites Rapid Data Pre- Processing with NVIDIA DALI,.

Speeding up Model Loading with fastsafetensors Rapid Data Pre- Processing with NVIDIA DALI,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.800106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:06.584427Z digest=sha256:b79ffdb39bf27d088b7cc1eeca09d575ad668e8ec2503a6bb251212c5f45bc6c

Observation b828a218-c17b-406d-b38d-c4422a76126c · outbound

This paper cites Accelerate AI and ML workloads with OCI, NVIDIA Magnum IO GPUDirect Storage, and IBM Storage Scale,.

Speeding up Model Loading with fastsafetensors Accelerate AI and ML workloads with OCI, NVIDIA Magnum IO GPUDirect Storage, and IBM Storage Scale,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.649672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:06.654325Z digest=sha256:1c8d37e73e3fd9640262eb287deae91367d6bbae5f8f269f2a8a2694e33d1538

Observation 84dd1e58-3951-47ed-a478-0fe6cd855632 · outbound

This paper cites FP8 Formats for Deep Learning.

Speeding up Model Loading with fastsafetensors FP8 Formats for Deep Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:07.253100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:07.253100Z digest=sha256:46022d67f7d48036a206904428a61e45c8a5936effb793db1f369b94aa7db7e5

Observation 151eca64-8839-4a13-bc73-a709b13deba7 · outbound

This paper cites Efficient Post-training Quantization with FP8 Formats.

Speeding up Model Loading with fastsafetensors Efficient Post-training Quantization with FP8 Formats

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:07.344437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:07.344437Z digest=sha256:4393b7d97bc7854527e2f7dd70e8670602fa4e1f958c5ebc7b047dcda744cf1e

Observation 8514445a-46b1-414c-9313-606aeec79f5a · outbound

This paper cites FP8 Quantization: The Power of the Exponent.

Speeding up Model Loading with fastsafetensors FP8 Quantization: The Power of the Exponent

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:07.427165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:07.427165Z digest=sha256:4c80490c7ed8a688e9fe9d12496bed32b460b9609878cdad2235e911636a4418

Observation 0ce17f57-773f-4db8-a49a-e450f646dbc0 · outbound

This paper cites Column Cache: Buffer Cache for Columnar Storage on HDFS,.

Speeding up Model Loading with fastsafetensors Column Cache: Buffer Cache for Columnar Storage on HDFS,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.008794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:06.941011Z digest=sha256:981c28ed292778c605bbc11533290470e9b6c2b3379eb63f766db5a2ff2b9930

Observation 4525d906-06ac-473e-a464-382a9ac9027a · outbound

This paper cites Teraheap: Reducing memory pressure in managed big data frameworks,.

Speeding up Model Loading with fastsafetensors Teraheap: Reducing memory pressure in managed big data frameworks,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:07.862265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:07.049645Z digest=sha256:c0fa5cdc7b7626d41a29bdc08ee161fd05937914417267cd7748717d4162b943

Observation fc358a47-a979-4eb3-9fd4-c52b5fc52e0b · outbound

This paper cites Accelerating multilingual applications with in-memory array sharing,.

Speeding up Model Loading with fastsafetensors Accelerating multilingual applications with in-memory array sharing,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:07.752961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:07.153431Z digest=sha256:2ec676b02c7ea006fc96dd2031efdc96020fff306d81c505511a5cf1b49ace9f

Observation 5d4bbd67-4c41-4ccf-be99-f739761dd62b · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Speeding up Model Loading with fastsafetensors PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.511142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.511142Z digest=sha256:2faf260c1143629356027ae187eec0737f0ef7b6e10baf62626fe0893b0d0a5c

Observation 465ebe04-efce-4d36-a082-c9493210a02f · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence.

Speeding up Model Loading with fastsafetensors Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.032742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.032742Z digest=sha256:7778db5c1c0d0aa7b5e5efb49d30470df95aca504c6773361a5b8318a5894a74

Observation 842d19cb-7c3f-4cbd-a9ff-095b167ac77c · outbound

This paper cites Available: https://docs .nvidia.com/gpudirect-storage/ design-guide/index.html.

Speeding up Model Loading with fastsafetensors Available: https://docs .nvidia.com/gpudirect-storage/ design-guide/index.html

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.368234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:58:05.911032Z digest=sha256:5eab3b84f8862de4f5680563ae10c622ea64ef56dda5e5d703623fc5e2af8494

Pith citing papers

Observation 42ecb65a-73c0-4442-bca7-7388a2662378 · inbound

RTP-LLM: High-Performance Alibaba LLM Inference Engine cites this paper.

RTP-LLM: High-Performance Alibaba LLM Inference Engine Speeding up Model Loading with fastsafetensors

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:52:49.146323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T23:52:40.763228Z digest=sha256:8818e81953b4ae72143c6cb9ec69e8d8d68c0b0745f1f29300e63397dd3622ea

Observation 7d5d1c20-ccf9-4905-a4bf-701289fb01aa · inbound

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing cites this paper.

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing Speeding up Model Loading with fastsafetensors

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-26T22:10:09.429560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T06:39:04.161585Z digest=sha256:d8791947fce7910178793fcfd5ec734093c2e5e18bfa0532a89ffcb0c610cbd7

Observation c7989ea8-90ee-4b21-9d22-1c174033d54c · inbound

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing cites this paper.

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing Speeding up Model Loading with fastsafetensors

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T07:20:36.786149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:20:36.786149Z digest=sha256:a28304df579b50be2081b60503596d448485ec7252f9788da7cc6af2a23b4c3c

Observation 28bcb40f-0c63-43a1-96ad-41b4b97a7e89 · inbound

InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata cites this paper.

InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata Speeding up Model Loading with fastsafetensors

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T13:52:43.151100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:52:43.151100Z digest=sha256:6b4e6fdfbf2eedf8285854c314469c51138c985e22bafe20823072fb21596123