Pith. sign in

Paper Citation Record · LEDGER

Speeding up Model Loading with fastsafetensors

As of 8 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 4 inbound Pith citation observations for arXiv:2505.23072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23072 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:07.427165Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:52:43.151100Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1eb807ef-31d1-4c1d-9bdf-4ecf60c58ed7 · outbound

This paper cites Introducing ChatGPT,.

Speeding up Model Loading with fastsafetensors Introducing ChatGPT,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.656524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:03.832875Z digest=sha256:0fa9b59e2b4d22d2a26c28595b8d36e14064971bb07531734fae90f25e11114b

Observation f5b86204-bff4-4e3f-ac2b-9f89a0b3f10e · outbound

This paper cites Introducing Gemini: our largest and most capable AI model,.

Speeding up Model Loading with fastsafetensors Introducing Gemini: our largest and most capable AI model,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.492922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:03.902873Z digest=sha256:e3b199e06b741ea6c03c35112984912370b34b61e9892b25c10355c7706871e0

Observation 9a68980f-964e-44a0-9507-846aeffe203f · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence,.

Speeding up Model Loading with fastsafetensors Granite Code Models: A Family of Open Foundation Models for Code Intelligence,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.366624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:03.973943Z digest=sha256:51ac07f7afaff2dcbd0146596cca80a728d8ac22a9e7d2efac30a667289c36f5

Observation 25defce5-5c40-4fd9-a40e-73806d7154ac · outbound

This paper cites The Shift from Models to Compound AI Systems,.

Speeding up Model Loading with fastsafetensors The Shift from Models to Compound AI Systems,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.176639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:04.077312Z digest=sha256:1e6a487e5941fec7acf0138f93b1e2ccc8fd9a657682d88c4a9015376c7481e0

Observation 87dc08d5-980d-4bd8-b7d9-83a130d7e7d3 · outbound

This paper cites FlashAttention: Fast and memory-efficient exact attention with IO-awareness,.

Speeding up Model Loading with fastsafetensors FlashAttention: Fast and memory-efficient exact attention with IO-awareness,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:12.059280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:04.145099Z digest=sha256:ad801247c68f0a7272f2d1731232edbf52a99fa23c47e619badc1dbed2d59e0f

Observation fd73d497-5dc8-42c4-9eb6-1b2a8d87bdc7 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning,.

Speeding up Model Loading with fastsafetensors FlashAttention-2: Faster attention with better parallelism and work partitioning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.895787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:04.214691Z digest=sha256:b7f06233364295e777f37e85868223a4ee937f5bb1b553badc09e222d53bde6c

Observation 12cc850f-9c15-4005-87b1-9ae7c7d25ae2 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention,.

Speeding up Model Loading with fastsafetensors Efficient Memory Management for Large Language Model Serving with PagedAttention,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.758834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:04.304079Z digest=sha256:76da8c21442f3bf2d5a48ee60bd77445523b7f8e7d8e5a6fc7cfbe0d0e6bbb81

Observation ae9808a8-4a8a-421d-9b8b-1cd773232cd5 · outbound

This paper cites Orca: A Distributed Serving System for Transformer-Based Generative Models,.

Speeding up Model Loading with fastsafetensors Orca: A Distributed Serving System for Transformer-Based Generative Models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.611414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:04.350083Z digest=sha256:15aa084223374276576aafd859edb7f5a0ef49740c65a8aea9a84badb0749d20

Observation e30d29f1-ab5e-4cdb-9cc2-d96cb0e427d8 · outbound

This paper cites Accelerating Production LLMs with Combined Token/Embedding Speculators.

Speeding up Model Loading with fastsafetensors Accelerating Production LLMs with Combined Token/Embedding Speculators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.425862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.425862Z digest=sha256:fc4174164d1b611c5b9f0c5c24b4d31f06d408db9612560a50a85448f7c3590b

Observation e64cecdc-7e37-422b-9e32-862ec995d3f4 · outbound

This paper cites DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving.

Speeding up Model Loading with fastsafetensors DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.496281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.496281Z digest=sha256:852a159a11a4658f888d66f6200826eaf23a4be9cd991b1e322567327f9c1d35

Observation c2718510-a082-46bb-aa53-00fac5a49f0d · outbound

This paper cites Taming throughput-latency tradeoff in llm inference with sarathi-serve,.

Speeding up Model Loading with fastsafetensors Taming throughput-latency tradeoff in llm inference with sarathi-serve,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.472887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:04.561087Z digest=sha256:d5ea6f2bfe91566d325f69cd81f713b83a18f2559a4b6878750e5e2b23399936

Observation 910f4ac6-e5ae-4023-9856-08bdf6a695fc · outbound

This paper cites Decrease PyTorch Model Load Times with CoreWeave’s Tensorizer,.

Speeding up Model Loading with fastsafetensors Decrease PyTorch Model Load Times with CoreWeave’s Tensorizer,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.309831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:04.639769Z digest=sha256:6538aa87fb2cc0df532fd31d54bf902d48f4ea1a750e5383cf24365e61024873

Observation cda45ea4-2e7e-4ab7-952f-e4009f79e25f · outbound

This paper cites ServerlessLLM: Low-Latency Serverless Inference for Large Language Models.

Speeding up Model Loading with fastsafetensors ServerlessLLM: Low-Latency Serverless Inference for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.729125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.729125Z digest=sha256:52b604b399b3e0144ea1db16f72eaa42c0eae0da00d58c64dbf8a89a57c9a2aa

Observation 2b1adefe-1e6a-4e8c-8fe1-fb2f309c00e3 · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

Speeding up Model Loading with fastsafetensors Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.791744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.791744Z digest=sha256:43bc62804ea4cacc4b6f14b124542d2de88fd108a900d37b8c98b0f2f7bca9b3

Observation 89ee66e9-bbba-4b12-ab18-81eab9f57679 · outbound

This paper cites Safetensors,.

Speeding up Model Loading with fastsafetensors Safetensors,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:11.115727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:04.876060Z digest=sha256:206086e5055c2b595a1573067093f8e30651285edb768920992f859f079438e2

Observation bc3dc6c4-e620-4184-8408-f68174e86e7b · outbound

This paper cites Models – Hugging Face,.

Speeding up Model Loading with fastsafetensors Models – Hugging Face,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.934865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:04.943400Z digest=sha256:9f17193907f415d8b1768621df59286a8efc6521346f48a938ed38f21febe2b9

Observation 8718fe95-bf6e-4e3d-8bfd-c7ac573d3059 · outbound

This paper cites pickle — Python object Serialization,.

Speeding up Model Loading with fastsafetensors pickle — Python object Serialization,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.697345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:05.012673Z digest=sha256:28fac71a287053ef8dfa22e006027bb34a6bd4aa9195598022b8f3a824623a22

Observation a4160a11-7b76-4e06-ac30-f8ef8dab91ce · outbound

This paper cites Pytorch,.

Speeding up Model Loading with fastsafetensors Pytorch,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.287909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:05.160118Z digest=sha256:182e17d9088bd05aeba76942f3c1cd453e280d2d6247c0947e5cc037ba0f4f02

Observation d5df08d8-f0c5-4082-a97e-39e272121d0e · outbound

This paper cites Available: https://docs .python.org/3/library/pickle.html.

Speeding up Model Loading with fastsafetensors Available: https://docs .python.org/3/library/pickle.html

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.472708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:05.086172Z digest=sha256:45803073ee94d031fe8389e8edfd12f0597cb03e99be9444c00bcfe38278cfd9

Observation eac066d1-55ad-4914-9134-18f059fbd4c8 · outbound

This paper cites ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning.

Speeding up Model Loading with fastsafetensors ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.322950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.322950Z digest=sha256:4b56f6d05fb2bfb7cdc9f6dd9e478ca3409c9e712bacbe81b9b04d7aaa23adf6

Observation 1c0e8d69-88b5-4a1d-9c5e-c55ee495787a · outbound

This paper cites TensorFlow: A System for Large-Scale Machine Learning,.

Speeding up Model Loading with fastsafetensors TensorFlow: A System for Large-Scale Machine Learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:10.162396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:05.256431Z digest=sha256:9842b32c3ab350a5136dda401b5ec5b383d07f61d818adcd9df08b9421682d1c

Observation f87c9321-eb6f-4360-a83c-21e94e87b89e · outbound

This paper cites ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development,.

Speeding up Model Loading with fastsafetensors ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.974543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:05.583944Z digest=sha256:b69fcf06bffc9b71d4a6a2dac4475936483d9d6fbd9e937f58270cc85d7f4554

Observation da9d9610-51ee-4e8f-8d8d-fd5c57aa3058 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel,.

Speeding up Model Loading with fastsafetensors PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.415353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.415353Z digest=sha256:307cb0bb5e1888c2b8877b219373b1d5e0efbcff3a37ffcf78fd8de75ead3dee

Observation f5fdf0db-d9a9-45f0-a3a6-9d4f12f43a1b · outbound

This paper cites NVIDIA Magnum IO GPUDirect Storage Design Guide,.

Speeding up Model Loading with fastsafetensors NVIDIA Magnum IO GPUDirect Storage Design Guide,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.563051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:05.795124Z digest=sha256:97f04611b1efd262f7bd742b729b498fc3807d9a0084a0e5d81719b10bc5c1ae

Observation ea280cc0-7b24-4930-8870-e6faefdd52ae · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Speeding up Model Loading with fastsafetensors Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.994728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.994728Z digest=sha256:b989171803e6fef1d6f70a636e0979b7817eb7f429f9f5c32753439682d3c4ab

Observation 17d7f8ea-1c41-42aa-b6e3-5e2dfc00f514 · outbound

This paper cites ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development.

Speeding up Model Loading with fastsafetensors ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.667602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.667602Z digest=sha256:7b137d47b2bf467327a48ec1073b9dcaf024dc5190c68cc5f8aad6e3b1af25b6

Observation bcd5b2ae-657d-4b94-bf1a-6bd50878fd60 · outbound

This paper cites Welcome to DLPack’s documentation! — DLPack 0.6.0 documentation,.

Speeding up Model Loading with fastsafetensors Welcome to DLPack’s documentation! — DLPack 0.6.0 documentation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.782963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:05.733164Z digest=sha256:c3828230a480518ad7d97123fbc4888bdb9c4e1fc91cc4220e172d82e07f4276

Observation 44d43244-a860-49cb-a83a-d032e69233e5 · outbound

This paper cites huggingface/text-generation-inference: Large Language Model Text Generation Inference,.

Speeding up Model Loading with fastsafetensors huggingface/text-generation-inference: Large Language Model Text Generation Inference,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.177797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:06.274780Z digest=sha256:13334753c4206571aa6948dd283ffac3006b95022d3a1c4596fd43919c452631

Observation 6356c0c6-767c-49d9-842a-31481fc4d10e · outbound

This paper cites vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs,.

Speeding up Model Loading with fastsafetensors vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.966023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:06.378960Z digest=sha256:703f0cd1ed9f73bbf4a48e980885157cdfa580351303eefacc7c106ed8965edf

Observation 1145e70a-359e-48e6-99fa-8aa855cee725 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs,.

Speeding up Model Loading with fastsafetensors SGLang: Efficient Execution of Structured Language Model Programs,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.449698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.449698Z digest=sha256:ab9926cadf00680f2378172640d38a2206fcf813e84102b075f86151cc3f29ce

Observation 11caba6e-a8c6-410f-a6e8-0f52e01635ec · outbound

This paper cites The Falcon Series of Open Language Models.

Speeding up Model Loading with fastsafetensors The Falcon Series of Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.100629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.100629Z digest=sha256:90316b39a5b9b1370a74a412fa81e007a2b7413a09cc8c8549b6f8d501584dde

Observation 911f8b35-c10d-40d5-bf03-ef139478ad66 · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Speeding up Model Loading with fastsafetensors BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.183400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.183400Z digest=sha256:6e92665d35202131badb1de0b55c1bc79cf2b4c0d386aa8a74169328319d9e9c

Observation 0471cae3-d345-484f-a461-0339ac493343 · outbound

This paper cites PEP 703 — Making the Global Interpreter Lock Optional in CPython,.

Speeding up Model Loading with fastsafetensors PEP 703 — Making the Global Interpreter Lock Optional in CPython,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.485679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:06.692824Z digest=sha256:1ec7f7cd9646a4304f41c2ecf7524486ebe7e76d70616619cdbd5cbbbefa55c8

Observation beddeef5-3e13-41b1-9354-49ff7190c4d2 · outbound

This paper cites SPIN: Seam- less Operating System Integration of Peer-to-Peer DMA Between SSDs and GPUs,.

Speeding up Model Loading with fastsafetensors SPIN: Seam- less Operating System Integration of Peer-to-Peer DMA Between SSDs and GPUs,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.306983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:06.791585Z digest=sha256:e1d3f9c747367ef3cd389ea54ea023c4708c6f33c35da2d74fd117e6244cf777

Observation 55590a15-344c-4e65-857b-a27540e8f8c1 · outbound

This paper cites How beneficial is peer-to-peer DMA?.

Speeding up Model Loading with fastsafetensors How beneficial is peer-to-peer DMA?

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.178224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:06.863437Z digest=sha256:3a8cf7cc2dc5d4e7652c015a5f7fd8198bf60add61f17224510c958d8943f4ac

Observation 3ee878dd-2212-47c6-ba76-acdff335049c · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Speeding up Model Loading with fastsafetensors SGLang: Efficient Execution of Structured Language Model Programs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:06.513769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:06.513769Z digest=sha256:48a5333f26df6ca5f5f12b249137c8afdba1c247c7e996cf75a6d017947b534d

Observation bee2d220-7aeb-4748-819b-4dc82321b264 · outbound

This paper cites Rapid Data Pre- Processing with NVIDIA DALI,.

Speeding up Model Loading with fastsafetensors Rapid Data Pre- Processing with NVIDIA DALI,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.800106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:06.584427Z digest=sha256:fe3573fcf50b70561c28638f0776743d829013958a33abfe6b384383bbb531a2

Observation b828a218-c17b-406d-b38d-c4422a76126c · outbound

This paper cites Accelerate AI and ML workloads with OCI, NVIDIA Magnum IO GPUDirect Storage, and IBM Storage Scale,.

Speeding up Model Loading with fastsafetensors Accelerate AI and ML workloads with OCI, NVIDIA Magnum IO GPUDirect Storage, and IBM Storage Scale,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.649672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:06.654325Z digest=sha256:760ab05a7c3932acc56b9efad03272b5c3bc88376c6a3bd539ba469f972d4bdf

Observation 84dd1e58-3951-47ed-a478-0fe6cd855632 · outbound

This paper cites FP8 Formats for Deep Learning.

Speeding up Model Loading with fastsafetensors FP8 Formats for Deep Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:07.253100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:07.253100Z digest=sha256:42b1c8a8a1f5b480bd6424b6a76bf1145371ae9681b2b7b104b8e2c1477bfd9b

Observation 151eca64-8839-4a13-bc73-a709b13deba7 · outbound

This paper cites Efficient Post-training Quantization with FP8 Formats.

Speeding up Model Loading with fastsafetensors Efficient Post-training Quantization with FP8 Formats

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:07.344437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:07.344437Z digest=sha256:c78f275cc2e7305bd8813596db8f9d5645c36b195b43620459312812e6a4c490

Observation 8514445a-46b1-414c-9313-606aeec79f5a · outbound

This paper cites FP8 Quantization: The Power of the Exponent.

Speeding up Model Loading with fastsafetensors FP8 Quantization: The Power of the Exponent

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:07.427165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:07.427165Z digest=sha256:7d513af2f488358c33661ae164d9e068fee9afbe951363b85c16ea9dffe450a7

Observation 0ce17f57-773f-4db8-a49a-e450f646dbc0 · outbound

This paper cites Column Cache: Buffer Cache for Columnar Storage on HDFS,.

Speeding up Model Loading with fastsafetensors Column Cache: Buffer Cache for Columnar Storage on HDFS,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:08.008794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:06.941011Z digest=sha256:5e9307eb7088a7b39ee7abe5053c3c6984d47d198a39ef2d2c8c9a53912213bc

Observation 4525d906-06ac-473e-a464-382a9ac9027a · outbound

This paper cites Teraheap: Reducing memory pressure in managed big data frameworks,.

Speeding up Model Loading with fastsafetensors Teraheap: Reducing memory pressure in managed big data frameworks,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:07.862265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:07.049645Z digest=sha256:7bbea44313e9f93ef350bdf499f40cb2a00c506e5f836c611564efcf17a6bd62

Observation fc358a47-a979-4eb3-9fd4-c52b5fc52e0b · outbound

This paper cites Accelerating multilingual applications with in-memory array sharing,.

Speeding up Model Loading with fastsafetensors Accelerating multilingual applications with in-memory array sharing,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:07.752961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:07.153431Z digest=sha256:e2042acad7021fb20a691e82bc733d80617d3605ef035f52e3bf75012584c77a

Observation 5d4bbd67-4c41-4ccf-be99-f739761dd62b · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Speeding up Model Loading with fastsafetensors PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:05.511142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:05.511142Z digest=sha256:d974906431420f39b9920868dab110b0a8fdb7781804f8ca3ea8b41d6e1d9581

Observation 465ebe04-efce-4d36-a082-c9493210a02f · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence.

Speeding up Model Loading with fastsafetensors Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:04.032742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:04.032742Z digest=sha256:f9d868589a08bd3b28b20858f19ad2b7783ffee46311dcaffcf584b2db0bcefb

Observation 842d19cb-7c3f-4cbd-a9ff-095b167ac77c · outbound

This paper cites Available: https://docs .nvidia.com/gpudirect-storage/ design-guide/index.html.

Speeding up Model Loading with fastsafetensors Available: https://docs .nvidia.com/gpudirect-storage/ design-guide/index.html

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:58:09.368234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:05.911032Z digest=sha256:4ec0565c779dd95a0cd92bfb07c8ba41f6523559c683a677929220cd347e2ec5

Pith citing papers

Observation 42ecb65a-73c0-4442-bca7-7388a2662378 · inbound

RTP-LLM: High-Performance Alibaba LLM Inference Engine cites this paper.

RTP-LLM: High-Performance Alibaba LLM Inference Engine Speeding up Model Loading with fastsafetensors

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:52:49.146323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:52:40.763228Z digest=sha256:072f4f2997adccb6ea89c9b96a7b0ec3c5f926ad1475334d9b7986d9610128c8

Observation 7d5d1c20-ccf9-4905-a4bf-701289fb01aa · inbound

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing cites this paper.

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing Speeding up Model Loading with fastsafetensors

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-26T22:10:09.429560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T06:39:04.161585Z digest=sha256:f41e89dfd8853f4f10d8fdb4c5c55f5d309953f90aa166f83e3a2ec55e523b3b

Observation c7989ea8-90ee-4b21-9d22-1c174033d54c · inbound

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing cites this paper.

The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing Speeding up Model Loading with fastsafetensors

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T07:20:36.786149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:20:36.786149Z digest=sha256:62c3973ef35dfae58a238ea0f7138663f48e8a6bdc5a51db6aba4dd634c74104

Observation 28bcb40f-0c63-43a1-96ad-41b4b97a7e89 · inbound

InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata cites this paper.

InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata Speeding up Model Loading with fastsafetensors

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T13:52:43.151100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:52:43.151100Z digest=sha256:f40eb2f65ce97722a06166c63cea596c65095e10641514fd16354c3730d2824a