Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:07.427165Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 4 inbound Pith citation observations for arXiv:2505.23072.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:07.427165Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T13:52:43.151100Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
47 of 47 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 1eb807ef-31d1-4c1d-9bdf-4ecf60c58ed7 · outbound
Speeding up Model Loading with fastsafetensors Introducing ChatGPT,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5b86204-bff4-4e3f-ac2b-9f89a0b3f10e · outbound
Speeding up Model Loading with fastsafetensors Introducing Gemini: our largest and most capable AI model,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a68980f-964e-44a0-9507-846aeffe203f · outbound
Speeding up Model Loading with fastsafetensors Granite Code Models: A Family of Open Foundation Models for Code Intelligence,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25defce5-5c40-4fd9-a40e-73806d7154ac · outbound
Speeding up Model Loading with fastsafetensors The Shift from Models to Compound AI Systems,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87dc08d5-980d-4bd8-b7d9-83a130d7e7d3 · outbound
Speeding up Model Loading with fastsafetensors FlashAttention: Fast and memory-efficient exact attention with IO-awareness,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd73d497-5dc8-42c4-9eb6-1b2a8d87bdc7 · outbound
Speeding up Model Loading with fastsafetensors FlashAttention-2: Faster attention with better parallelism and work partitioning,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12cc850f-9c15-4005-87b1-9ae7c7d25ae2 · outbound
Speeding up Model Loading with fastsafetensors Efficient Memory Management for Large Language Model Serving with PagedAttention,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae9808a8-4a8a-421d-9b8b-1cd773232cd5 · outbound
Speeding up Model Loading with fastsafetensors Orca: A Distributed Serving System for Transformer-Based Generative Models,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e30d29f1-ab5e-4cdb-9cc2-d96cb0e427d8 · outbound
Speeding up Model Loading with fastsafetensors Accelerating Production LLMs with Combined Token/Embedding Speculators
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e64cecdc-7e37-422b-9e32-862ec995d3f4 · outbound
Speeding up Model Loading with fastsafetensors DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2718510-a082-46bb-aa53-00fac5a49f0d · outbound
Speeding up Model Loading with fastsafetensors Taming throughput-latency tradeoff in llm inference with sarathi-serve,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 910f4ac6-e5ae-4023-9856-08bdf6a695fc · outbound
Speeding up Model Loading with fastsafetensors Decrease PyTorch Model Load Times with CoreWeave’s Tensorizer,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cda45ea4-2e7e-4ab7-952f-e4009f79e25f · outbound
Speeding up Model Loading with fastsafetensors ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b1adefe-1e6a-4e8c-8fe1-fb2f309c00e3 · outbound
Speeding up Model Loading with fastsafetensors Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89ee66e9-bbba-4b12-ab18-81eab9f57679 · outbound
Speeding up Model Loading with fastsafetensors Safetensors,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc3dc6c4-e620-4184-8408-f68174e86e7b · outbound
Speeding up Model Loading with fastsafetensors Models – Hugging Face,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8718fe95-bf6e-4e3d-8bfd-c7ac573d3059 · outbound
Speeding up Model Loading with fastsafetensors pickle — Python object Serialization,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4160a11-7b76-4e06-ac30-f8ef8dab91ce · outbound
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5df08d8-f0c5-4082-a97e-39e272121d0e · outbound
Speeding up Model Loading with fastsafetensors Available: https://docs .python.org/3/library/pickle.html
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eac066d1-55ad-4914-9134-18f059fbd4c8 · outbound
Speeding up Model Loading with fastsafetensors ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c0e8d69-88b5-4a1d-9c5e-c55ee495787a · outbound
Speeding up Model Loading with fastsafetensors TensorFlow: A System for Large-Scale Machine Learning,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f87c9321-eb6f-4360-a83c-21e94e87b89e · outbound
Speeding up Model Loading with fastsafetensors ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da9d9610-51ee-4e8f-8d8d-fd5c57aa3058 · outbound
Speeding up Model Loading with fastsafetensors PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5fdf0db-d9a9-45f0-a3a6-9d4f12f43a1b · outbound
Speeding up Model Loading with fastsafetensors NVIDIA Magnum IO GPUDirect Storage Design Guide,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea280cc0-7b24-4930-8870-e6faefdd52ae · outbound
Speeding up Model Loading with fastsafetensors Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17d7f8ea-1c41-42aa-b6e3-5e2dfc00f514 · outbound
Speeding up Model Loading with fastsafetensors ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcd5b2ae-657d-4b94-bf1a-6bd50878fd60 · outbound
Speeding up Model Loading with fastsafetensors Welcome to DLPack’s documentation! — DLPack 0.6.0 documentation,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44d43244-a860-49cb-a83a-d032e69233e5 · outbound
Speeding up Model Loading with fastsafetensors huggingface/text-generation-inference: Large Language Model Text Generation Inference,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6356c0c6-767c-49d9-842a-31481fc4d10e · outbound
Speeding up Model Loading with fastsafetensors vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1145e70a-359e-48e6-99fa-8aa855cee725 · outbound
Speeding up Model Loading with fastsafetensors SGLang: Efficient Execution of Structured Language Model Programs,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11caba6e-a8c6-410f-a6e8-0f52e01635ec · outbound
Speeding up Model Loading with fastsafetensors The Falcon Series of Open Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 911f8b35-c10d-40d5-bf03-ef139478ad66 · outbound
Speeding up Model Loading with fastsafetensors BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0471cae3-d345-484f-a461-0339ac493343 · outbound
Speeding up Model Loading with fastsafetensors PEP 703 — Making the Global Interpreter Lock Optional in CPython,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation beddeef5-3e13-41b1-9354-49ff7190c4d2 · outbound
Speeding up Model Loading with fastsafetensors SPIN: Seam- less Operating System Integration of Peer-to-Peer DMA Between SSDs and GPUs,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55590a15-344c-4e65-857b-a27540e8f8c1 · outbound
Speeding up Model Loading with fastsafetensors How beneficial is peer-to-peer DMA?
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ee878dd-2212-47c6-ba76-acdff335049c · outbound
Speeding up Model Loading with fastsafetensors SGLang: Efficient Execution of Structured Language Model Programs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee2d220-7aeb-4748-819b-4dc82321b264 · outbound
Speeding up Model Loading with fastsafetensors Rapid Data Pre- Processing with NVIDIA DALI,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b828a218-c17b-406d-b38d-c4422a76126c · outbound
Speeding up Model Loading with fastsafetensors Accelerate AI and ML workloads with OCI, NVIDIA Magnum IO GPUDirect Storage, and IBM Storage Scale,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84dd1e58-3951-47ed-a478-0fe6cd855632 · outbound
Speeding up Model Loading with fastsafetensors FP8 Formats for Deep Learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 151eca64-8839-4a13-bc73-a709b13deba7 · outbound
Speeding up Model Loading with fastsafetensors Efficient Post-training Quantization with FP8 Formats
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8514445a-46b1-414c-9313-606aeec79f5a · outbound
Speeding up Model Loading with fastsafetensors FP8 Quantization: The Power of the Exponent
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ce17f57-773f-4db8-a49a-e450f646dbc0 · outbound
Speeding up Model Loading with fastsafetensors Column Cache: Buffer Cache for Columnar Storage on HDFS,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4525d906-06ac-473e-a464-382a9ac9027a · outbound
Speeding up Model Loading with fastsafetensors Teraheap: Reducing memory pressure in managed big data frameworks,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc358a47-a979-4eb3-9fd4-c52b5fc52e0b · outbound
Speeding up Model Loading with fastsafetensors Accelerating multilingual applications with in-memory array sharing,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d4bbd67-4c41-4ccf-be99-f739761dd62b · outbound
Speeding up Model Loading with fastsafetensors PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 465ebe04-efce-4d36-a082-c9493210a02f · outbound
Speeding up Model Loading with fastsafetensors Granite Code Models: A Family of Open Foundation Models for Code Intelligence
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 842d19cb-7c3f-4cbd-a9ff-095b167ac77c · outbound
Speeding up Model Loading with fastsafetensors Available: https://docs .nvidia.com/gpudirect-storage/ design-guide/index.html
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42ecb65a-73c0-4442-bca7-7388a2662378 · inbound
RTP-LLM: High-Performance Alibaba LLM Inference Engine Speeding up Model Loading with fastsafetensors
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d5d1c20-ccf9-4905-a4bf-701289fb01aa · inbound
The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing Speeding up Model Loading with fastsafetensors
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c7989ea8-90ee-4b21-9d22-1c174033d54c · inbound
The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing Speeding up Model Loading with fastsafetensors
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28bcb40f-0c63-43a1-96ad-41b4b97a7e89 · inbound
InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata Speeding up Model Loading with fastsafetensors
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.