Pith. sign in

Paper Citation Record · LEDGER

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers

As of 20 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2601.16956.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.16956 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T08:30:57.861433Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-25T05:52:30.402478Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:55:24.657761Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d605c9a5-7744-434a-be3f-d4ee79ae7a19 · outbound

This paper cites Towards an AI co-scientist.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Towards an AI co-scientist

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.044625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.044625Z digest=sha256:38887f922af0be7b78d139818e78720d8998aa21487539c9ebb390869f1ae5af

Observation 408b28a6-ed92-4010-a027-ecea474b4c25 · outbound

This paper cites Scaling llama 3 training with efficient parallelism strategies,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Scaling llama 3 training with efficient parallelism strategies,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.125489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.125489Z digest=sha256:e70870a6d7ebfd5615ce7cb75e8b7fd54e51c14c8b0a795221bd387c465f8e87

Observation 84f44455-5f6d-4334-9f65-ed77f287944e · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers BLOOM: A 176B-Parameter Open-Access Multilingual Language Model,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.200566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.200566Z digest=sha256:679f24ff113d2965d31b3ffd496b30e479f302d4c6f53b96f6b1014642023289

Observation a5a9ef2c-86dc-4768-b88c-c86f9ba68e29 · outbound

This paper cites DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.281621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.281621Z digest=sha256:d46daaacb284115b7eb6eda3790877b561d4ff5996d80b1707383f47c1676d85

Observation 92d54919-66cd-4825-a8e5-b7a6181c66e8 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.350781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.350781Z digest=sha256:1837c77f5311cb6ce36b194391b31f6477d735102f39b316049b09b2f8b14b2b

Observation 3fa05e02-df14-4849-bd47-c134ef97c201 · outbound

This paper cites Robust llm training infrastructure at bytedance,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Robust llm training infrastructure at bytedance,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.395602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.395602Z digest=sha256:6a0e16e803643c1af59b2cd068f27385266b8f9099c33bdd768430ba05a543e8

Observation b2aed746-885f-4953-a74d-4220c88282ac · outbound

This paper cites Unicron: Economizing self-healing llm training at scale,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Unicron: Economizing self-healing llm training at scale,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.469818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.469818Z digest=sha256:51bb2c0684767834a3cae5cea6bbc4f44e64b7c37559136d9a64db9464115071

Observation c63a0b42-d534-471a-9761-203a00acd1b7 · outbound

This paper cites Spike No More: Stabilizing the Pre-training of Large Language Models.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.527546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.527546Z digest=sha256:6914797928bb54e6d35519635e587c646e21fd1ed97ddf68aba4ff180c3441fc

Observation 7fc18c1c-7f90-4ad5-86e0-a9d0a804edd4 · outbound

This paper cites FastPersist: Accelerating Model Checkpointing in Deep Learning.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers FastPersist: Accelerating Model Checkpointing in Deep Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.610643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.610643Z digest=sha256:b0ef022e9dadb2b24c1752ff8347c664ce272cab9f22a46a5b6c286012711839

Observation c9f1de20-8b8f-4fbe-b57d-225d92fc3d7a · outbound

This paper cites Datastates-llm: Lazy asynchronous checkpointing for large language models,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Datastates-llm: Lazy asynchronous checkpointing for large language models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.703124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.703124Z digest=sha256:ac96f6eef0445fafd59ec518bd643d3bdbcddfd8c2b1e9d863de98a1ea986131

Observation 51f573d2-766b-4c3b-add4-9a99b71594b4 · outbound

This paper cites Welcome to the torchsnapshot documentation,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Welcome to the torchsnapshot documentation,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.778591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.778591Z digest=sha256:46b95cf60fe2dac87c19662763d87ac4039300e483fea8c8e5a2c061c6b0d62b

Observation d28926fd-ef96-4cd1-9080-94b2ae94240b · outbound

This paper cites CheckFreq: Frequent, Fine-Grained DNN checkpointing,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers CheckFreq: Frequent, Fine-Grained DNN checkpointing,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.871968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.871968Z digest=sha256:16fc66ad9ba1b8e9b1553f6ef3fc72d3a57b2f97e15946708f058281fe480cb2

Observation 82835cb8-493f-4363-8f87-5484abfd7394 · outbound

This paper cites Gemini: Fast failure recovery in distributed training with in-memory checkpoints,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Gemini: Fast failure recovery in distributed training with in-memory checkpoints,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.919233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.919233Z digest=sha256:43109730f87d2693ab9e5ed56e633a37e84ca6c56c668d1843c4e8f7bb820749

Observation 3c97120e-1d6f-4131-aaff-86f91a5f9f31 · outbound

This paper cites DeepFreeze: Towards Scalable Asynchronous Checkpointing of Deep Learning Models,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers DeepFreeze: Towards Scalable Asynchronous Checkpointing of Deep Learning Models,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.983736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.983736Z digest=sha256:03de2de4b81eb9972e33fd3272c9cde6d727ebf1f05a15e0e742befc56d5c3d7

Observation 8215de6d-c59d-4afb-96cb-00aae19ec539 · outbound

This paper cites Reliable and efficient in-memory fault tolerance of large language model pretraining,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Reliable and efficient in-memory fault tolerance of large language model pretraining,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.040591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.040591Z digest=sha256:b3859d8d1886252f93327d3bd6541ef810c5b664df616c8d44e6e4a69d45882d

Observation ae48085a-f023-4995-a895-81031b58143f · outbound

This paper cites Optimize Checkpoint Performance for Large Models - Azure Machine Learning,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Optimize Checkpoint Performance for Large Models - Azure Machine Learning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.096554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.096554Z digest=sha256:d9c5b6781c2ce52019c85e36f1631c4cd791dd0c21f0821e6f00c7a15adc1b36

Observation 06e409a5-02f4-4eb3-8a9a-e9f613a64cea · outbound

This paper cites Zero- infinity: breaking the gpu memory wall for extreme scale deep learning,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Zero- infinity: breaking the gpu memory wall for extreme scale deep learning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.153666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.153666Z digest=sha256:f72727ea6d8e041d88bde6e089d6ec572501732c6f0be771abe5d957481fb7d3

Observation 8f83c0ba-35da-4f22-b9ca-68e46dc9ae71 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Mod- els Using Model Parallelism,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Megatron-LM: Training Multi-Billion Parameter Language Mod- els Using Model Parallelism,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.254023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.254023Z digest=sha256:d052b70278294de9ffc72f7b5c70a5ff307164ddac2139a65f91333438701b9d

Observation 2ae6e332-9d87-40fa-92a3-2d5207d3e625 · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers ZeRO: Memory Optimizations Toward Training Trillion Parameter Models,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.352544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.352544Z digest=sha256:59d443d876fb0877f436698fe586a17fdbee662cecd19c4385e199a7d61af0bf

Observation e8fde21f-36ea-4ae7-a3df-4345f259b720 · outbound

This paper cites Understanding llm checkpoint/restore i/o strategies and patterns,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Understanding llm checkpoint/restore i/o strategies and patterns,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.449326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.449326Z digest=sha256:c103fd98f439705ca9582d719c0f3fa60bdbb0631acb94fed525606cdd66041d

Observation 413dc35e-aba8-4037-849b-aba4f18523e6 · outbound

This paper cites A cost-efficient failure-tolerant scheme for distributed dnn training,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers A cost-efficient failure-tolerant scheme for distributed dnn training,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.535282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.535282Z digest=sha256:6937b66b38df4e68c3b10d69d5b8b8e1b0928898ef3d3d5344435f9269935872

Observation 347f7a01-18c5-4cd1-9588-62fd6da8892c · outbound

This paper cites Transom: An efficient fault-tolerant system for training llms,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Transom: An efficient fault-tolerant system for training llms,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.597408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.597408Z digest=sha256:d6125c2ec39fd411a642e9e09885096013981fc4a06e725dac39ae06ec6adc3d

Observation 59c1728f-cd2d-446c-83c7-39cad25b5f66 · outbound

This paper cites Berkeley lab checkpoint/restart (blcr) for linux clusters,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Berkeley lab checkpoint/restart (blcr) for linux clusters,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.656575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.656575Z digest=sha256:6da9076fdfea409e913faa2a3caee571313c18c0c2ef3529766e035ba255c5a5

Observation 9f7c1e1f-8c4d-400e-a40a-7a66bb9a2c78 · outbound

This paper cites Checuda: A checkpoint/restart tool for cuda applications,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Checuda: A checkpoint/restart tool for cuda applications,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.740816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.740816Z digest=sha256:fb02f40551cca207bb6d905fc0af00f7bac49d141b14f2d47661050f14f90dd2

Observation 937fca11-d27c-4518-966f-59980c61fac8 · outbound

This paper cites VeloC: Towards High Performance Adaptive Asynchronous Check- pointing at Large Scale,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers VeloC: Towards High Performance Adaptive Asynchronous Check- pointing at Large Scale,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.876667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.876667Z digest=sha256:e40688929ed206fe2955be689fd089621eff4e7a71e02313f1608ecd077df039

Observation 434eb340-2dce-453b-9863-493b5a62673f · outbound

This paper cites Towards Efficient Cache Allocation for High-Frequency Checkpointing,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Towards Efficient Cache Allocation for High-Frequency Checkpointing,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.011067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.011067Z digest=sha256:e0eeb68ca5424c80a2c36e8cdd88cbb196d5279af2f710f700b16dc8ee5209ad

Observation 333af9dc-f4cd-434f-b808-ab0fc767a0a6 · outbound

This paper cites GPU-Enabled Asynchronous Multi-level Checkpoint Caching and Prefetching,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers GPU-Enabled Asynchronous Multi-level Checkpoint Caching and Prefetching,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.079129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.079129Z digest=sha256:b3ea16d9643abdeb78855300ff363d0273e00607477859efa6172e28898885ea

Observation 1c341511-4424-4445-96bd-507ea5834b8c · outbound

This paper cites Checkpoint restart support for heterogeneous hpc applications,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Checkpoint restart support for heterogeneous hpc applications,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.142287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.142287Z digest=sha256:6caad42bf37707f63e42e80d54d259084870fa6a6a2df8f3ea2f323021a40be7

Observation c72bbc8a-17ef-427d-bc60-a73dc6a72564 · outbound

This paper cites Adios 2: The adaptable input output system. a framework for high-performance data management,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Adios 2: The adaptable input output system. a framework for high-performance data management,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.208517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.208517Z digest=sha256:41594d192b86c0696e77dcd6150270f0ade52a31f35fc967f532cbb28fd9e3ec

Observation e3f0cd8b-0f5c-477d-9128-36a56a691d98 · outbound

This paper cites An overview of gradient descent optimization algorithms,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers An overview of gradient descent optimization algorithms,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.276827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.276827Z digest=sha256:32caad7b88896f02911ac2b13889b4063431c38c825e4a205fd2f0c4d7cd76e1

Observation 4d951ac7-d575-42a4-9883-708b82301d30 · outbound

This paper cites Adam: A method for stochastic optimization,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Adam: A method for stochastic optimization,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.450739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.450739Z digest=sha256:40b108fed5444215ec9c17ed65e2080829b64a54d7bf3fcea195169a2f5961d3

Observation 79ac9fa6-930f-436d-9db4-42135a710bbc · outbound

This paper cites Mixed Precision Training.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Mixed Precision Training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.589590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.589590Z digest=sha256:ba3603130ea4fbee652e395f62316965aa400eb3235525c16e682d6f972a9e81

Observation f47b20f2-689e-457a-a4d7-222ee2868f08 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Adam: A Method for Stochastic Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.516419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.516419Z digest=sha256:b3b13fd8a107326010f3e7b16d1e7805a4919c3dd67c24d27ff9eb3e32f0c1a3

Observation b5726e1c-a469-4c55-9868-aa892988d9a3 · outbound

This paper cites Polaris,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Polaris,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.671452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.671452Z digest=sha256:d3711da41bc353ea12b475de67853b75f30502534f7c587255f809c78f4d1b53

Observation bffa5578-05b1-4d36-8146-29c46a95ee6b · outbound

This paper cites Asynccheckpointio– pytorch lightning,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Asynccheckpointio– pytorch lightning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.610290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.610290Z digest=sha256:2bcc1dbf2ad277aa9dbcbc208f3c52b79455c658da83a03c4dbf4ec4a7f90097

Observation 9d1dde8f-43cb-4a54-85ed-c1e81e725bc8 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Llama 2: Open Foundation and Fine-Tuned Chat Models,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.861433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.861433Z digest=sha256:857f928f53a517443e168c425a378f9eabbe3a972587effc54a6234a69d02956

Observation 94c7494d-6f77-431b-a8b3-03bfd525090c · outbound

This paper cites Lustre: Building a file system for 1000-node clusters,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Lustre: Building a file system for 1000-node clusters,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.769330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.769330Z digest=sha256:9ec6c98a1b4f12d76adb63452ae1a1584e3974c7a32ac1c76a4c95b26a0b79c2

Observation d4d5a879-a7d5-44bc-b7a3-fe72d6472b1c · outbound

This paper cites An overview of gradient descent optimization algorithms.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers An overview of gradient descent optimization algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.370717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.370717Z digest=sha256:c57e400784b143f863643d8ae7f83b80a9b7e1c508591d23118dc88e7b189196

Pith citing papers

Observation 62c009ef-6082-45fa-b036-27c0b91d5df8 · inbound

ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload cites this paper.

ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T01:14:28.007359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T01:59:28.408860Z digest=sha256:553ea18e2fe65f78d639f733d2bab9434ba0eef649da18156cf66da1df142edb

Observation 16ac29bb-b626-4a74-95e5-39a9cdfd2880 · inbound

ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload cites this paper.

ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T01:14:28.007359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:52:30.402478Z digest=sha256:b27369cfff5ca6a1ad64a14c9e4a6b0c06a68587024ad6e8ef15f8c90a641b8d