Pith. sign in

Paper Citation Record · LEDGER

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers

As of 7 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2601.16956.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.16956 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T08:30:57.861433Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-25T05:52:30.402478Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:55:24.657761Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d605c9a5-7744-434a-be3f-d4ee79ae7a19 · outbound

This paper cites Towards an AI co-scientist.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Towards an AI co-scientist

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.044625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.044625Z digest=sha256:2c09ed8b9dc42b9f2f11f5d3ea93f151287fbd45183a39ff8c7efed21622d720

Observation 408b28a6-ed92-4010-a027-ecea474b4c25 · outbound

This paper cites Scaling llama 3 training with efficient parallelism strategies,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Scaling llama 3 training with efficient parallelism strategies,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.125489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.125489Z digest=sha256:09aa402fe67721b3910c6659d6b31da61bd0640c41d332bab626917394a5b4b2

Observation 84f44455-5f6d-4334-9f65-ed77f287944e · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers BLOOM: A 176B-Parameter Open-Access Multilingual Language Model,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.200566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.200566Z digest=sha256:b5146c9507ad606f8d0406ed88404086348e41887143abe6703da96741e4c442

Observation a5a9ef2c-86dc-4768-b88c-c86f9ba68e29 · outbound

This paper cites DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.281621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.281621Z digest=sha256:39091650968c5c87522949caf08e401320760bd072d8f5f154ab7c1e3ac98d56

Observation 92d54919-66cd-4825-a8e5-b7a6181c66e8 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.350781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.350781Z digest=sha256:5cfb298050630ceccdfa35286a9897442c9b5fc525cc5c9f22699442356f94fe

Observation 3fa05e02-df14-4849-bd47-c134ef97c201 · outbound

This paper cites Robust llm training infrastructure at bytedance,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Robust llm training infrastructure at bytedance,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.395602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.395602Z digest=sha256:7bde48b79f35f1883c0ee4437f16be8c114fc639d2af4093bbc05bae4aca09cf

Observation b2aed746-885f-4953-a74d-4220c88282ac · outbound

This paper cites Unicron: Economizing self-healing llm training at scale,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Unicron: Economizing self-healing llm training at scale,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.469818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.469818Z digest=sha256:1960c35be812258c7bea15e8b6e21cc72c50cf1ab32fbb63cbd85004f0ed53f9

Observation c63a0b42-d534-471a-9761-203a00acd1b7 · outbound

This paper cites Spike No More: Stabilizing the Pre-training of Large Language Models.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.527546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.527546Z digest=sha256:6a6e76a712dbb69340066af2a7f1bb669767b37027ead385de2e4852044e14cf

Observation 7fc18c1c-7f90-4ad5-86e0-a9d0a804edd4 · outbound

This paper cites FastPersist: Accelerating Model Checkpointing in Deep Learning.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers FastPersist: Accelerating Model Checkpointing in Deep Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.610643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.610643Z digest=sha256:bb9f71a44304d582b4c8e935dba0f4b77cb4a5c19ed6d59e3bc50c8cf6f33c0e

Observation c9f1de20-8b8f-4fbe-b57d-225d92fc3d7a · outbound

This paper cites Datastates-llm: Lazy asynchronous checkpointing for large language models,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Datastates-llm: Lazy asynchronous checkpointing for large language models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.703124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.703124Z digest=sha256:dc240592292643122151e3d26fa9c374a5bb40ac60a141847122b1963c321ee3

Observation 51f573d2-766b-4c3b-add4-9a99b71594b4 · outbound

This paper cites Welcome to the torchsnapshot documentation,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Welcome to the torchsnapshot documentation,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.778591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.778591Z digest=sha256:b9d819fb31d8121e4a2421bc621b3802b1702fe8e37a2c54a43b79685d83c125

Observation d28926fd-ef96-4cd1-9080-94b2ae94240b · outbound

This paper cites CheckFreq: Frequent, Fine-Grained DNN checkpointing,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers CheckFreq: Frequent, Fine-Grained DNN checkpointing,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.871968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.871968Z digest=sha256:b52505b7ae22a9ea89f199f1d59eeb7edbdc4b006b30e272913045eae489e6c9

Observation 82835cb8-493f-4363-8f87-5484abfd7394 · outbound

This paper cites Gemini: Fast failure recovery in distributed training with in-memory checkpoints,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Gemini: Fast failure recovery in distributed training with in-memory checkpoints,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.919233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.919233Z digest=sha256:47d71ce849ed7044542bd355511bbb49d592c61a5404fdbf8727e8b5aa4b1144

Observation 3c97120e-1d6f-4131-aaff-86f91a5f9f31 · outbound

This paper cites DeepFreeze: Towards Scalable Asynchronous Checkpointing of Deep Learning Models,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers DeepFreeze: Towards Scalable Asynchronous Checkpointing of Deep Learning Models,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.983736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.983736Z digest=sha256:5cb551a428a7c8df50715e0d48f20637ee830e0cd9213e2cd22a39f699419f62

Observation 8215de6d-c59d-4afb-96cb-00aae19ec539 · outbound

This paper cites Reliable and efficient in-memory fault tolerance of large language model pretraining,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Reliable and efficient in-memory fault tolerance of large language model pretraining,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.040591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.040591Z digest=sha256:c30d96b5c71a13df13ed5209170758bbd04b6cd02cc654043b62789b3c94ddad

Observation ae48085a-f023-4995-a895-81031b58143f · outbound

This paper cites Optimize Checkpoint Performance for Large Models - Azure Machine Learning,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Optimize Checkpoint Performance for Large Models - Azure Machine Learning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.096554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.096554Z digest=sha256:be9fc540d2c957015ee8b3c97ab48489c7e2544e21becc910771e5a17348f197

Observation 06e409a5-02f4-4eb3-8a9a-e9f613a64cea · outbound

This paper cites Zero- infinity: breaking the gpu memory wall for extreme scale deep learning,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Zero- infinity: breaking the gpu memory wall for extreme scale deep learning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.153666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.153666Z digest=sha256:a7ecca3b352bae4e3ddd54e2746212073a76511cc69c82ae3a151bd129a51a34

Observation 8f83c0ba-35da-4f22-b9ca-68e46dc9ae71 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Mod- els Using Model Parallelism,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Megatron-LM: Training Multi-Billion Parameter Language Mod- els Using Model Parallelism,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.254023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.254023Z digest=sha256:80a663568bbea7ab57a27694b509a5b2019929deec44db3d1bd63bdbbfddb399

Observation 2ae6e332-9d87-40fa-92a3-2d5207d3e625 · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers ZeRO: Memory Optimizations Toward Training Trillion Parameter Models,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.352544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.352544Z digest=sha256:105126bfeffad2751a9d0062ad3051bd0063d8c9f593a7e9a337041d982ab642

Observation e8fde21f-36ea-4ae7-a3df-4345f259b720 · outbound

This paper cites Understanding llm checkpoint/restore i/o strategies and patterns,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Understanding llm checkpoint/restore i/o strategies and patterns,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.449326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.449326Z digest=sha256:1d67c45a3e0a946c3af549214de23c4bdcdb227869be187c51cd67a1b130aece

Observation 413dc35e-aba8-4037-849b-aba4f18523e6 · outbound

This paper cites A cost-efficient failure-tolerant scheme for distributed dnn training,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers A cost-efficient failure-tolerant scheme for distributed dnn training,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.535282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.535282Z digest=sha256:98f7f2d2d33a77b3c65ac0ca04398b4a6f159030dd8d3a85e7ad19fad3b7ffcc

Observation 347f7a01-18c5-4cd1-9588-62fd6da8892c · outbound

This paper cites Transom: An efficient fault-tolerant system for training llms,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Transom: An efficient fault-tolerant system for training llms,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.597408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.597408Z digest=sha256:5d845177542b2e40399c8c37ee17deca1f9ba7a564cc26e353e0ab6a1020176d

Observation 59c1728f-cd2d-446c-83c7-39cad25b5f66 · outbound

This paper cites Berkeley lab checkpoint/restart (blcr) for linux clusters,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Berkeley lab checkpoint/restart (blcr) for linux clusters,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.656575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.656575Z digest=sha256:80d561d292dcaae0c51bfd12199c7b8fe62a8d573325915cb9eba9a9e03ac092

Observation 9f7c1e1f-8c4d-400e-a40a-7a66bb9a2c78 · outbound

This paper cites Checuda: A checkpoint/restart tool for cuda applications,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Checuda: A checkpoint/restart tool for cuda applications,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.740816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.740816Z digest=sha256:e2f0ca88515c106dbf5a8849f9ede00d629649082cf12d89ceed5a10924422fc

Observation 937fca11-d27c-4518-966f-59980c61fac8 · outbound

This paper cites VeloC: Towards High Performance Adaptive Asynchronous Check- pointing at Large Scale,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers VeloC: Towards High Performance Adaptive Asynchronous Check- pointing at Large Scale,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:56.876667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:56.876667Z digest=sha256:a78eacbfa03e24cd7aef8945e434ce93c958009e2147a6884b12132006c03a26

Observation 434eb340-2dce-453b-9863-493b5a62673f · outbound

This paper cites Towards Efficient Cache Allocation for High-Frequency Checkpointing,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Towards Efficient Cache Allocation for High-Frequency Checkpointing,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.011067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.011067Z digest=sha256:90f24d774dc13c4c94ac498f3736f75a87c38ad8f20e171ddc1fe1d4ec00da33

Observation 333af9dc-f4cd-434f-b808-ab0fc767a0a6 · outbound

This paper cites GPU-Enabled Asynchronous Multi-level Checkpoint Caching and Prefetching,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers GPU-Enabled Asynchronous Multi-level Checkpoint Caching and Prefetching,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.079129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.079129Z digest=sha256:be7a8067b169687a0f732d27816131117a8ce0147dddb79f817db8e7f23eb1b9

Observation 1c341511-4424-4445-96bd-507ea5834b8c · outbound

This paper cites Checkpoint restart support for heterogeneous hpc applications,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Checkpoint restart support for heterogeneous hpc applications,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.142287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.142287Z digest=sha256:a6153edef9ced47c00ae350cb05fa7c118567275de73536d78bf85edf2655a13

Observation c72bbc8a-17ef-427d-bc60-a73dc6a72564 · outbound

This paper cites Adios 2: The adaptable input output system. a framework for high-performance data management,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Adios 2: The adaptable input output system. a framework for high-performance data management,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.208517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.208517Z digest=sha256:8c1fcf3c35c7e096053516f8d4af1ee0c6b1ba21ee74f8578695f2c3de16fd85

Observation e3f0cd8b-0f5c-477d-9128-36a56a691d98 · outbound

This paper cites An overview of gradient descent optimization algorithms,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers An overview of gradient descent optimization algorithms,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.276827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.276827Z digest=sha256:dd668727d531ceb10dba7870b05273c98d684047e9c8ed9b110072c2a6ed8fb6

Observation 4d951ac7-d575-42a4-9883-708b82301d30 · outbound

This paper cites Adam: A method for stochastic optimization,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Adam: A method for stochastic optimization,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.450739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.450739Z digest=sha256:5b173bb201da18d863e1f18e00dc0fcc528eff5426fc00200f87686db7928a93

Observation 79ac9fa6-930f-436d-9db4-42135a710bbc · outbound

This paper cites Mixed Precision Training.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Mixed Precision Training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.589590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.589590Z digest=sha256:a2dcb4a52191e7393c9828aa978b98f12cc8bad349e75bbb2265256ada4a9d79

Observation f47b20f2-689e-457a-a4d7-222ee2868f08 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Adam: A Method for Stochastic Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.516419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.516419Z digest=sha256:c4e2eef4fb0db658222b6419a4a625113a277feb2269127d1a217127c806887f

Observation b5726e1c-a469-4c55-9868-aa892988d9a3 · outbound

This paper cites Polaris,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Polaris,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.671452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.671452Z digest=sha256:ec16f3485bfa839def169c6d9db2ae03cc3430f4a973414cd04bba9b00bc2835

Observation bffa5578-05b1-4d36-8146-29c46a95ee6b · outbound

This paper cites Asynccheckpointio– pytorch lightning,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Asynccheckpointio– pytorch lightning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.610290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.610290Z digest=sha256:bc793054c75acccda6e6a3723aef5bb093d45eac0b70fff8cf2230d0a08bca94

Observation 9d1dde8f-43cb-4a54-85ed-c1e81e725bc8 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Llama 2: Open Foundation and Fine-Tuned Chat Models,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.861433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.861433Z digest=sha256:77e20895c0ef97ef4e1da874371e740eb9385cd1c83564dd1f94792269d2b9bf

Observation 94c7494d-6f77-431b-a8b3-03bfd525090c · outbound

This paper cites Lustre: Building a file system for 1000-node clusters,.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Lustre: Building a file system for 1000-node clusters,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.769330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.769330Z digest=sha256:3a8b8f1d5878911497b6da9d997b319e4117eccb5076de3b724ba8938b1adf12

Observation d4d5a879-a7d5-44bc-b7a3-fe72d6472b1c · outbound

This paper cites An overview of gradient descent optimization algorithms.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers An overview of gradient descent optimization algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:57.370717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:57.370717Z digest=sha256:28120efd9f8c2ca7bec975413fbbb9ba54fb13bba894c34a1e01f69d2d52194f

Pith citing papers

Observation 62c009ef-6082-45fa-b036-27c0b91d5df8 · inbound

ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload cites this paper.

ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T01:14:28.007359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T01:59:28.408860Z digest=sha256:d63c778496089c6cc979871868e5a08d57478905df53ab995babe78114d7e456

Observation 16ac29bb-b626-4a74-95e5-39a9cdfd2880 · inbound

ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload cites this paper.

ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T01:14:28.007359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T05:52:30.402478Z digest=sha256:f6749d98e5322725ff480d3f0473b030450877126d3f34621b556940f427bd31