Pith. sign in

Paper Citation Record · LEDGER

ZeRO-Offload: Democratizing Billion-Scale Model Training

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2101.06840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2101.06840 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:49:10.102076Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T02:29:25.467389Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6f09a7bd-06a9-4200-b46f-e61178dc17e7 · inbound

Movie Gen: A Cast of Media Foundation Models cites this paper.

Movie Gen: A Cast of Media Foundation Models ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:25.884336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:8d0de881ab94229ff535dff9f9826ef7ab26b9714f8cacc5b16cbc48d1a6d7c1

Observation 1a7350ec-65db-48e0-8758-936a44e8f0c5 · inbound

Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models cites this paper.

Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:15.517148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:44:15.517148Z digest=sha256:db0509f956b775b174b44418484e7f08be438d5cfeae853ecea988022438744e

Observation c62f4314-1d2e-4b3d-8812-0bfdfcd3d386 · inbound

TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain cites this paper.

TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:04:15.936377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:04:15.936377Z digest=sha256:c8a31cda7fbc0660e74ec1162a26d70ed04434aab3e64104924bdeb030d4c7e7

Observation 87cd75ca-9024-4b70-9036-d4edcb65e7c0 · inbound

Adjoint sharding for very long context training of state space models cites this paper.

Adjoint sharding for very long context training of state space models ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:28.765359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:28.765359Z digest=sha256:ab2d4151a8e87270aa5db469ed5202508319be9767071cb535773df4fb5af920

Observation 24ae9a37-3de0-472f-ac5b-de0e4f945a3c · inbound

Practical Design and Benchmarking of Generative AI Applications for Surgical Billing and Coding cites this paper.

Practical Design and Benchmarking of Generative AI Applications for Surgical Billing and Coding ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:48:15.226897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:48:15.226897Z digest=sha256:54d19c0797d9f9dd84a65b5fd1c726792589d91ceeddee75e209919f5facb0eb

Observation f8781437-3a20-4210-8d3c-598c1a342079 · inbound

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference cites this paper.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.357295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.357295Z digest=sha256:d734da11282346a009f49b8588facab3311f8e46fd4d7edc00d50c9d3f54f3cf

Observation ee17e96a-63e9-4342-a638-70190e5af613 · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:10.102076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:10.102076Z digest=sha256:d869ba5a63bf1d7f3ea438a272ecfcba2ed010629d4fa2b09aac5486d3bd259f

Observation 0f8ab8ff-fc6c-48c9-a51e-7f85312a48f0 · inbound

INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning cites this paper.

INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:25:53.111264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:25:53.111264Z digest=sha256:de2a7a8128388418287880280c6029995b97d82074f4020797d9d4c161379329

Observation 9d289a50-00d3-4dce-b04c-517a15cf9229 · inbound

StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs cites this paper.

StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:17:36.966763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:17:36.966763Z digest=sha256:100e44e44d4db78755175b9a0ae8f102f4e1599d398080355ba4e5b05019f8ae

Observation 2110afe0-67c5-4140-9e9b-df5b4b7e3dc5 · inbound

Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments cites this paper.

Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:28.214083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:28.214083Z digest=sha256:b1ad029826bac6b5904a549489b787d8b0ff231125ec8923c7f81a53daa2038f

Observation 7d326a68-20bb-44e4-b59a-0e462889b87f · inbound

Steering the Noise: Turning Random Perturbations into Effective Descent for Memory-Efficient LLM Fine-Tuning cites this paper.

Steering the Noise: Turning Random Perturbations into Effective Descent for Memory-Efficient LLM Fine-Tuning ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T12:03:26.187156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:03:26.187156Z digest=sha256:df04592b15d52438d7bfd965abc86b3497ad949e3fce6fb51f8e85e557f2e95f

Observation 9672136a-518e-4c30-851e-7bb5d7ebe336 · inbound

Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning cites this paper.

Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:21:09.679054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T12:03:58.279499Z digest=sha256:ad241cb45a39adac2b53160adea5fecedee05510acdf3087e7fd90796922c78f

Observation 6e4d8598-dae2-4139-a146-7a330834968d · inbound

Efficient Training on Multiple Consumer GPUs with RoundPipe cites this paper.

Efficient Training on Multiple Consumer GPUs with RoundPipe ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:26.806007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T10:37:22.251566Z digest=sha256:0ee5e5119e47e72edec546091519f7db56d63817c72726b7bae449def84d8336

Observation 4ee168a2-5bd8-4ea7-aee4-6535d948333d · inbound

PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR cites this paper.

PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:29:25.470695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T02:24:48.872065Z digest=sha256:9f330e48ba955a41803dce6e2ace5ed29322285ba6226b109665294bc65efc2b

Observation 4f838c85-c32c-4f75-95d2-955780b43977 · inbound

Memory-Efficient Activation Checkpointing with Sliding Window and Hirschberg's Algorithm for 0/1 Knapsack Solving in PyTorch cites this paper.

Memory-Efficient Activation Checkpointing with Sliding Window and Hirschberg's Algorithm for 0/1 Knapsack Solving in PyTorch ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:49.767449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:31:49.767449Z digest=sha256:e66f79b3c4058a71dce6cebb904c9766fcc56bbb6e0e38976691bfc0d4ac8b4d

Observation e7e33618-595a-4f2c-8e6d-b48460f01c7a · inbound

Interpreting Language Model Hidden States at Scale cites this paper.

Interpreting Language Model Hidden States at Scale ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:17.067120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:16:17.067120Z digest=sha256:2b41eccee9b62cde70bf2777b64c8d64b2080216d892c040675235e8bd90a41b