Pith. sign in

Paper Citation Record · LEDGER

Measuring the Effects of Data Parallelism on Neural Network Training

As of 26 July 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:1811.03600.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1811.03600 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-26T06:30:07.085553+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-25T14:19:33.774814Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-25T14:20:55.408139Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e066515e-4a52-4133-b1a8-7756e8e343b7 · inbound

Large Batch Optimization for Deep Learning: Training BERT in 76 minutes cites this paper.

Large Batch Optimization for Deep Learning: Training BERT in 76 minutes Measuring the Effects of Data Parallelism on Neural Network Training

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:39:00.043658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-21T21:38:59.970003Z digest=sha256:95e80a42b31153f6c4823cf79e82b29406a6ce8493f5bb66e53df5b1ec7fc944

Observation ad32e816-9d1b-4f0c-b9f9-9dcc26032a84 · inbound

Fast Training of Sparse Graph Neural Networks on Dense Hardware cites this paper.

Fast Training of Sparse Graph Neural Networks on Dense Hardware Measuring the Effects of Data Parallelism on Neural Network Training

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-25T14:20:55.411323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-25T14:19:33.774814Z digest=sha256:fb0d85e28c4070377cd772c7a1e00db1f1df7687a0d5e96e8b2c9d3c302ad65c

Observation f54c10a9-b956-4034-9cfa-811986e8bba2 · inbound

FinBERT: Financial Sentiment Analysis with Pre-trained Language Models cites this paper.

FinBERT: Financial Sentiment Analysis with Pre-trained Language Models Measuring the Effects of Data Parallelism on Neural Network Training

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:26:05.054403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-15T20:26:04.766701Z digest=sha256:36e3da6fdc0cc7f84e11f4ec6fce5088fabd54a29825e437a3438ef2dd26e3f6

Observation 18c7010e-f418-48c2-99ab-756ad1d4fadb · inbound

Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification cites this paper.

Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification Measuring the Effects of Data Parallelism on Neural Network Training

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T17:37:07.768937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=arxiv_source observed=2026-05-17T17:37:07.719640Z digest=sha256:27af0dcdda791d84364583b35599e622a4136dbad067ef736963012935048788

Observation 3699be57-3190-4231-9656-4a60e1056e7f · inbound

Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer cites this paper.

Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer Measuring the Effects of Data Parallelism on Neural Network Training

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:37:55.929334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=pdf_text observed=2026-05-12T05:37:55.083206Z digest=sha256:29e59975e350ce2d9e0d87c3fd4b5e9ee72f1c283d0cdb12dc3c1cf69500ed79

Observation 4cb72880-3f4a-4b97-ba58-7aeca5d2ea9d · inbound

Scaling Laws for Transfer cites this paper.

Scaling Laws for Transfer Measuring the Effects of Data Parallelism on Neural Network Training

Reference 106

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:58:13.782418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=arxiv_source observed=2026-05-18T00:58:13.116663Z digest=sha256:3cab39b407f80c0a912d2b24c00df2c224c62e6941c664de6c0f20504c62d45d

Observation ec5c4739-6616-4ddb-874b-7592b01828f1 · inbound

A General Language Assistant as a Laboratory for Alignment cites this paper.

A General Language Assistant as a Laboratory for Alignment Measuring the Effects of Data Parallelism on Neural Network Training

Reference 148

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:22:59.652392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=arxiv_source observed=2026-05-11T14:22:57.925354Z digest=sha256:6951a9bba6e1a92769988438b48c43463607c006d27082507bbd55564f389c58

Observation 94ae50a6-440f-41f4-9c87-0c76996aef80 · inbound

Language Models (Mostly) Know What They Know cites this paper.

Language Models (Mostly) Know What They Know Measuring the Effects of Data Parallelism on Neural Network Training

Reference 225

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T15:42:47.891665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=arxiv_source observed=2026-05-10T15:42:47.274448Z digest=sha256:14a043f66f0028dc599e42eb095e306d6b7ea1f16e401ee8bd6aa42bc0f2abb8

Observation 1d9ff0b1-1624-45bf-afe7-f145a134d1b7 · inbound

The Recurrent Transformer: Greater Effective Depth and Efficient Decoding cites this paper.

The Recurrent Transformer: Greater Effective Depth and Efficient Decoding Measuring the Effects of Data Parallelism on Neural Network Training

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:04.562340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=arxiv_source observed=2026-05-09T22:03:35.260373Z digest=sha256:023518aa826ff0391369fa67ff8df6dea62e2b3de53f05f4e5cb334254f7db4a

Observation 54706af1-3a7f-4199-847d-ddb8b5cdefe5 · inbound

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization cites this paper.

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization Measuring the Effects of Data Parallelism on Neural Network Training

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-15T04:49:44.826609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-26T06:30:07.085553+00:00.

source=arxiv_source observed=2026-05-15T04:45:20.091598Z digest=sha256:1afb691e08c8256c6c78a17335b9e5e7df040d48c1ccf6929e4f249abc92f477