Pith. sign in

Paper Citation Record · LEDGER

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

As of 24 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 5 inbound Pith citation observations for arXiv:2411.13055.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13055 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:57:08.280854Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:34:50.039718Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:49:35.673815Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f14b286-ad0b-4956-8309-f61943e7eb3b · outbound

This paper cites The framework tax: Disparities between inference efficiency in nlp research and deployment.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training The framework tax: Disparities between inference efficiency in nlp research and deployment

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:08.639044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T16:57:08.199477Z digest=sha256:bf26070d16584e598661366b41023b2ecadbceb639f510d4a796a097007a1665

Observation 2cb708ee-016e-4178-93f9-67e759f2c1b0 · outbound

This paper cites ISBN 9781450357999.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training ISBN 9781450357999

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.203305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.203305Z digest=sha256:a8fe63155d829166165f13ab7437d58182bfdcf6e987803479876acdbb7e57e0

Observation de09c840-e31d-410c-ac48-d355e0211be1 · outbound

This paper cites PipeDream: Fast and Efficient Pipeline Parallel DNN Training.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training PipeDream: Fast and Efficient Pipeline Parallel DNN Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.215331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.215331Z digest=sha256:d3060e02994af397165d3f17d089f1fc506fa6fa74f4f310fe8a0f5f660a1f8e

Observation 39cc2a9b-c2e3-4e21-9ffb-25c7fb214878 · outbound

This paper cites Counting Carbon: A Survey of Factors Influencing the Emissions of Machine Learning.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Counting Carbon: A Survey of Factors Influencing the Emissions of Machine Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.227649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.227649Z digest=sha256:fd1dac31fc8f4771d922dc68f8aaf8c45abf36977b1181701d71b46f7cc24722

Observation 059f9918-1a17-47ce-846b-a1478e1f7e8b · outbound

This paper cites Fully sharded data parallel: faster ai training with fewer gpus — engineering.fb.com.https://engineering.fb.com/2021/07/15/open-source/fsdp/,.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Fully sharded data parallel: faster ai training with fewer gpus — engineering.fb.com.https://engineering.fb.com/2021/07/15/open-source/fsdp/,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:08.616127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T16:57:08.232052Z digest=sha256:0de4c45f7a96edf3a67ebe4ffc975f9bf81854385dec3f3736530d7afe463b2d

Observation 379dd7f7-b779-49b2-95a0-f95aa70eec70 · outbound

This paper cites Resolving Discrepancies in Compute-Optimal Scaling of Language Models.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.240003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.240003Z digest=sha256:6d9b5e7a059dbfbdc9fb93d7d620cc2c64334be48bee48a6613a336bf7a84b69

Observation f8ef9593-f213-420f-a0a6-3c43dadb2040 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.243679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.243679Z digest=sha256:098d90249ec4f84dd7397710fccacbc07b8eb84dd19d291b0c535c4732149957

Observation 57c7f08b-967d-4e82-8dfb-8a8cb581cd13 · outbound

This paper cites Mesh-TensorFlow: Deep Learning for Supercomputers.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Mesh-TensorFlow: Deep Learning for Supercomputers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.248160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.248160Z digest=sha256:af5bb687dc35054eb96be624b08139d7f192888d961df71500fa846eebec50e4

Observation c8b49888-7de9-4398-852c-a9162dee8530 · outbound

This paper cites Local SGD Converges Fast and Communicates Little.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Local SGD Converges Fast and Communicates Little

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.252150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.252150Z digest=sha256:4461252c3488fe5188cf99dc90d7cca868810d80393e305f88bdeff245be2225

Observation 7b02bdda-a1f3-40ff-8d77-212258cf9285 · outbound

This paper cites doi: 10.18653/v1/P19-1355.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training doi: 10.18653/v1/P19-1355

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.256095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.256095Z digest=sha256:7f8183ad89608e0d9376dafeb03255de16c5a5f7becab076432dfa60d061ac39

Observation d9f5a1a4-afd7-4fb6-9944-7b3bb43c13de · outbound

This paper cites Accessed: 2023-05-05.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Accessed: 2023-05-05

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:08.588509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T16:57:08.259740Z digest=sha256:869bcbe5e2011766a17839bc0837e740d3561d3bbfb0ad182fdf810e6a20a9ed

Observation 6be5e19d-562b-4638-acf6-6bd1cfe95de2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.263714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.263714Z digest=sha256:359a3b3ee92435be67dd689381a0bf0d0582d34305e268af27effe1b4b8aac6a

Observation 66ae41e1-028a-4b2f-a0b6-98292fad1375 · outbound

This paper cites Context Parallelism for Scalable Million-Token Inference.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Context Parallelism for Scalable Million-Token Inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.267655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.267655Z digest=sha256:d749aa8d44bcd8bfe829ebd9cc199769f170fc4b5257f09f1684da10bdf6f8f8

Observation 83872ab0-d920-4d2e-bf4a-7778ac86523b · outbound

This paper cites Accelerating the training of large language models using efficient activation rematerialization and optimal hybrid parallelism.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Accelerating the training of large language models using efficient activation rematerialization and optimal hybrid parallelism

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:08.576955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T16:57:08.271858Z digest=sha256:4f0996ab4932ecc00b0d0f3a1bafbe88e191bf0564f98e4eccacb12be662b0b3

Observation eda0858f-df94-4e63-8458-1d40fb7b2a6b · outbound

This paper cites For our primary experiments, we trained models using PyTorch 2.3.1 built with CUDA 12.1, with attention implementation provided by XFormers 0.27.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training For our primary experiments, we trained models using PyTorch 2.3.1 built with CUDA 12.1, with attention implementation provided by XFormers 0.27

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:08.565904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T16:57:08.276496Z digest=sha256:f105b6d155a16243ee8191ee323f7cbdd01e6e8b6277bf8e8c722500f6815f52

Observation e2fc97e2-8491-43b8-bfa2-b63868813885 · outbound

This paper cites Nodes within the V100 cluster consist of 8-GPU setups connected with first-generation NVLink in a Hybrid Cube Mesh (HCM) topology.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Nodes within the V100 cluster consist of 8-GPU setups connected with first-generation NVLink in a Hybrid Cube Mesh (HCM) topology

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:08.554992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T16:57:08.280854Z digest=sha256:99ef9ca1d48a0750b12f9cccbb886c5e39285b7553ee5a5926cbaa5791274179

Observation c75461ec-bd9b-41bd-961b-2f835c97dea3 · outbound

This paper cites Efficient Parallelization Layouts for Large-Scale Distributed Model Training.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Efficient Parallelization Layouts for Large-Scale Distributed Model Training

Reference 1988

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:57:08.428335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T16:57:08.210789Z digest=sha256:14407c642e592d711f7fd08216895f7d1c7faf1ada0acc63565606847087fa59

Observation bdfeb6c4-754e-48a2-97e5-9caac51d19c9 · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training OLMo: Accelerating the Science of Language Models

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.206900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.206900Z digest=sha256:7907e8a9696c306020a8386b0e0f2b17f14bc8572c83a3ebcc88b958731f4146

Observation fb2eb854-d209-452e-9b05-46688aed7cae · outbound

This paper cites Zhenkun Cai, Xiao Yan, Kaihao Ma, Yidi Wu, Yuzhen Huang, James Cheng, Teng Su, and Fan Yu.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Zhenkun Cai, Xiao Yan, Kaihao Ma, Yidi Wu, Yuzhen Huang, James Cheng, Teng Su, and Fan Yu

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:08.650641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T16:57:08.182207Z digest=sha256:9c7bae26f3b9b60acad65121e6cfd017e6adab20f5feebf0dada3764fc4dfd4e

Observation e45cb14d-3e74-4c48-82a9-448c8fcf7ceb · outbound

This paper cites Scaling Laws for Neural Language Models.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Scaling Laws for Neural Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.219364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.219364Z digest=sha256:a54976febe48dfeafde28ccd8f3c6eb789bdc1f4d4f6a75ee5bd2a4780a5af7b

Observation 4943cd3b-1d50-40ba-b9b8-7516805e7260 · outbound

This paper cites Branch-train-merge: Embarrassingly parallel training of expert language models.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Branch-train-merge: Embarrassingly parallel training of expert language models

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:08.628106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T16:57:08.223363Z digest=sha256:a21d8908f1fda31b6035d29bb14f7d7e42fae9f0cf4a4e4aaa47a183452d3393

Observation aa226829-b6cf-4b00-b9c7-c58ebcd8511b · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Training Deep Nets with Sublinear Memory Cost

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.187106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.187106Z digest=sha256:7d7f15c16a7bf79e3519d5f38b16d3993d154e57de3e81d293e13c9f986b4acb

Observation 0a63f94a-d0b3-4149-871b-8cff7a9af167 · outbound

This paper cites DiLoCo: Distributed Low-Communication Training of Language Models.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training DiLoCo: Distributed Low-Communication Training of Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.191511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.191511Z digest=sha256:fdfc6694c52f1231d62a454d20305460c958414fc1b0108bbec7727aa8ba9e07

Observation 2e391d26-2c8f-4134-9c6c-cd6c54f0acc1 · outbound

This paper cites The Llama 3 Herd of Models.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training The Llama 3 Herd of Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.195645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.195645Z digest=sha256:ed0c06b82bcbdae1a2d81f54b1128f5ce8c8d98ef2939e6acd61b80fc153c3e8

Observation 67ad680d-8f13-440b-84fc-d6b538cbf2f7 · outbound

This paper cites Saptadeep Pal, Eiman Ebrahimi, Arslan Zulfiqar, Yaosheng Fu, Victor Zhang, Szymon Migacz, David Nellans, and Puneet Gupta.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Saptadeep Pal, Eiman Ebrahimi, Arslan Zulfiqar, Yaosheng Fu, Victor Zhang, Szymon Migacz, David Nellans, and Puneet Gupta

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:08.602676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T16:57:08.236177Z digest=sha256:74366442312a93a45ee9eb46d9b727abcf78791242c17546f13ecf8e403051f9

Pith citing papers

Observation 5e7c425e-9cf5-4e10-ac35-a024e27d7583 · inbound

Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations cites this paper.

Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:14:19.312416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:14:19.312416Z digest=sha256:dcc57da7c95b7e84d0897be48f16b88aa5023a72ec84e75ab7ffe0276279572e

Observation b50beeb8-3ed8-495d-89f3-06fbd20e7bea · inbound

Estudio de la eficiencia en la escalabilidad de GPUs para el entrenamiento de Inteligencia Artificial cites this paper.

Estudio de la eficiencia en la escalabilidad de GPUs para el entrenamiento de Inteligencia Artificial Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T16:34:50.039718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:34:50.039718Z digest=sha256:8cea9e37126cf2f75c10e753760a8e5b2ecb30cdf3955cbdfbd864197cd87ca8

Observation b1206a0a-57a3-49b7-aa5e-75bd1b9e54e9 · inbound

Dynamic Core Allocation for Malleable Jobs with Unknown Speed-up Parameters cites this paper.

Dynamic Core Allocation for Malleable Jobs with Unknown Speed-up Parameters Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:49:35.675890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T16:43:13.430673Z digest=sha256:5144ddd54f3bfbec2d80ff3dfa9cae624485aed56ed317e981dd66fb7d6dc4d3

Observation 93a46424-4d30-46fa-9e19-b47043b6880f · inbound

Evaluating MFU as a Proxy for GPU Power for Energy-Aware Simulation of LLM Training cites this paper.

Evaluating MFU as a Proxy for GPU Power for Energy-Aware Simulation of LLM Training Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:32.042400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:32.042400Z digest=sha256:50452631e3bbe03c734bd587a2907eff9794a0fc5432d702d6356063571d27bb

Observation a0f4deb9-0b83-40f3-85b5-0ccb76c475eb · inbound

Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts cites this paper.

Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T21:08:06.703997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:08:06.703997Z digest=sha256:4e041cf5671e239c1432901b3d52e4a96f05ed906aee9d82802bf22d78456f4f