Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:57:08.280854Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 5 inbound Pith citation observations for arXiv:2411.13055.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:57:08.280854Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:34:50.039718Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T04:49:35.673815Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9f14b286-ad0b-4956-8309-f61943e7eb3b · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training The framework tax: Disparities between inference efficiency in nlp research and deployment
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2cb708ee-016e-4178-93f9-67e759f2c1b0 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training ISBN 9781450357999
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de09c840-e31d-410c-ac48-d355e0211be1 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training PipeDream: Fast and Efficient Pipeline Parallel DNN Training
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39cc2a9b-c2e3-4e21-9ffb-25c7fb214878 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Counting Carbon: A Survey of Factors Influencing the Emissions of Machine Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 059f9918-1a17-47ce-846b-a1478e1f7e8b · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Fully sharded data parallel: faster ai training with fewer gpus — engineering.fb.com.https://engineering.fb.com/2021/07/15/open-source/fsdp/,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 379dd7f7-b779-49b2-95a0-f95aa70eec70 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Resolving Discrepancies in Compute-Optimal Scaling of Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8ef9593-f213-420f-a0a6-3c43dadb2040 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57c7f08b-967d-4e82-8dfb-8a8cb581cd13 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Mesh-TensorFlow: Deep Learning for Supercomputers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8b49888-7de9-4398-852c-a9162dee8530 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Local SGD Converges Fast and Communicates Little
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b02bdda-a1f3-40ff-8d77-212258cf9285 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training doi: 10.18653/v1/P19-1355
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9f5a1a4-afd7-4fb6-9944-7b3bb43c13de · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Accessed: 2023-05-05
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6be5e19d-562b-4638-acf6-6bd1cfe95de2 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66ae41e1-028a-4b2f-a0b6-98292fad1375 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Context Parallelism for Scalable Million-Token Inference
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83872ab0-d920-4d2e-bf4a-7778ac86523b · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Accelerating the training of large language models using efficient activation rematerialization and optimal hybrid parallelism
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation eda0858f-df94-4e63-8458-1d40fb7b2a6b · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training For our primary experiments, we trained models using PyTorch 2.3.1 built with CUDA 12.1, with attention implementation provided by XFormers 0.27
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e2fc97e2-8491-43b8-bfa2-b63868813885 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Nodes within the V100 cluster consist of 8-GPU setups connected with first-generation NVLink in a Hybrid Cube Mesh (HCM) topology
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c75461ec-bd9b-41bd-961b-2f835c97dea3 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Efficient Parallelization Layouts for Large-Scale Distributed Model Training
Reference 1988
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bdfeb6c4-754e-48a2-97e5-9caac51d19c9 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training OLMo: Accelerating the Science of Language Models
Reference 2000
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb2eb854-d209-452e-9b05-46688aed7cae · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Zhenkun Cai, Xiao Yan, Kaihao Ma, Yidi Wu, Yuzhen Huang, James Cheng, Teng Su, and Fan Yu
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e45cb14d-3e74-4c48-82a9-448c8fcf7ceb · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Scaling Laws for Neural Language Models
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4943cd3b-1d50-40ba-b9b8-7516805e7260 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Branch-train-merge: Embarrassingly parallel training of expert language models
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation aa226829-b6cf-4b00-b9c7-c58ebcd8511b · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Training Deep Nets with Sublinear Memory Cost
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a63f94a-d0b3-4149-871b-8cff7a9af167 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training DiLoCo: Distributed Low-Communication Training of Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e391d26-2c8f-4134-9c6c-cd6c54f0acc1 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training The Llama 3 Herd of Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ad680d-8f13-440b-84fc-d6b538cbf2f7 · outbound
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Saptadeep Pal, Eiman Ebrahimi, Arslan Zulfiqar, Yaosheng Fu, Victor Zhang, Szymon Migacz, David Nellans, and Puneet Gupta
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5e7c425e-9cf5-4e10-ac35-a024e27d7583 · inbound
Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b50beeb8-3ed8-495d-89f3-06fbd20e7bea · inbound
Estudio de la eficiencia en la escalabilidad de GPUs para el entrenamiento de Inteligencia Artificial Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1206a0a-57a3-49b7-aa5e-75bd1b9e54e9 · inbound
Dynamic Core Allocation for Malleable Jobs with Unknown Speed-up Parameters Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 93a46424-4d30-46fa-9e19-b47043b6880f · inbound
Evaluating MFU as a Proxy for GPU Power for Energy-Aware Simulation of LLM Training Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0f4deb9-0b83-40f3-85b5-0ccb76c475eb · inbound
Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.