Pith. sign in

Paper Citation Record · LEDGER

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning

As of 11 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2501.04266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04266 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:42:38.233849Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact2
  • verified fuzzy12
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 556f303e-7734-4b11-b9d8-2844dd8417d2 · outbound

This paper cites Introducing the next generation of Claude — anthropic.com,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Introducing the next generation of Claude — anthropic.com,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.665918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:42:38.092503Z digest=sha256:4bf7dbaa72ef3b8f7b7b887596da1776290ef1afadab1b3dbabfe1d3de9472b3

Observation d4e413ca-aabc-45e0-92b2-55fb4cbdfb9e · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Gemma: Open Models Based on Gemini Research and Technology

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.097420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.097420Z digest=sha256:214c8dcdc6a086f8a49c15d3540e20bafcbe570424eb150bb5ac7ff793d64b7d

Observation b15f156a-3aac-4470-a570-fb3a0d78245b · outbound

This paper cites Llama 3 model card,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Llama 3 model card,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.102113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.102113Z digest=sha256:f86913c1d83e5b31891fbf240bf84eb4ea4b94d5f56f809d3cb459294910c92d

Observation 9dbe729f-675d-4788-ba68-023caeb7369a · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.106390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.106390Z digest=sha256:df45441de929d8d01f15bf6bfd370f5eea7952902c80ea4bd83fe038fbf622d9

Observation 0d7b9409-785c-44b4-912a-902f96bb8725 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.110991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.110991Z digest=sha256:0dfca57b45c0c92b826d73df762656ccd2ec8c857dc40df988ddb96cfc43dfbe

Observation 3b8c2180-fd1c-459e-9437-ea320f8ebbd7 · outbound

This paper cites Measuring massive multitask language understanding,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Measuring massive multitask language understanding,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.115345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.115345Z digest=sha256:c9f1ff2480e8ec06095ca8f5564053cbf5d94b00578ab2b8491078667dc575e8

Observation bcbdff8a-f26d-4bc1-b441-fb0214fe0104 · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.123613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.123613Z digest=sha256:75d0cff3ec51375e3a88ea9464a2b936e438006347a67732c1b2977e73d4d564

Observation b3fe7a0d-a675-4545-a842-7d5c137f3e71 · outbound

This paper cites The mvapich project: Transforming research into high-performance mpi library for hpc community,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning The mvapich project: Transforming research into high-performance mpi library for hpc community,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.640073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:42:38.127811Z digest=sha256:3f19f686c3f4a288c45fd663b3fd2d33d8c5e17b77367baacfd1c83084a90521

Observation 064a4e3b-3562-49de-9495-bfe45d86a64f · outbound

This paper cites Scalable Distributed DNN Training using TensorFlow and CUDA-Aware MPI: Characterization, Designs, and Performance Evaluation,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Scalable Distributed DNN Training using TensorFlow and CUDA-Aware MPI: Characterization, Designs, and Performance Evaluation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.627707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:42:38.132220Z digest=sha256:5ae454651bdbb492a5c443908dd878c5af841ffe6674c8ed3bf0c8797b6a168a

Observation 954b19cf-6e28-4fac-b409-9ff30e210e7b · outbound

This paper cites NVIDIA Collective Communications Library (NCCL),.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning NVIDIA Collective Communications Library (NCCL),

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.614921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:42:38.135902Z digest=sha256:59d34a7c17ed914998984c16bd7b9ad6080e852d7fa2c737978387cff08e871c

Observation 02bcaf6a-9e55-4078-a489-a98d61c9662c · outbound

This paper cites AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:42:38.400265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:42:38.139962Z digest=sha256:0d6952377a50d95571c6ae3bc2a61a8cc81c1a35238b991e8ab0f82373a6f69b

Observation 3da99280-6328-4b4c-a270-735321a48d59 · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.144213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.144213Z digest=sha256:1b846aff467207151be5ab10caf646ef3f006db2a94743190e4438d76f15cf4c

Observation a9c65893-895e-43d9-84ff-ba1a9ebdca58 · outbound

This paper cites Fairscale: A general purpose modular pytorch li- brary for high performance and large scale training,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Fairscale: A general purpose modular pytorch li- brary for high performance and large scale training,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.595578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:42:38.157012Z digest=sha256:2dbf128eaa3d786b34dfed67967eac31a7e2535f4644d5217838c87db3e44e60

Observation 7d60d192-106f-499a-bdca-3fb29ecc7704 · outbound

This paper cites Megatron-LM: Ongoing research training transformer models at scale,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Megatron-LM: Ongoing research training transformer models at scale,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.583069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:42:38.160347Z digest=sha256:7a33076efbc34bb0a365bb66a7d728d5aecdb7d12ce1b50a7d276ce64b15cd90

Observation 29ad6087-e169-4d2c-85c3-f1d09dcbbedc · outbound

This paper cites Frontier - HPE Cray EX235a, AMD Optimized 3rd Generation EPYC 64C 2GHz, AMD Instinct MI250X, Slingshot-11 | TOP500 — top500.org,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Frontier - HPE Cray EX235a, AMD Optimized 3rd Generation EPYC 64C 2GHz, AMD Instinct MI250X, Slingshot-11 | TOP500 — top500.org,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.571146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:42:38.164101Z digest=sha256:da375945d5c23340c5fb52667515065efce6d5f0e5dbbc429ef62a3cbe9f0abb

Observation f524f000-a4fe-4496-a184-36f228f07747 · outbound

This paper cites An in-depth analysis of the slingshot interconnect,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning An in-depth analysis of the slingshot interconnect,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.558817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:42:38.167716Z digest=sha256:0fea1482bc88ef9a3312d070766e5cc70f2dd5617b43bf71ae46e83f020baa75

Observation e6cee94d-9069-47d4-9335-4375fccf66f0 · outbound

This paper cites ZeRO++: Extremely Efficient Collective Communication for Giant Model Training.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning ZeRO++: Extremely Efficient Collective Communication for Giant Model Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.171907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.171907Z digest=sha256:28dec39571f800bd2056b3dd97bf869283f1ae38242e130237a3893c05623836

Observation 18d3800b-6188-415e-b7f8-9f91ee29305c · outbound

This paper cites Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.175813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.175813Z digest=sha256:04f405a2cd6e0a6419d452fe51415963ee81b7b8e0e23c6a27cd02790a316dab

Observation 9b89cf4b-5e8f-44d7-9ef6-76c09e6094e6 · outbound

This paper cites Scaling single- image super-resolution training on modern hpc clusters: Early experi- ences,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Scaling single- image super-resolution training on modern hpc clusters: Early experi- ences,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.545770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:42:38.179854Z digest=sha256:d170aa6cb9a30cdf3e9fea2619b4ac22e6fc2573132aeb28f330d93cc1d0900c

Observation 660dbb0d-91ae-41fb-bdd4-68fbb0370391 · outbound

This paper cites Adam: A method for stochastic optimization,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Adam: A method for stochastic optimization,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.183548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.183548Z digest=sha256:9d65ca3e750d173825ee108c897cb56a4cc00920c6a58da36d4fcff7a4fe4460

Observation a4623ac3-e50c-431e-8457-f0b7523e542d · outbound

This paper cites 8-bit Optimizers via Block-wise Quantization.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning 8-bit Optimizers via Block-wise Quantization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.187634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.187634Z digest=sha256:55c2af0c1339dfd2517938fc01a437425a93be19865314fd0995debff3226459

Observation e42ab56d-3118-4ed5-8023-88c78fc161a3 · outbound

This paper cites Decoupled weight decay regularization,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Decoupled weight decay regularization,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.191520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.191520Z digest=sha256:56bf45856e1ddf84ad4b09e8e89d10b50b1dd6716e41dd22cbde1716c67aefb0

Observation b26f6e6b-5293-47ed-8b43-ed9f35407131 · outbound

This paper cites GPT-NeoX-20B: An open-source autoregressive language model,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning GPT-NeoX-20B: An open-source autoregressive language model,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.515169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:42:38.195131Z digest=sha256:444a130f53c5369befa02f8843194381adbd713a1383861b0602ea6b23f531ce

Observation bfdced74-66cc-4749-82cc-914fa6cd931a · outbound

This paper cites Language models are few-shot learners,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Language models are few-shot learners,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.198731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.198731Z digest=sha256:31cca88adbaa6a6d603fc6734b7e3e89743eb9bf54d22cf5a5e745373b14acf5

Observation fa9efb2b-2a5e-4b36-a55e-8d563f5b3058 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning RWKV: Reinventing RNNs for the Transformer Era

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.202557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.202557Z digest=sha256:ef27aa6f26ed0e02c8f7edf974b256430302d6d28cecb56e2820f4efcb25b140

Observation 71fdde9f-5416-476f-9f80-77c0ee2a3147 · outbound

This paper cites Pytorch fsdp: Experiences on scaling fully sharded data parallel,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Pytorch fsdp: Experiences on scaling fully sharded data parallel,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.206521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.206521Z digest=sha256:7add9c1ab63fa91c1a778bb3a598e85952037bd3e4ca26869060042a9f160b07

Observation 960dfec8-e83b-4ee4-8161-11a0675d1929 · outbound

This paper cites Mics: Near-linear scaling for training gigantic model on public cloud,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Mics: Near-linear scaling for training gigantic model on public cloud,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.487222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:42:38.214329Z digest=sha256:42fb80d57e42748ca60b11a743d76d9608dd248d1ae4bb321597221f1dea999a

Observation 432033b2-76e4-4c6b-839f-cbbce1021da7 · outbound

This paper cites MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.218267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.218267Z digest=sha256:02fd493b494e20921770252b5900ace16eccc556c0276046c196df62b189350c

Observation 2b629f5e-1470-412e-a753-13dd96880d73 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.210400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.210400Z digest=sha256:845a816ee90e80824971a6d182b237dceb3a51d4a56887a3af6579ac426a3741

Observation 12f5fa30-71b4-4dd6-8ae2-27c76559ba6e · outbound

This paper cites Optimizing Distributed Training on Frontier for Large Language Models.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Optimizing Distributed Training on Frontier for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.225511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.225511Z digest=sha256:bdb59dcd409213d87d138e1c2b1329ff32870c0a81f02ffbae96f7720cc58a49

Observation 5e608e03-f496-4b9a-8945-d3c111500d6a · outbound

This paper cites Comparative Study of Large Language Model Architectures on Frontier.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Comparative Study of Large Language Model Architectures on Frontier

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:42:38.286343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:42:38.229469Z digest=sha256:acd18e3de6784528c3662097871415e96cc4964eefd05a8154dc23f1a58a7af0

Observation 75297a7e-1ce8-4d1a-8f66-b58553ba9c54 · outbound

This paper cites Accelerating large language model training with hybrid gpu-based compression,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Accelerating large language model training with hybrid gpu-based compression,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.472809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:42:38.222134Z digest=sha256:92e19b519b953444a77bbcf312a28f416cb509b96c77d5684c1ea9e25e64fb41

Observation 70cee47f-4482-4411-b2cf-e42347847ace · outbound

This paper cites Hilfer fractional advection-diffusion equations with power-law initial condition; a Numerical study using variational iteration method.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Hilfer fractional advection-diffusion equations with power-law initial condition; a Numerical study using variational iteration method

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.233849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.233849Z digest=sha256:f47d60a318a3cada7a60ee3e1c5a17cc44895e87484d92a3ff980bc927cc9bf3

Observation c3c524c6-630f-477b-b15e-2c946f85c94a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Measuring Massive Multitask Language Understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.119113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.119113Z digest=sha256:ef92c385fc4067d2dd43dc5e352a0de32bbce63565cfca0a3e9867814ca8b014

Pith citing papers

No inbound Pith citation observations are available.