Pith. sign in

Paper Citation Record · LEDGER

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning

As of 12 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2501.04266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04266 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:42:38.233849Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact2
  • verified fuzzy12
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 556f303e-7734-4b11-b9d8-2844dd8417d2 · outbound

This paper cites Introducing the next generation of Claude — anthropic.com,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Introducing the next generation of Claude — anthropic.com,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.665918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:42:38.092503Z digest=sha256:049006b3443c350bf3d093720ed7f266efa3e482c88c0c51a7c4b0bdb3c75743

Observation d4e413ca-aabc-45e0-92b2-55fb4cbdfb9e · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Gemma: Open Models Based on Gemini Research and Technology

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.097420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.097420Z digest=sha256:242c62169d413724e4af564b62e1b8499e29c116d88b0fde582cf285ea763bae

Observation b15f156a-3aac-4470-a570-fb3a0d78245b · outbound

This paper cites Llama 3 model card,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Llama 3 model card,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.102113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.102113Z digest=sha256:bbd0c921ee6270c5bed83d18145be4f22810b2bace770cf7836dce46df7e15d3

Observation 9dbe729f-675d-4788-ba68-023caeb7369a · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.106390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.106390Z digest=sha256:1a1e68d33a99db18047ea24a512764dffe091d3b48514b8dbe0e50afb88f52ab

Observation 0d7b9409-785c-44b4-912a-902f96bb8725 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.110991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.110991Z digest=sha256:583c2a4f6f9bb17f9c6b4ae860f0c378d7c79f7ab88536e9ecff74a5410f4840

Observation 3b8c2180-fd1c-459e-9437-ea320f8ebbd7 · outbound

This paper cites Measuring massive multitask language understanding,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Measuring massive multitask language understanding,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.115345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.115345Z digest=sha256:8499f8e9e120621cec164b995701924df22936c6848a130b1955888488688c6b

Observation bcbdff8a-f26d-4bc1-b441-fb0214fe0104 · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.123613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.123613Z digest=sha256:5ba694ecf69f43bbfc76a3bc00f22c045d54ee59b6198dd678993a2690bfc980

Observation b3fe7a0d-a675-4545-a842-7d5c137f3e71 · outbound

This paper cites The mvapich project: Transforming research into high-performance mpi library for hpc community,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning The mvapich project: Transforming research into high-performance mpi library for hpc community,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.640073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:42:38.127811Z digest=sha256:227c2cdbb34d01901045ba3397f190581ff94a06180a47b0d9e3ecf7c8a2eca7

Observation 064a4e3b-3562-49de-9495-bfe45d86a64f · outbound

This paper cites Scalable Distributed DNN Training using TensorFlow and CUDA-Aware MPI: Characterization, Designs, and Performance Evaluation,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Scalable Distributed DNN Training using TensorFlow and CUDA-Aware MPI: Characterization, Designs, and Performance Evaluation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.627707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:42:38.132220Z digest=sha256:94b27fd1ccda5b11ada255cdfa777035fb458332260c75fbbfc0f18f27302561

Observation 954b19cf-6e28-4fac-b409-9ff30e210e7b · outbound

This paper cites NVIDIA Collective Communications Library (NCCL),.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning NVIDIA Collective Communications Library (NCCL),

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.614921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:42:38.135902Z digest=sha256:70bb1926b501a32f5d68f3fb33a8ab734926685f4ad30aefca09c02662bd7503

Observation 02bcaf6a-9e55-4078-a489-a98d61c9662c · outbound

This paper cites AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:42:38.400265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:42:38.139962Z digest=sha256:6b229d6634e6af6a387277db3d7e4c538453006bc1ffbd9c31c2b3b025249a41

Observation 3da99280-6328-4b4c-a270-735321a48d59 · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.144213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.144213Z digest=sha256:935065fcc87a8df072e51cbf38f1476ace968ba6ca4ce4b19ecf90fb0d0c6b2a

Observation a9c65893-895e-43d9-84ff-ba1a9ebdca58 · outbound

This paper cites Fairscale: A general purpose modular pytorch li- brary for high performance and large scale training,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Fairscale: A general purpose modular pytorch li- brary for high performance and large scale training,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.595578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:42:38.157012Z digest=sha256:e89ba600f8a5f5f6952731767d297767c9a39b27aee66fdb9a7991025ebe31c4

Observation 7d60d192-106f-499a-bdca-3fb29ecc7704 · outbound

This paper cites Megatron-LM: Ongoing research training transformer models at scale,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Megatron-LM: Ongoing research training transformer models at scale,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.583069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:42:38.160347Z digest=sha256:996fac388bdabf2f03fb623c397dc779e92c194d520a5d0bfd572fd74b262118

Observation 29ad6087-e169-4d2c-85c3-f1d09dcbbedc · outbound

This paper cites Frontier - HPE Cray EX235a, AMD Optimized 3rd Generation EPYC 64C 2GHz, AMD Instinct MI250X, Slingshot-11 | TOP500 — top500.org,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Frontier - HPE Cray EX235a, AMD Optimized 3rd Generation EPYC 64C 2GHz, AMD Instinct MI250X, Slingshot-11 | TOP500 — top500.org,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.571146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:42:38.164101Z digest=sha256:812a442cdfc238fff1a96d717f964f5947b3d0b5b2b5ca769e7b9964a9e87ac8

Observation f524f000-a4fe-4496-a184-36f228f07747 · outbound

This paper cites An in-depth analysis of the slingshot interconnect,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning An in-depth analysis of the slingshot interconnect,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.558817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:42:38.167716Z digest=sha256:72f3e7dab02d31a031da7b67c3f311f3b9adf0b686ce5b90f42e7d264ccd9058

Observation e6cee94d-9069-47d4-9335-4375fccf66f0 · outbound

This paper cites ZeRO++: Extremely Efficient Collective Communication for Giant Model Training.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning ZeRO++: Extremely Efficient Collective Communication for Giant Model Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.171907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.171907Z digest=sha256:9ec00d0bdc8a02370b55304074a544c6166f758e1d98e08de71e002c29645544

Observation 18d3800b-6188-415e-b7f8-9f91ee29305c · outbound

This paper cites Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.175813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.175813Z digest=sha256:7ec05ef6f818c9e053da1110cfedd8053f668944f14cf148cc679e886162c59f

Observation 9b89cf4b-5e8f-44d7-9ef6-76c09e6094e6 · outbound

This paper cites Scaling single- image super-resolution training on modern hpc clusters: Early experi- ences,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Scaling single- image super-resolution training on modern hpc clusters: Early experi- ences,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.545770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:42:38.179854Z digest=sha256:42d8166e4764d64f0002952514b0bf819c91f8d17a4d9ce919d4e77de2482117

Observation 660dbb0d-91ae-41fb-bdd4-68fbb0370391 · outbound

This paper cites Adam: A method for stochastic optimization,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Adam: A method for stochastic optimization,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.183548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.183548Z digest=sha256:d0171bf88122c751f8e01b880a11b412799ebcfbb8a840647d2ddb10540c00c5

Observation a4623ac3-e50c-431e-8457-f0b7523e542d · outbound

This paper cites 8-bit Optimizers via Block-wise Quantization.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning 8-bit Optimizers via Block-wise Quantization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.187634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.187634Z digest=sha256:8edfd71b3063243e7dd695342592db87c39981285d4067e8b34396fbb45b678c

Observation e42ab56d-3118-4ed5-8023-88c78fc161a3 · outbound

This paper cites Decoupled weight decay regularization,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Decoupled weight decay regularization,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.191520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.191520Z digest=sha256:7f843ff036504c3dd8cce4a6b2f265613c4b6d1667f98baaf4d16764f8d3c0cc

Observation b26f6e6b-5293-47ed-8b43-ed9f35407131 · outbound

This paper cites GPT-NeoX-20B: An open-source autoregressive language model,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning GPT-NeoX-20B: An open-source autoregressive language model,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.515169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:42:38.195131Z digest=sha256:96c021b96c8e4669a1fe9d9a0d54ea0dd8fd1a6397c8586b79383f96a8074192

Observation bfdced74-66cc-4749-82cc-914fa6cd931a · outbound

This paper cites Language models are few-shot learners,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Language models are few-shot learners,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.198731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.198731Z digest=sha256:dad9b7206879fd44e21a4e4f0cdc67a07dbf47ff2785dd2bc051a99f42eeed9b

Observation fa9efb2b-2a5e-4b36-a55e-8d563f5b3058 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning RWKV: Reinventing RNNs for the Transformer Era

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.202557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.202557Z digest=sha256:95a3fbca71685c6a5a4b8b50967bc5dc7849ea1cbb8a07ea8e5fd663e05ed0a3

Observation 71fdde9f-5416-476f-9f80-77c0ee2a3147 · outbound

This paper cites Pytorch fsdp: Experiences on scaling fully sharded data parallel,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Pytorch fsdp: Experiences on scaling fully sharded data parallel,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.206521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.206521Z digest=sha256:6fbcaa1f798094841eeeb42432595c75ee2a09be4e75e8956a23ee8692bab930

Observation 960dfec8-e83b-4ee4-8161-11a0675d1929 · outbound

This paper cites Mics: Near-linear scaling for training gigantic model on public cloud,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Mics: Near-linear scaling for training gigantic model on public cloud,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.487222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:42:38.214329Z digest=sha256:dc0c73fabac01257312374ed8d43f3597fd65a970a38ca94ad97290e1d077a94

Observation 432033b2-76e4-4c6b-839f-cbbce1021da7 · outbound

This paper cites MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.218267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.218267Z digest=sha256:19d7d421b8502dae7820784c1c1ffbfc471b34208514ee5ffc055069051dbb42

Observation 2b629f5e-1470-412e-a753-13dd96880d73 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.210400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.210400Z digest=sha256:d3304d9fd542636171dc765385832bd4d1803d72ac87cde037b880014171cbbc

Observation 12f5fa30-71b4-4dd6-8ae2-27c76559ba6e · outbound

This paper cites Optimizing Distributed Training on Frontier for Large Language Models.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Optimizing Distributed Training on Frontier for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.225511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.225511Z digest=sha256:6ad4d727823cf817d5553da53d5b69efb5e00b6f8d647d72ea8352027b9a1c29

Observation 5e608e03-f496-4b9a-8945-d3c111500d6a · outbound

This paper cites Comparative Study of Large Language Model Architectures on Frontier.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Comparative Study of Large Language Model Architectures on Frontier

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:42:38.286343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:42:38.229469Z digest=sha256:04b0509ed2f2cb2b1427c59113206debac9c4b0a3deec2414ced5abd57809724

Observation 75297a7e-1ce8-4d1a-8f66-b58553ba9c54 · outbound

This paper cites Accelerating large language model training with hybrid gpu-based compression,.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Accelerating large language model training with hybrid gpu-based compression,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:42:38.472809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:42:38.222134Z digest=sha256:1941c5e88deacb2eb8710b589d7c633bb092477fa02d23968732473ba9307b71

Observation 70cee47f-4482-4411-b2cf-e42347847ace · outbound

This paper cites Hilfer fractional advection-diffusion equations with power-law initial condition; a Numerical study using variational iteration method.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Hilfer fractional advection-diffusion equations with power-law initial condition; a Numerical study using variational iteration method

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.233849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.233849Z digest=sha256:1cf781b505e946357d0400196a9d2314994dd446aa6c48a355de4c73bf67949d

Observation c3c524c6-630f-477b-b15e-2c946f85c94a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Measuring Massive Multitask Language Understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:38.119113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:38.119113Z digest=sha256:39075f955eae01efb040b7d2f67bf406ac3035f716fa04f005443177c61ada20

Pith citing papers

No inbound Pith citation observations are available.