Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:42:38.233849Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2501.04266.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:42:38.233849Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 556f303e-7734-4b11-b9d8-2844dd8417d2 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Introducing the next generation of Claude — anthropic.com,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d4e413ca-aabc-45e0-92b2-55fb4cbdfb9e · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Gemma: Open Models Based on Gemini Research and Technology
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b15f156a-3aac-4470-a570-fb3a0d78245b · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Llama 3 model card,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dbe729f-675d-4788-ba68-023caeb7369a · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d7b9409-785c-44b4-912a-902f96bb8725 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b8c2180-fd1c-459e-9437-ea320f8ebbd7 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Measuring massive multitask language understanding,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcbdff8a-f26d-4bc1-b441-fb0214fe0104 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3fe7a0d-a675-4545-a842-7d5c137f3e71 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning The mvapich project: Transforming research into high-performance mpi library for hpc community,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 064a4e3b-3562-49de-9495-bfe45d86a64f · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Scalable Distributed DNN Training using TensorFlow and CUDA-Aware MPI: Characterization, Designs, and Performance Evaluation,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 954b19cf-6e28-4fac-b409-9ff30e210e7b · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning NVIDIA Collective Communications Library (NCCL),
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 02bcaf6a-9e55-4078-a489-a98d61c9662c · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3da99280-6328-4b4c-a270-735321a48d59 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9c65893-895e-43d9-84ff-ba1a9ebdca58 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Fairscale: A general purpose modular pytorch li- brary for high performance and large scale training,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7d60d192-106f-499a-bdca-3fb29ecc7704 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Megatron-LM: Ongoing research training transformer models at scale,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 29ad6087-e169-4d2c-85c3-f1d09dcbbedc · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Frontier - HPE Cray EX235a, AMD Optimized 3rd Generation EPYC 64C 2GHz, AMD Instinct MI250X, Slingshot-11 | TOP500 — top500.org,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f524f000-a4fe-4496-a184-36f228f07747 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning An in-depth analysis of the slingshot interconnect,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e6cee94d-9069-47d4-9335-4375fccf66f0 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning ZeRO++: Extremely Efficient Collective Communication for Giant Model Training
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18d3800b-6188-415e-b7f8-9f91ee29305c · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b89cf4b-5e8f-44d7-9ef6-76c09e6094e6 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Scaling single- image super-resolution training on modern hpc clusters: Early experi- ences,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 660dbb0d-91ae-41fb-bdd4-68fbb0370391 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Adam: A method for stochastic optimization,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4623ac3-e50c-431e-8457-f0b7523e542d · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning 8-bit Optimizers via Block-wise Quantization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e42ab56d-3118-4ed5-8023-88c78fc161a3 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Decoupled weight decay regularization,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26f6e6b-5293-47ed-8b43-ed9f35407131 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning GPT-NeoX-20B: An open-source autoregressive language model,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bfdced74-66cc-4749-82cc-914fa6cd931a · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Language models are few-shot learners,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa9efb2b-2a5e-4b36-a55e-8d563f5b3058 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning RWKV: Reinventing RNNs for the Transformer Era
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71fdde9f-5416-476f-9f80-77c0ee2a3147 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Pytorch fsdp: Experiences on scaling fully sharded data parallel,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 960dfec8-e83b-4ee4-8161-11a0675d1929 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Mics: Near-linear scaling for training gigantic model on public cloud,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 432033b2-76e4-4c6b-839f-cbbce1021da7 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b629f5e-1470-412e-a753-13dd96880d73 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12f5fa30-71b4-4dd6-8ae2-27c76559ba6e · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Optimizing Distributed Training on Frontier for Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e608e03-f496-4b9a-8945-d3c111500d6a · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Comparative Study of Large Language Model Architectures on Frontier
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 75297a7e-1ce8-4d1a-8f66-b58553ba9c54 · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Accelerating large language model training with hybrid gpu-based compression,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 70cee47f-4482-4411-b2cf-e42347847ace · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Hilfer fractional advection-diffusion equations with power-law initial condition; a Numerical study using variational iteration method
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3c524c6-630f-477b-b15e-2c946f85c94a · outbound
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning Measuring Massive Multitask Language Understanding
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.