Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2104.04473.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T11:25:09.562026Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
9
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation a4d62a99-e85b-4e0d-86d5-17278d726ffe · inbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 50c00f4f-b34c-4082-acab-2ff8b53dd1d2 · inbound
OPT: Open Pre-trained Transformer Language Models Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 142
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 98915eea-eaf6-4490-9f89-fa2b7b9db928 · inbound
SpikingBrain: Spiking Brain-inspired Large Models Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fdb78c4e-ce1a-4b43-aee2-89eaf78d9c06 · inbound
Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5016bb78-fabd-40f9-90ce-84c305794c98 · inbound
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 24b707cb-bacc-4ff0-948d-eaa90040a26f · inbound
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 19c9e4ea-19e8-4fd7-82ab-82d278e94d87 · inbound
Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 328733c0-12be-44c1-8352-58b97e7d75af · inbound
Kimi K2.5: Visual Agentic Intelligence Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 177b40ef-d995-4946-a80f-645b9db7e617 · inbound
Attention Residuals Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c9817349-2284-48aa-8094-16363d79dabe · inbound
AEGIS: Scaling Long-Sequence Homomorphic Encrypted Transformer Inference via Hybrid Parallelism on Multi-GPU Systems Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6a0912e7-1778-41aa-9a95-863b705e390c · inbound
An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b02ba0ef-36db-49ff-9907-d9d5d8aa0421 · inbound
Nautilus: An Auto-Scheduling Tensor Compiler for Efficient Tiled GPU Kernels Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b98fd996-021b-4bf4-bd1b-e32939204f3b · inbound
GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fbf8fc09-41b9-4bdb-b3b7-64885625d174 · inbound
Efficient Training on Multiple Consumer GPUs with RoundPipe Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4459bbc7-997e-4ae8-bc91-8ddb8ad82895 · inbound
A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d8715644-38f9-4b6f-9f37-30e957cda438 · inbound
Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e29785d-d1b4-451b-88e8-0ea4b476cde0 · inbound
MinT: Managed Infrastructure for Training and Serving Millions of LLMs Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 66d0774e-14e5-4333-ad20-0313dfe62708 · inbound
MinT: Managed Infrastructure for Training and Serving Millions of LLMs Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 98a325aa-8eea-4233-abb5-1b7f161c9428 · inbound
Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 74ea8c6e-298e-4381-a6e7-301a64b72879 · inbound
Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 79497a54-12f9-423d-aafb-370042858c82 · inbound
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 12c7c42b-e7ee-4c9e-a26f-e5612c9b589e · inbound
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 18dd1c7d-1d54-4fb3-b41d-d0ad1f22525a · inbound
Heterogeneous Parallelism for Multimodal Large Language Model Training Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 702f9e44-561d-495f-a5ef-ae30dc9bc281 · inbound
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c6e3757-cf93-446d-827b-e7eeb1a5facb · inbound
Model Multiplicity for Adversarial Detection in Small Language Model Training on Edge Devices Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8ac83b28-42d8-4410-8f26-1dd54aa5ebbe · inbound
Piper: A Programmable Distributed Training System Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 132e3e92-7590-4d9e-87c2-8672827fad05 · inbound
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 225
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation db29d1e9-7eb4-4784-8506-b0ee07621a42 · inbound
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 211
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78b96fde-1b39-415a-b5e3-11a9ee3918da · inbound
PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e857b2bb-5386-442e-8cc5-6b2d1d92bfee · inbound
PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e919841b-30d3-4e66-a2d1-e594d06b6753 · inbound
Design-CP: Context Parallelism for Design of Protein Nanoparticles Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ed52c58-f009-466d-870c-16cd34446578 · inbound
GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 00d01956-3547-4166-88a1-c427c5c4440e · inbound
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73494bae-8a1d-4c9f-a9ac-3a3c3d90d1a0 · inbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.