Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:34:44.807534Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 100 inbound Pith citation observations for arXiv:1909.08053.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:34:44.807534Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:04:51.559842Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
28 of 28 outbound references displayed
External citation measurements
826
pith, observed 2026-08-05T02:28:24.338817Z
Observation 0f8bb4d1-bb09-4ad7-b06d-538763b9e70a · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Layer Normalization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f426e9da-c57e-49f4-9d85-5dbfbb5dfa5d · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Training Deep Nets with Sublinear Memory Cost
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f594b516-3b61-4480-b3ca-ce64c3b4a0de · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 23f673ee-a086-4f61-bc9d-7679243b0416 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f99c712-d1bb-448e-a4d6-9732757e5846 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism PipeDream: Fast and Efficient Pipeline Parallel DNN Training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8c71e07a-a91d-4f96-ba2a-4b8ad9d56872 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Gaussian Error Linear Units (GELUs)
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e96f00de-6d16-4b2a-b9a5-91d6e5fcf5ba · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cbe37ed2-4c6e-457f-b06f-8b114db5ebc3 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism SpanBERT: Improving Pre-training by Representing and Predicting Spans
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0e5b8535-7186-4bda-9b80-a426ecdc938a · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Generalization through Memorization: Nearest Neighbor Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e80f894-356e-4ec8-8050-a30bb979415b · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Adam: A Method for Stochastic Optimization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 791dbb77-68b4-4464-8f52-f02c49a7d646 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism RACE: Large-scale ReAding Comprehension Dataset From Examinations
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e1a008e8-a390-4f8b-a586-78eec5e7c04f · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1755b56c-2157-4548-a092-a5ac8c2c3a2b · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Multi-Task Deep Neural Networks for Natural Language Understanding
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 28f73f52-0e3e-45bc-a273-fa66dcf955da · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Learned in Translation: Contextualized Word Vectors
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b449abbb-d6af-47a5-b95c-319adf15feac · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Pointer Sentinel Mixture Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7aee95bf-9046-4216-9b46-28e3acbc8b56 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Distributed Representations of Words and Phrases and their Compositionality
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fbd2b8c6-fdfb-4b3d-a7a0-61bac372cd73 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism The LAMBADA dataset: Word prediction requiring a broad discourse context
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 55968434-565c-492f-84f4-b9bfcd18854d · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Deep contextualized word representations
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6864d9ea-a007-44d1-afe5-972e74820c55 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 885fe66a-f0f7-4535-a22c-4f687f799def · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Unsupervised Pretraining for Sequence to Sequence Learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81e151ba-ea87-47c3-8ea4-fe6194a3d476 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism A Simple Method for Commonsense Reasoning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f290751-f657-4337-8785-ee0cc7b54b86 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Attention Is All You Need
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 89e3d504-5514-4529-a0d2-a9a705de9edd · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism XLNet: Generalized Autoregressive Pretraining for Language Understanding
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 13f72592-4f20-43a5-ab61-c1a3ff0cccf8 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ef5858c0-3f2a-4b7d-ae2b-5c2f722ede88 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Defending Against Neural Fake News
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb91cd72-2056-425d-8517-fc49774eaaa5 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d89a31a-bfcd-4ebc-aa1c-1a722f09c7c4 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Megatron-LM: With a broad scope, the conference ad- dresses the challenges and opportunities in machine learning for practitioners and researchers
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 079037d4-9cfe-42f9-b596-f0f3cafb1b32 · outbound
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 389f37ae-a7e7-43ca-8ce5-fc04e5d6d850 · inbound
HuggingFace's Transformers: State-of-the-art Natural Language Processing Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 181
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 75df99bf-b7de-470f-8b0e-5154a04a4374 · inbound
DeBERTa: Decoding-enhanced BERT with Disentangled Attention Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3cdc61d5-2f4f-4dd7-890d-d19e2ebc8c02 · inbound
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6059e413-b9db-4ba6-b45c-280509a840d6 · inbound
The Pile: An 800GB Dataset of Diverse Text for Language Modeling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ff2d19a5-333a-40db-91de-4cb9141bc090 · inbound
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7050cae9-5728-4145-b7c9-0ff1e6b8b9a8 · inbound
GSPMD: General and Scalable Parallelization for ML Computation Graphs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4a9a2c26-21ba-4384-9fbf-d7d9e3617763 · inbound
MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9fb6bc31-c5dc-4850-8bfa-047be6f058f4 · inbound
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aa35fa9b-c5ea-475f-84da-285eec412023 · inbound
Improving language models by retrieving from trillions of tokens Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 223dae09-93cc-48c5-a3ea-bd3020d8071e · inbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ea891404-2b92-42a4-bffa-bb49cf382547 · inbound
ST-MoE: Designing Stable and Transferable Sparse Expert Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e1b85adc-9af5-4395-97ab-b5e450bb0c5a · inbound
GPT-NeoX-20B: An Open-Source Autoregressive Language Model Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2722e53d-f30b-46c4-9266-f3fc840b341f · inbound
OPT: Open Pre-trained Transformer Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 141
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 74e63fb4-b2d3-429b-a758-bc9ebd576f79 · inbound
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ab57fee7-4f0e-4dce-b269-101f1b48e102 · inbound
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed668f91-b14f-41ff-af84-2f0bb5c72654 · inbound
PaLI: A Jointly-Scaled Multilingual Language-Image Model Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 01f93159-9b7e-4a3c-a612-250d99f198b3 · inbound
Language Is Not All You Need: Aligning Perception with Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 94bf987a-8a5d-4b5d-9d0d-fa318d7a604d · inbound
Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d240440c-9387-4056-9696-722de6a5387c · inbound
BloombergGPT: A Large Language Model for Finance Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5ffdfcdb-11fd-4139-bb69-0a575b0cb37c · inbound
A Survey of Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1e419a38-bd4e-459d-931c-d5f8d48ada3f · inbound
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6c1b631f-a8f0-47c8-9292-23e6d2c9f796 · inbound
StarCoder: may the source be with you! Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 162
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5d248830-30f7-4ef7-8b8d-dbafa6164844 · inbound
Scaling Data-Constrained Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e2758a1-29b0-475c-a3fa-98b450602b02 · inbound
A Comprehensive Overview of Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4d5e1dc2-610b-452f-b15e-cdc0e1093114 · inbound
Retentive Network: A Successor to Transformer for Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 370a420a-6b78-4548-8b44-ded2bda78287 · inbound
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a8d9d635-39ef-4241-8838-2d62a9677aad · inbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3cb0556b-ffb3-46a6-a7b9-f38e2f5426be · inbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 251524cd-c8d7-4865-821e-de0d462d0fb2 · inbound
DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c38da3a3-dc33-4426-ad05-d412605ed2ef · inbound
Demystifying CLIP Data Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4e48f560-9d38-4459-859f-884de148e8c2 · inbound
Ring Attention with Blockwise Transformers for Near-Infinite Context Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a0d90def-7f63-4854-8278-14a52c9f5640 · inbound
Llemma: An Open Language Model For Mathematics Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a04fcf6-8bc0-4098-ae6a-9d3ed5b5d8f6 · inbound
BitNet: Scaling 1-bit Transformers for Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b0cd216-b39c-457e-90d1-776d007642af · inbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation afc40ad2-a557-48b6-9475-9a272e659673 · inbound
DA-Cramming: Enhancing Cost-Effective Language Model Pretraining with Dependency Agreement Integration Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 07887c75-ef06-4fb9-af96-67573407701f · inbound
MEDITRON-70B: Scaling Medical Pretraining for Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e4d5640-482b-42da-a438-e5fa3f194924 · inbound
The Falcon Series of Open Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b9d0f233-c377-4ed9-8dfa-7df0dfe79328 · inbound
SGLang: Efficient Execution of Structured Language Model Programs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6a126239-b469-4f5a-b9d3-f1d1d07c6928 · inbound
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cfda3446-ac63-42b3-a9ed-9e8c4ac20d4d · inbound
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3ff9f499-89a9-46ad-9ed4-a3bdf03cfa41 · inbound
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e2b0f7d1-b5ba-46c8-9814-668b885c69de · inbound
Yi: Open Foundation Models by 01.AI Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6852170-6426-413d-bbd0-33776e220ae5 · inbound
DeepSeek-VL: Towards Real-World Vision-Language Understanding Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9160f9dd-6a1f-416d-b99a-d93e8edfe5d5 · inbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6373220f-520c-4288-8962-0666c4ec80a0 · inbound
InternLM2 Technical Report Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cfb546d7-68fa-432f-9030-7f2b14d9efae · inbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8026b487-4382-4610-9809-6d3fa8e08cd3 · inbound
PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b08bb3ce-8834-47b5-b6c5-8c80c58f6bbb · inbound
Lessons from the Trenches on Reproducible Evaluation of Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ceaa746-df02-42ae-935b-600396963f68 · inbound
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 335f17ad-e390-4060-8c26-752624ce4bc8 · inbound
Scaling and evaluating sparse autoencoders Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 77e44c69-3974-4f3f-9753-4600f4efe9c1 · inbound
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3451654-3f18-4fd8-bf63-f92420301084 · inbound
An Empirical Study of Mamba-based Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6196ac6d-2aca-4646-8cbb-3d072772567b · inbound
ProTrain: Efficient LLM Training via Memory-Aware Techniques Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cc674576-a2db-44d4-83a3-0a3e329d3f22 · inbound
PaliGemma: A versatile 3B VLM for transfer Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 755bcfab-8b96-441c-98ba-e66b641fc83a · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 871aeb76-d8d1-44a5-bdfc-d6499bad7856 · inbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bb16271e-5596-45ec-a8ee-a277a1b0fdef · inbound
HybridFlow: A Flexible and Efficient RLHF Framework Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e2a7c3e8-d3c0-4ab9-8695-9d710bf77841 · inbound
Movie Gen: A Cast of Media Foundation Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fb5c6eb2-2b62-45cf-ac0f-ad96dadad573 · inbound
On the Convergence Theory of Pipeline Gradient-based Analog In-memory Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 89cc7654-e2eb-4627-99f1-1b3a57559c73 · inbound
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 75907270-ba0a-48fd-912c-f78279ef2539 · inbound
Open-Sora Plan: Open-Source Large Video Generation Model Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5652ae3-f68e-4c14-84de-88f417e5fa37 · inbound
HunyuanVideo: A Systematic Framework For Large Video Generative Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 946b9d6b-8289-4cb5-97e6-8bbc958d2725 · inbound
TrainMover: An Interruption-Resilient Runtime for ML Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 59c4eda9-a28f-4e02-866b-173f20e5c922 · inbound
Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 185
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b6d43cd-efef-4b6d-b1de-ca28eab11800 · inbound
Cosmos World Foundation Model Platform for Physical AI Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8d15554c-d71a-4ef7-8c1c-fdf1cccd2c44 · inbound
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 002410d4-50e7-4c87-953f-205dc12df36f · inbound
MoBA: Mixture of Block Attention for Long-Context LLMs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 00bd2f63-38f4-42ae-bee6-f32946035640 · inbound
Green Prompting: Characterizing Prompt-driven Energy Costs of LLM Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d04ed5ec-f072-463e-b8b1-04f823f5bc93 · inbound
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 697b8262-a874-4ed8-b1d6-093e5e211755 · inbound
Advancing Intelligent Sequence Modeling: Evolution, Trade-offs, and Applications of State- Space Architectures from S4 to Mamba Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0e5c7601-8bbc-46ce-ad11-5ff741ea30e1 · inbound
Wan: Open and Advanced Large-Scale Video Generative Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8ccd6446-f1b5-496e-884f-fe66475fdb6a · inbound
Seed1.5-VL Technical Report Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a26afc32-6622-45ba-bdde-a04dee61efd2 · inbound
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9b490396-40b7-48ba-8703-c5eb8ac69ba5 · inbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 45c9ac0f-4a4b-45db-a568-ef3cddad83c8 · inbound
Geminet: Learning the Duality-based Iterative Process for Lightweight Traffic Engineering in Changing Topologies Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5e77505-c9de-4905-866d-c18ae043affe · inbound
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 169
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 19100ab3-2068-482c-8842-02ec7c733e0c · inbound
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 43daee91-c8b3-49af-be81-80a9a82f575a · inbound
WebSailor: Navigating Super-human Reasoning for Web Agent Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b0ccf34-b347-441f-8bc9-9dbc276bfc8e · inbound
The Serial Scaling Hypothesis Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 750ca42e-09b8-4e63-b44a-5a3822f002b0 · inbound
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 84dbb7d1-f4ee-41aa-b0a5-294368c681bb · inbound
Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a6dc1feb-bd98-4ba2-8553-c4c2cdd45b4c · inbound
Qwen-Image Technical Report Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9ba23427-8128-4646-99b4-5e853ce4bb3b · inbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78357e2c-19ba-4088-9ab8-930972083114 · inbound
Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28f25519-5370-4395-94e5-818027fc61df · inbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 643b31f4-664d-413d-b0c8-932a5bc5b3e5 · inbound
Power Stabilization for AI Training Datacenters Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 626d1528-c68b-4089-badd-9827848560c0 · inbound
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 69d8ba3e-4312-4cc6-a132-ffc9930d9915 · inbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1881a62-7eeb-4f60-8776-01134f216dbb · inbound
HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ca940f0a-03b5-4b78-ad19-859c10ef6793 · inbound
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5678c3e7-8d6f-494f-b41c-ad52a791efca · inbound
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3fbcb5b-5d85-41fa-b4a8-1d941a91c40c · inbound
Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55a7683b-3298-4207-b693-dc4fc84b9e36 · inbound
LobRA: Multi-tenant Fine-tuning over Heterogeneous Data Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 705d33f3-c8df-4c0f-81e4-0d44d97bb250 · inbound
MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b9b11c8-8cbe-43be-bb07-3eb2c68644b8 · inbound
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eaecbbb-1ed9-43e1-af8f-d6be9c80a590 · inbound
SpikingBrain: Spiking Brain-inspired Large Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f209819-099f-4cb4-9fa6-f9f7e7de0fd1 · inbound
veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03e444b3-3b0c-4c29-94c4-be4a7f71b65e · inbound
Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ff59f64d-5fb3-4901-82dd-7dd9e8368205 · inbound
Safe and Certifiable AI Systems: Concepts, Challenges, and Lessons Learned Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c005d75-46d8-4b02-8e66-8267716ce092 · inbound
HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.