Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-24T12:10:49.690618Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 55 inbound Pith citation observations for arXiv:2201.11990.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-24T12:10:49.690618Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:53:13.792040Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
78 of 78 outbound references displayed
External citation measurements
299
pith, observed 2026-08-05T02:28:24.338817Z
Observation c4a63d83-df42-47fd-aedc-fa7568234c1f · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model https://www.nvidia.com/en-us/data-center/a100/
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ddbc2dc-1fce-44b4-8cee-8c4ab5c08f44 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model https://www.top500.org/system/179842/
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5683aae2-8d70-469f-b2cf-48b660dc388c · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model https://www.nvidia.com/en-us/data-center/nvlink/
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec0bffb2-6f4a-440d-9eb8-f509b0b8dfff · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model https://www.microsoft.com/en-us/research/blog/ turing-nlg-a-17-billion-parameter-language-model-by-microsoft/
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ad2fdb52-dd7d-4acc-a816-91e5e9f271bb · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model https://wudaoai.cn/home
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5301a898-11bc-4d8e-bede-e78fe10659a5 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Piqa: Reasoning about physical commonsense in natural language
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 061dbb82-7aa5-433c-b3c6-f8ce0a0c28f4 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Zou, Venkatesh Saligrama, and Adam Tauman Kalai
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c3e9e4bd-8b95-4171-85e9-f3bfb7104923 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model On the Opportunities and Risks of Foundation Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 38e91e00-9566-4234-95cd-2e52230dc9e2 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 40fc54e9-1170-48a4-99a6-4c7894c4bedf · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model BoolQ: Exploring the surprising difficulty of natural yes/no questions
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0b3b913b-a9c5-48f8-8f97-1a96a0bc2af1 · outbound
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 97094877-c5db-418b-9804-aba9314b625d · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Bert: Pre-training of deep bidirectional transformers for language understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6bedb14c-37e8-4b98-9fc1-23b6561965bf · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Language and Gender
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0a2cab10-6e8a-4661-aaeb-31438aa8795f · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Improving Gender Fairness of Pre-Trained Language Models without Catastrophic Forgetting
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 188c7c7c-c2a4-4ac2-963e-06c2c914aecc · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf493b06-6eb2-490c-9a52-ed8b034d707e · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 95c74833-2394-4f52-9d56-76c3bd30f954 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 341446bc-7976-4cf0-a74f-82c12d1ad3f3 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model A framework for few-shot language model evaluation, September 2021
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 453a68ec-41cd-486a-8e1f-f183bea12a38 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Realtoxici- typrompts: Evaluating neural toxic degeneration in language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 231f2458-a0c7-4cc1-8d17-84179d2b88ee · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 81fc2180-f439-4102-b57d-87094c398071 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Bowman, and Noah A
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9ee7a11c-92cf-48bd-8e3d-fb43f4301d09 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Pretrained transformers improve out-of-distribution robustness
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3cd25986-83c6-45e4-a23e-a24a8d18f49b · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Gpipe: Efficient training of giant neural networks using pipeline parallelism
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a6dcb3a6-f388-4e35-99d4-376a46e86c7a · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Improving Machine Reading Comprehension with Single-choice Decision and Transfer Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a077427a-6faf-4744-8c8f-9acfd78bdd1c · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Weld, and Luke Zettlemoyer
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6e4ad3f1-64b6-4807-a5da-19b94718b99b · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Exploring the Limits of Language Modeling
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2f0e1829-96f9-43e7-8f7c-2136f1f5e30d · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Adam: A Method for Stochastic Optimization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1f4ce556-5d82-4d5b-ab90-36152535f120 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Gedi: Generative discriminator guided sequence generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 086efc00-3fc5-40d9-a52c-5b1a0567564b · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Dense-captioning events in videos
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dcf2d155-6709-455d-a93f-61efb86e1fba · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 94744b36-ea3e-4f7c-91b4-7485e170bee5 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model RACE: Large-scale ReAd- ing comprehension dataset from examinations
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7e4d6e39-99c8-4864-8525-f548a01cda0d · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d3b7ae9c-cd6d-447f-87fd-05b91b1129f0 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model The Power of Scale for Parameter-Efficient Prompt Tuning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 553acc1b-489e-4991-afe9-96627f8d4c91 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Jurassic-1: Technical details and evaluation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3f8851c0-dbbc-416f-9e39-fb38c754b508 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model M6: A Chinese Multimodal Pretrainer
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a6709e2-8ed1-478c-90d0-66943b39cc94 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model M6-10t: A sharing-delinking paradigm for efficient multi-trillion parameter pretraining
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4eb41430-3e27-4438-a046-df04adb8ab8e · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5092cdbb-bd20-43ab-8fd5-6bf253f1598e · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Unicorn on rainbow: A universal commonsense reasoning model on a new multitask benchmark
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5bc79588-2db3-4cdc-8f64-e82e16247f79 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Black is to criminal as cau- casian is to police: Detecting and removing multiclass bias in word embeddings
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 347773d3-4b81-4e67-a5de-1a5419679380 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Right for the wrong reasons: Diagnosing syntactic heuris- tics in natural language inference
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1a9df28b-c310-4be6-a43c-706ca4c966dc · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Mixed Precision Training
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0751dc8d-27f2-4f36-8e07-e6b601f4b40c · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Pipedream: generalized pipeline parallelism for dnn training
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a4d62a99-e85b-4e0d-86d5-17278d726ffe · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 171e47ce-a5d5-4175-8455-3b458f4b96e8 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Mitigating harm in language models with conditional-likelihood filtration
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b62bf687-faa3-4741-ba03-c1d63af740f4 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Transformers without Tears: Improving the Normalization of Self-Attention
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f5747286-3a03-4bc4-8303-b01f73c50a65 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Adversarial NLI: A New Benchmark for Natural Language Understanding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae6b4f09-0d1d-4406-8aa4-beb6478167e7 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Adversarial NLI: A new benchmark for natural language understanding
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 37000ecf-05fd-42dd-8bd6-cad015d9c07b · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 93a0fac9-b3a5-4b67-98ca-e34f5830cb80 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model The LAMBADA dataset: Word prediction requiring a broad discourse context
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 30d48419-ff8a-4163-a22b-5f828b5a9e04 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Wic: the word-in-context dataset for evalu- ating context-sensitive meaning representations
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e53022d9-3366-426c-8955-983383eb12fa · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Sentiwordnet
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1307093c-0a35-483f-9cdf-d657d3999aab · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Language models are unsupervised multitask learners
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 27b1282f-8fbe-4c2b-96f1-7128dba985ad · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c9ac42dd-898d-42f7-8216-4503ac9d465f · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5e5972b9-1778-4b5a-a8c5-5f7a752eb030 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Zero: Memory optimizations toward training trillion parameter models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 47c79d72-47c9-41b2-83e8-d4f6d4285e35 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 316adef9-851c-42ba-a34e-58244a5fe2ab · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Deepspeed: System optimiza- tions enable training deep learning models with over 100 billion parameters
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7f9cbea4-faac-407d-a3eb-90e30057f7dd · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Winogrande: An adver- sarial winograd schema challenge at scale
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 743884ad-3549-4c40-858f-14df0c2a1ed7 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Multitask Prompted Training Enables Zero-Shot Task Generalization
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c478c972-aa07-4af8-a28f-a2ed657972ba · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Self-Diagnosis and Self-Debiasing: A Proposal for Reducing Corpus-Based Bias in NLP
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 20423737-5423-46f1-8b2a-9bb7f7a3ffc8 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 627d563c-aaa2-40af-a782-20c3e82b85c1 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model The woman worked as a babysitter: On biases in language generation
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 223dae09-93cc-48c5-a3ea-bd3020d8071e · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 51722e09-3ae3-4119-bcde-b368d00c03c6 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9fbf2154-9722-4172-8a71-144dceb3d283 · outbound
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6cc1852f-d374-463e-924c-5a52840607e3 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model A Simple Method for Commonsense Reasoning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d6fef71-9e00-4580-be9f-9e29bd36edcc · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Attention Is All You Need
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a8585808-b885-480e-b8ee-e0a9b1efe396 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model InfoBERT: Improving Robustness of Language Models from An Information Theoretic Perspective
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 30bd5f71-2927-41e8-8907-6617bc661170 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Towards Zero-Label Language Learning
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f9efcea2-65d2-437e-adf1-d3b999de156d · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Finetuned Language Models Are Zero-Shot Learners
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 55f60e0e-276c-47c1-a39f-c3e23ec02e06 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Ethical and social risks of harm from Language Models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 186a5027-f14d-4885-85f4-8e785660a588 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Challenges in Detoxifying Language Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4cac1c9b-8be2-4863-b4da-6c735cd43960 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model A broad-coverage challenge corpus for sen- tence understanding through inference
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 724c78f5-694e-4ffe-afd8-5311deda88c0 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Yuan 1.0: Large-Scale Pre-trained Language Model in Zero-Shot and Few-Shot Learning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 807f0376-60e0-4910-ba28-ecc2a63210d5 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Learning and Evaluating General Linguistic Intelligence
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cdba02d7-1457-4bf5-839e-e6058e4d5a9d · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Hellaswag: Can a machine really finish your sentence? In ACL
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 41f90178-98a2-4fef-8b57-31015b5cf302 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Defending Against Neural Fake News
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a29f3d72-3466-4234-b011-c3164ddc53b9 · outbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model PanGu-$\alpha$: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 50ef9e93-dc16-4959-b069-f5a51622021d · inbound
PaLM: Scaling Language Modeling with Pathways Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 145
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a486b441-112b-4452-a22c-c8f449738476 · inbound
GPT-NeoX-20B: An Open-Source Autoregressive Language Model Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1b0f78c3-7254-406c-901b-7c0785a3cdbd · inbound
MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c1aa1546-e474-4276-80dd-173584b680f7 · inbound
OPT: Open Pre-trained Transformer Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 280
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ecb73857-13d3-4ed8-8a0d-c644e03c1c90 · inbound
Large Language Models are Zero-Shot Reasoners Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cd4c12af-5bee-44b1-ba33-16490830667f · inbound
Efficient Training of Language Models to Fill in the Middle Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8231ca83-a3a7-4cac-92ed-93c06e062dbb · inbound
Atlas: Few-shot Learning with Retrieval Augmented Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 099e2817-021f-4d81-bc17-e29e820820e7 · inbound
Atlas: Few-shot Learning with Retrieval Augmented Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 249
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a29b721d-82b8-4ac6-aa87-a8b8a8b4a1f0 · inbound
FP8 Formats for Deep Learning Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 52037305-cac3-4fb5-b16d-67b0121c8f9f · inbound
Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 63ab85ea-6a4e-4d41-a39b-2b6e521f4da2 · inbound
Language Models can Solve Computer Tasks Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 19dcd158-70b9-47a3-9321-656e5ac5fbf9 · inbound
BloombergGPT: A Large Language Model for Finance Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation faad2191-00ce-4e8b-8718-803da0a818a0 · inbound
A Survey of Large Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf6ab664-31cc-48a9-9a6b-e8cee9a02e2b · inbound
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d2bd2c8-6aed-4f51-814a-19248819316f · inbound
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 09e1623f-97a6-4ca5-99fd-2cdfb366953a · inbound
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 39312b7b-8b85-403a-82a7-e380fe41f8a8 · inbound
Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7f1f77ed-ed3c-49c1-b48b-820e49dda987 · inbound
StarCoder: may the source be with you! Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3ed52d22-268b-4b83-91c8-185a940931b4 · inbound
Scaling Data-Constrained Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 559e2ffa-4344-4b44-8ed1-33d5930b0bf0 · inbound
A Comprehensive Overview of Large Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6174c909-ccf6-4d60-8777-767adc0c54d6 · inbound
DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9ebc7711-7d7b-4a73-9054-aa5224e22477 · inbound
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6e875dfa-0357-496f-80de-72367519a446 · inbound
The Falcon Series of Open Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a7a3a822-d535-4b32-8004-9a7c8f3e918d · inbound
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 164
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ca5a0c81-b324-4ea3-b347-a8a154f6ebfe · inbound
Large Language Models: A Survey Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7af723dd-e793-4bde-b21d-2bc5f83e8c60 · inbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 162
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 95d34cb4-cc34-4d1c-b590-c1067d1ebeb8 · inbound
Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 68750bdd-ab4f-4e91-b674-5265d232da8f · inbound
SEDD: Scalable and Efficient Dataset Deduplication with GPUs Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4b54afb2-b5b9-488b-8db0-17fa88158b11 · inbound
MiniMax-01: Scaling Foundation Models with Lightning Attention Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d2031268-7be2-48ea-80f6-85d7535c0870 · inbound
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 247
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8526b1fe-2dca-4403-8acc-9cf99df5dfa8 · inbound
Language Models Improve When Pretraining Data Matches Target Tasks Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab3e7fe-6e51-40fb-9964-594166fb7d91 · inbound
The Carbon Cost of Conversation, Sustainability in the Age of Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90d1014f-f333-4cb1-9c37-0284d893ebb7 · inbound
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3356bce0-fe8f-4732-ac13-4893bee81912 · inbound
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 48c4c3d1-1bc5-4c7a-9ab4-81383c8a8a4c · inbound
Integrating Large Language Models with Network Optimization for Interactive and Explainable Supply Chain Planning: A Real-World Case Study Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ebde26c-f639-4595-bec5-41cf8ff961ba · inbound
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce99e3c-fdb0-4d9b-a830-b1114f96a3b9 · inbound
MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e4b7d6-fc68-497d-a019-8808945a2ab6 · inbound
Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 92cb1cd5-ba43-4c9d-b888-2ead73b2ced4 · inbound
Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50005be3-08b2-4d3c-8d04-6b26200316c7 · inbound
veScale-FSDP: Flexible and High-Performance FSDP at Scale Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9e152b6e-7278-46ae-af21-b87a2560c74c · inbound
M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9d979ed2-cb97-4d67-a004-6fafd0d4b770 · inbound
Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5ddb0c30-bb4e-4b90-924d-f362dcb321eb · inbound
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2e41a25a-1204-4377-b17e-0fd4c0a8fa19 · inbound
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0b7d393b-f9b7-4505-92e8-b080fc17d68f · inbound
Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8d33e553-0885-4dc1-b5ee-8dcc54f613a9 · inbound
A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9c53ad3d-384c-41fe-8417-da812e31e5ce · inbound
Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f24b1559-05c5-411b-98ef-1bc7bf1b1f74 · inbound
Phoenix-VL 1.5 Medium Technical Report Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e19f136e-48f6-4eb9-8d69-e1d27dc92593 · inbound
Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 529ce6db-d9d7-4fbf-a375-380928af73f6 · inbound
Large Language Model Selection with Limited Annotations Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 16bc8f80-e4bd-46ec-bb85-73c17165dfd3 · inbound
Heterogeneous Parallelism for Multimodal Large Language Model Training Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f451607f-3301-470c-b5c4-d2f56466ada0 · inbound
Specific Domain Ontology Construction Using Large Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 41110de5-87f0-4039-b6f1-6564292b6913 · inbound
CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dae08bfb-ab49-4177-b12c-1d65d6e46998 · inbound
SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2504b164-29fb-4361-bbf5-ae4a643ac76f · inbound
QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 248
Source-reported events for the cited work
Unavailable: canonical work link unavailable.