Pith. sign in

Paper Citation Record · LEDGER

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

As of 6 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 55 inbound Pith citation observations for arXiv:2201.11990.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2201.11990 v3

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-24T12:10:49.690618Z

measured 133 of 133 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 55 of 55 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:53:13.792040Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact30
  • verified fuzzy40
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

299
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c4a63d83-df42-47fd-aedc-fa7568234c1f · outbound

This paper cites https://www.nvidia.com/en-us/data-center/a100/.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model https://www.nvidia.com/en-us/data-center/a100/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:14:28.214658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:4e9e462cdaa7db62ae25a73caa34ee12e72ca63adc0bb9092277134b644554f9

Observation 4ddbc2dc-1fce-44b4-8cee-8c4ab5c08f44 · outbound

This paper cites https://www.top500.org/system/179842/.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model https://www.top500.org/system/179842/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.674946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:887a0024bbe00b2603815dc9da972fbb75ff278d6b87e3ed592ea2e02e81d873

Observation 5683aae2-8d70-469f-b2cf-48b660dc388c · outbound

This paper cites https://www.nvidia.com/en-us/data-center/nvlink/.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model https://www.nvidia.com/en-us/data-center/nvlink/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.652197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:5b637081e34aa3b1ca88a9849b4f610b68218b7186428727b27e257dd79d1478

Observation ec0bffb2-6f4a-440d-9eb8-f509b0b8dfff · outbound

This paper cites https://www.microsoft.com/en-us/research/blog/ turing-nlg-a-17-billion-parameter-language-model-by-microsoft/.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model https://www.microsoft.com/en-us/research/blog/ turing-nlg-a-17-billion-parameter-language-model-by-microsoft/

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.623132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:a80bb204db7725e2258cccacaebfc3923d225c4ca6bcc121d7a735414dcbc175

Observation ad2fdb52-dd7d-4acc-a816-91e5e9f271bb · outbound

This paper cites https://wudaoai.cn/home.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model https://wudaoai.cn/home

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.632939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:b216dfd4231c22e97a158a8d52529befe6d6e648a6c38a15896e7a83c97551c0

Observation 5301a898-11bc-4d8e-bede-e78fe10659a5 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Piqa: Reasoning about physical commonsense in natural language

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.615814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:f0711832e0d9f90bd08e263c81bb62f747ab8221cdfc0892497230559248ada3

Observation 061dbb82-7aa5-433c-b3c6-f8ce0a0c28f4 · outbound

This paper cites Zou, Venkatesh Saligrama, and Adam Tauman Kalai.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Zou, Venkatesh Saligrama, and Adam Tauman Kalai

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.711307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:790a96648e2a39e4ccaff9af26881665b6f69d0072c37541aee184849e2c72ce

Observation c3e9e4bd-8b95-4171-85e9-f3bfb7104923 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model On the Opportunities and Risks of Foundation Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.642555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:ed0653dd17cedd228e5ae95ed9eb25b601fb00451ca62ec2a6166470957853fe

Observation 38e91e00-9566-4234-95cd-2e52230dc9e2 · outbound

This paper cites an unresolved cited work.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-24T12:16:11.701886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:5435d4c223abc7530cfbe2e1d1c0ac27b322aac6b4aa334366510f199dab53d5

Observation 40fc54e9-1170-48a4-99a6-4c7894c4bedf · outbound

This paper cites BoolQ: Exploring the surprising difficulty of natural yes/no questions.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model BoolQ: Exploring the surprising difficulty of natural yes/no questions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.699086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:9dc49017a50a72d027022dbeda9b2be3b249ee607f9c34dec202102232e1c9a9

Observation 0b3b913b-a9c5-48f8-8f97-1a96a0bc2af1 · outbound

This paper cites langdetect.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model langdetect

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.695843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:1fa608f62d424b5237623d18ec8da79d8c95066219f00d991e6173a50fdb5569

Observation 97094877-c5db-418b-9804-aba9314b625d · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.692449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:c22aec5df0378a9462a41aa2fd35c9e126768b990b79c9e5038a0e445d2d9796

Observation 6bedb14c-37e8-4b98-9fc1-23b6561965bf · outbound

This paper cites Language and Gender.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Language and Gender

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.685180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:118897cb0b75fbf126f63fdf57900ec9ad8074d39e867e012573f6f5d8597353

Observation 0a2cab10-6e8a-4661-aaeb-31438aa8795f · outbound

This paper cites Improving Gender Fairness of Pre-Trained Language Models without Catastrophic Forgetting.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Improving Gender Fairness of Pre-Trained Language Models without Catastrophic Forgetting

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.529058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:cd58e1890edc65800984aeb91dcc4038840a7595b227366823e500965b10bcbe

Observation 188c7c7c-c2a4-4ac2-963e-06c2c914aecc · outbound

This paper cites Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.554412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:aee6ffd0308713f4910b0c8c4858047268fffec5e848214ccbe51bfd9ac30f48

Observation bf493b06-6eb2-490c-9a52-ed8b034d707e · outbound

This paper cites an unresolved cited work.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-24T12:16:11.678098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:6b0af8b774e5aacbf6bab651466303c06e82149d88132b02c57916f716762074

Observation 95c74833-2394-4f52-9d56-76c3bd30f954 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.617941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:4879068cf322be555c0b7c7030e8c6ffc475beb800921afddeb968c083a05619

Observation 341446bc-7976-4cf0-a74f-82c12d1ad3f3 · outbound

This paper cites A framework for few-shot language model evaluation, September 2021.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model A framework for few-shot language model evaluation, September 2021

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.681607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:65667d28a53443ca7e0c4887836e89de909f821db8961f40f4502bf607eb9285

Observation 453a68ec-41cd-486a-8e1f-f183bea12a38 · outbound

This paper cites Realtoxici- typrompts: Evaluating neural toxic degeneration in language models.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Realtoxici- typrompts: Evaluating neural toxic degeneration in language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.666278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:b8687c1425ab17f4b63f0a49f37455d49ccbcc446cf6d9139b68d56b6d1b0563

Observation 231f2458-a0c7-4cc1-8d17-84179d2b88ee · outbound

This paper cites an unresolved cited work.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-24T12:16:11.669759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:64a084209110cb59163d835c8b10b13bb0ba3c04f29721b6a92d794d6cf2372f

Observation 81fc2180-f439-4102-b57d-87094c398071 · outbound

This paper cites Bowman, and Noah A.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Bowman, and Noah A

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.612862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:4813ec563d50a11c2f4ceb243171c784d916b3bda9dfcd69db3ddfe48d9e1017

Observation 9ee7a11c-92cf-48bd-8e3d-fb43f4301d09 · outbound

This paper cites Pretrained transformers improve out-of-distribution robustness.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Pretrained transformers improve out-of-distribution robustness

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.655722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:b8dc6fa12657457ef3fa552bf6992d4c12ee6b6e00121f2282381b37138b658b

Observation 3cd25986-83c6-45e4-a23e-a24a8d18f49b · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Gpipe: Efficient training of giant neural networks using pipeline parallelism

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.618954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:c457335c43f2be98c7f119b8b20605d2bda1a3e747e92e948140f530b8a09ca2

Observation a6dcb3a6-f388-4e35-99d4-376a46e86c7a · outbound

This paper cites Improving Machine Reading Comprehension with Single-choice Decision and Transfer Learning.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Improving Machine Reading Comprehension with Single-choice Decision and Transfer Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.548306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:94b89acc3075ed30c3898ab1f42a928bffb0175abad39691df182f67456593f1

Observation a077427a-6faf-4744-8c8f-9acfd78bdd1c · outbound

This paper cites Weld, and Luke Zettlemoyer.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Weld, and Luke Zettlemoyer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.629590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:4923e07e5ebbaf6c34aa0bf63917b47b0d77a1674cf10191a78b8f285fd48ab1

Observation 6e4ad3f1-64b6-4807-a5da-19b94718b99b · outbound

This paper cites Exploring the Limits of Language Modeling.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Exploring the Limits of Language Modeling

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T12:14:26.671198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:5c4a4ff6935a0f6eb6a17aed4565709f872d8e78f45adc5e89f486f3ac8454af

Observation 2f0e1829-96f9-43e7-8f7c-2136f1f5e30d · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Adam: A Method for Stochastic Optimization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.590744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:6b1f873235bd9c08f6f8269bd7e1ae6f02e3ef65e8d325925c0bfef876a6fd6c

Observation 1f4ce556-5d82-4d5b-ab90-36152535f120 · outbound

This paper cites Gedi: Generative discriminator guided sequence generation.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Gedi: Generative discriminator guided sequence generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.638662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:fb3669a65dfc0b4fc92b40d5c2b64f14b5f316b41e0a1f4716abaa2e8594e79f

Observation 086efc00-3fc5-40d9-a52c-5b1a0567564b · outbound

This paper cites Dense-captioning events in videos.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Dense-captioning events in videos

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.626484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:19d6127f14ddef67f46f6c3cf377d30f278b9eecdb3e31f782d273ad8ca62d37

Observation dcf2d155-6709-455d-a93f-61efb86e1fba · outbound

This paper cites Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.642577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:86194c37ec57b51a23c4fa5c3b1a2826ce960b537e844a509e42c5e2d17fccf7

Observation 94744b36-ea3e-4f7c-91b4-7485e170bee5 · outbound

This paper cites RACE: Large-scale ReAd- ing comprehension dataset from examinations.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model RACE: Large-scale ReAd- ing comprehension dataset from examinations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.648792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:d103865b97a46c3027a472dbd6d2e6b4133a084a690f9fbc78a9e02695ee566d

Observation 7e4d6e39-99c8-4864-8525-f548a01cda0d · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.566178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:bf97b910d06b909dc0b6ebb1d23a2130bc368135727f59fbf2e6717ff970a931

Observation d3b7ae9c-cd6d-447f-87fd-05b91b1129f0 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.637272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:23a5151205f7205fd87f7eb4db1a66021f83c9a1db14f65ca5b8edaf8d68007a

Observation 553acc1b-489e-4991-afe9-96627f8d4c91 · outbound

This paper cites Jurassic-1: Technical details and evaluation.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Jurassic-1: Technical details and evaluation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.659837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:4716041b6b5714fa246f6051d6be4152faca88364a0802ba660ff990277b0216

Observation 3f8851c0-dbbc-416f-9e39-fb38c754b508 · outbound

This paper cites M6: A Chinese Multimodal Pretrainer.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model M6: A Chinese Multimodal Pretrainer

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.631436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:72e9f4571f74ce9b05c4390489f94d207063ee5368515f7611b1f59bf7b23f7c

Observation 8a6709e2-8ed1-478c-90d0-66943b39cc94 · outbound

This paper cites M6-10t: A sharing-delinking paradigm for efficient multi-trillion parameter pretraining.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model M6-10t: A sharing-delinking paradigm for efficient multi-trillion parameter pretraining

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.663013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:34c8f1e447f099dd9d6b7a893f0d4a5f10f62d29795d85056af17fdff3c1b10e

Observation 4eb41430-3e27-4438-a046-df04adb8ab8e · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.713341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:b7a10d7bd92c8fbdc78f2703977014790d23fd890053f32e7ddde966591c728e

Observation 5092cdbb-bd20-43ab-8fd5-6bf253f1598e · outbound

This paper cites Unicorn on rainbow: A universal commonsense reasoning model on a new multitask benchmark.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Unicorn on rainbow: A universal commonsense reasoning model on a new multitask benchmark

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.609969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:57a7922df4e39da9f910e8cda732bc19cfa961b25e3d43ba5b82325e9dacacc9

Observation 5bc79588-2db3-4cdc-8f64-e82e16247f79 · outbound

This paper cites Black is to criminal as cau- casian is to police: Detecting and removing multiclass bias in word embeddings.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Black is to criminal as cau- casian is to police: Detecting and removing multiclass bias in word embeddings

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.609179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:f017c57c9f732227983233d6e6bd17d6d8d7eaa310d406dc09dd1c4795709f55

Observation 347773d3-4b81-4e67-a5de-1a5419679380 · outbound

This paper cites Right for the wrong reasons: Diagnosing syntactic heuris- tics in natural language inference.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Right for the wrong reasons: Diagnosing syntactic heuris- tics in natural language inference

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.602015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:eefa153cdf71a190c77b32103d8326db0b240d9ab3eceba14cf43fd11840b65b

Observation 1a9df28b-c310-4be6-a43c-706ca4c966dc · outbound

This paper cites Mixed Precision Training.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Mixed Precision Training

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.665627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:d2d4a9ac1383dcf94d7192ed499f06848d592151cba4338364df96654aa9a345

Observation 0751dc8d-27f2-4f36-8e07-e6b601f4b40c · outbound

This paper cites Pipedream: generalized pipeline parallelism for dnn training.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Pipedream: generalized pipeline parallelism for dnn training

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.708368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:663eced32b2acaa43bc0bc1833e08019544be83edc9718371f1bda040503ecdd

Observation a4d62a99-e85b-4e0d-86d5-17278d726ffe · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.541851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:83a19689434ef6c88478b63f10d7e0ad2ed453d34afa5df4797f180e3a2fc664

Observation 171e47ce-a5d5-4175-8455-3b458f4b96e8 · outbound

This paper cites Mitigating harm in language models with conditional-likelihood filtration.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Mitigating harm in language models with conditional-likelihood filtration

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.660621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:aabb275c867fa496a57a4e61b921dcb237eb1f8e707fd876e9ff855768cc209b

Observation b62bf687-faa3-4741-ba03-c1d63af740f4 · outbound

This paper cites Transformers without Tears: Improving the Normalization of Self-Attention.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Transformers without Tears: Improving the Normalization of Self-Attention

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.535552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:c5fa45d003d9a40065e70be5995731492c1faa081e802b95e0024dbc11e62b94

Observation f5747286-3a03-4bc4-8303-b01f73c50a65 · outbound

This paper cites Adversarial NLI: A New Benchmark for Natural Language Understanding.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Adversarial NLI: A New Benchmark for Natural Language Understanding

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.572461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:4f9db5ccdc4d36e877b24b422a6a4a93f04396e33e72d06772e778c2b48d82ac

Observation ae6b4f09-0d1d-4406-8aa4-beb6478167e7 · outbound

This paper cites Adversarial NLI: A new benchmark for natural language understanding.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Adversarial NLI: A new benchmark for natural language understanding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:16:11.705336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:39938d51defdd8ad76f0fcd92ba55441bf372289f72730adb60a32ea061698a9

Observation 37000ecf-05fd-42dd-8bd6-cad015d9c07b · outbound

This paper cites Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:14:28.265669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:aef3533b8476da761024a4528197f2dfc0bad256270aeeed390f31bf97afbbd0

Observation 93a0fac9-b3a5-4b67-98ca-e34f5830cb80 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:14:28.261856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:ec25de0c48906d594071fcc48e6dcce8af4ca5f4dae96084c65e41a93cb98c51

Observation 30d48419-ff8a-4163-a22b-5f828b5a9e04 · outbound

This paper cites Wic: the word-in-context dataset for evalu- ating context-sensitive meaning representations.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Wic: the word-in-context dataset for evalu- ating context-sensitive meaning representations

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:14:28.257883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:b6b877f9debe9c22f1f7fb2422fe51beaccef09cd4b7e19e9aa0926bc64eecd5

Observation e53022d9-3366-426c-8955-983383eb12fa · outbound

This paper cites Sentiwordnet.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Sentiwordnet

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:14:28.253471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:47c2f5c891f8155069c28d0489373f3f692ef372dfda41d7b69277b4115e6024

Observation 1307093c-0a35-483f-9cdf-d657d3999aab · outbound

This paper cites Language models are unsupervised multitask learners.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Language models are unsupervised multitask learners

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:14:28.249270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:06ef120259222907c83dde0b4c1096e9cb1d6146354d006710eee48880a2fc55

Observation 27b1282f-8fbe-4c2b-96f1-7128dba985ad · outbound

This paper cites an unresolved cited work.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-24T12:14:28.244481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:f624e2fc6b58914bff26e1f49a0c983688c08fd3f1e4829382daafc956f3fbbb

Observation c9ac42dd-898d-42f7-8216-4503ac9d465f · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.516228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:1f66e76e52c7511ebaecb72e43c0e8cf5a1a03fc564ed0858506279d2af1a5b3

Observation 5e5972b9-1778-4b5a-a8c5-5f7a752eb030 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Zero: Memory optimizations toward training trillion parameter models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:14:28.240753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:38db417fb820919383630da2ffffdaa48559d0527868bfc15d5eb5d9e537b63e

Observation 47c79d72-47c9-41b2-83e8-d4f6d4285e35 · outbound

This paper cites ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.603797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:7216d403f01dd74691bd53db6f299900f4d7051a2e029f843b2707f8be7c1736

Observation 316adef9-851c-42ba-a34e-58244a5fe2ab · outbound

This paper cites Deepspeed: System optimiza- tions enable training deep learning models with over 100 billion parameters.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Deepspeed: System optimiza- tions enable training deep learning models with over 100 billion parameters

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:14:28.236509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:6238ed9a8ef65c8e303791d212316d429d650ad26d32b752f35c23937cade49a

Observation 7f9cbea4-faac-407d-a3eb-90e30057f7dd · outbound

This paper cites Winogrande: An adver- sarial winograd schema challenge at scale.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Winogrande: An adver- sarial winograd schema challenge at scale

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:14:28.230869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:e0544d726cc51d088c9849510258f31bf4d4c49ff8b50021af070b49187b63eb

Observation 743884ad-3549-4c40-858f-14df0c2a1ed7 · outbound

This paper cites Multitask Prompted Training Enables Zero-Shot Task Generalization.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Multitask Prompted Training Enables Zero-Shot Task Generalization

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.597618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:f0600a62e090fdfa40b7773ee4118d09298ef074cad7ec274968144a9499da28

Observation c478c972-aa07-4af8-a28f-a2ed657972ba · outbound

This paper cites Self-Diagnosis and Self-Debiasing: A Proposal for Reducing Corpus-Based Bias in NLP.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Self-Diagnosis and Self-Debiasing: A Proposal for Reducing Corpus-Based Bias in NLP

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.584905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:037fbd5d98e722c870a4394881e34e1d3d2a837e117a43f3d0ace3b2fb223998

Observation 20423737-5423-46f1-8b2a-9bb7f7a3ffc8 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T12:14:26.578410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:09525c030d1b7269a5710f93a37f9e0c983463fb09555e774912c3a23bda0563

Observation 627d563c-aaa2-40af-a782-20c3e82b85c1 · outbound

This paper cites The woman worked as a babysitter: On biases in language generation.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model The woman worked as a babysitter: On biases in language generation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:14:28.226417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:240687e6498a1a080b650ac8d751b27d49e2de54333dc518b509e1a641e42925

Observation 223dae09-93cc-48c5-a3ea-bd3020d8071e · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.694969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:3fbb6b05c16044a27ee8f13ecec48190c20d203cd58e884617cab0af619b2ff7

Observation 51722e09-3ae3-4119-bcde-b368d00c03c6 · outbound

This paper cites an unresolved cited work.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-24T12:14:28.222473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:139b3375c23ca0581a7192fe28c5670663c039e77d6afaeb48a3b408e2195b90

Observation 9fbf2154-9722-4172-8a71-144dceb3d283 · outbound

This paper cites DeepSpeed.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model DeepSpeed

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:14:28.218694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:d33a6bdc39e201d5ab2a0e006b319ce7db4b6db3a60bb4bcabd58ed0c01ee3a1

Observation 6cc1852f-d374-463e-924c-5a52840607e3 · outbound

This paper cites A Simple Method for Commonsense Reasoning.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model A Simple Method for Commonsense Reasoning

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.560624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:05ea7094d311288a7cd242954f11068a970903667308011c247f4cbc0fdc798c

Observation 1d6fef71-9e00-4580-be9f-9e29bd36edcc · outbound

This paper cites Attention Is All You Need.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Attention Is All You Need

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.653511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:8c74c375e17ea3ac696a75dd44c721996ef7b0c969357c3ae5fb8c2aa40cc61b

Observation a8585808-b885-480e-b8ee-e0a9b1efe396 · outbound

This paper cites InfoBERT: Improving Robustness of Language Models from An Information Theoretic Perspective.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model InfoBERT: Improving Robustness of Language Models from An Information Theoretic Perspective

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.624252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:a38abf377c92cf936d2070f592702db2aeb95eb525d599b72bd2beefb9f104f2

Observation 30bd5f71-2927-41e8-8907-6617bc661170 · outbound

This paper cites Towards Zero-Label Language Learning.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Towards Zero-Label Language Learning

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.610742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:7e7ec84ef4f26e5e258df1f6e88dacd5f9da36230cb144dd17cb1ad54307f916

Observation f9efcea2-65d2-437e-adf1-d3b999de156d · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Finetuned Language Models Are Zero-Shot Learners

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.522602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:69d50fad5b509d0667bfccafe79ffc49bfd8a4e1c2dded8cbbb3d2530b5db610

Observation 55f60e0e-276c-47c1-a39f-c3e23ec02e06 · outbound

This paper cites Ethical and social risks of harm from Language Models.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Ethical and social risks of harm from Language Models

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.682178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:fde2a4d9ab29f0e9cad1263b83b69707afd16713699fdf7621c1a3c327f0adb3

Observation 186a5027-f14d-4885-85f4-8e785660a588 · outbound

This paper cites Challenges in Detoxifying Language Models.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Challenges in Detoxifying Language Models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.647929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:962124f54c7a3fa8d90c3731a60c26aa28ffa0fa1bdd11e1cc7de1649aecc687

Observation 4cac1c9b-8be2-4863-b4da-6c735cd43960 · outbound

This paper cites A broad-coverage challenge corpus for sen- tence understanding through inference.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model A broad-coverage challenge corpus for sen- tence understanding through inference

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:14:28.209694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:fa98932d909bb357f03453deb8894b655995ab1b16c10131f5ef51552e639fae

Observation 724c78f5-694e-4ffe-afd8-5311deda88c0 · outbound

This paper cites Yuan 1.0: Large-Scale Pre-trained Language Model in Zero-Shot and Few-Shot Learning.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Yuan 1.0: Large-Scale Pre-trained Language Model in Zero-Shot and Few-Shot Learning

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.689281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:243355fdeabc39683a4528e076b07ce97120662930fc26ab7eb621ea295098ea

Observation 807f0376-60e0-4910-ba28-ecc2a63210d5 · outbound

This paper cites Learning and Evaluating General Linguistic Intelligence.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Learning and Evaluating General Linguistic Intelligence

Reference 75

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T12:14:26.676273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:1cfea2782978dd04281eea6689ace032edf92719ca5837f43be26ad14e953e12

Observation cdba02d7-1457-4bf5-839e-e6058e4d5a9d · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In ACL.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Hellaswag: Can a machine really finish your sentence? In ACL

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:14:28.204980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:3f316b24d30ae8a81f30873dbb90cdcf18dfcef0805649d21041f8ece4508a37

Observation 41f90178-98a2-4fef-8b57-31015b5cf302 · outbound

This paper cites Defending Against Neural Fake News.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Defending Against Neural Fake News

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.707955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:090ca0d9fdba7432e49a687c69f4a1c7e17d55c6574a2fea39f6dcd4cde09b4e

Observation a29f3d72-3466-4234-b011-c3164ddc53b9 · outbound

This paper cites PanGu-$\alpha$: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model PanGu-$\alpha$: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.701333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:79ec385abde6bef5a328abe67cd30ca507c8bfbbf7f62b24140343933a18911c

Pith citing papers

Observation 50ef9e93-dc16-4959-b069-f5a51622021d · inbound

PaLM: Scaling Language Modeling with Pathways cites this paper.

PaLM: Scaling Language Modeling with Pathways Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:07.412275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T23:45:06.755839Z digest=sha256:7953eb3e3eed5dd708a4a1f2e894a9740c01bfcd6fbb7edb33386d899fcf1480

Observation a486b441-112b-4452-a22c-c8f449738476 · inbound

GPT-NeoX-20B: An Open-Source Autoregressive Language Model cites this paper.

GPT-NeoX-20B: An Open-Source Autoregressive Language Model Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:34:28.453047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-24T12:33:37.701655Z digest=sha256:a0375d00d82d30a5726b85b3bc7f74b6a677a98b8f43cf9cdd04d128eafd0927

Observation 1b0f78c3-7254-406c-901b-7c0785a3cdbd · inbound

MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning cites this paper.

MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:31:08.360935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:31:08.266737Z digest=sha256:d0cf20cb3de622d84949466c5f3ab17539fff3cf365fcec240502f678793711b

Observation c1aa1546-e474-4276-80dd-173584b680f7 · inbound

OPT: Open Pre-trained Transformer Language Models cites this paper.

OPT: Open Pre-trained Transformer Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 280

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:53:17.867906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T20:53:16.720145Z digest=sha256:a48a066425a6a0fba37c32297a100ae25f26df6afd5018bb5b780ab828aa4638

Observation ecb73857-13d3-4ed8-8a0d-c644e03c1c90 · inbound

Large Language Models are Zero-Shot Reasoners cites this paper.

Large Language Models are Zero-Shot Reasoners Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:04:13.821923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T19:04:13.778347Z digest=sha256:9ecab5bd393e1cd736b531d35d001cf8ab2f398dcadb3ae9702c74afdbc31c32

Observation cd4c12af-5bee-44b1-ba33-16490830667f · inbound

Efficient Training of Language Models to Fill in the Middle cites this paper.

Efficient Training of Language Models to Fill in the Middle Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:40:41.818938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T00:40:41.647820Z digest=sha256:7ce3d1224769604da88296d06ad60ea4effa57abb9c39f7629ac3811d11e5379

Observation 8231ca83-a3a7-4cac-92ed-93c06e062dbb · inbound

Atlas: Few-shot Learning with Retrieval Augmented Language Models cites this paper.

Atlas: Few-shot Learning with Retrieval Augmented Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T13:48:43.160081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T13:48:43.024120Z digest=sha256:976e99754a9960255404b9c9e5f952da4ebb6e4989952b50724dd9b181b111d5

Observation 099e2817-021f-4d81-bc17-e29e820820e7 · inbound

Atlas: Few-shot Learning with Retrieval Augmented Language Models cites this paper.

Atlas: Few-shot Learning with Retrieval Augmented Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 249

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:48:43.486276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T13:48:43.024120Z digest=sha256:58252d328ffaa22577d4f20dcbead412fec66c11162ce47ecb5c7594f22e4b98

Observation a29b721d-82b8-4ac6-aa87-a8b8a8b4a1f0 · inbound

FP8 Formats for Deep Learning cites this paper.

FP8 Formats for Deep Learning Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:47:03.818455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T09:47:03.766814Z digest=sha256:10caabd8a129acad8a49d9dafc09fa1c94d3426719eb480f8ec31b6aa110967c

Observation 52037305-cac3-4fb5-b16d-67b0121c8f9f · inbound

Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks cites this paper.

Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:48:28.168142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T16:48:27.918334Z digest=sha256:72be953d1e47e15b8ad80dec160790758195659dfbe881450852f122ca8b36f7

Observation 63ab85ea-6a4e-4d41-a39b-2b6e521f4da2 · inbound

Language Models can Solve Computer Tasks cites this paper.

Language Models can Solve Computer Tasks Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-17T12:17:26.821735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T12:17:26.602361Z digest=sha256:20f66273130b41702e3b159e322df8160e7a838d4f927a92bc1399f6ea601630

Observation 19dcd158-70b9-47a3-9321-656e5ac5fbf9 · inbound

BloombergGPT: A Large Language Model for Finance cites this paper.

BloombergGPT: A Large Language Model for Finance Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:19:46.785310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T23:19:46.231145Z digest=sha256:6de838cbeb69053409ca0996250c3f6b5c78fecd58ff28816d9c5b71cf2a4270

Observation faad2191-00ce-4e8b-8718-803da0a818a0 · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:46:40.117344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:a14495029e805b00713ae796f3cacff776b613320739d801babb1ad5a81786ea

Observation bf6ab664-31cc-48a9-9a6b-e8cee9a02e2b · inbound

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling cites this paper.

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 113

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:45:17.729178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T17:45:17.540282Z digest=sha256:22109e8d4118a185ae0ae233ecd90a7c471be16be827d098ddeddbf3ed5e84e4

Observation 1d2bd2c8-6aed-4f51-814a-19248819316f · inbound

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment cites this paper.

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 112

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:46:56.769745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T00:46:56.664582Z digest=sha256:c452c1b84d9576cafe81ad75e3fdeb08a594e8041fd0dff5ff532e1bdc7c1c14

Observation 09e1623f-97a6-4ca5-99fd-2cdfb366953a · inbound

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models cites this paper.

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:37:01.953355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T20:37:01.617345Z digest=sha256:d91338e80a98b76ad83ba4b9887c6739e174d1fac1466eaff95927c33b3ce4a5

Observation 39312b7b-8b85-403a-82a7-e380fe41f8a8 · inbound

Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes cites this paper.

Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:50:09.469443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T20:50:09.265838Z digest=sha256:612c86db3ccb99e5bd6871b134e8f0e979cc5d0624153a3750792bc70eb25931

Observation 7f1f77ed-ed3c-49c1-b48b-820e49dda987 · inbound

StarCoder: may the source be with you! cites this paper.

StarCoder: may the source be with you! Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:33:00.632672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T23:32:59.517389Z digest=sha256:a927a2afd0fec3bb533c0526fe5e92e8ea1b48668c5fc2324a2e5c3f2c32ef3f

Observation 3ed52d22-268b-4b83-91c8-185a940931b4 · inbound

Scaling Data-Constrained Language Models cites this paper.

Scaling Data-Constrained Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 107

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:35:21.363963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T01:35:21.150772Z digest=sha256:79895678a9a4ed41face358148486acac1a6adb5c362abd88c11a36a7bb274a9

Observation 559e2ffa-4344-4b44-8ed1-33d5930b0bf0 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 117

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:28:39.395686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:b9cae83efa51e9cf48b604388c9ddb9ab0f868a80bc71000b746b7fd3ff84d79

Observation 6174c909-ccf6-4d60-8777-767adc0c54d6 · inbound

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models cites this paper.

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:07:22.310566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T01:07:22.166595Z digest=sha256:69528b711f71abab074b20f4604efa75ca5d52c07277ccbed54eb74c816e9b3a

Observation 9ebc7711-7d7b-4a73-9054-aa5224e22477 · inbound

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning cites this paper.

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:13:08.963276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T07:13:08.867745Z digest=sha256:b8aa72ab626db2ff2051ed0bcd61cf2843cc62e58765cb5bfe45ad181c961fde

Observation 6e875dfa-0357-496f-80de-72367519a446 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:46:10.039408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:10087a63387840b6368d252e50e8373c22f0d1dd3a191ad2536e1f4521621ef2

Observation a7a3a822-d535-4b32-8004-9a7c8f3e918d · inbound

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism cites this paper.

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 164

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:08:06.314335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T06:08:05.550346Z digest=sha256:8ee182aec569ceb05f0940007907323a2e123abead881c599c2ee561b72d9afb

Observation ca5a0c81-b324-4ea3-b347-a8a154f6ebfe · inbound

Large Language Models: A Survey cites this paper.

Large Language Models: A Survey Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:22:55.404021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T15:22:54.023279Z digest=sha256:f58c0035c72b0b31e6990f21f338edcfd49d4a5bfe8405e994ef3a094808578a

Observation 7af723dd-e793-4bde-b21d-2bc5f83e8c60 · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 162

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:36:26.849180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:608700a3c5c35a1877e35e879d7248c08472aa2ee30f0b9753a44104f1eb46ab

Observation 95d34cb4-cc34-4d1c-b590-c1067d1ebeb8 · inbound

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs cites this paper.

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:51:29.509284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T15:51:29.022336Z digest=sha256:71e9f0af07f59be47ed7acd0c82117c0b5469dc6a88abf7b8fc831bf250ed06a

Observation 68750bdd-ab4f-4e91-b674-5265d232da8f · inbound

SEDD: Scalable and Efficient Dataset Deduplication with GPUs cites this paper.

SEDD: Scalable and Efficient Dataset Deduplication with GPUs Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T06:47:39.580914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T06:45:53.685914Z digest=sha256:ad6cb3e553a1cd3437c490d710868b1555b28e90f98ee35978c1d90ddad968c1

Observation 4b54afb2-b5b9-488b-8db0-17fa88158b11 · inbound

MiniMax-01: Scaling Foundation Models with Lightning Attention cites this paper.

MiniMax-01: Scaling Foundation Models with Lightning Attention Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:26:38.632248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T06:26:38.569394Z digest=sha256:98623e4818f58e36ed865d50a67cb98b252bdc2236a456a3d9eb0c6591824f1d

Observation d2031268-7be2-48ea-80f6-85d7535c0870 · inbound

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning cites this paper.

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 247

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T01:01:10.436989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T01:01:09.840919Z digest=sha256:9e713d19447969dc772300a504f66631dd96556ddd50a602a4e19d69b45e8131

Observation 8526b1fe-2dca-4403-8acc-9cf99df5dfa8 · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:13.792040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:13.792040Z digest=sha256:fd1b10a584747044956754710a968f2dee62030ecabd00abd96f8e52e5ab16cf

Observation 7ab3e7fe-6e51-40fb-9964-594166fb7d91 · inbound

The Carbon Cost of Conversation, Sustainability in the Age of Language Models cites this paper.

The Carbon Cost of Conversation, Sustainability in the Age of Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:17.516037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:17.516037Z digest=sha256:092cfbbd0713b72cad3f762ad5630eccbf6edc34642b37cd99702cc29a0e702a

Observation 90d1014f-f333-4cb1-9c37-0284d893ebb7 · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:07.250063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:07.250063Z digest=sha256:a5ad4c5d15c53a1b2b7daf79b3c0f6040993273e59fd7e1ccb7a2142d032f84f

Observation 3356bce0-fe8f-4732-ac13-4893bee81912 · inbound

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection cites this paper.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.646394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:bb261bebc7df59a03c7925fc7891a3d47531851db0d53ac6f882d6cc540f40e9

Observation 48c4c3d1-1bc5-4c7a-9ab4-81383c8a8a4c · inbound

Integrating Large Language Models with Network Optimization for Interactive and Explainable Supply Chain Planning: A Real-World Case Study cites this paper.

Integrating Large Language Models with Network Optimization for Interactive and Explainable Supply Chain Planning: A Real-World Case Study Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:12:37.071923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:12:37.071923Z digest=sha256:90170830dbcf4cf01feea7322e8c8a0d59641377ee5af78d9b242ec8dcc2a19e

Observation 7ebde26c-f639-4595-bec5-41cf8ff961ba · inbound

Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training cites this paper.

Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T11:18:08.081946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:18:08.081946Z digest=sha256:614eb6c64b37dbab8084c0e7d7f387666b11602f3584267f8cefe53d1269ac0e

Observation 1ce99e3c-fdb0-4d9b-a830-b1114f96a3b9 · inbound

MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS cites this paper.

MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:49:31.251317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:49:31.251317Z digest=sha256:8fdeb976d45c8e5a455bec5ec84e6109dbbe7359e6b70a4f665c38931ace041b

Observation f7e4b7d6-fc68-497d-a019-8808945a2ab6 · inbound

Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector cites this paper.

Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:42:47.433805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T17:39:17.456350Z digest=sha256:34e9842ee4cf2991c19895470bd0277b469b9537b3d9df0c3e8fc6a8070cebfe

Observation 92cb1cd5-ba43-4c9d-b888-2ead73b2ced4 · inbound

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency cites this paper.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.565155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.565155Z digest=sha256:ec890b9ea39e33348e83af7e79371005d38aa57c0ce558dc29b01357de098afa

Observation 50005be3-08b2-4d3c-8d04-6b26200316c7 · inbound

veScale-FSDP: Flexible and High-Performance FSDP at Scale cites this paper.

veScale-FSDP: Flexible and High-Performance FSDP at Scale Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:06:30.823510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T19:03:22.142671Z digest=sha256:89e291c28f1ebce3e78577f1e5d85a2e0c370b8993e8ff1b97a951f0ff674a5a

Observation 9e152b6e-7278-46ae-af21-b87a2560c74c · inbound

M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling cites this paper.

M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-15T11:39:59.287357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T11:35:43.088803Z digest=sha256:c7d52755d2f1e1ad4e1988ab6030fc7d612ed9aa39d2ac3060423cbf0021f222

Observation 9d979ed2-cb97-4d67-a004-6fafd0d4b770 · inbound

Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods cites this paper.

Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:38:14.387153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:37:49.558450Z digest=sha256:26b7b89774472c5ca9ee516a098d6f6437344c905fd9a1b498773197b409f68c

Observation 5ddb0c30-bb4e-4b90-924d-f362dcb321eb · inbound

SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention cites this paper.

SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:55:28.011720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:51:03.591878Z digest=sha256:5c8f115a5156d359d38c073c8a9b15338811cc3b539eb0be1b77f4f5b781eebf

Observation 2e41a25a-1204-4377-b17e-0fd4c0a8fa19 · inbound

TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training cites this paper.

TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:06:20.780351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T01:36:41.804171Z digest=sha256:dda56f6445db36be1a7ddaf10579cecc4e596cbc5c8c0c22ff929aee0f3825d2

Observation 0b7d393b-f9b7-4505-92e8-b080fc17d68f · inbound

Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips cites this paper.

Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-09T21:58:47.272285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:58:52.425995Z digest=sha256:c101d52cdfd74995e7d58d206c46abbef0503d02dc70b4322421dab1609b29fd

Observation 8d33e553-0885-4dc1-b5ee-8dcc54f613a9 · inbound

A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models cites this paper.

A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:45:56.344897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T02:21:17.841592Z digest=sha256:32522985d4b2f0be38dbb0b7f31b38daee32d317f6c18c8bbfac52a4d3f0b6fe

Observation 9c53ad3d-384c-41fe-8417-da812e31e5ce · inbound

Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction cites this paper.

Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:24.088540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:14:17.332353Z digest=sha256:8b37ea45f004a16ae5901c872de74579048ed29a530782baba03f5a453125d1d

Observation f24b1559-05c5-411b-98ef-1bc7bf1b1f74 · inbound

Phoenix-VL 1.5 Medium Technical Report cites this paper.

Phoenix-VL 1.5 Medium Technical Report Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:01:24.844809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:41:27.144815Z digest=sha256:70f7f1f4c1b478b67d189b4fe36c582b40b78737f2d8537c5c6e17f50ab170d9

Observation e19f136e-48f6-4eb9-8d69-e1d27dc92593 · inbound

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference cites this paper.

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:13:21.224636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T14:11:58.106397Z digest=sha256:f615000914d1a48c212c8a5dd948b18d8f847df593912d4d54c0495b622633ad

Observation 529ce6db-d9d7-4fbf-a375-380928af73f6 · inbound

Large Language Model Selection with Limited Annotations cites this paper.

Large Language Model Selection with Limited Annotations Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-06-30T12:14:39.001090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T12:13:26.167487Z digest=sha256:fbfb5a57fc1132c95ef7dc2948fde9497c676f80eb7379679f0e801a5fcddf0d

Observation 16bc8f80-e4bd-46ec-bb85-73c17165dfd3 · inbound

Heterogeneous Parallelism for Multimodal Large Language Model Training cites this paper.

Heterogeneous Parallelism for Multimodal Large Language Model Training Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:43:50.407336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T18:42:09.591282Z digest=sha256:a820b9d24817de1845590e81f150182e4049730928765a07837aae117c0b5944

Observation f451607f-3301-470c-b5c4-d2f56466ada0 · inbound

Specific Domain Ontology Construction Using Large Language Models cites this paper.

Specific Domain Ontology Construction Using Large Language Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:58:47.538594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T03:26:15.805999Z digest=sha256:2db9faeda2ab46c7da31f3f3f8daef3f95d05f78851404bd9b309d8922a48040

Observation 41110de5-87f0-4039-b6f1-6564292b6913 · inbound

CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems cites this paper.

CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-10T16:37:23.050768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-10T16:30:09.561233Z digest=sha256:05e3ed4aa26ff74cde7590d12162439aa14174e540530d7cac780fd2225ddfc3

Observation dae08bfb-ab49-4177-b12c-1d65d6e46998 · inbound

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text cites this paper.

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T13:37:00.195526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:37:00.195526Z digest=sha256:77c2be314c467eead54c7976eb425d86320a5a65fc014c7122cfb63a12f30670

Observation 2504b164-29fb-4361-bbf5-ae4a643ac76f · inbound

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization cites this paper.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 248

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:52.175260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:52.175260Z digest=sha256:9934b55d822f0e468e239ded226c91fac941c1766054c7016da5d8785eb3f7ca