Pith. sign in

Paper Citation Record · LEDGER

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

As of 6 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 100 inbound Pith citation observations for arXiv:1909.08053.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1909.08053 v4

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:34:44.807534Z

measured 128 of 128 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 440 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:04:51.559842Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact10
  • verified fuzzy1
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch15

External citation measurements

826
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 0f8bb4d1-bb09-4ad7-b06d-538763b9e70a · outbound

This paper cites Layer Normalization.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Layer Normalization

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T18:34:44.824148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:68d879307e346685c8cf8c030c30957946279ff7fa8fed8a212d2f8dfdffd7f2

Observation f426e9da-c57e-49f4-9d85-5dbfbb5dfa5d · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Training Deep Nets with Sublinear Memory Cost

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:44:07.405726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:5df22141f0c80141aaf4967b9478cd0d0f6b06c303c7010acdc95a8818630691

Observation f594b516-3b61-4480-b3ca-ce64c3b4a0de · outbound

This paper cites Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:34:44.830228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:f966ed35447a01faec65cce89477072ea1eefddb704bbea5d40dd7d340c533c5

Observation 23f673ee-a086-4f61-bc9d-7679243b0416 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:20:31.604247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:85bb508a4c69a862aa08e5b0f291cef09b36ebc1803d1d5b0499433c2f50fee8

Observation 8f99c712-d1bb-448e-a4d6-9732757e5846 · outbound

This paper cites PipeDream: Fast and Efficient Pipeline Parallel DNN Training.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism PipeDream: Fast and Efficient Pipeline Parallel DNN Training

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:34:44.836255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:9dff2c1eb868280eca073068e9e462c9c468d126612b2d895dab49d15de57105

Observation 8c71e07a-a91d-4f96-ba2a-4b8ad9d56872 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Gaussian Error Linear Units (GELUs)

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:34:44.838702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:ee5edbc6a3f4f9c8de6c9920c4172f53793dea65702cbd3ff5357279ede7d9e2

Observation e96f00de-6d16-4b2a-b9a5-91d6e5fcf5ba · outbound

This paper cites GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:34:44.841573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:28e5966d8c46bbe37b8bd747c3db89ab6029776c640ac4184966f1226a307c35

Observation cbe37ed2-4c6e-457f-b06f-8b114db5ebc3 · outbound

This paper cites SpanBERT: Improving Pre-training by Representing and Predicting Spans.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism SpanBERT: Improving Pre-training by Representing and Predicting Spans

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:34:44.844288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:d6a80646f528869a175702d1c06ca92d8558ed9c264a548a8f4c80b6279ad78f

Observation 0e5b8535-7186-4bda-9b80-a426ecdc938a · outbound

This paper cites Generalization through Memorization: Nearest Neighbor Language Models.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Generalization through Memorization: Nearest Neighbor Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:34:44.847208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:56fd6bba9510ae148084bdce97a89699f20e7415b48a084561ecd71e80e9b013

Observation 5e80f894-356e-4ec8-8050-a30bb979415b · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Adam: A Method for Stochastic Optimization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:34:44.849784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:4dc4371494ee44bc4cdf123a9b868fe43971caaf387042778dfff77241e59b89

Observation 791dbb77-68b4-4464-8f52-f02c49a7d646 · outbound

This paper cites RACE: Large-scale ReAding Comprehension Dataset From Examinations.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism RACE: Large-scale ReAding Comprehension Dataset From Examinations

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:34:44.853068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:857339639392188af3deafb747f4aa254720c4d6006f27a18ea18650a36f60c1

Observation e1a008e8-a390-4f8b-a586-78eec5e7c04f · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:26:58.164513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:58f1ead25649f61bb23c611a1c4ef8ddeff6fb7a9cd0e24bf7bba490e2341453

Observation 1755b56c-2157-4548-a092-a5ac8c2c3a2b · outbound

This paper cites Multi-Task Deep Neural Networks for Natural Language Understanding.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Multi-Task Deep Neural Networks for Natural Language Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:34:44.858886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:480a94b3a348f0c9d79791c8adf20668ed1af8157099e82a0b8addbf56f77d33

Observation 28f73f52-0e3e-45bc-a273-fa66dcf955da · outbound

This paper cites Learned in Translation: Contextualized Word Vectors.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Learned in Translation: Contextualized Word Vectors

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:34:44.861716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:b8bf813de02fc86c01cf18377bb75872d7b07f8764f5a116eb10ffd8dddb8f17

Observation b449abbb-d6af-47a5-b95c-319adf15feac · outbound

This paper cites Pointer Sentinel Mixture Models.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Pointer Sentinel Mixture Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:10:32.318844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:c3f77d3171dcf2f7d50afad26dcc14db287a96bdaa2ba8fb6ab4f0401cdd6bdf

Observation 7aee95bf-9046-4216-9b46-28e3acbc8b56 · outbound

This paper cites Distributed Representations of Words and Phrases and their Compositionality.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Distributed Representations of Words and Phrases and their Compositionality

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:34:44.867681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:4e0e622634bf061dd97faa562f3b1bfda268cb59f12ad6c19547d189c52aa27e

Observation fbd2b8c6-fdfb-4b3d-a7a0-61bac372cd73 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:34:44.870493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:331256871a33678f945feea6424eaa604542f044de0a672756c1a3927bf01081

Observation 55968434-565c-492f-84f4-b9bfcd18854d · outbound

This paper cites Deep contextualized word representations.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Deep contextualized word representations

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:34:44.873282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:98b3cb7faa10ab7f9478d014a602709b54b299ce86b65ee70bd31dd03d40c803

Observation 6864d9ea-a007-44d1-afe5-972e74820c55 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:37:56.023799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:dff5c36b4e0e3b28dde95f6abce9a8d4b900c7791b229d0c9db2a6919f5fb012

Observation 885fe66a-f0f7-4535-a22c-4f687f799def · outbound

This paper cites Unsupervised Pretraining for Sequence to Sequence Learning.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Unsupervised Pretraining for Sequence to Sequence Learning

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:34:44.878753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:55d415c2b0d2b11e8e62ed34fa0db529f2fd92dc26ae9c669297db7927e49684

Observation 81e151ba-ea87-47c3-8ea4-fe6194a3d476 · outbound

This paper cites A Simple Method for Commonsense Reasoning.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism A Simple Method for Commonsense Reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:34:44.881528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:77b52fd25c33a3027c831aed785d9937b83844d9098cb70d09993ea1a5cbc33d

Observation 8f290751-f657-4337-8785-ee0cc7b54b86 · outbound

This paper cites Attention Is All You Need.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Attention Is All You Need

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T18:34:44.884749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:f960142dda84b356ee80cd2b8b7232374483675162b70288876cf533d504128b

Observation 89e3d504-5514-4529-a0d2-a9a705de9edd · outbound

This paper cites XLNet: Generalized Autoregressive Pretraining for Language Understanding.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism XLNet: Generalized Autoregressive Pretraining for Language Understanding

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:29:27.629409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:85788052a3e96951b5cb54262150ea82413eaa3302259aecf4fde565886d4a4f

Observation 13f72592-4f20-43a5-ab61-c1a3ff0cccf8 · outbound

This paper cites Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:39:00.104164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:eb809488c0182b76a8ca45d254f8c05980b772f56d6180255b13e83601df9924

Observation ef5858c0-3f2a-4b7d-ae2b-5c2f722ede88 · outbound

This paper cites Defending Against Neural Fake News.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Defending Against Neural Fake News

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:34:44.893015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:711a17d5c23c4de8ca4a2c2dc772cc2156d9b5b4f758e5dca856ed5fc7149fb4

Observation cb91cd72-2056-425d-8517-fc49774eaaa5 · outbound

This paper cites an unresolved cited work.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:34:44.894685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:1bfd47c130ad154c8a3bffe133a07377bda813f0756fea33c8e85b927b4f8bd8

Observation 0d89a31a-bfcd-4ebc-aa1c-1a722f09c7c4 · outbound

This paper cites Megatron-LM: With a broad scope, the conference ad- dresses the challenges and opportunities in machine learning for practitioners and researchers.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Megatron-LM: With a broad scope, the conference ad- dresses the challenges and opportunities in machine learning for practitioners and researchers

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:34:44.896612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:0e70330809bfdb6429a291c04347a995acf717c5a23985e22edac80e7a5c40c9

Observation 079037d4-9cfe-42f9-b596-f0f3cafb1b32 · outbound

This paper cites an unresolved cited work.

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:34:44.898263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:34:44.807534Z digest=sha256:e072478ac64a3e6c3eeab5c4a871e9d55cf860aa10f7d5005edda51959d674d0

Pith citing papers

Observation 389f37ae-a7e7-43ca-8ce5-fc04e5d6d850 · inbound

HuggingFace's Transformers: State-of-the-art Natural Language Processing cites this paper.

HuggingFace's Transformers: State-of-the-art Natural Language Processing Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 181

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:53:59.845595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T14:53:58.963468Z digest=sha256:2babdd48847783c3179ac27911812f2158cf2f923e33690c51ad7538bbe4407e

Observation 75df99bf-b7de-470f-8b0e-5154a04a4374 · inbound

DeBERTa: Decoding-enhanced BERT with Disentangled Attention cites this paper.

DeBERTa: Decoding-enhanced BERT with Disentangled Attention Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-13T04:50:53.664901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T04:50:53.587891Z digest=sha256:b3cd5349789fee3b74158623c93a9d4307b4f055c4e1772df424f5c199468a65

Observation 3cdc61d5-2f4f-4dd7-890d-d19e2ebc8c02 · inbound

GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding cites this paper.

GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-11T02:26:44.764583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:26:44.624137Z digest=sha256:eeae9b4cd7b8110057956cda2d5a2984ce51a39e83e57e4b73626833571d0135

Observation 6059e413-b9db-4ba6-b45c-280509a840d6 · inbound

The Pile: An 800GB Dataset of Diverse Text for Language Modeling cites this paper.

The Pile: An 800GB Dataset of Diverse Text for Language Modeling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-10T21:35:18.965346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T21:35:18.513342Z digest=sha256:eacb71d2d18e7981fe58d123844a0628f37324208a9d1b0a6c94f954edab6189

Observation ff2d19a5-333a-40db-91de-4cb9141bc090 · inbound

Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity cites this paper.

Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:57:11.118507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:57:10.953617Z digest=sha256:0fc0188d3f085458b9fb1901c19b38bf97508f6d8a181bbec8876b6a80e9dda1

Observation 7050cae9-5728-4145-b7c9-0ff1e6b8b9a8 · inbound

GSPMD: General and Scalable Parallelization for ML Computation Graphs cites this paper.

GSPMD: General and Scalable Parallelization for ML Computation Graphs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-18T12:36:36.438390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T12:36:36.359390Z digest=sha256:53d2db61dff5d8c659d4541d18260f2e7f983a337e7689c9298f18cc8fcfe490

Observation 4a9a2c26-21ba-4384-9fbf-d7d9e3617763 · inbound

MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer cites this paper.

MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:46:35.166152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T20:46:35.073600Z digest=sha256:9ec9eb1f5a171d44b57799adc1cfde26f1f0f42998fabd055de5191642182c22

Observation 9fb6bc31-c5dc-4850-8bfa-047be6f058f4 · inbound

DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing cites this paper.

DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:49:19.437534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T12:49:19.418548Z digest=sha256:5e6bf09c31034bb49ad6ab00250b4248b5ae95ab3d82994f397dd3fb338996b2

Observation aa35fa9b-c5ea-475f-84da-285eec412023 · inbound

Improving language models by retrieving from trillions of tokens cites this paper.

Improving language models by retrieving from trillions of tokens Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-17T12:55:41.678018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T12:55:41.551478Z digest=sha256:b100754f51be13aaca0b97f7fa7d944ee2c430d2f2cdae9c1bace8718712af55

Observation 223dae09-93cc-48c5-a3ea-bd3020d8071e · inbound

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model cites this paper.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.694969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:3bceb2b83850097d6b2945de240e709885d6364cd330153b30c664298c5a0da4

Observation ea891404-2b92-42a4-bffa-bb49cf382547 · inbound

ST-MoE: Designing Stable and Transferable Sparse Expert Models cites this paper.

ST-MoE: Designing Stable and Transferable Sparse Expert Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T23:14:25.821498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T23:14:25.431471Z digest=sha256:a57f310e9b4119772f317a0e5096fcf22064950df7be942a4459297b4332ae67

Observation e1b85adc-9af5-4395-97ab-b5e450bb0c5a · inbound

GPT-NeoX-20B: An Open-Source Autoregressive Language Model cites this paper.

GPT-NeoX-20B: An Open-Source Autoregressive Language Model Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:34:28.447930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-24T12:33:37.701655Z digest=sha256:024408ff1dff586f434933054c0a2597e719784f3c968ddea9413b026118e6d8

Observation 2722e53d-f30b-46c4-9266-f3fc840b341f · inbound

OPT: Open Pre-trained Transformer Language Models cites this paper.

OPT: Open Pre-trained Transformer Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 141

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T20:53:17.529283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T20:53:16.720145Z digest=sha256:e8483905d858d8aa52114c45b51e85213623cd8ce0023e7485c24264c8700857

Observation 74e63fb4-b2d3-429b-a758-bc9ebd576f79 · inbound

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness cites this paper.

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-12T16:22:09.001664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T16:22:08.801066Z digest=sha256:a0adae91ac7c8371210d523aacbf8c5bb6132f9f5d2811440b21ec3e784bf4e6

Observation ab57fee7-4f0e-4dce-b269-101f1b48e102 · inbound

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale cites this paper.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T13:35:36.190379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T13:35:35.972596Z digest=sha256:c1bb7475fa89f2279503ea9e6a483c6ac8bc220869fc75239f8ef9bf431ef782

Observation ed668f91-b14f-41ff-af84-2f0bb5c72654 · inbound

PaLI: A Jointly-Scaled Multilingual Language-Image Model cites this paper.

PaLI: A Jointly-Scaled Multilingual Language-Image Model Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:29:06.185047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T09:29:05.956863Z digest=sha256:6049ca5b332b63ca3614090f5c79e9660d914e158ed8e46cded598afd189f57f

Observation 01f93159-9b7e-4a3c-a612-250d99f198b3 · inbound

Language Is Not All You Need: Aligning Perception with Language Models cites this paper.

Language Is Not All You Need: Aligning Perception with Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:32:22.987326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T18:32:22.813668Z digest=sha256:60352ff22130d95a3fb141c68e49038881b670c43e11dc5c81e41bf0127fd070

Observation 94bf987a-8a5d-4b5d-9d0d-fa318d7a604d · inbound

Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training cites this paper.

Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:16:06.447856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T09:15:33.705393Z digest=sha256:9740f2c0506d03bcd3453096a389dc6121d4728d350ab388eaf63f598c1d3e78

Observation d240440c-9387-4056-9696-722de6a5387c · inbound

BloombergGPT: A Large Language Model for Finance cites this paper.

BloombergGPT: A Large Language Model for Finance Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:19:46.757909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T23:19:46.231145Z digest=sha256:3bb6ba0de7eafb5e2e92fb0f849460d373997c2e201fc08b308fa9a3a6b7a6cd

Observation 5ffdfcdb-11fd-4139-bb69-0a575b0cb37c · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:46:40.797972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:82089ac43212c1f4d0c23935e7c6340a6ed02687635101adf92e28c18dbc693f

Observation 1e419a38-bd4e-459d-931c-d5f8d48ada3f · inbound

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling cites this paper.

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 112

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:45:17.723599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T17:45:17.540282Z digest=sha256:77b003dad8f41a8f4728eac4425f0c9880f7f696df3eece78f01ef9f56dbb656

Observation 6c1b631f-a8f0-47c8-9292-23e6d2c9f796 · inbound

StarCoder: may the source be with you! cites this paper.

StarCoder: may the source be with you! Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 162

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:33:00.812389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T23:32:59.517389Z digest=sha256:2c18e4f47963b9e566f8545a6ce0eedc0cdeed81d586438738427cb77b6acb3d

Observation 5d248830-30f7-4ef7-8b8d-dbafa6164844 · inbound

Scaling Data-Constrained Language Models cites this paper.

Scaling Data-Constrained Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 106

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:35:21.358946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T01:35:21.150772Z digest=sha256:838f847384044f1b9e95342c1acecc44538784f22bf5e10f91269773468ff671

Observation 9e2758a1-29b0-475c-a3fa-98b450602b02 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:28:39.288878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:4b5e6918f447ceec57ccc495726fd1d24f94729a0fc8b20b96aba998cc064ab5

Observation 4d5e1dc2-610b-452f-b15e-cdc0e1093114 · inbound

Retentive Network: A Successor to Transformer for Large Language Models cites this paper.

Retentive Network: A Successor to Transformer for Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:29:59.707631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T20:29:59.633357Z digest=sha256:099dcec1f2527666dc9cdef22a41efe53ad4cfd88bf9522aa3b5cba09d3c4430

Observation 370a420a-6b78-4548-8b44-ded2bda78287 · inbound

FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning cites this paper.

FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T02:39:44.903701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:39:44.770344Z digest=sha256:b6bb23217ed75a2e223de3c867dddc0cbc0db17223527cb8e552f551f7a93c2c

Observation a8d9d635-39ef-4241-8838-2d62a9677aad · inbound

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills cites this paper.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:31:47.377211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:f5f9a18febfb89fa24da675c4433bfce36a7dea7307ffffb36a310ec74867697

Observation 3cb0556b-ffb3-46a6-a7b9-f38e2f5426be · inbound

Efficient Memory Management for Large Language Model Serving with PagedAttention cites this paper.

Efficient Memory Management for Large Language Model Serving with PagedAttention Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-12T15:03:07.812810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:d9b2502f30947e69366d6d2656c23acb58416ef5a79d88036728e869e64b2b2a

Observation 251524cd-c8d7-4865-821e-de0d462d0fb2 · inbound

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models cites this paper.

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:07:22.339357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:07:22.166595Z digest=sha256:3d5573cfcb17776536102b032bd9fef3c21ae6154821b29c646467bb417bd15c

Observation c38da3a3-dc33-4426-ad05-d412605ed2ef · inbound

Demystifying CLIP Data cites this paper.

Demystifying CLIP Data Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:20:20.294219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T09:20:20.143143Z digest=sha256:84688ca74067a76c75dd6fe8a2d8346d5342ec57c43d57a48e3d0aece05c5c63

Observation 4e48f560-9d38-4459-859f-884de148e8c2 · inbound

Ring Attention with Blockwise Transformers for Near-Infinite Context cites this paper.

Ring Attention with Blockwise Transformers for Near-Infinite Context Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-12T19:28:28.295957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T19:28:28.201789Z digest=sha256:7e7caf197e40041ab5f5c1e53d34b33cf97e3aa598bf1b42a9f7e0199285c611

Observation a0d90def-7f63-4854-8278-14a52c9f5640 · inbound

Llemma: An Open Language Model For Mathematics cites this paper.

Llemma: An Open Language Model For Mathematics Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 178

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:17:46.236990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:17:46.055279Z digest=sha256:bdd3868f1912f2fab874f27c2ce5b93408b9165b2cf03c04f6ffb266dff3c610

Observation 9a04fcf6-8bc0-4098-ae6a-9d3ed5b5d8f6 · inbound

BitNet: Scaling 1-bit Transformers for Large Language Models cites this paper.

BitNet: Scaling 1-bit Transformers for Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-24T05:43:56.587719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T05:41:15.544164Z digest=sha256:8dc4ea9d8a911c5d5c17cc6299c584312a201250ba239e985e4e82e05273e5a9

Observation 3b0cd216-b39c-457e-90d1-776d007642af · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.744387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:31fc25cb38b5bdc6ef88908518989e77e18132ce69588537eea4afebc2014047

Observation afc40ad2-a557-48b6-9475-9a272e659673 · inbound

DA-Cramming: Enhancing Cost-Effective Language Model Pretraining with Dependency Agreement Integration cites this paper.

DA-Cramming: Enhancing Cost-Effective Language Model Pretraining with Dependency Agreement Integration Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-24T05:33:56.703757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-24T05:31:30.519530Z digest=sha256:d52819bb27071979853412d9e7c6528c75206dfe9c82688a707250b8767831e8

Observation 07887c75-ef06-4fb9-af96-67573407701f · inbound

MEDITRON-70B: Scaling Medical Pretraining for Large Language Models cites this paper.

MEDITRON-70B: Scaling Medical Pretraining for Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:12:08.549057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T14:12:08.478384Z digest=sha256:753345445c43a8af5ffdf9e6666adc438c8933db29cdd21cb6f76fb2d6a68a02

Observation 2e4d5640-482b-42da-a438-e5fa3f194924 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 105

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:46:09.836472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:a78d0dfb82ce9e4334648adea95bfb89043aa90fc2784e95db5be020ff210cbf

Observation b9d0f233-c377-4ed9-8dfa-7df0dfe79328 · inbound

SGLang: Efficient Execution of Structured Language Model Programs cites this paper.

SGLang: Efficient Execution of Structured Language Model Programs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:20:01.100447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:bad9cd40142c73f97839b4f657fa3f480aafa5242aefdcf37050a215f8e7321d

Observation 6a126239-b469-4f5a-b9d3-f1d1d07c6928 · inbound

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism cites this paper.

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 93

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T06:08:06.052350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T06:08:05.550346Z digest=sha256:3c4f14b5278cd5af8360b7ed7fa1a3980cda71a46588b5fd3df0ee02820bda65

Observation cfda3446-ac63-42b3-a9ed-9e8c4ac20d4d · inbound

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models cites this paper.

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 116

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T22:50:14.250875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T22:50:06.399707Z digest=sha256:d1c058c450955f4ed2d476edff466f3a726bc743f2eca4e94e3032ff55bdcb33

Observation 3ff9f499-89a9-46ad-9ed4-a3bdf03cfa41 · inbound

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models cites this paper.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.592017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:d80a0005e24faf1725c7752cdc4178215a32a239edef3607431fa1b7b7b7b573

Observation e2b0f7d1-b5ba-46c8-9814-668b885c69de · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:47:27.906266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:31873bfe57060f9ea11da18ff582da10b919642c350d559409de22a38870e6e3

Observation f6852170-6426-413d-bbd0-33776e220ae5 · inbound

DeepSeek-VL: Towards Real-World Vision-Language Understanding cites this paper.

DeepSeek-VL: Towards Real-World Vision-Language Understanding Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:58:54.781496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T17:58:54.177359Z digest=sha256:295f850f2db6d993b8fd1b486d6be693bd5d66b9679f5f62c6b7bc8f2f752289

Observation 9160f9dd-6a1f-416d-b99a-d93e8edfe5d5 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 102

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.241377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:fa9b2535b685a3db1c4370e054830ca89ee73df4b63c26f2a0995bc60b4cf024

Observation 6373220f-520c-4288-8962-0666c4ec80a0 · inbound

InternLM2 Technical Report cites this paper.

InternLM2 Technical Report Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 123

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T11:44:38.305794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T11:44:38.066501Z digest=sha256:a859f8d90cf7f9a51052a3741bd9b2f537caafcde56eccf191b46c828ca53dfd

Observation cfb546d7-68fa-432f-9030-7f2b14d9efae · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 82

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:36:27.124128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:bb1eed74ab098552bf26288c14831a2b102764cc758e197efa3551df5bf1e001

Observation 8026b487-4382-4610-9809-6d3fa8e08cd3 · inbound

PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference cites this paper.

PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-24T01:18:42.389726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T01:17:11.261301Z digest=sha256:35aec741653393d86aea2bbd61f306595900429a4dba416ee895e125d9fb64d6

Observation b08bb3ce-8834-47b5-b6c5-8c80c58f6bbb · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T18:44:49.742107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:24303d8ec11580f935c0e7d96c3be7ea9fb07689df2a1d690fd83e9ee05a97b2

Observation 6ceaa746-df02-42ae-935b-600396963f68 · inbound

Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality cites this paper.

Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:16:25.952829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T12:16:25.390683Z digest=sha256:f711f996bb96b20ef6d6f5a0410fdd610adec5354fd4930bd142c90fcab1bf03

Observation 335f17ad-e390-4060-8c26-752624ce4bc8 · inbound

Scaling and evaluating sparse autoencoders cites this paper.

Scaling and evaluating sparse autoencoders Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-12T17:47:23.252126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T17:47:23.089288Z digest=sha256:9a1ff508e94296f40d6198d63d236e83d654b83254b36039b8176386f21ced9e

Observation 77e44c69-3974-4f3f-9753-4600f4efe9c1 · inbound

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation cites this paper.

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:09:17.034483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T22:09:16.622717Z digest=sha256:7eb755a521b28925e7dec0a031f8bd62ed563bae37d6f0a60a5ae592e2ae533f

Observation b3451654-3f18-4fd8-bf63-f92420301084 · inbound

An Empirical Study of Mamba-based Language Models cites this paper.

An Empirical Study of Mamba-based Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:31:03.924982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T10:31:03.777169Z digest=sha256:46c3b1675141be84c3dbde14b161650a10712582905e545bedc8619d875b254f

Observation 6196ac6d-2aca-4646-8cbb-3d072772567b · inbound

ProTrain: Efficient LLM Training via Memory-Aware Techniques cites this paper.

ProTrain: Efficient LLM Training via Memory-Aware Techniques Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-23T23:53:39.447712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:48:40.669335Z digest=sha256:fbad91825e22534ccf579883eb2775414979c8e5358937fbb796400585d08881

Observation cc674576-a2db-44d4-83a3-0a3e329d3f22 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 118

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:10:20.996073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:7196b3f439946bdf434fddb51e18197d58ec3630473f82622ec0ab743f4a4c07

Observation 755bcfab-8b96-441c-98ba-e66b641fc83a · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 109

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.485479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:2d1b113fe456bbfc86d7c9a3a25c6b1f66b4d97ad3aefbd965a37e8935ba4ff4

Observation 871aeb76-d8d1-44a5-bdfc-d6499bad7856 · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.546496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:ef948c8d6051c80799663f7ab45a793a1e907b00d06d2cb38f3aec11069fd85e

Observation bb16271e-5596-45ec-a8ee-a277a1b0fdef · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:53:38.884441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:274a89bf390695506d960b0308f0b5ec75219edea1855a92e2981c209b3a89f1

Observation e2a7c3e8-d3c0-4ab9-8695-9d710bf77841 · inbound

Movie Gen: A Cast of Media Foundation Models cites this paper.

Movie Gen: A Cast of Media Foundation Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:25.965837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:76558cab39d77197dcdef442eb5b5522ad79e491eba24f281d247a701b518846

Observation fb5c6eb2-2b62-45cf-ac0f-ad96dadad573 · inbound

On the Convergence Theory of Pipeline Gradient-based Analog In-memory Training cites this paper.

On the Convergence Theory of Pipeline Gradient-based Analog In-memory Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:18:20.779225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T19:15:54.807005Z digest=sha256:9d32ce836b8827b668dcf1d237319f2d2dd1b4859229be21bb9c5a64826ac18f

Observation 89cc7654-e2eb-4627-99f1-1b3a57559c73 · inbound

Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading cites this paper.

Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:43:23.637603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T19:40:53.056745Z digest=sha256:ba2a215019386af5f6ae76fd6aac9109203003f9b72d07cb7895db2c62e1d1ab

Observation 75907270-ba0a-48fd-912c-f78279ef2539 · inbound

Open-Sora Plan: Open-Source Large Video Generation Model cites this paper.

Open-Sora Plan: Open-Source Large Video Generation Model Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:42:45.182411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T08:38:27.946746Z digest=sha256:d6ab23c3a19168604ed0d9401d6f7b73f1a8fc26f636f7ad4be29f5b1fb887a1

Observation a5652ae3-f68e-4c14-84de-88f417e5fa37 · inbound

HunyuanVideo: A Systematic Framework For Large Video Generative Models cites this paper.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.467254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:9f34aa4dcc8cfa544f043d11826a2a1d39927ffc5988887766097ee38a9464e9

Observation 946b9d6b-8289-4cb5-97e6-8bbc958d2725 · inbound

TrainMover: An Interruption-Resilient Runtime for ML Training cites this paper.

TrainMover: An Interruption-Resilient Runtime for ML Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:15:28.689460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:13:16.130354Z digest=sha256:b8f0a8b02190301a976021fbd5f870d8b691efd5cf441bdeb31ed6e84894c7f1

Observation 59c4eda9-a28f-4e02-866b-173f20e5c922 · inbound

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference cites this paper.

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 185

Resolution
verified exact
local_arxiv, observed 2026-05-20T17:46:47.053057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T17:46:46.845424Z digest=sha256:ba8374da8a9a3a68ca50d42205d2f8b40fef5c354e7ce64ba736993937104bda

Observation 3b6d43cd-efef-4b6d-b1de-ca28eab11800 · inbound

Cosmos World Foundation Model Platform for Physical AI cites this paper.

Cosmos World Foundation Model Platform for Physical AI Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 178

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:38:45.900922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T23:38:44.933410Z digest=sha256:2d6aca520b01008bb147ceab5a264d79d14fab4e03fee73b9f3b9fe7766cd66d

Observation 8d15554c-d71a-4ef7-8c1c-fdf1cccd2c44 · inbound

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws cites this paper.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.022377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:93b751442ce19cd7eea093276d68bb40a8ec6ac3901c0d9145fdedb4662841c3

Observation 002410d4-50e7-4c87-953f-205dc12df36f · inbound

MoBA: Mixture of Block Attention for Long-Context LLMs cites this paper.

MoBA: Mixture of Block Attention for Long-Context LLMs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.206948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:ac9449e658452de0fdf41f4378a1c8a07a77b39bb617585a798827228f67fc4c

Observation 00bd2f63-38f4-42ae-bee6-f32946035640 · inbound

Green Prompting: Characterizing Prompt-driven Energy Costs of LLM Inference cites this paper.

Green Prompting: Characterizing Prompt-driven Energy Costs of LLM Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:07:17.295068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T00:05:26.205947Z digest=sha256:028d0350956753ab21dea4a3152ffd8c96bc873924aa7bcd2302b1b0b4eab4aa

Observation d04ed5ec-f072-463e-b8b1-04f823f5bc93 · inbound

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning cites this paper.

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:47:10.263743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T12:47:10.146795Z digest=sha256:dc11d3c1f44d1f12e5acb68ba62dfbcd6380b04adf33bad8c23afa078d8b0550

Observation 697b8262-a874-4ed8-b1d6-093e5e211755 · inbound

Advancing Intelligent Sequence Modeling: Evolution, Trade-offs, and Applications of State- Space Architectures from S4 to Mamba cites this paper.

Advancing Intelligent Sequence Modeling: Evolution, Trade-offs, and Applications of State- Space Architectures from S4 to Mamba Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:12:11.781629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T22:07:12.198095Z digest=sha256:5f5db0fb26bcb9a3599d8ff9ea53ae9bdbcd56e63931924aa20362c508a7e122

Observation 0e5c7601-8bbc-46ce-ad11-5ff741ea30e1 · inbound

Wan: Open and Advanced Large-Scale Video Generative Models cites this paper.

Wan: Open and Advanced Large-Scale Video Generative Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:07:14.464274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T23:05:32.595632Z digest=sha256:cde1656696cfd8cb9910883967c67aea13e9467a380c7507ac167201ea061273

Observation 8ccd6446-f1b5-496e-884f-fe66475fdb6a · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 124

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.478506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:bb552fc12df0f48cb34fc9d83cc2ff8db8ec70ba7fb0eceed886713ff1262c21

Observation a26afc32-6622-45ba-bdde-a04dee61efd2 · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:44:58.099014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:2aa14838369994f0b748da708650f62a5419e15704db0579acd43f36e402299b

Observation 9b490396-40b7-48ba-8703-c5eb8ac69ba5 · inbound

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning cites this paper.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.046981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:c908735241ec2191f01985ead53e8dbbf23768062a804ebf151a22c59b3bbf21

Observation 45c9ac0f-4a4b-45db-a568-ef3cddad83c8 · inbound

Geminet: Learning the Duality-based Iterative Process for Lightweight Traffic Engineering in Changing Topologies cites this paper.

Geminet: Learning the Duality-based Iterative Process for Lightweight Traffic Engineering in Changing Topologies Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:09.184836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:52:01.501848Z digest=sha256:e2f21379751050f220ada2bfca600e46fb8797101690f771a9b74cfe0d244e50

Observation a5e77505-c9de-4905-866d-c18ae043affe · inbound

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning cites this paper.

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 169

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T01:01:10.306238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T01:01:09.840919Z digest=sha256:24783229149e1b1975c96798067510a8bfbd98df562c361dba2a12a916b2f808

Observation 19100ab3-2068-482c-8842-02ec7c733e0c · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:08:35.324172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:b772b088af110837f9992df747353f522cc6d3545509cea3fc1c327d04846c80

Observation 43daee91-c8b3-49af-be81-80a9a82f575a · inbound

WebSailor: Navigating Super-human Reasoning for Web Agent cites this paper.

WebSailor: Navigating Super-human Reasoning for Web Agent Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:37:09.701354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T15:37:09.572241Z digest=sha256:ae98849d04d5bda7dcf888f9fad2ee685d54bcd5210f8b14514d3a26e5f5f1e5

Observation 3b0ccf34-b347-441f-8bc9-9dbc276bfc8e · inbound

The Serial Scaling Hypothesis cites this paper.

The Serial Scaling Hypothesis Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-05-19T04:12:02.387783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T04:08:11.344622Z digest=sha256:0169f2ad7ac2f7b48ae326ff221e0a3e2fbd6d023076e5cc757dbeb7e623b0c2

Observation 750ca42e-09b8-4e63-b44a-5a3822f002b0 · inbound

Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training cites this paper.

Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T03:37:00.897231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:36:50.366757Z digest=sha256:b69c29c3f2108a6f9d9202cf4de200907b0c4bf60695a58035692709ed9f1444

Observation 84dbb7d1-f4ee-41aa-b0a5-294368c681bb · inbound

Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving cites this paper.

Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:11:43.920084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T15:08:15.328986Z digest=sha256:cd43312c263c4e2deb53b95a5511efe8f8f18bdee051cfdd2525d1b2ffa8c3a5

Observation a6dc1feb-bd98-4ba2-8553-c4c2cdd45b4c · inbound

Qwen-Image Technical Report cites this paper.

Qwen-Image Technical Report Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:34:44.898910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:29:06.883874Z digest=sha256:8a42381bcaad0f560eff3b73c603a183bca990bbb9570ae0d1ad0a06fa72da50

Observation 9ba23427-8128-4646-99b4-5e853ce4bb3b · inbound

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap cites this paper.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.559842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.559842Z digest=sha256:cb6f6ceead192abb4c308b5e7d88f457928e5080e7828c2dbc56005f6f57ae04

Observation 78357e2c-19ba-4088-9ab8-930972083114 · inbound

Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking cites this paper.

Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:39:13.733149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:39:13.733149Z digest=sha256:c5fc7a78e1dcf62b08b4e07d3781ae66a26c537531257165aa1572a4eb1f289e

Observation 28f25519-5370-4395-94e5-818027fc61df · inbound

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement cites this paper.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:52:51.942278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:08876573e4d6d8176316951ccfa95a16a73d0816cf0e45c38f3b90a91785443e

Observation 643b31f4-664d-413d-b0c8-932a5bc5b3e5 · inbound

Power Stabilization for AI Training Datacenters cites this paper.

Power Stabilization for AI Training Datacenters Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T18:42:45.378838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:42:45.378838Z digest=sha256:b2055180032b9ce172e62680de4beb0abeac6b986a576f8e41a4ad760ff4f12a

Observation 626d1528-c68b-4089-badd-9827848560c0 · inbound

LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind cites this paper.

LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:20:42.111816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T22:16:07.638869Z digest=sha256:b7254976f4ea37708bbc5dc16f2b932a89539a80de52838459122d13e9ffdd3a

Observation 69d8ba3e-4312-4cc6-a132-ffc9930d9915 · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:10.957633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:10.957633Z digest=sha256:db1930df81591c3dd3dd127354cf064eb613864a6f5c179c0c40fd8fea57b7af

Observation d1881a62-7eeb-4f60-8776-01134f216dbb · inbound

HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling cites this paper.

HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-18T21:41:51.675079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:39:02.560962Z digest=sha256:da1b6251b361ac0d88d06068da0f89191a81d8c867100182b7241b37281f922e

Observation ca940f0a-03b5-4b78-ad19-859c10ef6793 · inbound

ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive cites this paper.

ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T16:18:06.947838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:18:06.947838Z digest=sha256:2cce3b5a8e34f0887555fdd227c857af587c619524667073c2cfaf2f9a6f30b8

Observation 5678c3e7-8d6f-494f-b41c-ad52a791efca · inbound

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models cites this paper.

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:25.667622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:23:25.667622Z digest=sha256:9f77cc5eb06c593e0d052dd5a766c67ef976453be217d571503971aa7f3572aa

Observation e3fbcb5b-5d85-41fa-b4a8-1d941a91c40c · inbound

Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling cites this paper.

Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:36.817108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:36.817108Z digest=sha256:9cb0efb1378affb2259700701aff4936d9f931d9ff247021224eb11bc1f678a3

Observation 55a7683b-3298-4207-b693-dc4fc84b9e36 · inbound

LobRA: Multi-tenant Fine-tuning over Heterogeneous Data cites this paper.

LobRA: Multi-tenant Fine-tuning over Heterogeneous Data Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T12:56:43.214169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:56:43.214169Z digest=sha256:597c9a26797f8c54f560a528fbb4b58917eb7a48fefb2a28d887d71751ddf8da

Observation 705d33f3-c8df-4c0f-81e4-0d44d97bb250 · inbound

MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall cites this paper.

MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T11:38:40.899514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:38:40.899514Z digest=sha256:02c42df3eaaaaad7752d732f1fbf41079876d45b69f7cc5ce02757e124123faa

Observation 8b9b11c8-8cbe-43be-bb07-3eb2c68644b8 · inbound

Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training cites this paper.

Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T11:18:07.601248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:18:07.601248Z digest=sha256:8412467d2a7831bbf7cbb6206b25c78cee1640e70ea9d180a2f9253034d7e552

Observation 2eaecbbb-1ed9-43e1-af8f-d6be9c80a590 · inbound

SpikingBrain: Spiking Brain-inspired Large Models cites this paper.

SpikingBrain: Spiking Brain-inspired Large Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:51:45.597948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T18:51:06.243305Z digest=sha256:f4171e86f3544ccd23bd240e4617ff5efb0a4dc78970daf90cc8042351f62daf

Observation 9f209819-099f-4cb4-9fa6-f9f7e7de0fd1 · inbound

veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD cites this paper.

veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T05:29:35.924742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:29:35.924742Z digest=sha256:bb37d5f156a68aee354107a88ec12301f0f9c722ed7822a5e27b2544c65cac1f

Observation 03e444b3-3b0c-4c29-94c4-be4a7f71b65e · inbound

Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector cites this paper.

Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:42:47.450549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T17:39:17.456350Z digest=sha256:06b2293fc4cecd8bf748608a87671e4b2cd61ea960764b953d3d35b91565f0f1

Observation ff59f64d-5fb3-4901-82dd-7dd9e8368205 · inbound

Safe and Certifiable AI Systems: Concepts, Challenges, and Lessons Learned cites this paper.

Safe and Certifiable AI Systems: Concepts, Challenges, and Lessons Learned Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:28.570950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:28.570950Z digest=sha256:f6a818c9f545f988359e7a7eb43abe12a9de675f90bf929b72e816e90c18a362

Observation 8c005d75-46d8-4b02-8e66-8267716ce092 · inbound

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing cites this paper.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:54.346068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:54.346068Z digest=sha256:39495b031b5f8da8ad30971952be7ffc81604ad810e75b01152dc830d88bee35