Pith. sign in

Paper Citation Record · LEDGER

Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2104.04473.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.04473 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:25:09.562026Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a4d62a99-e85b-4e0d-86d5-17278d726ffe · inbound

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model cites this paper.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.541851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:a8fd1a993893a103757af5002cebe9ab5d6db0653f30505ffcc22082608deaaa

Observation 50c00f4f-b34c-4082-acab-2ff8b53dd1d2 · inbound

OPT: Open Pre-trained Transformer Language Models cites this paper.

OPT: Open Pre-trained Transformer Language Models Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 142

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T20:53:17.533584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T20:53:16.720145Z digest=sha256:b99108d15e9a94960bc975b8a82b8889dea0433b252b6ad5dd72f3624adc9719

Observation 98915eea-eaf6-4490-9f89-fa2b7b9db928 · inbound

SpikingBrain: Spiking Brain-inspired Large Models cites this paper.

SpikingBrain: Spiking Brain-inspired Large Models Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:51:45.811840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T18:51:06.243305Z digest=sha256:a7baed81a532e82a9fdf429b92e0b02782a1a8a6fed00c04db5a78e08e62298a

Observation fdb78c4e-ce1a-4b43-aee2-89eaf78d9c06 · inbound

Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector cites this paper.

Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:42:47.536706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T17:39:17.456350Z digest=sha256:5bc961c30f827e6725c30fcbd686fcc922393d38e24ddb0ce165152d13d1b024

Observation 5016bb78-fabd-40f9-90ce-84c305794c98 · inbound

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining cites this paper.

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T04:49:03.022290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T04:46:33.641714Z digest=sha256:9ecef9077b753e3a82a86618a3b57dfcffb0df5166a46a729ed5f2a5e045163a

Observation 24b707cb-bacc-4ff0-948d-eaa90040a26f · inbound

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning cites this paper.

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:43:37.893546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T22:43:01.937642Z digest=sha256:843f788260a83f0e80119960ba07199e403cb743db86ddc4060e9407346436d3

Observation 19c9e4ea-19e8-4fd7-82ab-82d278e94d87 · inbound

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency cites this paper.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.562026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.562026Z digest=sha256:9b00c501923264cb586fe624becb9099a5ccc4e0d0b1d58f7be09e7f2c1a9401

Observation 328733c0-12be-44c1-8352-58b97e7d75af · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:09:05.345041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:2cdfb9b6fc132e8c0d9bf8a85685fa82e0e9527ba116a03095bef7c159140e68

Observation 177b40ef-d995-4946-a80f-645b9db7e617 · inbound

Attention Residuals cites this paper.

Attention Residuals Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:39:04.522153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T06:39:04.312270Z digest=sha256:0037bed28a17191bae03f96d1b882b4dac9dfab036a53839228fc28a10c75b55

Observation c9817349-2284-48aa-8094-16363d79dabe · inbound

AEGIS: Scaling Long-Sequence Homomorphic Encrypted Transformer Inference via Hybrid Parallelism on Multi-GPU Systems cites this paper.

AEGIS: Scaling Long-Sequence Homomorphic Encrypted Transformer Inference via Hybrid Parallelism on Multi-GPU Systems Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:23:09.283474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T19:21:26.158870Z digest=sha256:7c7aabbb67d7dc9e5b0f0e9e5ea2b5f8305b91f8787d21dedcceb69decfbc0f7

Observation 6a0912e7-1778-41aa-9a95-863b705e390c · inbound

An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience cites this paper.

An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:20:30.878284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:16:02.816822Z digest=sha256:ecc49fbb1278027cd57d00be937f39f6ee8d6d686387d1119bd10f1f29c26012

Observation b02ba0ef-36db-49ff-9907-d9d5d8aa0421 · inbound

Nautilus: An Auto-Scheduling Tensor Compiler for Efficient Tiled GPU Kernels cites this paper.

Nautilus: An Auto-Scheduling Tensor Compiler for Efficient Tiled GPU Kernels Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:21.492412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T10:21:05.219519Z digest=sha256:fd37400f2aef8f69b8a34e0ef580de2f7147aa1899763fb239d2da148e7aef22

Observation b98fd996-021b-4bf4-bd1b-e32939204f3b · inbound

GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion cites this paper.

GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:51:04.200107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T04:30:44.728872Z digest=sha256:b82c88dbea790b01c2d5958769451aaac390d76f06d851e29d9298a4ba869052

Observation fbf8fc09-41b9-4bdb-b3b7-64885625d174 · inbound

Efficient Training on Multiple Consumer GPUs with RoundPipe cites this paper.

Efficient Training on Multiple Consumer GPUs with RoundPipe Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:26.784462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T10:37:22.251566Z digest=sha256:ae0a76e4c02cacf629a329474288d019ce71f781e2315a06ded6b7cd3bc9c24a

Observation 4459bbc7-997e-4ae8-bc91-8ddb8ad82895 · inbound

A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models cites this paper.

A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:45:56.331268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:21:17.841592Z digest=sha256:d6ecfd2711c8b24facabc0e756b99ddc942a06c4c236c6138a70f71e0cb617c4

Observation d8715644-38f9-4b6f-9f37-30e957cda438 · inbound

Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers cites this paper.

Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:45.669017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:01:32.057022Z digest=sha256:61c05e988f941c5ebdb5445b199607bdbe7b6454a16639af7691a73a50aaaf5a

Observation 5e29785d-d1b4-451b-88e8-0ea4b476cde0 · inbound

MinT: Managed Infrastructure for Training and Serving Millions of LLMs cites this paper.

MinT: Managed Infrastructure for Training and Serving Millions of LLMs Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:27:51.896687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:25:12.407148Z digest=sha256:36d9f1c947b4900e77a02551e427825f8d619b2dcb9dfcc3777cb3f946228ce2

Observation 66d0774e-14e5-4333-ad20-0313dfe62708 · inbound

MinT: Managed Infrastructure for Training and Serving Millions of LLMs cites this paper.

MinT: Managed Infrastructure for Training and Serving Millions of LLMs Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:05:06.246888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:47:00.295144Z digest=sha256:1c20506ff939a6147a6b1a57d63f0220eebbc567815a644ad41aaa34a6ebde7d

Observation 98a325aa-8eea-4233-abb5-1b7f161c9428 · inbound

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference cites this paper.

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:13:21.202970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T14:11:58.106397Z digest=sha256:39e619447746d9529f2a83febebe6848461d44c663a0f54fca2f9dac79a6bc61

Observation 74ea8c6e-298e-4381-a6e7-301a64b72879 · inbound

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference cites this paper.

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:59:55.795270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T08:55:31.298030Z digest=sha256:b8c26a327ae2836bf2efe0066e5cf0e40f4055be442ddf96a03efea3ae2888a5

Observation 79497a54-12f9-423d-aafb-370042858c82 · inbound

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability cites this paper.

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:38:09.493332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T07:35:32.225708Z digest=sha256:591b4fd36d4cdb5985480d7ae9d49f1cba5254f991e2b6761dac7d20d68f2ed7

Observation 12c7c42b-e7ee-4c9e-a26f-e5612c9b589e · inbound

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks cites this paper.

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:44:48.308088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T15:42:21.405913Z digest=sha256:a03f88276ca4c190346879b86b354f6c1edb41962f3c2050220de617f3536515

Observation 18dd1c7d-1d54-4fb3-b41d-d0ad1f22525a · inbound

Heterogeneous Parallelism for Multimodal Large Language Model Training cites this paper.

Heterogeneous Parallelism for Multimodal Large Language Model Training Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:43:50.412459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:42:09.591282Z digest=sha256:01b77dfa7c62ef0bf06f6c4ca3baf093998070e1d259a39e8edae37c4b91d025

Observation 702f9e44-561d-495f-a5ef-ae30dc9bc281 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.872984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:121aeef42d8e3f994bf23635d975ad4e5f069b9dcdeb7f501c20eaa0c3513a5a

Observation 5c6e3757-cf93-446d-827b-e7eeb1a5facb · inbound

Model Multiplicity for Adversarial Detection in Small Language Model Training on Edge Devices cites this paper.

Model Multiplicity for Adversarial Detection in Small Language Model Training on Edge Devices Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:37:19.322887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T21:26:08.955055Z digest=sha256:fac4ed6a431d84e90d9968c4b16ceea84a462e38058e6eb4276a9a193fd6d69f

Observation 8ac83b28-42d8-4410-8f26-1dd54aa5ebbe · inbound

Piper: A Programmable Distributed Training System cites this paper.

Piper: A Programmable Distributed Training System Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:57:44.655438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T11:34:02.562929Z digest=sha256:0afa3ff06448ef362034da254f9174f78b0a977ea57154b014bca6eb978acea2

Observation 132e3e92-7590-4d9e-87c2-8672827fad05 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 225

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.461495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:515cfc8c2625e31f88db270d5b2d2c53b2393f12789215055ff434842344bb17

Observation db29d1e9-7eb4-4784-8506-b0ee07621a42 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 211

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:18.508664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:18.508664Z digest=sha256:4accbbaab603c60652a2aeac03c66b717217e9d9ec81ea3f702149dffc1454ea

Observation 78b96fde-1b39-415a-b5e3-11a9ee3918da · inbound

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint cites this paper.

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:28:43.916831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T17:26:07.870260Z digest=sha256:32d184d7c568e061309c67915f9c04cfa31b5326d53d37bf246765456f6b3e4a

Observation e857b2bb-5386-442e-8cc5-6b2d1d92bfee · inbound

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint cites this paper.

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T08:40:29.554554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:40:29.554554Z digest=sha256:bf2926f74d2ee05ab1dbf355e1439bd431ac03eb6fcdf2370b52f722f2e641b3

Observation e919841b-30d3-4e66-a2d1-e594d06b6753 · inbound

Design-CP: Context Parallelism for Design of Protein Nanoparticles cites this paper.

Design-CP: Context Parallelism for Design of Protein Nanoparticles Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-12T02:29:52.764344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:29:52.764344Z digest=sha256:17f7cd275be52b8359a8a4c8eb70f86ae47106900e1a6fd74261a019892be4fe

Observation 7ed52c58-f009-466d-870c-16cd34446578 · inbound

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining cites this paper.

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T09:16:06.525993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T09:15:12.214083Z digest=sha256:f95273fa77dea3d09bd61bb332733a5ab244e39874cff4780f2ee3eb4b870f95

Observation 00d01956-3547-4166-88a1-c427c5c4440e · inbound

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget cites this paper.

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T00:41:45.468565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:41:45.468565Z digest=sha256:21348d3515de4b73720011f531901c5fbb70938ac59815bd7b9bc25408a8de99

Observation 73494bae-8a1d-4c9f-a9ac-3a3c3d90d1a0 · inbound

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix cites this paper.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.511773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.511773Z digest=sha256:5f84eae1cde5d07bb8e661715712d4c16f812d3dbc3ded7c609d2a8e419e1b4d