Pith. sign in

Paper Citation Record · LEDGER

MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2402.15627.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.15627 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T04:42:46.167810Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

24
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c9623257-d2ec-477c-912e-e570e9abb52d · inbound

PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference cites this paper.

PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:18:42.456861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T01:17:11.261301Z digest=sha256:64f177acaff1599fb816dbef46b8d3b646e934ed01573cca54375dd0a2d6424e

Observation 1e7d53a5-b4b8-4a87-ba54-3f17715a0d3d · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:53:38.974569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:ca57307c82f750e6ca78fec019653d8e5f5debbb22fcd3964ff14318045032c7

Observation a0ac690c-53b7-4fe9-a033-64c61ee2c191 · inbound

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection cites this paper.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.682623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:1be2410cb944fa4d84364302629304892e39d476ecb66611d97100093d57e5ac

Observation e6b0403f-0105-4ae9-9b66-2307729c82ea · inbound

TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training cites this paper.

TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:06:20.803716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T01:36:41.804171Z digest=sha256:f28ba5513cef82eaacd3c3df302f75df46d2170f54f940e5f926fc209fa3c0df

Observation 1419dc6c-744b-40fa-b6f4-b5082c8d4f49 · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:06:14.978055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:fbf66916df33100e434b4d2d4ec82062d8698536bcf90deefa42d3fe65dfcdf9

Observation 5d072b5f-85c8-4d9e-b9b7-2b701f313c91 · inbound

Instant GPU Efficiency Visibility at Fleet Scale cites this paper.

Instant GPU Efficiency Visibility at Fleet Scale MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:43:54.770415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T02:43:09.077121Z digest=sha256:2fd64ae35986f967e603e07eb44ab372e28216393153327f4d24f0affd03a9d2

Observation 5bbf6308-359d-4b6c-aba3-505a1ec8a955 · inbound

The Cost and Network Limits of Space-Based AI Compute cites this paper.

The Cost and Network Limits of Space-Based AI Compute MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T04:42:46.167810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:42:46.167810Z digest=sha256:56a894e752a93a23ba27592314d90c4883b0513d42cc1fb2dc13d74823888646