Pith. sign in

Paper Citation Record · LEDGER

Scaling Deep Learning Training with MPMD Pipeline Parallelism

As of 20 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2412.14374.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14374 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:21:38.881968Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cca14982-b4a1-4f08-983e-2a7c1e0bf6ca · outbound

This paper cites URL https://www.top500.org/system/180239.

Scaling Deep Learning Training with MPMD Pipeline Parallelism URL https://www.top500.org/system/180239

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.343082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.800001Z digest=sha256:e091e3936f791dc5ca3302bb8743c8be8fe6f420d4e6d0e6b55580b5c75d102d

Observation 2b7404cb-0d72-4344-9513-d0b078a2de96 · outbound

This paper cites PartIR: Composing SPMD Partitioning Strategies for Machine Learning.

Scaling Deep Learning Training with MPMD Pipeline Parallelism PartIR: Composing SPMD Partitioning Strategies for Machine Learning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T12:21:39.241140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.803149Z digest=sha256:0bbcec2f006c160b7bea03fc24f151ad8a78d531d4203f85c51aa01f31585bc7

Observation 583b702d-4480-4e21-b82d-241700fc0f2d · outbound

This paper cites E., Thekkath, C.

Scaling Deep Learning Training with MPMD Pipeline Parallelism E., Thekkath, C

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.335359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.806478Z digest=sha256:23ba04f8318508f1ee8a954237b08f4bd16f27fa0b9ee66cb573761cf3b4a5f6

Observation 169ba462-9729-485e-9869-e24430089a78 · outbound

This paper cites J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne , S., and Zhang, Q.

Scaling Deep Learning Training with MPMD Pipeline Parallelism J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne , S., and Zhang, Q

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.327203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.809473Z digest=sha256:7a1672d28a9af5c99b7d20114c7e91bd0d7be3351aea77b04bedc5812ade9369

Observation 6bd4f832-9cf5-4502-b5ba-978ad4580e4f · outbound

This paper cites an unresolved cited work.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:21:39.318466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.811689Z digest=sha256:0cb357abc9170414493677e8bbdb073e922811585fc34c07161ab658924da606

Observation 50454ee1-f96d-4d8b-8d0a-553ed39ff100 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Training Deep Nets with Sublinear Memory Cost

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.814422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.814422Z digest=sha256:85e4f4d046494da88b6962db47738d08dd7dc071c47af84d5be0a143630e905a

Observation 521085b4-422a-4ef0-b961-4f06d2a84a86 · outbound

This paper cites cuDNN: Efficient Primitives for Deep Learning.

Scaling Deep Learning Training with MPMD Pipeline Parallelism cuDNN: Efficient Primitives for Deep Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.817499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.817499Z digest=sha256:e750b7202448ed5aa9f64928221b9c67182b089e56ed71dac2513dc74792eea3

Observation cfe67fb1-a97d-430e-a0da-89b941f13668 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Scaling Deep Learning Training with MPMD Pipeline Parallelism PaLM: Scaling Language Modeling with Pathways

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.820248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.820248Z digest=sha256:58b5ea4bf8c4a91445068f9fec69ed9b567e8bea62d099ad9db54617e5971815

Observation f474a3f4-9cc6-4e65-9e4b-e784cdc0c0dd · outbound

This paper cites An Image is Worth 16x16 Words : Transformers for Image Recognition at Scale.

Scaling Deep Learning Training with MPMD Pipeline Parallelism An Image is Worth 16x16 Words : Transformers for Image Recognition at Scale

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.311729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.823696Z digest=sha256:5ca5608bda7bf8b605777c383a6d68a242a786d15ba94034436fbfd846c86ca8

Observation eee7ab37-6dc2-4689-a7e0-0f1f5b09b8e3 · outbound

This paper cites The Llama 3 Herd of Models.

Scaling Deep Learning Training with MPMD Pipeline Parallelism The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.826145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.826145Z digest=sha256:90cb6e7ae6f55f7ad7e5db82c61c8294ff833722f3ebec7202b8dbc4c58b9efb

Observation 6e0a0107-32a3-4dbc-9bdf-94719833b9ae · outbound

This paper cites Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.829004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.829004Z digest=sha256:b99a78ad02e8daffca9e2912a27983e00d95ea413501fb68d543ddd9c7e9ca05

Observation 2b855f9a-56f2-4370-b895-1f31fc336153 · outbound

This paper cites NeMo: a toolkit for Conversational AI and Large Language Models.

Scaling Deep Learning Training with MPMD Pipeline Parallelism NeMo: a toolkit for Conversational AI and Large Language Models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.304871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.831819Z digest=sha256:912bf194c4eda3c4182f774acded59c01f11f4cd69fbe54882b8cb69fb9d0e58

Observation ce303d1b-1216-4043-960c-62ef0b257403 · outbound

This paper cites The Hardware Lottery.

Scaling Deep Learning Training with MPMD Pipeline Parallelism The Hardware Lottery

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.834591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.834591Z digest=sha256:4379450dc074b9c9c722db82de2c90db86670e1fae5a4dee28d2619cb18dc871

Observation 16c969e5-03c0-4f46-a657-3984bd7abf54 · outbound

This paper cites DISTMM : Accelerating distributed multimodal model training.

Scaling Deep Learning Training with MPMD Pipeline Parallelism DISTMM : Accelerating distributed multimodal model training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.297798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.837340Z digest=sha256:b05763085cc387cf398d8aee2f28c9b93f2ffc7b8292380de3bc9c4a847d56b5

Observation e11d4ac1-9a8f-4e44-a1eb-fa043afc8243 · outbound

This paper cites X., Lee, H., Ngiam, J., Le, Q.

Scaling Deep Learning Training with MPMD Pipeline Parallelism X., Lee, H., Ngiam, J., Le, Q

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.290838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.840111Z digest=sha256:5c07324a3bd4eb34178737114467b3b23520e3b248fcc3283f5664332f57746e

Observation da6ae688-feb0-4673-833b-5932ec6ad8a1 · outbound

This paper cites Megascale: Scaling large language model training to more than 10,000 gpus.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Megascale: Scaling large language model training to more than 10,000 gpus

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.283713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.842540Z digest=sha256:7d09a0b7fb576460e1095a698ed74604c1b2a7c2b8156e99eb1b8ee5d7bc8507

Observation d633020d-f208-47aa-bfcf-919c4a64c8ca · outbound

This paper cites Breadth-First Pipeline Parallelism.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Breadth-First Pipeline Parallelism

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.844910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.844910Z digest=sha256:58806a9b412993f2417f28a0c7c5a77b76e1138659067b5125d42e2dacc74d68

Observation 25aee9dc-f7e2-40fd-aa0f-a5b604a316ee · outbound

This paper cites Mlir: Scaling compiler infrastructure for domain specific computation.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Mlir: Scaling compiler infrastructure for domain specific computation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.847921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.847921Z digest=sha256:62827c329c33fec69cb18d1afad228367ac94f792494b4e8e7768d298f8f312d

Observation 2b3b2526-6801-492d-8f4f-685660219c67 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

Scaling Deep Learning Training with MPMD Pipeline Parallelism GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.850831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.850831Z digest=sha256:6530b55a653df09d317eb6adb996e9f3be70d7ef76e79d514233b89e9cbf247e

Observation 7e5b2747-743f-40fd-b377-275a8e5b7bb4 · outbound

This paper cites TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training.

Scaling Deep Learning Training with MPMD Pipeline Parallelism TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.853492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.853492Z digest=sha256:af271f2c8767564a700845bc934c079ae38d1615574c4e44096e3037be5b5e92

Observation 05d3bdda-1b07-48be-be7b-1baa835d2375 · outbound

This paper cites nnScaler : Constraint-Guided parallelization plan generation for deep learning training.

Scaling Deep Learning Training with MPMD Pipeline Parallelism nnScaler : Constraint-Guided parallelization plan generation for deep learning training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.275275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.855609Z digest=sha256:4533cebbe19bb152caceb8e89f862f752bc20e35435e4442fa5bac73107bd71e

Observation 69ef8b5f-b024-4bc5-97b8-d9d8565df6c8 · outbound

This paper cites I., and Stoica, I.

Scaling Deep Learning Training with MPMD Pipeline Parallelism I., and Stoica, I

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.268253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.857809Z digest=sha256:aefe88fc781d375b8907ef6720b2411b17a166108fd9f58ad813ea192a31038c

Observation 415839ce-169e-4de3-ac1a-a2e0c9f7c9b5 · outbound

This paper cites R., Ganger, G.

Scaling Deep Learning Training with MPMD Pipeline Parallelism R., Ganger, G

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.859640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.859640Z digest=sha256:84fb7cc28ddd66f5d9c86ef405dd7e77c68ec2001a57dcaa82af17ed335615a2

Observation 506b34d0-76bc-43cb-9d32-bd63359d76c2 · outbound

This paper cites Efficient large-scale language model training on GPU clusters using megatron- LM.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Efficient large-scale language model training on GPU clusters using megatron- LM

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.861312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.861312Z digest=sha256:320f246af533e52c5498b9cd67a0913a334d6b33019ae2b2c05762a59fa8aa8c

Observation 7e71a091-1a37-45d6-80f0-d925ea1e9ac6 · outbound

This paper cites Efficiently Scaling Transformer Inference.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Efficiently Scaling Transformer Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.863523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.863523Z digest=sha256:4bdd2aaf54b0f5b36e5edaed1586cf21e1a4e6bdd39d42cbfeca0dda020d7b0b

Observation 23434b05-17ef-4201-a699-71dc0ea260c2 · outbound

This paper cites Zero bubble (almost) pipeline parallelism.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Zero bubble (almost) pipeline parallelism

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.261101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.865394Z digest=sha256:ed08bc4c6d3f00e9230fdf26cd18d1f98ba0702318ad1eb921a0995af378adda

Observation b2250dfa-f9ba-4c79-b19d-84aa37f179a4 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.867300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.867300Z digest=sha256:df5e2cd2f91501a9bf520c8a89ddef3610c6d83b4287e8891cbf66c305d3d0af

Observation 0d66b898-5d83-450e-a330-79555137b4ee · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.871089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.871089Z digest=sha256:bb08c5f708104a07e3df4a62fa03a9a51e56ea69f5c1eaddc9a65309294da1d6

Observation 35c41776-aef2-47f9-b8e2-de1b42291142 · outbound

This paper cites Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.873748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.873748Z digest=sha256:da44b5a8922278f9391ce7d8f113832d8514d78a5c196caeace5690d42ec1d03

Observation 84b218d4-88fd-44d4-bd43-7e7d3c8c1b09 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.875932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.875932Z digest=sha256:d087dbee2e66f3a1af0bac96d3fb2848a339f545c75a45339b5cdd24eb19d471

Observation 7a1aa935-be64-4cc3-a3ad-1ded6930fefc · outbound

This paper cites GSPMD: General and Scalable Parallelization for ML Computation Graphs.

Scaling Deep Learning Training with MPMD Pipeline Parallelism GSPMD: General and Scalable Parallelization for ML Computation Graphs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.877852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.877852Z digest=sha256:11032c7e8ad4f3235ebf9faeecbfd69699d1c5702069d7c9f6b161f7d3d83d62

Observation 19f5bb7c-b741-4659-8308-a0c83c856c75 · outbound

This paper cites P., Gonzalez, J.

Scaling Deep Learning Training with MPMD Pipeline Parallelism P., Gonzalez, J

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.254328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.880211Z digest=sha256:b571dca799a443d2f2d06d2ec7ac473944de54ade417e10640b256ab3da30098

Observation a8926a3b-9286-4308-9bc3-b35537842853 · outbound

This paper cites write newline.

Scaling Deep Learning Training with MPMD Pipeline Parallelism write newline

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.881968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.881968Z digest=sha256:21dc2348abc4983c33af13dd091a83e3fa1b5ae2552bd9d4a7fbdcb7f935b483

Pith citing papers

No inbound Pith citation observations are available.