Pith. sign in

Paper Citation Record · LEDGER

Scaling Deep Learning Training with MPMD Pipeline Parallelism

As of 20 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2412.14374.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14374 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:21:38.881968Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cca14982-b4a1-4f08-983e-2a7c1e0bf6ca · outbound

This paper cites URL https://www.top500.org/system/180239.

Scaling Deep Learning Training with MPMD Pipeline Parallelism URL https://www.top500.org/system/180239

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.343082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.800001Z digest=sha256:a730b5ddf2dddc82583061d3100f307bae46426ede9e83ed12e19bae739ba6f5

Observation 2b7404cb-0d72-4344-9513-d0b078a2de96 · outbound

This paper cites PartIR: Composing SPMD Partitioning Strategies for Machine Learning.

Scaling Deep Learning Training with MPMD Pipeline Parallelism PartIR: Composing SPMD Partitioning Strategies for Machine Learning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T12:21:39.241140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.803149Z digest=sha256:bdccbf0a919eb6f3167fe8ea55282f9303c3afb2b96001afc2b6b2be9d4c7221

Observation 583b702d-4480-4e21-b82d-241700fc0f2d · outbound

This paper cites E., Thekkath, C.

Scaling Deep Learning Training with MPMD Pipeline Parallelism E., Thekkath, C

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.335359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.806478Z digest=sha256:0bc709d093cb46cd53c76f77bf26ab141f1a394da18dc5ca9c0dbc86633555f1

Observation 169ba462-9729-485e-9869-e24430089a78 · outbound

This paper cites J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne , S., and Zhang, Q.

Scaling Deep Learning Training with MPMD Pipeline Parallelism J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne , S., and Zhang, Q

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.327203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.809473Z digest=sha256:531251fffc2e6285508b30dcff00e770137a139487090e582ae4996187046fee

Observation 6bd4f832-9cf5-4502-b5ba-978ad4580e4f · outbound

This paper cites an unresolved cited work.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:21:39.318466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.811689Z digest=sha256:cba9c721dff2a3fc3e15b756221b80a3ebd308d259cd28b799e2f3e4f5a0a0fe

Observation 50454ee1-f96d-4d8b-8d0a-553ed39ff100 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Training Deep Nets with Sublinear Memory Cost

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.814422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.814422Z digest=sha256:f86a382541d78376a10aa57729c757ed01d57bef6bd676d00bc9ea08ffdb1937

Observation 521085b4-422a-4ef0-b961-4f06d2a84a86 · outbound

This paper cites cuDNN: Efficient Primitives for Deep Learning.

Scaling Deep Learning Training with MPMD Pipeline Parallelism cuDNN: Efficient Primitives for Deep Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.817499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.817499Z digest=sha256:4edc75063847de31ef2c7325bc7a19c183bfd8176957f0a231024e7a1c412763

Observation cfe67fb1-a97d-430e-a0da-89b941f13668 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Scaling Deep Learning Training with MPMD Pipeline Parallelism PaLM: Scaling Language Modeling with Pathways

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.820248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.820248Z digest=sha256:f9d8a6f7df1ff4f67db75e45bffc6d8fc57c70a3bb78e4d8d911fb0f96637fe9

Observation f474a3f4-9cc6-4e65-9e4b-e784cdc0c0dd · outbound

This paper cites An Image is Worth 16x16 Words : Transformers for Image Recognition at Scale.

Scaling Deep Learning Training with MPMD Pipeline Parallelism An Image is Worth 16x16 Words : Transformers for Image Recognition at Scale

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.311729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.823696Z digest=sha256:25d7f8d3c1fe055e88da41d3144bcdc5e527d2d796899569f03ac8e951a2521e

Observation eee7ab37-6dc2-4689-a7e0-0f1f5b09b8e3 · outbound

This paper cites The Llama 3 Herd of Models.

Scaling Deep Learning Training with MPMD Pipeline Parallelism The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.826145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.826145Z digest=sha256:0c340addd2982b16a13e8b0121f5ea7d5f21d6a7f6379840dd2bcdd2a6a3815f

Observation 6e0a0107-32a3-4dbc-9bdf-94719833b9ae · outbound

This paper cites Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.829004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.829004Z digest=sha256:c766baa87fbdf2fa564e3246e41b6e2ff76092ea0f0b4e00652145c80b46de5b

Observation 2b855f9a-56f2-4370-b895-1f31fc336153 · outbound

This paper cites NeMo: a toolkit for Conversational AI and Large Language Models.

Scaling Deep Learning Training with MPMD Pipeline Parallelism NeMo: a toolkit for Conversational AI and Large Language Models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.304871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.831819Z digest=sha256:b564eaa6408fef609458d9889e74c68242555f34fd23cd10d512a36de7672383

Observation ce303d1b-1216-4043-960c-62ef0b257403 · outbound

This paper cites The Hardware Lottery.

Scaling Deep Learning Training with MPMD Pipeline Parallelism The Hardware Lottery

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.834591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.834591Z digest=sha256:6222722c48bac90b589eff9fd06b67e70f516eae30671ac0dfa126430d0134cd

Observation 16c969e5-03c0-4f46-a657-3984bd7abf54 · outbound

This paper cites DISTMM : Accelerating distributed multimodal model training.

Scaling Deep Learning Training with MPMD Pipeline Parallelism DISTMM : Accelerating distributed multimodal model training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.297798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.837340Z digest=sha256:4151fe37824e47a4aaf11516b8dfb62ff0c3e408841881c082a98d99fa363794

Observation e11d4ac1-9a8f-4e44-a1eb-fa043afc8243 · outbound

This paper cites X., Lee, H., Ngiam, J., Le, Q.

Scaling Deep Learning Training with MPMD Pipeline Parallelism X., Lee, H., Ngiam, J., Le, Q

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.290838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.840111Z digest=sha256:2f4a67280280d03b88ac9400f37cf3b05dde2f7f33cffcdc8508c9c6660f9ad0

Observation da6ae688-feb0-4673-833b-5932ec6ad8a1 · outbound

This paper cites Megascale: Scaling large language model training to more than 10,000 gpus.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Megascale: Scaling large language model training to more than 10,000 gpus

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.283713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.842540Z digest=sha256:27c9d5d243aa525c9fd7454b0ec4101830715e37d7015e63bdb87ad0e96e0f43

Observation d633020d-f208-47aa-bfcf-919c4a64c8ca · outbound

This paper cites Breadth-First Pipeline Parallelism.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Breadth-First Pipeline Parallelism

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.844910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.844910Z digest=sha256:68cdea0c1c744983c939f73c7a57773d2243eb28db15f046d71ddafe22a113cc

Observation 25aee9dc-f7e2-40fd-aa0f-a5b604a316ee · outbound

This paper cites Mlir: Scaling compiler infrastructure for domain specific computation.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Mlir: Scaling compiler infrastructure for domain specific computation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.847921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.847921Z digest=sha256:8ac81a98ac69e3d632f283efce0d3a5055243a173f307e1bfc36ac6ccfac2602

Observation 2b3b2526-6801-492d-8f4f-685660219c67 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

Scaling Deep Learning Training with MPMD Pipeline Parallelism GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.850831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.850831Z digest=sha256:3554c9d3cd749adc9a76be4579a0045d7a3010013344cee978641430665636fb

Observation 7e5b2747-743f-40fd-b377-275a8e5b7bb4 · outbound

This paper cites TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training.

Scaling Deep Learning Training with MPMD Pipeline Parallelism TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.853492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.853492Z digest=sha256:cb62d4abade2d6aa0939364e857ed50c5e322da8ade9c276b0e50a19351d935c

Observation 05d3bdda-1b07-48be-be7b-1baa835d2375 · outbound

This paper cites nnScaler : Constraint-Guided parallelization plan generation for deep learning training.

Scaling Deep Learning Training with MPMD Pipeline Parallelism nnScaler : Constraint-Guided parallelization plan generation for deep learning training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.275275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.855609Z digest=sha256:34aa2853520987a614d2e737075fa4d7b7f07f143e5e358f37f57ac54fec9200

Observation 69ef8b5f-b024-4bc5-97b8-d9d8565df6c8 · outbound

This paper cites I., and Stoica, I.

Scaling Deep Learning Training with MPMD Pipeline Parallelism I., and Stoica, I

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.268253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.857809Z digest=sha256:aecd538237e212cafe25345a301e555b8edcde3e48931be6d2f591fb5b347f41

Observation 415839ce-169e-4de3-ac1a-a2e0c9f7c9b5 · outbound

This paper cites R., Ganger, G.

Scaling Deep Learning Training with MPMD Pipeline Parallelism R., Ganger, G

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.859640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.859640Z digest=sha256:380d7287a80d1e0d085d277881b6283e72e9eab38c854f9fff19d021dff774ed

Observation 506b34d0-76bc-43cb-9d32-bd63359d76c2 · outbound

This paper cites Efficient large-scale language model training on GPU clusters using megatron- LM.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Efficient large-scale language model training on GPU clusters using megatron- LM

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.861312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.861312Z digest=sha256:5743fee3dba3e61796cb1ecb248bf646c6058b969d097b120549b6fd9b7ec7e3

Observation 7e71a091-1a37-45d6-80f0-d925ea1e9ac6 · outbound

This paper cites Efficiently Scaling Transformer Inference.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Efficiently Scaling Transformer Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.863523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.863523Z digest=sha256:1149d578b59d1af096c306f1665645407a139454de9c6aef05e0f4914cd13465

Observation 23434b05-17ef-4201-a699-71dc0ea260c2 · outbound

This paper cites Zero bubble (almost) pipeline parallelism.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Zero bubble (almost) pipeline parallelism

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.261101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.865394Z digest=sha256:10565bc0e753750b6780ed07fbe81b428c822b589be33c2a8d660165431da4a3

Observation b2250dfa-f9ba-4c79-b19d-84aa37f179a4 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.867300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.867300Z digest=sha256:db9e8ea5ca8899e01ff2931e0143079bde0cc1773eea166d344e4123df556e37

Observation 0d66b898-5d83-450e-a330-79555137b4ee · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.871089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.871089Z digest=sha256:922ec1401f6706b60b2604ee8835262bedeb02a6b5ce796ea3189125e1fceefb

Observation 35c41776-aef2-47f9-b8e2-de1b42291142 · outbound

This paper cites Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.873748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.873748Z digest=sha256:ba82d841e52bd7b559ad6ebc58f4669060c105784f1a3b1f054d5b0e3abcd468

Observation 84b218d4-88fd-44d4-bd43-7e7d3c8c1b09 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.875932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.875932Z digest=sha256:2562bb06ef8853d1353fb8c11fb3ae79a71231a9ff93ab10dba38bb918010e70

Observation 7a1aa935-be64-4cc3-a3ad-1ded6930fefc · outbound

This paper cites GSPMD: General and Scalable Parallelization for ML Computation Graphs.

Scaling Deep Learning Training with MPMD Pipeline Parallelism GSPMD: General and Scalable Parallelization for ML Computation Graphs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.877852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.877852Z digest=sha256:ac4fdc84c02f890392d5e0f573438f1c3459157adb3f9a4295c44082f4ca8993

Observation 19f5bb7c-b741-4659-8308-a0c83c856c75 · outbound

This paper cites P., Gonzalez, J.

Scaling Deep Learning Training with MPMD Pipeline Parallelism P., Gonzalez, J

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:21:39.254328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T12:21:38.880211Z digest=sha256:92d90a74b75f035ac0fbf822c5347dd79f9a18081744db2dad3b18da4f4b33b3

Observation a8926a3b-9286-4308-9bc3-b35537842853 · outbound

This paper cites write newline.

Scaling Deep Learning Training with MPMD Pipeline Parallelism write newline

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.881968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.881968Z digest=sha256:6da4fb34cfba5446ecd7297d979cafbfa7e96c72ffdb76548d1ab01ca945d0f0

Pith citing papers

No inbound Pith citation observations are available.