Pith. sign in

Paper Citation Record · LEDGER

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism

As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2506.22175.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22175 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:15:28.228455Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact3
  • verified fuzzy20
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 80e302e1-70b3-48e4-9fb2-efa3d16c2011 · outbound

This paper cites On the optimization of deep networks: Implicit acceleration by overparameterization,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism On the optimization of deep networks: Implicit acceleration by overparameterization,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:32.788377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:25.501953Z digest=sha256:059c2a963f0280192cd6fc9e0a3a25d3f842b21db4bbdfcf48329e7730526a76

Observation da7e2b1f-e19f-4471-8b74-013bcc6e3fac · outbound

This paper cites Exploring the limits of weakly supervised pretraining,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Exploring the limits of weakly supervised pretraining,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:32.565005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:25.611239Z digest=sha256:7e4792c551c77a646588b44d9cd235076f27c52b9b6a492e4c9f02e82713a3e7

Observation c0832550-adb5-4c73-999e-1b1107862a00 · outbound

This paper cites Antman: Dynamic scaling on gpu clusters for deep learning,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Antman: Dynamic scaling on gpu clusters for deep learning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:32.311875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:25.755050Z digest=sha256:755f9e562aee061a59ee3e3b4ed68400fea1fb1dcfb7b0f56ca6f03c3a3a7c95

Observation b4e4428f-2f87-4a99-a327-3399924419b9 · outbound

This paper cites Whale: Efficient giant model training over heterogeneous gpus,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Whale: Efficient giant model training over heterogeneous gpus,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:32.114853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:25.901684Z digest=sha256:e888772928367eb8b8fe55b27dd0d1d9719dccd109d8a6fee6584bb235b809a0

Observation 4a45c5aa-fed0-4ad2-903a-aa8ea74b4ac7 · outbound

This paper cites Axonn: An asynchronous, message-driven parallel framework for extreme-scale deep learning,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Axonn: An asynchronous, message-driven parallel framework for extreme-scale deep learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:31.870218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:26.037612Z digest=sha256:df164e00e7086e325f0137bf9cf1e990daefd01bf13d2568531c9bb6c797c4bc

Observation 202fd7a0-346f-4d78-b43d-6d7a2e0321a8 · outbound

This paper cites An efficient and non-intrusive gpu schedul- ing framework for deep learning training systems,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism An efficient and non-intrusive gpu schedul- ing framework for deep learning training systems,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:31.651693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:26.141093Z digest=sha256:2ed2919dd9b8b92a401a9dde52dbf335c620fe1072a67fbd2ac9353efff0b3d9

Observation f4fe10b9-0f1b-4450-9e60-b23b2571da6f · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:31.543830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:26.269023Z digest=sha256:1d26ca4ad77343f3cd34dbfa776057a4ef5fb4318575e371dcd65e753b8f1bed

Observation ace6e3a0-ef11-45f4-a95e-78859d2aefd2 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:26.398427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:26.398427Z digest=sha256:8d6a4896d81860ea091371b6d47dcb4d2b03c0aaf2bb95d2402a7d74b55adda5

Observation 5ea1faec-ca1e-4b18-8d4d-900f39b7ebea · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:31.337050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:26.502541Z digest=sha256:fdeb49a8a285171d938d74e5d209ef96887a4a92b97c0f621085b4b7d206c7cb

Observation c0368665-aa1a-4d71-a103-10d3923d409e · outbound

This paper cites Language mod- els are few-shot learners,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Language mod- els are few-shot learners,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:26.600395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:26.600395Z digest=sha256:f1dd29b514b8fcd9ad57d05350cdac082c0675e3bf2416feb53fa0f842e1eb66

Observation 9e9724b0-cf24-495c-a834-667601713bd7 · outbound

This paper cites BLISS: Robust Sequence-to-Sequence Learning via Self-Supervised Input Representation.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism BLISS: Robust Sequence-to-Sequence Learning via Self-Supervised Input Representation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:15:28.679030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:26.715460Z digest=sha256:59c39fb2f46cd8f967f22414695b46fefc79dc47e45161a9884960a5630b5b0f

Observation 9129e399-e0bb-4626-8d95-36dbff9f54fe · outbound

This paper cites E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and Generation.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and Generation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:15:28.537261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:26.817024Z digest=sha256:08d8891bebcb2120bad921353f0f162e5c0a63a57d4cc72accf5d5eee29362c6

Observation a588efc2-b4fd-4dd4-8bdc-8dc75e06d433 · outbound

This paper cites Unsu- pervised cross-lingual representation learning at scale,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Unsu- pervised cross-lingual representation learning at scale,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:31.199126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:26.871985Z digest=sha256:4bcf12afbb5db564003d0e4c19e2e36f3054cb65ac6be3189fff5ee7bcf79a2e

Observation c51cf98b-e571-425f-bbf9-cb073dfb0325 · outbound

This paper cites Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:26.922616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:26.922616Z digest=sha256:4879dd54d1ef573ca4e174aeaaaf2535911514479c9ddc4513655c94acbd5dbb

Observation 45b1a33d-616a-487a-bccb-fb8ccd773284 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:26.970238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:26.970238Z digest=sha256:de397c69606d012dbc905d78271f2f5fe4d9118d015368b3f948715c2c0037d5

Observation 1e365e47-038f-4936-851c-975894fcde0d · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:27.043950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:27.043950Z digest=sha256:9d6c9ea6e5e1556783e587bbdf7be585137681b4a2ef3b88c58260b56ab2b651

Observation 25121131-c2c5-4c52-98d2-7cfd717c1894 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:30.997680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:27.110599Z digest=sha256:38935fd9384a7817ab0cc5d25091f84952d8f957348b0ca356d71171b14c0432

Observation 52d7527b-358b-4360-af2c-cbe2b11d13a6 · outbound

This paper cites PAD-Net: An Efficient Framework for Dynamic Networks.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism PAD-Net: An Efficient Framework for Dynamic Networks

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:15:28.401357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:27.198838Z digest=sha256:a2801027fdeb5c6b953d65b54ec8eaa62a88520601456d692134a250b1d0b62b

Observation d1ab2883-4be9-4ebd-8bcc-18f2bfd914b5 · outbound

This paper cites Base layers: Simplifying training of large, sparse models,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Base layers: Simplifying training of large, sparse models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:30.739187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:27.260718Z digest=sha256:0ef31332ef5b7021278c4e94a851ead5721df49e72f078918caaa2a161509cf1

Observation 81121fdc-711e-4918-8173-af010a61e654 · outbound

This paper cites Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:30.504450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:27.323226Z digest=sha256:0789360b8678a835794ee4712cdd801ab60454b30d08b12876ace4bda139ea33

Observation f5762c01-f99e-42bd-903b-cb895bac91df · outbound

This paper cites Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation AI scale,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation AI scale,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:30.256375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:27.397140Z digest=sha256:fed782b883ffb6914dff31868de2e4ed13ad71b38cf7cb24c7b02194d89b0fc8

Observation 10ad3246-ce10-4f75-8cf0-b201c7eae5b4 · outbound

This paper cites Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:30.040462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:27.451658Z digest=sha256:b200d616ac0cccde80271d49a399cc09638fa98fdc03b17179401f41a15fb09e

Observation da501c63-0f1f-4977-a420-cf8acc594f93 · outbound

This paper cites Scalable distributed dl training: Batching communication and computation,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Scalable distributed dl training: Batching communication and computation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:29.822345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:27.495731Z digest=sha256:22745eada7f47d57f0e48e342a8b10cf77062609e9b2bf130e1d35b150f89b53

Observation 77bc7860-4971-447f-867d-c7965f20e622 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Zero: Memory optimizations toward training trillion parameter models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:29.593956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:27.557041Z digest=sha256:487c75526ea4b3b4e444493fee24449000a948b0d481b6519a9a8702b56a7a44

Observation 89d2c24a-c7ce-4c6e-b494-0c6e08c1f0db · outbound

This paper cites Scalable and Efficient MoE Training for Multitask Multilingual Models.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Scalable and Efficient MoE Training for Multitask Multilingual Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:27.618354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:27.618354Z digest=sha256:448ff7dc36d522bbc714cb05120f1ee97727d8317f2f2a954a806779f211f5d4

Observation e3157efc-d1e2-4f18-928f-81e7d59f0558 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Training Deep Nets with Sublinear Memory Cost

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:27.666690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:27.666690Z digest=sha256:fc4848c8e86cb5b56c7a13593a2ec3ad44af55fafb7d584c0e3a3d81cd833629

Observation 61131dc7-3ae8-4b15-b543-4815bf7e6cda · outbound

This paper cites vdnn: Virtualized deep neural networks for scalable, memory-efficient neural network design,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism vdnn: Virtualized deep neural networks for scalable, memory-efficient neural network design,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:29.248516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:27.733735Z digest=sha256:db549111cb0307416ffe48484fde9ae3489b154ba9e1de7b11eeffee70ade05d

Observation 954749e8-cf15-4840-8bcd-07b8fddcad2b · outbound

This paper cites Buddy compression: Enabling larger memory for deep learning and hpc workloads on gpus,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Buddy compression: Enabling larger memory for deep learning and hpc workloads on gpus,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:29.046713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:27.805975Z digest=sha256:94d6f373c0fd54b7fd91e86351d5d061b82353058992b74986eca8304ec43249

Observation 0637980d-29df-4b1c-b823-c91db9e0624f · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Efficient large-scale language model training on gpu clusters using megatron-lm,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:28.945997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:27.873864Z digest=sha256:7fb2b3d0ff725c2071e6c779b3aaa5977170d9b4187d29852452f76de0274aff

Observation 5f77f9e4-68da-46a4-9a5b-fe15d0db598c · outbound

This paper cites Adam: A method for stochastic optimization,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Adam: A method for stochastic optimization,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:27.918794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:27.918794Z digest=sha256:2ad76c7be112f953828b3f0b6f7984b4fde0d5aed3d6d31738e7ec5219aeffc3

Observation 1df6a646-9d37-49d9-94b6-61dbb4a70cfb · outbound

This paper cites Gpipe: Efficient training of giant neu- ral networks using pipeline parallelism,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Gpipe: Efficient training of giant neu- ral networks using pipeline parallelism,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:27.975062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:27.975062Z digest=sha256:cba7fd7db1ac20e06876c6775b8ac048b98eea92a828b6ba6bb81bdf3931e289

Observation cc6a1c04-8cc7-41c1-9dfd-bd462064a2e0 · outbound

This paper cites Tutel: Adaptive Mixture-of-Experts at Scale.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Tutel: Adaptive Mixture-of-Experts at Scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:28.050477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:28.050477Z digest=sha256:68d0d15c6a1d84bee2d34b913ba336f3708118b99a6cf1719bbecce2cc1a2d22

Observation 74773f93-26a2-4e1a-bccc-250a63346df8 · outbound

This paper cites Mesh-tensorflow: Deep learning for supercomputers,.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Mesh-tensorflow: Deep learning for supercomputers,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:28.801078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:28.110425Z digest=sha256:b0eb6f4d005ed1487d4f8e203c2458b4dba2d5b4ee491d5f3771e806aa7418e1

Observation 0cc0be77-5179-45f1-b922-db552ce61723 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:28.158421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:28.158421Z digest=sha256:468038a9b09dc15191edc410467a6d962d5a5e3eee0820bae398cc128f86069e

Observation fcc388e8-10f4-48cc-a89d-78c841b56b78 · outbound

This paper cites PipeDream: Fast and Efficient Pipeline Parallel DNN Training.

MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism PipeDream: Fast and Efficient Pipeline Parallel DNN Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:28.228455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:28.228455Z digest=sha256:a9371ebe7548f37ba23ece6be7855206bf6b1df1b8c52b86fee37fc4defc4a23

Pith citing papers

No inbound Pith citation observations are available.