Pith. sign in

Paper Citation Record · LEDGER

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models

As of 20 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2411.10003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10003 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:10:00.382168Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e36d548d-3952-47a7-a449-edec124c2c0a · outbound

This paper cites Gshard: Scaling giant models with conditional computation and automatic sharding,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Gshard: Scaling giant models with conditional computation and automatic sharding,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.199938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.199938Z digest=sha256:23a0b62311e28507c6417dd1f287b28a4394975b18c67e64308dadcc56a640dc

Observation b452d13e-0844-444c-9281-3036ffbb069a · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Glam: Efficient scaling of language models with mixture-of-experts,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.203536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.203536Z digest=sha256:5e8ea7b2232062753d188ff2196022a9c340db01586b5bb0d4ecc47a690e835e

Observation 7f798d6e-bf4f-47ca-b80d-96dc7b51556f · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.206774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.206774Z digest=sha256:0631873f97af4cfe90774312eae2a6112cc73f848647949fdcca35737cb631d3

Observation cc6d02f9-b1e9-4827-a2fe-338b2737372e · outbound

This paper cites Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.210378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.210378Z digest=sha256:a87c9d587ca4811520af6630d998edd11c33008aeee54376180e9dfe7fa8cb71

Observation 05b77c6d-d58f-4c68-99c6-487f56dee8dd · outbound

This paper cites Tutel: Adaptive mixture-of-experts at scale,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Tutel: Adaptive mixture-of-experts at scale,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.213587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.213587Z digest=sha256:9d96c982d3b9116fedff7552d02188763bf7f14d7b43feee46aec23c8fa579b2

Observation 9dfd5072-c6ac-4d8d-b664-59294ae5c7ce · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.216540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.216540Z digest=sha256:ecfcc75a7cd8eb44554b7a79bc9308a001b59664139643485b7f87107da61b79

Observation c9efb5b9-d7e9-4890-bca6-8868a99fa61e · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.219726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.219726Z digest=sha256:72b15b1a7116c6ab9c48265fdc7dbfa3e337235b24554b3ecbc2cceb2f748dc8

Observation 09976f2d-8bef-4494-8040-2d022e3473b2 · outbound

This paper cites Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.788777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.222415Z digest=sha256:def8e1022a7cb5777b6560c2ffcf0b887f78f1184ae81254f6630e58c30c2598

Observation c36279c3-9609-4d50-a18a-153a4be382b1 · outbound

This paper cites Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.224968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.224968Z digest=sha256:37a72264f31aa41f0742a0eec8c8c385c2785bf5536b7ca2b32bbe9cc77bf221

Observation 29dd0267-7124-4cae-8baa-246793263d54 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Zero: Memory optimizations toward training trillion parameter models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.227609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.227609Z digest=sha256:08780cd05a4274fb03cb860e29aca5e1e06f76a2c8d4059ce0b60b90eb3ee9fc

Observation 49b1da81-cc16-4eb7-ae87-f6c63a864304 · outbound

This paper cites Scaling Laws for Neural Language Models.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Scaling Laws for Neural Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.231213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.231213Z digest=sha256:6fad528d24576e7e5010384134e23e639b2efdc949554a2fad6ddaf3b1bad17f

Observation c9d435d4-ea4a-4a82-bbef-b908c81fa615 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.234422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.234422Z digest=sha256:147ae026e3fcc6fce0d86f1e1cd7c8ed513c04b5b0d0eaf74f7be3c8d609d0e6

Observation 2e8315ff-cb5d-4199-b7e1-652a144dcd75 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.237219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.237219Z digest=sha256:63670ba5c40b7f3893b9541238d0700cc125d48784e15a0457a04cae1f5e1ca5

Observation 4440c7bc-d19d-4091-a2cf-65cfecebc704 · outbound

This paper cites Xlnet: Generalized autoregressive pretraining for language understanding,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Xlnet: Generalized autoregressive pretraining for language understanding,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.240658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.240658Z digest=sha256:16de6f5fed0484c8bacdb89aee13302d2a285b2ae32da375320d148a9c22c3a9

Observation 34821e7a-2354-4cba-853e-7974fa80fe15 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.243213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.243213Z digest=sha256:8c8daf15ab5dcf6596a7c098a90a2e67c71002371ef32ffffd79e6286b436a75

Observation d145997b-c7d3-446b-80e6-fb8425bac0f5 · outbound

This paper cites Language models are unsupervised multitask learners,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Language models are unsupervised multitask learners,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.246182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.246182Z digest=sha256:830d799ff9b8d12fbf29bc9fdb59fa068bc6848cdcadda44b4a6df801fa4a33d

Observation 21a68326-a51d-4cf2-9ef0-cb58780bbe26 · outbound

This paper cites Language mod- els are few-shot learners,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Language mod- els are few-shot learners,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.248941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.248941Z digest=sha256:f47ab9eb59fc3e295d5108c15593e1401da24736ec76f2ffb277c23cb6c1a9ec

Observation aff407bc-51ed-4917-b305-4a238d19ef69 · outbound

This paper cites Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.252688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.252688Z digest=sha256:112bc145b86b72714d1ed313c6b24c218fa944579c3b71204578a31261e56905

Observation 3809ba1a-448b-47f8-b717-3d5b47fde7ba · outbound

This paper cites Mixture of A Million Experts.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Mixture of A Million Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.257132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.257132Z digest=sha256:de2ca119be81bb976bef56ff112034a27efd283d7df2cf719c5a483182325f65

Observation 52602f3f-2653-4d8f-b858-a60f96c09ede · outbound

This paper cites Taming Sparsely Activated Transformer with Stochastic Experts.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Taming Sparsely Activated Transformer with Stochastic Experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.260529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.260529Z digest=sha256:f5a6d45de97b7a315c17088fcbd20c6102554308a4bae2f270f308d77a2cb9e0

Observation fc4f260d-5b54-4dd9-a9aa-07fe5e8455b5 · outbound

This paper cites Scaling laws for fine-grained mixture of experts,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Scaling laws for fine-grained mixture of experts,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.752028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.263578Z digest=sha256:9b0a6bf648d81c7dec17ba05a584d7a5bb6aa19f723a565e00495a1d5cb0c0fa

Observation e9b2b9de-b84d-42fa-82f9-8122320186ae · outbound

This paper cites Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.266892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.266892Z digest=sha256:7bf79f63a091fe11e2c08147d6db1b40837e3847f076ed016998c9d8a6ae99c1

Observation 290384aa-3190-4961-b0ae-38e50cd62f5c · outbound

This paper cites Go wider instead of deeper,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Go wider instead of deeper,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.270242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.270242Z digest=sha256:02f431b550e7f3af78e053173906264b0edc8607c87eee0b146a40975494d450

Observation cc186b9a-7c60-4240-9bf8-99c451518807 · outbound

This paper cites One Student Knows All Experts Know: From Sparse to Dense.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models One Student Knows All Experts Know: From Sparse to Dense

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.272918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.272918Z digest=sha256:ffb56e5a0eaad070c3b3b6148c358f0fb06b20208d0a984eec646adeb3be2816

Observation 38ee513d-0e03-42ba-b6f2-90914a5cf819 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.275906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.275906Z digest=sha256:e5741053cb0ab7c3777edc59530051b2f7a719fc66660120ab930fdd22115584

Observation 2d53ee6d-218c-4ed4-aa53-ae7d224fd501 · outbound

This paper cites GPT-4 Technical Report.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models GPT-4 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.279105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.279105Z digest=sha256:a139c0f567ef0d4865d769c15be9495d6e0d9347a3873e116cb82faa887daa83

Observation 699f02f3-d1f9-4b75-8533-38a0486d714c · outbound

This paper cites Stabilization of planar collective motion: All-to-all communication,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Stabilization of planar collective motion: All-to-all communication,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.736945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.282279Z digest=sha256:909ef9efca27444495340300595f913252f3d1b96f3f9e4b1fe301e24ad7afc5

Observation 8c2d3890-f7a6-4d17-bd70-7f01b8af434f · outbound

This paper cites Optimization of all-to-all communication on the blue gene/l supercomputer,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Optimization of all-to-all communication on the blue gene/l supercomputer,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.727086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.285105Z digest=sha256:616ac180d63b0c372628fcb56351956530701f300ba5f7e2ff0da9ec411bc6a1

Observation f37a21b1-5643-4dd9-9d6c-4b0286378d15 · outbound

This paper cites The hierarchical factor algorithm for all-to-all communication,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models The hierarchical factor algorithm for all-to-all communication,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.718422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.288038Z digest=sha256:dfde0bb08227ce2aa883cb07a8d67dc66637bc001a99d5bfa5fcd122781215d5

Observation af6da1f4-89f5-4d0c-8aec-55ac85df5188 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.290731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.290731Z digest=sha256:191c02a8fe5498633212f6db7daa7deb5b0aec689b5e66dae664d45b1478144b

Observation 7ebe3cf2-a672-4597-8f35-550841df79cc · outbound

This paper cites HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.293880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.293880Z digest=sha256:4d1156642a87c3114d8e8edd45e095272b5bac87337e1e63b575cd73c74212aa

Observation 826989a0-251e-4486-b7c8-e460c0678026 · outbound

This paper cites FastMoE: A Fast Mixture-of-Expert Training System.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models FastMoE: A Fast Mixture-of-Expert Training System

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.296977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.296977Z digest=sha256:5b4798eba4fe5378844a8d259de27288ea3965ad2b7fc3c20be81395d31ba983

Observation 80806519-2428-4747-949b-66856a7bfbb3 · outbound

This paper cites Moesys: A distributed and efficient mixture-of-experts training and inference system for internet services,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Moesys: A distributed and efficient mixture-of-experts training and inference system for internet services,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.300195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.300195Z digest=sha256:c1c01542c480e300173457f0fa1d74cdd33e040f6d8c8fea9319a88aac529b75

Observation 9f26843c-db03-4bab-92db-f7371fc18d86 · outbound

This paper cites Accelerating distributed moe training and inference with lina,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Accelerating distributed moe training and inference with lina,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.704177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.302914Z digest=sha256:f70203e08ca601cf61bc28063d1019054335592b5778de69d994d83d94a17862

Observation dd018d99-45b1-47b4-9e39-d869372a3e2e · outbound

This paper cites Janus: A unified distributed training framework for sparse mixture-of-experts models,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Janus: A unified distributed training framework for sparse mixture-of-experts models,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.695266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.305921Z digest=sha256:ad44102ce2fc0f8dc16727c34203f2e1dafd9f6d2e572d4ddf03c3de7d0cf2e6

Observation 7e2bfbb1-5918-4be1-88c8-a3c798fc9ed9 · outbound

This paper cites Amp: Automatically finding model parallel strategies with heterogeneity awareness,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Amp: Automatically finding model parallel strategies with heterogeneity awareness,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.686137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.308883Z digest=sha256:262a108eab80eac8c092183626b945dca8e34b9c0d56a82e1d9c9145c9ebcc2f

Observation 9f5c84a0-ef36-470f-be87-de44ef55257a · outbound

This paper cites Merak: An efficient distributed dnn training framework with automated 3d parallelism for giant foundation models,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Merak: An efficient distributed dnn training framework with automated 3d parallelism for giant foundation models,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.311581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.311581Z digest=sha256:33b478886cf82a9df2387717e9e995f7f40b10b8b51eb3bdf37a362883fd8516

Observation 46c22993-8ac0-47e4-a2ca-434d3d1db28e · outbound

This paper cites Hippie: A data-paralleled pipeline approach to improve memory-efficiency and scalability for large dnn training,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Hippie: A data-paralleled pipeline approach to improve memory-efficiency and scalability for large dnn training,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.672088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.314245Z digest=sha256:e37806df53ffab5ae47cbec5a7b9b058f50e0959a25eb4e6eab1ffa8df8c2dbc

Observation 360af553-4124-49e6-bedf-1f840cc44f51 · outbound

This paper cites Colossal-ai: A unified deep learning system for large-scale parallel training,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Colossal-ai: A unified deep learning system for large-scale parallel training,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.317161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.317161Z digest=sha256:c3adc62042590c8d1ad4a584e246e0af717abbd6d785907d60cebac0510a8c41

Observation 6e0242cb-1274-4526-880d-b762bbc4d5af · outbound

This paper cites Piper: Mul- tidimensional planner for dnn parallelization,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Piper: Mul- tidimensional planner for dnn parallelization,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.658455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.320208Z digest=sha256:175ff194e15696d7cc46d620e26c13a8f4a6e84a95fb41b7c2a7d6cfefb02425

Observation b951c0a5-908f-4ecc-bc98-899fa3472acf · outbound

This paper cites Parallel intelligent computing: development and challenges,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Parallel intelligent computing: development and challenges,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.649130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.322933Z digest=sha256:30041b3d25e88a769b1ce58d24389023447ad859df8ce82aa639b9832a750337

Observation 8d3a6b4a-c378-4a98-b349-19dc4cfa6be5 · outbound

This paper cites Zero- infinity: Breaking the gpu memory wall for extreme scale deep learning,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Zero- infinity: Breaking the gpu memory wall for extreme scale deep learning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.639818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.325797Z digest=sha256:4bf8cb786cc882edcb8f4a0e48c2ee613c0733c653282244c4a1f595ffc527f2

Observation 72044f93-4f03-47a2-b08f-2c859365b88a · outbound

This paper cites Zero-offload: Democratizing billion-scale model training,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Zero-offload: Democratizing billion-scale model training,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.630265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.329512Z digest=sha256:306d1a8e57a2762dacb49433c9092f442ea8dda8a4acf8b1caea977460a59c87

Observation 10ce9216-144c-4931-bc69-a483136a8b34 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.332429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.332429Z digest=sha256:bcf66c6c81920b9b70799b6f0678215d0c85298f4ecdfe7c0c436afdc8cb912f

Observation 51bd6a73-b9cd-4b03-8802-1d55e0763222 · outbound

This paper cites MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.335691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.335691Z digest=sha256:6b3b5475345115bb3f4a612dff667451a91041a18d88682130e72791391517b7

Observation c05e997e-1f81-4404-a78f-22131273c31c · outbound

This paper cites Maximizing Parallelism in Distributed Training for Huge Neural Networks.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Maximizing Parallelism in Distributed Training for Huge Neural Networks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.338766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.338766Z digest=sha256:c477bc92e8f4d2d8926c70bfdd54c92aed32fcacffb33982a7853d437cfbf79e

Observation 9b97b545-503f-4084-9ef2-f3196c6a566f · outbound

This paper cites Autopipe: A fast pipeline parallelism approach with balanced partitioning and micro- batch slicing,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Autopipe: A fast pipeline parallelism approach with balanced partitioning and micro- batch slicing,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.620566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.342063Z digest=sha256:ccd973f149af0d9396a3d4c5a8d9d070a1f52545ca4497832e8f873b1b5ef35a

Observation 24de2ed7-f199-4900-a52d-3a1c7e66e0e0 · outbound

This paper cites Hph: Hybrid parallelism on heterogeneous clusters for accelerating large-scale dnns training,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Hph: Hybrid parallelism on heterogeneous clusters for accelerating large-scale dnns training,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.611602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.345381Z digest=sha256:cbc1894db5962497b49f66a59555ae2f092404401093c5cf99f4eebfb5636c8a

Observation 85eae24b-9be0-41e4-9ced-6facda989775 · outbound

This paper cites Sequence Parallelism: Long Sequence Training from System Perspective.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Sequence Parallelism: Long Sequence Training from System Perspective

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.348204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.348204Z digest=sha256:4eb83163c720853024024d100d41c64b664713c0c60ed638b8bb1c4f89526426

Observation ae5eb7df-23fc-48f6-86c1-912b5b2efea5 · outbound

This paper cites Reducing activation recomputation in large transformer models,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Reducing activation recomputation in large transformer models,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.351017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.351017Z digest=sha256:0cde9cccd1b5e81ff2e4bab6766fda0abfea8c8fe0499858f8187789f8e32f40

Observation 459f212d-2063-40f8-a3d1-cb5d50a95469 · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.354500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.354500Z digest=sha256:726b92ca2f77e22f41f44267594a7183484d9546e5f3b94fc3e4029bca24d366

Observation c2d30e42-31d6-44f3-ace7-11f1858ea94a · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.357780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.357780Z digest=sha256:ab7d7ecc017fc284308fa52074f3d4ee646f0b6273f520aa0247448efabdd050

Observation 0b245745-c6a7-491f-bf00-0ca0fe2d9f0c · outbound

This paper cites Bagualu: targeting brain scale pretrained models with over 37 million cores,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Bagualu: targeting brain scale pretrained models with over 37 million cores,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.361017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.361017Z digest=sha256:3fe7324e73d1be656f4c83bb0ee216347d61305d1f31e7c0bbb68d307e21a91e

Observation 90b3f15f-2ac1-461d-a45f-6634f43a9305 · outbound

This paper cites Parm: Efficient training of large sparsely-activated models with dedicated schedules,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Parm: Efficient training of large sparsely-activated models with dedicated schedules,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.593618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.363917Z digest=sha256:1d1c17ceead2a81136b614eb8fa0aa54dddcaaa1256d4521371d77a9c008fef7

Observation 170fcd03-b062-4ae9-8564-7850af247d9e · outbound

This paper cites A hybrid tensor-expert-data parallelism approach to optimize mixture- of-experts training,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models A hybrid tensor-expert-data parallelism approach to optimize mixture- of-experts training,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.367178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.367178Z digest=sha256:ef376ba71a497505fb92e15b1706ef6a19956127e779680afc5889fd22dc7388

Observation 8b1933bb-90f3-4d74-878b-0bbc965268e7 · outbound

This paper cites Mg-wfbp: Efficient data communication for distributed synchronous sgd algorithms,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Mg-wfbp: Efficient data communication for distributed synchronous sgd algorithms,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.579108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.370350Z digest=sha256:1bc24590e1e654f65e40cda660a88abcc84763fe2fa05e1af6d81d339fb11d33

Observation 330a99b6-790c-430b-ace4-199aeb2e85b3 · outbound

This paper cites PipeTransformer: Automated Elastic Pipelining for Distributed Training of Transformers.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models PipeTransformer: Automated Elastic Pipelining for Distributed Training of Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.372977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.372977Z digest=sha256:6475dc74479062e119ce7c721b66fb505a7294e4070769ecbca09a307b8dc3c5

Observation fe168bc5-30cb-4bf0-a885-1718896f5399 · outbound

This paper cites A multidimensional communication scheduling method for hybrid parallel dnn training,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models A multidimensional communication scheduling method for hybrid parallel dnn training,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:10:00.570645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T20:10:00.376113Z digest=sha256:789f4054d1834d83c13cf8894addc50a2b174fa8b57674e91c9cce3c41bf953b

Observation b1c76ec9-0a12-47b8-96bd-a305aff7edc3 · outbound

This paper cites Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.379017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.379017Z digest=sha256:de99c7b97dc80627a536875324a40fdbcc39a1be00ac98939cc5efca29e9d206

Observation 7b64b832-e524-4522-b65d-dd22fe32f106 · outbound

This paper cites Pipemoe: Accelerating mixture- of-experts through adaptive pipelining,.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Pipemoe: Accelerating mixture- of-experts through adaptive pipelining,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.382168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.382168Z digest=sha256:20f1fa25f2e97009dfd6912ffae2cfbb82f503d056ff7ebf5acffb4e8a95f14a

Pith citing papers

No inbound Pith citation observations are available.