Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T20:10:00.382168Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2411.10003.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T20:10:00.382168Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e36d548d-3952-47a7-a449-edec124c2c0a · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Gshard: Scaling giant models with conditional computation and automatic sharding,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b452d13e-0844-444c-9281-3036ffbb069a · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Glam: Efficient scaling of language models with mixture-of-experts,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f798d6e-bf4f-47ca-b80d-96dc7b51556f · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc6d02f9-b1e9-4827-a2fe-338b2737372e · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b77c6d-d58f-4c68-99c6-487f56dee8dd · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Tutel: Adaptive mixture-of-experts at scale,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dfd5072-c6ac-4d8d-b664-59294ae5c7ce · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9efb5b9-d7e9-4890-bca6-8868a99fa61e · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09976f2d-8bef-4494-8040-2d022e3473b2 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c36279c3-9609-4d50-a18a-153a4be382b1 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29dd0267-7124-4cae-8baa-246793263d54 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Zero: Memory optimizations toward training trillion parameter models,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49b1da81-cc16-4eb7-ae87-f6c63a864304 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Scaling Laws for Neural Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d435d4-ea4a-4a82-bbef-b908c81fa615 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e8315ff-cb5d-4199-b7e1-652a144dcd75 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Exploring the limits of transfer learning with a unified text-to-text transformer,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4440c7bc-d19d-4091-a2cf-65cfecebc704 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Xlnet: Generalized autoregressive pretraining for language understanding,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34821e7a-2354-4cba-853e-7974fa80fe15 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d145997b-c7d3-446b-80e6-fb8425bac0f5 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Language models are unsupervised multitask learners,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21a68326-a51d-4cf2-9ef0-cb58780bbe26 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Language mod- els are few-shot learners,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aff407bc-51ed-4917-b305-4a238d19ef69 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3809ba1a-448b-47f8-b717-3d5b47fde7ba · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Mixture of A Million Experts
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52602f3f-2653-4d8f-b858-a60f96c09ede · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Taming Sparsely Activated Transformer with Stochastic Experts
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc4f260d-5b54-4dd9-a9aa-07fe5e8455b5 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Scaling laws for fine-grained mixture of experts,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e9b2b9de-b84d-42fa-82f9-8122320186ae · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 290384aa-3190-4961-b0ae-38e50cd62f5c · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Go wider instead of deeper,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc186b9a-7c60-4240-9bf8-99c451518807 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models One Student Knows All Experts Know: From Sparse to Dense
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38ee513d-0e03-42ba-b6f2-90914a5cf819 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models ST-MoE: Designing Stable and Transferable Sparse Expert Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d53ee6d-218c-4ed4-aa53-ae7d224fd501 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models GPT-4 Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 699f02f3-d1f9-4b75-8533-38a0486d714c · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Stabilization of planar collective motion: All-to-all communication,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8c2d3890-f7a6-4d17-bd70-7f01b8af434f · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Optimization of all-to-all communication on the blue gene/l supercomputer,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f37a21b1-5643-4dd9-9d6c-4b0286378d15 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models The hierarchical factor algorithm for all-to-all communication,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation af6da1f4-89f5-4d0c-8aec-55ac85df5188 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ebe3cf2-a672-4597-8f35-550841df79cc · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 826989a0-251e-4486-b7c8-e460c0678026 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models FastMoE: A Fast Mixture-of-Expert Training System
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80806519-2428-4747-949b-66856a7bfbb3 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Moesys: A distributed and efficient mixture-of-experts training and inference system for internet services,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f26843c-db03-4bab-92db-f7371fc18d86 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Accelerating distributed moe training and inference with lina,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dd018d99-45b1-47b4-9e39-d869372a3e2e · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Janus: A unified distributed training framework for sparse mixture-of-experts models,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7e2bfbb1-5918-4be1-88c8-a3c798fc9ed9 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Amp: Automatically finding model parallel strategies with heterogeneity awareness,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9f5c84a0-ef36-470f-be87-de44ef55257a · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Merak: An efficient distributed dnn training framework with automated 3d parallelism for giant foundation models,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46c22993-8ac0-47e4-a2ca-434d3d1db28e · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Hippie: A data-paralleled pipeline approach to improve memory-efficiency and scalability for large dnn training,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 360af553-4124-49e6-bedf-1f840cc44f51 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Colossal-ai: A unified deep learning system for large-scale parallel training,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e0242cb-1274-4526-880d-b762bbc4d5af · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Piper: Mul- tidimensional planner for dnn parallelization,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b951c0a5-908f-4ecc-bc98-899fa3472acf · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Parallel intelligent computing: development and challenges,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8d3a6b4a-c378-4a98-b349-19dc4cfa6be5 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Zero- infinity: Breaking the gpu memory wall for extreme scale deep learning,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 72044f93-4f03-47a2-b08f-2c859365b88a · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Zero-offload: Democratizing billion-scale model training,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 10ce9216-144c-4931-bc69-a483136a8b34 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51bd6a73-b9cd-4b03-8802-1d55e0763222 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c05e997e-1f81-4404-a78f-22131273c31c · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Maximizing Parallelism in Distributed Training for Huge Neural Networks
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b97b545-503f-4084-9ef2-f3196c6a566f · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Autopipe: A fast pipeline parallelism approach with balanced partitioning and micro- batch slicing,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 24de2ed7-f199-4900-a52d-3a1c7e66e0e0 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Hph: Hybrid parallelism on heterogeneous clusters for accelerating large-scale dnns training,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 85eae24b-9be0-41e4-9ced-6facda989775 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Sequence Parallelism: Long Sequence Training from System Perspective
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae5eb7df-23fc-48f6-86c1-912b5b2efea5 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Reducing activation recomputation in large transformer models,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 459f212d-2063-40f8-a3d1-cb5d50a95469 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2d30e42-31d6-44f3-ace7-11f1858ea94a · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Ring Attention with Blockwise Transformers for Near-Infinite Context
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b245745-c6a7-491f-bf00-0ca0fe2d9f0c · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Bagualu: targeting brain scale pretrained models with over 37 million cores,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90b3f15f-2ac1-461d-a45f-6634f43a9305 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Parm: Efficient training of large sparsely-activated models with dedicated schedules,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 170fcd03-b062-4ae9-8564-7850af247d9e · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models A hybrid tensor-expert-data parallelism approach to optimize mixture- of-experts training,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b1933bb-90f3-4d74-878b-0bbc965268e7 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Mg-wfbp: Efficient data communication for distributed synchronous sgd algorithms,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 330a99b6-790c-430b-ace4-199aeb2e85b3 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models PipeTransformer: Automated Elastic Pipelining for Distributed Training of Transformers
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe168bc5-30cb-4bf0-a885-1718896f5399 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models A multidimensional communication scheduling method for hybrid parallel dnn training,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b1c76ec9-0a12-47b8-96bd-a305aff7edc3 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b64b832-e524-4522-b65d-dd22fe32f106 · outbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Pipemoe: Accelerating mixture- of-experts through adaptive pipelining,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.