Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:05:55.300522Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2501.10714.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:05:55.300522Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6eca6e43-5a8a-4632-be01-fc69e718477d · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models https://developer.nvidia.com/blog/doubling-all2all- performance-with-nvidia-collective-communication-library-2-12/
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7ae7939d-1a8c-450f-ae8a-9912a56d556d · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Deepspeed-inference: enabling efficient infer- ence of transformer models at unprecedented scale
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a8bf19b0-5037-4016-a15e-e11f7b484f37 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Language models are few-shot learners
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dfe1503-fe16-4ecd-9e8c-9c6b07ef8dd9 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07bbda01-9b0a-45de-a31f-80cd7be94982 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Centauri: Enabling efficient sched- uling for communication-computation overlap in large model train- ing via communication partitioning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2c3471e6-22bb-4817-a501-5d26b5371634 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models On the representation collapse of sparse mixture of experts
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 91b8be49-ef8e-47b9-87da-b812022e2b72 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Palm: Scaling language modeling with pathways
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96093689-98ff-4c7c-82c8-acc315e8ae4f · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Stablemoe: Stable routing strategy for mixture of experts
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a9ed7357-2285-4a41-83fc-f6ff33610287 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Large scale distributed deep networks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27b22620-c9b6-4b58-bc38-3cf0ad4ac1fd · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b705c846-265e-4b57-8a8f-d58dbca2b99f · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models PaLM-E: An Embodied Multimodal Language Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f150b62a-8c17-4517-afe3-92c62b7d4306 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e984d50e-df31-4780-aa12-a2adcb3a7427 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models FastMoE: A Fast Mixture-of-Expert Training System
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f4daa76-9012-4571-9e41-e1ff91bc8ae1 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 18eda69a-b686-4d42-ba83-a3c27739cd61 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Gpipe: Efficient training of giant neural networks using pipeline parallelism
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f4a3b7-a753-4059-98a5-2809dbed33f5 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Experts Weights Averaging: A New General Training Scheme for Vision Transformers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 634bc3ae-50d3-473e-a721-10f5ad0a058d · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Tutel: Adaptive mixture-of-experts at scale
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 83246816-fbbf-4c31-8cf4-f1b71791528f · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Breaking the computation and communication abstraction barrier in distributed machine learning workloads
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7088d145-0197-4f83-ab97-0eee12afdced · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Highly scalable deep learning training system with mixed-precision: Training ImageNet in four minutes
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e201b5e5-da61-4f10-9622-d1c5cf2c216a · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Mixtral of Experts
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd46062d-05f2-4e73-b91b-3ebd3e54bb9d · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Lancet: Accelerating mixture-of-experts training by over- lapping weight gradient computation and all-to-all communication
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 339ab4a9-001b-4c33-a978-aa7fe1876c84 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Gshard: Scaling giant models with conditional compu- tation and automatic sharding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b95e5bc8-9bad-4eae-b600-eddb1a94b54b · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models BASE layers: Simplifying training of large, sparse models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8fb7b99b-61d0-48b7-882b-b28145780031 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Acceler- ating distributed{MoE} training and inference with lina
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 71b22f36-7a78-4d02-b6df-3910757c8aae · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Janus: A unified dis- tributed training framework for sparse mixture-of-experts models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 36092380-1151-4b27-8fc5-ccf28957205b · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Gating dropout: Communication-efficient regularization for sparsely activated transformers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a3148a43-8b3e-42db-a45f-4a910214841e · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Modeling task relationships in multi-task learning with multi- gate mixture-of-experts
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ee9cd83a-4fed-4246-ad23-f565fc339481 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Bagualu: targeting brain scale pretrained models with over 37 million cores
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f4051df7-4d54-4b72-a120-664d48e2a369 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Efficient large- scale language model training on GPU clusters using Megatron-LM
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation df6aacf6-b731-4fa8-a8e7-6483b5eca643 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7a7c8003-d358-44a5-8c97-ee152cf73e38 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e0c7257-5456-482f-b6bc-294482416b5e · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Springer, 1999
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b9e643d-5a85-4770-8283-463e0c0ee288 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Parm: Efficient training of large sparsely-activated models with dedicated schedules
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 637c6623-802f-4d7d-94e8-908e7294bf6b · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Sinclair
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dcf067a7-1b92-4851-b2f1-48df2593038a · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Differential evolution
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 97fbd66a-aa5f-4700-a6e9-6a3ee8637a87 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models From Sparse to Soft Mixtures of Experts
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9527e6fe-90cd-4b97-be65-8a0bddcc8609 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Beckmann
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 63a85fec-f237-48cf-a2aa-62c79a08a610 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Language models are unsupervised multitask learners
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 936a2d2d-802f-491f-bae0-7c05ef6573c6 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bb688046-91ee-43d3-96bf-356c9dabf994 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8829b396-86f2-4407-b2a7-edefdf954dd6 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Exploiting simultaneous communications to accelerate data parallel distributed deep learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6cb210e4-3597-4aea-b8ef-2c3ed4f88006 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models PipeMoE: Ac- celerating mixture-of-experts through adaptive pipelining
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5ade87aa-a439-476f-97d5-c54b1ed5f19d · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Schemoe: An ex- tensible mixture-of-experts distributed training system with tasks scheduling
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3fa568d8-b289-4def-bee2-4f77473dc4e9 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models A hybrid tensor-expert- data parallelism approach to optimize mixture-of-experts training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e9e45f2c-f3cd-4126-b2c9-30360a6bdf08 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Attention is all you need
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b2112e4-8c85-4dc1-85ab-63011375483b · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Overlap communication with dependent compu- tation via decomposition in large deep learning models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d8952115-49a6-4eb3-8d5f-a75eb3d24060 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Large batch optimization for deep learning: Training BERT in 76 minutes
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3c0f0013-3795-4faa-acdd-7e56b87e99e7 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models SpeechMoE: Scaling to Large Acoustic Models with Dynamic Routing Mixture of Experts
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c5b9b8b-b453-497a-906e-b1a03ab170a3 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models SmartMoE: Efficiently training Sparsely-Activated mod- els through combining offline and online parallelization
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d9c12ac6-24df-4160-81c0-2d54e0dd9b73 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Pit: Optimization of dynamic sparse deep learning models via permutation invariant transformation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 487fa7b9-974a-45bc-9f95-830f85be3280 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Mixture-of-Experts with Expert Choice Routing
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4b27ede-16fc-4098-bbbe-a828271fcf63 · outbound
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Taming sparsely activated transformer with stochastic experts
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
No inbound Pith citation observations are available.