Pith. sign in

Paper Citation Record · LEDGER

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models

As of 11 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2501.10714.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10714 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:05:55.300522Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6eca6e43-5a8a-4632-be01-fc69e718477d · outbound

This paper cites https://developer.nvidia.com/blog/doubling-all2all- performance-with-nvidia-collective-communication-library-2-12/.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models https://developer.nvidia.com/blog/doubling-all2all- performance-with-nvidia-collective-communication-library-2-12/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.906955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.113340Z digest=sha256:5f42bc630f491b191e172aee812c210c4cb9fc22437080e6a50e610327ad2c54

Observation 7ae7939d-1a8c-450f-ae8a-9912a56d556d · outbound

This paper cites Deepspeed-inference: enabling efficient infer- ence of transformer models at unprecedented scale.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Deepspeed-inference: enabling efficient infer- ence of transformer models at unprecedented scale

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.895967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.117819Z digest=sha256:f3918924762cbb71a04bc8dc41d3151abaa0f6893414df35ad57bb6fbae2850e

Observation a8bf19b0-5037-4016-a15e-e11f7b484f37 · outbound

This paper cites Language models are few-shot learners.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Language models are few-shot learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.121971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.121971Z digest=sha256:2fbfa3bf8cb6e36a1f7a656a1b5b4a0bc6e1b3cf5e5fce914ccd6862466377e1

Observation 0dfe1503-fe16-4ecd-9e8c-9c6b07ef8dd9 · outbound

This paper cites FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.126143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.126143Z digest=sha256:90736e3d7cab2095cbc8500c7ec08f29c856f8d6bbc157af0267ca81c11ef877

Observation 07bbda01-9b0a-45de-a31f-80cd7be94982 · outbound

This paper cites Centauri: Enabling efficient sched- uling for communication-computation overlap in large model train- ing via communication partitioning.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Centauri: Enabling efficient sched- uling for communication-computation overlap in large model train- ing via communication partitioning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.877549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.130548Z digest=sha256:589a067f179cb7fc1bf4015ce4a8b9f7e711ce2fd7ecd946e8fbdb36a91df498

Observation 2c3471e6-22bb-4817-a501-5d26b5371634 · outbound

This paper cites On the representation collapse of sparse mixture of experts.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models On the representation collapse of sparse mixture of experts

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.866473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.134272Z digest=sha256:acb6c63e8b07d1ff8f495ae7ec8e9406936fa59b8db0d725f1e563938992101a

Observation 91b8be49-ef8e-47b9-87da-b812022e2b72 · outbound

This paper cites Palm: Scaling language modeling with pathways.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Palm: Scaling language modeling with pathways

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.138161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.138161Z digest=sha256:e1546043a1520860652ae9fe08fa653abeb69165f271966760e5b308cc553916

Observation 96093689-98ff-4c7c-82c8-acc315e8ae4f · outbound

This paper cites Stablemoe: Stable routing strategy for mixture of experts.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Stablemoe: Stable routing strategy for mixture of experts

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.849207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.141579Z digest=sha256:a320329ed667954516e0da7b8ee69f30a2d9b2bd99eccca14e67330e0f99fc84

Observation a9ed7357-2285-4a41-83fc-f6ff33610287 · outbound

This paper cites Large scale distributed deep networks.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Large scale distributed deep networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.145012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.145012Z digest=sha256:d3045433bab387e129f79c8359cdfeea11a151f3820c56118a75c49eae3951c5

Observation 27b22620-c9b6-4b58-bc38-3cf0ad4ac1fd · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.148417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.148417Z digest=sha256:7f6481fa15d37869a62f4894810e1d7ffd173ecd4f8cc049b0af7f55591d9e92

Observation b705c846-265e-4b57-8a8f-d58dbca2b99f · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.151848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.151848Z digest=sha256:fd43629be977f16eb6b93dbaaa3e5feb8ef2165142d1e3e5c9634b77061c90eb

Observation f150b62a-8c17-4517-afe3-92c62b7d4306 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.155744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.155744Z digest=sha256:b3e0cd5a31a8843e6b82e020cc2c6b8dbc90d4e2814ae39032e9c7e81ef621d5

Observation e984d50e-df31-4780-aa12-a2adcb3a7427 · outbound

This paper cites FastMoE: A Fast Mixture-of-Expert Training System.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models FastMoE: A Fast Mixture-of-Expert Training System

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.159063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.159063Z digest=sha256:dad839e7b3376b4c6d6474766a99388549d230ac79c0293f07133d2d56781a51

Observation 3f4daa76-9012-4571-9e41-e1ff91bc8ae1 · outbound

This paper cites FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.815293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.163115Z digest=sha256:9c2d672a3e9d723c739b97ed5fa08e48e6fbcdde012594a6ef75b6c70d623e0d

Observation 18eda69a-b686-4d42-ba83-a3c27739cd61 · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Gpipe: Efficient training of giant neural networks using pipeline parallelism

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.166393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.166393Z digest=sha256:8052409ae373b4f4d9467452434de5b7eb812dee686bbdc34b7a180a90c87549

Observation 78f4a3b7-a753-4059-98a5-2809dbed33f5 · outbound

This paper cites Experts Weights Averaging: A New General Training Scheme for Vision Transformers.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Experts Weights Averaging: A New General Training Scheme for Vision Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.169824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.169824Z digest=sha256:e15bbec93a27186f9ae3eda0f46c9105b0f831974c701aa9443132c5247107ad

Observation 634bc3ae-50d3-473e-a721-10f5ad0a058d · outbound

This paper cites Tutel: Adaptive mixture-of-experts at scale.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Tutel: Adaptive mixture-of-experts at scale

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.797394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.173791Z digest=sha256:3939d0b6ebdc16571002c0f5c08342fe2472d9581fbb3d7815c913f6050c13c9

Observation 83246816-fbbf-4c31-8cf4-f1b71791528f · outbound

This paper cites Breaking the computation and communication abstraction barrier in distributed machine learning workloads.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Breaking the computation and communication abstraction barrier in distributed machine learning workloads

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.786294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.177255Z digest=sha256:4d303e6339cc73076a3e86dcc3801340ed5ec5f5a8a3391debb691165d557041

Observation 7088d145-0197-4f83-ab97-0eee12afdced · outbound

This paper cites Highly scalable deep learning training system with mixed-precision: Training ImageNet in four minutes.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Highly scalable deep learning training system with mixed-precision: Training ImageNet in four minutes

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.775383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.180968Z digest=sha256:f2c4fea015deb2b72006e63b31fbf295fd3e54a67e5260298080a98ca71d9f8c

Observation e201b5e5-da61-4f10-9622-d1c5cf2c216a · outbound

This paper cites Mixtral of Experts.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Mixtral of Experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.184623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.184623Z digest=sha256:6b26014e857e084ecb510bfa7a3dc39db178337eb3bb88529dd1151dea233607

Observation fd46062d-05f2-4e73-b91b-3ebd3e54bb9d · outbound

This paper cites Lancet: Accelerating mixture-of-experts training by over- lapping weight gradient computation and all-to-all communication.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Lancet: Accelerating mixture-of-experts training by over- lapping weight gradient computation and all-to-all communication

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.763414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.188654Z digest=sha256:47df9ce09de5c27aa50e10cd1e042f83e6b94ba035115bc65d6da796fc2e8338

Observation 339ab4a9-001b-4c33-a978-aa7fe1876c84 · outbound

This paper cites Gshard: Scaling giant models with conditional compu- tation and automatic sharding.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Gshard: Scaling giant models with conditional compu- tation and automatic sharding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.752070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.192429Z digest=sha256:aa354953948c9b896292164c82d2c343bb7e6fcbf66296ca012bb4c72dad9876

Observation b95e5bc8-9bad-4eae-b600-eddb1a94b54b · outbound

This paper cites BASE layers: Simplifying training of large, sparse models.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models BASE layers: Simplifying training of large, sparse models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.741240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.196205Z digest=sha256:2e1ac50ea579e05b147f55d9e6413564fcdb2b0cf91a7ab87a35ca8842c1c545

Observation 8fb7b99b-61d0-48b7-882b-b28145780031 · outbound

This paper cites Acceler- ating distributed{MoE} training and inference with lina.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Acceler- ating distributed{MoE} training and inference with lina

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.730249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.199901Z digest=sha256:4ebd4d108af14ed4eded6bff3df27ffc072b55a2037efa5dc5fbf22612c6e626

Observation 71b22f36-7a78-4d02-b6df-3910757c8aae · outbound

This paper cites Janus: A unified dis- tributed training framework for sparse mixture-of-experts models.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Janus: A unified dis- tributed training framework for sparse mixture-of-experts models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.717844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.203250Z digest=sha256:2339e86f15712c921ba94d094be7d4613989f6c0f96685eb8354d70bab936b2d

Observation 36092380-1151-4b27-8fc5-ccf28957205b · outbound

This paper cites Gating dropout: Communication-efficient regularization for sparsely activated transformers.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Gating dropout: Communication-efficient regularization for sparsely activated transformers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.704357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.206956Z digest=sha256:e8f57bac45ff035666a39ed07b8b7d38b020a6dd88ec20bfad4afa5b424888bc

Observation a3148a43-8b3e-42db-a45f-4a910214841e · outbound

This paper cites Modeling task relationships in multi-task learning with multi- gate mixture-of-experts.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Modeling task relationships in multi-task learning with multi- gate mixture-of-experts

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.691636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.210485Z digest=sha256:e7a154070b5bc2917515f705c424bc4346f150a0519bad74425ef7b611f41ace

Observation ee9cd83a-4fed-4246-ad23-f565fc339481 · outbound

This paper cites Bagualu: targeting brain scale pretrained models with over 37 million cores.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Bagualu: targeting brain scale pretrained models with over 37 million cores

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.678285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.214438Z digest=sha256:6e3ae2dbb553cc0da63690a1ba5c0153690e782f7c96c461294e634a517349dd

Observation f4051df7-4d54-4b72-a120-664d48e2a369 · outbound

This paper cites Efficient large- scale language model training on GPU clusters using Megatron-LM.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Efficient large- scale language model training on GPU clusters using Megatron-LM

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.664058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.218354Z digest=sha256:1fed3048ee2a4ba8ad6a11008614c56adbb8105e5d37399c696098ec2b2600ff

Observation df6aacf6-b731-4fa8-a8e7-6483b5eca643 · outbound

This paper cites Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.645363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.222298Z digest=sha256:7987be2bc976eba82702b0da425586d4b3b54fb55ffdee80d635c068f29ef340

Observation 7a7c8003-d358-44a5-8c97-ee152cf73e38 · outbound

This paper cites HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.225891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.225891Z digest=sha256:288f0134943ea25f58f115144ea18439bcb43903f42a49df056b69b08bf14d89

Observation 5e0c7257-5456-482f-b6bc-294482416b5e · outbound

This paper cites Springer, 1999.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Springer, 1999

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.229870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.229870Z digest=sha256:b5fa6b49ead701ed295e9432901ce3b21c8e38337d3f56f9ff341161a1de6390

Observation 1b9e643d-5a85-4770-8283-463e0c0ee288 · outbound

This paper cites Parm: Efficient training of large sparsely-activated models with dedicated schedules.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Parm: Efficient training of large sparsely-activated models with dedicated schedules

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.623764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.233478Z digest=sha256:6fa52bca37c87df63b064cc45ce0db33305a3aa32efff59864d37c3431ed4325

Observation 637c6623-802f-4d7d-94e8-908e7294bf6b · outbound

This paper cites Sinclair.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Sinclair

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.610245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.236995Z digest=sha256:5a7af241f8bb7fa1a20149e7eb06c3b3175c3982964a835485b5b764d2d6c867

Observation dcf067a7-1b92-4851-b2f1-48df2593038a · outbound

This paper cites Differential evolution.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Differential evolution

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.597154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.240768Z digest=sha256:db37f6d8b4bbce74ccc03632c64dca0135ba177aa4055572d8a8f8b6ac2c296c

Observation 97fbd66a-aa5f-4700-a6e9-6a3ee8637a87 · outbound

This paper cites From Sparse to Soft Mixtures of Experts.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models From Sparse to Soft Mixtures of Experts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.244443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.244443Z digest=sha256:9b3b1d617ea9f19f87b7017b4f8f3a3d614135958bf2c0e41977c9d9209b650d

Observation 9527e6fe-90cd-4b97-be65-8a0bddcc8609 · outbound

This paper cites Beckmann.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Beckmann

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.583586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.248468Z digest=sha256:ea665394f2aeb301d92c283d3dc7ae50b8522f7b0a58224bf9302cfae250d1a0

Observation 63a85fec-f237-48cf-a2aa-62c79a08a610 · outbound

This paper cites Language models are unsupervised multitask learners.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Language models are unsupervised multitask learners

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.252114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.252114Z digest=sha256:70232f18ad44f391659a06797efb13c71f2795bcb84f783fd19c7248aed8b2f0

Observation 936a2d2d-802f-491f-bae0-7c05ef6573c6 · outbound

This paper cites Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.560292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.255661Z digest=sha256:02cd7f43a266c664f36e4cc12a39ffa52af392a67d0e3e51c6a5f3cafb255c44

Observation bb688046-91ee-43d3-96bf-356c9dabf994 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.546293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.259055Z digest=sha256:1013858b9165f3da22ab1f013b8deaa4f9f36b85468cc19265eff77e301c79ad

Observation 8829b396-86f2-4407-b2a7-edefdf954dd6 · outbound

This paper cites Exploiting simultaneous communications to accelerate data parallel distributed deep learning.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Exploiting simultaneous communications to accelerate data parallel distributed deep learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.533093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.263102Z digest=sha256:9bfd4ab9e7d7c37ad375c050fb2f63f8d8b5af5b74edf7ec00a4b5d681136584

Observation 6cb210e4-3597-4aea-b8ef-2c3ed4f88006 · outbound

This paper cites PipeMoE: Ac- celerating mixture-of-experts through adaptive pipelining.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models PipeMoE: Ac- celerating mixture-of-experts through adaptive pipelining

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.518217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.266430Z digest=sha256:2e42271ad4de2a9fd544a6dcd44adbf4b582122a806369c61c93d38f254faff8

Observation 5ade87aa-a439-476f-97d5-c54b1ed5f19d · outbound

This paper cites Schemoe: An ex- tensible mixture-of-experts distributed training system with tasks scheduling.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Schemoe: An ex- tensible mixture-of-experts distributed training system with tasks scheduling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.504495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.269557Z digest=sha256:66d810e105e7d188553ae1904af55912f0c5a7b858866bacbf7db1cf0ea90108

Observation 3fa568d8-b289-4def-bee2-4f77473dc4e9 · outbound

This paper cites A hybrid tensor-expert- data parallelism approach to optimize mixture-of-experts training.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models A hybrid tensor-expert- data parallelism approach to optimize mixture-of-experts training

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.485399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.273159Z digest=sha256:95813654356ce9a1f4be335687e0e7f1f604f7cb4c27bf522b13cb86b2baff48

Observation e9e45f2c-f3cd-4126-b2c9-30360a6bdf08 · outbound

This paper cites Attention is all you need.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Attention is all you need

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.276366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.276366Z digest=sha256:1a59bf7d02d2c419a2680bea4f767001b4c7c8c06031c36deef79da6fe1948e5

Observation 2b2112e4-8c85-4dc1-85ab-63011375483b · outbound

This paper cites Overlap communication with dependent compu- tation via decomposition in large deep learning models.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Overlap communication with dependent compu- tation via decomposition in large deep learning models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.468097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.279486Z digest=sha256:6e03692b02ba0b3ecabbb8454be8f74316d194547dc4c216837109447a98b49f

Observation d8952115-49a6-4eb3-8d5f-a75eb3d24060 · outbound

This paper cites Large batch optimization for deep learning: Training BERT in 76 minutes.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Large batch optimization for deep learning: Training BERT in 76 minutes

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.457249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.282689Z digest=sha256:0ca5c9e33a3192984a36f50d1556ab7ce5e59cfa9367d48644e8e1be8bff9e15

Observation 3c0f0013-3795-4faa-acdd-7e56b87e99e7 · outbound

This paper cites SpeechMoE: Scaling to Large Acoustic Models with Dynamic Routing Mixture of Experts.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models SpeechMoE: Scaling to Large Acoustic Models with Dynamic Routing Mixture of Experts

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.285944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.285944Z digest=sha256:9405d37f9b675e06b1ff0a471aad839c614bdad5ab2a8f23009f0a8634e4365a

Observation 9c5b9b8b-b453-497a-906e-b1a03ab170a3 · outbound

This paper cites SmartMoE: Efficiently training Sparsely-Activated mod- els through combining offline and online parallelization.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models SmartMoE: Efficiently training Sparsely-Activated mod- els through combining offline and online parallelization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.446707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.289368Z digest=sha256:1b5f014546835fb5622d48faf1f6927875c434a8b843490513405773d96a6c41

Observation d9c12ac6-24df-4160-81c0-2d54e0dd9b73 · outbound

This paper cites Pit: Optimization of dynamic sparse deep learning models via permutation invariant transformation.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Pit: Optimization of dynamic sparse deep learning models via permutation invariant transformation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.435346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.292629Z digest=sha256:315cff90f0fe30a5e7e9a986b5bedbab8798eaef81fbc5e0c0f3a114a16ba432

Observation 487fa7b9-974a-45bc-9f95-830f85be3280 · outbound

This paper cites Mixture-of-Experts with Expert Choice Routing.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Mixture-of-Experts with Expert Choice Routing

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.296446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.296446Z digest=sha256:c0f970e2d9e633582067d1b35d695b28e77af1dbe0f48a371369eff8a013bb23

Observation f4b27ede-16fc-4098-bbbe-a828271fcf63 · outbound

This paper cites Taming sparsely activated transformer with stochastic experts.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models Taming sparsely activated transformer with stochastic experts

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:55.423266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T19:05:55.300522Z digest=sha256:1e1f98a6e80142a78265ec0945c631f925b95c6280c3fe0bc620b2ee9352d322

Pith citing papers

No inbound Pith citation observations are available.