Pith. sign in

Paper Citation Record · LEDGER

DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2201.05596.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2201.05596 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:49:10.088007Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

55
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f1ba0c67-19ca-41da-a15f-04e42cd467ac · inbound

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines cites this paper.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.741039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.741039Z digest=sha256:a6767c7a79e3f396b2920cdddd821bf4685364817e856052dd36d267573edb3a

Observation 65049aef-5c7d-47a0-8808-3a526f13763e · inbound

MixGCN: Scalable GCN Training by Mixture of Parallelism and Mixture of Accelerators cites this paper.

MixGCN: Scalable GCN Training by Mixture of Parallelism and Mixture of Accelerators DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:20:59.102053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:20:59.102053Z digest=sha256:fa2b537fc4c4e79abd7976286a73b55092b46280fa53299b473f5cbea61cced2

Observation 92503e35-092a-481a-8410-d52ecb943a1e · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:10.088007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:10.088007Z digest=sha256:08a0eedb275c6dd9b9bfc8744d32675f2420a3f1daf3432c6f4bcf188abb2bba

Observation 4a6e3fea-d3f0-420e-8fe3-141a6ee723d2 · inbound

THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine Translation cites this paper.

THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine Translation DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:40.349198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:40.349198Z digest=sha256:1531fa98883b98abda88cc8b9b18553a538f80ddf953bd5e907a58b8599cae2e

Observation f142a64e-c741-441d-9f0f-2b83526640db · inbound

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource cites this paper.

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:05:47.665184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T00:05:08.916339Z digest=sha256:c285e97295f4101eb21b0ee9f54ef3cef11676df986d0934f0302f74fed7194c

Observation eedc4595-aabf-4780-b92a-0a605933d8d6 · inbound

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing cites this paper.

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:26.969481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:26.969481Z digest=sha256:a22e866607635821f2532d63f46c6a388768e923638576b178abf199448411c8

Observation e60790ab-ee19-4e09-a1ef-42a9d02db857 · inbound

Hecto: Modular Sparse Experts for Adaptive and Interpretable Reasoning cites this paper.

Hecto: Modular Sparse Experts for Adaptive and Interpretable Reasoning DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:51.170050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:51.170050Z digest=sha256:ceaa43a78911cb377a66b5d7ee36c21715e24718fd2ba3816d6db2a622628357

Observation 1efa5c55-1c32-41a8-a2fd-f94b669458ff · inbound

Lilith: Developmental Modular LLMs with Chemical Signaling cites this paper.

Lilith: Developmental Modular LLMs with Chemical Signaling DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:48:44.601247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:48:44.601247Z digest=sha256:63863963dd297bf15aa76de6fa5700923ffacac89d1eb1af164df0b58f7ad523

Observation e9f8e3a1-bdbe-4adc-94fd-9bb5e037feab · inbound

MindSpeed RL: Distributed Dataflow for Scalable and Efficient RL Training on Ascend NPU Cluster cites this paper.

MindSpeed RL: Distributed Dataflow for Scalable and Efficient RL Training on Ascend NPU Cluster DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T18:07:35.028705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:07:35.028705Z digest=sha256:2f66a610af1ddaf649704fff84c032f31a09ce596077dd86c54e51120b608e09

Observation b8bfb265-4748-4fe9-9caa-a729f3db85ad · inbound

MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices cites this paper.

MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:03:42.700380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:03:42.700380Z digest=sha256:7026f40be011f57a998269fae85e4fd2cffce6df24156d3ad93bd6e86fea8ff2

Observation add94ac0-d2e8-450f-beb4-56798e2c52e7 · inbound

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference cites this paper.

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:39.094750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:39.094750Z digest=sha256:6f0deca6aac00cc31c99670a2e6f010767829b6544f9db00ac67db2e32a8c5d5

Observation b2da6820-2894-4665-8d46-9c877e452924 · inbound

InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models cites this paper.

InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:55:59.273378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:23:40.872418Z digest=sha256:a44a86e6496ab0b77d26b2fbc8d18bedd19bf71d0eec79d3d4a3db1bf8ce69df

Observation d66fa626-19c4-4c42-ac0c-12dac90fa6be · inbound

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns cites this paper.

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:36:10.855194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T08:29:17.149710Z digest=sha256:dad8f6b9143de7849563f97a5b4b0e8561cd62802c0cc08d60a293c5994cf2cf

Observation 7ad34d7a-9293-4320-8ee3-62a0b1adc42d · inbound

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE cites this paper.

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T18:28:57.552698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T18:26:58.696936Z digest=sha256:95c19975a1be0b0301ea766d19fe7b50af01e3cb7cb043c8e19d81a9cd51ca58

Observation 2b3f315d-a421-439c-9a50-5f12bf42227b · inbound

Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism cites this paper.

Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T17:13:38.704148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T17:04:02.418499Z digest=sha256:25a452e69c591517581a571a1442bf70b0345b20143a3f3713e24c67ec2e22e8

Observation 107b3812-4841-44d2-82cf-c27f45b52dcc · inbound

UniPool: A Globally Shared Expert Pool for Mixture-of-Experts cites this paper.

UniPool: A Globally Shared Expert Pool for Mixture-of-Experts DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:26:09.286609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T11:56:21.623709Z digest=sha256:68546bbcbd3b7c4ab0530e17d5ff0f98dfa1e4b78d83722ccadbb83eaf4a586f

Observation d0eb442a-f1cf-44d7-805a-660f5595b62f · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:06:15.015003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:480af1c73072459fcc4a2a51fdf4697142ef5a8d55cde699d9042ed2e79a4dd9

Observation a1849866-3667-4f10-885c-d955ce1e7070 · inbound

Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference cites this paper.

Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:11:24.719889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:30:58.729425Z digest=sha256:4f56c186e3639e182ec169237a622e0e9a7473329b7df804f7808c69215ad883

Observation 9e642b10-4d14-4661-9f30-87781363a849 · inbound

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload cites this paper.

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:13:03.707662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T05:08:07.318040Z digest=sha256:dbfbb006a416ef3217864738ddf045677cba1ee1c038461a0c4a108dbb90987c

Observation 3015ac02-6974-4910-8eef-975fc1f85fd0 · inbound

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models cites this paper.

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:26:00.064149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T22:43:33.929871Z digest=sha256:1d97b1bd20a6ff8af648961d4cbbc4e89f3ac12ec53bb4db9795a86c2961b3b7

Observation 96cb56dc-3ea7-4769-a6d4-f553726eeb32 · inbound

PCCL: Process Group-Aware Scalable and Generic Collective Algorithm Synthesizer cites this paper.

PCCL: Process Group-Aware Scalable and Generic Collective Algorithm Synthesizer DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.216865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T21:08:19.410192Z digest=sha256:5acd3d54f7f4883b1e80c8f176a14f0147e893512a99c156f016dedc2dccd65a

Observation bde65f2a-46c9-434d-b8be-446d7b679644 · inbound

EPnG: Adaptive Expert Prune-and-Grow for Parameter-Efficient MoE Fine-tuning cites this paper.

EPnG: Adaptive Expert Prune-and-Grow for Parameter-Efficient MoE Fine-tuning DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:28:44.017643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T17:25:19.699552Z digest=sha256:815759df1489072d7504cf7892a8de2bacbdac0b4cca09df1f37d5c54fb21d24

Observation 77eefc14-4106-4b35-b2b6-885d49af2fd6 · inbound

Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference cites this paper.

Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T08:35:22.347459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:35:22.347459Z digest=sha256:8b9c6e80571cc808897f812c61dec634194887446f9e87f2357da211b9250ec0