Pith. sign in

Paper Citation Record · LEDGER

DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2201.05596.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2201.05596 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:40.349198Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

55
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4a6e3fea-d3f0-420e-8fe3-141a6ee723d2 · inbound

THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine Translation cites this paper.

THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine Translation DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:40.349198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:40.349198Z digest=sha256:f4f4acea9ad876bce734bc5a218e9bc1d7e2d71f229d2bc9218c0f5b669bdf16

Observation f142a64e-c741-441d-9f0f-2b83526640db · inbound

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource cites this paper.

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:05:47.665184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T00:05:08.916339Z digest=sha256:fde35b4f66515b9cda6bddc5e780febd4786494e128f4d0dcb7444345990c6c5

Observation eedc4595-aabf-4780-b92a-0a605933d8d6 · inbound

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing cites this paper.

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:26.969481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:26.969481Z digest=sha256:7094fa3d51d5c322eb09892c43a5f396be4fbbacc6bca839d72b82197ea4765b

Observation e60790ab-ee19-4e09-a1ef-42a9d02db857 · inbound

Hecto: Modular Sparse Experts for Adaptive and Interpretable Reasoning cites this paper.

Hecto: Modular Sparse Experts for Adaptive and Interpretable Reasoning DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:51.170050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:51.170050Z digest=sha256:164f648432d26e491113f09c514204192d13dbfdfc59f4b1496ef067b78c00fe

Observation 1efa5c55-1c32-41a8-a2fd-f94b669458ff · inbound

Lilith: Developmental Modular LLMs with Chemical Signaling cites this paper.

Lilith: Developmental Modular LLMs with Chemical Signaling DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:48:44.601247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:48:44.601247Z digest=sha256:3f3ed7ed0493ddb88a6664b378c9af8b9e219d3438822c14b6f840d7e561ab99

Observation b8bfb265-4748-4fe9-9caa-a729f3db85ad · inbound

MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices cites this paper.

MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:03:42.700380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:03:42.700380Z digest=sha256:80f05c49ceae488274f411864e363de21c507baf3fb5aaf6e06e40b0f0b6ef5e

Observation add94ac0-d2e8-450f-beb4-56798e2c52e7 · inbound

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference cites this paper.

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:39.094750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:39.094750Z digest=sha256:10416bd553b2522cc27ff5f083a079f0182dbcd71b6833becec656152841325d

Observation b2da6820-2894-4665-8d46-9c877e452924 · inbound

InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models cites this paper.

InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:55:59.273378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:23:40.872418Z digest=sha256:64215fcd4321cbf40add0d36c915ce3f955e74565d44796359975f7545b0d2c8

Observation d66fa626-19c4-4c42-ac0c-12dac90fa6be · inbound

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns cites this paper.

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:36:10.855194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T08:29:17.149710Z digest=sha256:dc8efc4a0766d26bf3c46f405fca3a44e55a5a941bb0ccc1078f9bd6bda0dfea

Observation 7ad34d7a-9293-4320-8ee3-62a0b1adc42d · inbound

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE cites this paper.

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T18:28:57.552698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T18:26:58.696936Z digest=sha256:e6a198cf2a777f7d35319b38950c0672b49115fcfafe8885167d06d8386840a9

Observation 2b3f315d-a421-439c-9a50-5f12bf42227b · inbound

Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism cites this paper.

Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T17:13:38.704148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:04:02.418499Z digest=sha256:a898997c1d59eaed45a7449041d7b4cd97ecc4ea134852a2721568a5acc9a17b

Observation 107b3812-4841-44d2-82cf-c27f45b52dcc · inbound

UniPool: A Globally Shared Expert Pool for Mixture-of-Experts cites this paper.

UniPool: A Globally Shared Expert Pool for Mixture-of-Experts DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:26:09.286609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T11:56:21.623709Z digest=sha256:10f343dc5310e08933629b2a6fbe4d3ac9e9cccf0cde249bbc34566bed2cc253

Observation d0eb442a-f1cf-44d7-805a-660f5595b62f · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:06:15.015003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:a3d9f72d282327acff2148cdfca8f136a72ea9468e6184e2deee0386dcac956b

Observation a1849866-3667-4f10-885c-d955ce1e7070 · inbound

Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference cites this paper.

Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:11:24.719889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:30:58.729425Z digest=sha256:ea27cf6e98340a0caf882c05871c974e8aff6d29883001463182724d9e7a413a

Observation 9e642b10-4d14-4661-9f30-87781363a849 · inbound

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload cites this paper.

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:13:03.707662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:08:07.318040Z digest=sha256:2d6a34396b1d942fb41a6702d5aeb732d24cc84bbfadd2b7582505bd5b2edaa9

Observation 3015ac02-6974-4910-8eef-975fc1f85fd0 · inbound

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models cites this paper.

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:26:00.064149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:43:33.929871Z digest=sha256:5bf79ea4ce5310d0ec625f6262924511bf081adb115e527c17b52cfc0e1f4e8b

Observation 96cb56dc-3ea7-4769-a6d4-f553726eeb32 · inbound

PCCL: Process Group-Aware Scalable and Generic Collective Algorithm Synthesizer cites this paper.

PCCL: Process Group-Aware Scalable and Generic Collective Algorithm Synthesizer DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.216865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:08:19.410192Z digest=sha256:c382d04bc429de64fc3f31d83e3708c94e834d2711c35bad8f49fbc96d49371c

Observation bde65f2a-46c9-434d-b8be-446d7b679644 · inbound

EPnG: Adaptive Expert Prune-and-Grow for Parameter-Efficient MoE Fine-tuning cites this paper.

EPnG: Adaptive Expert Prune-and-Grow for Parameter-Efficient MoE Fine-tuning DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:28:44.017643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T17:25:19.699552Z digest=sha256:b2bd2252c1868bf26958441460e445f785874b46cb4a410593bf6d8fb6dda01e

Observation 77eefc14-4106-4b35-b2b6-885d49af2fd6 · inbound

Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference cites this paper.

Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T08:35:22.347459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:35:22.347459Z digest=sha256:66cfea56ca269f650533a5587e09c24f3d4f899adf9a40a589b7a6ed7c6b8962