Pith. sign in

Paper Citation Record · LEDGER

HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2203.14685.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.14685 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:10:00.293880Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

11
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7ebe3cf2-a672-4597-8f35-550841df79cc · inbound

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models cites this paper.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.293880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.293880Z digest=sha256:4d1156642a87c3114d8e8edd45e095272b5bac87337e1e63b575cd73c74212aa

Observation 59943dbd-fe5b-468a-bb19-f13bdc852270 · inbound

Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference cites this paper.

Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:11:39.214770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:11:39.214770Z digest=sha256:3f66349f42dcba27f557898269bfa74aacde80e65ddd8d1de760f1f802487586

Observation c0ebdc64-942b-4ce3-ae81-a3302f4fd3f4 · inbound

A Survey on Inference Optimization Techniques for Mixture of Experts Models cites this paper.

A Survey on Inference Optimization Techniques for Mixture of Experts Models HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.826482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.826482Z digest=sha256:11bf400d24aca35db86ee66e1d1ed0b8426dc224ffd56fde84294fd72d578f05

Observation 7a7c8003-d358-44a5-8c97-ee152cf73e38 · inbound

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models cites this paper.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.225891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.225891Z digest=sha256:c83cad569591c28f680f3268c4041876ef68f1ef10f33c9c23311bec848ad65b

Observation 4785830a-468c-44e2-9326-b239560043ea · inbound

Mixture of Experts (MoE): A Big Data Perspective cites this paper.

Mixture of Experts (MoE): A Big Data Perspective HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 131

Resolution
unresolved
no resolver link, observed 2026-08-10T18:56:37.658955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:56:37.658955Z digest=sha256:689c11022c9cb03d3f8c532cfd7c0896e69d07f0bc8343636365d6a824d28a4c

Observation cc7bfa8b-6dc4-47de-9f03-c8c844c7093f · inbound

Hecto: Modular Sparse Experts for Adaptive and Interpretable Reasoning cites this paper.

Hecto: Modular Sparse Experts for Adaptive and Interpretable Reasoning HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:50.741606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:50.741606Z digest=sha256:4d57dea2fe90e9f71c20365ee3435dd005fcb403045053880dadab82019aca5f

Observation fdbe7140-bb2b-40c8-b365-33682c0c10ea · inbound

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs cites this paper.

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:31:17.580741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-09T19:46:13.015064Z digest=sha256:371b79ae1a15bbed62ba3437ff8d85d7fa1e431ab310016a8a55e2b0d2769031

Observation 7bf4f180-ae02-4dfb-b571-5400bba21d81 · inbound

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs cites this paper.

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:21:30.061160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T05:17:09.793360Z digest=sha256:23b2484ae1becfc4e70fa2582c2d3e2a3210b2b5ed94a1a4431326432970290d

Observation 57c4bd64-3ab0-423a-a8d5-5ccc39f8405e · inbound

Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism cites this paper.

Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T17:13:38.771953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T17:04:02.418499Z digest=sha256:4cdca6dcef2525109f94d6b3b8689e734a81dffaaf1c56efdb30ea4b41747361

Observation dd5fa225-7254-41be-acb6-3c24b9cf0e3d · inbound

Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs cites this paper.

Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:36:17.246414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T04:46:17.028437Z digest=sha256:b6c4bfe243ca4ed6f9cdea80cdbbaafa59653aca5deca5f2e971d0a35cd167ea

Observation f41717dd-3c21-4f74-aba6-9b0211cc8618 · inbound

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning cites this paper.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:14.809929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:6f0ab81512bf2163534d67e32a9f186ccfe5b5a169b8beab6f139c3711afeb3d

Observation a778404f-af0b-416c-a5a1-5df933cc78c9 · inbound

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization cites this paper.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 294

Resolution
unresolved
no resolver link, observed 2026-08-01T06:49:00.092774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:49:00.092774Z digest=sha256:4f2a09c591361a4a9fbd2139e9f2a1bedb5a754d7b8f4995357ea2b24b5e6fe5