Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:04:51.571500Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 3 inbound Pith citation observations for arXiv:2508.09591.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:04:51.571500Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T01:06:46.426582Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T01:07:44.209146Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1cb75582-0214-4189-b99b-4aefa2ff262d · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bd7a2b6a-3511-405f-9107-436b990b1f71 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Gshard: Scaling giant models with conditional computation and automatic sharding,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 97ee3d29-b70b-4ec7-a2f6-2e4424528411 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Mixtral of Experts
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fe12440-a22f-44c7-8e7f-4c35e2ab370c · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Qwen3 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4dd5e64-513d-4883-b2d9-6e0ba4258259 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7dfca171-e8f1-4f89-8175-4a349e2d0941 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap JetMoE: Reaching Llama2 Performance with 0.1M Dollars
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 430868f8-87b8-4556-9d26-4990572789cb · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap DeepSeek-V3 Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98f84a7c-d2dd-4677-9c33-c5124866eb79 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Tutel: Adaptive mixture-of-experts at scale,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1dbc42e1-3ad6-4281-a3ac-c575fe6d2756 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation edf06073-732f-481a-bcd5-0f03954ef7b7 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Accelerating distributed {MoE} training and inference with lina,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1d81ebaf-5a67-4147-9d8b-bd9b15b36c14 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fc967ab-b6c7-41ab-ab82-f029d9c621f5 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Fsmoe: A flexible and scalable training system for sparse mixture- of-experts models,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e678a3c3-96b4-4565-926c-a2cb7352bde7 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap BASE layers: Simplifying training of large, sparse models,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dcb4c457-3d14-4810-a4f7-d4f4eb7b5871 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Mixture-of-Experts with Expert Choice Routing
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b04c30a7-1cee-46c6-9058-e4f228c2d156 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fca64878-a51f-4498-84d0-b9e636da083e · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap On the representation collapse of sparse mixture of experts,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a9b30a42-7534-45c3-a6b9-c85f4cd04e63 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap From Sparse to Soft Mixtures of Experts
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afecca68-c281-4afa-bcc9-02bfd9cc36f1 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c507ad06-8e9b-46b5-aff5-8393a810ed2e · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4b445470-b14e-47b9-9cb9-7e3d78bb1883 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Bagualu: targeting brain scale pretrained models with over 37 million cores,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation de2d75d7-950d-4dd5-9a92-4c387e5f0251 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Faster- MoE: modeling and optimizing training of large-scale dynamic pre- trained models,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0da1faa3-4a03-407e-873e-fc637cc5a4a1 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap PipeMoE: Accelerating mixture- of-experts through adaptive pipelining,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3be2d0c8-b444-4139-ac1e-7be7179b9241 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap SmartMoE: Efficiently training Sparsely-Activated models through combining offline and online parallelization,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f28ca483-ff35-4eb9-9518-2f6a846ce2b8 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Janus: A unified distributed training framework for sparse mixture-of-experts models,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d700887f-d7c4-47c4-88e5-1275501fc589 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Parm: Efficient training of large sparsely-activated models with dedicated schedules,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e83c30be-0e0e-4733-b8b0-d1a15fbda6a4 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Cen- tauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f78ca91-8299-40b5-a666-3600dde75cd9 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Lancet: Accelerating mixture-of-experts training by overlapping weight gradient computation and all-to-all communication,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 90e812b6-c3eb-474b-b727-262729d7fd02 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Sp- moe: Expediting mixture-of-experts training with optimized pipelining planning,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2e5e7af2-e2d4-494e-bb9d-d8e6afcc504a · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Mitigating contention in stream multiprocessors for pipelined mixture of experts: An sm-aware scheduling approac,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a5eef799-1339-48b1-9d7b-3df5b2c750bb · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Mast: Efficient training of mixture-of-experts transformers with task pipelining and ordering,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cb13c37a-e101-4934-a63c-04708b943edf · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Scheinfer: Efficient inference of large language models with task scheduling on moderate gpus
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a549a193-b524-4e4a-ad17-7e8d658ebdc4 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Data center tcp (dctcp),
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ae1b1ddc-afa8-41cd-956e-31be9d4bd519 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap A scalable, commodity data center network architecture,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 469e70d0-d164-4a92-9ac2-e8330ed989a4 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap {TopoOpt}: Co-optimizing network topol- ogy and parallelization strategy for distributed training jobs,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 338d7657-7a2d-4e0e-b118-399e3b83cc26 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Insights into deepseek-v3: Scaling challenges and reflections on hardware for ai architectures,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0647cf00-4a4f-47e2-af54-efd24a285ebd · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Large scale distributed deep networks,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9ba23427-8128-4646-99b4-5e853ce4bb3b · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 752da9a0-9f57-42cf-bd63-9ff8691ede32 · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdaebd7e-20d1-46c9-b8bc-29853dd642da · outbound
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap \ell 1, p-norm regularization: Error bounds and convergence rate analysis of first-order methods,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3d3fa414-138c-4c95-845e-f8435ce927ce · inbound
ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1b663c6f-23fb-4918-84fa-57d389092cb0 · inbound
UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e1c042cc-f0e7-43ee-87af-608aeeacf1a7 · inbound
UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.