Pith. sign in

Paper Citation Record · LEDGER

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap

As of 22 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 3 inbound Pith citation observations for arXiv:2508.09591.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09591 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:04:51.571500Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T01:06:46.426582Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:07:44.209146Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1cb75582-0214-4189-b99b-4aefa2ff262d · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.466256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.307636Z digest=sha256:2924071e539f6ba493a6b5d9f0add080093cac04d0bf5a08885e0bf1bbcd60b1

Observation bd7a2b6a-3511-405f-9107-436b990b1f71 · outbound

This paper cites Gshard: Scaling giant models with conditional computation and automatic sharding,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Gshard: Scaling giant models with conditional computation and automatic sharding,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.411606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.313212Z digest=sha256:dd7dec2ebdc88fbdd5681af7b40c52d51d209e1eee36a53cb239ad57bb95a8fe

Observation 97ee3d29-b70b-4ec7-a2f6-2e4424528411 · outbound

This paper cites Mixtral of Experts.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Mixtral of Experts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.318582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.318582Z digest=sha256:f33b80c31d8b0e528b455bc7b90f35c09d31fe51e245c1416faacf637e4b66a5

Observation 0fe12440-a22f-44c7-8e7f-4c35e2ab370c · outbound

This paper cites Qwen3 Technical Report.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Qwen3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.324759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.324759Z digest=sha256:73723d767c0432af3d4eb879c76cdaa343a6cd0b7a867c469e009420f5ca3db4

Observation a4dd5e64-513d-4883-b2d9-6e0ba4258259 · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.372176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.331758Z digest=sha256:6ff8ebdd478df17456140af32d9ca09d1a47d4ca76af8f566b9c4f5458494ffd

Observation 7dfca171-e8f1-4f89-8175-4a349e2d0941 · outbound

This paper cites JetMoE: Reaching Llama2 Performance with 0.1M Dollars.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap JetMoE: Reaching Llama2 Performance with 0.1M Dollars

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.339338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.339338Z digest=sha256:ef81c2e4939ba87a741eb0b7fbbaf0a57fca9356168b617f0133214c5fe0ac75

Observation 430868f8-87b8-4556-9d26-4990572789cb · outbound

This paper cites DeepSeek-V3 Technical Report.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap DeepSeek-V3 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.347798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.347798Z digest=sha256:a6bfaed21515bf8134826de4dbb811d1ca4493fb6e0281c9be06f6b63960f506

Observation 98f84a7c-d2dd-4677-9c33-c5124866eb79 · outbound

This paper cites Tutel: Adaptive mixture-of-experts at scale,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Tutel: Adaptive mixture-of-experts at scale,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.340459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.354950Z digest=sha256:68916bde0b40ca8fb03c3dae04bbb515705b4aad3bc1e5635ee4de0dd358aff1

Observation 1dbc42e1-3ad6-4281-a3ac-c575fe6d2756 · outbound

This paper cites Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.312401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.361002Z digest=sha256:d267008c7c4ffca1dadac4cb968d460565949ac842f539db040393a5510ccf85

Observation edf06073-732f-481a-bcd5-0f03954ef7b7 · outbound

This paper cites Accelerating distributed {MoE} training and inference with lina,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Accelerating distributed {MoE} training and inference with lina,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.288439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.367266Z digest=sha256:0d23a6b1483f7e5319ce1c224ab5a766bd613f786f81edf274785288220c3d31

Observation 1d81ebaf-5a67-4147-9d8b-bd9b15b36c14 · outbound

This paper cites Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.374061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.374061Z digest=sha256:7d079ff1d75ff0349960305a4919030794ef50c8ca7776eba4c882fad36e6fce

Observation 2fc967ab-b6c7-41ab-ab82-f029d9c621f5 · outbound

This paper cites Fsmoe: A flexible and scalable training system for sparse mixture- of-experts models,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Fsmoe: A flexible and scalable training system for sparse mixture- of-experts models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.249782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.381629Z digest=sha256:9f9c82fc4a1f1ea489e0c835713bb69bba5b336185395e7b4831c3fdb4a712cd

Observation e678a3c3-96b4-4565-926c-a2cb7352bde7 · outbound

This paper cites BASE layers: Simplifying training of large, sparse models,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap BASE layers: Simplifying training of large, sparse models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.229381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.387518Z digest=sha256:3af4724e437ab8811a994aa19eecb2d431b49142f576eaf694238edb7e22bed5

Observation dcb4c457-3d14-4810-a4f7-d4f4eb7b5871 · outbound

This paper cites Mixture-of-Experts with Expert Choice Routing.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Mixture-of-Experts with Expert Choice Routing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.395288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.395288Z digest=sha256:3cf74b3eb9dcefd195788a8ae88152861bf2fe901dc50f21aec240824989f4c8

Observation b04c30a7-1cee-46c6-9058-e4f228c2d156 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.208055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.403226Z digest=sha256:97ff3610b74d72741f62c43a88a64d70a9543f4b852e94b6b945d8a16e297692

Observation fca64878-a51f-4498-84d0-b9e636da083e · outbound

This paper cites On the representation collapse of sparse mixture of experts,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap On the representation collapse of sparse mixture of experts,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.181226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.408738Z digest=sha256:92e5dc0a6d974173bb3d0a2868447e2bfcc7032e0b9198a077d5764da4a10e95

Observation a9b30a42-7534-45c3-a6b9-c85f4cd04e63 · outbound

This paper cites From Sparse to Soft Mixtures of Experts.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap From Sparse to Soft Mixtures of Experts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.415891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.415891Z digest=sha256:d3068724a3c502f80347c114c5001ba1cc0dcf8f72f58d6f3a21c6ba0ef26229

Observation afecca68-c281-4afa-bcc9-02bfd9cc36f1 · outbound

This paper cites Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.151690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.423312Z digest=sha256:defef58cc5455dfc8ccdc543d3855aa6588ccc0944baea8376eacbfb47b937f3

Observation c507ad06-8e9b-46b5-aff5-8393a810ed2e · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.132386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.431842Z digest=sha256:b19e8366bab5499e79e64b6379e58f1a8fada6f4531bb398c9277e82c983ab17

Observation 4b445470-b14e-47b9-9cb9-7e3d78bb1883 · outbound

This paper cites Bagualu: targeting brain scale pretrained models with over 37 million cores,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Bagualu: targeting brain scale pretrained models with over 37 million cores,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.110429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.439613Z digest=sha256:f05c2bb87df67088602d339b3db3f4b8889e86a2d39762aa5f106851558524f8

Observation de2d75d7-950d-4dd5-9a92-4c387e5f0251 · outbound

This paper cites Faster- MoE: modeling and optimizing training of large-scale dynamic pre- trained models,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Faster- MoE: modeling and optimizing training of large-scale dynamic pre- trained models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.086717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.446137Z digest=sha256:fe99b68d31bc4460d59ff301af7a8ff76c52319f467fa82b9b4db82a5b70b608

Observation 0da1faa3-4a03-407e-873e-fc637cc5a4a1 · outbound

This paper cites PipeMoE: Accelerating mixture- of-experts through adaptive pipelining,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap PipeMoE: Accelerating mixture- of-experts through adaptive pipelining,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.060197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.455108Z digest=sha256:55938936ad56421c73088a64089244cdc23228625476cef2f3ab095f9f3cd7bb

Observation 3be2d0c8-b444-4139-ac1e-7be7179b9241 · outbound

This paper cites SmartMoE: Efficiently training Sparsely-Activated models through combining offline and online parallelization,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap SmartMoE: Efficiently training Sparsely-Activated models through combining offline and online parallelization,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.034297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.462127Z digest=sha256:e34fef349d5ece95081a20e57b0da6b25d52ac0e70ac061a6722edb123189b62

Observation f28ca483-ff35-4eb9-9518-2f6a846ce2b8 · outbound

This paper cites Janus: A unified distributed training framework for sparse mixture-of-experts models,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Janus: A unified distributed training framework for sparse mixture-of-experts models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.014066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.474427Z digest=sha256:fd4011f46872afb267c140517c939d88ff0dda4bbbc519283455c623dbdc142c

Observation d700887f-d7c4-47c4-88e5-1275501fc589 · outbound

This paper cites Parm: Efficient training of large sparsely-activated models with dedicated schedules,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Parm: Efficient training of large sparsely-activated models with dedicated schedules,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.994739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.482427Z digest=sha256:e6ca29e960ecdb2625a0e591475f541baf8c2c8f8329ede71d052821453b8f30

Observation e83c30be-0e0e-4733-b8b0-d1a15fbda6a4 · outbound

This paper cites Cen- tauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Cen- tauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.494649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.494649Z digest=sha256:2e8f473c77ec7958d8e640c0c622c01e61f196cd0f9c103fb80275b532e1b9a2

Observation 1f78ca91-8299-40b5-a666-3600dde75cd9 · outbound

This paper cites Lancet: Accelerating mixture-of-experts training by overlapping weight gradient computation and all-to-all communication,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Lancet: Accelerating mixture-of-experts training by overlapping weight gradient computation and all-to-all communication,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.961310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.500807Z digest=sha256:7d97458ad39458456b1bca38f00fcbb3ade1badd796ca0e5798731d50f0c08b4

Observation 90e812b6-c3eb-474b-b727-262729d7fd02 · outbound

This paper cites Sp- moe: Expediting mixture-of-experts training with optimized pipelining planning,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Sp- moe: Expediting mixture-of-experts training with optimized pipelining planning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.944284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.509667Z digest=sha256:dd86057761ba9eb1bc6de782b592bfd29c861d39f816c2eef39d5e582fc835b4

Observation 2e5e7af2-e2d4-494e-bb9d-d8e6afcc504a · outbound

This paper cites Mitigating contention in stream multiprocessors for pipelined mixture of experts: An sm-aware scheduling approac,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Mitigating contention in stream multiprocessors for pipelined mixture of experts: An sm-aware scheduling approac,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.928112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.515751Z digest=sha256:219a560266b8f2d39ae0decf7208d58697fc25758053b2131bf50972da73c0ca

Observation a5eef799-1339-48b1-9d7b-3df5b2c750bb · outbound

This paper cites Mast: Efficient training of mixture-of-experts transformers with task pipelining and ordering,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Mast: Efficient training of mixture-of-experts transformers with task pipelining and ordering,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.911072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.521028Z digest=sha256:7b086453c59cad2ea69bcb7f6e3088cc144b3f8d6d65533213965642483ae1e0

Observation cb13c37a-e101-4934-a63c-04708b943edf · outbound

This paper cites Scheinfer: Efficient inference of large language models with task scheduling on moderate gpus.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Scheinfer: Efficient inference of large language models with task scheduling on moderate gpus

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.893405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.525705Z digest=sha256:346da4f80d869446e87b6a03e64dd184be0a3530e92c8ae83fedbff9c95dd9a9

Observation a549a193-b524-4e4a-ad17-7e8d658ebdc4 · outbound

This paper cites Data center tcp (dctcp),.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Data center tcp (dctcp),

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.874271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.531999Z digest=sha256:6c4a480ba674ab272eda669ee3f995ee7a58fe05407c9273977d46970202770c

Observation ae1b1ddc-afa8-41cd-956e-31be9d4bd519 · outbound

This paper cites A scalable, commodity data center network architecture,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap A scalable, commodity data center network architecture,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.856828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.538271Z digest=sha256:ade20f6c811c6096c391ec6fd1cf7c8ec366b102519e217319a282df3e4cacc8

Observation 469e70d0-d164-4a92-9ac2-e8330ed989a4 · outbound

This paper cites {TopoOpt}: Co-optimizing network topol- ogy and parallelization strategy for distributed training jobs,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap {TopoOpt}: Co-optimizing network topol- ogy and parallelization strategy for distributed training jobs,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.836229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.544173Z digest=sha256:a09624af80831806b258e232d8488420886fa6bf1e29bee6888a8ea91ae3f3e4

Observation 338d7657-7a2d-4e0e-b118-399e3b83cc26 · outbound

This paper cites Insights into deepseek-v3: Scaling challenges and reflections on hardware for ai architectures,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Insights into deepseek-v3: Scaling challenges and reflections on hardware for ai architectures,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.814898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.549688Z digest=sha256:405fe8d4da29728922d415f3051ac51acf6c79562002c2c5c8a26a01918d5852

Observation 0647cf00-4a4f-47e2-af54-efd24a285ebd · outbound

This paper cites Large scale distributed deep networks,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Large scale distributed deep networks,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.791727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.554736Z digest=sha256:4807292ba3f8ac7adfbdc348ae7a824332b9b9c41a835c84197a0b057a54acab

Observation 9ba23427-8128-4646-99b4-5e853ce4bb3b · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.559842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.559842Z digest=sha256:589a11c0495fc65f1c371e208b0bb910e46444c7fd4a7edc9eb57461dbe82d5d

Observation 752da9a0-9f57-42cf-bd63-9ff8691ede32 · outbound

This paper cites Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.565480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.565480Z digest=sha256:3644a3b66049a4f6b4860a163bdd51dbe609a7a44bb378e10a8387d9c2584ea9

Observation fdaebd7e-20d1-46c9-b8bc-29853dd642da · outbound

This paper cites \ell 1, p-norm regularization: Error bounds and convergence rate analysis of first-order methods,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap \ell 1, p-norm regularization: Error bounds and convergence rate analysis of first-order methods,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.754885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T21:04:51.571500Z digest=sha256:590475386d35f7de119bf91fbd47c9d9e266c044590c61327b94ab95032cb54c

Pith citing papers

Observation 3d3fa414-138c-4c95-845e-f8435ce927ce · inbound

ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training cites this paper.

ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.633632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T05:22:35.487173Z digest=sha256:aec87136e2828a58ae134fae7746e2dfc9eddd1e7699380ef60c3e0efeecb345

Observation 1b663c6f-23fb-4918-84fa-57d389092cb0 · inbound

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods cites this paper.

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:44:54.933618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T13:43:17.950000Z digest=sha256:3ff9f637c820e0cd5e56724bf29bb51ae54ae5b201fdbac599a2521c5a894986

Observation e1c042cc-f0e7-43ee-87af-608aeeacf1a7 · inbound

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods cites this paper.

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:07:44.237527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-11T01:06:46.426582Z digest=sha256:8650469436dbdb8dc01ba696a892db567df43a925bec5fa155d93f5e63353844