Pith. sign in

Paper Citation Record · LEDGER

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap

As of 11 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 3 inbound Pith citation observations for arXiv:2508.09591.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09591 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:04:51.571500Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T01:06:46.426582Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:07:44.209146Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1cb75582-0214-4189-b99b-4aefa2ff262d · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.466256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.307636Z digest=sha256:5ab87f17e7348df33cef261d83c8adfee4a0c31f0a758ea79b9d44d65378d411

Observation bd7a2b6a-3511-405f-9107-436b990b1f71 · outbound

This paper cites Gshard: Scaling giant models with conditional computation and automatic sharding,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Gshard: Scaling giant models with conditional computation and automatic sharding,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.411606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.313212Z digest=sha256:76341ca7b4218485cacc4eb798ae22c23dfafbff5ce400af10bbd0057c09926f

Observation 97ee3d29-b70b-4ec7-a2f6-2e4424528411 · outbound

This paper cites Mixtral of Experts.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Mixtral of Experts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.318582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.318582Z digest=sha256:e854f4ab8ffb9382265db560fc351a5dec4a53048f0c60e8fe631688a97ed5ee

Observation 0fe12440-a22f-44c7-8e7f-4c35e2ab370c · outbound

This paper cites Qwen3 Technical Report.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Qwen3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.324759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.324759Z digest=sha256:c76c21645f4fb6fee83c7f1a04fd218069e251fac6332b6255778e585c4222a4

Observation a4dd5e64-513d-4883-b2d9-6e0ba4258259 · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.372176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.331758Z digest=sha256:434fe15db1e2847b98f8dfbac903f5a322b896746735e036ec2d778cbbec35e9

Observation 7dfca171-e8f1-4f89-8175-4a349e2d0941 · outbound

This paper cites JetMoE: Reaching Llama2 Performance with 0.1M Dollars.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap JetMoE: Reaching Llama2 Performance with 0.1M Dollars

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.339338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.339338Z digest=sha256:008b3e5c243e617df9682c6f702e40fb276622dc3892607a88a912f4c3389e02

Observation 430868f8-87b8-4556-9d26-4990572789cb · outbound

This paper cites DeepSeek-V3 Technical Report.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap DeepSeek-V3 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.347798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.347798Z digest=sha256:eb11b75ce83f33089ec5fa39043b906969f14fbf903fb0f9c80afe544dbbaaac

Observation 98f84a7c-d2dd-4677-9c33-c5124866eb79 · outbound

This paper cites Tutel: Adaptive mixture-of-experts at scale,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Tutel: Adaptive mixture-of-experts at scale,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.340459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.354950Z digest=sha256:13ddea108c39f697484f7222b69fe441c37e6ee90885e4ab1353fdc43a5e8e28

Observation 1dbc42e1-3ad6-4281-a3ac-c575fe6d2756 · outbound

This paper cites Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.312401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.361002Z digest=sha256:834342b73061754d256c01ce7be34763189a54bdb4a3a28fafa71cdf91a4dd71

Observation edf06073-732f-481a-bcd5-0f03954ef7b7 · outbound

This paper cites Accelerating distributed {MoE} training and inference with lina,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Accelerating distributed {MoE} training and inference with lina,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.288439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.367266Z digest=sha256:a65368267c586a03bc828b5f182c5e7314c5e1cf00d6773a65de054147b25faa

Observation 1d81ebaf-5a67-4147-9d8b-bd9b15b36c14 · outbound

This paper cites Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.374061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.374061Z digest=sha256:423921097345462dc3a270076a613dfc130ddf97b42925616da5e38e0ed99a5a

Observation 2fc967ab-b6c7-41ab-ab82-f029d9c621f5 · outbound

This paper cites Fsmoe: A flexible and scalable training system for sparse mixture- of-experts models,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Fsmoe: A flexible and scalable training system for sparse mixture- of-experts models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.249782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.381629Z digest=sha256:a14fd265dc39a935521a8e3ae81865e6f92a3b4f7846b20b83acf63d4a269a45

Observation e678a3c3-96b4-4565-926c-a2cb7352bde7 · outbound

This paper cites BASE layers: Simplifying training of large, sparse models,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap BASE layers: Simplifying training of large, sparse models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.229381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.387518Z digest=sha256:425fd377c6c44dd51dab3c1349f2919628eb9786598be04e9d9fbf92dd04eba3

Observation dcb4c457-3d14-4810-a4f7-d4f4eb7b5871 · outbound

This paper cites Mixture-of-Experts with Expert Choice Routing.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Mixture-of-Experts with Expert Choice Routing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.395288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.395288Z digest=sha256:266d2710cb941b8bfce85f562364f2b7a7d85db874f4c23bacc494d6d78ff9ae

Observation b04c30a7-1cee-46c6-9058-e4f228c2d156 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.208055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.403226Z digest=sha256:4361e39ba555f028f33d13ffad98c7b0361767f74cdac6d5098337ba286c09f3

Observation fca64878-a51f-4498-84d0-b9e636da083e · outbound

This paper cites On the representation collapse of sparse mixture of experts,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap On the representation collapse of sparse mixture of experts,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.181226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.408738Z digest=sha256:f0b0db74b2e477cffb00258037d8266a21ee00e0125626c63d816c5ff06980c8

Observation a9b30a42-7534-45c3-a6b9-c85f4cd04e63 · outbound

This paper cites From Sparse to Soft Mixtures of Experts.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap From Sparse to Soft Mixtures of Experts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.415891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.415891Z digest=sha256:2def29f52df6e7e44fc45f5ae5b56d001db2822c3c7033cbdd775abb5234d9ab

Observation afecca68-c281-4afa-bcc9-02bfd9cc36f1 · outbound

This paper cites Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.151690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.423312Z digest=sha256:51d3e47d6a7a82f721ea93e409d2848151964fa7acaee4aa2c1d31d36f7f22c6

Observation c507ad06-8e9b-46b5-aff5-8393a810ed2e · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.132386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.431842Z digest=sha256:e6fee58905025c66462b9a81eb614fe47ab69107285d4205a98bf7bcc1df2c8d

Observation 4b445470-b14e-47b9-9cb9-7e3d78bb1883 · outbound

This paper cites Bagualu: targeting brain scale pretrained models with over 37 million cores,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Bagualu: targeting brain scale pretrained models with over 37 million cores,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.110429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.439613Z digest=sha256:fd9cabc926cd7f76cb17cf3bcb1d5f43a2f74ccccc4012cddd37cbe7061e9d64

Observation de2d75d7-950d-4dd5-9a92-4c387e5f0251 · outbound

This paper cites Faster- MoE: modeling and optimizing training of large-scale dynamic pre- trained models,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Faster- MoE: modeling and optimizing training of large-scale dynamic pre- trained models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.086717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.446137Z digest=sha256:ef8ca386b2a814bad25ed2fe2ee7166dd4f73ae2c21ee3f1b6f050b1638df2f0

Observation 0da1faa3-4a03-407e-873e-fc637cc5a4a1 · outbound

This paper cites PipeMoE: Accelerating mixture- of-experts through adaptive pipelining,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap PipeMoE: Accelerating mixture- of-experts through adaptive pipelining,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.060197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.455108Z digest=sha256:2df39da9774459eaa3e5648572b2cafefaa8e16da660185236937fb9e705577a

Observation 3be2d0c8-b444-4139-ac1e-7be7179b9241 · outbound

This paper cites SmartMoE: Efficiently training Sparsely-Activated models through combining offline and online parallelization,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap SmartMoE: Efficiently training Sparsely-Activated models through combining offline and online parallelization,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.034297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.462127Z digest=sha256:0de7d7b772f530cbf99b27d5630751f4623d9d60c7808483ccb6b08ab2a6d6a2

Observation f28ca483-ff35-4eb9-9518-2f6a846ce2b8 · outbound

This paper cites Janus: A unified distributed training framework for sparse mixture-of-experts models,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Janus: A unified distributed training framework for sparse mixture-of-experts models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:52.014066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.474427Z digest=sha256:dad5540da609babb8d45d8b8e54b76448cecf93b652325011f2d2bf688f5d4e1

Observation d700887f-d7c4-47c4-88e5-1275501fc589 · outbound

This paper cites Parm: Efficient training of large sparsely-activated models with dedicated schedules,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Parm: Efficient training of large sparsely-activated models with dedicated schedules,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.994739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.482427Z digest=sha256:1803bced3ec25bf6816d9331d9c390a2143b2d750c0f72d1182e58e51dedbee6

Observation e83c30be-0e0e-4733-b8b0-d1a15fbda6a4 · outbound

This paper cites Cen- tauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Cen- tauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.494649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.494649Z digest=sha256:f22eba8445bb81189d9dbc86346db4d9decab80b3a32b2cfdd86a97b4b56117b

Observation 1f78ca91-8299-40b5-a666-3600dde75cd9 · outbound

This paper cites Lancet: Accelerating mixture-of-experts training by overlapping weight gradient computation and all-to-all communication,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Lancet: Accelerating mixture-of-experts training by overlapping weight gradient computation and all-to-all communication,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.961310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.500807Z digest=sha256:b0862409cf49f31e971f1524144787e6e0593d2f9dd69e1c0f23e3c3af4a347d

Observation 90e812b6-c3eb-474b-b727-262729d7fd02 · outbound

This paper cites Sp- moe: Expediting mixture-of-experts training with optimized pipelining planning,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Sp- moe: Expediting mixture-of-experts training with optimized pipelining planning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.944284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.509667Z digest=sha256:0ba076d4c384ff6f5c4bbae1f5241bf8d9f54fb642eec83731929084fa96d7e2

Observation 2e5e7af2-e2d4-494e-bb9d-d8e6afcc504a · outbound

This paper cites Mitigating contention in stream multiprocessors for pipelined mixture of experts: An sm-aware scheduling approac,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Mitigating contention in stream multiprocessors for pipelined mixture of experts: An sm-aware scheduling approac,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.928112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.515751Z digest=sha256:ac54b328cfb021a85bd69029f7dcee046987a4fa499505f2b2806b3730e66869

Observation a5eef799-1339-48b1-9d7b-3df5b2c750bb · outbound

This paper cites Mast: Efficient training of mixture-of-experts transformers with task pipelining and ordering,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Mast: Efficient training of mixture-of-experts transformers with task pipelining and ordering,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.911072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.521028Z digest=sha256:32d9c9baf7341725f81c8bc4e237cc9c11f9ae1a2b6bf97e4f765fe6d1a7c02a

Observation cb13c37a-e101-4934-a63c-04708b943edf · outbound

This paper cites Scheinfer: Efficient inference of large language models with task scheduling on moderate gpus.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Scheinfer: Efficient inference of large language models with task scheduling on moderate gpus

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.893405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.525705Z digest=sha256:eda9c96a91fdf2629e39fbbda45c933c77d2ec06972bbb5900cf7e024b1b7799

Observation a549a193-b524-4e4a-ad17-7e8d658ebdc4 · outbound

This paper cites Data center tcp (dctcp),.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Data center tcp (dctcp),

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.874271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.531999Z digest=sha256:078e2a98959358a25bdf4e86e2f63d1f0cf3f3812005eecaa2753f07527225af

Observation ae1b1ddc-afa8-41cd-956e-31be9d4bd519 · outbound

This paper cites A scalable, commodity data center network architecture,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap A scalable, commodity data center network architecture,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.856828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.538271Z digest=sha256:83eea8bf5bad5891e99998a0c790049be019f8c1fdbb005a397882468dc2d033

Observation 469e70d0-d164-4a92-9ac2-e8330ed989a4 · outbound

This paper cites {TopoOpt}: Co-optimizing network topol- ogy and parallelization strategy for distributed training jobs,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap {TopoOpt}: Co-optimizing network topol- ogy and parallelization strategy for distributed training jobs,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.836229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.544173Z digest=sha256:a1c3354c41a396e35262e89c4439db534e693476270ccf53eb0a4a3743296f13

Observation 338d7657-7a2d-4e0e-b118-399e3b83cc26 · outbound

This paper cites Insights into deepseek-v3: Scaling challenges and reflections on hardware for ai architectures,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Insights into deepseek-v3: Scaling challenges and reflections on hardware for ai architectures,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.814898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.549688Z digest=sha256:845aaa90a234c4d0ba975e6f301dd7a3890e546b72ee9883fc554d46171156fd

Observation 0647cf00-4a4f-47e2-af54-efd24a285ebd · outbound

This paper cites Large scale distributed deep networks,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Large scale distributed deep networks,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.791727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.554736Z digest=sha256:8db36afb30770be7451e327b69f6b34fe4a29d170733b65ce1c5b95fa48eefc1

Observation 9ba23427-8128-4646-99b4-5e853ce4bb3b · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.559842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.559842Z digest=sha256:605d27cd616052a7543d9934c9eaa6342cbedb768b8c8a535ecfa48f9571185f

Observation 752da9a0-9f57-42cf-bd63-9ff8691ede32 · outbound

This paper cites Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:51.565480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:04:51.565480Z digest=sha256:9d594c323d6fb56b4a1932e6cf45b23fcefa35af3730e057f1d8d95a50020e8c

Observation fdaebd7e-20d1-46c9-b8bc-29853dd642da · outbound

This paper cites \ell 1, p-norm regularization: Error bounds and convergence rate analysis of first-order methods,.

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap \ell 1, p-norm regularization: Error bounds and convergence rate analysis of first-order methods,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:04:51.754885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T21:04:51.571500Z digest=sha256:2e95627a15446eaeb8a68a077dc9daa84ae52f75d590ffe79ec06ca4a107ba2d

Pith citing papers

Observation 3d3fa414-138c-4c95-845e-f8435ce927ce · inbound

ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training cites this paper.

ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.633632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T05:22:35.487173Z digest=sha256:c39e2641847680cc9f9b6d5d189de24696bb00c8933e33a8c8b473cc4664ecac

Observation 1b663c6f-23fb-4918-84fa-57d389092cb0 · inbound

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods cites this paper.

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:44:54.933618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T13:43:17.950000Z digest=sha256:8e57213d79e8836713096771d4bf577b0b5b2df9cac36fcf0b0324e5536ae6dc

Observation e1c042cc-f0e7-43ee-87af-608aeeacf1a7 · inbound

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods cites this paper.

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:07:44.237527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-11T01:06:46.426582Z digest=sha256:6279505a83f4c7346ec5c513941f98e80b0e4db2e8576b5dee8436cbb2e391e1