Pith. sign in

Paper Citation Record · LEDGER

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

As of 20 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 16 inbound Pith citation observations for arXiv:2501.11873.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.11873 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:52:40.058197Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:52:49.484653Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:29:38.231668Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4bf2f323-164a-497e-9800-7b593a2331db · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models Training Verifiers to Solve Math Word Problems

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:39.991278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:39.991278Z digest=sha256:c1e4a73880322649928ab7fca77a4cf28e6861243d6ea466df72d3fb41ea3969

Observation c7ab2795-4ea1-413b-b981-c340372fe0e7 · outbound

This paper cites Mixtral of Experts.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models Mixtral of Experts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:40.009414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:40.009414Z digest=sha256:318e03f3a257b883b899221ff761e9b767a0b9542cd4d6bf33564fdc9047119a

Observation eaa8bfc5-f426-4971-ba7e-2167ff294735 · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:40.014024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:40.014024Z digest=sha256:ad44dd9151487e419913fd8c58cb8560780e5dd9f1a4b0c53ce8b5309cff0b5e

Observation 1c32c31c-3782-4a90-a4ef-5015b9945b92 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:40.019023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:40.019023Z digest=sha256:d0b2508dbc4903e539b0d15593d7985faa587f3f8fcb52b9919948e829812596

Observation 8c2629c3-2d37-482a-86b1-05d609fd0e2a · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:40.033283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:40.033283Z digest=sha256:18f725b861c94d68db8a7144f893965c572e85021ffe5d7f48652c839af93d22

Observation 59a592fc-3e1a-4abb-ad23-b8907b8352fa · outbound

This paper cites Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:40.037472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:40.037472Z digest=sha256:15b24e93bf41dc54cdddaa9661ddb7222103341fef78ec3b6c5b609b68dd01c7

Observation 18ced0b9-e159-4c30-bb57-738b72c1a44e · outbound

This paper cites GW-MoE: Resolving Uncertainty in MoE Router with Global Workspace Theory.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models GW-MoE: Resolving Uncertainty in MoE Router with Global Workspace Theory

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-10T17:52:40.143268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T17:52:40.041845Z digest=sha256:df8cb051b21376bc381a5c4a28556976c6f29a3a687855bcba12156e6566e6b6

Observation 17d70408-3f98-41f6-9a8b-20090de6b482 · outbound

This paper cites OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:40.045727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:40.045727Z digest=sha256:3b159dd204cbcfbeb62dbf7754331e33d2a91327505d117ed4a425705043b535

Observation ed9380e0-389a-4e8d-abcd-2e3b55c1fd38 · outbound

This paper cites Qwen2.5 Technical Report.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models Qwen2.5 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:40.049755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:40.049755Z digest=sha256:f5d5bffd71d14d8f620b1aefedcf15c04e9f58b2eb89bdb224278725febc4f13

Observation ee0167fc-5425-4966-a27d-f180b7c770e9 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:40.053393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:40.053393Z digest=sha256:88af1da1d2b6ae74012733e1fd3a5d230167a6a0f46eecdf428227d127b68736

Observation 2f7b7390-f2bc-45ac-ad92-acc719afa316 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:40.058197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:40.058197Z digest=sha256:90a418868c1b513ff86396456462e7861437d8b45f795a776814dab633043397

Observation 3357eeef-6a9c-432c-95d6-ee2732a49980 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:40.028539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:40.028539Z digest=sha256:f2b31c3cf4f650c7d26c1315ac1ce1eecc107bcd835269b4cb7d5585c6520eb1

Observation bf8ebc0e-3107-416e-b1d5-dd880f4ad923 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:39.995962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:39.995962Z digest=sha256:31a86ed56ef5ab5626fa957d3e9bc1833a35500f960556619840d9e0eebf386a

Observation 095965d9-f4fb-4162-acfa-aa2f23bc5cf5 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:40.024022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:40.024022Z digest=sha256:a4f73187d92cc5eb04af733f2c5773d0c7bcb34c88e4c07aa1b4a1d47b6a073b

Observation 6ebe77be-79c4-4605-8b17-34f279130158 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models Measuring Massive Multitask Language Understanding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:40.004698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:40.004698Z digest=sha256:2b28bfe60913b05fe8500abf4b031ff030334f3f89afd8548e5c6b5bdc1cc6ba

Observation d594c9a4-4720-4ecb-bede-1cd6ac1b492a · outbound

This paper cites Fewer Truncations Improve Language Modeling.

Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models Fewer Truncations Improve Language Modeling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T17:52:40.000179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:52:40.000179Z digest=sha256:047ca6432a8df95bec1eb1694381ad4112e0f9887b18321b3e4479062a54229e

Pith citing papers

Observation 1fb44ae8-7558-427a-b060-8dc3ce7f5184 · inbound

Neural network task specialization via domain constraining cites this paper.

Neural network task specialization via domain constraining Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:49.484653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:49.484653Z digest=sha256:2cd4e3d239984f773c994bbd7333fdf342cbdfe60d8d4c3669d0e85b6f6d7ee8

Observation 654efa76-a041-45e8-8d27-b452aba0b77b · inbound

Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs cites this paper.

Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T23:31:19.556401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:31:19.556401Z digest=sha256:2644d0eef26fa4b931a91f485851972908df18047ffac18722940a0edca57030

Observation e8f5c733-853d-4c28-a2db-c92d08ce8e6e · inbound

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free cites this paper.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:04:34.915567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:b633d2d12fea023bac38cfe32a963a5186d871c739aa6dd54bcae574fba3cd42

Observation 3fb5fb12-d4c3-4807-a719-3be2e722a6dd · inbound

Qwen3 Technical Report cites this paper.

Qwen3 Technical Report Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:28.577251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T06:35:27.813995Z digest=sha256:83e5920200127dac1315a0efb433653bb4a76f509934c36774f95611c9df1b40

Observation a55a3517-79b9-4275-87dd-907dd808aa42 · inbound

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications cites this paper.

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T17:26:28.891630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:26:28.891630Z digest=sha256:277edfcabdab1a1964730434393e979867ba7e123a7193c4bcb9fc5959cbeccf

Observation 37873acc-e3d7-48bd-8a28-adbbf0a517ba · inbound

ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution cites this paper.

ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 246

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:58:59.050128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T13:58:58.627748Z digest=sha256:6ac4caa591e59eb1e8b99ffd70aa6a1e7d87e62b659f126e6bf55b255a699bd8

Observation c0f700ef-2573-4db8-80a5-f9fbd5a97679 · inbound

ProPhy: Progressive Physical Alignment for Dynamic World Simulation cites this paper.

ProPhy: Progressive Physical Alignment for Dynamic World Simulation Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:08:47.973460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T01:05:08.087136Z digest=sha256:8a092f1d32967c7da94d47cc0d86c985e2e2ae6f151e6c665ec818ab4ff51117

Observation f4f170e1-5c34-4fc9-82d0-295786056337 · inbound

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training cites this paper.

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:28.419367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T03:34:10.370956Z digest=sha256:a2bee7ccc5578343e2e2b316b17202200921b7592b9fd9ef99bb34d59d90571e

Observation a4d7779a-5cac-47e0-9902-a895b78ff088 · inbound

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training cites this paper.

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:23:51.260199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T23:22:51.808346Z digest=sha256:477425dd94ada21ccd36867711117bd47732cdfc72225e352a8355d4a423daed

Observation 099bf210-82e4-4240-8e4b-af340a8dc30a · inbound

UB-SMoE: Universally Balanced Sparse Mixture-of-Experts for Resource-adaptive Federated Fine-tuning of Foundation Models cites this paper.

UB-SMoE: Universally Balanced Sparse Mixture-of-Experts for Resource-adaptive Federated Fine-tuning of Foundation Models Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:08:54.313335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T19:06:08.951475Z digest=sha256:81bf96269e227ac2f15cb45be9f1bd6349e23654d740287d6a50fbdd232a7093

Observation 3e42793d-2843-4518-a6e6-d98f4fe431bc · inbound

PithTrain: A Compact and Agent-Native MoE Training System cites this paper.

PithTrain: A Compact and Agent-Native MoE Training System Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.524678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T23:12:14.193089Z digest=sha256:24f18df0d930edbd1ece71b5f3a0a6fb8d71fe5b33e874bda2ddcf2b3b8ca83a

Observation 58a1a656-1b11-4128-b2d9-0c3fb1c34225 · inbound

DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts cites this paper.

DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:16:14.421840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T17:14:53.648013Z digest=sha256:ee63e2e134d4b6bc434c606e9155da1c1e9011470921b1d23afa4d3b6a327bc8

Observation 8e376c96-e784-4433-b402-e7650c3cbfc6 · inbound

STAR: Rethinking MoE Routing as Structure-Aware Subspace Learning cites this paper.

STAR: Rethinking MoE Routing as Structure-Aware Subspace Learning Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:07:26.934730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T18:28:35.162934Z digest=sha256:6378a982a004d92be0b3fdbc7c437f23f71276e1f21e62de20e466ffb22d2c73

Observation 0c5e3505-ab35-4928-9905-18b88a6de820 · inbound

Sakana Fugu Technical Report cites this paper.

Sakana Fugu Technical Report Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 207

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:29:38.233271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T14:22:37.596720Z digest=sha256:5c6845921052f8632ee65b70528dae847f30eb9c80407f44cf68926fef294b4b

Observation 834c0362-1b70-4fbe-9cf8-2149a4c4dd79 · inbound

Fisher-Routed Mixture of Experts for Federated Class-Incremental Learning cites this paper.

Fisher-Routed Mixture of Experts for Federated Class-Incremental Learning Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:04:39.217496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T10:19:16.463961Z digest=sha256:6a6205945b0053e769f2875607327c2642d76cf317025fa58020fd4966a92591

Observation a6bb9d1a-5b2b-484e-952f-d80ec9951353 · inbound

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts cites this paper.

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T00:38:33.233904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:38:33.233904Z digest=sha256:09b84e2ddf86a0055946c18fb74e63661eca639cf0de083d880d97cca9ee029a