Pith. sign in

Paper Citation Record · LEDGER

Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2404.05567.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.05567 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:13.192754Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T19:08:50.558628Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 17cd2fc4-25f7-46dc-b552-0baf645a894d · inbound

Monet: Mixture of Monosemantic Experts for Transformers cites this paper.

Monet: Mixture of Monosemantic Experts for Transformers Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:49:17.400132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:49:17.400132Z digest=sha256:ecad15433a66e3a017d20ff2b5e9cfa96e1ddb9011994c246a6b0276e7639d5e

Observation 731951e7-5796-4419-93dc-ecec60a589d5 · inbound

A Survey on Inference Optimization Techniques for Mixture of Experts Models cites this paper.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 130

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.848329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.848329Z digest=sha256:7b9b1d2d6e68f13c1489bf665d6139fa45e2a2c976e93aa357466165b15788c5

Observation b8c98053-70df-430a-ae69-2e0d3592da0e · inbound

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch cites this paper.

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:32.427965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:32.427965Z digest=sha256:4fac8a52544e7e9dc84bc18fcab4e86f7c785bbc6e783e5394bc9a616d1b0855

Observation 427b3fff-99d2-4f59-8d94-36d76c0383d9 · inbound

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging cites this paper.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.509397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.509397Z digest=sha256:9bb877250ad351aff82e0418499b5a38b54e704cec3fc282f8879d98ed0d4e22

Observation baacc1a6-b538-4d4e-bf01-7025b4327245 · inbound

Dynamic Chain-of-Thought: Towards Adaptive Deep Reasoning cites this paper.

Dynamic Chain-of-Thought: Towards Adaptive Deep Reasoning Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T20:51:16.713812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:51:16.713812Z digest=sha256:eb724a81c2a1af05b5c18d958fd97330a26e704beede63ee0b8dd615e8c849e9

Observation c3f672a2-a990-4bce-a299-880423aa4676 · inbound

EfficientLLM: Efficiency in Large Language Models cites this paper.

EfficientLLM: Efficiency in Large Language Models Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 155

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:36.622539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:36.622539Z digest=sha256:0646a8c9b774a29376e22f7596051cd99896166ea4af45cc1c7aa85763e4cf79

Observation 129e737e-87b6-446a-90f2-d248528f26ab · inbound

Industrial brain: a human-like autonomous neuro-symbolic cognitive decision-making system cites this paper.

Industrial brain: a human-like autonomous neuro-symbolic cognitive decision-making system Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:31:50.478054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:31:50.478054Z digest=sha256:ff7f5a35ecfb8d316156f234e6dcc0b7e07abdf34883418f0326ec1e5c5e962d

Observation 45314fb4-18bb-41b1-927d-d722a3d2ef11 · inbound

Universal Pansharpening Model cites this paper.

Universal Pansharpening Model Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T19:05:48.549327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:05:48.549327Z digest=sha256:aa357a4ea7a11f12e379e0302e760f65a7b610a5d4aecb0d1caaf0b3dc5f176a

Observation 307bd995-80ed-4b41-9a09-f9327d153578 · inbound

Does a Global Perspective Help Prune Sparse MoEs Elegantly? cites this paper.

Does a Global Perspective Help Prune Sparse MoEs Elegantly? Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:50.751478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T19:06:25.626026Z digest=sha256:05a16a6ebc0d92f941e9f6bfa35a93c8b417d57408f5c3df5743abd6d0736c37

Observation 598b2a3e-c1b4-4a67-9eec-3ebfbaa60759 · inbound

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning cites this paper.

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 152

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:01:10.646889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T17:19:59.247074Z digest=sha256:80d29ca0571ec66363ac0f37e2f685c640096479ba173e0a9850992afcd9a5b3

Observation 9939682b-2787-4f99-a1e4-0f849141381c · inbound

Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts cites this paper.

Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.855594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T05:20:32.087436Z digest=sha256:ad7d73731d3e5f960655d8df0e3bbab8e59a429b2dc96e1f8b1483368cce927d

Observation 764cd259-ac5d-41df-b6f7-18966e1e1d4f · inbound

SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs cites this paper.

SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:08:50.560200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T01:58:01.787208Z digest=sha256:6be5d0bd3e97380030edaf959fe3776ebd175ad5d8ea8d432a56dc22834245f5

Observation f5c7008d-2b31-435d-bf08-e1f8ea2e907f · inbound

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining cites this paper.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.192754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.192754Z digest=sha256:0ee50be0d08ea5f490a9c4a583d8aca1970d14c5ec8d5b794480c2ecd77bb02a