Pith. sign in

Paper Citation Record · LEDGER

MH-MoE: Multi-Head Mixture-of-Experts

As of 13 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2411.16205.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16205 v3

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:36:22.161515Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T07:09:48.239662Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T07:11:53.241949Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f772b9f-db1b-4702-acc0-211d23f566d6 · outbound

This paper cites Unified Scaling Laws for Routed Language Models.

MH-MoE: Multi-Head Mixture-of-Experts Unified Scaling Laws for Routed Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.060160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.060160Z digest=sha256:de7cce0714130c1da8c151daf234fc91aa4a77430f34fac20ad277de01d634bd

Observation 55be5405-1829-435d-ab92-8b65b5b0a414 · outbound

This paper cites On the representation collapse of sparse mixture of experts.

MH-MoE: Multi-Head Mixture-of-Experts On the representation collapse of sparse mixture of experts

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:36:22.538281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T13:36:22.066880Z digest=sha256:47236e36ea009e8b9db883a893424184bf475cc184b8c3be38405c685bc2161b

Observation dbaff43e-891c-4416-8036-fd6580c6873e · outbound

This paper cites Redpajama: An open source recipe to reproduce llama training dataset, 2023.

MH-MoE: Multi-Head Mixture-of-Experts Redpajama: An open source recipe to reproduce llama training dataset, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.074453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.074453Z digest=sha256:d7392505352a944ec3acdc0d0ea6a9f9a9bfa5dcabe02dec5f54f3082af448e9

Observation 92bc78df-b855-43ef-9d3a-6da1f4159854 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

MH-MoE: Multi-Head Mixture-of-Experts DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.080452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.080452Z digest=sha256:8a0f5f6b86775d6808e8fb79049bbf402abbe6584ee503ebe4e81869487c8c5a

Observation 3512a424-659b-4d05-8c2c-aeccb0abaf4f · outbound

This paper cites GLaM: Efficient Scaling of Language Models with Mixture-of-Experts.

MH-MoE: Multi-Head Mixture-of-Experts GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.087205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.087205Z digest=sha256:8e7ee3663cfcdc8b89c41eca1baf5bc9ef0b0248b65ee699cd5e03907aa62c1e

Observation e40960bd-62b0-4b0e-8396-f6a4bb71df0f · outbound

This paper cites Mixtral of Experts.

MH-MoE: Multi-Head Mixture-of-Experts Mixtral of Experts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.094853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.094853Z digest=sha256:85a76cb03de3546bd3cf28b084867137892572c4bb3556f5667a6752b46cfff2

Observation c5d00a8c-1b1a-463e-aefc-eada142594db · outbound

This paper cites Building a great multi-lingual teacher with sparsely-gated mixture of experts for speech recognition.

MH-MoE: Multi-Head Mixture-of-Experts Building a great multi-lingual teacher with sparsely-gated mixture of experts for speech recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.102485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.102485Z digest=sha256:9f42348960bd46c34bc6f1ef86f4aafa1b26a2c7666e2a6042bc700ef02d6472

Observation ace33193-78a4-4200-afb9-43bb5dc90c2a · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

MH-MoE: Multi-Head Mixture-of-Experts GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.110704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.110704Z digest=sha256:4f91e7f6f2d3c6f871c25dabbedca21f0d7240b7da1af5dfca0b7c5b4c6ed1b8

Observation db7bfbcd-d3b1-4bb7-a5b9-42a4fd43a481 · outbound

This paper cites The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits.

MH-MoE: Multi-Head Mixture-of-Experts The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.116961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.116961Z digest=sha256:7e665fe30168bfe1dca98ecaed71e9314c39b7428b65a806f59f059f3d1009f4

Observation 008f5eab-48bc-4fc2-a7bc-761799eff79c · outbound

This paper cites Task-Based MoE for Multitask Multilingual Machine Translation.

MH-MoE: Multi-Head Mixture-of-Experts Task-Based MoE for Multitask Multilingual Machine Translation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:36:22.214087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T13:36:22.122709Z digest=sha256:5f134c11e2357526ad1f8f1d897069d3dca29250e9d79028f04d56edf6dcd0d3

Observation 317e6bd1-2d98-465c-a436-3c3c3e857a66 · outbound

This paper cites Improving language understanding by generative pre-training.

MH-MoE: Multi-Head Mixture-of-Experts Improving language understanding by generative pre-training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.128726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.128726Z digest=sha256:e61d3cf8e5796fecebb5e7916403d1441ced25b08b8520cadd3a120affa22ed6

Observation f2fe8c18-f4c9-4840-b3c8-5dd8a333a288 · outbound

This paper cites Language models are unsupervised multitask learners.

MH-MoE: Multi-Head Mixture-of-Experts Language models are unsupervised multitask learners

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:36:22.487038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T13:36:22.135594Z digest=sha256:d1a6b07314d937cfe02fbdbec9ad2ee0e564a7d9046c4416b7866c220c24afac

Observation 81c405ce-8d53-4abb-b2d8-484bd882e0a7 · outbound

This paper cites Glu variants improve transformer, 2020.

MH-MoE: Multi-Head Mixture-of-Experts Glu variants improve transformer, 2020

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.140800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.140800Z digest=sha256:a3e081edc8e3431b8221df5d09ab7ea1ef2d9083ce6b7b79b5c4cbd38b648098

Observation 6ead7d61-f527-4df4-b62a-9c54c92cab1b · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

MH-MoE: Multi-Head Mixture-of-Experts Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:36:22.454163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T13:36:22.145987Z digest=sha256:0fe1be34f93040b72c6fc8067f7652f02533d0493604beeadfd29d02418bc939

Observation 21a5d4f9-0246-4e28-b23e-fda064011781 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

MH-MoE: Multi-Head Mixture-of-Experts Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.150888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.150888Z digest=sha256:c2eabfc06dd2307911170de9822a4f58e036218ecb73f079f1c1dd14d5ea93b4

Observation 27c9b5cb-6aff-4ab6-a9ce-a943f1d78f36 · outbound

This paper cites Multi-head mixture-of-experts, 2024.

MH-MoE: Multi-Head Mixture-of-Experts Multi-head mixture-of-experts, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:36:22.423691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T13:36:22.156005Z digest=sha256:d4f706cc7bab8f3f0df335fc94ac73e8c9377cf3108ef4aded0e6d314daa9606

Observation cf163fbb-fc4b-4516-949e-008540e77a9e · outbound

This paper cites Sparse moe with language guided routing for multilingual machine translation.

MH-MoE: Multi-Head Mixture-of-Experts Sparse moe with language guided routing for multilingual machine translation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:36:22.404450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T13:36:22.161515Z digest=sha256:cc1eda099fa35d7b9a33a27a2029518038f093d561e2bfcdbee6b47f53e0463a

Pith citing papers

Observation 3a1c93c9-6a87-46ad-9b8d-ebeba8ea395b · inbound

CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering cites this paper.

CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering MH-MoE: Multi-Head Mixture-of-Experts

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.243173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T07:09:48.239662Z digest=sha256:c3c6e72eddf2268b585ec46674838df5d497d1e52904acd66bb86960f66188f9