Pith. sign in

Paper Citation Record · LEDGER

MH-MoE: Multi-Head Mixture-of-Experts

As of 14 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2411.16205.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16205 v3

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:36:22.161515Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T07:09:48.239662Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T07:11:53.241949Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f772b9f-db1b-4702-acc0-211d23f566d6 · outbound

This paper cites Unified Scaling Laws for Routed Language Models.

MH-MoE: Multi-Head Mixture-of-Experts Unified Scaling Laws for Routed Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.060160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.060160Z digest=sha256:78c777cffa31d7ed93f2973a0eb3c2b7a3221a462212f86736f41a4312dcc890

Observation 55be5405-1829-435d-ab92-8b65b5b0a414 · outbound

This paper cites On the representation collapse of sparse mixture of experts.

MH-MoE: Multi-Head Mixture-of-Experts On the representation collapse of sparse mixture of experts

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:36:22.538281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T13:36:22.066880Z digest=sha256:f5250404241b52649809700cab6d9cab8cc60952a2d8efa2eb6ebecf2af19dc3

Observation dbaff43e-891c-4416-8036-fd6580c6873e · outbound

This paper cites Redpajama: An open source recipe to reproduce llama training dataset, 2023.

MH-MoE: Multi-Head Mixture-of-Experts Redpajama: An open source recipe to reproduce llama training dataset, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.074453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.074453Z digest=sha256:fbd1c82e5ef2a6c8458416d291a00dbdb3ff83ef1191b5d909317196c66d17ff

Observation 92bc78df-b855-43ef-9d3a-6da1f4159854 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

MH-MoE: Multi-Head Mixture-of-Experts DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.080452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.080452Z digest=sha256:8cb364ef0752dccf39014dadbad7ed78466985ea0889a3d5171af2aecd46c1ff

Observation 3512a424-659b-4d05-8c2c-aeccb0abaf4f · outbound

This paper cites GLaM: Efficient Scaling of Language Models with Mixture-of-Experts.

MH-MoE: Multi-Head Mixture-of-Experts GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.087205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.087205Z digest=sha256:d75c7f30cb33b376e5419958d24b996f5a13887b1eb834ffbcf6b4252c93da25

Observation e40960bd-62b0-4b0e-8396-f6a4bb71df0f · outbound

This paper cites Mixtral of Experts.

MH-MoE: Multi-Head Mixture-of-Experts Mixtral of Experts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.094853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.094853Z digest=sha256:c1aabddcf2d8e6ac8c3ca9057c41d49b24b830f012d8878a0bb66206245c041d

Observation c5d00a8c-1b1a-463e-aefc-eada142594db · outbound

This paper cites Building a great multi-lingual teacher with sparsely-gated mixture of experts for speech recognition.

MH-MoE: Multi-Head Mixture-of-Experts Building a great multi-lingual teacher with sparsely-gated mixture of experts for speech recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.102485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.102485Z digest=sha256:80df6c7d255c980ec9f6a350de1e2768112a1408549e2d6206f6c805f70d3901

Observation ace33193-78a4-4200-afb9-43bb5dc90c2a · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

MH-MoE: Multi-Head Mixture-of-Experts GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.110704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.110704Z digest=sha256:ad59a493bdc11917916cde3d0d8da17a9785a39cdb46475ab467a96b0576eaf1

Observation db7bfbcd-d3b1-4bb7-a5b9-42a4fd43a481 · outbound

This paper cites The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits.

MH-MoE: Multi-Head Mixture-of-Experts The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.116961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.116961Z digest=sha256:dfa5452de4dd32df3b35a0837cbdf155d242f8f53fc56678645615ce21cbc816

Observation 008f5eab-48bc-4fc2-a7bc-761799eff79c · outbound

This paper cites Task-Based MoE for Multitask Multilingual Machine Translation.

MH-MoE: Multi-Head Mixture-of-Experts Task-Based MoE for Multitask Multilingual Machine Translation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:36:22.214087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T13:36:22.122709Z digest=sha256:539a6433fbcfb522354544fc14e686bb30dfc65ea911780e9a44be458bc2d441

Observation 317e6bd1-2d98-465c-a436-3c3c3e857a66 · outbound

This paper cites Improving language understanding by generative pre-training.

MH-MoE: Multi-Head Mixture-of-Experts Improving language understanding by generative pre-training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.128726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.128726Z digest=sha256:45c9aaffd40f49c1132845cb39e7cd1def7f14688db9e700e62a9083d31e04a7

Observation f2fe8c18-f4c9-4840-b3c8-5dd8a333a288 · outbound

This paper cites Language models are unsupervised multitask learners.

MH-MoE: Multi-Head Mixture-of-Experts Language models are unsupervised multitask learners

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:36:22.487038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T13:36:22.135594Z digest=sha256:af0fd80ec7c7883611fdd1436c11deb1123dad371355ac91346c3245883a24f5

Observation 81c405ce-8d53-4abb-b2d8-484bd882e0a7 · outbound

This paper cites Glu variants improve transformer, 2020.

MH-MoE: Multi-Head Mixture-of-Experts Glu variants improve transformer, 2020

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.140800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.140800Z digest=sha256:74c0208482d5d4481cbec6b899c66155b92698d3557ff223a11da315b8991f8b

Observation 6ead7d61-f527-4df4-b62a-9c54c92cab1b · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

MH-MoE: Multi-Head Mixture-of-Experts Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:36:22.454163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T13:36:22.145987Z digest=sha256:637bded6cc260b5aae63ffed369f8c3ab5b25202fd0d321d2108fef6998b408e

Observation 21a5d4f9-0246-4e28-b23e-fda064011781 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

MH-MoE: Multi-Head Mixture-of-Experts Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:22.150888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:22.150888Z digest=sha256:9b14b5c0d2c277d4e9a7b1780b0c8a6a709f2c6def96cdc8134e1735163fb106

Observation 27c9b5cb-6aff-4ab6-a9ce-a943f1d78f36 · outbound

This paper cites Multi-head mixture-of-experts, 2024.

MH-MoE: Multi-Head Mixture-of-Experts Multi-head mixture-of-experts, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:36:22.423691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T13:36:22.156005Z digest=sha256:8ad9a60ef914741643086bb584ea5a5d3acf24bf70875cad6445396a5724b59c

Observation cf163fbb-fc4b-4516-949e-008540e77a9e · outbound

This paper cites Sparse moe with language guided routing for multilingual machine translation.

MH-MoE: Multi-Head Mixture-of-Experts Sparse moe with language guided routing for multilingual machine translation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:36:22.404450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T13:36:22.161515Z digest=sha256:d194911454c991d17a65e7e6e8df22e7031926c343375500e6415d001633d1de

Pith citing papers

Observation 3a1c93c9-6a87-46ad-9b8d-ebeba8ea395b · inbound

CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering cites this paper.

CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering MH-MoE: Multi-Head Mixture-of-Experts

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.243173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T07:09:48.239662Z digest=sha256:3ecce8b777a3ffdf36d3620d7e3f36f0fb6fe7f4b4ce88c53e600b33885a84eb