Pith. sign in

Paper Citation Record · LEDGER

MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2211.15841.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.15841 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:20:23.608575Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T15:24:49.924239Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 405b044b-1f8b-4e02-ab51-dfb4f02a968f · inbound

Mixtral of Experts cites this paper.

Mixtral of Experts MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:13:53.846793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-24T04:09:15.921778Z digest=sha256:69b63e7e3d4c10220d05b4188f8c13e6f913537d775597a6fd22027b2fd6cf8c

Observation 107736f1-d318-43c3-ba93-bbc2de294613 · inbound

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models cites this paper.

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:48:44.988800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T02:48:44.900467Z digest=sha256:bc717aadc6472bbc92e90903034b374d710a0587f5bf139ab8221729eb2314d2

Observation 99384180-b8a9-4ab8-a0b7-1a9605377599 · inbound

Training Sparse Mixture Of Experts Text Embedding Models cites this paper.

Training Sparse Mixture Of Experts Text Embedding Models MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T11:20:23.608575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:20:23.608575Z digest=sha256:3279ac6944df4bd26874451c8a098453b90bc034a07528d018a28cdef2346147

Observation 55dafae2-588d-4e18-b4ee-64b9a1a0224f · inbound

Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights cites this paper.

Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:58.768850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:58.768850Z digest=sha256:bff4ffb3a3e518dda490bb512222478e48d3b05b31ec8cf8a03a902a43f24405

Observation afe6b634-b2a1-4f08-ba14-063b67a85fe5 · inbound

Apple Intelligence Foundation Language Models: Tech Report 2025 cites this paper.

Apple Intelligence Foundation Language Models: Tech Report 2025 MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:26:58.583406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:26:58.583406Z digest=sha256:6d81aa920efc05d8a1dd537db0e0a31e43c76dec7c6976375f7acae2d55eb436

Observation 9eb5650d-b3eb-45be-924d-73e7de6ca05d · inbound

Maximum Score Routing For Mixture-of-Experts cites this paper.

Maximum Score Routing For Mixture-of-Experts MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T19:23:15.344590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T19:23:15.344590Z digest=sha256:ba5c8123f9ecf3dac5b4e68a6d8d0e3f9178f9046434c231831774cc02f9f8c6

Observation 479a75b0-ca0a-40cb-aac4-948b76d48d76 · inbound

When Does Sparsity Mitigate the Curse of Depth in LLMs cites this paper.

When Does Sparsity Mitigate the Curse of Depth in LLMs MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T20:29:33.439034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:29:33.439034Z digest=sha256:c7e4ab1f81b0a2f495831ceb07c496c154d8fecf7754b04fcf26497ec8cb8b1a

Observation c71d9f5f-e52f-4b4b-924c-2810a0ffde9c · inbound

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns cites this paper.

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:10.790930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T08:29:17.149710Z digest=sha256:b83d4dc274cf593f875f9ba27a71598eddd8026c0620cad2e4d9afc813eb8f43

Observation 52548489-bca6-4a80-a3ec-91609d656bed · inbound

TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training cites this paper.

TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:06:20.725388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T01:36:41.804171Z digest=sha256:e1e0fad90fb4adcdc577721ba1a30d608219e43acd44672ab8cf1164283e3534

Observation 4171cd97-1054-4062-be7b-654dd935ea7f · inbound

RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts cites this paper.

RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 10

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T23:36:31.194142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T16:39:35.128879Z digest=sha256:6a05f7b3c11a66feacdb357a3776d4027b700cfb1ddb03134477cee3d97441b1

Observation 9575abc1-549e-4c21-b324-4d33f24d60a2 · inbound

Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference cites this paper.

Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:11:24.699038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:30:58.729425Z digest=sha256:ec3e43d34106af6bcdd637ef2950d6c32c0981862a10f7c6c0ad077fd7e8028b

Observation 1235cce1-ca82-4a06-910f-d8a7456eb1cb · inbound

Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts cites this paper.

Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:24:49.925986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T15:21:34.127680Z digest=sha256:dac20d77ea5e17a3f2a3291ef483e1b26b98b34c2dcbbc6e4732c95661cd5fb2

Observation ba00ad91-0f43-49e1-944c-510d54818d01 · inbound

Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference cites this paper.

Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T08:35:22.347459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:35:22.347459Z digest=sha256:809be4c410b61c20bd328e463f9e852de029296666648ada7f2472e33d31bbff

Observation fb0b9cf3-8307-4248-9806-c49d1a1b584e · inbound

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix cites this paper.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.115161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.115161Z digest=sha256:6e550ede12429a2a598a6c5ce0d3b4b701be483d7f6ed27f8505e315eda159fe

Observation ce1b3d73-d720-4578-998b-fb5c930107d3 · inbound

Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study cites this paper.

Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T00:15:48.390607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:15:48.390607Z digest=sha256:e5ddf965ee560045693f297329a4eaa7537ea80f56188fda19eb8c74fb2b9265

Observation abfb2628-53e8-4d81-a6a7-fcac483bb81d · inbound

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts cites this paper.

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T00:38:32.637398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:38:32.637398Z digest=sha256:ff0bae6378c27cd5cf1fb12b35fbe6daf5e10c511950e0ebc69027ede5ba92a3