Pith. sign in

Paper Citation Record · LEDGER

Efficient Large Scale Language Modeling with Mixtures of Experts

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2112.10684.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.10684 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:53:28.504969Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:36:16.721505Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 32a6ba29-5386-4c8e-a5f7-0db809e90907 · inbound

Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? cites this paper.

Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:51:46.877478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T09:51:46.701149Z digest=sha256:109700c5cbcbaf02bdc92d3231a1947466f0682d26bad80a17f150393e8b9864

Observation f55a99eb-5a16-4349-8d7c-42307fc0760b · inbound

InCoder: A Generative Model for Code Infilling and Synthesis cites this paper.

InCoder: A Generative Model for Code Infilling and Synthesis Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:21:20.464819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T02:21:20.438666Z digest=sha256:4b183fe7b22df92c8eb546d4025125df0ee8cd77ce0716a050748c9fb8599e83

Observation 87fc9f96-7a52-45ad-a387-688048f16105 · inbound

GPT-NeoX-20B: An Open-Source Autoregressive Language Model cites this paper.

GPT-NeoX-20B: An Open-Source Autoregressive Language Model Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:34:28.200651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-24T12:33:37.701655Z digest=sha256:9701262e3a05bbd030c655dea65d6ecab6f236eae33d458e286983aefad7d1c0

Observation 1cb42d52-0636-42b9-97ed-e67f246027a3 · inbound

OPT: Open Pre-trained Transformer Language Models cites this paper.

OPT: Open Pre-trained Transformer Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 287

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T20:53:17.887823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T20:53:16.720145Z digest=sha256:5cb3bbdb5cc323639074af5549e97cfee346d902bc12a6d8f5386792a274d805

Observation f487086e-f9af-4a06-a760-1fba15f9477d · inbound

Emergent Abilities of Large Language Models cites this paper.

Emergent Abilities of Large Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:38:38.018560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T07:38:37.734402Z digest=sha256:2f1e8760089b2b549d2cc69c8d400e16173a31603a29d4d1c90a274c7f4726a9

Observation 33f09b31-2c0a-4d40-8320-75c228c7512c · inbound

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale cites this paper.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T13:35:36.034981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T13:35:35.972596Z digest=sha256:791f8a4229d4950c6b4b22fab63f51886d65e07f0e164b1e2c166172f0235d45

Observation 4cfcdc35-95fd-4e6d-9bea-35bfe4e126f3 · inbound

eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers cites this paper.

eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:44:22.789288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T01:44:22.710205Z digest=sha256:08d440bd0cafd979f9f551bdb73bfbd92ca670b64f1c06e033fe051550971de0

Observation 9582cd58-99b0-4787-bf69-af2e15d3b682 · inbound

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only cites this paper.

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:43:45.814699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T20:43:45.770157Z digest=sha256:805c7a4f563c0ed55e0c974c51e145039ed0ae301434143e3b3088b979018b23

Observation de023224-8b76-4114-9561-e053f1dcdced · inbound

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models cites this paper.

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 170

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:07:22.300478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:07:22.166595Z digest=sha256:29ff4805e048e65263e5d39667bbb87e71ff853a7a244bf9df7356928f7877dc

Observation 4209e00a-e48b-4842-8231-8b05048a10d6 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 242

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:46:09.982830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:12d3749488995d976aafff7d157f4f690bd018489c4d4c11e169032d18426e00

Observation 78aa2a76-a5c6-4126-907b-2a602040a306 · inbound

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference cites this paper.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:28.504969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:28.504969Z digest=sha256:d6b512b15eaed1886955010910faa14ce6ac2a367a45bb3d3360f9b5a0328fbe

Observation f403ef49-7024-4c19-a63d-0f28c93c1b59 · inbound

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference cites this paper.

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:51:42.224496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T17:47:58.019030Z digest=sha256:5f7d1a8d0956d2bea957edb4390cb322e932618cab83648fdce537d27b59dce5

Observation af5eecf4-66be-463d-a200-6a4e94ef4ba4 · inbound

Tracing the ongoing emergence of human-like reasoning in Large Language Models cites this paper.

Tracing the ongoing emergence of human-like reasoning in Large Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:04:37.198324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T05:04:30.829386Z digest=sha256:77379c3434e31ea42e75a206173459dc1968ff9a45f5ec34c2bf776abacc7221

Observation 54776406-cfa2-4997-b3f1-36e6e72426e5 · inbound

Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs cites this paper.

Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.034915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T13:22:40.017922Z digest=sha256:dd621a4f2230a067be643008418c263db3278cb88eea53df87de2c4963f32a6f

Observation 072673fc-84d4-4fd9-8b66-1fd0f9e22efa · inbound

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models cites this paper.

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:16.722971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T15:19:26.983760Z digest=sha256:c103d5ef85d5ac5f3a605b84d896b8edd2dabc52d6af5c13cfca955a26d418bc