Pith. sign in

Paper Citation Record · LEDGER

BASE Layers: Simplifying Training of Large, Sparse Models

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2103.16716.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2103.16716 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:29:49.048895Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T10:03:17.812816Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 422d557e-f2d9-4302-9520-2311b527c97e · inbound

ST-MoE: Designing Stable and Transferable Sparse Expert Models cites this paper.

ST-MoE: Designing Stable and Transferable Sparse Expert Models BASE Layers: Simplifying Training of Large, Sparse Models

Reference 170

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:14:25.641570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T23:14:25.431471Z digest=sha256:fff8c89befaeae1c0aabde90e904016e1ecf53d48182744da708d78442fd7219

Observation 016edb7f-bd71-4238-9d6b-453ca3201168 · inbound

OPT: Open Pre-trained Transformer Language Models cites this paper.

OPT: Open Pre-trained Transformer Language Models BASE Layers: Simplifying Training of Large, Sparse Models

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:53:17.632763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T20:53:16.720145Z digest=sha256:2a628c2d3baf676ebf84fb3063445b9605a28aba4d1df514c01556274b0dc696

Observation fb0dd2ee-7687-41a9-b492-3a32e9295655 · inbound

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale cites this paper.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale BASE Layers: Simplifying Training of Large, Sparse Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:35:36.151655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T13:35:35.972596Z digest=sha256:2c5ddf909a8b3afbbaa1184bd2f84916606fd875e4345965d4d7d022aadcc18f

Observation 66d12c9b-7f6f-4926-8994-39f71ca8ff5a · inbound

A Minimal Bifurcation Model of Load Imbalance in a Softmax Mixture-of-Experts Router cites this paper.

A Minimal Bifurcation Model of Load Imbalance in a Softmax Mixture-of-Experts Router BASE Layers: Simplifying Training of Large, Sparse Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:03:17.814167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T09:28:02.669459Z digest=sha256:72ffb2508d1ec69dfa8e9f4710c9fde68cf97f4876b02e4a21c47842a3a495da

Observation a98cf91f-94b7-4bdc-afd9-d0719badc588 · inbound

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism cites this paper.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism BASE Layers: Simplifying Training of Large, Sparse Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.299078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.299078Z digest=sha256:4b38cd9d459fcf597b2e18b5236a7754aeb6cee2d9bae78992813017ea18cb54

Observation d6a0a8a9-c102-4e9d-9621-8e5174d5fa11 · inbound

Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean Resolver to Hyperbolic-Spherical Geometry cites this paper.

Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean Resolver to Hyperbolic-Spherical Geometry BASE Layers: Simplifying Training of Large, Sparse Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:29:49.048895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:29:49.048895Z digest=sha256:75e30000143381aa6e70d266ae4f09a1add0131807add55699b264be0253252e