Pith. sign in

Paper Citation Record · LEDGER

Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2501.12370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12370 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:43:28.592940Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:17:25.673360Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2b1402f3-23ec-4a99-b4c5-b2d7da5810f7 · inbound

Scaling Inference-Efficient Language Models cites this paper.

Scaling Inference-Efficient Language Models Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.592940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.592940Z digest=sha256:30caecf67e4bc8a8f7c9ca1e42393142bbc25ba6e20356416f3b66c085fa9375

Observation d8853291-31a6-4438-838c-e8177ea277fc · inbound

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging cites this paper.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.391563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.391563Z digest=sha256:676780bbbadd25343004cae262b59bc6835afd4758247f1318e873a13fec8b5d

Observation 6361a183-bd8a-407d-b52b-e72d7762cd38 · inbound

Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient cites this paper.

Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:23.465646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:06:23.465646Z digest=sha256:92d99d5bf56d1b8e37e6de0dce80ba8594797a5f1ba841caf82b3572a22bcd1f

Observation 5936b078-f0f5-4a30-b5cf-48ee2b83b68a · inbound

Towards Foundational Models for Dynamical System Reconstruction: Hierarchical Meta-Learning via Mixture of Experts cites this paper.

Towards Foundational Models for Dynamical System Reconstruction: Hierarchical Meta-Learning via Mixture of Experts Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T19:48:11.228581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T19:48:11.228581Z digest=sha256:3cf29924f621c3905343cb0c6f17535a1aa01a78c5e81efc19066174224ce4b3

Observation 61bd56b4-0238-4bd8-9835-d8ebbd3b0c15 · inbound

Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models cites this paper.

Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:29.247011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:29.247011Z digest=sha256:01ac2d7067eae7d90fd827ed854eb14a887c295040682a89514621d512436ace

Observation 67bd1f81-0502-42bc-852f-28eb8af6e605 · inbound

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource cites this paper.

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:05:47.697749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T00:05:08.916339Z digest=sha256:adae8057d23fdf5bc23898b5f98ef15d9e7f326e00f8604849c17e01d8d20a71

Observation 709620e1-3cdb-448e-ae3e-15838ae8e350 · inbound

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs cites this paper.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:54.974819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:0139afd405d33050d8ca62d0af3f69d2c8b75d1ae766fd92d9ea7afa6642c7b5

Observation f13dd8c7-eedf-4c8b-ad1b-8a75355d77f7 · inbound

When Does Sparsity Mitigate the Curse of Depth in LLMs cites this paper.

When Does Sparsity Mitigate the Curse of Depth in LLMs Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T20:29:33.439034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:29:33.439034Z digest=sha256:91fea09966a45ffae8a172bf7880fce5da316b083f1d38686f2c968a550485cb

Observation 3caf1dfb-fc80-4345-8074-e8733f1aad62 · inbound

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts cites this paper.

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:21.516826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T03:29:16.555166Z digest=sha256:2b045186a109f75c5af2887f27241f43127fc3d027baf49a4d7981e2f7a3ab17

Observation ab36c26a-dc41-4aca-9625-6c996c4dc159 · inbound

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts cites this paper.

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:06:15.405336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:03:02.654035Z digest=sha256:9be3e7a63cd6867a1b811d5f2fb27abaed1357ee129501ef178d563ec86183e6

Observation e4cf1cc1-725f-4acd-a839-6a14408f17bd · inbound

DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts cites this paper.

DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:16:14.400239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T17:14:53.648013Z digest=sha256:40b848229dc1daf14901ae53e03376b4d96aa34c03a02899d7135709cf404d0d

Observation b6bbcd40-c9a2-4a63-9faa-e2e8350ceee5 · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:17:25.674860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:5649aa2ad9d36895b16f3e4fcd29c718c731e161ae2e96a5f01607740b744aea