Pith. sign in

Paper Citation Record · LEDGER

Scaling Laws for Upcycling Mixture-of-Experts Language Models

As of 9 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 2 inbound Pith citation observations for arXiv:2502.03009.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.03009 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:21:00.944082Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T02:03:02.654035Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T02:06:15.344429Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85da0116-e440-45da-9b37-6451776f526e · outbound

This paper cites write newline.

Scaling Laws for Upcycling Mixture-of-Experts Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.617897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.617897Z digest=sha256:162aab4edbdc14bcb0f33e4b0d1059fc1230fdb8ce60429edc82d1ac78c435c7

Observation a65fcdad-47f3-467c-b873-bb3ceeb51c9a · outbound

This paper cites GPT-4 Technical Report.

Scaling Laws for Upcycling Mixture-of-Experts Language Models GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.625802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.625802Z digest=sha256:bb563c08c4b097ca2fa0eac2cafd6beb562c136d25b6d7a75d1251828f7d87cd

Observation 47c36e10-15d0-4875-8aea-3400615cb5c4 · outbound

This paper cites Explaining neural scaling laws.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Explaining neural scaling laws

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:02.309774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.631698Z digest=sha256:070f12f185a5fd9e3a3e984c056d8fbecb64ced247291d2ac7395e35b36f5fbf

Observation b8a2c956-b33b-42f2-9e52-67bae9d3752b · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.636556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.636556Z digest=sha256:d71fe084df2788b68d0be176b2fc97800642ec6e8f7f34962eec121c695f6533

Observation 87b7a5ac-8255-4fb7-8bb3-19f496aea454 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Scaling Laws for Upcycling Mixture-of-Experts Language Models DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.641993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.641993Z digest=sha256:74cc97f1a8b44815c093e5510a1cb76100b52d6e516f4ee85f0e246a217e7bc2

Observation eda704f4-73e9-4697-9b78-1a0f8a949dd2 · outbound

This paper cites G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M.

Scaling Laws for Upcycling Mixture-of-Experts Language Models G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:02.294206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.647339Z digest=sha256:b3989b8727af14b8b485be78d293fc919241269b6d139366d7fb1ef7723540f7

Observation 4729b4da-ba19-4475-ba84-e1897fd071eb · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Piqa: Reasoning about physical commonsense in natural language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.652424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.652424Z digest=sha256:df6cb0f43ab853f2ebb50809667d0f782b2fcbaa898716180cc8d1f1687b08df

Observation 53b3d6b3-ade0-4e68-86da-341a02cd031d · outbound

This paper cites A dynamical model of neural scaling laws.

Scaling Laws for Upcycling Mixture-of-Experts Language Models A dynamical model of neural scaling laws

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:02.269975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.657831Z digest=sha256:2001bf156008894368f08ae3a760e4d9bd86dff8f32551ed1212c850f9c648cc

Observation 33224728-8591-4d2a-8024-1a4b53e80b2a · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

Scaling Laws for Upcycling Mixture-of-Experts Language Models A Survey on Mixture of Experts in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.662324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.662324Z digest=sha256:8b3c757e526ef18b9087981b1d426eec059b40769996d16fe7cd304fd74d9abc

Observation 7cb399b9-10f1-4eef-a836-8f9184ab22d9 · outbound

This paper cites Net2Net: Accelerating Learning via Knowledge Transfer.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Net2Net: Accelerating Learning via Knowledge Transfer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.667613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.667613Z digest=sha256:90e31764d459edef603451e17c167360f5d264545ba36fbaf031180f8d6a52fa

Observation 883287d9-ebfe-4c77-99bf-7a90a331fba5 · outbound

This paper cites W., Sutton, C., Gehrmann, S., et al.

Scaling Laws for Upcycling Mixture-of-Experts Language Models W., Sutton, C., Gehrmann, S., et al

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.672484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.672484Z digest=sha256:6fa9da812e14b64fc4df1ade1442cb2017c12284c1952a653c81cb2b0608a25c

Observation c4567530-195c-4bec-9dce-4af996636eca · outbound

This paper cites Unified scaling laws for routed language models.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Unified scaling laws for routed language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:02.244755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.677087Z digest=sha256:5ab3f4e55909427fe8143b293d2c9a5578ca76e2bd001575dca508049375dc60

Observation 5fd4668b-d383-4d62-a2d4-d3c3dc6075dc · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.681881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.681881Z digest=sha256:538cdf5fca86e803204ed76ffdb3aca06e155963b30177d5e55bfe765d189108

Observation a1bb4647-828f-44e9-bedb-d218e06b35db · outbound

This paper cites Redpajama: an open dataset for training large language models, October 2023.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Redpajama: an open dataset for training large language models, October 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:02.210857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.686705Z digest=sha256:00b52abec8074b04cae2326d0a804f728b723af0fe0e7082ab237c1d5754159c

Observation 6b07a1c1-4c52-4f7a-bc3b-29f2dd701fcf · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

Scaling Laws for Upcycling Mixture-of-Experts Language Models DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.691362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.691362Z digest=sha256:92f493225ebf05782ab020c70544a7e631f906908ca4a29edb59bcfb97666a09

Observation fda904c6-dc05-4204-95cd-fcc9d1026e5e · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.696339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.696339Z digest=sha256:0277e6d9db0d533836f139e3d6fe5037a9750231a96cb572d726436a2d7af59f

Observation 2a90d32c-9fd5-4821-a316-a649dcc32ee8 · outbound

This paper cites Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.701046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.701046Z digest=sha256:01db632a19e516f4bb23a6c11d95c7fe0566192af0e8e8182936be3a3ad13aea

Observation 87e5af3c-c95e-4256-a545-24f066e8d2df · outbound

This paper cites Learning Factored Representations in a Deep Mixture of Experts.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Learning Factored Representations in a Deep Mixture of Experts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.706164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.706164Z digest=sha256:175fd496d9a3ba0055fde0c8084b4ebdf1360ea50add115a6daaea1611aac170

Observation 3f3fdee1-9260-4c4f-a023-e5dbee9e29e9 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.711390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.711390Z digest=sha256:719191d1c50f35658c126a2404b45a5ea95459baf71f329a5bc8ec30dd6a323f

Observation 13336a41-e767-4b52-8e44-3e9a29b8e27f · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Scaling Laws for Upcycling Mixture-of-Experts Language Models A framework for few-shot language model evaluation, 07 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.716108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.716108Z digest=sha256:67917d0d570bf5f25a93c759e3735887f7ff7efe87b789ac0ca336ca14ace197

Observation ffa0b637-a104-430b-b604-a831ef976192 · outbound

This paper cites Scaling laws and compute-optimal training beyond fixed training durations.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Scaling laws and compute-optimal training beyond fixed training durations

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:02.164114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.720812Z digest=sha256:d909b1a67f19bb0ee980c8cd9d9cd7755c415e5a34067ae36e43ed9d3d871d5c

Observation 08b928ef-fcaa-4d83-8a1b-653c58669297 · outbound

This paper cites Upcycling Large Language Models into Mixture of Experts.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Upcycling Large Language Models into Mixture of Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.725253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.725253Z digest=sha256:2606020e0c3b0746a7617bfb78242657291f7d13b949427d9f8305636b5246c9

Observation 5b185aa1-0e58-4ccd-b21a-9a70429a31b7 · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Scaling Laws for Autoregressive Generative Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.730107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.730107Z digest=sha256:570b65fa1449750d287e22a2a130a983cfee11032c289c154119c224a1c87f34

Observation 9f7a8bd6-ae3d-4d58-b576-cd5ec6a8ec88 · outbound

This paper cites Scaling Laws for Transfer.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Scaling Laws for Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.735319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.735319Z digest=sha256:9ceb5ecf9ba9daf991424e185be8e658c14f49720131adcb4c6b6c9d2fe380a4

Observation 2f638896-1c91-4e00-bec3-70e5df8927ad · outbound

This paper cites Deep Learning Scaling is Predictable, Empirically.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Deep Learning Scaling is Predictable, Empirically

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.740322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.740322Z digest=sha256:7fb5500002542bd2e41428ab2058160be306482af5cdd8cd1b1bb863fd3301a3

Observation f37ac61d-49b3-45bc-98d2-d691cae9f443 · outbound

This paper cites Beyond human-level accuracy: Computational challenges in deep learning.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Beyond human-level accuracy: Computational challenges in deep learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:02.148518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.745347Z digest=sha256:17227fb8f1718e749e793a00853f6d7c38bbfca496a096c37601ad9511c0c825

Observation b47847ef-140e-41cc-9a4b-96f7f70dc5bb · outbound

This paper cites A., Welbl, J., Clark, A., et al.

Scaling Laws for Upcycling Mixture-of-Experts Language Models A., Welbl, J., Clark, A., et al

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:02.133361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.750248Z digest=sha256:f9b4039deda68f0f05df586909c43538e506949061f38fc25a315245ebedde44

Observation 1361b2c7-5d92-4014-8058-8aa869935150 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Scaling Laws for Upcycling Mixture-of-Experts Language Models MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.754953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.754953Z digest=sha256:5c0dcf139eac5d2f0a730ee41915b989207012fff07f978cb3f6c4611334082c

Observation 54430451-229d-4bbf-bab9-6a08dda57ccc · outbound

This paper cites Learning Curve Theory.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Learning Curve Theory

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.760063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.760063Z digest=sha256:38f1f0551400d2908dc5f94b6fffe329828bc64ac502b38ee124f06c7b35528f

Observation 1fda8a08-bd22-4934-b618-5ba91f6f76ec · outbound

This paper cites Mixtral of Experts.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Mixtral of Experts

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.764838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.764838Z digest=sha256:ac3492f8c91e750b7ee6e40dbd6bd2c1ed5cc5bb8c8323d5818ead3c18e845ab

Observation 50eb29d8-3277-47fd-94f3-a8aa9674764f · outbound

This paper cites Scaling Laws for Neural Language Models.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Scaling Laws for Neural Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.769659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.769659Z digest=sha256:d09a61113a391c62dfa119ad13c47303b3a421367fb3448b67b6bc60966ab3ca

Observation 50699452-3054-4c60-b654-854daaaab24d · outbound

This paper cites M., Hughes, S., Wolf, T., Bahdanau, D., et al.

Scaling Laws for Upcycling Mixture-of-Experts Language Models M., Hughes, S., Wolf, T., Bahdanau, D., et al

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:02.117618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.774674Z digest=sha256:a847f8d1fcd57435665345b9ea495c688a25980eec52b5e949c8e86e129369c9

Observation d5fcd860-a71a-42d2-9b18-4446f8d085c5 · outbound

This paper cites R., Mustafa, B., Ainslie, J., Tay, Y., Dehghani, M., and Houlsby, N.

Scaling Laws for Upcycling Mixture-of-Experts Language Models R., Mustafa, B., Ainslie, J., Tay, Y., Dehghani, M., and Houlsby, N

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:02.101081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.779697Z digest=sha256:fa79d7db127af7bca48582f15a57c47ac3658a9cbc9076639bd2ed7a51d670c5

Observation a6d2e450-2f8a-41b6-90cb-3a8fd444da18 · outbound

This paper cites A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B.

Scaling Laws for Upcycling Mixture-of-Experts Language Models A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.784808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.784808Z digest=sha256:fe98f63e0f8d2236421ad0cace733c590de3e80370265c8ce8b3d05c02009b52

Observation d2574811-c62c-4e89-af7d-ca2f6aad0326 · outbound

This paper cites Scaling Laws for Fine-Grained Mixture of Experts.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Scaling Laws for Fine-Grained Mixture of Experts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.789646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.789646Z digest=sha256:5391cc17d548bacaf37d4b040256322be09de8002d8f9f634d515738b93e1287

Observation 3c7a0ded-65cc-4836-b541-50bc857eacb2 · outbound

This paper cites S., Biderman, S., Elsahar, H., Muennighoff, N., Phang, J., Press, O., Raffel, C., Sanh, V., Shen, S., Sutawika, L., Tae, J., Yong, Z.

Scaling Laws for Upcycling Mixture-of-Experts Language Models S., Biderman, S., Elsahar, H., Muennighoff, N., Phang, J., Press, O., Raffel, C., Sanh, V., Shen, S., Sutawika, L., Tae, J., Yong, Z

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.795115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.795115Z digest=sha256:4f89636f91cbe1bdd8ff30aa8951378836133325fcac3f6713f21a8b19cb4f6b

Observation 4e2caa37-9260-42eb-9d83-df591fd26608 · outbound

This paper cites Gshard: Scaling giant models with conditional computation and automatic sharding.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Gshard: Scaling giant models with conditional computation and automatic sharding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:02.067364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.800048Z digest=sha256:78713905b4d15f90a69395f1257c9d97e60163ee3b721bc80dcd5b78c64de3d1

Observation 5de7d31e-7e08-4b32-a44d-8a24eff3d3a2 · outbound

This paper cites M., Bartlett, P., and Lee, J.

Scaling Laws for Upcycling Mixture-of-Experts Language Models M., Bartlett, P., and Lee, J

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:02.036098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.805002Z digest=sha256:83cc31e7c850ecd14e696a2f235e708a9bcfefd7446171cf579361a15ee2b60a

Observation b58d31b1-339a-4bd6-b826-84c72b41dbea · outbound

This paper cites Logiqa: a challenge dataset for machine reading comprehension with logical reasoning.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Logiqa: a challenge dataset for machine reading comprehension with logical reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:02.021522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.809795Z digest=sha256:b10fba5c00afb4f4e91c1a216d4077834a3509fdbb420919e9b4f1124ce3ca74

Observation 4ee4bebb-f43a-44a5-8107-e7a2c1a05720 · outbound

This paper cites GRIN: GRadient-INformed MoE.

Scaling Laws for Upcycling Mixture-of-Experts Language Models GRIN: GRadient-INformed MoE

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.814433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.814433Z digest=sha256:fde50a31c55c64ab1d84707fd513bfdd1032be1b529f4ba0011f87084be09e7c

Observation 86a159af-e51a-4f7d-8607-56031be5fcbe · outbound

This paper cites A Closer Look into Mixture-of-Experts in Large Language Models.

Scaling Laws for Upcycling Mixture-of-Experts Language Models A Closer Look into Mixture-of-Experts in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.819871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.819871Z digest=sha256:fcd026c4b16c45f047f76a52ace956533b2b3e1686d9ec7e5251c3de34cd7dae

Observation 1cdda9a7-164e-42c4-9021-44270f66114f · outbound

This paper cites Decoupled Weight Decay Regularization.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Decoupled Weight Decay Regularization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.824954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.824954Z digest=sha256:4ddf2c88b0a83f304f2645c91be9231b9586495b6101a8d836af28de292a7bda

Observation 2dc43d41-29b2-4d25-8cd5-f5c519b3891f · outbound

This paper cites A Solvable Model of Neural Scaling Laws.

Scaling Laws for Upcycling Mixture-of-Experts Language Models A Solvable Model of Neural Scaling Laws

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.829988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.829988Z digest=sha256:d9981c6a0c71ef2eccbdccdf15e5bc06bc4825d1bb734f6ab69a9644b1e12387

Observation f01c82c9-9c9c-43e3-b1b8-2988dfc586ad · outbound

This paper cites A scaling law for syn2real transfer: How much is your pre-training effective? pp.\ 477--492.

Scaling Laws for Upcycling Mixture-of-Experts Language Models A scaling law for syn2real transfer: How much is your pre-training effective? pp.\ 477--492

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:02.005441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.834868Z digest=sha256:e7cbb379c468d429a3ce8c48b0993f4ff95a5a3d2bf522f5ee9026e600a01d0c

Observation 8be729b1-2ceb-4a77-8185-3372d3f8c5b0 · outbound

This paper cites an unresolved cited work.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.839985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.839985Z digest=sha256:2abe6cd8bdc5bac1ba1f54fb2fc2b72372afca646aa50235420ed2afd94ab2d2

Observation 680872fe-0a7c-4365-9f05-cf91e01f32a7 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

Scaling Laws for Upcycling Mixture-of-Experts Language Models OLMoE: Open Mixture-of-Experts Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.844581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.844581Z digest=sha256:d6157468632b317bf062ead3205be46d4497fb9cb0d4fd23934915c3f135b9b6

Observation a2d039d1-370e-4106-b220-37765d9b164e · outbound

This paper cites The lambada dataset: Word prediction requiring a broad discourse context.

Scaling Laws for Upcycling Mixture-of-Experts Language Models The lambada dataset: Word prediction requiring a broad discourse context

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:01.981117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.849362Z digest=sha256:3d12445ba18df75c1720d18f33fe31529471f8c3d84eb43d71d704d33903966d

Observation 280992d9-4f18-4209-ab8e-d10ca4f736dd · outbound

This paper cites 4+ 3 phases of compute-optimal neural scaling laws.

Scaling Laws for Upcycling Mixture-of-Experts Language Models 4+ 3 phases of compute-optimal neural scaling laws

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:01.964523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.853726Z digest=sha256:7b741a1115fd336b61605526cc36b8f908f1449e450544a6731dd5cd75781be3

Observation 40cfef44-92d1-423f-9e14-20219f79c683 · outbound

This paper cites Resolving discrepancies in compute-optimal scaling of language models.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Resolving discrepancies in compute-optimal scaling of language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:01.947453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.858849Z digest=sha256:2bc57e3f683fe5f4e63d88ba3b0b115a52a8414a4498d9f93668eb3f2650531f

Observation f3b76803-d364-4546-93a5-74a6e24070d9 · outbound

This paper cites S., Rosenfeld, A., Belinkov, Y., and Shavit, N.

Scaling Laws for Upcycling Mixture-of-Experts Language Models S., Rosenfeld, A., Belinkov, Y., and Shavit, N

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:01.917112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.863756Z digest=sha256:4fe4b9f251cd7f83a7fffda67ff40fa19ee18f74ee3a6dc1b644406d54f20eb3

Observation 99732fc2-1166-484f-b7f2-71649a3a555a · outbound

This paper cites L., Bhagavatula, C., and Choi, Y.

Scaling Laws for Upcycling Mixture-of-Experts Language Models L., Bhagavatula, C., and Choi, Y

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.868372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.868372Z digest=sha256:c167bad01c945acf8fd8c4a7b5d5a996f53e6776907883b06d36315f6b0e8bef

Observation bc71ab9e-b12e-429d-b91e-87c3b666b0f6 · outbound

This paper cites GLU Variants Improve Transformer.

Scaling Laws for Upcycling Mixture-of-Experts Language Models GLU Variants Improve Transformer

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.872817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.872817Z digest=sha256:2043d81ac3d67aceac780351cc30f852188072c69a5af12b142f7e061f1cd9ab

Observation 6816f0a3-8d18-4162-ae58-85bcfa0696ab · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:01.874187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.877615Z digest=sha256:67c86b6c356a8f9509c8b41d897d99bd539ff2f5b2cba0cbaaf1318d92517de7

Observation d20a0011-25d6-411d-a42c-e1aed4ad5f61 · outbound

This paper cites SlimPajama-DC: Understanding Data Combinations for LLM Training.

Scaling Laws for Upcycling Mixture-of-Experts Language Models SlimPajama-DC: Understanding Data Combinations for LLM Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.882166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.882166Z digest=sha256:a8e3a15166baa7cbe55395fe7de9f4a6c34ad965f37b58603ce1454509ace789

Observation 3faeec56-b100-4767-b61e-50d2fa3cf0ca · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.886921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.886921Z digest=sha256:a079415ef8a221fa6cf91f190cf67f42277ffaddea71ceaa2a0970d9b0b3eb9a

Observation dc091f55-5163-4f97-9248-7fb4d3c0fcf2 · outbound

This paper cites R., Hestness, J., and Dey, N.

Scaling Laws for Upcycling Mixture-of-Experts Language Models R., Hestness, J., and Dey, N

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.892147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.892147Z digest=sha256:14ce3579f8e60b861112c5c594cc3c1f1db50c385f9f6114ac8818256ea4fb39

Observation 2acef10c-e539-4861-81a4-9229639810bc · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Roformer: Enhanced transformer with rotary position embedding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.896782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.896782Z digest=sha256:450c537d99b4f06a648305f949ebd30ea3f400d0ed7565f6efbeca50a5ddcbf6

Observation 40b78dfb-ca84-4a26-a305-2859f1f528e0 · outbound

This paper cites Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.901515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.901515Z digest=sha256:73e0ee00bcc1d746735ad83df335b09e7721d0d193a39ac2f9ecce5737ae5b6a

Observation 129ddd88-8d01-4c95-a87b-4cf01335da08 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.906340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.906340Z digest=sha256:174b9d10edcab3ae5f8f5aa59004c6e1560fae9af7c12b20188b6ad58342c020

Observation f98b996f-06a2-4953-8bfc-228e16830031 · outbound

This paper cites N., Kaiser, ., and Polosukhin, I.

Scaling Laws for Upcycling Mixture-of-Experts Language Models N., Kaiser, ., and Polosukhin, I

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.911209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.911209Z digest=sha256:8da9a1a3f6d170c8476e5d2e5a9906e78ea8d42848341a2320c8adc629488304

Observation 786be296-2a3e-4063-892b-bc080c2e5855 · outbound

This paper cites Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.915837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.915837Z digest=sha256:bae2dc7f282b27008788a9211f4867d8cfc8e473be07d45f9e0d61aaf0eba84c

Observation c1d2843b-91d9-40cd-ba17-3cc691608f57 · outbound

This paper cites F., and Gardner, M.

Scaling Laws for Upcycling Mixture-of-Experts Language Models F., and Gardner, M

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:01.821093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.920741Z digest=sha256:4b8340261b1c3d76c5fb75c759125635a40b4b0bd3275154678528b07ed115a2

Observation 73235337-3961-4fff-b80c-60d1b45d0480 · outbound

This paper cites Qwen2 Technical Report.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Qwen2 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.925327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.925327Z digest=sha256:9f3c20102f905ef3b4d2e18839a41737cf7026e812e77d7757712dfd7374f42d

Observation 7085dff3-2b7b-4c6e-a993-ab51154e4655 · outbound

This paper cites GLM-130B: An Open Bilingual Pre-trained Model.

Scaling Laws for Upcycling Mixture-of-Experts Language Models GLM-130B: An Open Bilingual Pre-trained Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.930068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.930068Z digest=sha256:a27a748128cce679bcf5de31ca5f1d2de08d998bd23238b9f0d48c38d8cba878

Observation 4e3656cd-e3ed-4ab5-ad68-2a6bf38c3847 · outbound

This paper cites Scaling vision transformers.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Scaling vision transformers

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:01.804747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.934916Z digest=sha256:0fccc236b1e033021b5f2cd572c4332d3782bd1c1880165f0c8ec2a10120100a

Observation ef2ec21b-60e4-446b-ad46-4afe06ae3ff9 · outbound

This paper cites When scaling meets llm finetuning: The effect of data, model and finetuning method.

Scaling Laws for Upcycling Mixture-of-Experts Language Models When scaling meets llm finetuning: The effect of data, model and finetuning method

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:21:01.785357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:21:00.939456Z digest=sha256:ab205c7ba332100a68634e7c899e8097d8e3da61e576117776bae25739edf6c2

Observation 1c69fbcc-5255-4efe-b4de-12c7ad6ebc77 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Scaling Laws for Upcycling Mixture-of-Experts Language Models TinyLlama: An Open-Source Small Language Model

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.944082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.944082Z digest=sha256:c1f0fdb99c424c8e361a349ce8b86ad64c4a7872ed9c76eed75ddf2167c1b605

Pith citing papers

Observation a66c818a-7cbf-4daf-8efb-cc02b96989f7 · inbound

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts cites this paper.

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts Scaling Laws for Upcycling Mixture-of-Experts Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:21.480155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:29:16.555166Z digest=sha256:1dc4653149ca2087419f9cdd6e347e0ab3fcf827ef601c6bb7aee2391df64f7f

Observation ada55683-84c5-462b-a6ac-97427250ae1b · inbound

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts cites this paper.

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts Scaling Laws for Upcycling Mixture-of-Experts Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:06:15.346925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:03:02.654035Z digest=sha256:6f7454db2ce33ad9d771e0025f76b4ed6e165501c9d19105134075473db4b259