Pith. sign in

Paper Citation Record · LEDGER

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

As of 5 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 2 inbound Pith citation observations for arXiv:2607.20548.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.20548 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T06:45:18.400599Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:19:37.877565Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24b3c119-0f44-45f3-b04d-4950669c0edd · outbound

This paper cites Adam: A Method for Stochastic Optimization.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Adam: A Method for Stochastic Optimization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:09.667912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:09.667912Z digest=sha256:3c1349c77c738fa7ae664fb328cd66d6241afe7f8aa3094cd56802c3c1dad092

Observation 04c785ae-fa76-4867-a04f-12418d339f05 · outbound

This paper cites Decoupled Weight Decay Regularization.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Decoupled Weight Decay Regularization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:09.783704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:09.783704Z digest=sha256:d23158fcb3b71142410c7149ad07a8d32e223257e135f4b8042274f0b917e0af

Observation 1b5f3881-521d-4e8a-adc2-e3e225d3ac3d · outbound

This paper cites Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:09.875431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:09.875431Z digest=sha256:2f0895870d751f24ac0dc4d40029d85dac00acecfa9dbbe8979ca89116d51d36

Observation 583768f7-0d65-4b9d-8ab3-6787c2ab6141 · outbound

This paper cites LaProp: Separating Momentum and Adaptivity in Adam.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales LaProp: Separating Momentum and Adaptivity in Adam

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:09.935650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:09.935650Z digest=sha256:658f793cd7d1cd0f93213d04f5185e5571f1a6ca0113081d9ca0d0c7e36443ec

Observation f6fa9006-bd39-4a99-92b5-100a5156d058 · outbound

This paper cites Fast approximate natural gradient descent in a kronecker factored eigenbasis.Advances in neural information processing systems, 31, 2018.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Fast approximate natural gradient descent in a kronecker factored eigenbasis.Advances in neural information processing systems, 31, 2018

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.034846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.034846Z digest=sha256:14bc2ce5938c47bd9c9b02b6a663cd1584ade1f01a1624bb71ff21ddd8466614

Observation 960149df-9156-47f5-a4a5-d4dbc1eb8b57 · outbound

This paper cites Optimizing neural networks with kronecker-factored approxi- mate curvature.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Optimizing neural networks with kronecker-factored approxi- mate curvature

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.085591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.085591Z digest=sha256:51e1e70c34e7d941f5cafeb9ed2d8feb67fc520ef4799913917c4fe797e31847

Observation dbe0683b-1f7c-4898-ad95-351fa82d782a · outbound

This paper cites A progressive batching l-bfgs method for machine learning.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales A progressive batching l-bfgs method for machine learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.245738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.245738Z digest=sha256:1c86e4d7548ec762c15ba00a1b8552fc0a1aef8ea0bbcce3c496aa6a5b768e21

Observation 6471271c-25a1-4623-bd6d-4e239ae7ec23 · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales SOAP: Improving and Stabilizing Shampoo using Adam

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.411447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.411447Z digest=sha256:8e95c0a9bf5d8b6aaba24320f349cfdf0a4d8b55cf823883d2e0aa0e46141491

Observation c4a5c059-ffe8-43a6-9b96-6a3883b0f0a0 · outbound

This paper cites Purifying shampoo: Investigating shampoo’s heuristics by decomposing its preconditioner.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Purifying shampoo: Investigating shampoo’s heuristics by decomposing its preconditioner

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.561041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.561041Z digest=sha256:081d3af09ce282a5a28b4a403947df2dd49fc893c603d7fe20ca9cdb4f6cc08a

Observation 6b60bb62-a945-4e99-9272-47654e70fdfe · outbound

This paper cites Understanding and improving the shampoo optimizer via kullback-leibler minimization.arXiv e-prints, pages arXiv–2509, 2025.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Understanding and improving the shampoo optimizer via kullback-leibler minimization.arXiv e-prints, pages arXiv–2509, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.707170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.707170Z digest=sha256:3894680c716f687371860ffddc986e33f645f4b1611c891cbeaf20ff57b3aa18

Observation 36bb432b-31dc-4628-9626-3b30782906b5 · outbound

This paper cites Training Deep Learning Models with Norm-Constrained LMOs.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Training Deep Learning Models with Norm-Constrained LMOs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.812240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.812240Z digest=sha256:22c634a50209496c5342939e0eba0fafdab77b5b21283f8a4680361eb0548bd1

Observation e615af54-4239-4f12-a5f5-0e2681fa0570 · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Shampoo: Preconditioned stochastic tensor optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.907906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.907906Z digest=sha256:2c9a4d360338888cf1c117af9d86bc2ae1665d25ebd2fc3a8e8ba32224f4a678

Observation 9a69c054-a9b9-485f-b4ce-e7bc2b0c8908 · outbound

This paper cites Muon is Scalable for LLM Training.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Muon is Scalable for LLM Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:11.096240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:11.096240Z digest=sha256:a0f4592e67985de1ea6217f33fc29e877d5d983fe80ae4e8f267447730731c5c

Observation 7910fb49-bf75-4774-be98-e01007f6edad · outbound

This paper cites Practical Efficiency of Muon for Pretraining.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Practical Efficiency of Muon for Pretraining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:11.255675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:11.255675Z digest=sha256:d9b9f76b516c24423ed603cb6be0bef53e826bbdf25d7edca580d667335e3216

Observation 1b3c3efd-595e-4bda-9410-8e66ec6475e9 · outbound

This paper cites The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:11.360068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:11.360068Z digest=sha256:19b7ce3c665b5863bb73161b635f1c96750e8bf452326efe23390d24802da476

Observation 8d6bc1ba-0e54-491d-8fae-534617199d7d · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks, 2024b.URL https://kellerjordan.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Muon: An optimizer for hidden layers in neural networks, 2024b.URL https://kellerjordan

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:11.491221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:11.491221Z digest=sha256:de6f4ede1a40824efc875d77fcd0653456480f7cbf6881ab4eafe6fd6f2636dc

Observation f6ed0731-c1ab-403f-ab7f-6be77ca1c507 · outbound

This paper cites DeepSeek-V3 Technical Report.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales DeepSeek-V3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:11.623090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:11.623090Z digest=sha256:5fe97cb7dda2515e8887520a5f181e82e42c53f8e5b4af7db8b640851618b2c2

Observation 7cb5d772-0a91-4b0e-a2a1-11c3498113fa · outbound

This paper cites Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:11.904635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:11.904635Z digest=sha256:cff1d6fb17e275e173a08123ac48f52d25f1303b6bef6495d788c9f1e00789fb

Observation ae0008f2-8e98-464f-b4ca-30bcf5ad6dec · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.033235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.033235Z digest=sha256:13ebed757836e1b4aa406fdf3e215cf874ced868b039088d82691a3839a82ae3

Observation e33c0488-b8fd-4afe-bbc3-ae038b6a67ba · outbound

This paper cites On the sdes and scaling rules for adaptive gradient algorithms.Advances in Neural Information Processing Systems, 35:7697–7711, 2022.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales On the sdes and scaling rules for adaptive gradient algorithms.Advances in Neural Information Processing Systems, 35:7697–7711, 2022

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.193929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.193929Z digest=sha256:b0f696db9d2579ee2ca11a9c840e2b206d303019699e1a75148b848f29fbcad3

Observation 7d6cd56e-b842-46b8-9f51-76f9f692011d · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.318217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.318217Z digest=sha256:5dd7037825cea6000ba87cdd1d9ba50ceb8e675edd1411c18dcb89e4af040201

Observation 1c8cb48c-1b66-414f-9d7e-629157238217 · outbound

This paper cites Zero: Memory optimiza- tions toward training trillion parameter models.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Zero: Memory optimiza- tions toward training trillion parameter models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.447125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.447125Z digest=sha256:046daa8a3fbd61f97310b618d27b18b58f8ac36f1b36abfadcc8a4d5c3af249c

Observation 8a20dc7e-1b59-41eb-9a72-b22682ab6864 · outbound

This paper cites Scalable Second Order Optimization for Deep Learning.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scalable Second Order Optimization for Deep Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.628217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.628217Z digest=sha256:19b4a7cf850bf016ef1e37058f360fd83ecbc9cc4f9d4adb543970086624a919

Observation 143f7807-a46c-4659-a567-62ad324b12cd · outbound

This paper cites DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.772594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.772594Z digest=sha256:e00094023dd0a7e3eaa3ddaf0d793b8768e056837363f2d3275cec862dacd3f3

Observation 5d724d58-fcc3-450b-bf1c-691b1d8f3b2a · outbound

This paper cites Precondi- tioned spectral descent for deep learning.Advances in neural information processing systems, 28, 2015.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Precondi- tioned spectral descent for deep learning.Advances in neural information processing systems, 28, 2015

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.853115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.853115Z digest=sha256:15f544b51cd524dd35965315746ae456b2f55cf48617fc19fa21789224dc985d

Observation 196eabd2-3d09-436e-8ffa-c5989a880cfa · outbound

This paper cites Stochastic spectral descent for restricted boltzmann machines.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Stochastic spectral descent for restricted boltzmann machines

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.936009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.936009Z digest=sha256:6f50d9a5486513c9917ac4730d21d29cba4602296602477b3116b63a7fab7242

Observation 4ddaba03-2b0d-494e-8568-8e5adb2e9b38 · outbound

This paper cites Stochastic spectral descent for discrete graphical models.IEEE Journal of Selected Topics in Signal Processing, 10(2):296–311, 2015.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Stochastic spectral descent for discrete graphical models.IEEE Journal of Selected Topics in Signal Processing, 10(2):296–311, 2015

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:13.072219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:13.072219Z digest=sha256:29ff6c53434acfacc80f63b84621bcd450f2ea9c5b410926b7aaae56c3f9e66f

Observation 7f0d3422-b192-46fb-8d8d-e6784f27a2cc · outbound

This paper cites The duality structure gradient descent algorithm: analysis and applications to neural networks.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales The duality structure gradient descent algorithm: analysis and applications to neural networks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:13.215435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:13.215435Z digest=sha256:5176d5037142dbe609022f29fa88010ca74fda51432205d5f1ecca2fadf4f9ca

Observation 9344d4b5-482f-4669-bde0-23ea357f164f · outbound

This paper cites Modular Duality in Deep Learning.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Modular Duality in Deep Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:13.333343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:13.333343Z digest=sha256:4abf7a87676df12fb8501e65c6c625858199fbcddc7e79ba5721e8a3344ceaf4

Observation fe6db98e-2778-4a76-817f-3e3b684d676b · outbound

This paper cites Old Optimizer, New Norm: An Anthology.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Old Optimizer, New Norm: An Anthology

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:13.459657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:13.459657Z digest=sha256:858db6700197c097fc82780494c32eec0a08920712467b4881435ffdac04974f

Observation c5dd4445-3ab7-4f60-abd2-b1ec2f7ce1bf · outbound

This paper cites Scalable Optimization in the Modular Norm.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scalable Optimization in the Modular Norm

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:13.650206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:13.650206Z digest=sha256:f6eab9e1bf857e0dccdb7b976b5f5f37473a5d276c134da14b808f9282fd363d

Observation 8f0c4daa-6a9b-42a1-b7c7-9a0712126ce9 · outbound

This paper cites MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:13.858281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:13.858281Z digest=sha256:e4e68735a9aa29bae0ff541708d678e7219e52a9a0c4188c3c86cb74b77b3d57

Observation cb74dc77-c2f2-4d48-8721-054ed188f737 · outbound

This paper cites Towards a principled muon under𝜇p: Ensuring spectral conditions throughout training.arXiv preprint arXiv:2601.01306, 2026.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Towards a principled muon under𝜇p: Ensuring spectral conditions throughout training.arXiv preprint arXiv:2601.01306, 2026

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.041176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.041176Z digest=sha256:20cddbc1b94527e0f5be1b36653b6c5afd3b13b0076c8bbd92d9521147f6173d

Observation 12fc0053-0d22-46f7-a0d1-55c94af31181 · outbound

This paper cites Adamuon: Adaptive muon optimizer.arXiv preprint arXiv:2507.11005, 2025.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Adamuon: Adaptive muon optimizer.arXiv preprint arXiv:2507.11005, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.198277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.198277Z digest=sha256:f2ffc757b125d4fd3d58d869b289777859f36872f771fd448b41bba0749f4fa8

Observation cb874508-99c0-40b3-902f-5457386e0d67 · outbound

This paper cites Normuon: Making muon more efficient and scalable.arXiv preprint arXiv:2510.05491, 2025.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Normuon: Making muon more efficient and scalable.arXiv preprint arXiv:2510.05491, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.357745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.357745Z digest=sha256:957a9942be72241bfc1e537c20d429ec10d8a2379c3c59bb017ef386cba7f41c

Observation 0355528b-3d71-4247-b20c-0a44613f9d01 · outbound

This paper cites Fantastic pretraining optimizers and where to find them ii: From weight decay to hyperball optimization, 12 2025.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Fantastic pretraining optimizers and where to find them ii: From weight decay to hyperball optimization, 12 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.446539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.446539Z digest=sha256:d47d08989a5aabadca8786cf8227fb8aea8b66cc8d1bfdaccd1c858cf3ef8ed4

Observation 0f139f67-6d74-4086-8116-65a3081c0abb · outbound

This paper cites Controlled llm training on spectral sphere.arXiv preprint arXiv:2601.08393, 2026.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Controlled llm training on spectral sphere.arXiv preprint arXiv:2601.08393, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.558228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.558228Z digest=sha256:54211684d4b23552e6cfc13bbec1eae10dde435a09c0f09b17483969c825af54

Observation 6ff5efe9-4b21-46db-afbb-c0a149f7f567 · outbound

This paper cites Adam improves muon: Adaptive moment estimation with orthogonalized momentum.arXiv preprint arXiv:2602.17080, 2026.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Adam improves muon: Adaptive moment estimation with orthogonalized momentum.arXiv preprint arXiv:2602.17080, 2026

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.680021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.680021Z digest=sha256:5f429b24be9f88011408ee34244808b713a87a431e49c8695322f480faf0195c

Observation 90fbaf49-3678-4346-a1d8-701290c5a6de · outbound

This paper cites Manifold constrained steepest descent.arXiv preprint arXiv:2601.21487, 2026.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Manifold constrained steepest descent.arXiv preprint arXiv:2601.21487, 2026

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.814783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.814783Z digest=sha256:2c3ce1b53b0ca642f32302be2c3d0d9e68111da0600ae9e28c40266c8b15e326

Observation d5abe931-8959-4242-a595-eb7d7302b4bb · outbound

This paper cites The newton-muon optimizer.arXiv preprint arXiv:2604.01472, 2026.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales The newton-muon optimizer.arXiv preprint arXiv:2604.01472, 2026

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.970929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.970929Z digest=sha256:208f9dedac7935f6d3062eba90bd16c62d6d2020e7304d17cc8cf480f08c681f

Observation e2e59c56-13ea-4b79-9f26-ca532089e751 · outbound

This paper cites Mousse: Rectifying the geometry of muon with curvature-aware preconditioning.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Mousse: Rectifying the geometry of muon with curvature-aware preconditioning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.138354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.138354Z digest=sha256:0b33300e73db6314f24cac414ed0edb3b4b7ce790f3c72fb01d499c7ab8d1b1d

Observation afdba9c2-f340-4d9f-8127-a4445be7afed · outbound

This paper cites veScale-FSDP: Flexible and High-Performance FSDP at Scale.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales veScale-FSDP: Flexible and High-Performance FSDP at Scale

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.274057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.274057Z digest=sha256:9cd5048fab2bf0ed6c46fd04d41bc6a549e4a3ae9f5a2c45e0a595abc3202ec2

Observation 3a78d3df-ded1-4a62-8f36-e0e22de95125 · outbound

This paper cites Canzona: A unified, asynchronous, and load-balanced framework for distributed matrix-based optimizers.arXiv preprint arXiv:2602.06079, 2026.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Canzona: A unified, asynchronous, and load-balanced framework for distributed matrix-based optimizers.arXiv preprint arXiv:2602.06079, 2026

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.365593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.365593Z digest=sha256:0530732a34de7f1f1b918d46c53cb8056238b139e3c5b467a1e40e2e10e9fee7

Observation f4aeb0c7-eeb3-4584-9e46-caac917c807e · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.479625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.479625Z digest=sha256:21b5fb30ba8b48b3dcd6d1fd58e809f7490e19901048d4a7c4610bba71b9fc35

Observation 79c929b9-1e0b-4fe6-8127-e129ef6366e9 · outbound

This paper cites Qwen3 Technical Report.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Qwen3 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.554716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.554716Z digest=sha256:37548b6b7f7fadf45c5abb4857a7355dc465717364be545547d0997715e42397

Observation b852a0db-a10b-4374-b161-84fe4122c301 · outbound

This paper cites Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.647299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.647299Z digest=sha256:529bb4f36dc2142a3edb5f303a9de55bfef35036e350f4c62acabc98f50af919

Observation 6afbda3e-b356-4586-b040-b4bc1235ad81 · outbound

This paper cites Scaling laws and compute-optimal training beyond fixed training durations.Advances in Neural Information Processing Systems, 37:76232–76264, 2024.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scaling laws and compute-optimal training beyond fixed training durations.Advances in Neural Information Processing Systems, 37:76232–76264, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.766177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.766177Z digest=sha256:32947cc5842f21c247a3a2f39be9b241c24c98f52d054c03dde9b26312d358d6

Observation e8a9c316-bb80-4cc9-8888-88f47dfbb6ec · outbound

This paper cites AdamW Weight RMS.https: // kexue.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales AdamW Weight RMS.https: // kexue

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.860483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.860483Z digest=sha256:b85bb9efd20a0108ff0cf5505671a9677001573380fc82fc52d0008bef11f8d0

Observation 2d5699f8-5d94-44e8-91a1-0395eb3fe87f · outbound

This paper cites Stochastic hessian fittings with lie groups.arXiv preprint arXiv:2402.11858, 2024.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Stochastic hessian fittings with lie groups.arXiv preprint arXiv:2402.11858, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.968785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.968785Z digest=sha256:15d196a3d1d502ee374f5b60c41354812eccb08bb2ac37ecf83cf615fecb56f2

Observation 878d8978-c3a8-4f8b-9b0a-410cf6543df5 · outbound

This paper cites Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.141224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.141224Z digest=sha256:6b38c59e15aacdc20988a3020a65e87dcbe2dd0be401f9557b8a366f529daaca

Observation 9a1e34b9-f4ff-46f4-ae65-0ee446bb34cf · outbound

This paper cites On the importance of initialization and momentum in deep learning.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales On the importance of initialization and momentum in deep learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.231667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.231667Z digest=sha256:1c3a003151e240f56d2160e1eaa40685711ec5286d50be2b486fc628a3c02bb0

Observation 091b542e-1483-4e6a-8cf8-956074cec15d · outbound

This paper cites Incorporating nesterov momentum into adam.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Incorporating nesterov momentum into adam

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.328886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.328886Z digest=sha256:a8798139fd53972d2759aeaf64a959cf23821a1c39ad3d5964ed60a2317e6e61

Observation 03a165c4-03da-46cf-bcfc-fbc1d85880fa · outbound

This paper cites An Empirical Model of Large-Batch Training.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales An Empirical Model of Large-Batch Training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.428609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.428609Z digest=sha256:ae7b2d372631a0dba53b7be076450ea60b8d62b118f49661ff6fd7786de6a95e

Observation 3c457dcc-ced8-4f86-8e6d-43d37c92b5be · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.519251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.519251Z digest=sha256:a01ef556149550919c8301ce1e3e96cbb83d764038bb09789037d2d5a0bea862

Observation 7afbe3e8-33bc-483f-ab41-728fb29e6a88 · outbound

This paper cites Large batch optimization for deep learning: Training bert in 76 minutes.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Large batch optimization for deep learning: Training bert in 76 minutes

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.597763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.597763Z digest=sha256:8d27f41172a682efc9fe2557be8215ce3684fa1537f94d80ba804f8ecc108be2

Observation d3a4cf97-6a11-4c93-a5f7-4ba4d6371442 · outbound

This paper cites Train longer, generalize better: closing the generalization gap in large batch training of neural networks.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Train longer, generalize better: closing the generalization gap in large batch training of neural networks

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.690586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.690586Z digest=sha256:2ccf9d1cb6b59b5c943bf59e60ef531d424d6dee62e0fb4e0ef7e3fae9227938

Observation f33cc214-6673-4722-a1a7-a122be0ab08c · outbound

This paper cites Large Batch Training of Convolutional Networks.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Large Batch Training of Convolutional Networks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.783035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.783035Z digest=sha256:899fad370800344421ace6dba01f2f341ce735fe4f4f0da5c16f3c7673165114

Observation ba1434f9-71b8-4ebd-93fc-f5b258b7ca97 · outbound

This paper cites One weird trick for parallelizing convolutional neural networks.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales One weird trick for parallelizing convolutional neural networks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.934273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.934273Z digest=sha256:976cb31851af2121d5e019eafd3b3f9dea4d0f572ac47182eb33117a0db865f4

Observation b94154f1-d860-4bd8-a50c-b1197948a0dd · outbound

This paper cites Better & faster large language models via multi-token prediction.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Better & faster large language models via multi-token prediction

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.032023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.032023Z digest=sha256:009e0481e5c7b838cec007ffec9e7850cf3eb50a1f318a043ddb1b4ec2357dad

Observation 9ca1e866-9879-44e1-8c69-1d70f4252a2c · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.122273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.122273Z digest=sha256:55ad4cb714714b2774c0b0cf792bcab344352480cfac2ccec18a768ec4c47507

Observation 0ca474ff-08db-4668-8d38-4b65e51d4f23 · outbound

This paper cites Scaling Exponents Across Parameterizations and Optimizers.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scaling Exponents Across Parameterizations and Optimizers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.238514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.238514Z digest=sha256:a3bd38f942539c28a6e76f23d019bd728c51f06c0300df4db695abc97f78cc1b

Observation 7736d6ae-3bda-448a-8cda-86d9d4c88fce · outbound

This paper cites Small-scale proxies for large-scale Transformer training instabilities.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Small-scale proxies for large-scale Transformer training instabilities

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.325825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.325825Z digest=sha256:77d33c377157a26d4327b40e7d2c54ce1834d53847406ea19b5e51c396914a17

Observation cb5521bc-be78-4a00-bd29-7497c9c6c2de · outbound

This paper cites Tensor Programs IVb: Adaptive Optimization in the Infinite-Width Limit.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Tensor Programs IVb: Adaptive Optimization in the Infinite-Width Limit

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.441355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.441355Z digest=sha256:638eb6ebb561ddc61256f6ca30a1e8ad18e30d029fe168fb621f1e11016ebb26

Observation a0cd8fc0-c97a-492d-b98b-10299d4474ff · outbound

This paper cites an unresolved cited work.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.502177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.502177Z digest=sha256:d32f9a6fdf80060c6b09f2e324d5c2b0d88e67ab21f2eef5715badeccc2ad960

Observation f0f33c12-0406-4afd-9e1d-95246b8dcd82 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.620629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.620629Z digest=sha256:0412f0a8fc5a1a461cf41bfdbd69c33b012b2f6d296b55cca378993bf960e2f4

Observation a247749e-4341-4344-9723-5365be0dad2a · outbound

This paper cites Stabilizing Native Low-Rank LLM Pretraining.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Stabilizing Native Low-Rank LLM Pretraining

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.674394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.674394Z digest=sha256:89d292c1386ee5a7c51a2ccfa372eafbee0ca85101bd040a3f10ada4d6a615dc

Observation 82daf934-0776-4118-a6b7-f616e56bda32 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.759660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.759660Z digest=sha256:66a029eab0048ac4700acf4732bf039401ca187ad730cd3500fc9a0f01f7b234

Observation 92ff9c0d-a7ff-4c5c-9f4c-26e6bab97084 · outbound

This paper cites Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.820219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.820219Z digest=sha256:5d5c45426afe41e732549ed4519e244d927e41339fd8e55f7154eb6d6f126d0d

Observation 30005dc9-09e8-4498-b90e-1a5701d11666 · outbound

This paper cites On the role of batch size in stochastic conditional gradient methods.arXiv preprint arXiv:2603.21191, 2026.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales On the role of batch size in stochastic conditional gradient methods.arXiv preprint arXiv:2603.21191, 2026

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.872461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.872461Z digest=sha256:e7f7f5c7efde5c6545dab4ea1f4578ef73d890de813f92635f2d8d5322132f9f

Observation 806beb69-1116-4fc1-936b-d60b22816623 · outbound

This paper cites Spectral Scaling Laws of Muon.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Spectral Scaling Laws of Muon

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.932930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.932930Z digest=sha256:18f674f1d662212b012548694bc92b9ae9e7d49676e6ae878b5fed620c366ef4

Observation de41cb6e-f64a-4414-beba-cdea8e55294d · outbound

This paper cites Dissecting adam: The sign, magnitude and variance of stochastic gradients.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Dissecting adam: The sign, magnitude and variance of stochastic gradients

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:18.077028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:18.077028Z digest=sha256:302c1a31325c7616e360e2d0ee6222eff37ee75ab6b1bf31cf514f4f1cc56ba7

Observation e96956c5-1e58-4b1b-9a58-788abb95e819 · outbound

This paper cites A New Perspective on Shampoo's Preconditioner.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales A New Perspective on Shampoo's Preconditioner

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:18.187657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:18.187657Z digest=sha256:b63bcb0cce8fb2064538ad5d88dea76f2beb3e50eefcfef77c516d6ac308da4b

Observation cc3f4a9d-f5f0-4d3c-9c6d-a1906fda7d6a · outbound

This paper cites Recipes for Pre-training LLMs with MXFP8.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Recipes for Pre-training LLMs with MXFP8

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:18.277482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:18.277482Z digest=sha256:5f588384f609810757ce1fa0fc32b125f3c5ac77de71380a91d7bcf866ad3c10

Observation a77b5641-96a2-4418-9b83-a5f60b0352f6 · outbound

This paper cites Symbolic discovery of optimization algorithms.Advances in neural information processing systems, 36:49205–49233, 2023.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Symbolic discovery of optimization algorithms.Advances in neural information processing systems, 36:49205–49233, 2023

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:18.400599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:18.400599Z digest=sha256:ed65bbccd7391cc83ac2123afc47472100fc3d2ff2fb0c703e5c65e5bd7a0fa5

Pith citing papers

Observation d2249c7c-3bcb-425f-b743-688c4ff23520 · inbound

When Does Muon Help Agentic Reinforcement Learning? cites this paper.

When Does Muon Help Agentic Reinforcement Learning? SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T21:13:38.825691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:13:38.825691Z digest=sha256:77a435462717ae81678770b22a665dd9ed00172bcfea9024037534adc18b0f4b

Observation f0a3e926-f1ad-4016-8abd-f1b2aeac5ec5 · inbound

When Does Muon Help Agentic Reinforcement Learning? cites this paper.

When Does Muon Help Agentic Reinforcement Learning? SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:37.877565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:37.877565Z digest=sha256:571edf470337f68deedc2ec53e70470812292d2b85e933d61d719760afc04625