Pith. sign in

Paper Citation Record · LEDGER

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

As of 22 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 2 inbound Pith citation observations for arXiv:2607.20548.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.20548 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T06:45:18.400599Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:19:37.877565Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24b3c119-0f44-45f3-b04d-4950669c0edd · outbound

This paper cites Adam: A Method for Stochastic Optimization.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Adam: A Method for Stochastic Optimization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:09.667912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:09.667912Z digest=sha256:c188857e7bab51b6891d3d12991eb3a771d53f1b2e58cd5a7b2cdc3a4088ad1e

Observation 04c785ae-fa76-4867-a04f-12418d339f05 · outbound

This paper cites Decoupled Weight Decay Regularization.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Decoupled Weight Decay Regularization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:09.783704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:09.783704Z digest=sha256:bf1aa751d557e02cc7d8ff41edd894f5b68621ea9d0aed2c47413f2e1bc5fd35

Observation 1b5f3881-521d-4e8a-adc2-e3e225d3ac3d · outbound

This paper cites Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:09.875431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:09.875431Z digest=sha256:387eb950fc3e73c987ff9226057d4926bfd15a214c040e659ac31ac4ac9a6941

Observation 583768f7-0d65-4b9d-8ab3-6787c2ab6141 · outbound

This paper cites LaProp: Separating Momentum and Adaptivity in Adam.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales LaProp: Separating Momentum and Adaptivity in Adam

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:09.935650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:09.935650Z digest=sha256:2e34a1923e7ea7fb79b58f99fc1466b72ff70c290ce8ed25d1200db46a34d4d4

Observation f6fa9006-bd39-4a99-92b5-100a5156d058 · outbound

This paper cites Fast approximate natural gradient descent in a kronecker factored eigenbasis.Advances in neural information processing systems, 31, 2018.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Fast approximate natural gradient descent in a kronecker factored eigenbasis.Advances in neural information processing systems, 31, 2018

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.034846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.034846Z digest=sha256:1c216e27029ed4b4181ff45d3d1e06d6633edf864706fa782b77c68f0ed669eb

Observation 960149df-9156-47f5-a4a5-d4dbc1eb8b57 · outbound

This paper cites Optimizing neural networks with kronecker-factored approxi- mate curvature.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Optimizing neural networks with kronecker-factored approxi- mate curvature

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.085591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.085591Z digest=sha256:9224e92f0a5d97606c5eb54ecee282b6b957ecd6be87d4801570e920d1c85b35

Observation dbe0683b-1f7c-4898-ad95-351fa82d782a · outbound

This paper cites A progressive batching l-bfgs method for machine learning.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales A progressive batching l-bfgs method for machine learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.245738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.245738Z digest=sha256:e4554e3025b15891fbb53ed4f334310753788dd8c882580ce7ca98ca60e49834

Observation 6471271c-25a1-4623-bd6d-4e239ae7ec23 · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales SOAP: Improving and Stabilizing Shampoo using Adam

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.411447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.411447Z digest=sha256:f07fe4e5bc0bcde64ac82b1b2801ab5a9c14c6e90a262c6a09847a48cdef4c44

Observation c4a5c059-ffe8-43a6-9b96-6a3883b0f0a0 · outbound

This paper cites Purifying shampoo: Investigating shampoo’s heuristics by decomposing its preconditioner.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Purifying shampoo: Investigating shampoo’s heuristics by decomposing its preconditioner

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.561041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.561041Z digest=sha256:17a21c584ceb956e68505e61508ea9716af2b1fe8abfff5e867f0262ab6d6f92

Observation 6b60bb62-a945-4e99-9272-47654e70fdfe · outbound

This paper cites Understanding and improving the shampoo optimizer via kullback-leibler minimization.arXiv e-prints, pages arXiv–2509, 2025.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Understanding and improving the shampoo optimizer via kullback-leibler minimization.arXiv e-prints, pages arXiv–2509, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.707170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.707170Z digest=sha256:b6633d4b5cda6d08d89223815857f3a8f8e122290b9064b72ca761ce6d567728

Observation 36bb432b-31dc-4628-9626-3b30782906b5 · outbound

This paper cites Training Deep Learning Models with Norm-Constrained LMOs.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Training Deep Learning Models with Norm-Constrained LMOs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.812240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.812240Z digest=sha256:f2373b92c1b3f45ebc44fc475c4a04b2296c4fdda88a37e197c394f081154f95

Observation e615af54-4239-4f12-a5f5-0e2681fa0570 · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Shampoo: Preconditioned stochastic tensor optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:10.907906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:10.907906Z digest=sha256:f72b7d45f718335b696fd0132024d808bca4eb911eceaecc1a71a570606f61c0

Observation 9a69c054-a9b9-485f-b4ce-e7bc2b0c8908 · outbound

This paper cites Muon is Scalable for LLM Training.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Muon is Scalable for LLM Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:11.096240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:11.096240Z digest=sha256:c8722ece8cd9d7ba452a1f1deb5b2e7b2179d274b5031c0c52a48fb6c67091c2

Observation 7910fb49-bf75-4774-be98-e01007f6edad · outbound

This paper cites Practical Efficiency of Muon for Pretraining.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Practical Efficiency of Muon for Pretraining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:11.255675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:11.255675Z digest=sha256:696a5fd0447fcec5b2f39a89bd17bbf5214b242d247eb56af4b5faafa2d83c4b

Observation 1b3c3efd-595e-4bda-9410-8e66ec6475e9 · outbound

This paper cites The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:11.360068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:11.360068Z digest=sha256:1cc19a71af83925c3d4fab2af8e5f3895d93911bf41b1e395a84f7596613956a

Observation 8d6bc1ba-0e54-491d-8fae-534617199d7d · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks, 2024b.URL https://kellerjordan.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Muon: An optimizer for hidden layers in neural networks, 2024b.URL https://kellerjordan

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:11.491221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:11.491221Z digest=sha256:d218d740ae42899bd8977a0c04071f582634022b1719a16bd6b6215595a4f9fb

Observation f6ed0731-c1ab-403f-ab7f-6be77ca1c507 · outbound

This paper cites DeepSeek-V3 Technical Report.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales DeepSeek-V3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:11.623090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:11.623090Z digest=sha256:a77e9fbfd67f5631f4991b2d33aebad7e422ce9733c2858abbcf5144dfb0e9ca

Observation 7cb5d772-0a91-4b0e-a2a1-11c3498113fa · outbound

This paper cites Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:11.904635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:11.904635Z digest=sha256:f1488ecd4bd520e381ab644d92801da46ba186b821fba93812334df80328af2a

Observation ae0008f2-8e98-464f-b4ca-30bcf5ad6dec · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.033235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.033235Z digest=sha256:845a28fb4c374412f4430f19771cf8c2d6c764244e30d4cca24bfe1fba72e0e0

Observation e33c0488-b8fd-4afe-bbc3-ae038b6a67ba · outbound

This paper cites On the sdes and scaling rules for adaptive gradient algorithms.Advances in Neural Information Processing Systems, 35:7697–7711, 2022.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales On the sdes and scaling rules for adaptive gradient algorithms.Advances in Neural Information Processing Systems, 35:7697–7711, 2022

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.193929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.193929Z digest=sha256:dbd37991945569c36393784aa8623fd88b63d86d5f3d1f73c884400a20999132

Observation 7d6cd56e-b842-46b8-9f51-76f9f692011d · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.318217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.318217Z digest=sha256:78399f6ab4d9bec1d8032c26c750678eb57e4ae4bdeeb9c5f9b9342046aa452d

Observation 1c8cb48c-1b66-414f-9d7e-629157238217 · outbound

This paper cites Zero: Memory optimiza- tions toward training trillion parameter models.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Zero: Memory optimiza- tions toward training trillion parameter models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.447125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.447125Z digest=sha256:b15dbcad3917d6f036bf73d457eb960376a04839d0abb41632c8cb459c7299da

Observation 8a20dc7e-1b59-41eb-9a72-b22682ab6864 · outbound

This paper cites Scalable Second Order Optimization for Deep Learning.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scalable Second Order Optimization for Deep Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.628217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.628217Z digest=sha256:4c3548bf162d649d782f4d71f1c48f9e7ada6a084be0f1828df82381db04a916

Observation 143f7807-a46c-4659-a567-62ad324b12cd · outbound

This paper cites DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.772594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.772594Z digest=sha256:cdc0d0372aeb7bd5b19e7f96c22fb07332cd6a3ba65c116c0c0f366fc7429f0b

Observation 5d724d58-fcc3-450b-bf1c-691b1d8f3b2a · outbound

This paper cites Precondi- tioned spectral descent for deep learning.Advances in neural information processing systems, 28, 2015.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Precondi- tioned spectral descent for deep learning.Advances in neural information processing systems, 28, 2015

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.853115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.853115Z digest=sha256:1468dd936793c62e6ef0f3a4c185430c9b63d963e5c72faec30f0c3c869e9cf3

Observation 196eabd2-3d09-436e-8ffa-c5989a880cfa · outbound

This paper cites Stochastic spectral descent for restricted boltzmann machines.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Stochastic spectral descent for restricted boltzmann machines

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:12.936009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:12.936009Z digest=sha256:507498f3a2dfc594c442db18df52928efa744ae4341404fef6e6b33ec1e9cf39

Observation 4ddaba03-2b0d-494e-8568-8e5adb2e9b38 · outbound

This paper cites Stochastic spectral descent for discrete graphical models.IEEE Journal of Selected Topics in Signal Processing, 10(2):296–311, 2015.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Stochastic spectral descent for discrete graphical models.IEEE Journal of Selected Topics in Signal Processing, 10(2):296–311, 2015

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:13.072219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:13.072219Z digest=sha256:be2d81d685097ebf87fb38a860797e84f5ee099c3d99bf7fb385b609a960c508

Observation 7f0d3422-b192-46fb-8d8d-e6784f27a2cc · outbound

This paper cites The duality structure gradient descent algorithm: analysis and applications to neural networks.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales The duality structure gradient descent algorithm: analysis and applications to neural networks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:13.215435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:13.215435Z digest=sha256:118fa84d453df3901954aa5b5b65901980fab4ec97d82a465f687995e21a0500

Observation 9344d4b5-482f-4669-bde0-23ea357f164f · outbound

This paper cites Modular Duality in Deep Learning.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Modular Duality in Deep Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:13.333343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:13.333343Z digest=sha256:3e4c14e38cda536018d2dcf9f400a87acca1eebdd7887291b957a614355f46a2

Observation fe6db98e-2778-4a76-817f-3e3b684d676b · outbound

This paper cites Old Optimizer, New Norm: An Anthology.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Old Optimizer, New Norm: An Anthology

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:13.459657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:13.459657Z digest=sha256:f073b207ae00d02848cadfeff9ea2bbe609aa7a505e9038869211980dfab1cc3

Observation c5dd4445-3ab7-4f60-abd2-b1ec2f7ce1bf · outbound

This paper cites Scalable Optimization in the Modular Norm.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scalable Optimization in the Modular Norm

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:13.650206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:13.650206Z digest=sha256:a7fab7238214ad699db6d8bfb67a557878112cb3a875e3a78ddab18ba3bcd201

Observation 8f0c4daa-6a9b-42a1-b7c7-9a0712126ce9 · outbound

This paper cites MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:13.858281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:13.858281Z digest=sha256:ec8f98c61d1a70d797fea0f6c08572e6077a45af8e66bc1226442dbe6f7464d9

Observation cb74dc77-c2f2-4d48-8721-054ed188f737 · outbound

This paper cites Towards a principled muon under𝜇p: Ensuring spectral conditions throughout training.arXiv preprint arXiv:2601.01306, 2026.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Towards a principled muon under𝜇p: Ensuring spectral conditions throughout training.arXiv preprint arXiv:2601.01306, 2026

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.041176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.041176Z digest=sha256:eada432b0bee473f0b5c9e69b2e1e9bd82e744d266274be428762bf7b83614f2

Observation 12fc0053-0d22-46f7-a0d1-55c94af31181 · outbound

This paper cites Adamuon: Adaptive muon optimizer.arXiv preprint arXiv:2507.11005, 2025.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Adamuon: Adaptive muon optimizer.arXiv preprint arXiv:2507.11005, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.198277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.198277Z digest=sha256:86f27af5fe91811c03823b3e39e9ec57c8d7dfc8b02f954859453017bad35092

Observation cb874508-99c0-40b3-902f-5457386e0d67 · outbound

This paper cites Normuon: Making muon more efficient and scalable.arXiv preprint arXiv:2510.05491, 2025.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Normuon: Making muon more efficient and scalable.arXiv preprint arXiv:2510.05491, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.357745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.357745Z digest=sha256:8f9a4020a33cd9129e3e3d8802a5e2e1bbcedfb7d86831d861d2dccd89db67f8

Observation 0355528b-3d71-4247-b20c-0a44613f9d01 · outbound

This paper cites Fantastic pretraining optimizers and where to find them ii: From weight decay to hyperball optimization, 12 2025.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Fantastic pretraining optimizers and where to find them ii: From weight decay to hyperball optimization, 12 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.446539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.446539Z digest=sha256:b934c6f0822c018ff241f67769e4a6ffbb439c8f2956e5020189bde2e8d49bdf

Observation 0f139f67-6d74-4086-8116-65a3081c0abb · outbound

This paper cites Controlled llm training on spectral sphere.arXiv preprint arXiv:2601.08393, 2026.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Controlled llm training on spectral sphere.arXiv preprint arXiv:2601.08393, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.558228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.558228Z digest=sha256:f76b2cc013ee3ff82b37d105ec5392df6f88505b6628657bfc250ea84627ee8e

Observation 6ff5efe9-4b21-46db-afbb-c0a149f7f567 · outbound

This paper cites Adam improves muon: Adaptive moment estimation with orthogonalized momentum.arXiv preprint arXiv:2602.17080, 2026.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Adam improves muon: Adaptive moment estimation with orthogonalized momentum.arXiv preprint arXiv:2602.17080, 2026

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.680021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.680021Z digest=sha256:0bbe093963e615f36b03db9df0302d4abf5829cee59fb97666ee7f867068ded3

Observation 90fbaf49-3678-4346-a1d8-701290c5a6de · outbound

This paper cites Manifold constrained steepest descent.arXiv preprint arXiv:2601.21487, 2026.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Manifold constrained steepest descent.arXiv preprint arXiv:2601.21487, 2026

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.814783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.814783Z digest=sha256:61d093f34951bc89698cda0a304b70788312379afa8685f82468598e47fab76f

Observation d5abe931-8959-4242-a595-eb7d7302b4bb · outbound

This paper cites The newton-muon optimizer.arXiv preprint arXiv:2604.01472, 2026.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales The newton-muon optimizer.arXiv preprint arXiv:2604.01472, 2026

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:14.970929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:14.970929Z digest=sha256:e647f3997256e3a46d8ab4e2d8c0e3fa7bbc69b273e92cda3d163aea83ffb905

Observation e2e59c56-13ea-4b79-9f26-ca532089e751 · outbound

This paper cites Mousse: Rectifying the geometry of muon with curvature-aware preconditioning.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Mousse: Rectifying the geometry of muon with curvature-aware preconditioning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.138354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.138354Z digest=sha256:4a872e80028f8b41d86450808e142efde1b4dc0798b2391257780e09743e018d

Observation afdba9c2-f340-4d9f-8127-a4445be7afed · outbound

This paper cites veScale-FSDP: Flexible and High-Performance FSDP at Scale.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales veScale-FSDP: Flexible and High-Performance FSDP at Scale

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.274057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.274057Z digest=sha256:cf5f0cc38f3fe3ca2eae032cc55fac1bdb22e0197e975817918bd0a6d48e59b3

Observation 3a78d3df-ded1-4a62-8f36-e0e22de95125 · outbound

This paper cites Canzona: A unified, asynchronous, and load-balanced framework for distributed matrix-based optimizers.arXiv preprint arXiv:2602.06079, 2026.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Canzona: A unified, asynchronous, and load-balanced framework for distributed matrix-based optimizers.arXiv preprint arXiv:2602.06079, 2026

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.365593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.365593Z digest=sha256:f2fac33e2ff476eebef24ae86ad079b9325740f5e5950f0c1c0c852f61008ed2

Observation f4aeb0c7-eeb3-4584-9e46-caac917c807e · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.479625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.479625Z digest=sha256:7dd9728930f04d326cb6c234ac438ac776f18943ab53151ca74a864be3752260

Observation 79c929b9-1e0b-4fe6-8127-e129ef6366e9 · outbound

This paper cites Qwen3 Technical Report.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Qwen3 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.554716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.554716Z digest=sha256:d4904200c93602a5e2c25b0daed460b1845470075e1fe000d4f1def379f9077c

Observation b852a0db-a10b-4374-b161-84fe4122c301 · outbound

This paper cites Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.647299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.647299Z digest=sha256:7a690e4b5124151e100b936c8b50185e70e515c5abda6c27247af87fa4ac107e

Observation 6afbda3e-b356-4586-b040-b4bc1235ad81 · outbound

This paper cites Scaling laws and compute-optimal training beyond fixed training durations.Advances in Neural Information Processing Systems, 37:76232–76264, 2024.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scaling laws and compute-optimal training beyond fixed training durations.Advances in Neural Information Processing Systems, 37:76232–76264, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.766177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.766177Z digest=sha256:ac2aa61c7a670aec1ce0f73cc8968120f41fd34c50b1d06ed5a78e4731a27c1c

Observation e8a9c316-bb80-4cc9-8888-88f47dfbb6ec · outbound

This paper cites AdamW Weight RMS.https: // kexue.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales AdamW Weight RMS.https: // kexue

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.860483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.860483Z digest=sha256:70775884c7647527004a538a2f70993d6a9a7976cb33051639c0800c66388fea

Observation 2d5699f8-5d94-44e8-91a1-0395eb3fe87f · outbound

This paper cites Stochastic hessian fittings with lie groups.arXiv preprint arXiv:2402.11858, 2024.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Stochastic hessian fittings with lie groups.arXiv preprint arXiv:2402.11858, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:15.968785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:15.968785Z digest=sha256:7bdeac6ad32f02630a1cb27e0e53e4bb083c32351d09793bef9d8c1536cdb053

Observation 878d8978-c3a8-4f8b-9b0a-410cf6543df5 · outbound

This paper cites Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.141224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.141224Z digest=sha256:0ad1f874ab911a45ee0c40616590597eba222a257470de327371fabbd5a55221

Observation 9a1e34b9-f4ff-46f4-ae65-0ee446bb34cf · outbound

This paper cites On the importance of initialization and momentum in deep learning.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales On the importance of initialization and momentum in deep learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.231667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.231667Z digest=sha256:2fae9ccde9180b7cd998801479f26ed121fff9e04a291de9fbffcd6d553da838

Observation 091b542e-1483-4e6a-8cf8-956074cec15d · outbound

This paper cites Incorporating nesterov momentum into adam.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Incorporating nesterov momentum into adam

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.328886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.328886Z digest=sha256:3aa2e3053580fa18317c3ed69acbac215de11f8c7ddc4b8c811774f8e5d03b9c

Observation 03a165c4-03da-46cf-bcfc-fbc1d85880fa · outbound

This paper cites An Empirical Model of Large-Batch Training.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales An Empirical Model of Large-Batch Training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.428609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.428609Z digest=sha256:a99da4d993e57bc44905e6a7de28c5547485da139fa5abe6a11192d9da68047d

Observation 3c457dcc-ced8-4f86-8e6d-43d37c92b5be · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.519251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.519251Z digest=sha256:c7beef7af5703a4f8f52f47f0cc7e6fbb14f0c0b91e2d14c976207975c5834d8

Observation 7afbe3e8-33bc-483f-ab41-728fb29e6a88 · outbound

This paper cites Large batch optimization for deep learning: Training bert in 76 minutes.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Large batch optimization for deep learning: Training bert in 76 minutes

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.597763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.597763Z digest=sha256:32d82ca64aeb01092e2c79ed620f5eacc57aa4396bd6e3c03ca2b39c10032f07

Observation d3a4cf97-6a11-4c93-a5f7-4ba4d6371442 · outbound

This paper cites Train longer, generalize better: closing the generalization gap in large batch training of neural networks.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Train longer, generalize better: closing the generalization gap in large batch training of neural networks

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.690586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.690586Z digest=sha256:9c3ef991faf6f448726d40ea13546691a5dcfaeba376bada135ec0f910b4824e

Observation f33cc214-6673-4722-a1a7-a122be0ab08c · outbound

This paper cites Large Batch Training of Convolutional Networks.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Large Batch Training of Convolutional Networks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.783035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.783035Z digest=sha256:d5ef5a526f83a1fa239f091bd520af8eac882d1d0a59067ff4e2a62f67f8a202

Observation ba1434f9-71b8-4ebd-93fc-f5b258b7ca97 · outbound

This paper cites One weird trick for parallelizing convolutional neural networks.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales One weird trick for parallelizing convolutional neural networks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:16.934273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:16.934273Z digest=sha256:4511edb38ec1a44302f0be1b9464bc05cb8a84535289ad60e74603af03c6e074

Observation b94154f1-d860-4bd8-a50c-b1197948a0dd · outbound

This paper cites Better & faster large language models via multi-token prediction.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Better & faster large language models via multi-token prediction

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.032023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.032023Z digest=sha256:720b0f9e229ec6149cb6a06750de41ddadc0dda12944e9b1978c5921447671b8

Observation 9ca1e866-9879-44e1-8c69-1d70f4252a2c · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.122273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.122273Z digest=sha256:18692a0e25ff20d51e22afb0ba3d6c59a3c8eb07a1eba7994bacef171efd563d

Observation 0ca474ff-08db-4668-8d38-4b65e51d4f23 · outbound

This paper cites Scaling Exponents Across Parameterizations and Optimizers.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scaling Exponents Across Parameterizations and Optimizers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.238514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.238514Z digest=sha256:caf6212f8cb392132cf1227b6e8b4a2b957a83215bc5b456fe974dfdaec18408

Observation 7736d6ae-3bda-448a-8cda-86d9d4c88fce · outbound

This paper cites Small-scale proxies for large-scale Transformer training instabilities.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Small-scale proxies for large-scale Transformer training instabilities

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.325825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.325825Z digest=sha256:07bb726b3fe7963c1ecf59126e2dcc834e1f67aa5ac2971309208109f3a68f38

Observation cb5521bc-be78-4a00-bd29-7497c9c6c2de · outbound

This paper cites Tensor Programs IVb: Adaptive Optimization in the Infinite-Width Limit.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Tensor Programs IVb: Adaptive Optimization in the Infinite-Width Limit

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.441355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.441355Z digest=sha256:639e1771ca5c7b2166d2c6fdef35d40d298d5694c660e4d11382d5df10dfaac2

Observation a0cd8fc0-c97a-492d-b98b-10299d4474ff · outbound

This paper cites an unresolved cited work.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.502177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.502177Z digest=sha256:81e5874b056e0d49a9991817f04586dccab3a87f4e1be209be794884be3b34e4

Observation f0f33c12-0406-4afd-9e1d-95246b8dcd82 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.620629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.620629Z digest=sha256:5fac7e0d73a519de9f743426fb8630ca68037b186998fa055d0ec85d763e656b

Observation a247749e-4341-4344-9723-5365be0dad2a · outbound

This paper cites Stabilizing Native Low-Rank LLM Pretraining.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Stabilizing Native Low-Rank LLM Pretraining

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.674394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.674394Z digest=sha256:d265c3d81db435a9bcd3a3d7471ad60b108c0c546b5342b2cb8e5d136c95b04f

Observation 82daf934-0776-4118-a6b7-f616e56bda32 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.759660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.759660Z digest=sha256:84305978bb56ccf6f8955ec3e1e276acfba766b5a28c93cac9e9ed673bd88be4

Observation 92ff9c0d-a7ff-4c5c-9f4c-26e6bab97084 · outbound

This paper cites Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.820219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.820219Z digest=sha256:5f9aff64ee77c02bdc6cc97d2b88d6f7d23693b316630150c45db9b5b0614587

Observation 30005dc9-09e8-4498-b90e-1a5701d11666 · outbound

This paper cites On the role of batch size in stochastic conditional gradient methods.arXiv preprint arXiv:2603.21191, 2026.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales On the role of batch size in stochastic conditional gradient methods.arXiv preprint arXiv:2603.21191, 2026

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.872461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.872461Z digest=sha256:ea1be271e3b4fd32bc44c4d73c55be5b59f470ef7237aa5cc752786603cb97bd

Observation 806beb69-1116-4fc1-936b-d60b22816623 · outbound

This paper cites Spectral Scaling Laws of Muon.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Spectral Scaling Laws of Muon

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.932930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.932930Z digest=sha256:0bdb716d55ea82e300f1d348e9f233dba77bf21c814abb9c3f96c18d208e22d5

Observation de41cb6e-f64a-4414-beba-cdea8e55294d · outbound

This paper cites Dissecting adam: The sign, magnitude and variance of stochastic gradients.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Dissecting adam: The sign, magnitude and variance of stochastic gradients

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:18.077028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:18.077028Z digest=sha256:376d3a19b40988f6acd8032a6155097137af41a0d0ec3b466293e773def66b91

Observation e96956c5-1e58-4b1b-9a58-788abb95e819 · outbound

This paper cites A New Perspective on Shampoo's Preconditioner.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales A New Perspective on Shampoo's Preconditioner

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:18.187657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:18.187657Z digest=sha256:cc0735c473d2273b7d20d68c9dc55977ce76f7af88f46a58e7b326b746b31037

Observation cc3f4a9d-f5f0-4d3c-9c6d-a1906fda7d6a · outbound

This paper cites Recipes for Pre-training LLMs with MXFP8.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Recipes for Pre-training LLMs with MXFP8

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:18.277482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:18.277482Z digest=sha256:13a3dda6185d2a566b1a3fc5c852eca4a8b1b31ff4c684aa5c799f30b0beca3b

Observation a77b5641-96a2-4418-9b83-a5f60b0352f6 · outbound

This paper cites Symbolic discovery of optimization algorithms.Advances in neural information processing systems, 36:49205–49233, 2023.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Symbolic discovery of optimization algorithms.Advances in neural information processing systems, 36:49205–49233, 2023

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:18.400599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:18.400599Z digest=sha256:45d54ab78a9a463b24b1aeae09cce730f502d304ac9ce75b0b4096d4f9f9543b

Pith citing papers

Observation d2249c7c-3bcb-425f-b743-688c4ff23520 · inbound

When Does Muon Help Agentic Reinforcement Learning? cites this paper.

When Does Muon Help Agentic Reinforcement Learning? SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T21:13:38.825691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:13:38.825691Z digest=sha256:b6f15e26cd5059377aefb0d538973b91cf5243527f50059ffa0d38e6f98c4e10

Observation f0a3e926-f1ad-4016-8abd-f1b2aeac5ec5 · inbound

When Does Muon Help Agentic Reinforcement Learning? cites this paper.

When Does Muon Help Agentic Reinforcement Learning? SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:37.877565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:37.877565Z digest=sha256:1654cd94146e85ea07c5a9745c3bbe93317a7364c6c26584b169a5aee203c0b3