Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T06:45:18.400599Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 2 inbound Pith citation observations for arXiv:2607.20548.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T06:45:18.400599Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T04:19:37.877565Z
A source-named dated measurement, never combined with another source.
Source: cited_works
74 of 74 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 24b3c119-0f44-45f3-b04d-4950669c0edd · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Adam: A Method for Stochastic Optimization
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04c785ae-fa76-4867-a04f-12418d339f05 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Decoupled Weight Decay Regularization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b5f3881-521d-4e8a-adc2-e3e225d3ac3d · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 583768f7-0d65-4b9d-8ab3-6787c2ab6141 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales LaProp: Separating Momentum and Adaptivity in Adam
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6fa9006-bd39-4a99-92b5-100a5156d058 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Fast approximate natural gradient descent in a kronecker factored eigenbasis.Advances in neural information processing systems, 31, 2018
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 960149df-9156-47f5-a4a5-d4dbc1eb8b57 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Optimizing neural networks with kronecker-factored approxi- mate curvature
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbe0683b-1f7c-4898-ad95-351fa82d782a · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales A progressive batching l-bfgs method for machine learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6471271c-25a1-4623-bd6d-4e239ae7ec23 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales SOAP: Improving and Stabilizing Shampoo using Adam
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4a5c059-ffe8-43a6-9b96-6a3883b0f0a0 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Purifying shampoo: Investigating shampoo’s heuristics by decomposing its preconditioner
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b60bb62-a945-4e99-9272-47654e70fdfe · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Understanding and improving the shampoo optimizer via kullback-leibler minimization.arXiv e-prints, pages arXiv–2509, 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36bb432b-31dc-4628-9626-3b30782906b5 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Training Deep Learning Models with Norm-Constrained LMOs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e615af54-4239-4f12-a5f5-0e2681fa0570 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Shampoo: Preconditioned stochastic tensor optimization
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a69c054-a9b9-485f-b4ce-e7bc2b0c8908 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Muon is Scalable for LLM Training
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7910fb49-bf75-4774-be98-e01007f6edad · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Practical Efficiency of Muon for Pretraining
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b3c3efd-595e-4bda-9410-8e66ec6475e9 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d6bc1ba-0e54-491d-8fae-534617199d7d · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Muon: An optimizer for hidden layers in neural networks, 2024b.URL https://kellerjordan
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6ed0731-c1ab-403f-ab7f-6be77ca1c507 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales DeepSeek-V3 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb5d772-0a91-4b0e-a2a1-11c3498113fa · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae0008f2-8e98-464f-b4ca-30bcf5ad6dec · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e33c0488-b8fd-4afe-bbc3-ae038b6a67ba · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales On the sdes and scaling rules for adaptive gradient algorithms.Advances in Neural Information Processing Systems, 35:7697–7711, 2022
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d6cd56e-b842-46b8-9f51-76f9f692011d · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c8cb48c-1b66-414f-9d7e-629157238217 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Zero: Memory optimiza- tions toward training trillion parameter models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a20dc7e-1b59-41eb-9a72-b22682ab6864 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scalable Second Order Optimization for Deep Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 143f7807-a46c-4659-a567-62ad324b12cd · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d724d58-fcc3-450b-bf1c-691b1d8f3b2a · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Precondi- tioned spectral descent for deep learning.Advances in neural information processing systems, 28, 2015
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 196eabd2-3d09-436e-8ffa-c5989a880cfa · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Stochastic spectral descent for restricted boltzmann machines
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ddaba03-2b0d-494e-8568-8e5adb2e9b38 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Stochastic spectral descent for discrete graphical models.IEEE Journal of Selected Topics in Signal Processing, 10(2):296–311, 2015
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f0d3422-b192-46fb-8d8d-e6784f27a2cc · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales The duality structure gradient descent algorithm: analysis and applications to neural networks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9344d4b5-482f-4669-bde0-23ea357f164f · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Modular Duality in Deep Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe6db98e-2778-4a76-817f-3e3b684d676b · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Old Optimizer, New Norm: An Anthology
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5dd4445-3ab7-4f60-abd2-b1ec2f7ce1bf · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scalable Optimization in the Modular Norm
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f0c4daa-6a9b-42a1-b7c7-9a0712126ce9 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb74dc77-c2f2-4d48-8721-054ed188f737 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Towards a principled muon under𝜇p: Ensuring spectral conditions throughout training.arXiv preprint arXiv:2601.01306, 2026
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12fc0053-0d22-46f7-a0d1-55c94af31181 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Adamuon: Adaptive muon optimizer.arXiv preprint arXiv:2507.11005, 2025
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb874508-99c0-40b3-902f-5457386e0d67 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Normuon: Making muon more efficient and scalable.arXiv preprint arXiv:2510.05491, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0355528b-3d71-4247-b20c-0a44613f9d01 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Fantastic pretraining optimizers and where to find them ii: From weight decay to hyperball optimization, 12 2025
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f139f67-6d74-4086-8116-65a3081c0abb · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Controlled llm training on spectral sphere.arXiv preprint arXiv:2601.08393, 2026
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ff5efe9-4b21-46db-afbb-c0a149f7f567 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Adam improves muon: Adaptive moment estimation with orthogonalized momentum.arXiv preprint arXiv:2602.17080, 2026
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90fbaf49-3678-4346-a1d8-701290c5a6de · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Manifold constrained steepest descent.arXiv preprint arXiv:2601.21487, 2026
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5abe931-8959-4242-a595-eb7d7302b4bb · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales The newton-muon optimizer.arXiv preprint arXiv:2604.01472, 2026
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2e59c56-13ea-4b79-9f26-ca532089e751 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Mousse: Rectifying the geometry of muon with curvature-aware preconditioning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afdba9c2-f340-4d9f-8127-a4445be7afed · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales veScale-FSDP: Flexible and High-Performance FSDP at Scale
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a78d3df-ded1-4a62-8f36-e0e22de95125 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Canzona: A unified, asynchronous, and load-balanced framework for distributed matrix-based optimizers.arXiv preprint arXiv:2602.06079, 2026
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4aeb0c7-eeb3-4584-9e46-caac917c807e · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79c929b9-1e0b-4fe6-8127-e129ef6366e9 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Qwen3 Technical Report
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b852a0db-a10b-4374-b161-84fe4122c301 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6afbda3e-b356-4586-b040-b4bc1235ad81 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scaling laws and compute-optimal training beyond fixed training durations.Advances in Neural Information Processing Systems, 37:76232–76264, 2024
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8a9c316-bb80-4cc9-8888-88f47dfbb6ec · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales AdamW Weight RMS.https: // kexue
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d5699f8-5d94-44e8-91a1-0395eb3fe87f · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Stochastic hessian fittings with lie groups.arXiv preprint arXiv:2402.11858, 2024
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 878d8978-c3a8-4f8b-9b0a-410cf6543df5 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a1e34b9-f4ff-46f4-ae65-0ee446bb34cf · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales On the importance of initialization and momentum in deep learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 091b542e-1483-4e6a-8cf8-956074cec15d · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Incorporating nesterov momentum into adam
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03a165c4-03da-46cf-bcfc-fbc1d85880fa · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales An Empirical Model of Large-Batch Training
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c457dcc-ced8-4f86-8e6d-43d37c92b5be · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales MiniMax-01: Scaling Foundation Models with Lightning Attention
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7afbe3e8-33bc-483f-ab41-728fb29e6a88 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Large batch optimization for deep learning: Training bert in 76 minutes
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3a4cf97-6a11-4c93-a5f7-4ba4d6371442 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f33cc214-6673-4722-a1a7-a122be0ab08c · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Large Batch Training of Convolutional Networks
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba1434f9-71b8-4ebd-93fc-f5b258b7ca97 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales One weird trick for parallelizing convolutional neural networks
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b94154f1-d860-4bd8-a50c-b1197948a0dd · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Better & faster large language models via multi-token prediction
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ca1e866-9879-44e1-8c69-1d70f4252a2c · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ca474ff-08db-4668-8d38-4b65e51d4f23 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scaling Exponents Across Parameterizations and Optimizers
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7736d6ae-3bda-448a-8cda-86d9d4c88fce · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Small-scale proxies for large-scale Transformer training instabilities
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb5521bc-be78-4a00-bd29-7497c9c6c2de · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Tensor Programs IVb: Adaptive Optimization in the Infinite-Width Limit
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0cd8fc0-c97a-492d-b98b-10299d4474ff · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Unresolved cited work
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f33c12-0406-4afd-9e1d-95246b8dcd82 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a247749e-4341-4344-9723-5365be0dad2a · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Stabilizing Native Low-Rank LLM Pretraining
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82daf934-0776-4118-a6b7-f616e56bda32 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92ff9c0d-a7ff-4c5c-9f4c-26e6bab97084 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30005dc9-09e8-4498-b90e-1a5701d11666 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales On the role of batch size in stochastic conditional gradient methods.arXiv preprint arXiv:2603.21191, 2026
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 806beb69-1116-4fc1-936b-d60b22816623 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Spectral Scaling Laws of Muon
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de41cb6e-f64a-4414-beba-cdea8e55294d · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Dissecting adam: The sign, magnitude and variance of stochastic gradients
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e96956c5-1e58-4b1b-9a58-788abb95e819 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales A New Perspective on Shampoo's Preconditioner
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc3f4a9d-f5f0-4d3c-9c6d-a1906fda7d6a · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Recipes for Pre-training LLMs with MXFP8
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a77b5641-96a2-4418-9b83-a5f60b0352f6 · outbound
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Symbolic discovery of optimization algorithms.Advances in neural information processing systems, 36:49205–49233, 2023
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2249c7c-3bcb-425f-b743-688c4ff23520 · inbound
When Does Muon Help Agentic Reinforcement Learning? SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0a3e926-f1ad-4016-8abd-f1b2aeac5ec5 · inbound
When Does Muon Help Agentic Reinforcement Learning? SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.