Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 100 of 114 outbound references and 100 inbound Pith citation observations for arXiv:2502.16982.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T21:51:05.313853Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T21:57:38.615639Z
100 of 114 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fb514fb3-d163-4219-bfbc-d7c18563d097 · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 886ce2dd-d23f-4fb6-a33b-514a8a83d9c3 · outbound
Muon is Scalable for LLM Training 2017 , eprint=
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d3b8bed8-e591-44bb-ba0f-544b61312d4e · outbound
Muon is Scalable for LLM Training The effective rank: A measure of effective dimensionality , year=
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d64caa3a-8fa7-4841-9acf-f4b88581c6e3 · outbound
Muon is Scalable for LLM Training Brown and David Botstein , title =
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4b8188d2-9275-4586-a29b-376a6ae44505 · outbound
Muon is Scalable for LLM Training 2023 , eprint=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aef93166-9bd0-4f18-8085-1907ad5e7baa · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7a941ef3-59e7-4c5c-aa90-6df531766d9b · outbound
Muon is Scalable for LLM Training Advances in Neural Information Processing Systems , volume=
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2beb3b2b-a8b7-4807-96ee-9f12ce2719da · outbound
Muon is Scalable for LLM Training Advances in Neural Information Processing Systems , volume=
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 56f944c8-504e-425a-ba49-eec402a91bac · outbound
Muon is Scalable for LLM Training Advances in Neural Information Processing Systems , volume=
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3038a8e9-a940-44eb-aa91-1161e362c1e7 · outbound
Muon is Scalable for LLM Training Advances in Neural Information Processing Systems , volume=
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b219efbb-de3e-40cd-9cd7-17550a6aeb17 · outbound
Muon is Scalable for LLM Training YaRN: Efficient Context Window Extension of Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5432d18c-db72-404c-b234-3c7b78834898 · outbound
Muon is Scalable for LLM Training MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation , year=
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation efe04cd2-d40e-49cc-9c53-0a2e03168609 · outbound
Muon is Scalable for LLM Training IEEE transactions on Systems Science and Cybernetics , volume=
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b655815f-1d16-4b1d-9f0e-61cbff7a30fd · outbound
Muon is Scalable for LLM Training International conference on computers and games , pages=
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 19db4135-72eb-4705-bff6-7954e08497f8 · outbound
Muon is Scalable for LLM Training European conference on machine learning , pages=
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ba6d9d32-e677-40a0-8f37-023d298fe290 · outbound
Muon is Scalable for LLM Training Advances in Neural Information Processing Systems , volume=
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b77600a2-1578-4996-92e0-53f2ea8bc468 · outbound
Muon is Scalable for LLM Training Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 454dbe6f-7d6d-45cd-9182-ee68d466bed7 · outbound
Muon is Scalable for LLM Training Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 25fbcfea-0db0-43f8-935f-5fa23cb5b2e9 · outbound
Muon is Scalable for LLM Training Advances in neural information processing systems , volume=
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a97ed35f-381e-45e4-853b-52b45d252804 · outbound
Muon is Scalable for LLM Training Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee4b0765-b7e8-4e73-bcb9-54cbd5bbbb37 · outbound
Muon is Scalable for LLM Training Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c08fb611-51ed-4a05-81fc-0e9df8b6a122 · outbound
Muon is Scalable for LLM Training International Conference on Machine Learning , pages=
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f8f6113d-ac82-4ac8-b223-d044059720c0 · outbound
Muon is Scalable for LLM Training Proceedings of the 28th International Joint Conference on Artificial Intelligence , pages=
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ce92ea2b-0577-448c-bb91-92038e927280 · outbound
Muon is Scalable for LLM Training Mirror Descent Policy Optimization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bff0a9b5-3de0-4a08-b0ab-e8d6a307c037 · outbound
Muon is Scalable for LLM Training Advances in neural information processing systems , volume=
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9584fd9d-7dd7-4646-ac66-67c7309c143e · outbound
Muon is Scalable for LLM Training Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c572c383-e3a0-4b64-81aa-9aa0b40dcbae · outbound
Muon is Scalable for LLM Training Neurocomputing , volume=
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2ecc027d-356b-4301-9e74-1e1d66aed3b6 · outbound
Muon is Scalable for LLM Training 2024 , url=
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 745d6af4-01ba-4076-a8f9-7ee54fc918f8 · outbound
Muon is Scalable for LLM Training 2020 , eprint=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 47e734c1-68d0-4c6d-b134-557405cd62ac · outbound
Muon is Scalable for LLM Training Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 218fdfe9-dd08-4663-92fb-e71f40e9b444 · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 51fe9497-b28b-405b-bd37-7ee887bcc684 · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0f347ae0-d2b3-4a06-8852-81fadffebc0f · outbound
Muon is Scalable for LLM Training Attention is All you Need , url =
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 47715c98-443d-4357-b9aa-1935978d4d4e · outbound
Muon is Scalable for LLM Training ArXiv , year=
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3884b18c-2c95-41e0-b61c-e6016444b9a3 · outbound
Muon is Scalable for LLM Training North American Chapter of the Association for Computational Linguistics , year=
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3ef53054-c7c7-4b06-bebe-003e2e18272a · outbound
Muon is Scalable for LLM Training ArXiv , year=
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9c4206c1-73ef-4da8-8d4f-1125511f16b5 · outbound
Muon is Scalable for LLM Training 2024 , journal=
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 42a8f0af-be7a-4117-b39b-4648b0f9b409 · outbound
Muon is Scalable for LLM Training International Conference on Computational Linguistics , year=
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ca864ecf-6c9e-45f8-b69b-bf1aea58544f · outbound
Muon is Scalable for LLM Training ArXiv , year=
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e5abf61a-7a9d-43ce-9284-3e19c05b52a9 · outbound
Muon is Scalable for LLM Training ArXiv , year=
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 259b7364-c77d-4c74-a145-98f984338d56 · outbound
Muon is Scalable for LLM Training ArXiv , year=
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1a7366c6-3d29-474c-93af-70d0191b028e · outbound
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b18c9154-cbfe-4a1c-b7dc-8debf06b2534 · outbound
Muon is Scalable for LLM Training Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 46e1fea2-eee6-4e06-a972-6b49dd525922 · outbound
Muon is Scalable for LLM Training Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 837a51a9-5e67-4a82-bdc8-4dab344e566f · outbound
Muon is Scalable for LLM Training MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c6e22c65-b264-4d57-9f4c-6d554820de07 · outbound
Muon is Scalable for LLM Training Bag of Tricks for Efficient Text Classification
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c0fe3862-babb-4845-928e-3732197b3774 · outbound
Muon is Scalable for LLM Training M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 49884703-9a55-46ec-ace3-0babb8b1dc27 · outbound
Muon is Scalable for LLM Training The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2b92e572-5f33-4239-b90a-7faa28f0ed05 · outbound
Muon is Scalable for LLM Training DataComp-LM: In search of the next generation of training sets for language models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 69ad38c9-2a10-47e0-8daf-ff886b1b8b93 · outbound
Muon is Scalable for LLM Training OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4817a92b-4d58-4eb8-8dbf-5b32f884855f · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 71cd9f16-bac3-42cc-bad7-e63b2ab8adbf · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3cfd9307-426a-4237-a64b-10268a5adc5c · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5309a0e4-bbe1-46cf-b595-07a34d1ce3b0 · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 863a3ad5-6f54-4f78-b28c-fb028a5e851e · outbound
Muon is Scalable for LLM Training MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b11cfc50-8406-40d3-a048-37ecaeeb235a · outbound
Muon is Scalable for LLM Training Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6a6ce351-a034-489c-ab0e-dcb263f03aba · outbound
Muon is Scalable for LLM Training Reinforced Self-Training (ReST) for Language Modeling
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 24648568-03a9-427e-8527-ae940339d12f · outbound
Muon is Scalable for LLM Training nature , volume=
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d226672c-2fbc-4a49-953c-54051a936ea9 · outbound
Muon is Scalable for LLM Training nature , volume=
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d1e62c1-c4bc-451e-833a-0b6314859c9a · outbound
Muon is Scalable for LLM Training Dota 2 with Large Scale Deep Reinforcement Learning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0cb61bc2-57fb-430f-8858-cf56bb15c7d1 · outbound
Muon is Scalable for LLM Training Advances in neural information processing systems , volume=
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a4982843-a78a-4e39-a2db-6254f9a8c684 · outbound
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6d2d216c-a8ec-45ad-85e5-abcf105a04e0 · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 43dbd1b0-49b2-45ac-b3d2-78b5687f14df · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4ba30ccc-ff60-4982-9a27-e58d5ab7f6d8 · outbound
Muon is Scalable for LLM Training Advances in Neural Information Processing Systems , volume=
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2a9b8a1f-0e12-46e7-a21d-d42e47d42237 · outbound
Muon is Scalable for LLM Training Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8d3ec19a-861c-4a7f-8993-442add9c6c59 · outbound
Muon is Scalable for LLM Training 2020 , eprint=
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0edcdcdd-c3fb-4829-bf26-4dbb9eac454f · outbound
Muon is Scalable for LLM Training 2022 , eprint=
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9247de9c-dbca-4cfc-9353-d3ebd6eda934 · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cf9cb028-ce10-4708-b546-9eac009236a6 · outbound
Muon is Scalable for LLM Training 2023 , eprint=
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0b73bebb-7e9d-4dae-9ec4-25d3d96d9260 · outbound
Muon is Scalable for LLM Training General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c04232aa-f25a-4a92-8760-f1cb83beb0a3 · outbound
Muon is Scalable for LLM Training 2021 , eprint=
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation be3f4f2f-1e36-4708-a436-c187dfd25a42 · outbound
Muon is Scalable for LLM Training International Conference on Learning Representations , year=
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d92228bd-9519-474b-a53d-5786dc40b4b8 · outbound
Muon is Scalable for LLM Training What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8185dbc5-4548-4b35-b7aa-1c461762f290 · outbound
Muon is Scalable for LLM Training From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e9f0f70c-a7dd-4908-904b-d89ee30a0c47 · outbound
Muon is Scalable for LLM Training Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 65ca70d8-8452-4428-aaca-1b1968a32574 · outbound
Muon is Scalable for LLM Training 2024 , url =
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2cc933f4-5fa3-481e-b972-5e83562bf20b · outbound
Muon is Scalable for LLM Training International Conference on Learning Representations , year=
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 62acc690-6490-460c-aef6-6368b83111e4 · outbound
Muon is Scalable for LLM Training 2024 , month = Oct, url =
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eabc162d-b034-4ac0-84f3-174f869b1f5d · outbound
Muon is Scalable for LLM Training 2024 , url =
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1cc1be1f-15d2-44e0-8557-bb616d4d9d4e · outbound
Muon is Scalable for LLM Training Kingma and Jimmy Ba , editor =
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d1790b3d-761c-4076-b2bf-c736a36bb187 · outbound
Muon is Scalable for LLM Training The Twelfth International Conference on Learning Representations , year=
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eab95519-6437-4218-91a6-f8ce0b193ccf · outbound
Muon is Scalable for LLM Training Kakade , booktitle=
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7164ae32-d75e-4f41-90be-c777bf4b3a31 · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6e9ae587-8672-4b42-8429-ede093b82551 · outbound
Muon is Scalable for LLM Training Generalized Slow Roll for Tensors
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3f2a0681-8893-47c9-9788-51b10763fb44 · outbound
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3042fe40-f88e-48de-9612-f6d904d6ba39 · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5beb1a16-9548-41de-a50a-a8df575e0e2b · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ca70984-6f90-45b4-848d-da4d5cef47e0 · outbound
Muon is Scalable for LLM Training 2024 , email =
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 267dfd4b-1a3a-4ed6-9c47-00928948093e · outbound
Muon is Scalable for LLM Training Unresolved cited work
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9d90a456-eb1b-4a75-bc4c-878908725abb · outbound
Muon is Scalable for LLM Training 2025 , eprint=
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c8c2a863-01b1-4057-afbf-66ab9330d08f · outbound
Muon is Scalable for LLM Training 2025 , url=
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 085eef7e-99e0-4653-a839-7f65fa8a4361 · outbound
Muon is Scalable for LLM Training DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 19daf070-9e86-4e00-9fc2-cad4bbef2b63 · outbound
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0099226d-dcb8-4b66-915a-790e5a06eba1 · outbound
Muon is Scalable for LLM Training 2019 , eprint=
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 28e75e17-90e1-45b1-8283-28ba35d04394 · outbound
Muon is Scalable for LLM Training Gemma 2: Improving Open Language Models at a Practical Size
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 239e3013-ac07-4a2c-aadc-ebace1e1b24e · outbound
Muon is Scalable for LLM Training 2024 , month=
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e9cfec47-2c93-48e6-a9e4-c66e6aa25b00 · outbound
Muon is Scalable for LLM Training 2021 , eprint=
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b078db38-f633-4b5e-a137-46673d6bfed2 · outbound
Muon is Scalable for LLM Training 2024 , eprint=
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 53530e26-5765-491b-91d6-199387f19311 · outbound
Muon is Scalable for LLM Training 2022 , eprint=
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1c8d5c47-a6a5-40c7-adad-c9db58cfed8d · inbound
Eliciting Latent Predictions from Transformers with the Tuned Lens Muon is Scalable for LLM Training
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bd0c1398-2ded-4a28-bf0f-31c707348370 · inbound
GWT: Scalable Optimizer State Compression for Large Language Model Training Muon is Scalable for LLM Training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1abb37ca-9b1e-4a01-a0ab-727623b95114 · inbound
Training Deep Learning Models with Norm-Constrained LMOs Muon is Scalable for LLM Training
Reference 195
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bb2dc52c-4fe7-4999-8f96-38a57c4de888 · inbound
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning Muon is Scalable for LLM Training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 779f8ea5-0cab-4f2c-abc8-edb41048c54d · inbound
Kimi-Audio Technical Report Muon is Scalable for LLM Training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 028443af-ccb7-41c9-b7b5-5de65298e426 · inbound
On the Convergence Analysis of Muon Muon is Scalable for LLM Training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ac2c9a3-d5ce-4a81-bdec-76d82c467ae7 · inbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Muon is Scalable for LLM Training
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 23c85ed5-aa98-44c1-b5fb-1e0a6e8d7cc8 · inbound
Kimi K2: Open Agentic Intelligence Muon is Scalable for LLM Training
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dd200900-e693-451a-8f30-4c6a5327215a · inbound
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models Muon is Scalable for LLM Training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d84da147-e5ab-4ced-ac32-2eb0ea9b1545 · inbound
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? Muon is Scalable for LLM Training
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8106248e-29ca-4ddb-a302-3578382c0ab9 · inbound
Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training Muon is Scalable for LLM Training
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f20acbef-4f18-4bf9-95e6-66ca5b658d1c · inbound
On the Convergence of Muon and Beyond Muon is Scalable for LLM Training
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5dbd9f93-54bc-45a2-8270-a0b344d78fe4 · inbound
LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers Muon is Scalable for LLM Training
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 53da14d2-3c98-4737-a976-d6c49bdd4168 · inbound
Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization Muon is Scalable for LLM Training
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b202a901-a872-4140-81f3-f3b9e9cc3410 · inbound
Evolutionary Profiles for Protein Fitness Prediction Muon is Scalable for LLM Training
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 51d4eab7-67fa-46b7-8794-9dbc0d13dfbf · inbound
Kimi Linear: An Expressive, Efficient Attention Architecture Muon is Scalable for LLM Training
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bd615397-34e6-4f3d-b93b-624e56dccaee · inbound
Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates Muon is Scalable for LLM Training
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d63cf5f7-9c32-4fc6-8221-433fc555cfa4 · inbound
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning Muon is Scalable for LLM Training
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 77f8b84a-73f9-4ea2-bf6f-7b336a2505e0 · inbound
Turbo-Muon: Almost-Orthogonal Pre-Conditioning for Fast Muon Updates Muon is Scalable for LLM Training
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3bfe19c-93fa-486a-bef9-452a91363307 · inbound
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization Muon is Scalable for LLM Training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3176bd42-ff3e-405f-ac86-649b7b23ba6e · inbound
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization Muon is Scalable for LLM Training
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d33cb1ce-5500-497c-80be-7f05f59d008a · inbound
KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual Matrices Muon is Scalable for LLM Training
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dafce45-9263-48c8-9906-d5bac2f8e6ed · inbound
Kimi K2.5: Visual Agentic Intelligence Muon is Scalable for LLM Training
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation efd02246-9016-4e27-94dc-59067dfd3c17 · inbound
SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning Muon is Scalable for LLM Training
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc4694a4-06e3-4f69-9336-3ca7353ecb1b · inbound
Muon in Associative Memory Learning: Training Dynamics and Scaling Laws Muon is Scalable for LLM Training
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82d1aae8-36a2-4341-9f32-4e429bb9a959 · inbound
Decoupling Variance and Scale-Invariant Updates in Adaptive Gradient Descent for Unified Vector and Matrix Optimization Muon is Scalable for LLM Training
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8bbbba-21aa-46eb-9ad5-cb83fd061767 · inbound
The Implicit Bias of Steepest Descent with Mini-batch Stochastic Gradient Muon is Scalable for LLM Training
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef59ff3e-048b-4fa1-b43d-0c1cabffa28b · inbound
MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training Muon is Scalable for LLM Training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4c27bc63-3c62-43be-8c8e-ed9c3682717f · inbound
Spectral Condition for $\mu$P under Width-Depth Scaling Muon is Scalable for LLM Training
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9d93a0da-45d1-4a23-be1a-c0325b060e90 · inbound
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 613478ff-c5db-4e36-84fc-b923408a9f61 · inbound
GLENN: Neural network-enhanced computation of Ginzburg-Landau energy minimizers Muon is Scalable for LLM Training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c2d0b97b-bb88-456b-a333-5b8b7d0bf01c · inbound
RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization Muon is Scalable for LLM Training
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 515d4ac9-d67d-4f32-a8e1-883ce029ccd2 · inbound
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory Muon is Scalable for LLM Training
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7a6a9290-0319-47ea-8ab9-8e7ade9fba11 · inbound
MuonEq: Balancing Before Orthogonalization with Lightweight Equilibration Muon is Scalable for LLM Training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b1a5ce80-7005-4e1a-9c30-b216974c0f81 · inbound
Optimal Projection-Free Adaptive SGD for Matrix Optimization Muon is Scalable for LLM Training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 40a15ea3-c03a-4718-9a26-b2a4e97e996c · inbound
Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning Muon is Scalable for LLM Training
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13fd3e5e-6e3a-4c2d-8109-fc4bf2e882e1 · inbound
A Muon-Accelerated Algorithm for Low Separation Rank Tensor Generalized Linear Models Muon is Scalable for LLM Training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4e2e3e8a-9e20-4273-a49b-7d7e5ac2f63b · inbound
Fast Spatial Memory with Elastic Test-Time Training Muon is Scalable for LLM Training
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3c1893ac-f410-4a89-9ac8-1be1d80cfc0d · inbound
PRAGMA: Revolut Foundation Model Muon is Scalable for LLM Training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 10dd01c6-e6fe-46d2-8ad5-ba3210215024 · inbound
Communication-Efficient Gluon in Federated Learning Muon is Scalable for LLM Training
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4486f7b9-e735-4ca7-acae-986cb115f93a · inbound
ResBM: Residual Bottleneck Models for Low-Bandwidth Pipeline Parallelism Muon is Scalable for LLM Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b17e7849-5a25-49a3-82d1-0dd5458a9638 · inbound
Benchmarking Optimizers for MLPs in Tabular Deep Learning Muon is Scalable for LLM Training
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ebe768c2-dc29-43a0-ab34-d5be14132dbe · inbound
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation Muon is Scalable for LLM Training
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f6ccf5c5-5801-4fec-99cc-a2cdc32107ad · inbound
In-context modeling as a retrain-free paradigm for foundation models in computational science Muon is Scalable for LLM Training
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f4b85007-3a3c-4b57-aba4-e5f7eef53e6e · inbound
SUDA-Muon: Structural Design Principles and Boundaries for Fully Decentralized Muon Muon is Scalable for LLM Training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 52f23ffd-fd96-4273-ac76-49e58547b722 · inbound
SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs Muon is Scalable for LLM Training
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f0339a5b-2a53-436f-8ba0-958d64dc4e3f · inbound
Model Merging: Foundations and Algorithms Muon is Scalable for LLM Training
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 72d92fad-fd70-4650-b4bc-2c7a29ee42b5 · inbound
DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs Muon is Scalable for LLM Training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7345c383-cb32-498a-bdc1-381d3660127f · inbound
Nora: Normalized Orthogonal Row Alignment for Scalable Matrix Optimizer Muon is Scalable for LLM Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a8c2ae73-8b1d-4dcc-8fc9-bdcec59f3105 · inbound
Budget-aware Auto Optimizer Configurator Muon is Scalable for LLM Training
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2c3cf7ab-8c86-423c-a0d9-653851ae2b25 · inbound
Reference 199
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation adb25634-3da3-41c4-82f9-6ccb3d6c1d4b · inbound
Autoregressive One-Step Generative Modeling for Dynamical System Forecasting Muon is Scalable for LLM Training
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation def732c4-3ec1-4ca9-ba36-44d39dbae4b7 · inbound
Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio Muon is Scalable for LLM Training
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f1bafc29-79a4-4fea-83ef-1a532689648b · inbound
MDN: Parallelizing Stepwise Momentum for Delta Linear Attention Muon is Scalable for LLM Training
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2eeae874-359d-4a75-986c-9e046e734539 · inbound
The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks Muon is Scalable for LLM Training
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4ac87229-543a-4618-82ed-a76d92920cdc · inbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Muon is Scalable for LLM Training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fd89786f-fda9-40b4-830d-05dc7b5b46ed · inbound
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition Muon is Scalable for LLM Training
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9a27d841-7790-445e-8685-b04aa426701e · inbound
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs Muon is Scalable for LLM Training
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6ad88bc1-6ada-4333-bd01-da2f5a247fd2 · inbound
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs Muon is Scalable for LLM Training
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1f996a58-1a1a-4d0b-90ff-130a5dc0fcec · inbound
OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling Muon is Scalable for LLM Training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9feeee59-8ee3-4639-9e50-378fba61f74e · inbound
ZAYA1-VL-8B Technical Report Muon is Scalable for LLM Training
Reference 131
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 439cb60a-8571-43c0-b3a0-1804fe07baeb · inbound
MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI Muon is Scalable for LLM Training
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6f01e845-5ad5-4b83-8a4d-e6fbf1b43089 · inbound
MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI Muon is Scalable for LLM Training
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4157f7d6-3182-4674-9842-4694dd87c5f1 · inbound
MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI Muon is Scalable for LLM Training
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 082b43f8-c05d-4c48-8c34-e220d98336f2 · inbound
Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning Muon is Scalable for LLM Training
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ed53cae7-0045-4aea-a096-2cf0aa219e72 · inbound
Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning Muon is Scalable for LLM Training
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 489a2f61-411b-402e-9ed1-1d257368a913 · inbound
Accelerating Zeroth-Order Spectral Optimization with Partial Orthogonalization from Power Iteration Muon is Scalable for LLM Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 588618a4-18ba-487b-b379-860d76eff338 · inbound
Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers Muon is Scalable for LLM Training
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cd84be3b-6c64-4028-b4bf-6b52a458b634 · inbound
Dimension-Free Saddle-Point Escape in Muon Muon is Scalable for LLM Training
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 07adcf44-0419-4878-ae91-8d2211613b70 · inbound
Phases of Muon: When Muon Eclipses SignSGD Muon is Scalable for LLM Training
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6cf5ea67-3b49-4e40-9ffd-6a597ae35eb4 · inbound
Can Muon Fine-tune Adam-Pretrained Models? Muon is Scalable for LLM Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c7a76cb0-7019-417f-8036-e07bb5e72db2 · inbound
Mela: Test-Time Memory Consolidation based on Transformation Hypothesis Muon is Scalable for LLM Training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 83500b89-85a4-4ee5-939f-70c314f66592 · inbound
Uniform Scaling Limits in AdamW-Trained Transformers Muon is Scalable for LLM Training
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d81fa253-3fc1-43f7-bad7-fd6cadbb506d · inbound
MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization Muon is Scalable for LLM Training
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5f798e7d-9033-465e-a6c4-aad6ff2ee2a9 · inbound
MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization Muon is Scalable for LLM Training
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b127981-0878-415d-b3c2-ea2a82116c00 · inbound
Elastic Attention Cores for Scalable Vision Transformers Muon is Scalable for LLM Training
Reference 165
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8709b56d-e829-4d26-9752-9b5e0612d388 · inbound
Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation Muon is Scalable for LLM Training
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 475a0326-70d3-4ead-bdf2-0facbf298f23 · inbound
Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence Muon is Scalable for LLM Training
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3b2c0eb8-8870-4042-ae4a-d987ab929428 · inbound
$\phi$-Balancing for Mixture-of-Experts Training Muon is Scalable for LLM Training
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3496cfcc-8c8d-40a8-b583-6e4fe1791b34 · inbound
Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered Muon is Scalable for LLM Training
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e53c553a-4335-4195-8e83-eb13908704c2 · inbound
Towards Human-Level Book-Writing Capability Muon is Scalable for LLM Training
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 98b1ed54-b265-4213-b9a4-c1ab2b4cac64 · inbound
Towards Human-Level Book-Writing Capability Muon is Scalable for LLM Training
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2f910442-65a7-4948-ba07-b30c69b04633 · inbound
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers Muon is Scalable for LLM Training
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ae9030a1-d890-4c16-9497-c7003ec2333d · inbound
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers Muon is Scalable for LLM Training
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 40c85e73-8832-41fb-883f-bd2ddc3c322f · inbound
Scale-Invariant Neural Network Optimization: Norm Geometry and Heavy-Tailed Noise Muon is Scalable for LLM Training
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0da511cb-f4d0-4551-8171-f95bd4a1c45d · inbound
Distance-Aware Muon: Adaptive Step Scaling for Normalized Optimization Muon is Scalable for LLM Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e7a935f2-eb42-431a-a562-2c26e1fa6b52 · inbound
Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR Muon is Scalable for LLM Training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7e5e0ac7-5f69-452b-b20c-7eea6194db59 · inbound
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models Muon is Scalable for LLM Training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8568f2cf-0fef-488a-9257-39c79db255ef · inbound
LionMuon: Alternating Spectral and Sign Descent for Efficient Training Muon is Scalable for LLM Training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9230e505-d52a-4291-a4c4-cd2e6caf4c2d · inbound
LionMuon: Alternating Spectral and Sign Descent for Efficient Training Muon is Scalable for LLM Training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2d95b3b5-d4f1-44d4-b3f1-e1c51dc799c2 · inbound
Toto 2.0: Time Series Forecasting Enters the Scaling Era Muon is Scalable for LLM Training
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 66207ba8-a4e4-4d7a-984d-2f4aea0ddd59 · inbound
Toto 2.0: Time Series Forecasting Enters the Scaling Era Muon is Scalable for LLM Training
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4483ef33-44ec-4757-b69a-f26d42275e14 · inbound
TBP-mHC: full expressivity for manifold-constrained hyper connections through transportation polytopes Muon is Scalable for LLM Training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ca2c268b-d528-4cc1-a53c-6b96195ed89b · inbound
Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws Muon is Scalable for LLM Training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 42973ad1-9341-4004-aa92-3496f450b54c · inbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Muon is Scalable for LLM Training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 400b49a7-d530-4f35-8c50-7b2f34679cef · inbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Muon is Scalable for LLM Training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b2bd861a-3e82-474f-a63a-6073cd79519e · inbound
Anytime Training with Schedule-Free Spectral Optimization Muon is Scalable for LLM Training
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a6d0dd5a-0a70-42c3-b00a-5f59378a7a05 · inbound
Training-Free Looped Transformers Muon is Scalable for LLM Training
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 454d169b-00e5-4c05-bb1d-dd07503770a9 · inbound
Learning Laplacian Eigenspace with Mass-Aware Neural Operators on Point Clouds Muon is Scalable for LLM Training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f7482c3e-6e45-4439-a5fa-357a90e8cd11 · inbound
Mapping the Schedule x Bit-Width Boundary in Sub-100M Quantisation-Aware Training Muon is Scalable for LLM Training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.