Pith. sign in

Paper Citation Record · LEDGER

Muon is Scalable for LLM Training

As of 5 August 2026, this Paper Citation Record lists 100 of 114 outbound references and 100 inbound Pith citation observations for arXiv:2502.16982.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.16982 v1

Coverage vector

measured 100 of 114 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 100 of 183 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:51:05.313853Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T21:57:38.615639Z

Reference resolution

100 of 114 outbound references displayed

  • verified exact4
  • verified fuzzy68
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch26

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fb514fb3-d163-4219-bfbc-d7c18563d097 · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.778741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:9ab5e00b0a67a4e4401747e966d044e6f2cf36afe1cbbbb35cc06e198e0e3e12

Observation 886ce2dd-d23f-4fb6-a33b-514a8a83d9c3 · outbound

This paper cites 2017 , eprint=.

Muon is Scalable for LLM Training 2017 , eprint=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.793592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:57f41253f54d3264b82b2249500892c4f275d715d873d322ae3d4b2128c61390

Observation d3b8bed8-e591-44bb-ba0f-544b61312d4e · outbound

This paper cites The effective rank: A measure of effective dimensionality , year=.

Muon is Scalable for LLM Training The effective rank: A measure of effective dimensionality , year=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.799753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:625307f5a79f59091c4d1c89aac75b9555cae939a295ed40500c48347a0f7ea7

Observation d64caa3a-8fa7-4841-9acf-f4b88581c6e3 · outbound

This paper cites Brown and David Botstein , title =.

Muon is Scalable for LLM Training Brown and David Botstein , title =

Reference 4

Resolution
verified exact
doi, observed 2026-05-11T23:02:51.860430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:b0a64544227b0c7eac7c322662f939340a1a1cdfaf32296e097b90e6a195a1c1

Observation 4b8188d2-9275-4586-a29b-376a6ae44505 · outbound

This paper cites 2023 , eprint=.

Muon is Scalable for LLM Training 2023 , eprint=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.804186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:73f53e8e1adb26614914ce319d48dc2d0a25383da6e73d3a34bb6aa9497977c3

Observation aef93166-9bd0-4f18-8085-1907ad5e7baa · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.807812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:639a1ec24064e72584c9fe43d72ff4b436899092daf13557cc36dc3e46b3f6e8

Observation 7a941ef3-59e7-4c5c-aa90-6df531766d9b · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Muon is Scalable for LLM Training Advances in Neural Information Processing Systems , volume=

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.812664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:52e0f29bc9de3b4ab2be8233116c0b994f710b797c6e8ae7fd8ce51565ad02a0

Observation 2beb3b2b-a8b7-4807-96ee-9f12ce2719da · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Muon is Scalable for LLM Training Advances in Neural Information Processing Systems , volume=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.820688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:d58201ef9104009818bf53bd93821c91c7dc350d800879c06c49c02a2c329c83

Observation 56f944c8-504e-425a-ba49-eec402a91bac · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Muon is Scalable for LLM Training Advances in Neural Information Processing Systems , volume=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.826851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:1621ba5d7a791abc034ff54d07990d8161a70f60d6caa9ced5fadbd928e8846a

Observation 3038a8e9-a940-44eb-aa91-1161e362c1e7 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Muon is Scalable for LLM Training Advances in Neural Information Processing Systems , volume=

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.831674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:93b78e396c09b88535a550000aad181085c72bd9227c8beff5c3b81569f8e1fc

Observation b219efbb-de3e-40cd-9cd7-17550a6aeb17 · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

Muon is Scalable for LLM Training YaRN: Efficient Context Window Extension of Large Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:46:51.373347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:99c3e767779a5239e70aaff9168ca4b473908a7af99065ab24c5ba8d3fb2c9dd

Observation 5432d18c-db72-404c-b234-3c7b78834898 · outbound

This paper cites MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation , year=.

Muon is Scalable for LLM Training MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation , year=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.837217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:def22517bfa227346083d636d376c4d3d5f6183aa9211b0efd8bbb26da3a4289

Observation efe04cd2-d40e-49cc-9c53-0a2e03168609 · outbound

This paper cites IEEE transactions on Systems Science and Cybernetics , volume=.

Muon is Scalable for LLM Training IEEE transactions on Systems Science and Cybernetics , volume=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.848348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:913ae4b52c0fd89f3df7d4a8413fc6fd9cb078f4c60e16a8a1b9810c0949384d

Observation b655815f-1d16-4b1d-9f0e-61cbff7a30fd · outbound

This paper cites International conference on computers and games , pages=.

Muon is Scalable for LLM Training International conference on computers and games , pages=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.853325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:cc765e2e3687691ad1bc399cfa2571fd7f408fb55ce5f8346a6860306afd0feb

Observation 19db4135-72eb-4705-bff6-7954e08497f8 · outbound

This paper cites European conference on machine learning , pages=.

Muon is Scalable for LLM Training European conference on machine learning , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.857885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:12e62ad802d7790961577c394fad2c5fa948fd46a80b6b3d15e529f8276daba3

Observation ba6d9d32-e677-40a0-8f37-023d298fe290 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Muon is Scalable for LLM Training Advances in Neural Information Processing Systems , volume=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.869987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:7281b570697e50743fc3ef5681f08ca1c7d3872897b1be172d8e18ecbc11f40d

Observation b77600a2-1578-4996-92e0-53f2ea8bc468 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Muon is Scalable for LLM Training Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:38:37.536547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:9074be07f1d018dead8d07fa5fb0666df9f87b428734f8a326ef0a5384b35d57

Observation 454dbe6f-7d6d-45cd-9182-ee68d466bed7 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Muon is Scalable for LLM Training Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:02:52.253433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:02ce91887c3ba0dd2e69496edc786b99cab7685e6777a34d3e6ccbc405988690

Observation 25fbcfea-0db0-43f8-935f-5fa23cb5b2e9 · outbound

This paper cites Advances in neural information processing systems , volume=.

Muon is Scalable for LLM Training Advances in neural information processing systems , volume=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.878144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:3a2b0ad70a41a0d156aa4de7b77291138b9c23afee76c9187877dec1b6b0b967

Observation a97ed35f-381e-45e4-853b-52b45d252804 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Muon is Scalable for LLM Training Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.312751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:2ab10fb50d9e063110a85da89a65602deac7acf0f88358eaf8c4f0de366ddd66

Observation ee4b0765-b7e8-4e73-bcb9-54cbd5bbbb37 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Muon is Scalable for LLM Training Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:42:22.769824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:c0fc7f19a94dc3ef54c13276d42d8db54c7947c62a91f65a84f4a21ef710595b

Observation c08fb611-51ed-4a05-81fc-0e9df8b6a122 · outbound

This paper cites International Conference on Machine Learning , pages=.

Muon is Scalable for LLM Training International Conference on Machine Learning , pages=

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.882834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:1cd68010e65778920300c88b54b8522cef3d210749499081235661ebfc8c75af

Observation f8f6113d-ac82-4ac8-b223-d044059720c0 · outbound

This paper cites Proceedings of the 28th International Joint Conference on Artificial Intelligence , pages=.

Muon is Scalable for LLM Training Proceedings of the 28th International Joint Conference on Artificial Intelligence , pages=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.887700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:07c7fb2154f3ae37324f9e150667f62cdfaa46ccd18e708b06176d3d878582f2

Observation ce92ea2b-0577-448c-bb91-92038e927280 · outbound

This paper cites Mirror Descent Policy Optimization.

Muon is Scalable for LLM Training Mirror Descent Policy Optimization

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:02:52.240735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:7a206951a7de5f0dd8e430890dc8387c57ecfe305dc1dc4406626950f6630440

Observation bff0a9b5-3de0-4a08-b0ab-e8d6a307c037 · outbound

This paper cites Advances in neural information processing systems , volume=.

Muon is Scalable for LLM Training Advances in neural information processing systems , volume=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.903038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:6cdc4d7a7534903017ae5775e8957f41ee7f989adc1f488f37220057e24a9ecc

Observation 9584fd9d-7dd7-4646-ac66-67c7309c143e · outbound

This paper cites an unresolved cited work.

Muon is Scalable for LLM Training Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-11T23:02:52.908300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:b31cafdf3f47a33025617c521ac1258dca4cc2b4069cd4f5d48705c7cd5c20fd

Observation c572c383-e3a0-4b64-81aa-9aa0b40dcbae · outbound

This paper cites Neurocomputing , volume=.

Muon is Scalable for LLM Training Neurocomputing , volume=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.915322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:eb5f7fee5426a16c535730a608ab53276014849d19c8a698c7a252020bd7d584

Observation 2ecc027d-356b-4301-9e74-1e1d66aed3b6 · outbound

This paper cites 2024 , url=.

Muon is Scalable for LLM Training 2024 , url=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.920360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:ef31d1eb8e47a36605041d9fd7020410a1cebd3079875c8524ade0282178f434

Observation 745d6af4-01ba-4076-a8f9-7ee54fc918f8 · outbound

This paper cites 2020 , eprint=.

Muon is Scalable for LLM Training 2020 , eprint=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.925202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:87f1a3e833b53a643be04045a8cd17b51e841e49df249685c112b00649e3497d

Observation 47e734c1-68d0-4c6d-b134-557405cd62ac · outbound

This paper cites Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=.

Muon is Scalable for LLM Training Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.930560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:5643b7520e9227911cc138f9fe053ab42f961be2b88e870a990279de3704342e

Observation 218fdfe9-dd08-4663-92fb-e71f40e9b444 · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.937454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:0ac43e25880cb34835485c402f97f8baf094f53fc5af0f7578109f91523714b5

Observation 51fe9497-b28b-405b-bd37-7ee887bcc684 · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.944713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:cfe74b09c9d1e5734234ef5df543b700dcf0c7b87b6788caeaf5f6bea7bf9e57

Observation 0f347ae0-d2b3-4a06-8852-81fadffebc0f · outbound

This paper cites Attention is All you Need , url =.

Muon is Scalable for LLM Training Attention is All you Need , url =

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.954370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:c5d9b592eed3a30244a432e090b6b8836c24609cea03be386b6bcc1dcec209c2

Observation 47715c98-443d-4357-b9aa-1935978d4d4e · outbound

This paper cites ArXiv , year=.

Muon is Scalable for LLM Training ArXiv , year=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.959643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:97dadae4aeff07838ef60624cf72cbff6ab81fba88e35d771104e224953d0940

Observation 3884b18c-2c95-41e0-b61c-e6016444b9a3 · outbound

This paper cites North American Chapter of the Association for Computational Linguistics , year=.

Muon is Scalable for LLM Training North American Chapter of the Association for Computational Linguistics , year=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.964869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:a3b864c9a24107b82153004b28e81c2dd183ada80cf1d66f7fc8fc117245b404

Observation 3ef53054-c7c7-4b06-bebe-003e2e18272a · outbound

This paper cites ArXiv , year=.

Muon is Scalable for LLM Training ArXiv , year=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.970007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:b3a93e475111c38a2ffe35e602b6ac83b5dc378b6301926226c814dfc85ac703

Observation 9c4206c1-73ef-4da8-8d4f-1125511f16b5 · outbound

This paper cites 2024 , journal=.

Muon is Scalable for LLM Training 2024 , journal=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.974239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:7b326a2b5bd3f78431a4e45321d110ff24d197d87676e64b207da141fd41c497

Observation 42a8f0af-be7a-4117-b39b-4648b0f9b409 · outbound

This paper cites International Conference on Computational Linguistics , year=.

Muon is Scalable for LLM Training International Conference on Computational Linguistics , year=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.979536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:b170b32853ed50f33d7d49d87b6a1350be0ca965045f3ed4cd3aa8d370c7babe

Observation ca864ecf-6c9e-45f8-b69b-bf1aea58544f · outbound

This paper cites ArXiv , year=.

Muon is Scalable for LLM Training ArXiv , year=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.362347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:70c69a9e9ab3147e0ca92d2fa546810b9abd50cf8e3352d2ca5fa3bcb15999cb

Observation e5abf61a-7a9d-43ce-9284-3e19c05b52a9 · outbound

This paper cites ArXiv , year=.

Muon is Scalable for LLM Training ArXiv , year=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.373362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:25675419ebf6aa68ee515440a136f73be765fbc29b49e40c9b7cbae0c7ddc68d

Observation 259b7364-c77d-4c74-a145-98f984338d56 · outbound

This paper cites ArXiv , year=.

Muon is Scalable for LLM Training ArXiv , year=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.381475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:ec4d7c9ab388a27e5fda5aad9c66160d1278f5721badac26aa59350f78acaf0f

Observation 1a7366c6-3d29-474c-93af-70d0191b028e · outbound

This paper cites Let's Verify Step by Step.

Muon is Scalable for LLM Training Let's Verify Step by Step

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:02:51.885074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:acc27748ae211668530e5b1e132d8138b78151eabc957ee1ce4a32f9017ddc27

Observation b18c9154-cbfe-4a1c-b7dc-8debf06b2534 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

Muon is Scalable for LLM Training Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:45:37.909362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:a09d1b2654e36f6cef0e15ad4473245e89db63c20e1734b5641352cfef5221e3

Observation 46e1fea2-eee6-4e06-a972-6b49dd525922 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Muon is Scalable for LLM Training Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.391769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:357e2c0620c48077c656694da488eea2aabeabe0abef1b2a93264a3b3de99408

Observation 837a51a9-5e67-4a82-bdc8-4dab344e566f · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Muon is Scalable for LLM Training MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:02:51.940482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:1698bf6315f0488ebf6f79a25e393f6a88e0577edabe3e7de53e81080a5dbef9

Observation c6e22c65-b264-4d57-9f4c-6d554820de07 · outbound

This paper cites Bag of Tricks for Efficient Text Classification.

Muon is Scalable for LLM Training Bag of Tricks for Efficient Text Classification

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:02:52.041442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:5a17c8bf420f4b307a6c9bef70c67d281835cbdd39e3f7613dbe760993036be3

Observation c0fe3862-babb-4845-928e-3732197b3774 · outbound

This paper cites M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation.

Muon is Scalable for LLM Training M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:02:52.058712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:170bcb7cd8e2e932185fe476ef29838fc5ec12fd9cdfceb3e4996368c33120d1

Observation 49884703-9a55-46ec-ace3-0babb8b1dc27 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Muon is Scalable for LLM Training The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:36:45.776658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:8dbb694046309030b2b13ef71a35549daa188c66a50c4d162e39b3ea4ae7a827

Observation 2b92e572-5f33-4239-b90a-7faa28f0ed05 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Muon is Scalable for LLM Training DataComp-LM: In search of the next generation of training sets for language models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:58:17.761776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:a42cd9216b06032513f147c90e0c539e836b852fbb29753e89a66167f599f2be

Observation 69ad38c9-2a10-47e0-8daf-ff886b1b8b93 · outbound

This paper cites OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text.

Muon is Scalable for LLM Training OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:02:52.122348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:7df0d528b1f838212aa49a9c0a9176a9af63e470076cb69c552b21157d5411a0

Observation 4817a92b-4d58-4eb8-8dbf-5b32f884855f · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.409671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:fa8d2093680fe5402287537dcf9e613ac9cbd3aa8e12a574ede66060b117aa8f

Observation 71cd9f16-bac3-42cc-bad7-e63b2ab8adbf · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.419390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:a8978572e97ae231a062ae2ee0e70adbcd8752c9fae72c517871501c1bcfa5ff

Observation 3cfd9307-426a-4237-a64b-10268a5adc5c · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.429650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:10fb84a17b864ec7c89bc24b04341a234bc083ecefa45e0ad02c76e5bd125648

Observation 5309a0e4-bbe1-46cf-b595-07a34d1ce3b0 · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.439344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:aac4eb684351e1242c230d7bc8e8683120da33d7d067b14a7f9207b16c3d6ecd

Observation 863a3ad5-6f54-4f78-b28c-fb028a5e851e · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Muon is Scalable for LLM Training MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T23:46:39.773471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:d57f4eaf8cd15cad564162a5876134fb8fe318cf0851d3c6c711421e3ed8e3d6

Observation b11cfc50-8406-40d3-a048-37ecaeeb235a · outbound

This paper cites Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset.

Muon is Scalable for LLM Training Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.280166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:f40bf86e52ea28664d5cac8b23c6c182ba6cfc65d3828e74d49afa1d5966e833

Observation 6a6ce351-a034-489c-ab0e-dcb263f03aba · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Muon is Scalable for LLM Training Reinforced Self-Training (ReST) for Language Modeling

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:02:52.299431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:c4b4799120630c9a3974e617e744a5408bf966217921f042e82fad178e3c61e9

Observation 24648568-03a9-427e-8527-ae940339d12f · outbound

This paper cites nature , volume=.

Muon is Scalable for LLM Training nature , volume=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.446057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:c668638e25e497d51b173c1672bbdc7e3f00c51e08a2ef6c56d550f9476fa7ea

Observation d226672c-2fbc-4a49-953c-54051a936ea9 · outbound

This paper cites nature , volume=.

Muon is Scalable for LLM Training nature , volume=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.452706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:2a68a3168a43ae3056c29b1a1d5f0b4358758d4c1a8efeb01ad1e0bca9f98f53

Observation 7d1e62c1-c4bc-451e-833a-0b6314859c9a · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Muon is Scalable for LLM Training Dota 2 with Large Scale Deep Reinforcement Learning

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T22:18:16.018301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:9371ee43f84ba39f6f1a4ac5a88ab9884520bd3add02a77e0bb6b744ac3f546b

Observation 0cb61bc2-57fb-430f-8858-cf56bb15c7d1 · outbound

This paper cites Advances in neural information processing systems , volume=.

Muon is Scalable for LLM Training Advances in neural information processing systems , volume=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.462678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:1b681cae7ced2ae440bcf3a8463838b55154faa56eb740b1dbb35067d5c18147

Observation a4982843-a78a-4e39-a2db-6254f9a8c684 · outbound

This paper cites OpenAI o1 System Card.

Muon is Scalable for LLM Training OpenAI o1 System Card

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:02:51.896015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:1300b5a924a4dfc9e0c09a4737e9040d85bd2fcc59e13fed337ef32b29169349

Observation 6d2d216c-a8ec-45ad-85e5-abcf105a04e0 · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.473591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:37af579be499d843571432aba773dd5e4255ec8f760d8be5eac253447f559ac5

Observation 43dbd1b0-49b2-45ac-b3d2-78b5687f14df · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.486494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:76b553e59389fe08ad55865709479a6a6ac5181c30cf4febb677e093fdf31273

Observation 4ba30ccc-ff60-4982-9a27-e58d5ab7f6d8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Muon is Scalable for LLM Training Advances in Neural Information Processing Systems , volume=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.492668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:44acc1e91f59e3b7013e42c1b1588a0b0824714be339029175fa6c8cbfb3af4c

Observation 2a9b8a1f-0e12-46e7-a21d-d42e47d42237 · outbound

This paper cites Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities.

Muon is Scalable for LLM Training Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:16:05.189340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:d4f19876c5468514ea3c0f06c982886f99e81573672d6eca1e7aefb5d1114200

Observation 8d3ec19a-861c-4a7f-8993-442add9c6c59 · outbound

This paper cites 2020 , eprint=.

Muon is Scalable for LLM Training 2020 , eprint=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.500737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:f1900d1dd6513ca189cab04da5f0e78061a0c3186acfb01bb0b64d2d0f37f2ee

Observation 0edcdcdd-c3fb-4829-bf26-4dbb9eac454f · outbound

This paper cites 2022 , eprint=.

Muon is Scalable for LLM Training 2022 , eprint=

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.508726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:605311b1479b60256c71e28ba6c43cc4bcf78b2b7f60b233c79824d533cd2357

Observation 9247de9c-dbca-4cfc-9353-d3ebd6eda934 · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.520076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:c08ff403c11ffbda7e8fa00023330cf6be70c98dd81b2683afc775393a53d791

Observation cf9cb028-ce10-4708-b546-9eac009236a6 · outbound

This paper cites 2023 , eprint=.

Muon is Scalable for LLM Training 2023 , eprint=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.525423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:11a704dca1f542f2f3545dcea74d3503d4a1ebfc3fa6029c437b756f95d648de

Observation 0b73bebb-7e9d-4dae-9ec4-25d3d96d9260 · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

Muon is Scalable for LLM Training General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:c3a43bb5687fae9062bdb380ac6f095ac8a6c63068c86e3eb4ab6ad0f61f50b4

Observation c04232aa-f25a-4a92-8760-f1cb83beb0a3 · outbound

This paper cites 2021 , eprint=.

Muon is Scalable for LLM Training 2021 , eprint=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.534241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:bd9dcd01350eae316ab2647af546707648666812b70376f44c26fbc7ade07535

Observation be3f4f2f-1e36-4708-a436-c187dfd25a42 · outbound

This paper cites International Conference on Learning Representations , year=.

Muon is Scalable for LLM Training International Conference on Learning Representations , year=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.542398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:42f0eb6470e8784e4f294a86e7cdbc1c319d8ecd61e6615b95efdfff42f47583

Observation d92228bd-9519-474b-a53d-5786dc40b4b8 · outbound

This paper cites What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning.

Muon is Scalable for LLM Training What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:02:52.146183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:61e98f3e24406ca686d2b680b9d8702eb8285774f36b9e6d8aeeecb3b32b7712

Observation 8185dbc5-4548-4b35-b7aa-1c461762f290 · outbound

This paper cites From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning.

Muon is Scalable for LLM Training From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.179429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:a890d44a467cceada8f266556c19c0f79e369196b93b577d084d4e93a35e815a

Observation e9f0f70c-a7dd-4908-904b-d89ee30a0c47 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Muon is Scalable for LLM Training Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:51:30.429014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:f67eb7b9faaf5db3041cc1e6a9673f7d5e7d88f5e9391dfc972859e868042ca2

Observation 65ca70d8-8452-4428-aaca-1b1968a32574 · outbound

This paper cites 2024 , url =.

Muon is Scalable for LLM Training 2024 , url =

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.554669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:7a7c7155970da0bb46655d331998d3204321b16f1a3c437161ead86a74e65cea

Observation 2cc933f4-5fa3-481e-b972-5e83562bf20b · outbound

This paper cites International Conference on Learning Representations , year=.

Muon is Scalable for LLM Training International Conference on Learning Representations , year=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.563135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:2e180bcddef67896873369a175c32260a1a0a7bc7108cc892debca270412ed11

Observation 62acc690-6490-460c-aef6-6368b83111e4 · outbound

This paper cites 2024 , month = Oct, url =.

Muon is Scalable for LLM Training 2024 , month = Oct, url =

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.568447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:ac8b9490ac14aea72bdb1e2f1105db9bd0bc5e8789f7787bd27d0f72ed4bba56

Observation eabc162d-b034-4ac0-84f3-174f869b1f5d · outbound

This paper cites 2024 , url =.

Muon is Scalable for LLM Training 2024 , url =

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.580131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:263e1b278a84ed04b70abe2d6bfa662af37e207564f4c928598f2c70df52cb57

Observation 1cc1be1f-15d2-44e0-8557-bb616d4d9d4e · outbound

This paper cites Kingma and Jimmy Ba , editor =.

Muon is Scalable for LLM Training Kingma and Jimmy Ba , editor =

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.586463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:3df374c0c118eaf3c53894796a7efe8a7410596b443519de4f602a4f932a740f

Observation d1790b3d-761c-4076-b2bf-c736a36bb187 · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

Muon is Scalable for LLM Training The Twelfth International Conference on Learning Representations , year=

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.591632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:af8ca828b920b2047864f933f5d5275bed6972b54d54699496e5138ba0232ab6

Observation eab95519-6437-4218-91a6-f8ce0b193ccf · outbound

This paper cites Kakade , booktitle=.

Muon is Scalable for LLM Training Kakade , booktitle=

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.596864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:d29532b39a93bbf2e4f0b7e04f92e3b480d2eb40ddd05e00aab3c8864060cbc7

Observation 7164ae32-d75e-4f41-90be-c777bf4b3a31 · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.606188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:e67ff909adaee1c5d3e60284a4ad8a1248d1e9df692354cd7ec7b84f351a928a

Observation 6e9ae587-8672-4b42-8429-ede093b82551 · outbound

This paper cites Generalized Slow Roll for Tensors.

Muon is Scalable for LLM Training Generalized Slow Roll for Tensors

Reference 85

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:02:51.871689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-11T11:50:26.030339Z digest=sha256:5fb9b4c2b92619b0fe6c16c0a7b59bb0bbf4485361e9d14bae7e247a35d998d6

Observation 3f2a0681-8893-47c9-9788-51b10763fb44 · outbound

This paper cites Qwen2.5 Technical Report.

Muon is Scalable for LLM Training Qwen2.5 Technical Report

Reference 86

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:02:52.338525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:6ac23643b40d79b2fce3f51049e3c57207a643c80ca98bfde61994ba3aa911e8

Observation 3042fe40-f88e-48de-9612-f6d904d6ba39 · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.611223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:20f97b7fd8f07164413e26b9875e0f2c97ce77c487287447c6abd122c9334325

Observation 5beb1a16-9548-41de-a50a-a8df575e0e2b · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.622339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:4ca32c4e53688abe2e035f3e1bf532d209123925b3913bd19ff6377cc123fffa

Observation 8ca70984-6f90-45b4-848d-da4d5cef47e0 · outbound

This paper cites 2024 , email =.

Muon is Scalable for LLM Training 2024 , email =

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.631347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:f7d91368ca959740f8e2850c08e8594eb83b697a5bef5cceffba85f65b29d892

Observation 267dfd4b-1a3a-4ed6-9c47-00928948093e · outbound

This paper cites an unresolved cited work.

Muon is Scalable for LLM Training Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-05-11T23:02:52.636172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:7bf9a47a0a25644049ef1bd18a6f19abc99d6f6b3f364a66223afa55ad8014b9

Observation 9d90a456-eb1b-4a75-bc4c-878908725abb · outbound

This paper cites 2025 , eprint=.

Muon is Scalable for LLM Training 2025 , eprint=

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.646251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:30ea56703dad0ce538cfc0aa3294d5adb7f8244a1a1a1d26a9ec74b23671dc19

Observation c8c2a863-01b1-4057-afbf-66ab9330d08f · outbound

This paper cites 2025 , url=.

Muon is Scalable for LLM Training 2025 , url=

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.662230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:7eab0d14de424feca2552cec3b7e11c0088ae856f84b3b50136fefb99ee8500f

Observation 085eef7e-99e0-4653-a839-7f65fa8a4361 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Muon is Scalable for LLM Training DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 93

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:02:51.951346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:3a406d72ac09f8458b7cddfb6318a868114270944c85b00b1771fc9d8c609a2d

Observation 19daf070-9e86-4e00-9fc2-cad4bbef2b63 · outbound

This paper cites 2 OLMo 2 Furious.

Muon is Scalable for LLM Training 2 OLMo 2 Furious

Reference 94

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:02:51.983358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:8c0b8063e68dde0cd46001b610180f4ff01fb20319632f04f77ace9b1670c637

Observation 0099226d-dcb8-4b66-915a-790e5a06eba1 · outbound

This paper cites 2019 , eprint=.

Muon is Scalable for LLM Training 2019 , eprint=

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.670572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:e93e26a95e971837f6debc299e7619daeb87e82ed58fd981dd37b88ac2d71bae

Observation 28e75e17-90e1-45b1-8283-28ba35d04394 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Muon is Scalable for LLM Training Gemma 2: Improving Open Language Models at a Practical Size

Reference 96

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:02:52.022870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:25c0ff99eedd6e7bec31145cb81cd9585c6fcef2a7ca717a9748b669b67a0fd4

Observation 239e3013-ac07-4a2c-aadc-ebace1e1b24e · outbound

This paper cites 2024 , month=.

Muon is Scalable for LLM Training 2024 , month=

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.678187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:6dc3ebd9fe8e54c32d5e7de06a4bf6ea051b1582aff55670baeb27ebd21272a8

Observation e9cfec47-2c93-48e6-a9e4-c66e6aa25b00 · outbound

This paper cites 2021 , eprint=.

Muon is Scalable for LLM Training 2021 , eprint=

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.682511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:f9834ab869ad81348cc4f5fd1951787f374647c42e715638b873ec6f196871a6

Observation b078db38-f633-4b5e-a137-46673d6bfed2 · outbound

This paper cites 2024 , eprint=.

Muon is Scalable for LLM Training 2024 , eprint=

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.686683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:0f7ee24509f6fcb1be65ec4024cc66cd3d72aa0ff7d27079bea19094126da295

Observation 53530e26-5765-491b-91d6-199387f19311 · outbound

This paper cites 2022 , eprint=.

Muon is Scalable for LLM Training 2022 , eprint=

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T23:02:52.692691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:0e5a84d2ba56569f520851d4899b6f803bd9a499f3ed791c8763eaf5e5b59bc3

Pith citing papers

Observation 1c8d5c47-a6a5-40c7-adad-c9db58cfed8d · inbound

Eliciting Latent Predictions from Transformers with the Tuned Lens cites this paper.

Eliciting Latent Predictions from Transformers with the Tuned Lens Muon is Scalable for LLM Training

Reference 104

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T16:54:37.589737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T16:54:37.382049Z digest=sha256:5f8a747a216d5630bd2f45a2b3688dc47a1d3782c339ce543ad3e9cacc98bab6

Observation bd0c1398-2ded-4a28-bf0f-31c707348370 · inbound

GWT: Scalable Optimizer State Compression for Large Language Model Training cites this paper.

GWT: Scalable Optimizer State Compression for Large Language Model Training Muon is Scalable for LLM Training

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:57:36.698034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T05:57:09.276224Z digest=sha256:53cc62cf9c6ee15b591ca4b88b5aa08321c6f3ff3b94518987395fb1d57c4247

Observation 1abb37ca-9b1e-4a01-a0ab-727623b95114 · inbound

Training Deep Learning Models with Norm-Constrained LMOs cites this paper.

Training Deep Learning Models with Norm-Constrained LMOs Muon is Scalable for LLM Training

Reference 195

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:22:37.053910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T21:22:36.870292Z digest=sha256:086362b11c1f796bc26dda512e9c7272a7cfb4bbfdc18d1b6840cad46a67453b

Observation bb2dc52c-4fe7-4999-8f96-38a57c4de888 · inbound

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning cites this paper.

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning Muon is Scalable for LLM Training

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:18:43.909761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T15:18:43.724432Z digest=sha256:d1581baa86134b47c872c40a3fa0dd9aac176e34b1e2b2d0a9471432b9d1f867

Observation 779f8ea5-0cab-4f2c-abc8-edb41048c54d · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report Muon is Scalable for LLM Training

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:9075a3a403b3e8ff4e30ab2c02bcc3e04a69e2a4b2c3210c57e643ae3320ea11

Observation 028443af-ccb7-41c9-b7b5-5de65298e426 · inbound

On the Convergence Analysis of Muon cites this paper.

On the Convergence Analysis of Muon Muon is Scalable for LLM Training

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:22:19.180929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:19:37.035526Z digest=sha256:4ac6bd6bda86010c88afb674bbfb52eb69d630775bddf7c5a4e66aaf68568c6d

Observation 8ac2c9a3-d5ce-4a81-bdec-76d82c467ae7 · inbound

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design cites this paper.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Muon is Scalable for LLM Training

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T13:24:53.207881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:2661b65b016724de2f6a6b440fe1f7b4707a8f3d500a38f05616853dd64b2883

Observation 23c85ed5-aa98-44c1-b5fb-1e0a6e8d7cc8 · inbound

Kimi K2: Open Agentic Intelligence cites this paper.

Kimi K2: Open Agentic Intelligence Muon is Scalable for LLM Training

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:49:27.926646Z digest=sha256:28ef25416b5b69cbc4390b1ddb2062b95c15abe4d678edf5ed9f17b967d48e37

Observation dd200900-e693-451a-8f30-4c6a5327215a · inbound

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models cites this paper.

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models Muon is Scalable for LLM Training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T17:50:08.399160Z digest=sha256:b10e50659c036c7b314378a44c865541660a1b146b041468a0b0f8c1e9f2edd8

Observation d84da147-e5ab-4ced-ac32-2eb0ea9b1545 · inbound

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? cites this paper.

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? Muon is Scalable for LLM Training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T21:51:05.313853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:51:05.313853Z digest=sha256:f8dee6033023b7ad0fcd5edf2df1f2307c6e99d82699bc1fac5cf7e41d237389

Observation 8106248e-29ca-4ddb-a302-3578382c0ab9 · inbound

Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training cites this paper.

Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training Muon is Scalable for LLM Training

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-18T15:51:33.994807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T15:48:43.422716Z digest=sha256:be767f3b6c4b85fbf39d468d8698345044222163e12e9ddf55a3952531f8c3a2

Observation f20acbef-4f18-4bf9-95e6-66ca5b658d1c · inbound

On the Convergence of Muon and Beyond cites this paper.

On the Convergence of Muon and Beyond Muon is Scalable for LLM Training

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-18T15:56:34.068152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T15:56:30.602824Z digest=sha256:4e27279625b9733624b360d216a7fc3bce8fa9a0e36821f0dec355b04da66691

Observation 5dbd9f93-54bc-45a2-8270-a0b344d78fe4 · inbound

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers cites this paper.

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers Muon is Scalable for LLM Training

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-18T12:41:22.754671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T12:38:31.783807Z digest=sha256:5eec6c46209386a87a59b8fff2b1efe48b20522b52b597cdc2c6803d5c0d030f

Observation 53da14d2-3c98-4737-a976-d6c49bdd4168 · inbound

Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization cites this paper.

Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization Muon is Scalable for LLM Training

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-18T09:46:12.531052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T09:44:53.004293Z digest=sha256:ed7ae05bf13b5197fe044bcd8b7023eab79d6d8c5b1dd8a602d9eb6b7dad3190

Observation b202a901-a872-4140-81f3-f3b9e9cc3410 · inbound

Evolutionary Profiles for Protein Fitness Prediction cites this paper.

Evolutionary Profiles for Protein Fitness Prediction Muon is Scalable for LLM Training

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-18T09:01:09.675552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T08:56:44.544404Z digest=sha256:7f8c78f987c356482471e8470a15d487eb77625f6dcc533a9d53b2f22a11f41b

Observation 51d4eab7-67fa-46b7-8794-9dbc0d13dfbf · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture Muon is Scalable for LLM Training

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:49:11.220159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:f8eda41a01eb3bc2d088ab880c354e764356e9b035dc06964eed3651e10b9256

Observation bd615397-34e6-4f3d-b93b-624e56dccaee · inbound

Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates cites this paper.

Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates Muon is Scalable for LLM Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T22:20:05.578066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T22:20:05.578066Z digest=sha256:4dc0651fe0041ecf744b9c9bfe4b2f70568c239cfb3d11e3a369d1a7787d1e34

Observation d63cf5f7-9c32-4fc6-8221-433fc555cfa4 · inbound

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning cites this paper.

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning Muon is Scalable for LLM Training

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:40:14.556303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:38:30.169363Z digest=sha256:0c22f9785648108a88544691853214f02dfc154110368e1906dc12f75be06302

Observation 77f8b84a-73f9-4ea2-bf6f-7b336a2505e0 · inbound

Turbo-Muon: Almost-Orthogonal Pre-Conditioning for Fast Muon Updates cites this paper.

Turbo-Muon: Almost-Orthogonal Pre-Conditioning for Fast Muon Updates Muon is Scalable for LLM Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T18:37:49.596778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:37:49.596778Z digest=sha256:803b6a85426b2eb1e28cfc78061bd269014b67b352a8e248fb4fbe66a70eaff7

Observation f3bfe19c-93fa-486a-bef9-452a91363307 · inbound

On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization cites this paper.

On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization Muon is Scalable for LLM Training

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T15:40:18.938406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T15:36:44.154033Z digest=sha256:4721a152076057816d29fcc1bcc74263022828b9c461aa8a78c29b5f7fde4cce

Observation 3176bd42-ff3e-405f-ac86-649b7b23ba6e · inbound

On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization cites this paper.

On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization Muon is Scalable for LLM Training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T10:09:37.093609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:09:37.093609Z digest=sha256:6baeff0cd86615d06533cfc633f04dfc685eb07b0e54e2965826ba3974d0fa46

Observation d33cb1ce-5500-497c-80be-7f05f59d008a · inbound

KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual Matrices cites this paper.

KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual Matrices Muon is Scalable for LLM Training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T06:57:48.394471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:57:48.394471Z digest=sha256:d1e91f29eda005f542d4bf50031622908a5cc47fd7d741b90a883ecdaa41353a

Observation 6dafce45-9263-48c8-9906-d5bac2f8e6ed · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence Muon is Scalable for LLM Training

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:c431286cf10a99bb3bff804ff4b529282596b029e5564be1b9f692596d954cee

Observation efd02246-9016-4e27-94dc-59067dfd3c17 · inbound

SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning cites this paper.

SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning Muon is Scalable for LLM Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T05:27:10.427125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:27:10.427125Z digest=sha256:5163bdc3a1c3194fb6bd7917af180404f6bee0b3f80969cc02fe2bb209060eb3

Observation dc4694a4-06e3-4f69-9336-3ca7353ecb1b · inbound

Muon in Associative Memory Learning: Training Dynamics and Scaling Laws cites this paper.

Muon in Associative Memory Learning: Training Dynamics and Scaling Laws Muon is Scalable for LLM Training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T04:14:16.087028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:14:16.087028Z digest=sha256:712ad5ebe57669714502e6222b878bc2f3c95cc33d03e151588ba2603f09f6cf

Observation 82d1aae8-36a2-4341-9f32-4e429bb9a959 · inbound

Decoupling Variance and Scale-Invariant Updates in Adaptive Gradient Descent for Unified Vector and Matrix Optimization cites this paper.

Decoupling Variance and Scale-Invariant Updates in Adaptive Gradient Descent for Unified Vector and Matrix Optimization Muon is Scalable for LLM Training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:41.622106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:41.622106Z digest=sha256:f01447cdd9945f9a81f3a89df6008fc301b185c257f14a55f49eaf586e8c5b43

Observation de8bbbba-21aa-46eb-9ad5-cb83fd061767 · inbound

The Implicit Bias of Steepest Descent with Mini-batch Stochastic Gradient cites this paper.

The Implicit Bias of Steepest Descent with Mini-batch Stochastic Gradient Muon is Scalable for LLM Training

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-03T00:19:05.635066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:19:05.635066Z digest=sha256:de386d8c744492b88d09588225fce878efe2705ca7d6b4b83e6798ab1cde6da7

Observation ef59ff3e-048b-4fa1-b43d-0c1cabffa28b · inbound

MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training cites this paper.

MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training Muon is Scalable for LLM Training

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:26:31.348741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T19:24:22.699712Z digest=sha256:4e671ca8af144bbfd8c70ef8d3ea00e709324e715bddeeec5cede7239a6e250f

Observation 4c27bc63-3c62-43be-8c8e-ed9c3682717f · inbound

Spectral Condition for $\mu$P under Width-Depth Scaling cites this paper.

Spectral Condition for $\mu$P under Width-Depth Scaling Muon is Scalable for LLM Training

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:06:25.520938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T18:03:31.202134Z digest=sha256:61bafc584ae75324111ae36dc1a5bfcafe0d2953bafc08b22caa196f2c13aca4

Observation 9d93a0da-45d1-4a23-be1a-c0325b060e90 · inbound

Attention Residuals cites this paper.

Attention Residuals Muon is Scalable for LLM Training

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:39:04.506714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T06:39:04.312270Z digest=sha256:551a61090d0e27d72a8046522c672520c8171b7447255eac5ff38a424fa125a7

Observation 613478ff-c5db-4e36-84fc-b923408a9f61 · inbound

GLENN: Neural network-enhanced computation of Ginzburg-Landau energy minimizers cites this paper.

GLENN: Neural network-enhanced computation of Ginzburg-Landau energy minimizers Muon is Scalable for LLM Training

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:35:19.268190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T08:30:49.868638Z digest=sha256:a57f4c17c307c22e936b45b5279523b0fadcd723569d7013a02e9b1b34dcfa8d

Observation c2d0b97b-bb88-456b-a333-5b8b7d0bf01c · inbound

RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization cites this paper.

RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization Muon is Scalable for LLM Training

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:49:50.756272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T07:49:39.417566Z digest=sha256:3d9b16563cb07a53432d24fde0ca619c62cc7c190a0bbb5821892e7584a97fc7

Observation 515d4ac9-d67d-4f32-a8e1-883ce029ccd2 · inbound

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory cites this paper.

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory Muon is Scalable for LLM Training

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-14T23:38:16.547879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T23:37:33.106390Z digest=sha256:d0216137fb5313255055a66b19cfb27d943173feee913e3b5f66a7018ac43fd3

Observation 7a6a9290-0319-47ea-8ab9-8e7ade9fba11 · inbound

MuonEq: Balancing Before Orthogonalization with Lightweight Equilibration cites this paper.

MuonEq: Balancing Before Orthogonalization with Lightweight Equilibration Muon is Scalable for LLM Training

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:58.572251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:12:32.579136Z digest=sha256:76274eda3f65c4c7a961a14b2baf40619524eb6ee8c7e6eee5b5d3f0c85ece4d

Observation b1a5ce80-7005-4e1a-9c30-b216974c0f81 · inbound

Optimal Projection-Free Adaptive SGD for Matrix Optimization cites this paper.

Optimal Projection-Free Adaptive SGD for Matrix Optimization Muon is Scalable for LLM Training

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:58:15.738102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T20:56:41.466669Z digest=sha256:b17c3d32bc8fd4c30bde6ad6159baf12c86160c270127eac37e294d521cf6bf8

Observation 40a15ea3-c03a-4718-9a26-b2a4e97e996c · inbound

Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning cites this paper.

Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning Muon is Scalable for LLM Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T10:41:51.251736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:41:51.251736Z digest=sha256:fd7d4a54fbf87de81ff91efbe8ec91413bc4dbca4f9547b2561310c3a8a3dafd

Observation 13fd3e5e-6e3a-4c2d-8109-fc4bf2e882e1 · inbound

A Muon-Accelerated Algorithm for Low Separation Rank Tensor Generalized Linear Models cites this paper.

A Muon-Accelerated Algorithm for Low Separation Rank Tensor Generalized Linear Models Muon is Scalable for LLM Training

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T19:16:35.992399Z digest=sha256:6ed880dd31a9fddce37719a0e565d8514bd666f782c5ce850cc3f26e553ef8b5

Observation 4e2e3e8a-9e20-4273-a49b-7d7e5ac2f63b · inbound

Fast Spatial Memory with Elastic Test-Time Training cites this paper.

Fast Spatial Memory with Elastic Test-Time Training Muon is Scalable for LLM Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:27:24.750995Z digest=sha256:fb7370780a646a25cf1df702a112d86acbfaa954c8edfd6b771b1e97d8119d8b

Observation 3c1893ac-f410-4a89-9ac8-1be1d80cfc0d · inbound

PRAGMA: Revolut Foundation Model cites this paper.

PRAGMA: Revolut Foundation Model Muon is Scalable for LLM Training

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:37:04.381074Z digest=sha256:02cd6bca84e2caa72c1f34d1171ddd2840a01e39f3bdc509dbc1c6c69e1fac85

Observation 10dd01c6-e6fe-46d2-8ad5-ba3210215024 · inbound

Communication-Efficient Gluon in Federated Learning cites this paper.

Communication-Efficient Gluon in Federated Learning Muon is Scalable for LLM Training

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T15:16:15.739456Z digest=sha256:aa353cf12cf4431913b67045250e5221b36ef7e64a61dd1864e340b608166b68

Observation 4486f7b9-e735-4ca7-acae-986cb115f93a · inbound

ResBM: Residual Bottleneck Models for Low-Bandwidth Pipeline Parallelism cites this paper.

ResBM: Residual Bottleneck Models for Low-Bandwidth Pipeline Parallelism Muon is Scalable for LLM Training

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:36:15.760164Z digest=sha256:3e5fb5bd6d4797d9c16e4e65be09e5b569705e4d49a665365dc3a259d531f65c

Observation b17e7849-5a25-49a3-82d1-0dd5458a9638 · inbound

Benchmarking Optimizers for MLPs in Tabular Deep Learning cites this paper.

Benchmarking Optimizers for MLPs in Tabular Deep Learning Muon is Scalable for LLM Training

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T12:06:08.798712Z digest=sha256:40f85ca2b33c442541ac617e0839cbfd8532bd6a2920504803ed577cfcf79e28

Observation ebe768c2-dc29-43a0-ab34-d5be14132dbe · inbound

CityRAG: Stepping Into a City via Spatially-Grounded Video Generation cites this paper.

CityRAG: Stepping Into a City via Spatially-Grounded Video Generation Muon is Scalable for LLM Training

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T02:10:54.935053Z digest=sha256:0793aa77d9d46a97ba6ab3992c6db2973174e43f7a3eff168603cd6d997fde95

Observation f6ccf5c5-5801-4fec-99cc-a2cdc32107ad · inbound

In-context modeling as a retrain-free paradigm for foundation models in computational science cites this paper.

In-context modeling as a retrain-free paradigm for foundation models in computational science Muon is Scalable for LLM Training

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T07:06:51.187492Z digest=sha256:fcb76c1fa395f0551dcc3f0fc386a27180b23230ac379346d6ba79606650eec4

Observation f4b85007-3a3c-4b57-aba4-e5f7eef53e6e · inbound

SUDA-Muon: Structural Design Principles and Boundaries for Fully Decentralized Muon cites this paper.

SUDA-Muon: Structural Design Principles and Boundaries for Fully Decentralized Muon Muon is Scalable for LLM Training

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T02:48:35.070309Z digest=sha256:b39ae1e6c7861be456a8c3e7222ff5de4c54cf5d8739bd6faa1b3330038ba16a

Observation 52f23ffd-fd96-4273-ac76-49e58547b722 · inbound

SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs cites this paper.

SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs Muon is Scalable for LLM Training

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T04:41:52.098355Z digest=sha256:0e4bb1ce73fbae39723cedcd3f0c1b692a4149ef206c9b97c2abd2fe494c0ac7

Observation f0339a5b-2a53-436f-8ba0-958d64dc4e3f · inbound

Model Merging: Foundations and Algorithms cites this paper.

Model Merging: Foundations and Algorithms Muon is Scalable for LLM Training

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T14:47:58.710291Z digest=sha256:c2831b017063fa5a8bad99e1be0399ae3182cef25103713c9002a10d9fa2fb2c

Observation 72d92fad-fd70-4650-b4bc-2c7a29ee42b5 · inbound

DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs cites this paper.

DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs Muon is Scalable for LLM Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:57:35.983246Z digest=sha256:f9bc56dd260fae520b73db949c0c6cbce2b8f7704e257d6ac3993e47b67ef1e7

Observation 7345c383-cb32-498a-bdc1-381d3660127f · inbound

Nora: Normalized Orthogonal Row Alignment for Scalable Matrix Optimizer cites this paper.

Nora: Normalized Orthogonal Row Alignment for Scalable Matrix Optimizer Muon is Scalable for LLM Training

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:36:18.997820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:42:34.498806Z digest=sha256:60d6a335c099faaa92c7886853e80c976d2617420c222ea00aacb5a8e0bd3d24

Observation a8c2ae73-8b1d-4dcc-8fc9-bdcec59f3105 · inbound

Budget-aware Auto Optimizer Configurator cites this paper.

Budget-aware Auto Optimizer Configurator Muon is Scalable for LLM Training

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T17:46:09.673049Z digest=sha256:2c1ffaaf77e9e65222568900de495dbe556aeaf59547f9da278357c12b5a2622

Observation 2c3cf7ab-8c86-423c-a0d9-653851ae2b25 · inbound

ZAYA1-8B Technical Report cites this paper.

ZAYA1-8B Technical Report Muon is Scalable for LLM Training

Reference 199

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T17:36:37.182196Z digest=sha256:dd3730b4d46942481de9dd7c1a2e9b5e5f9b5bf09201bd20898493227cb1f8fb

Observation adb25634-3da3-41c4-82f9-6ccb3d6c1d4b · inbound

Autoregressive One-Step Generative Modeling for Dynamical System Forecasting cites this paper.

Autoregressive One-Step Generative Modeling for Dynamical System Forecasting Muon is Scalable for LLM Training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T14:48:38.107956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:48:38.107956Z digest=sha256:0601d3879dfcb09a8e28abbd77a7006df07d15e205a504b0756b0e7a040590ca

Observation def732c4-3ec1-4ca9-ba36-44d39dbae4b7 · inbound

Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio cites this paper.

Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio Muon is Scalable for LLM Training

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:39:51.611115Z digest=sha256:59a8a9cc9f3d18cddf519d685449544dd6392d19103a835a72a22d1064334c4e

Observation f1bafc29-79a4-4fea-83ef-1a532689648b · inbound

MDN: Parallelizing Stepwise Momentum for Delta Linear Attention cites this paper.

MDN: Parallelizing Stepwise Momentum for Delta Linear Attention Muon is Scalable for LLM Training

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T15:27:55.566795Z digest=sha256:9970048ec983b701f5f0149d16dbc687150b5aefb8f7d850c6fa8e2e1071d998

Observation 2eeae874-359d-4a75-986c-9e046e734539 · inbound

The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks cites this paper.

The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks Muon is Scalable for LLM Training

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T13:14:32.856570Z digest=sha256:e51a215515aa1c695b733e32f7f90595d2ee5b03ff9c78e2c44962ee061af26d

Observation 4ac87229-543a-4618-82ed-a76d92920cdc · inbound

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less cites this paper.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Muon is Scalable for LLM Training

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:0aa19c0216e5c65654f0c5e10496013d61a7fb3f799420672599f5d2982332b6

Observation fd89786f-fda9-40b4-830d-05dc7b5b46ed · inbound

Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition cites this paper.

Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition Muon is Scalable for LLM Training

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T00:54:23.683013Z digest=sha256:ca7b51dcb621596111e86f96c6dee3fbbbcc3ede29ff83e427ed3f3f2b5bc406

Observation 9a27d841-7790-445e-8685-b04aa426701e · inbound

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs cites this paper.

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs Muon is Scalable for LLM Training

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T03:40:04.692279Z digest=sha256:f71212b25fed52512b349b937c432625e91b5e025f082bd093724f97ca200ada

Observation 6ad88bc1-6ada-4333-bd01-da2f5a247fd2 · inbound

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs cites this paper.

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs Muon is Scalable for LLM Training

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.939073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:14:55.858466Z digest=sha256:12eee13f514c1876c9cff32b382be4f5b70edd8f5112a0c6f2e1e90f7c2d8b5c

Observation 1f996a58-1a1a-4d0b-90ff-130a5dc0fcec · inbound

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling cites this paper.

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling Muon is Scalable for LLM Training

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:47:08.766380Z digest=sha256:1ad121170273046c7666565b8678bc3eb87d7a0eaa7907b1addf5772eef50f86

Observation 9feeee59-8ee3-4639-9e50-378fba61f74e · inbound

ZAYA1-VL-8B Technical Report cites this paper.

ZAYA1-VL-8B Technical Report Muon is Scalable for LLM Training

Reference 131

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:21:23.438502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:15:16.607346Z digest=sha256:aece6aea3ad0f7cdf15f02433de459695e7403fd9dd3b3e209efbd574cb67164

Observation 439cb60a-8571-43c0-b3a0-1804fe07baeb · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI Muon is Scalable for LLM Training

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:21:25.238346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:13:35.990078Z digest=sha256:c02d19dc323187ea2375ca24fa61e486151719f2468cf086f58ae78c88d50a21

Observation 6f01e845-5ad5-4b83-8a4d-e6fbf1b43089 · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI Muon is Scalable for LLM Training

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:25:46.043477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T23:12:57.154537Z digest=sha256:d4c4af02828ce6491102725988c46b8c21469cb0aee00065cff0006bb670b4c2

Observation 4157f7d6-3182-4674-9842-4694dd87c5f1 · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI Muon is Scalable for LLM Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T17:14:49.310598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:14:49.310598Z digest=sha256:746bb831cf36feb2c38e0158e1d36d0de8397ace3fdb7edb6df0f7e9856cae79

Observation 082b43f8-c05d-4c48-8c34-e220d98336f2 · inbound

Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning cites this paper.

Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning Muon is Scalable for LLM Training

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:42:06.515065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:12:25.713592Z digest=sha256:ad23109e6e5dec328147198fa1a633b3f09db73d5df81268087bb5c791c6fb08

Observation ed53cae7-0045-4aea-a096-2cf0aa219e72 · inbound

Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning cites this paper.

Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning Muon is Scalable for LLM Training

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-19T15:12:37.512690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T15:11:43.827012Z digest=sha256:b124d4c3e9e57144f109abf34043bc8de57195bc1cac333eb0f2dff49c7464f9

Observation 489a2f61-411b-402e-9ed1-1d257368a913 · inbound

Accelerating Zeroth-Order Spectral Optimization with Partial Orthogonalization from Power Iteration cites this paper.

Accelerating Zeroth-Order Spectral Optimization with Partial Orthogonalization from Power Iteration Muon is Scalable for LLM Training

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:33:09.431089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T17:33:00.747846Z digest=sha256:4063150d09c7d2ddd18d030af99fb0dc26e28469b8797e3f1ac846c884e48376

Observation 588618a4-18ba-487b-b379-860d76eff338 · inbound

Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers cites this paper.

Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers Muon is Scalable for LLM Training

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:41:45.320350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:01:32.057022Z digest=sha256:bdd661af8027d924bad04c8fd574f07a60bfdb8da8bc607dce47a2da33c6f803

Observation cd84be3b-6c64-4028-b4bf-6b52a458b634 · inbound

Dimension-Free Saddle-Point Escape in Muon cites this paper.

Dimension-Free Saddle-Point Escape in Muon Muon is Scalable for LLM Training

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:26.345270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:38:25.480020Z digest=sha256:ba55d0fa1fc85ec5d74101ce17eb620f2402f74af1c668ad355fd51b72e97ea8

Observation 07adcf44-0419-4878-ae91-8d2211613b70 · inbound

Phases of Muon: When Muon Eclipses SignSGD cites this paper.

Phases of Muon: When Muon Eclipses SignSGD Muon is Scalable for LLM Training

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:41:29.316395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:05:08.899876Z digest=sha256:af5e2a13eec9ad8c40256a300583385ade9be3d7b7ae8d935fc821d216c1d01a

Observation 6cf5ea67-3b49-4e40-9ffd-6a597ae35eb4 · inbound

Can Muon Fine-tune Adam-Pretrained Models? cites this paper.

Can Muon Fine-tune Adam-Pretrained Models? Muon is Scalable for LLM Training

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T06:51:28.739353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T03:53:11.469583Z digest=sha256:35f59a66fd1eb0a30dcd659a00902af7a8902861a6a3edbd816a893d99fa7483

Observation c7a76cb0-7019-417f-8036-e07bb5e72db2 · inbound

Mela: Test-Time Memory Consolidation based on Transformation Hypothesis cites this paper.

Mela: Test-Time Memory Consolidation based on Transformation Hypothesis Muon is Scalable for LLM Training

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:46:33.348726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:01:01.993926Z digest=sha256:e781d94c67b77bba67bd7daec906531877647d4d26007ac7cb665010a5f9f51b

Observation 83500b89-85a4-4ee5-939f-70c314f66592 · inbound

Uniform Scaling Limits in AdamW-Trained Transformers cites this paper.

Uniform Scaling Limits in AdamW-Trained Transformers Muon is Scalable for LLM Training

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:27:01.950220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T01:24:52.640510Z digest=sha256:a619ccc77bbda4b1599cd8223b0e727bbe1c22531f730117a82feb39c7ee920a

Observation d81fa253-3fc1-43f7-bad7-fd6cadbb506d · inbound

MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization cites this paper.

MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization Muon is Scalable for LLM Training

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:12:07.087426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T02:11:43.022143Z digest=sha256:8f18c3d45877acc98b826284e4444d7bf4773d771f2fe09d1773b3f782f4ba80

Observation 5f798e7d-9033-465e-a6c4-aad6ff2ee2a9 · inbound

MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization cites this paper.

MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization Muon is Scalable for LLM Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T14:24:48.118761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:24:48.118761Z digest=sha256:00839567ade358bf028a7a944e5dbdee4bd277080023758f7e60610635e68d29

Observation 7b127981-0878-415d-b3c2-ea2a82116c00 · inbound

Elastic Attention Cores for Scalable Vision Transformers cites this paper.

Elastic Attention Cores for Scalable Vision Transformers Muon is Scalable for LLM Training

Reference 165

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:07:22.858225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T06:02:40.158866Z digest=sha256:46b0258e3caf5d58388f631d06a2bded7c538cf41aae3a5310236bb42c04d038

Observation 8709b56d-e829-4d26-9752-9b5e0612d388 · inbound

Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation cites this paper.

Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation Muon is Scalable for LLM Training

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-13T04:57:17.626506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T04:53:52.898843Z digest=sha256:f231e2d6ddfafcd74f02c9d3699755464e6e4c07eca255dc3840b15a6dc5d0c4

Observation 475a0326-70d3-4ead-bdf2-0facbf298f23 · inbound

Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence cites this paper.

Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence Muon is Scalable for LLM Training

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:57:53.130865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T19:56:23.124500Z digest=sha256:37e8de31a177c2b2417c46b0e861a3e267c9f507ebd085122d79598f9f24b3b9

Observation 3b2c0eb8-8870-4042-ae4a-d987ab929428 · inbound

$\phi$-Balancing for Mixture-of-Experts Training cites this paper.

$\phi$-Balancing for Mixture-of-Experts Training Muon is Scalable for LLM Training

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:22:40.144819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T16:16:38.731203Z digest=sha256:54f990f3a41610403b7e4e0e027fce0acfa643c11406486a8a6d2e5730b2d71c

Observation 3496cfcc-8c8d-40a8-b583-6e4fe1791b34 · inbound

Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered cites this paper.

Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered Muon is Scalable for LLM Training

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T21:23:44.548558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T21:19:55.074853Z digest=sha256:e3d87fe5b6421e16a3d3ef52ee5613ec70cbc923df5c01328569eb90d178abae

Observation e53c553a-4335-4195-8e83-eb13908704c2 · inbound

Towards Human-Level Book-Writing Capability cites this paper.

Towards Human-Level Book-Writing Capability Muon is Scalable for LLM Training

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:43:27.011512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:40:53.245419Z digest=sha256:56567240319ab403479c6615008299e83c701c277743be2b84c3a04c4e6c4bda

Observation 98b1ed54-b265-4213-b9a4-c1ab2b4cac64 · inbound

Towards Human-Level Book-Writing Capability cites this paper.

Towards Human-Level Book-Writing Capability Muon is Scalable for LLM Training

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.331587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:27.257098Z digest=sha256:8fbfda7456d842542470bd9baaa0a5b586af35f387f1be1df6c4ffd4c1fae10f

Observation 2f910442-65a7-4948-ba07-b30c69b04633 · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers Muon is Scalable for LLM Training

Reference 105

Resolution
verified exact
local_arxiv, observed 2026-05-20T09:38:11.113200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T09:34:45.186929Z digest=sha256:9fa8a86e62aeb08355e61ca99ad1e1619cc6f394201703d6aeebaad4383657ce

Observation ae9030a1-d890-4c16-9497-c7003ec2333d · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers Muon is Scalable for LLM Training

Reference 107

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.342917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T18:42:01.854481Z digest=sha256:6d9172d1d7be635d7688e56f926ddc0310933eb1fb82f5e6693aa2d22e278a56

Observation 40c85e73-8832-41fb-883f-bd2ddc3c322f · inbound

Scale-Invariant Neural Network Optimization: Norm Geometry and Heavy-Tailed Noise cites this paper.

Scale-Invariant Neural Network Optimization: Norm Geometry and Heavy-Tailed Noise Muon is Scalable for LLM Training

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.630939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-30T18:27:40.390908Z digest=sha256:19d5740d7826cd1fcf8fe4197ba8ac36e213c7b05f535fec4e8256b8d13b8a67

Observation 0da511cb-f4d0-4551-8171-f95bd4a1c45d · inbound

Distance-Aware Muon: Adaptive Step Scaling for Normalized Optimization cites this paper.

Distance-Aware Muon: Adaptive Step Scaling for Normalized Optimization Muon is Scalable for LLM Training

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:28:17.207016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T12:24:07.167770Z digest=sha256:b75d047a1af3876c5d4a1483b6a5eca6bbe397f23b03c11a162c8d955db3c97b

Observation e7a935f2-eb42-431a-a562-2c26e1fa6b52 · inbound

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR cites this paper.

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR Muon is Scalable for LLM Training

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:18:07.210100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T07:14:31.613251Z digest=sha256:eac73d84350159a6de3fd9369193b58be2daef19c0cc496a7830a974ff548d97

Observation 7e5e0ac7-5f69-452b-b20c-7eea6194db59 · inbound

MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models cites this paper.

MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models Muon is Scalable for LLM Training

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-20T08:13:08.317153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T08:10:14.711735Z digest=sha256:5f0dbb4b42b196df5eff843b39231c610cf2a7707f5bd21daad99e41b06430b5

Observation 8568f2cf-0fef-488a-9257-39c79db255ef · inbound

LionMuon: Alternating Spectral and Sign Descent for Efficient Training cites this paper.

LionMuon: Alternating Spectral and Sign Descent for Efficient Training Muon is Scalable for LLM Training

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.808665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T07:24:55.516803Z digest=sha256:08d625e3f9cb069142b49e4c3099afe6e49ad731b1ef3562214ec5c0c3a04fdf

Observation 9230e505-d52a-4291-a4c4-cd2e6caf4c2d · inbound

LionMuon: Alternating Spectral and Sign Descent for Efficient Training cites this paper.

LionMuon: Alternating Spectral and Sign Descent for Efficient Training Muon is Scalable for LLM Training

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:35:00.448741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T18:30:51.719396Z digest=sha256:2366eeee374aca4fef0d185ac08b8b15988359023a19a52d2301db16bde1cefb

Observation 2d95b3b5-d4f1-44d4-b3f1-e1c51dc799c2 · inbound

Toto 2.0: Time Series Forecasting Enters the Scaling Era cites this paper.

Toto 2.0: Time Series Forecasting Enters the Scaling Era Muon is Scalable for LLM Training

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:43:05.682831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T06:42:13.688042Z digest=sha256:abacc2a2bb72393cfabcbbd7c46d12e32d95b8411e3971a67295e3c298067953

Observation 66207ba8-a4e4-4d7a-984d-2f4aea0ddd59 · inbound

Toto 2.0: Time Series Forecasting Enters the Scaling Era cites this paper.

Toto 2.0: Time Series Forecasting Enters the Scaling Era Muon is Scalable for LLM Training

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:14:59.910573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T18:11:53.058900Z digest=sha256:83f0bb1c86698bd4590904eafc25795b07b0846d23eda6075bef56bfc23a5fd2

Observation 4483ef33-44ec-4757-b69a-f26d42275e14 · inbound

TBP-mHC: full expressivity for manifold-constrained hyper connections through transportation polytopes cites this paper.

TBP-mHC: full expressivity for manifold-constrained hyper connections through transportation polytopes Muon is Scalable for LLM Training

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:01:23.422546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T10:00:07.226252Z digest=sha256:21237193a0ef999af78f24f2584e09d2a9cb6fe6f156465cc7d5d475661a8978

Observation ca2c268b-d528-4cc1-a53c-6b96195ed89b · inbound

Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws cites this paper.

Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws Muon is Scalable for LLM Training

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:51:18.227053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T08:49:32.864065Z digest=sha256:2a4d1deb9df6bab15891cac47270067a92e9edf52fee2d71fcc24d16ff9e2b31

Observation 42973ad1-9341-4004-aa92-3496f450b54c · inbound

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs cites this paper.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Muon is Scalable for LLM Training

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:44:42.659994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T07:44:16.677054Z digest=sha256:c7844d000f05dd80ca2519285286d280ccdd19843bd34293aa8f43d4693fda47

Observation 400b49a7-d530-4f35-8c50-7b2f34679cef · inbound

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs cites this paper.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Muon is Scalable for LLM Training

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.604761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:f2e14e314bd5e385b67bdf57cc3d8fdc9c9bb2fddd359530b8d86d6f0c47eaad

Observation b2bd861a-3e82-474f-a63a-6073cd79519e · inbound

Anytime Training with Schedule-Free Spectral Optimization cites this paper.

Anytime Training with Schedule-Free Spectral Optimization Muon is Scalable for LLM Training

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:40:24.249511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T05:38:16.958574Z digest=sha256:75036c43f7ff1b72b22aec9cc789d4fff9fccef8222d8160b0530eb12428d5b4

Observation a6d0dd5a-0a70-42c3-b00a-5f59378a7a05 · inbound

Training-Free Looped Transformers cites this paper.

Training-Free Looped Transformers Muon is Scalable for LLM Training

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:36:36.720763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T04:36:10.278337Z digest=sha256:32c9680995ce116d2a8f2398dd7f1136ce1042624155a9725c581d8b493275cb

Observation 454d169b-00e5-4c05-bb1d-dd07503770a9 · inbound

Learning Laplacian Eigenspace with Mass-Aware Neural Operators on Point Clouds cites this paper.

Learning Laplacian Eigenspace with Mass-Aware Neural Operators on Point Clouds Muon is Scalable for LLM Training

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T14:54:45.497577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T14:50:49.368332Z digest=sha256:4b2d6610f7e62661e3d512677d6f746c9fe337c9555ff8c83b90a4ca4e9957bf

Observation f7482c3e-6e45-4439-a5fa-357a90e8cd11 · inbound

Mapping the Schedule x Bit-Width Boundary in Sub-100M Quantisation-Aware Training cites this paper.

Mapping the Schedule x Bit-Width Boundary in Sub-100M Quantisation-Aware Training Muon is Scalable for LLM Training

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T23:14:01.500246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T23:10:47.199537Z digest=sha256:ee99ba41d525fe6ee4e0e3f404f05fb5cca38ae82e041cafc18bbe11cbce28c7