Pith. sign in

Paper Citation Record · LEDGER

Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2503.12645.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.12645 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T01:10:04.430235Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:29:38.574824Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 71fcfa0c-e72a-4410-af0d-4f238ce0ef33 · inbound

Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs) cites this paper.

Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs) Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:20:53.821206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:20:53.821206Z digest=sha256:0b9f4bca8999a1966ad4274e6fb36e48a598fe64ce2fa5bca15fca46fb69d9a3

Observation a403ee15-b40b-461e-9cb3-2621c37bd61e · inbound

On the Convergence Analysis of Muon cites this paper.

On the Convergence Analysis of Muon Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:22:19.190615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T13:19:37.035526Z digest=sha256:722d0ab157329a5c6840b14d345e91648e96b7ba5a70429e2f47a0ff683b6415

Observation 9c7a278f-ce05-4ead-bf51-e1a07dabc265 · inbound

On the Convergence Analysis of Muon cites this paper.

On the Convergence Analysis of Muon Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:19.083274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:19.083274Z digest=sha256:01e24ad4fb838a7bac3a7f842d196c1d22a05b7174fecd27c13602df6cdb0da7

Observation 7e55803d-1621-4fa3-9926-13f92e574a0b · inbound

SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration cites this paper.

SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:08.503792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:08.503792Z digest=sha256:32dc73a74d1b84d44ff0e7c3fe26f0497262e0432f21bea303cd2900b9ae94a4

Observation 5404d0a2-7c9d-475b-8089-0d0e6fa8232f · inbound

Convergence Bound and Critical Batch Size of Muon Optimizer cites this paper.

Convergence Bound and Critical Batch Size of Muon Optimizer Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:10.260807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:57:10.260807Z digest=sha256:7ad3b23c48c13cc0a9269132c37af7e2d7a87d4aec9d1b232e12c64b30af0d63

Observation eee4ad59-1abc-4558-9daf-6ec7edd85f4e · inbound

Simple Stepsize for Quasi-Newton Methods with Global Convergence Guarantees cites this paper.

Simple Stepsize for Quasi-Newton Methods with Global Convergence Guarantees Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 1993

Resolution
unresolved
no resolver link, observed 2026-08-05T15:43:57.016425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:43:57.016425Z digest=sha256:1cdee0b83fbd4d0c3a88983a45f680d52880d57ebc9b196ebe224084f2645b24

Observation 2a56226a-725f-4d22-b5f4-0a573f71370d · inbound

AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates cites this paper.

AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T11:25:45.016558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:25:45.016558Z digest=sha256:7e98e2f8df0e065280f2bc7cc47de4591f4d37a7b7475f3a138c717ebab183b7

Observation b279ccfb-6373-4faf-8821-92c8a85b74d2 · inbound

Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training cites this paper.

Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:51:33.991661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T15:48:43.422716Z digest=sha256:494bcee18af0fa6beae440b332a2fb55c26875eb1c3e08645881d300829593f2

Observation b59ea046-1d03-4ce1-a722-85a675230cfc · inbound

LiMuon: Light and Fast Muon Optimizer for Large Models cites this paper.

LiMuon: Light and Fast Muon Optimizer for Large Models Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-15T15:57:38.243210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:57:38.243210Z digest=sha256:3bd44e95b6537c4498ab24491701c67cdb3d9645014e03bdc8d4703910bbcedd

Observation 57eb9d9d-2d5d-4dea-8cf8-b3e09d8adc77 · inbound

DeMuon: A Decentralized Muon for Matrix Optimization over Graphs cites this paper.

DeMuon: A Decentralized Muon for Matrix Optimization over Graphs Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T13:25:32.657347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:25:32.657347Z digest=sha256:d97bd2c025efd3f91319dd39d263873286a9083a656f8a55395849bef5f5ce13

Observation 31fac1bf-6fa3-4f1a-920a-03ea72ffd199 · inbound

Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates cites this paper.

Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T22:20:05.564818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T22:20:05.564818Z digest=sha256:31279768a7fd60323c736d2af7d97b3d78278a4ce2645f5ec282193374952447

Observation 32aafd05-c463-4274-82c8-177a95cc14d9 · inbound

Muon in Associative Memory Learning: Training Dynamics and Scaling Laws cites this paper.

Muon in Associative Memory Learning: Training Dynamics and Scaling Laws Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:14:15.087140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:14:15.087140Z digest=sha256:1a1d413d8a0485b24baa2aa84297a2948e7cf9ceefda5e54b18ef2674cb0a2e5

Observation 08bf6a57-2fa3-4747-afa6-add27c1398bb · inbound

MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training cites this paper.

MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:26:31.356452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T19:24:22.699712Z digest=sha256:72cd0995344d30baf7e8d158019026272831859435160017541ac0296b0644aa

Observation d4f427ad-ab0b-44fe-9324-88d058f61489 · inbound

Optimal Projection-Free Adaptive SGD for Matrix Optimization cites this paper.

Optimal Projection-Free Adaptive SGD for Matrix Optimization Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:58:15.724809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T20:56:41.466669Z digest=sha256:1dde4b81aea708536fb9b688c25d1ae4d7c85368c0a4002bc55946ae9ef6f171

Observation 5644bc27-0c90-4460-8b7f-13fdffb925a8 · inbound

Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning cites this paper.

Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T10:41:51.251736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:41:51.251736Z digest=sha256:eeea98134f0579d6de47ff17bb8dfc83a26c09627a45296359ddf173c979064c

Observation db3811a1-8987-466a-a078-189239afeba2 · inbound

Communication-Efficient Gluon in Federated Learning cites this paper.

Communication-Efficient Gluon in Federated Learning Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:56:03.238852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T15:16:15.739456Z digest=sha256:708c6d1b338867be290a6fea976091fbbc94bb6c3ddf903cb05b82ccd7ce37f0

Observation 68c33c2a-8056-46bc-818c-01ca6e0e7086 · inbound

Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition cites this paper.

Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:56.679583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T00:54:23.683013Z digest=sha256:cc4e723be0e8834b773753b909760b4b9844438f37a816ae9d28f697cea23f95

Observation c5e23c37-b1b4-4987-a14c-ec5eacfec255 · inbound

Dimension-Free Saddle-Point Escape in Muon cites this paper.

Dimension-Free Saddle-Point Escape in Muon Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:01:26.384083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:38:25.480020Z digest=sha256:3a65733460207a8fb4692f2212e36ec72d6f3ecb4ea75cf2247e52423164b12f

Observation f84eac50-3db7-4eda-937f-962f38444cd5 · inbound

Can Muon Fine-tune Adam-Pretrained Models? cites this paper.

Can Muon Fine-tune Adam-Pretrained Models? Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:28.816677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T03:53:11.469583Z digest=sha256:6c19ff847bd9cb99e26b476824431088ec38814ee19157fa0a4c0a7ee05ce239

Observation 2da3c9a3-fd34-4e5b-8432-e89c744aaf97 · inbound

Muon is Not That Special: Random or Inverted Spectra Work Just as Well cites this paper.

Muon is Not That Special: Random or Inverted Spectra Work Just as Well Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:17:13.744761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T04:16:48.332851Z digest=sha256:faf3da2245c940604131a8aa14b67a10ba154879e9ec79027fd05d8370bfb871

Observation 27324d75-d6d0-405f-a281-84a770635039 · inbound

Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives cites this paper.

Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:37:19.199758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T05:34:48.195468Z digest=sha256:6229faabfcfee9266b4b835e381e80f0ed3774a62cd0d60d60d7e3acaf82e3a1

Observation cc020d3c-5332-471f-b36b-22789052180e · inbound

Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence cites this paper.

Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:57:53.124895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T19:56:23.124500Z digest=sha256:350cb83c74b2f32a522b45adf52f196ef4a48fb773319542b7cba2148e0c0e1b

Observation fc703698-597a-43f9-96c4-149c0f5d0827 · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:38:11.144993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T09:34:45.186929Z digest=sha256:bed34035c057249485ab6b1efec029f86f957e5e2a8d17f5e7f941d1bace504e

Observation dd6c9bb4-3fd9-4e3d-a8c6-c59268ee4379 · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.379916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:42:01.854481Z digest=sha256:a68098ec4fd015049db511b941ed05be0509079ff10cd8c09db4b6ad18472599

Observation 3ae51170-44fa-4757-bf12-cac193dd8c65 · inbound

Ringmaster LMO: Asynchronous Linear Minimization Oracle Momentum Method cites this paper.

Ringmaster LMO: Asynchronous Linear Minimization Oracle Momentum Method Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 291

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:13:18.539911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T13:08:52.912250Z digest=sha256:ad166eeca9aa3eeb055f69ed6ee87a8a779e1c40d17bacfea242f6b98ea09665

Observation a48fa612-1c17-4b82-9c88-de12dc58630e · inbound

Scale-Invariant Neural Network Optimization: Norm Geometry and Heavy-Tailed Noise cites this paper.

Scale-Invariant Neural Network Optimization: Norm Geometry and Heavy-Tailed Noise Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:55:48.633794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-30T18:27:40.390908Z digest=sha256:7f314884681c0f400ec82cfa35cf52ad7d89e2934a8fea4d36ae8f4cb68db28f

Observation e5bdcb5e-c7a0-4ad1-9652-0246994ced76 · inbound

MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models cites this paper.

MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:13:08.332905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T08:10:14.711735Z digest=sha256:387aa1421e0a5f5799565c006341350edb06e9e158aee348b2861090fffd8dfc

Observation f21999a9-1d98-463a-8383-c83f55c8bfbf · inbound

LOSCAR-SGD: Local SGD with Communication-Computation Overlap and Delay-Corrected Sparse Model Averaging cites this paper.

LOSCAR-SGD: Local SGD with Communication-Computation Overlap and Delay-Corrected Sparse Model Averaging Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 292

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:49:40.639082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-21T05:49:28.713982Z digest=sha256:8ff3fb0a6d191c27dde4acb7a9f3d40e36a71da4b815d2f95f2e785e336c31b4

Observation a9e57df4-e912-4f7a-a6c6-13c2da87d407 · inbound

Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer cites this paper.

Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:56:33.542117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-25T02:56:18.035759Z digest=sha256:482f2c8ae5b7eac45c87f83d44ac20b0565d8dd3ded5cf5ca19b301ecb612bb0

Observation d1b21808-3b1a-460c-a8c0-61c11dbe8459 · inbound

Denoise First, Orthogonalize Later: Understanding Momentum in Muon via Spectral Filtering cites this paper.

Denoise First, Orthogonalize Later: Understanding Momentum in Muon via Spectral Filtering Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:56:27.803949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T11:24:38.292078Z digest=sha256:b579cff9a4208d4197a499886eeb5c95acc28c4697408a4cd5e04be605e6a90a

Observation d15c3da6-790d-4fde-9e38-517bc61fda1d · inbound

Why Muon Outperforms Adam: A Curvature Perspective cites this paper.

Why Muon Outperforms Adam: A Curvature Perspective Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 158

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:16:44.401048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T07:04:21.012269Z digest=sha256:798b93904512eb13969acfa1de88dca97e0ccc6327755880089baeda3f9c5d33

Observation a6d4c29e-6ceb-4b6d-b6e2-277b35a6c2aa · inbound

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss cites this paper.

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:56:55.972612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T02:35:39.845487Z digest=sha256:9e13999c1b6743892b7d83dbeb61b32cb72b0814246196e0870c186b1f896824

Observation 9a7b1be7-92a7-45b0-8396-c61ef473872c · inbound

Muon Learns More Robust and Transferable Features than Adam cites this paper.

Muon Learns More Robust and Transferable Features than Adam Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:30.190097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T17:08:30.717799Z digest=sha256:67e2c72fa5725e63ff4a78ae2fdf2d6fee5aa89d5db7345e0237d1960dd7e8d7

Observation a83a9502-bad2-4dd0-ac1c-c4ee563acc12 · inbound

Restart and Adaptive Acceleration in Stochastic Gradient Methods cites this paper.

Restart and Adaptive Acceleration in Stochastic Gradient Methods Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:59:37.581655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T14:00:19.385619Z digest=sha256:9537e5b2866a4bbb75690dfa270b42b3ba198e1596ca12b528c90c09e3cad9e2

Observation 27877e57-6a11-4f20-8e68-4525bcd92350 · inbound

Convergence Analysis of Muon-type Methods with Inexact LMO in the Degenerate Case cites this paper.

Convergence Analysis of Muon-type Methods with Inexact LMO in the Degenerate Case Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:29:38.576286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T13:28:36.618788Z digest=sha256:b911562f0f6fd8e0aba938fd3e555472a5634ad0aadd05fe7a4fb49bab51388a

Observation d81577a1-d2f5-4d0a-a6e7-1d7e99935950 · inbound

One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining cites this paper.

One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:44:18.874749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T06:41:55.732230Z digest=sha256:2d93893832f44feb90cf55902dcee45eeb522e6e157ea86c880d0a367f3321f1

Observation 90e3855b-ec5b-4461-bb0d-7df4849b25c5 · inbound

Muon as a Residual Connection cites this paper.

Muon as a Residual Connection Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:47:05.787269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T15:39:06.013130Z digest=sha256:89a769996403a441b3bd68a33216e93b71f10338ed105364b2de9be7c204aa36

Observation 0fdde877-86a3-400f-b3c7-77b547e89579 · inbound

How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size cites this paper.

How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:57.594636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-03T21:02:31.246432Z digest=sha256:85e10081c4535a70af114d58928266777eceec312e1f6c83d2500c19ac261714

Observation 286bd7cf-619e-4dde-b067-ce118df92e4b · inbound

Reassessing Muon for Matrix Factorization cites this paper.

Reassessing Muon for Matrix Factorization Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T05:49:20.095841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:49:20.095841Z digest=sha256:a8b7629b2af0b7912960e00bcb066954d444f6cc9bd064804f372568a756d67d

Observation 83163206-68a5-4987-a452-0f8b275f4bbc · inbound

Reassessing Muon for Matrix Factorization cites this paper.

Reassessing Muon for Matrix Factorization Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-02T05:49:29.127140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:49:29.127140Z digest=sha256:c3a98e6bccf168cc699d98999ece846c30b7efdccc479ab74a8b4f03639ef7a3

Observation f121ad90-f4bb-4347-be89-68b31afbfc2e · inbound

Reassessing Muon for Matrix Factorization cites this paper.

Reassessing Muon for Matrix Factorization Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T04:22:57.127991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:22:57.127991Z digest=sha256:d84ee5134c9985e3b8003563f048fe7e1c4ae4c959a2147447b01c33c040d0fd

Observation ce417a14-eb4d-400d-8945-4f731974707c · inbound

Reassessing Muon for Matrix Factorization cites this paper.

Reassessing Muon for Matrix Factorization Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T04:23:03.940072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:23:03.940072Z digest=sha256:1e3d78b9b80a3aabbb0cd2bb4da140a4e5f4b98d03f0aff6dd5c197ecc45da56

Observation 9fd80c70-3004-4e13-a6ce-3a669fbf97e4 · inbound

Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback cites this paper.

Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T02:09:01.948787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:09:01.948787Z digest=sha256:6289706e3e4ac81d42ad7fe75ef227122a777c19116daa271ec3bbc058e7d241

Observation cac7645d-81c6-4c70-a7d9-f8b90b295011 · inbound

A Continuous-Time Analysis of Smoothed Matrix-Polar Spectral Gradient Flows for Muon-Type Optimization cites this paper.

A Continuous-Time Analysis of Smoothed Matrix-Polar Spectral Gradient Flows for Muon-Type Optimization Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T18:34:20.692665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:34:20.692665Z digest=sha256:6e08413ca1ce942bbe97b5c702cbf976a301c5fc934bc8faf9f6db8c9d63128a

Observation 6192ee82-19e9-4df3-ac41-20cae44fb588 · inbound

Federated Compositional Muon Optimizer for Matrix-Wise Models cites this paper.

Federated Compositional Muon Optimizer for Matrix-Wise Models Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T01:10:04.430235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:10:04.430235Z digest=sha256:0281bdb08f8feb75f21975d8dc03d3125092d9fadb1f0a370c46d2041dedc3f0