Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2309.06497.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:35:57.949886Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 9215efe9-e270-4cbb-b857-68acc22d380f · inbound
Old Optimizer, New Norm: An Anthology A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 840b886c-4ca4-4976-9d25-026b41b288c8 · inbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1458026b-f75b-4f94-84e3-c5f9a354b5c9 · inbound
Materials Learning Algorithms (MALA): Scalable Machine Learning for Electronic Structure Calculations in Large-Scale Atomistic Simulations A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4da20aa6-8ab7-4c1e-8e21-8c47e36fb61a · inbound
Memory-Efficient 4-bit Preconditioned Stochastic Optimization A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50362495-f249-4a7d-bd41-0e1ed9c1767d · inbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ce83e7f-3c9c-4240-9c0a-ed5c29df1bf1 · inbound
On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ee00b1-8cf5-4761-bc55-477f7dd8b99e · inbound
Spectral-factorized Positive-definite Curvature Learning for NN Training A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b2fc8e5-cbe0-4a3f-9627-e9adc83f1ef1 · inbound
DHO$_2$: Accelerating Distributed Hybrid Order Optimization via Model Parallelism and ADMM A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 774ec820-ea2a-42c4-9cb4-74490d2ff5f8 · inbound
Learning by solving differential equations A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2ee34b3-b56c-4a37-a986-75b2883a309f · inbound
Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ec66fb5-b447-420e-8fef-c969e22a2760 · inbound
RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f7799bca-5bf9-4092-bb26-3a19778cd3b2 · inbound
Demystifying Manifold Constraints in LLM Pre-training A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation de53d088-01a2-481a-be0b-cd623329dde2 · inbound
Pro-KLShampoo: Projected KL-Shampoo with Whitening Recovered by Orthogonalization A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 29a8efb0-f48f-454f-8e0d-08b8ac786a8f · inbound
Phases of Muon: When Muon Eclipses SignSGD A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5e8147c0-a478-4d4c-a0a2-7757f8f6de36 · inbound
Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 27b85356-b987-4a75-ae0a-e606698c19aa · inbound
Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2e523952-ef52-4d32-9045-51b22cf1a7e0 · inbound
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 139
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8899af77-1701-440b-9ed2-a0f582976f4a · inbound
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 141
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 23a46ad9-1f96-4330-8661-9f5a74ec8462 · inbound
FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0b05e4df-5e09-4577-b05a-e3ff329b9e69 · inbound
Spectral Scaling Laws of Muon A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 14c2112f-286a-498c-ab05-649dc17b70d9 · inbound
Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0f602b61-e958-4214-81c3-0332ca7b31a6 · inbound
SOAP-Bubbles: Structured Weight Uncertainty for Neural Networks A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d4ed3575-bfbf-49d8-86df-f929d40420ba · inbound
DMuon: Efficient Distributed Muon Training with Near-Adam Overhead A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8afea9ee-df30-476b-8d85-8941f75027c9 · inbound
MatrixFSDP: communication-free matrix optimizers under ZeRO-3 parameter sharding A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6d826c50-3951-4413-bbf3-3c7d8feafd5f · inbound
Muse: Representation Geometry of Muon Beyond Normalized Momentum A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 891d468f-8722-410b-92a1-9391048fd58d · inbound
PoLoRA: A Preconditioned Orthogonalized LoRA Optimizer A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc89e38e-cb7b-4ec1-99c0-eda052520eaf · inbound
OneShot: Index-in-Ranking with Neural Scoring for Large-Scale Retrieval A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18859707-d211-499f-8c3e-d8fcb74962e9 · inbound
OneShot: Index-in-Ranking with Neural Scoring for Large-Scale Retrieval A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ec7bf56-7c37-4897-b8e9-256fc8d7d013 · inbound
Muon Meets Mamba: Spectral Optimization for State Space Models A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a0b3809-ca41-41cc-aa65-bcf638e00f14 · inbound
Second-Order Muon Done Right: A Principled Marriage of Spectral Geometry and Curvature A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.