Pith. sign in

Paper Citation Record · LEDGER

A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2309.06497.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.06497 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:35:57.949886Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9215efe9-e270-4cbb-b857-68acc22d380f · inbound

Old Optimizer, New Norm: An Anthology cites this paper.

Old Optimizer, New Norm: An Anthology A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:27:52.941479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-16T07:27:52.883335Z digest=sha256:6a1840c66daf82e02da77eadaeafd64a09e04fdd701d5b7bd2601a52cddfdcca

Observation 840b886c-4ca4-4976-9d25-026b41b288c8 · inbound

MARS: Unleashing the Power of Variance Reduction for Training Large Models cites this paper.

MARS: Unleashing the Power of Variance Reduction for Training Large Models A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.988790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.988790Z digest=sha256:c992258b9f9bc3a94d365894815123ab43b004ff1037e7137672009bc2741d4d

Observation 1458026b-f75b-4f94-84e3-c5f9a354b5c9 · inbound

Materials Learning Algorithms (MALA): Scalable Machine Learning for Electronic Structure Calculations in Large-Scale Atomistic Simulations cites this paper.

Materials Learning Algorithms (MALA): Scalable Machine Learning for Electronic Structure Calculations in Large-Scale Atomistic Simulations A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T06:03:49.313548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:03:49.313548Z digest=sha256:9697981ce506a78897f6610fdefce5cfb206efa9398f03b84dd12d485cd07d79

Observation 4da20aa6-8ab7-4c1e-8e21-8c47e36fb61a · inbound

Memory-Efficient 4-bit Preconditioned Stochastic Optimization cites this paper.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.267978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.267978Z digest=sha256:b736d94b15c23c183fbb125f49d8dbe7bf78d2463fa5b593db7a087f3f06fb21

Observation 50362495-f249-4a7d-bd41-0e1ed9c1767d · inbound

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training cites this paper.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.142246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.142246Z digest=sha256:17e99d9db0d82409b9f071677b15acc12fdc3c468db28ebcc27af54c7e718c53

Observation 5ce83e7f-3c9c-4240-9c0a-ed5c29df1bf1 · inbound

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning cites this paper.

On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:40.724257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:40.724257Z digest=sha256:2b1a821a7e94aec6847c0acb34d2158c02f53b81c78d6be7e140405cd64c4037

Observation a5ee00b1-8cf5-4761-bc55-477f7dd8b99e · inbound

Spectral-factorized Positive-definite Curvature Learning for NN Training cites this paper.

Spectral-factorized Positive-definite Curvature Learning for NN Training A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T16:15:07.336710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:15:07.336710Z digest=sha256:e39071b8641a2d003f01f9e43814543533d9816b907f6ff726a591a9d07ca205

Observation 9b2fc8e5-cbe0-4a3f-9627-e9adc83f1ef1 · inbound

DHO$_2$: Accelerating Distributed Hybrid Order Optimization via Model Parallelism and ADMM cites this paper.

DHO$_2$: Accelerating Distributed Hybrid Order Optimization via Model Parallelism and ADMM A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:35:57.949886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:35:57.949886Z digest=sha256:f8a4e15bd49e4449676c66d566add909e910e7c2fa4d6ecff71f3ad7981e88cf

Observation 774ec820-ea2a-42c4-9cb4-74490d2ff5f8 · inbound

Learning by solving differential equations cites this paper.

Learning by solving differential equations A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:22.869640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:18:22.869640Z digest=sha256:d0339110be758c1d364911ff0fd7162705d5911eeedc439e45c4d34d659b74cb

Observation a2ee34b3-b56c-4a37-a986-75b2883a309f · inbound

Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization cites this paper.

Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T11:03:40.394757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:03:40.394757Z digest=sha256:366138e3da77931b27640fbb6052d0796a8decd0b342d4da28a1759fb3e4ebe6

Observation 9ec66fb5-b447-420e-8fef-c969e22a2760 · inbound

RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization cites this paper.

RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:49:50.787770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T07:49:39.417566Z digest=sha256:35dd91ed15847388d05475beb2ab569b6457cefdf7ae2bddbd396a45c6199711

Observation f7799bca-5bf9-4092-bb26-3a19778cd3b2 · inbound

Demystifying Manifold Constraints in LLM Pre-training cites this paper.

Demystifying Manifold Constraints in LLM Pre-training A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:16:08.604600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T17:44:44.438637Z digest=sha256:fafb9e8d22c7c4b80cd9834baf8837620ed154d3ad5efd43ef49cdcc51bca22f

Observation de53d088-01a2-481a-be0b-cd623329dde2 · inbound

Pro-KLShampoo: Projected KL-Shampoo with Whitening Recovered by Orthogonalization cites this paper.

Pro-KLShampoo: Projected KL-Shampoo with Whitening Recovered by Orthogonalization A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:01:12.783437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T13:06:05.234216Z digest=sha256:8e0f3554261ce9e3e46cdb338d052a510ce04f8a63f359c104f19f368677a44c

Observation 29a8efb0-f48f-454f-8e0d-08b8ac786a8f · inbound

Phases of Muon: When Muon Eclipses SignSGD cites this paper.

Phases of Muon: When Muon Eclipses SignSGD A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:29.878617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:05:08.899876Z digest=sha256:b129b9a0144055fe1f6f900faf25e8349c6c51b68d199b0675ab652d60841f46

Observation 5e8147c0-a478-4d4c-a0a2-7757f8f6de36 · inbound

Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models cites this paper.

Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:03:39.552743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T19:02:02.266052Z digest=sha256:75928de45b28fb928461cc186d13fe46ce78dde44ea2dd547ca9d698eb665149

Observation 27b85356-b987-4a75-ae0a-e606698c19aa · inbound

Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training cites this paper.

Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:32:42.896458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T18:30:24.657592Z digest=sha256:9f4a1af0d4be0d99c50e9a5cba8d9014152bab77d60a239fd3dd26e0455ee053

Observation 2e523952-ef52-4d32-9045-51b22cf1a7e0 · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 139

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:38:11.084495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T09:34:45.186929Z digest=sha256:3fcae5bf92efda30e9f591103784e66688b579d5e1098b46328515ccaf7fd828

Observation 8899af77-1701-440b-9ed2-a0f582976f4a · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.341379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:42:01.854481Z digest=sha256:56ec7fe5d02fcfaaaa028ac213c4d1407254f172248214f4c2bc0dd481eb8701

Observation 23a46ad9-1f96-4330-8661-9f5a74ec8462 · inbound

FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo cites this paper.

FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:16.001137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T15:49:18.685160Z digest=sha256:42e282734ab167029cb405c44be913a1bb820f6d64873b3d4c5eca63ee9dc007

Observation 0b05e4df-5e09-4577-b05a-e3ff329b9e69 · inbound

Spectral Scaling Laws of Muon cites this paper.

Spectral Scaling Laws of Muon A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:06:26.379608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T11:19:05.939340Z digest=sha256:23475f02d245fbd07ed9e21c496e29d91252a91ea2f318d5ec4c7b2ff974ca02

Observation 14c2112f-286a-498c-ab05-649dc17b70d9 · inbound

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss cites this paper.

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.400219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T02:35:39.845487Z digest=sha256:ee48b2b7220cee805dccd0738b47ee7fb4893a74ff877d97f2d198b52ebdebaa

Observation 0f602b61-e958-4214-81c3-0332ca7b31a6 · inbound

SOAP-Bubbles: Structured Weight Uncertainty for Neural Networks cites this paper.

SOAP-Bubbles: Structured Weight Uncertainty for Neural Networks A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:19:47.219672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T09:00:50.315383Z digest=sha256:2f2db439b2a653a8a82f8cbdc48415fabb98451809f700fd813e921010d4d5ea

Observation d4ed3575-bfbf-49d8-86df-f929d40420ba · inbound

DMuon: Efficient Distributed Muon Training with Near-Adam Overhead cites this paper.

DMuon: Efficient Distributed Muon Training with Near-Adam Overhead A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:39:58.558807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T02:41:02.917064Z digest=sha256:cfbb7750402387540676316c77bf369d20f27dbc8b0b652274a63a735cfefefe

Observation 8afea9ee-df30-476b-8d85-8941f75027c9 · inbound

MatrixFSDP: communication-free matrix optimizers under ZeRO-3 parameter sharding cites this paper.

MatrixFSDP: communication-free matrix optimizers under ZeRO-3 parameter sharding A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:35:37.643914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T21:33:20.705837Z digest=sha256:1bb6d39f4579f79286e1340f4f033c567ad62e09d03526d8b411e15734810109

Observation 6d826c50-3951-4413-bbf3-3c7d8feafd5f · inbound

Muse: Representation Geometry of Muon Beyond Normalized Momentum cites this paper.

Muse: Representation Geometry of Muon Beyond Normalized Momentum A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T01:58:57.470749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:58:57.470749Z digest=sha256:7fa368c2acbdec2e7a1dcceabe153a80dc1cfae2305e1a20599656af4ff91573

Observation 891d468f-8722-410b-92a1-9391048fd58d · inbound

PoLoRA: A Preconditioned Orthogonalized LoRA Optimizer cites this paper.

PoLoRA: A Preconditioned Orthogonalized LoRA Optimizer A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:35.758665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:35.758665Z digest=sha256:60eddf2bfbb36a5cd262422bdc774f5060380bb4d7344e0b03f99e3121d4d379

Observation fc89e38e-cb7b-4ec1-99c0-eda052520eaf · inbound

OneShot: Index-in-Ranking with Neural Scoring for Large-Scale Retrieval cites this paper.

OneShot: Index-in-Ranking with Neural Scoring for Large-Scale Retrieval A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T07:21:46.180322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:21:46.180322Z digest=sha256:3b5718b503b348347b069e7c3cb852cfaae7aeb3792f6fd8316a0b5286a9c826

Observation 18859707-d211-499f-8c3e-d8fcb74962e9 · inbound

OneShot: Index-in-Ranking with Neural Scoring for Large-Scale Retrieval cites this paper.

OneShot: Index-in-Ranking with Neural Scoring for Large-Scale Retrieval A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T01:45:23.261226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:45:23.261226Z digest=sha256:9880cc8a2a4cc1ec99b7e7a355ce452fc9ef73a12ddc4d50a2f1320c7af5b082

Observation 6ec7bf56-7c37-4897-b8e9-256fc8d7d013 · inbound

Muon Meets Mamba: Spectral Optimization for State Space Models cites this paper.

Muon Meets Mamba: Spectral Optimization for State Space Models A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T05:27:17.313102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:27:17.313102Z digest=sha256:87472ef178f5890270d0077a60c04a8862e9bec186f95317f1721ac67c85e35c

Observation 4a0b3809-ca41-41cc-aa65-bcf638e00f14 · inbound

Second-Order Muon Done Right: A Principled Marriage of Spectral Geometry and Curvature cites this paper.

Second-Order Muon Done Right: A Principled Marriage of Spectral Geometry and Curvature A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:16:02.781870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:16:02.781870Z digest=sha256:f486ca4e1792a56e6f75e2327c8e8eb2c1ddf7e12980a6895537af7b5ba3f502