Pith. sign in

Paper Citation Record · LEDGER

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

As of 14 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 8 inbound Pith citation observations for arXiv:2412.13148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13148 v3

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:28:29.503764Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:37:23.643471Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T14:29:54.502956Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fad8d5b4-5ab3-480d-b0b1-ae043dcf7a3b · outbound

This paper cites write newline.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.403451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.403451Z digest=sha256:bbf78340ec3cb01d3d2ea0bbe8e9c159156cac9f2095df0df0263c516ff340c1

Observation 375f9e7d-20f4-4454-a14d-bc26b8c76ce9 · outbound

This paper cites Scalable Second Order Optimization for Deep Learning.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Scalable Second Order Optimization for Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.468777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.468777Z digest=sha256:e75478321c3e45f11012813a02e1ae8233443c48d7ada791b50d489dd48c50b9

Observation 1b6b617b-414f-4f52-8947-255aaf655679 · outbound

This paper cites Layer Normalization.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Layer Normalization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.542452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.542452Z digest=sha256:ab38b24f203f1ce8cf0ff610f9c36b6c58bb53635e1b223ae4ce563386f9d59e

Observation 130dbf43-533e-41eb-bfe6-67052f66da7b · outbound

This paper cites Qwen Technical Report.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.548347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.548347Z digest=sha256:2d4f0cce6389e6109f6272d16ade197eee9031582ddcf297be8d56fec95d99b6

Observation 8a29298a-891d-4d77-baca-217d986a2364 · outbound

This paper cites Modular Duality in Deep Learning.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Modular Duality in Deep Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.554180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.554180Z digest=sha256:e8b8195972e472b84eb8c41d50a9f626b19cdedf5fc66cf15e48870fee57f927

Observation 7d453c04-24ab-4e5c-b210-1cd9640bcbff · outbound

This paper cites Old Optimizer, New Norm: An Anthology.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Old Optimizer, New Norm: An Anthology

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.566252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.566252Z digest=sha256:8e31aa8151b4d7f3b64c54ac3a67c0c537216f4488b0cf68148090f1b85225d5

Observation 037d451a-1734-4c56-8dcf-c6aecf768210 · outbound

This paper cites signsgd: Compressed optimisation for non-convex problems.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training signsgd: Compressed optimisation for non-convex problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.571659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.571659Z digest=sha256:2a33c3069a45f34fd380042e1115e039de61cfdedd9c5b8eccf88f6c0f1400eb

Observation 880f976d-a632-4f37-a7c7-e7a7111df26f · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.576403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.576403Z digest=sha256:d7f53932b6449b79f8c195ebf8ecb956d947cd9638bdd2b0bf7efe7313aac3ea

Observation 483f8f95-5dd9-4282-80ec-d0140a041dc4 · outbound

This paper cites an unresolved cited work.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.582160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.582160Z digest=sha256:1b68f567fde528ca3eba1f2623957fe8d7bc3af2a218f0f8f7dcfdd66462ad48

Observation ad29e9b1-c6f9-46ff-bb4e-a0757af56314 · outbound

This paper cites Preconditioned spectral descent for deep learning.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Preconditioned spectral descent for deep learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.981713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.586420Z digest=sha256:0803d852bdd19433da07836380edd0646a49f7784cb2e47279b57b84f3ddec33

Observation 43bac816-83fd-4506-9aa9-ab833454669b · outbound

This paper cites an unresolved cited work.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:28:30.965966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.590821Z digest=sha256:157f70f3315f9b7da4a0dcb76c21bf8bf55b21722681b79eac752e412f5dc0f2

Observation 0b3d67e7-eba6-4e5a-8237-6fdfd2a0b217 · outbound

This paper cites Robustness to unbounded smoothness of generalized signsgd.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Robustness to unbounded smoothness of generalized signsgd

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.595571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.595571Z digest=sha256:2f016ea114795309ded22460d6fadee9c419aaf9410af0b63782c01cf1acc983

Observation 51bb9488-1910-4ced-a052-97dc5f16e618 · outbound

This paper cites Momentum improves normalized sgd.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Momentum improves normalized sgd

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.939348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.600029Z digest=sha256:f92fcc2c1f3d0c3449f116fdb3f27e433bca7df0177b190e52bb02b4bcce27c7

Observation 2d7a6c7f-9f4c-47e9-93b7-5bae4abb1bbc · outbound

This paper cites A general system of differential equations to model first-order adaptive algorithms.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training A general system of differential equations to model first-order adaptive algorithms

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.921513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.604662Z digest=sha256:cbea09bbbdf554df99270618ab5fabea11683f7f734d21cd95dfe52a832592ce

Observation 297e43fb-e6f9-4af6-a931-464318b052ae · outbound

This paper cites The Llama 3 Herd of Models.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.609620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.609620Z digest=sha256:800a76077261944a2490c4239f3ade419d0f8bf02301ed4d9962d4b802cc69a5

Observation 9f62afd6-8290-4170-b690-05f277f6044b · outbound

This paper cites Duchi, Elad Hazan, and Yoram Singer.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Duchi, Elad Hazan, and Yoram Singer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.904742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.614522Z digest=sha256:c6023534f69233da7c7a9609e352f9635e5e7eee0fd0e3ec608a04378904fa55

Observation 3fc39d8d-795a-4fb1-b5e7-0e16756eee74 · outbound

This paper cites Kronecker-factored approximate curvature for modern neural network architectures.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Kronecker-factored approximate curvature for modern neural network architectures

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.888396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.619816Z digest=sha256:2c87758375de50a01aa7f2870ce8637434f3901c6d95d0b3858a1262b08af416

Observation 6616db8b-1b10-40b1-bb1d-985e6fc574c7 · outbound

This paper cites A trace-restricted kronecker-factored approximation to natural gradient.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training A trace-restricted kronecker-factored approximation to natural gradient

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.830479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.624173Z digest=sha256:0df3b6f3e9808bbc6d469b595de3ab449334d01eaae1fd1bc248122e021d8657

Observation 1d1dc410-000d-4a2a-a607-84d73850b15d · outbound

This paper cites Eigenvalue-corrected natural gradient based on a new approximation.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Eigenvalue-corrected natural gradient based on a new approximation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.800073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.628843Z digest=sha256:f2e7039c78f96285f20f17569da2fd9135c169465374bbfe6843db25513ed40d

Observation 62e675b6-ce37-4d73-ad63-51ec6435c85b · outbound

This paper cites Fast approximate natural gradient descent in a kronecker factored eigenbasis.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Fast approximate natural gradient descent in a kronecker factored eigenbasis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.633800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.633800Z digest=sha256:790f3378c377732ae4a15f524f0f2db78f3671bb497a82c7aedfefa0db89ba5d

Observation 7d44018d-acef-4d0c-9fe3-adc9af8dc4df · outbound

This paper cites Shampoo: Preconditioned Stochastic Tensor Optimization.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Shampoo: Preconditioned Stochastic Tensor Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.638293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.638293Z digest=sha256:4d8e7b0be9313731f7d2db3a1e464a75480e37c563efce63a72f73d60dc13502

Observation 7660468d-b3b9-4def-a6d9-b55e2c2d2c3e · outbound

This paper cites Flora: Low-Rank Adapters Are Secretly Gradient Compressors.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Flora: Low-Rank Adapters Are Secretly Gradient Compressors

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.643348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.643348Z digest=sha256:7f4c69a5ece25c206290d48fdc00c2fb485c7223d80238791b2c3efdc5785c96

Observation 2135bc15-fa3c-4380-beb5-2e1983674a79 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training LoRA: Low-Rank Adaptation of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.648169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.648169Z digest=sha256:ebcb7b2584dc4660dc892d34ad81adec6f777873a344d3d57407b23b753cec57

Observation b346cecc-1dcb-4502-b4df-4f3f53f9ef8a · outbound

This paper cites Decorrelated batch normalization.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Decorrelated batch normalization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.773319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.681920Z digest=sha256:5e1499209122f635e2a036d0f487933f493bc9f76a5ef46c265be1209d123540

Observation 687ce154-7c3c-4382-833c-f198c4017be5 · outbound

This paper cites Iterative normalization: Beyond standardization towards efficient whitening.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Iterative normalization: Beyond standardization towards efficient whitening

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.756753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.780886Z digest=sha256:d3a5ec3a56ff500d8548d5d205175e923ef9bdd5465c201fa5a03eb0e129b856

Observation 9ddb77ce-a62d-475a-a7f8-239610115fa8 · outbound

This paper cites FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.812838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.812838Z digest=sha256:3498e5693593e40e7d4c1e0fa9eb0430e5215e7b0bb314454db5ac222eafe342

Observation 3e4cc3cf-6e69-4e06-9f35-4e65fa3a0560 · outbound

This paper cites An Isometric Stochastic Optimizer.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training An Isometric Stochastic Optimizer

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:28:29.997271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.841238Z digest=sha256:8826f798de0acb7c3039461e218628a8c5026176b99ddcc20272968e2000edb4

Observation 4b192fc3-cabc-4b70-913a-0c737c7d0a76 · outbound

This paper cites Three Factors Influencing Minima in SGD.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Three Factors Influencing Minima in SGD

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.874045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.874045Z digest=sha256:23d7b91efffd49b9fa7e0d08a0bb432a971d4063095fa967f65a06b88db2ad21

Observation 7b2dc94f-b4ed-4b0a-b1bd-e2fadb99c281 · outbound

This paper cites How does adaptive optimization impact local neural network geometry? Advances in Neural Information Processing Systems, 36, 2024.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training How does adaptive optimization impact local neural network geometry? Advances in Neural Information Processing Systems, 36, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.902305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.902305Z digest=sha256:44ca225399eab0f4af4083085d2d7ba6ce3139b924d2be2628fb3c6ac9fd3d0f

Observation 7cb5ec09-42f5-4035-9007-cfea6e913dcc · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks, 2024.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Muon: An optimizer for hidden layers in neural networks, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.728322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.907199Z digest=sha256:076ef91a6764dad983e17670638bc757ed7296be183a470dbf2aed3f0431ecea

Observation 48896bd7-6170-4848-b749-1ff378b5e378 · outbound

This paper cites No train no gain: Revisiting efficient training algorithms for transformer-based language models.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training No train no gain: Revisiting efficient training algorithms for transformer-based language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.712135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.911867Z digest=sha256:41dffb13972e5c46bb2dcdfa781e0834de907482c028d078a6ce0ad1dab4ab14

Observation f55ac609-a53b-47bb-b9c0-f03fa3f1cccb · outbound

This paper cites Exploring Low Rank Training of Deep Neural Networks.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Exploring Low Rank Training of Deep Neural Networks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.916986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.916986Z digest=sha256:80eddfdaf565f12f822befbce391344ac4e9b6b790a22f98131e2d7300df354a

Observation d6525654-fd46-4476-995f-46c04712cb89 · outbound

This paper cites Kingma and Jimmy Ba.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Kingma and Jimmy Ba

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.695028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.922325Z digest=sha256:9fbd09dcb64f5aded19fd3a583544db0d42cac52526faca20c89c5fbe5ac0838

Observation ed262b40-e69e-451b-a28b-995efa0b4727 · outbound

This paper cites Efficient Approximations of the Fisher Matrix in Neural Networks using Kronecker Product Singular Value Decomposition.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Efficient Approximations of the Fisher Matrix in Neural Networks using Kronecker Product Singular Value Decomposition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.926599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.926599Z digest=sha256:2502f96c240678fee2288fd9d423f5d60e6938ac54e89a7729a75682e6d1f78f

Observation c7cc8bfb-00d7-4254-b6a6-ab19965059c3 · outbound

This paper cites Reducing activation recomputation in large transformer models.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Reducing activation recomputation in large transformer models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.931607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.931607Z digest=sha256:7a43e8aefe1c1011e64eeabed22bd75eea54bbd2b15e480c0946ad4b26571c96

Observation a8342c03-3d74-48e8-83a4-86663878fd95 · outbound

This paper cites Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.936208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.936208Z digest=sha256:ac4e57a36b4b64cc1c9dd661b911a87d460fbe927d2397eddd07f8e98eab7c56

Observation 27594f91-5cf2-4a8b-b430-4fbd50d73a4e · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.941168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.941168Z digest=sha256:1ee638aabff8a6c6514e5b4d6bc3dbb0d32be8e329d9b4cfe5ff3828d16ce723

Observation 8534924f-03bf-4527-b9ad-9114558c45b8 · outbound

This paper cites Towards faster training of global covariance pooling networks by iterative matrix square root normalization.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Towards faster training of global covariance pooling networks by iterative matrix square root normalization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.668961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.945674Z digest=sha256:319f0fcc8f865ac0bc8f051a51abbff7fb69335a9350738f004a0ffc4256a616

Observation 5a127efb-8843-4917-87c3-3eab6bac0872 · outbound

This paper cites Preconditioned stochastic gradient descent.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Preconditioned stochastic gradient descent

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.652212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.950402Z digest=sha256:dfbd4976960b0600caa0869faa5d39584f294c44262fee7aa49aa2e669d5d236

Observation dbf2de18-9eaf-434f-aa26-493a00180d7f · outbound

This paper cites Relora: High-rank training through low-rank updates.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Relora: High-rank training through low-rank updates

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.558099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.954761Z digest=sha256:697d45c8dfbf6aabcc9fb0905eb9bc10103252dd71778c8fb831bdde559d529b

Observation b1c10969-318b-4b0e-9838-ba034372db6e · outbound

This paper cites Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.959819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.959819Z digest=sha256:fe16c3c7ac3274d447a12ef65ddad7207e646e4a6cc0e6c055024ee1ff7bf811

Observation 816c38e3-096c-4bc3-8bd0-0af2dfd509f5 · outbound

This paper cites Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.964729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.964729Z digest=sha256:4e0acb2abfb57d7fd50aa4aa2a0b9128d534aa2bfd687b4bc168885da5e10ad8

Observation 74acc1a3-0091-456c-b28e-b78b8d6d97d8 · outbound

This paper cites Decoupled weight decay regularization.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Decoupled weight decay regularization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.969926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.969926Z digest=sha256:5e99f7b2dbd565b73eb5f4c958105b66c0170381819739bcb764baa2361fbce4

Observation 90870da8-9c28-444b-a0d4-b680b9f28c53 · outbound

This paper cites Optimizing neural networks with kronecker-factored approximate curvature.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Optimizing neural networks with kronecker-factored approximate curvature

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.974524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.974524Z digest=sha256:705f6b7c673076763eff5d6766451db15c6d5c8b8d4788c1d318d2cf98ea9d69

Observation 9ea789dc-f0a1-47b1-804e-691e8b483c9b · outbound

This paper cites Kronecker-factored curvature approximations for recurrent neural networks.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Kronecker-factored curvature approximations for recurrent neural networks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.424554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.978832Z digest=sha256:4f19a98dda25f638f6295117d4f769f744c81d429722c72de2c20895ca52903c

Observation 31a91584-c553-4fa0-9209-2bf3f648c475 · outbound

This paper cites Kradagrad: kronecker approximation-domination gradient preconditioned stochastic optimization.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Kradagrad: kronecker approximation-domination gradient preconditioned stochastic optimization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.340119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:28.983562Z digest=sha256:3a73054760231c25174f0c9437e26fa83dcae3e6d6db72fc384c1f06c7f463ad

Observation d6f21902-5e7a-4295-96b8-7fee0b54d973 · outbound

This paper cites A Theory on Adam Instability in Large-Scale Machine Learning.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training A Theory on Adam Instability in Large-Scale Machine Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.988199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.988199Z digest=sha256:e99a6cb5618149bb24ffe7a958ceed93bb2c0a73837c089158152891a7840c00

Observation 373bb1a4-8514-41a3-bc60-c5f167e0e9b8 · outbound

This paper cites Introductory lectures on convex optimization: A basic course, volume 87.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Introductory lectures on convex optimization: A basic course, volume 87

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:28.993345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:28.993345Z digest=sha256:2ace0538ea6ffa0ec6ec09ff0155ee8ee6c6ae2847dcdc3c3c396849b8cea35a

Observation ce538f7e-81c3-4510-a8d7-9a0d02364d7d · outbound

This paper cites The AdEMAMix Optimizer: Better, Faster, Older.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training The AdEMAMix Optimizer: Better, Faster, Older

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.021881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.021881Z digest=sha256:0ff541da84f48bea532148daf3b848d06624ad93837108ca80d5d7fb31026222

Observation 6f93ca78-6eaf-4d4c-b50e-b54a6db651a8 · outbound

This paper cites Fishy: Layerwise fisher approximation for higher-order neural network optimization.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Fishy: Layerwise fisher approximation for higher-order neural network optimization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.313668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:29.053316Z digest=sha256:0309bed96048ea96f585eb7b82afeabd76fe0eec4e637308b8d34c6e3684f942

Observation 5721afeb-b1a6-4472-bfef-5dd3671c81f8 · outbound

This paper cites Curvature-Informed SGD via General Purpose Lie-Group Preconditioners.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Curvature-Informed SGD via General Purpose Lie-Group Preconditioners

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.100571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.100571Z digest=sha256:3c6e6f17cc0150c13225b0bfb4417418bc6393984948b39e5aad5363d57922d7

Observation 10beb0ba-f15c-48e7-a456-43d095ef6c0e · outbound

This paper cites an unresolved cited work.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.131888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.131888Z digest=sha256:8267c7acf604c24837aa526e15aa8ecb55b53c240f93d63c96cebca39885d24a

Observation 69378cf6-8b94-4ca7-9808-c9d3df570ddb · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Adafactor: Adaptive learning rates with sublinear memory cost

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.288834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:29.137351Z digest=sha256:6dd130dd6722ba44460a35d2a034ed7ec692911355f38112dcf159a103c2c75e

Observation 50362495-f249-4a7d-bd41-0e1ed9c1767d · outbound

This paper cites A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.142246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.142246Z digest=sha256:71026694b3a55c496431a14851917333eb1f773f3f566a01ec7e3a8c69629365

Observation e7722d98-497e-4693-8d11-9e5252a4494a · outbound

This paper cites Fast differentiable matrix square root and inverse square root.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Fast differentiable matrix square root and inverse square root

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.273781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:29.147087Z digest=sha256:212685fd06b9d59e287260e11770810216daa238fb2ef6d938380a664b674fa2

Observation e99ac229-0fe9-4cdf-a4fd-197fe7b6e2be · outbound

This paper cites JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.151686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.151686Z digest=sha256:43489c3c3ac514a3388cdc4d0f44394fa653ab884018f22aa4db943ce8bd2a88

Observation b2f6baea-eee4-487b-9d2a-3207e91d89f3 · outbound

This paper cites Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.257689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:29.157060Z digest=sha256:89fe6924e2ce5765ab1eab102f57d4127c320de3a9dd8334d278b0e1a4ef7b07

Observation 5f96d0d0-39bf-4a1f-8fcb-22feb894ac6c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training LLaMA: Open and Efficient Foundation Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.161798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.161798Z digest=sha256:19bb0ee3ae7f867a9ddaa460af9c0e9f9792f3e36d7301a542aa9f16069dcd90

Observation d981dbb4-8874-4227-884a-969946d17287 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.166823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.166823Z digest=sha256:476c0fd2b07124e37a177bc6ed6f3ae19077368008523b7c708d187530a05728

Observation 33285297-0730-4f02-98af-685372b2a3d2 · outbound

This paper cites Orthogonalising gradients to speed up neural network optimisation.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Orthogonalising gradients to speed up neural network optimisation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.172170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.172170Z digest=sha256:b997c6431a699f5bea1963f70c3bf9316e0d933cd0e4ff2332dd903475f90686

Observation 360f6a6e-38c1-4698-8e1b-8a7ce24975dd · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training SOAP: Improving and Stabilizing Shampoo using Adam

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.177761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.177761Z digest=sha256:d530c0819ee8005d9f75cda68691fbc5ef2633115d6ff69a12f7391a07bfa0a0

Observation d810eae4-64ff-4333-8d25-24231d45a8c2 · outbound

This paper cites 4-bit Shampoo for Memory-Efficient Network Training.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training 4-bit Shampoo for Memory-Efficient Network Training

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.182673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.182673Z digest=sha256:dd8c54c7233d4b2d136ff06eb9438a54200b13bb00398571cf72db5d41c61a51

Observation 42800392-4b12-4e39-aad4-2377e9ba907e · outbound

This paper cites No More Adam: Learning Rate Scaling at Initialization is All You Need.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.191879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.191879Z digest=sha256:364a51c549985b1b8ea427f3b99fdd380527409c1bc1d862c23cffd0225a53d6

Observation bd1cc1d3-5c7d-4581-a3b4-e827aa1919e5 · outbound

This paper cites Principal whitened gradient for information geometry.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Principal whitened gradient for information geometry

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:28:30.240879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T13:28:29.196373Z digest=sha256:b9a924cfae6c6a1299b3d7ab358e1f51917ec0c44862c4cc46a343ca7a57a428

Observation c3ce154d-4c22-4367-9712-890425faea5a · outbound

This paper cites Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.200920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.200920Z digest=sha256:323f16e9c8839a2f7a887bfe43fc988b55ea38546ccd3e20c1ac32adf5c8f9d3

Observation deaf2b4f-0d60-45c4-be8e-0e094577709d · outbound

This paper cites Root mean square layer normalization.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Root mean square layer normalization

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.206050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.206050Z digest=sha256:65153895990dc3d2b4e4757cf832a99581e3f258e2b6b78fb7012e49d97c3622

Observation 67a1919a-af36-4aa3-a511-2361c4d54eee · outbound

This paper cites Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.210374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.210374Z digest=sha256:32daa50f0d87946a1fde6c51582450b3b0cafdfde7d3f7b51f3123f535ca481e

Observation fe381d5a-8155-4bd1-8787-fb38cd144f1a · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training OPT: Open Pre-trained Transformer Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.232812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.232812Z digest=sha256:48d12ca6a476acd0093438a5835988c89975c3035f5ba96b2e89858be4efa0a2

Observation 10b9c2ff-b849-46d1-b73a-0bb669575807 · outbound

This paper cites Why Transformers Need Adam: A Hessian Perspective.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Why Transformers Need Adam: A Hessian Perspective

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.314973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.314973Z digest=sha256:cfe8a3c27d68909cc5d0cb694ddf7baef21e56ad016eb57b006a441885a8ed78

Observation c8ef6745-d8a0-4b4d-a5a6-c269d390b9ee · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Adam-mini: Use Fewer Learning Rates To Gain More

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.340070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.340070Z digest=sha256:d199d07df7479c9771d1c2e8c895c35d1f18a03881385c627e3566206e91320c

Observation 3de82ffc-8709-4fda-99e0-56ad1c91c1b2 · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.400702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.400702Z digest=sha256:40ed9d726f30885d50aa88fcc35da59e6508b55f7758d9bdf4f19b928bded6f2

Observation e65589d3-b5f4-4b86-9757-59787ff81fe7 · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Deconstructing What Makes a Good Optimizer for Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.479284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.479284Z digest=sha256:e762fcb8141493c46375e198af664eb5177532603a586d6b3b79c12f67b46cf7

Observation 4a38ff1a-5543-4da6-9d7f-a057b390944b · outbound

This paper cites APOLLO: SGD-like Memory, AdamW-level Performance.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training APOLLO: SGD-like Memory, AdamW-level Performance

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.483763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.483763Z digest=sha256:aa3b97056fb971483f3794427098fc6e2b394f644e3a67d518074476ca1feec1

Observation d48760b1-3cb5-477d-a3c9-2a160f539ca6 · outbound

This paper cites The Anisotropic Noise in Stochastic Gradient Descent: Its Behavior of Escaping from Sharp Minima and Regularization Effects.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training The Anisotropic Noise in Stochastic Gradient Descent: Its Behavior of Escaping from Sharp Minima and Regularization Effects

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.488725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.488725Z digest=sha256:ab924baa87178c9259405b76074c5287f75292772e41c7b35186e3b9228304b2

Observation 0ba84744-5121-4f25-9c76-0ff64573fdfb · outbound

This paper cites @esa (Ref.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training @esa (Ref

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.493867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.493867Z digest=sha256:96f504efa82dbb9291bbb934e9ed7e9a9e6cc2d145dc1432370cf9ab8cc3f5a5

Observation 2c0553f1-7cad-4189-8e31-66f0d1f92661 · outbound

This paper cites an unresolved cited work.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.498996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.498996Z digest=sha256:9b1fca737295f27d4330a73ae3acd2eab8e996ec4d7a1ff092e7cc2071c6a123

Observation 40712ffe-5059-4790-bab5-9e0e781f6725 · outbound

This paper cites an unresolved cited work.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.503764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.503764Z digest=sha256:53492cd4a8907764c8bdc0aa468a666ca4435ba1fae7f770c33a41d8859d878f

Pith citing papers

Observation 5897e39f-a8d8-44e2-863d-cbf31763c4dc · inbound

Gradient Multi-Normalization for Stateless and Scalable LLM Training cites this paper.

Gradient Multi-Normalization for Stateless and Scalable LLM Training SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.643471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.643471Z digest=sha256:39a162c2bf7c303476c0275573ce3603b967f2043fd26b08f4282d0064f5a252

Observation 860c6b52-a841-4eaf-9efe-f4947bd814be · inbound

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension cites this paper.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.751693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.751693Z digest=sha256:5fd6022e41b3235181dcd61da22fa84b5a4eba5beb4a1ce3844835eeb7f8310f

Observation ff737671-9801-4638-be91-b3d68b07fc22 · inbound

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design cites this paper.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:24:53.210558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:d0ca18099158225b8662893ae9e7d6d6e4a5fe8a52ae5f38e06c85d82f13f40b

Observation b33f2a32-f26b-4cab-bc92-e0b243e5deb4 · inbound

Low-rank Momentum Factorization for Memory Efficient Training cites this paper.

Low-rank Momentum Factorization for Memory Efficient Training SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.196266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.196266Z digest=sha256:a84756c638018ef21c9703eaf8c850d6c3fee971dcff7919490b5a97e8d0709d

Observation 7736c809-eda6-4a2d-846a-f5e47244863e · inbound

Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning cites this paper.

Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T10:41:51.251736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:41:51.251736Z digest=sha256:e564d4e5f1c804fbe3f80698e4cf4491b79bb0780bb1978da7e638c3c63ba136

Observation a7c1ad9e-fe70-49b7-82be-fedba31a45a1 · inbound

Demystifying Manifold Constraints in LLM Pre-training cites this paper.

Demystifying Manifold Constraints in LLM Pre-training SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:16:08.611550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T17:44:44.438637Z digest=sha256:9b6449125b5ee6573f7fb60268409daeea4ac99c488bfb14468a71ca175ee802

Observation a5ef100f-8376-4513-948d-4dc002dbe515 · inbound

Hierarchical Muon: Tiled Newton-Schulz Updates for Efficient Muon Optimization cites this paper.

Hierarchical Muon: Tiled Newton-Schulz Updates for Efficient Muon Optimization SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:29:54.505468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T03:33:34.662754Z digest=sha256:c6b5387090ae11c05bef8d4b59c4e12d6ab30a0b3432599d332f20378cdaaf0f

Observation 76fda286-b3b7-4c75-8c6d-04316cf8a1cb · inbound

Muse: Representation Geometry of Muon Beyond Normalized Momentum cites this paper.

Muse: Representation Geometry of Muon Beyond Normalized Momentum SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-02T01:58:56.668607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:58:56.668607Z digest=sha256:a6465137f206f681b19d1f9193bd49d2cb2a220f1a665bba847c603ed29ba7fb