Pith. sign in

Paper Citation Record · LEDGER

Gradient Multi-Normalization for Stateless and Scalable LLM Training

As of 22 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2502.06742.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06742 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:37:23.819526Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:35:40.221296Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T05:52:21.848754Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b08ad4a5-16fe-4a52-8ad6-c06e6506952c · outbound

This paper cites Layer Normalization.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.413471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.413471Z digest=sha256:0ecf649f4e43ac96fc6dcdb4aa4e4f4b2c2da82fef8aa5b18d36d08bf8ff837c

Observation f122e6d1-83f1-4c54-b2a7-b5a8d3177886 · outbound

This paper cites Iterative bregman projections for regularized transportation problems.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Iterative bregman projections for regularized transportation problems

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.906346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.451252Z digest=sha256:688d46eaeb42480c2f0aedcf5e1c815b7dfa13fe0e93a03e21ec8bbbed78c082

Observation 4b0d4c41-37f0-483c-943c-122dde9a1987 · outbound

This paper cites Old Optimizer, New Norm: An Anthology.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Old Optimizer, New Norm: An Anthology

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.455661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.455661Z digest=sha256:fb0a3504c88ae0bf169f8d7d5996d4b9f13745fe975cb667d82ef42e9557b75b

Observation 4a520790-e2ea-4bb8-8431-e9247760aaaa · outbound

This paper cites signsgd: Compressed optimisation for non-convex problems.

Gradient Multi-Normalization for Stateless and Scalable LLM Training signsgd: Compressed optimisation for non-convex problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.459883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.459883Z digest=sha256:606746dd7c7da0f47b7826f330b39a1f605afe2a608849a3cfb3046594c5a372

Observation 2c1c0c26-fcb7-4c9e-abc0-a04d482caf08 · outbound

This paper cites Proximal alternating linearized minimization for nonconvex and nonsmooth problems.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Proximal alternating linearized minimization for nonconvex and nonsmooth problems

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.888347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.463975Z digest=sha256:36795d1b2fd97729a832b64619af2d503cf0522696dd7275430578bcbebaecca

Observation 8c2f0518-867a-4c4a-9849-f4602c19c6fa · outbound

This paper cites Distributed optimization and statistical learning via the alternating direction method of multipliers.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Distributed optimization and statistical learning via the alternating direction method of multipliers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.876277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.468048Z digest=sha256:9f754b22273a5dd7d07d3caf782b592acd75d091b8cc310417d34bd45f812699

Observation 6cf160a8-e1af-4778-b2c9-83d0377eb545 · outbound

This paper cites Stochastic spectral descent for restricted boltzmann machines.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Stochastic spectral descent for restricted boltzmann machines

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.865195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.472410Z digest=sha256:81e3c86c1a13c30a78fd7d19e8cf6a94a32921f85248d7d45d97dbfcf566b3ab

Observation bbf7d13b-7496-4394-bab7-87a183480732 · outbound

This paper cites and Pock, T.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Pock, T

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.854101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.475776Z digest=sha256:ed0e7359f77d2c206d3c0711d70b410c86e02e0be3420359ae9e90a4af4181ba

Observation 4299bd35-014a-4e46-b469-d40de0536ef7 · outbound

This paper cites Fira: Can we achieve full-rank training of llms under low-rank constraint?, 2024 b.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Fira: Can we achieve full-rank training of llms under low-rank constraint?, 2024 b

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.482602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.482602Z digest=sha256:e63556bfd148c947020ff63058e566dc108bd8fe5181196026e358d0d49a9734

Observation daea1544-84d4-4f64-b7fd-68adce40a54d · outbound

This paper cites Symbolic discovery of optimization algorithms.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Symbolic discovery of optimization algorithms

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.843957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.486131Z digest=sha256:dd379ddeb91e61fad3662b8cf4e38b1602bb821a5f3ba02d1eeb490b76fb5266

Observation f3125a71-49a0-4a47-af5e-3d31edbf12e4 · outbound

This paper cites and Mehta, H.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Mehta, H

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.832640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.489575Z digest=sha256:f8674e7fc3e15f3520ee7cd8d7e254cea7d0eede47c8fb5b22e4255ecfb2b29e

Observation bed42d3e-476b-4ee1-bf47-18a02c229b59 · outbound

This paper cites On hilbert’s metric for simplices.

Gradient Multi-Normalization for Stateless and Scalable LLM Training On hilbert’s metric for simplices

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.821392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.492994Z digest=sha256:9593e00387075ae895a6749fffbea39d2fcb449432fb3aa4be9830b5b817a7bf

Observation 7bf6bb02-4816-494e-a22a-7d422156ceb6 · outbound

This paper cites The Llama 3 Herd of Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.496672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.496672Z digest=sha256:162a2a96b14c25119b4653118be2fbbde0fafc9b134f038d3c3539afac26b7e3

Observation 5a1036d1-bd39-485a-bbb4-42a88518df53 · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.810400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.500426Z digest=sha256:93876bce424582b996d5cb100e34544783ef0c9905229a5d5fbe43b891c86b29

Observation 807a43a5-376f-44f6-a5ff-ff3e66350ff1 · outbound

This paper cites and Lorenz, J.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Lorenz, J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.799506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.504111Z digest=sha256:9d621973f4e284bb8d92c4b5330528eab5414069c713ba9d39a90416c5e6ef97

Observation 0a3d308f-309a-4bda-8ce9-3dee363a3b86 · outbound

This paper cites Eigenvalue-corrected natural gradient based on a new approximation.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Eigenvalue-corrected natural gradient based on a new approximation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.787737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.507644Z digest=sha256:ada71e05a306401cb601f8ca903e376be871b4c392969b0f6a3e6674b4e3c24c

Observation f3a89ebb-4335-405f-ac92-2ae195999eb1 · outbound

This paper cites Shampoo: Preconditioned Stochastic Tensor Optimization.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Shampoo: Preconditioned Stochastic Tensor Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.510817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.510817Z digest=sha256:f82ad15049fc846cee592ddbe38d7544f203c349707ae58013d385c0edcd7a60

Observation bbb83fe5-3150-4ccf-81ac-2e61517b0bdc · outbound

This paper cites Flora: Low-Rank Adapters Are Secretly Gradient Compressors.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Flora: Low-Rank Adapters Are Secretly Gradient Compressors

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.514931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.514931Z digest=sha256:1ae34e4cb639ca523847042deb64290e847627bd30e9c2f3bf50ce83135818b2

Observation 6d529124-8416-4357-8a69-f4d6d4058af4 · outbound

This paper cites Beyond convexity: Stochastic quasi-convex optimization.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Beyond convexity: Stochastic quasi-convex optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.518833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.518833Z digest=sha256:646846648c1011e0ac40da9dedf506df82c1ef51bdd500e45d3f17f42ba83843

Observation 820b1f11-b874-4e60-a1fb-a7959e808908 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training LoRA: Low-Rank Adaptation of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.522219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.522219Z digest=sha256:5bbc544d46f62b34cf543761bde9f7f0ef003b56f9e9d68e9e40fec8c37eb012

Observation f6927afa-6bcc-4b13-ba94-53b1afe52fe4 · outbound

This paper cites Iterative normalization: Beyond standardization towards efficient whitening.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Iterative normalization: Beyond standardization towards efficient whitening

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.735222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.525891Z digest=sha256:1552b483753234c91b0547e6a9e86600739a8c4ced2997a12a102ad534bd6744

Observation b3245607-4341-4ef3-8480-8707711b816a · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks, 2024.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Muon: An optimizer for hidden layers in neural networks, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.681640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.529508Z digest=sha256:0772569ba91c83424d26b89912a0e6107ff816fbc09b0d4e4890b9f07592e69b

Observation add29701-088d-434a-b17a-76bad058594f · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.623494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.532971Z digest=sha256:8599140254daf8e2783060858978675811d395324abe1670dbcdc6cfb2ad4625

Observation 6080fe88-34be-44b1-8865-08a9657870ae · outbound

This paper cites A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B.

Gradient Multi-Normalization for Stateless and Scalable LLM Training A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.536830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.536830Z digest=sha256:0699bac7b1d477afcc804bd1d6ee3d17fbeb65efcf92b6c2b2390e3c843564c9

Observation f8d51ea9-6283-4132-a32b-95ef53b7d28e · outbound

This paper cites Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.540265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.540265Z digest=sha256:107c66477ee5f5a1b550946a845214984aa58940fffb09af12f22f3d3c99b622

Observation 34a40cd0-c212-4fd7-ab57-f29d81426256 · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.544235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.544235Z digest=sha256:5364161f970b450c223677da2ab7e39c91f707bf695daea2d6e5368643d2e194

Observation 56561ed3-f3da-4622-bb5c-cde77f1bfe3b · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.558644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.561000Z digest=sha256:ecc9e910c7d4f99f43815c15354e526f65b4fcea59ea631c3b04ec4b35ed1beb

Observation 02956aa1-6bba-4a64-b1e0-378a505245c6 · outbound

This paper cites Towards faster training of global covariance pooling networks by iterative matrix square root normalization.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Towards faster training of global covariance pooling networks by iterative matrix square root normalization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.540119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.586474Z digest=sha256:ccdb12e69cb396a3a27eb9de613642192bb912d0fdb9412102fbe72eef78a37c

Observation 91f59945-4410-454e-9dd0-6fce281b283b · outbound

This paper cites Relora: High-rank training through low-rank updates.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Relora: High-rank training through low-rank updates

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.529355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.613591Z digest=sha256:ed2d9bb76244dc867f1b3f6bf191e844bd1b0a5e4a2d9e85c192b63aac112eed

Observation 5897e39f-a8d8-44e2-863d-cbf31763c4dc · outbound

This paper cites SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training.

Gradient Multi-Normalization for Stateless and Scalable LLM Training SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.643471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.643471Z digest=sha256:1c2e36e767cf10cf06d8328949045e0269c2ea863952131ea7e6ed72549f242a

Observation 226a59d9-2e4a-4ced-a9f0-f4b43badc942 · outbound

This paper cites Decomposition through formalization in a product space.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Decomposition through formalization in a product space

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.518803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.669813Z digest=sha256:d88a08c7b52ba129270dcc595f5e71ec5ba9e3e1a1625e294520655325bf2823

Observation c6aa30c5-ec4d-443b-9791-a2c525b24b47 · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.508248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.685068Z digest=sha256:60620b9eb9f6e0aed9391c52d815884dafd68e827780888cf5bef07e78fcbf14

Observation 39ffadbf-194d-4c97-8e4e-3b406d8b9716 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Zero: Memory optimizations toward training trillion parameter models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.713139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.713139Z digest=sha256:c564cda3a1395d0f74ca553a75205bd1211b13d59283175d7de6bc57f94caebc

Observation 945ca777-367e-43f9-8627-88b80a777228 · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.492141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.744517Z digest=sha256:f0b6a29bd629c787b8a1f787fa980cafa4cca2128d6943c5db47a109af194ad3

Observation 5dd8a654-35d7-4b34-9180-f3bfd7b25fe4 · outbound

This paper cites A relationship between arbitrary positive matrices and doubly stochastic matrices.

Gradient Multi-Normalization for Stateless and Scalable LLM Training A relationship between arbitrary positive matrices and doubly stochastic matrices

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.757416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.757416Z digest=sha256:9aa9a99ce20be22e460eb0b3f325cb7fd55eb80c11ca9243ee188c8c76160bfa

Observation 8bbad764-6fb7-4ec2-8bdd-239f32ec1b26 · outbound

This paper cites and Knopp, P.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Knopp, P

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.761340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.761340Z digest=sha256:581b68fcd5931c5efbc4d99b63ca416b461d6ef5009ee8348b86a2a9812e1102

Observation b6e29277-8a95-484d-b646-8346a7ffa4bf · outbound

This paper cites Fast differentiable matrix square root and inverse square root.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Fast differentiable matrix square root and inverse square root

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.467626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.765130Z digest=sha256:8af3abea95eace87420150cbf15280ac1dd156976fa2c16b7ec290099cf45d96

Observation 1907b5da-4d76-4ef1-87d2-df117dc1f8e6 · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.456973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.768652Z digest=sha256:3d696150d3d0279c0c5598761b8cdfd0309e7a5e8d0a0bec9e36c4270884f211

Observation 5217120b-dfbe-4c6c-bb3c-8a46fd1c51cd · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.772296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.772296Z digest=sha256:22556a4d0e85b0b4c817c9bb630d5410625b22d9e2ef6b3e56a1b0a16eb2d0d4

Observation 555089af-2e8b-4842-91c8-0301ec813625 · outbound

This paper cites Functional operators: Measures and integrals, volume 1.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Functional operators: Measures and integrals, volume 1

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.446304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.776322Z digest=sha256:9eafedf9b586d3a93af5c39914ec07deb56cba20054f69cccd4b9296009cee1f

Observation 0943772a-8700-42f1-93ce-bd84ff49c3bc · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.434492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.779538Z digest=sha256:dd5a9dafcc32843ed9c7ea73d6bd0ec01835610c7d752eb176d5d985d374f77f

Observation f4028448-fec2-4f51-9294-9d575a3c015a · outbound

This paper cites No More Adam: Learning Rate Scaling at Initialization is All You Need.

Gradient Multi-Normalization for Stateless and Scalable LLM Training No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.786570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.786570Z digest=sha256:1d976433449b4da566574a738c602541381b4eaf7d6f814ac09e4a03161b029f

Observation c3dcb6c0-fd3a-4990-ae1b-f914be422b5b · outbound

This paper cites Large Batch Training of Convolutional Networks.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Large Batch Training of Convolutional Networks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.789854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.789854Z digest=sha256:4600bce8f443dba06c3270cb087c073430c90bea510dcc405b24157c11763a54

Observation e8841d1e-9bc4-4414-bfe0-5f8a69154ff8 · outbound

This paper cites Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.793583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.793583Z digest=sha256:08e0b41e5381fe20a4f17eadfc85743a4c00bf9e3d2d853117556c96772b1d62

Observation 8a9cfafa-3b90-4834-9d6b-fdada3dad04f · outbound

This paper cites and Sennrich, R.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Sennrich, R

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.797421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.797421Z digest=sha256:34f7aa73618fe65d78751a9a8e9975b9e785ba3960f24348e4ff0908194ed483

Observation 47a10613-9aab-45f7-a17a-6476b27fdf96 · outbound

This paper cites P., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S.

Gradient Multi-Normalization for Stateless and Scalable LLM Training P., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.417347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.801164Z digest=sha256:b57917b291260ba0b71707bb6a800998210813160a9135f6f1c0ecc3256d5c7f

Observation 6d77efdc-e81c-4111-b116-cb3594c20dbf · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Adam-mini: Use Fewer Learning Rates To Gain More

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.804426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.804426Z digest=sha256:a2f2a2ecbc6166e5238e5dd82e10baee228a1fe644de12067b1e20f6c31b9419

Observation 14e30325-302c-4d31-b315-0b89b822b86b · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Gradient Multi-Normalization for Stateless and Scalable LLM Training GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.808168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.808168Z digest=sha256:aaf84b7bfe429997959553ac6cb9a50b5e64e4226ad972a2fe75188d19ff3f8c

Observation 552f658d-4508-401a-9dea-25d0d0515d53 · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Deconstructing What Makes a Good Optimizer for Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.812206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.812206Z digest=sha256:a3bf96c7b46195ae8258f39697c2f59aa60379a97403810d1313c3a2c9c9bb91

Observation a6b8083a-4cd7-47bb-ad53-6e19cd7f3df8 · outbound

This paper cites APOLLO: SGD-like Memory, AdamW-level Performance.

Gradient Multi-Normalization for Stateless and Scalable LLM Training APOLLO: SGD-like Memory, AdamW-level Performance

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.815685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.815685Z digest=sha256:8e13fb4484a7fb7814e10b3dc1d076115efb51b96842afbb55670abf43eea2c9

Observation 69ac2474-aae8-426b-a383-655c9cc14123 · outbound

This paper cites write newline.

Gradient Multi-Normalization for Stateless and Scalable LLM Training write newline

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.819526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.819526Z digest=sha256:bd0182fd67a8ed959c82ee2de52e42061cd16c33ba384ff4c807a748cbc90285

Pith citing papers

Observation b23d3601-7f59-4904-8d2e-0039c1fa3197 · inbound

Low-rank Momentum Factorization for Memory Efficient Training cites this paper.

Low-rank Momentum Factorization for Memory Efficient Training Gradient Multi-Normalization for Stateless and Scalable LLM Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.221296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.221296Z digest=sha256:0c07ed50a2788f121e9800db1738e096f0d446c4d610c2c71829000a1c94f435

Observation 903bcba7-1617-42cb-b2b1-21f27dad111d · inbound

Demystifying Manifold Constraints in LLM Pre-training cites this paper.

Demystifying Manifold Constraints in LLM Pre-training Gradient Multi-Normalization for Stateless and Scalable LLM Training

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:16:08.860137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T17:44:44.438637Z digest=sha256:cf53f1ed4c6636716cc1a51cb238332b9a71e3a734d923a5525cdfd79be9c84d

Observation e25f7357-df6e-4089-bfd0-46d14d6b6413 · inbound

Optimistic Dual Averaging Unifies Modern Optimizers cites this paper.

Optimistic Dual Averaging Unifies Modern Optimizers Gradient Multi-Normalization for Stateless and Scalable LLM Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:21.850767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T05:52:16.805180Z digest=sha256:9d09fcae9cd407da3a6aa0135c486c64b1ec13936ad80c4289e6b3c0d09b3d15