Pith. sign in

Paper Citation Record · LEDGER

Insights from Gradient Dynamics: Gradient Autoscaled Normalization

As of 7 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2509.03677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.03677 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:51:39.015136Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved13
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e8a9f65-878d-44ad-9fdb-c2c37f4b8f47 · outbound

This paper cites Layer Normalization.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.901558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.901558Z digest=sha256:06a33b49cfd9e400df4ffdfc1f6a2b4160649e71b6daccab4dfdeb64f69cdc8b

Observation 1407376e-2cff-4088-9ec5-946660a7db70 · outbound

This paper cites Large-scale machine learning with stochastic gradient descent.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Large-scale machine learning with stochastic gradient descent

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.906194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.906194Z digest=sha256:1b2bffefd0e34dbff172d5a7441d6e1ab6cd7cce73a95df3e82954fb44423796

Observation 1cca347b-b28b-4c88-b834-82b04607df9a · outbound

This paper cites The tradeoffs of large scale learning.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization The tradeoffs of large scale learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.384027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.909953Z digest=sha256:98e1c50d713f43431c9a2293d94b430080504c061f558a4d37c4ed2db6e4f9f1

Observation 5eb220e4-b18d-4c94-87c2-831deb57219a · outbound

This paper cites Curtis, and Jorge Nocedal.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Curtis, and Jorge Nocedal

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.913948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.913948Z digest=sha256:f7656854f9a01cba6647d85cb6ea2f815793dfaedfa191d24dc7b819fd0bfa0b

Observation 508674f8-9780-4928-8f92-23f6ddc6c4fd · outbound

This paper cites Entropy-sgd: biasing gradient descent into wide valleys.Journal of Statistical Mechanics: Theory and Experiment, 2019 (12):124018, 2019.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Entropy-sgd: biasing gradient descent into wide valleys.Journal of Statistical Mechanics: Theory and Experiment, 2019 (12):124018, 2019

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.917822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.917822Z digest=sha256:cdce6b4573ee8da87da77c4be63e984cd29687423fb7559a25f5ad25a52efa85

Observation 50dd9eb9-c3a2-4dea-bfc4-8127695d03bd · outbound

This paper cites Gradnorm: Gra- dient normalization for adaptive loss balancing in deep multitask networks.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Gradnorm: Gra- dient normalization for adaptive loss balancing in deep multitask networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.372116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.921632Z digest=sha256:76fed71fa0251e10b4e5354acbb8e34980d64c2aedd20d439d5f1fb1fda8b9dd

Observation daf8b0b3-9c90-4ece-9d20-2d14d4ef2401 · outbound

This paper cites A Study of Gradient Variance in Deep Learning.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization A Study of Gradient Variance in Deep Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.925694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.925694Z digest=sha256:64e2894392a81b66dbc7961e80fd80fcbb6fdf0a675d818a5e55e0c69670c9d5

Observation 39c137a0-1bb3-4d63-a904-e524da2847fc · outbound

This paper cites Understanding the difficulty of training deep feedfor- ward neural networks.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Understanding the difficulty of training deep feedfor- ward neural networks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.358803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.929387Z digest=sha256:b30ba55f14070ca3cdf77923280bbef87b3b110db4e0b633fbfdd41d398f686d

Observation 2b9b9e64-56f9-45ad-82a7-5bf40561b76f · outbound

This paper cites Take a shortcut back: Mitigating the gradient vanishing for training spiking neural networks.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Take a shortcut back: Mitigating the gradient vanishing for training spiking neural networks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.346789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.933179Z digest=sha256:57f4a15e8fe832b687b763da54e79bf0ac70d2dd3cdb9a2573de41d09bbe7aed

Observation 0ded4a05-1d21-438d-876b-2a6eb660e9db · outbound

This paper cites Stable architectures for deep neural networks.Inverse Prob- lems, 34(1):014004, 2018.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Stable architectures for deep neural networks.Inverse Prob- lems, 34(1):014004, 2018

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.936717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.936717Z digest=sha256:3a7b3951925285f06f7a47a016375e31894872272df57d0c2f30a0c7084ab3e6

Observation 5b8fee62-d60a-4c24-982a-08ffbde2ad7a · outbound

This paper cites Deep residual learning for im- age recognition.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Deep residual learning for im- age recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.940301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.940301Z digest=sha256:2e31ce19c6a17d53e20d7dcf307276701e84d090d0903b8f48fd2f068ea49ada

Observation 277ef9dd-4430-4ac2-84ce-f950cd0255d2 · outbound

This paper cites Densely con- nected convolutional networks.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Densely con- nected convolutional networks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.326434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.943653Z digest=sha256:1838b8fbba65df2a9139360053024ba80cc7f7ed925a7685a876803505d7d395

Observation 8680cf61-593d-422f-a7ab-24d101621042 · outbound

This paper cites Batch normalization: accelerating deep network training by reducing internal covariate shift.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Batch normalization: accelerating deep network training by reducing internal covariate shift

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.314758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.947919Z digest=sha256:abf7cf5a3972ad1420c69d55126d9d228272701b22a5e717f1be5a98a4222cda

Observation 14022fec-0fee-4ade-b581-4aeadc1fc45d · outbound

This paper cites Adam: A method for stochastic optimization.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Adam: A method for stochastic optimization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.302393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.951799Z digest=sha256:e176a66ace22f7ecd4ddb1beeec94e87e31efb477338f891f56910ba3d0ee278

Observation 136da08e-f23b-47f1-9cbe-f0acc8558cf5 · outbound

This paper cites Learning multiple layers of features from tiny images.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Learning multiple layers of features from tiny images

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.955468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.955468Z digest=sha256:a442588dc4a91e4677b75bd5ea42d5a79368a1a920003b8aac0ddcb62af04557

Observation 0b18b826-dd66-4c09-9ecc-2c100c3d7a92 · outbound

This paper cites Decoupled weight decay regularization.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Decoupled weight decay regularization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.958937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.958937Z digest=sha256:30d93b6bcdbc136729c25cd75c3c1616af3732af2e8170876dd36103ec0bb946

Observation 30333d8f-2e0f-45f9-a5e4-418c21ce1513 · outbound

This paper cites an unresolved cited work.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:51:39.273980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.962218Z digest=sha256:2142b973bb395e444f5a1c309a2a51d2b3ac8e7e52a7ed3944e2161d528356ec

Observation d105d30b-bd0c-4d82-b4f0-7a71b96498ea · outbound

This paper cites ViT-CIFAR: PyTorch implementation for Vision Transformer on CIFAR datasets.https://github.com/omihub777/ViT-CIFAR, 2021.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization ViT-CIFAR: PyTorch implementation for Vision Transformer on CIFAR datasets.https://github.com/omihub777/ViT-CIFAR, 2021

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.965544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.965544Z digest=sha256:b908b8bbf1bd51af796d10e5488e1ed9a1236fe3e8cc471158ad5534b6d1f1b4

Observation cab82ace-f151-41f4-8b72-45abeeb5da04 · outbound

This paper cites On the difficulty of training recur- rent neural networks.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization On the difficulty of training recur- rent neural networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.253197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.969144Z digest=sha256:08ff94343aa7e3eb8bcfa847441e89bad6b11897d62695829e875a163e8c82c2

Observation 44fd442c-17b7-4fc6-a684-1050325c3feb · outbound

This paper cites How does batch normalization help optimization? InProceedings of the 32nd International Conference on Neural Information Processing Systems, page 2488–2498, 2018.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization How does batch normalization help optimization? InProceedings of the 32nd International Conference on Neural Information Processing Systems, page 2488–2498, 2018

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.240820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.972958Z digest=sha256:852b93fa061ecc02e89454c7c300b2f74183aeeede59b353fe9a9cba23861066

Observation 5e225c31-4624-4cfe-a521-1c1bde9e2861 · outbound

This paper cites Very deep convolutional networks for large-scale image recognition.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Very deep convolutional networks for large-scale image recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.976502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.976502Z digest=sha256:9dd7e1c8291f1728afa24939a33b71c8a53a562372a0525553d0803e484228f9

Observation eeccafcb-e5fc-4df6-aa39-dfcdf2de9fec · outbound

This paper cites Lecture 6.5 - RMSProp: Divide the gradient by a running average of its recent magnitude.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Lecture 6.5 - RMSProp: Divide the gradient by a running average of its recent magnitude

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.220382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.980056Z digest=sha256:e3843819952ff9d8da779aec0b8ac941a71512396628ca76db8efdf02a23b167

Observation 41d7e801-aebb-46f8-bee6-af6aa73c5519 · outbound

This paper cites Aggregated residual transformations for deep neural networks.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Aggregated residual transformations for deep neural networks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.199732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.987147Z digest=sha256:dde614bc6fef707df50e536401d8def72e4d8a0d9434fb9cb144de51f84de20c

Observation 043e8490-a3aa-401d-8c54-cea17189a075 · outbound

This paper cites Gradient centraliza- tion: A new optimization technique for deep neural networks.Proceedings of the Eu- ropean Conference on Computer Vision (ECCV), pages 635–651, 2020.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Gradient centraliza- tion: A new optimization technique for deep neural networks.Proceedings of the Eu- ropean Conference on Computer Vision (ECCV), pages 635–651, 2020

Reference 24

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T10:51:39.187647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.990904Z digest=sha256:4d15dc3fe6ddc82898be1786f936439bc4dea1f58d33981df2f0705c2af4a814

Observation 6e4fcd63-838c-4864-b798-6e7bacc0f7c0 · outbound

This paper cites Znorm: Z-score gradient normalization accelerating skip-connected network training without architectural modification.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Znorm: Z-score gradient normalization accelerating skip-connected network training without architectural modification

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.175325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.994138Z digest=sha256:d85aa315529f42482087b6432374c3f05eb3f0ccfc4a777c95906d1155bf8440

Observation 1f6f9682-e8f3-43ab-b0f5-5f048a893f88 · outbound

This paper cites Cutmix: Regularization strategy to train strong classifiers with localizable features.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Cutmix: Regularization strategy to train strong classifiers with localizable features

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.162901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:38.997441Z digest=sha256:7b035ce4b86a3507953b6b66f52c88e4566e4f18707f0916b5a96506e4a58476

Observation 03ba7a4c-fef0-45a3-bab9-bfdc4d3a0626 · outbound

This paper cites Wide residual networks.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Wide residual networks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.150168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:39.000807Z digest=sha256:ecb18dbbc9a32bb8abf5f1c2bc85505a656915b16bf47007aa0cf397930200ac

Observation b3896c69-dee6-45ef-87f8-c291c0611120 · outbound

This paper cites When will gradient regularization be harmful? In Forty-first International Conference on Machine Learning.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization When will gradient regularization be harmful? In Forty-first International Conference on Machine Learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.138025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:39.004136Z digest=sha256:6e50380c8984b35f35c83c46ed13c34430401160e1b94149eb92ec13c33c2c30

Observation a4417bfb-ccbe-4eef-b9b1-300ea2cf1166 · outbound

This paper cites Penalizing gradient norm for efficiently improving generalization in deep learning.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Penalizing gradient norm for efficiently improving generalization in deep learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.125921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:39.007740Z digest=sha256:a9da908abb55ecccb250382420abad95c96893fd8323e1594c60c2314a64c35d

Observation 4a462906-6865-4436-bb8f-37482170f4b7 · outbound

This paper cites Recurrent neural networks: vanishing and exploding gradients are not the end of the story.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Recurrent neural networks: vanishing and exploding gradients are not the end of the story

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.113707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:39.011129Z digest=sha256:90da062131da967d5e6005890794c10f013a7e1fab26e620e4143050092da49b

Observation cba70997-9462-43ef-88ca-6d9c72dc12f8 · outbound

This paper cites an unresolved cited work.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:51:39.100030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T10:51:39.015136Z digest=sha256:2d62a669429d00a81049bcfb6f75d2b17a45871c1d0d8608def3cf49ce50bf39

Observation dd4adf9b-d6e3-4efe-8629-5e515d7ffd98 · outbound

This paper cites an unresolved cited work.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Unresolved cited work

Reference 2012

Resolution
parse uncertain
no resolver link, observed 2026-08-05T10:51:38.983623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.983623Z digest=sha256:3ab3d51cdf2b5ed608ce0dbb4eda0c7e6dc5b589f388125b7663ed2035ab8cc4

Pith citing papers

No inbound Pith citation observations are available.