Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:51:39.015136Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2509.03677.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:51:39.015136Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2e8a9f65-878d-44ad-9fdb-c2c37f4b8f47 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Layer Normalization
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1407376e-2cff-4088-9ec5-946660a7db70 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Large-scale machine learning with stochastic gradient descent
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cca347b-b28b-4c88-b834-82b04607df9a · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization The tradeoffs of large scale learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5eb220e4-b18d-4c94-87c2-831deb57219a · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Curtis, and Jorge Nocedal
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 508674f8-9780-4928-8f92-23f6ddc6c4fd · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Entropy-sgd: biasing gradient descent into wide valleys.Journal of Statistical Mechanics: Theory and Experiment, 2019 (12):124018, 2019
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50dd9eb9-c3a2-4dea-bfc4-8127695d03bd · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Gradnorm: Gra- dient normalization for adaptive loss balancing in deep multitask networks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation daf8b0b3-9c90-4ece-9d20-2d14d4ef2401 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization A Study of Gradient Variance in Deep Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39c137a0-1bb3-4d63-a904-e524da2847fc · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Understanding the difficulty of training deep feedfor- ward neural networks
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b9b9e64-56f9-45ad-82a7-5bf40561b76f · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Take a shortcut back: Mitigating the gradient vanishing for training spiking neural networks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ded4a05-1d21-438d-876b-2a6eb660e9db · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Stable architectures for deep neural networks.Inverse Prob- lems, 34(1):014004, 2018
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b8fee62-d60a-4c24-982a-08ffbde2ad7a · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Deep residual learning for im- age recognition
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 277ef9dd-4430-4ac2-84ce-f950cd0255d2 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Densely con- nected convolutional networks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8680cf61-593d-422f-a7ab-24d101621042 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Batch normalization: accelerating deep network training by reducing internal covariate shift
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14022fec-0fee-4ade-b581-4aeadc1fc45d · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Adam: A method for stochastic optimization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 136da08e-f23b-47f1-9cbe-f0acc8558cf5 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Learning multiple layers of features from tiny images
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b18b826-dd66-4c09-9ecc-2c100c3d7a92 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Decoupled weight decay regularization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30333d8f-2e0f-45f9-a5e4-418c21ce1513 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d105d30b-bd0c-4d82-b4f0-7a71b96498ea · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization ViT-CIFAR: PyTorch implementation for Vision Transformer on CIFAR datasets.https://github.com/omihub777/ViT-CIFAR, 2021
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cab82ace-f151-41f4-8b72-45abeeb5da04 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization On the difficulty of training recur- rent neural networks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44fd442c-17b7-4fc6-a684-1050325c3feb · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization How does batch normalization help optimization? InProceedings of the 32nd International Conference on Neural Information Processing Systems, page 2488–2498, 2018
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e225c31-4624-4cfe-a521-1c1bde9e2861 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Very deep convolutional networks for large-scale image recognition
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeccafcb-e5fc-4df6-aa39-dfcdf2de9fec · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Lecture 6.5 - RMSProp: Divide the gradient by a running average of its recent magnitude
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41d7e801-aebb-46f8-bee6-af6aa73c5519 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Aggregated residual transformations for deep neural networks
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 043e8490-a3aa-401d-8c54-cea17189a075 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Gradient centraliza- tion: A new optimization technique for deep neural networks.Proceedings of the Eu- ropean Conference on Computer Vision (ECCV), pages 635–651, 2020
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e4fcd63-838c-4864-b798-6e7bacc0f7c0 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Znorm: Z-score gradient normalization accelerating skip-connected network training without architectural modification
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f6f9682-e8f3-43ab-b0f5-5f048a893f88 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Cutmix: Regularization strategy to train strong classifiers with localizable features
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03ba7a4c-fef0-45a3-bab9-bfdc4d3a0626 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Wide residual networks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3896c69-dee6-45ef-87f8-c291c0611120 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization When will gradient regularization be harmful? In Forty-first International Conference on Machine Learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4417bfb-ccbe-4eef-b9b1-300ea2cf1166 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Penalizing gradient norm for efficiently improving generalization in deep learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a462906-6865-4436-bb8f-37482170f4b7 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Recurrent neural networks: vanishing and exploding gradients are not the end of the story
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cba70997-9462-43ef-88ca-6d9c72dc12f8 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd4adf9b-d6e3-4efe-8629-5e515d7ffd98 · outbound
Insights from Gradient Dynamics: Gradient Autoscaled Normalization Unresolved cited work
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.