Pith. sign in

Paper Citation Record · LEDGER

Improving Adaptive Moment Optimization via Preconditioner Diagonalization

As of 9 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2502.07488.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07488 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:40:42.772228Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:16:58.221358Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:51:21.208174Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ff3673c-00bf-43f4-9a38-8e1ba1a324fa · outbound

This paper cites Fisher information and natural gradient learning in random deep networks.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Fisher information and natural gradient learning in random deep networks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.204735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.138133Z digest=sha256:b797aa9e8f0cc0047f9630a1010646968fcdabfe3b0ba9fcf0ab5a49fe3d56fa

Observation a1faf0f7-6720-4769-9149-368f6fb422e5 · outbound

This paper cites Scalable Second Order Optimization for Deep Learning.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Scalable Second Order Optimization for Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.155942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.155942Z digest=sha256:73c92313743eb9f11b14cdcb8ca8182007325abf9acf70db4fbade0b9906c623

Observation 0b42b222-d8a5-4361-90a6-d3e3bbef40fb · outbound

This paper cites Numerical optimization: theoretical and practical aspects.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Numerical optimization: theoretical and practical aspects

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.193082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.166592Z digest=sha256:2a2c8cf22340a7e78047cca2503ff95b825d2e7ccb54bf0c9918b248afafc630

Observation 19311ef6-1f96-40d6-8143-6e36827dd755 · outbound

This paper cites Practical gauss-newton optimisation for deep learning.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Practical gauss-newton optimisation for deep learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.181798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.181235Z digest=sha256:aaf13cfb87484d0f34d2246f3ab8123be2ed14e28ba32382c878e8ce4d9ab460

Observation d80d7340-773c-4fc8-947f-5f75f5656915 · outbound

This paper cites Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.189563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.189563Z digest=sha256:7b595e0dafdd76f26b840c87dc5ce4191e7aefed286eef011f3179e37312d7d0

Observation 6574f7c3-4993-4272-a8ee-6f89730b2b15 · outbound

This paper cites Symbolic discovery of optimization algorithms.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Symbolic discovery of optimization algorithms

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.171270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.201534Z digest=sha256:c5d0d02f2b859562b28c626ad43a9ba8704777a36a4760ca37cac229b6fc85c2

Observation 4ceb3334-28c6-41de-98e4-2de2786b8804 · outbound

This paper cites an unresolved cited work.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:40:43.160693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.208031Z digest=sha256:28ccd1fd682db04972b2188e7186a5fa46273552aef01915325785504c69c236

Observation 9c457a4a-7eaa-41c9-983f-a551320b8fd9 · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Adaptive subgradient methods for online learning and stochastic optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.218356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.218356Z digest=sha256:fc0e5b6682d266b92334c6dd70fdac38586cd0de79fd4dc5066d3ce2051f69f1

Observation 4d74f84b-e6f9-47f4-a236-c160e529e279 · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Model-agnostic meta-learning for fast adaptation of deep networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.228124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.228124Z digest=sha256:9b8f4d898dc95132552d58ef1435b73bab0d064b4699b24356588c9767f966b2

Observation 739f93e1-af6b-446d-b3c2-9cd88cdc9b72 · outbound

This paper cites Practical methods of optimization.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Practical methods of optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.233646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.233646Z digest=sha256:a15c61d8c43e59e061fc0d76e20b858a3e89a9f562177edd71d45e77c0443b8e

Observation db6573a3-3b95-4023-89ee-e5a2e404548b · outbound

This paper cites Global convergence of stochastic gradient hamiltonian monte carlo for nonconvex stochastic optimization: Nonasymptotic performance bounds and momentum-based acceleration.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Global convergence of stochastic gradient hamiltonian monte carlo for nonconvex stochastic optimization: Nonasymptotic performance bounds and momentum-based acceleration

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.134162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.239916Z digest=sha256:c5a8953a07e2a28ace89f9c808266a70021e0335962683eba4c49af0085cd44b

Observation ee8bbd5f-4502-4309-8150-a4e8a838e64e · outbound

This paper cites Fast approximate natural gradient descent in a kronecker factored eigenbasis.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Fast approximate natural gradient descent in a kronecker factored eigenbasis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.122673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.243228Z digest=sha256:79f3fec0ca9e3ea56c4a99b0fb9d0c1be77e391aee66292ec0475d6f8bda1a3d

Observation d6aa1a4c-2a00-4a35-947a-52939f3a1752 · outbound

This paper cites Low-rank gradient approximation for memory-efficient on-device training of deep neural network.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Low-rank gradient approximation for memory-efficient on-device training of deep neural network

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.111751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.246592Z digest=sha256:ae4b4bd72b0ef826fa685a48354df28411fa455843759f9eca4bd02e0376c1c1

Observation f1c07a84-fd80-4865-88ff-db80c54f2a29 · outbound

This paper cites A kronecker-factored approximate fisher matrix for convolution layers.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization A kronecker-factored approximate fisher matrix for convolution layers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.249746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.249746Z digest=sha256:01a7dc802391415a5e899ab3fb2ee84fb0a52c4440538db5371c1fa112cd4086

Observation 4362436c-8008-4cb4-a035-3faa035aafaa · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Shampoo: Preconditioned stochastic tensor optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.252968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.252968Z digest=sha256:98900f630c5c72ca2b724ddf61bf9eaa661e03dff6923880b5613e443b2ac457

Observation efc383d0-48fd-45d1-9082-178b65955821 · outbound

This paper cites Gradient Descent Happens in a Tiny Subspace.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Gradient Descent Happens in a Tiny Subspace

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.256046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.256046Z digest=sha256:79f146e432863e15550520a38ddf0614d412cfe2498132069d98d2a3e06ca71a

Observation cbd4949f-d459-4d3f-b64a-27ad3c22c975 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Adam: A Method for Stochastic Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.259693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.259693Z digest=sha256:accc397a059c3bf4652b35645654d5c07b6891c12f7685af5ec4f3ae9338a2a1

Observation 4a56218b-9d7a-4914-8271-0d1a1ef6e4bc · outbound

This paper cites Federated Optimization: Distributed Machine Learning for On-Device Intelligence.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Federated Optimization: Distributed Machine Learning for On-Device Intelligence

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.263049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.263049Z digest=sha256:756b6dd17134c107add76e3d59c34a1ba23c0877f0fe257b58c08332786c753c

Observation 4ee2ef16-3de8-498f-9667-616dc150ad73 · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.267822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.267822Z digest=sha256:5012d02784e535f5af1f948a510b20978be24859f38810e4f7b830d9a18cede8

Observation d301ac9e-e0c3-4471-8e02-27667e5696b4 · outbound

This paper cites Federated learning: Challenges, methods, and future directions.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Federated learning: Challenges, methods, and future directions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.087643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.271382Z digest=sha256:d9ec78cfe25a31fb958ea43a5db8034ea3476bcdbcb2f3715a1675173c661120

Observation 0e4a53f3-10b6-493e-b57e-f85a8a3f113d · outbound

This paper cites Memory-Efficient LLM Training with Online Subspace Descent.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Memory-Efficient LLM Training with Online Subspace Descent

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.274506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.274506Z digest=sha256:ffc5ad907ddc37e40ec55d7d4dc1b10875e09bff81f22bac5cbba24168fd315f

Observation 7e82254b-2dae-417a-a908-44bf23178cdc · outbound

This paper cites Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.278139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.278139Z digest=sha256:5ceebc299beb7895378444a6d690f8b757fbade4e6f773012e4862d22f513e56

Observation 3e7889d9-cf31-4f79-99d8-4259ffa823c1 · outbound

This paper cites Rotate your networks: Better weight consolidation and less catastrophic forgetting.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Rotate your networks: Better weight consolidation and less catastrophic forgetting

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.077676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.281195Z digest=sha256:494753ab35938704a896626218d8237b31e1337411f38b02ac0d5d4d913ed7e5

Observation 029ceef8-ab5d-4aca-b05b-4e2f4b180344 · outbound

This paper cites Decoupled Weight Decay Regularization.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Decoupled Weight Decay Regularization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.284368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.284368Z digest=sha256:8c4317aa5e85d6e5c0f48ef4499e97bea060dd549d99d728bcff2b8cd2398897

Observation e8cf7dd5-e2d9-4a41-a976-cdaca4afda18 · outbound

This paper cites Hamiltonian Descent Methods.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Hamiltonian Descent Methods

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.287659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.287659Z digest=sha256:03d735e75db0924b85c15b517066d46884ab013e25aa1367c8662c00bf6400d3

Observation eef32fcf-6c65-4a70-8af7-aff3bea8de8e · outbound

This paper cites New insights and perspectives on the natural gradient method.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization New insights and perspectives on the natural gradient method

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.312933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.312933Z digest=sha256:39d66c0ea89333ea82f6aabd7ff4533d1e3ad05eb9e0703bc439bda6e8657920

Observation e76b30be-105c-4f14-baba-1295bb6c8e89 · outbound

This paper cites Optimizing neural networks with kronecker-factored approximate curvature.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Optimizing neural networks with kronecker-factored approximate curvature

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.063560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.347622Z digest=sha256:c23f0e091af2aadc2baf30ac6e804ff44e8769cd2a8647768d649a5705a705e6

Observation 9bce9ba0-b666-4026-8659-30882c7e084f · outbound

This paper cites Memory-Efficient Optimization with Factorized Hamiltonian Descent.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Memory-Efficient Optimization with Factorized Hamiltonian Descent

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.374967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.374967Z digest=sha256:bae62ebc26c7c3f3a5c42367a1b82b76352f49bab0b80fa66f3fb6d043f343f5

Observation ebfacd26-ae79-436e-bd40-bb24d510028c · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.402836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.402836Z digest=sha256:30f63500f47139e78917e0d82cdabe97f6baa03a487460f8d534f663d7b30cb1

Observation 4c273979-9323-4434-b003-6b635919f17e · outbound

This paper cites On the Convergence of Adam and Beyond.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization On the Convergence of Adam and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.419046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.419046Z digest=sha256:1c52cfd92b2e261db5037a63c7b558f75ba8921c10d102904b1fb75f9f8a0cf7

Observation 14d8ddb4-6d87-478f-9b8b-32a33b1b6277 · outbound

This paper cites GLU Variants Improve Transformer.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization GLU Variants Improve Transformer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.449148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.449148Z digest=sha256:be9db9031a448d306e2a174d614c19a0e8fba9c5beedb11475965dbe6f1b557e

Observation 76ef8fd9-bd9d-4880-a84b-7cf82267a834 · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Adafactor: Adaptive learning rates with sublinear memory cost

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.485495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.485495Z digest=sha256:ecae59e182cef3d9e949ea40ea4beb4347cd65dfb0d0d93bc17dd178e1a79f97

Observation c03ac2ed-81f4-4cef-8001-8964a1fe084f · outbound

This paper cites On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.523157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.523157Z digest=sha256:9bb7b14ecd50a7d6d9eb84c8e5d5c13b396c4de433c03a73ffd6de72c496b7f0

Observation 07613736-b5c4-4731-872c-95b963c63c15 · outbound

This paper cites On the Origin of Implicit Regularization in Stochastic Gradient Descent.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization On the Origin of Implicit Regularization in Stochastic Gradient Descent

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.527482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.527482Z digest=sha256:c44530d6ac238e64b341f975a7ef4fe415a419eeb40b224d578e65fdc06fdf18

Observation aa28a64b-0392-4235-a35e-cf3d0223a4f9 · outbound

This paper cites Rethinking the inception architecture for computer vision.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Rethinking the inception architecture for computer vision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.531440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.531440Z digest=sha256:95ae6173d1fd960e18ba04d23a4187ff2a1ac79d486797b3ae88c5d1a7e7bc05

Observation da615b9c-8a5f-450f-ac68-904309ce080b · outbound

This paper cites Recent advances in stochastic gradient descent in deep learning.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Recent advances in stochastic gradient descent in deep learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.038273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.535692Z digest=sha256:1d882e4c3242ebd884800c9d5f4e49ca64aa8fae96d53ef0b4b5f2585fe699ef

Observation 7b21e59c-053e-42b9-99a5-bfdbe53db156 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.539056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.539056Z digest=sha256:6995189f49023cba798660d2338393429c732b803982f1ba54f900ab595df95a

Observation 85f06c53-e8cb-48e6-8f05-71f606f0d17c · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization SOAP: Improving and Stabilizing Shampoo using Adam

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.542090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.542090Z digest=sha256:c0d8236c9d4ff0ee601b89d9d3387c9f48b07676f221b08577299486172358ea

Observation c1c76ea2-26c3-4e6e-a6a4-a345fe1028d1 · outbound

This paper cites Root mean square layer normalization.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Root mean square layer normalization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.594954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.594954Z digest=sha256:c00ecb475ba976be76a04fe7b342f6075795faf2322d200e5bd4a5bb79fc8a67

Observation c919e508-7103-45e9-9caa-f3bf0e7be8d6 · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization mixup: Beyond Empirical Risk Minimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.712127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.712127Z digest=sha256:a82421a3a8ebff5d33c17238c374d821a5aefc7901d7a3e121abc3c538052f8f

Observation c6255f01-01fc-4fc8-804e-e9dc9e45100b · outbound

This paper cites Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.024078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.756600Z digest=sha256:b92bb2b33f709b348c9f5bd6e24fa991e7fa7fe7319d4b70f41c0715868cc590

Observation f8377d3a-7a91-4b77-8ca3-1083b32cc39e · outbound

This paper cites Adam can converge without any modification on update rules.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Adam can converge without any modification on update rules

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.013528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.759926Z digest=sha256:463543b0d3f1e95b0c4f6082b58804ff3aaaf3e8d48d5683f87a4baafc5ac212

Observation 7a7411e0-c16f-4ab8-84f3-1b891eecdaa3 · outbound

This paper cites Why Transformers Need Adam: A Hessian Perspective.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Why Transformers Need Adam: A Hessian Perspective

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.763594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.763594Z digest=sha256:96102278e8f005b3030ed9bc5267508c42310cdb4a72d8084fe0a1fbcaacc0f0

Observation 5d7ab80c-32a2-4b23-b766-31a4adc02603 · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.767627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.767627Z digest=sha256:4b26641cd89075861f01612867bd6eb74ade238ae389417874cb9690e1fd33cc

Observation 48ae89c5-412a-4f97-b9fe-72f49e1e412f · outbound

This paper cites Towards theoretically understanding why sgd generalizes better than adam in deep learning.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Towards theoretically understanding why sgd generalizes better than adam in deep learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.001783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.772228Z digest=sha256:dbb047f8d47a0605005a7e6799a96bcb182152269c86aa025e24432e6580efc7

Pith citing papers

Observation e1060d59-b93e-4c5d-ab10-06d97084637f · inbound

Evolution of Optimization Methods: Algorithms, Scenarios, and Evaluations cites this paper.

Evolution of Optimization Methods: Algorithms, Scenarios, and Evaluations Improving Adaptive Moment Optimization via Preconditioner Diagonalization

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:51:21.226714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:16:58.221358Z digest=sha256:cf19c1ab3cdbf08c408c01f704f98e1cef8648d1ba35bac52bd3ce764dad000b