Pith. sign in

Paper Citation Record · LEDGER

Improving Adaptive Moment Optimization via Preconditioner Diagonalization

As of 9 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2502.07488.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07488 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:40:42.772228Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:16:58.221358Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:51:21.208174Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ff3673c-00bf-43f4-9a38-8e1ba1a324fa · outbound

This paper cites Fisher information and natural gradient learning in random deep networks.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Fisher information and natural gradient learning in random deep networks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.204735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.138133Z digest=sha256:addbdb9a1cd5676bb4764da9b189f2bac1a8ffc1044d923b0005465bb9b625d6

Observation a1faf0f7-6720-4769-9149-368f6fb422e5 · outbound

This paper cites Scalable Second Order Optimization for Deep Learning.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Scalable Second Order Optimization for Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.155942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.155942Z digest=sha256:9e7c498ac19720e31e97981a566461d17fa17be5d3fd84ee23374905f9a4043b

Observation 0b42b222-d8a5-4361-90a6-d3e3bbef40fb · outbound

This paper cites Numerical optimization: theoretical and practical aspects.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Numerical optimization: theoretical and practical aspects

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.193082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.166592Z digest=sha256:31a1ec5ea8c285a0f00145094c7b92dad6c78335fc1a86244f498767945f5766

Observation 19311ef6-1f96-40d6-8143-6e36827dd755 · outbound

This paper cites Practical gauss-newton optimisation for deep learning.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Practical gauss-newton optimisation for deep learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.181798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.181235Z digest=sha256:cd286191e21b63556a173330dd627ac16777c26617457f6338f08202809d613d

Observation d80d7340-773c-4fc8-947f-5f75f5656915 · outbound

This paper cites Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.189563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.189563Z digest=sha256:ab255eaf51e2d07a07cd8a92584123f145f2e1963971f236110c0123c960ee50

Observation 6574f7c3-4993-4272-a8ee-6f89730b2b15 · outbound

This paper cites Symbolic discovery of optimization algorithms.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Symbolic discovery of optimization algorithms

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.171270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.201534Z digest=sha256:69cecd491dcb8b3eff6ef33145bed6d9197b61aa3237da07309b747c1840dea6

Observation 4ceb3334-28c6-41de-98e4-2de2786b8804 · outbound

This paper cites an unresolved cited work.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:40:43.160693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.208031Z digest=sha256:25453a57ed80fb3f5aa9f7cf0bc2939396ea991090eea96428353729cdc22178

Observation 9c457a4a-7eaa-41c9-983f-a551320b8fd9 · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Adaptive subgradient methods for online learning and stochastic optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.218356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.218356Z digest=sha256:58a18e316d396cbd5935df89d4c3dd3dae59acf554ac66bf4815091a121b29af

Observation 4d74f84b-e6f9-47f4-a236-c160e529e279 · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Model-agnostic meta-learning for fast adaptation of deep networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.228124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.228124Z digest=sha256:b9be4e0c206c6edfd75ddce501d06fbea8c24b48f91803bc0bc9505b8704ea7b

Observation 739f93e1-af6b-446d-b3c2-9cd88cdc9b72 · outbound

This paper cites Practical methods of optimization.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Practical methods of optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.233646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.233646Z digest=sha256:63808000b04840430c9322088d0faf0d42393e2eb2fd580eac53c22bf442b962

Observation db6573a3-3b95-4023-89ee-e5a2e404548b · outbound

This paper cites Global convergence of stochastic gradient hamiltonian monte carlo for nonconvex stochastic optimization: Nonasymptotic performance bounds and momentum-based acceleration.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Global convergence of stochastic gradient hamiltonian monte carlo for nonconvex stochastic optimization: Nonasymptotic performance bounds and momentum-based acceleration

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.134162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.239916Z digest=sha256:9bb02c2ad3f2ed401082beb52916c3208aa1540dfc66bd2c2b500758b560bdff

Observation ee8bbd5f-4502-4309-8150-a4e8a838e64e · outbound

This paper cites Fast approximate natural gradient descent in a kronecker factored eigenbasis.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Fast approximate natural gradient descent in a kronecker factored eigenbasis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.122673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.243228Z digest=sha256:53592278e9de62cc94203126c952ccf61ef193fe02cf7b83b4aad0da5c0ef2f7

Observation d6aa1a4c-2a00-4a35-947a-52939f3a1752 · outbound

This paper cites Low-rank gradient approximation for memory-efficient on-device training of deep neural network.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Low-rank gradient approximation for memory-efficient on-device training of deep neural network

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.111751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.246592Z digest=sha256:3c13d4bfefa60388f7fa9a51c636339a4acfdea232693c3bc88e5b53ba2ed54f

Observation f1c07a84-fd80-4865-88ff-db80c54f2a29 · outbound

This paper cites A kronecker-factored approximate fisher matrix for convolution layers.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization A kronecker-factored approximate fisher matrix for convolution layers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.249746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.249746Z digest=sha256:f8b1dd864aeabb22ab3e105068d621da846c6c45f0faf38ebe509cf1c48a93ef

Observation 4362436c-8008-4cb4-a035-3faa035aafaa · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Shampoo: Preconditioned stochastic tensor optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.252968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.252968Z digest=sha256:3cd333ff931fa203d50c7ae49f62c6372339a8bfb843d43086a52b73d10503de

Observation efc383d0-48fd-45d1-9082-178b65955821 · outbound

This paper cites Gradient Descent Happens in a Tiny Subspace.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Gradient Descent Happens in a Tiny Subspace

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.256046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.256046Z digest=sha256:8753e5b8e222f7c6ac89c2734b41f1dfec8f1a007a6cbb897f20b4af396149ca

Observation cbd4949f-d459-4d3f-b64a-27ad3c22c975 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Adam: A Method for Stochastic Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.259693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.259693Z digest=sha256:47bf456f81d9b8806a633e59126db0bfe3b11c6c92c10ddb8b2fc95123849b43

Observation 4a56218b-9d7a-4914-8271-0d1a1ef6e4bc · outbound

This paper cites Federated Optimization: Distributed Machine Learning for On-Device Intelligence.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Federated Optimization: Distributed Machine Learning for On-Device Intelligence

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.263049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.263049Z digest=sha256:ab8e4b2ad4762132b31857d131541068e4784ffa1da8691feea0d742695a6297

Observation 4ee2ef16-3de8-498f-9667-616dc150ad73 · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.267822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.267822Z digest=sha256:7226e7e91c2492f63ff3f15d9800673d4320dc0a3b197f1185e94aedb7d1bbb1

Observation d301ac9e-e0c3-4471-8e02-27667e5696b4 · outbound

This paper cites Federated learning: Challenges, methods, and future directions.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Federated learning: Challenges, methods, and future directions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.087643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.271382Z digest=sha256:7199123485fa4b7a60fa1f60cda0794bfa662200ed3cd3b6db487cd3896f8f27

Observation 0e4a53f3-10b6-493e-b57e-f85a8a3f113d · outbound

This paper cites Memory-Efficient LLM Training with Online Subspace Descent.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Memory-Efficient LLM Training with Online Subspace Descent

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.274506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.274506Z digest=sha256:d052bf8698a11203ba2c85dbced8dc999bee69d6585c3409716831f95a52197c

Observation 7e82254b-2dae-417a-a908-44bf23178cdc · outbound

This paper cites Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.278139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.278139Z digest=sha256:4e48bac78de978c437f734eb72e26d94a8dd3aa71da00209087d979f5d7d0a30

Observation 3e7889d9-cf31-4f79-99d8-4259ffa823c1 · outbound

This paper cites Rotate your networks: Better weight consolidation and less catastrophic forgetting.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Rotate your networks: Better weight consolidation and less catastrophic forgetting

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.077676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.281195Z digest=sha256:9f24f97288df56dbe0cd622ca8b5e9e0ecd6f69a2d9b220ae490052bfac245a5

Observation 029ceef8-ab5d-4aca-b05b-4e2f4b180344 · outbound

This paper cites Decoupled Weight Decay Regularization.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Decoupled Weight Decay Regularization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.284368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.284368Z digest=sha256:899430e9ad94784175bb5ba571a1ec0edd75ccfcf3cddf366efcb38762a9f9ba

Observation e8cf7dd5-e2d9-4a41-a976-cdaca4afda18 · outbound

This paper cites Hamiltonian Descent Methods.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Hamiltonian Descent Methods

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.287659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.287659Z digest=sha256:af5c4430bc9b5c91e3f294c1296b89aab148f3a77fb7447e98407e615ba419ad

Observation eef32fcf-6c65-4a70-8af7-aff3bea8de8e · outbound

This paper cites New insights and perspectives on the natural gradient method.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization New insights and perspectives on the natural gradient method

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.312933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.312933Z digest=sha256:d2d162f575d993b754379458a7965829562bdb47bc92f489693493e3a5076627

Observation e76b30be-105c-4f14-baba-1295bb6c8e89 · outbound

This paper cites Optimizing neural networks with kronecker-factored approximate curvature.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Optimizing neural networks with kronecker-factored approximate curvature

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.063560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.347622Z digest=sha256:92863f374a72b99c7c8f055bb211b0dce7adb363677e0add61c4ef511b65c54e

Observation 9bce9ba0-b666-4026-8659-30882c7e084f · outbound

This paper cites Memory-Efficient Optimization with Factorized Hamiltonian Descent.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Memory-Efficient Optimization with Factorized Hamiltonian Descent

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.374967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.374967Z digest=sha256:366f2789134bc8f8f9b818743533669560d096387cc02e2a18b86bb9cde7256a

Observation ebfacd26-ae79-436e-bd40-bb24d510028c · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.402836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.402836Z digest=sha256:230b3f3863359dd418420f2156f6201ca0b8cc66b8e079553e8103ea736e011c

Observation 4c273979-9323-4434-b003-6b635919f17e · outbound

This paper cites On the Convergence of Adam and Beyond.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization On the Convergence of Adam and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.419046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.419046Z digest=sha256:1130b0871a882e07c4f58fc36db5f4742e36f92c250bd08a4b6e509bb3ae4e31

Observation 14d8ddb4-6d87-478f-9b8b-32a33b1b6277 · outbound

This paper cites GLU Variants Improve Transformer.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization GLU Variants Improve Transformer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.449148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.449148Z digest=sha256:1d64e1bb1c7a4c62074a0cd2a2ae8d2e92f9634d963dcd10dc1aed088494b625

Observation 76ef8fd9-bd9d-4880-a84b-7cf82267a834 · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Adafactor: Adaptive learning rates with sublinear memory cost

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.485495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.485495Z digest=sha256:3277d7f731ba1124cdc217e8f024005d3f2e801bdd5dfb242b2aa181ed7188f5

Observation c03ac2ed-81f4-4cef-8001-8964a1fe084f · outbound

This paper cites On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.523157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.523157Z digest=sha256:e622c9de503de6bd87d06f2610c7063ab19decabc0c7696a337c8ca4a8355ef8

Observation 07613736-b5c4-4731-872c-95b963c63c15 · outbound

This paper cites On the Origin of Implicit Regularization in Stochastic Gradient Descent.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization On the Origin of Implicit Regularization in Stochastic Gradient Descent

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.527482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.527482Z digest=sha256:ce818f3bbbee2417fc50fb8a30a0c781af29ebbc334c5ebbab2300c05bbf3e4e

Observation aa28a64b-0392-4235-a35e-cf3d0223a4f9 · outbound

This paper cites Rethinking the inception architecture for computer vision.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Rethinking the inception architecture for computer vision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.531440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.531440Z digest=sha256:5980dc7eda3d9f99d6cc199ca991f703d3c780d7e506c5415851281ee75d35e7

Observation da615b9c-8a5f-450f-ac68-904309ce080b · outbound

This paper cites Recent advances in stochastic gradient descent in deep learning.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Recent advances in stochastic gradient descent in deep learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.038273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.535692Z digest=sha256:cee39f5e208f181fd281bab8544f36cf791546e4a06dc2484cfed88518965feb

Observation 7b21e59c-053e-42b9-99a5-bfdbe53db156 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.539056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.539056Z digest=sha256:228e03f6e31478716a9f3686a535046457f53bf4315d04c67f576cfe13ec0d00

Observation 85f06c53-e8cb-48e6-8f05-71f606f0d17c · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization SOAP: Improving and Stabilizing Shampoo using Adam

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.542090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.542090Z digest=sha256:c849453fe5e495da6812c04af7008bfe92f0e38f8ad40b06a8a776d931bada10

Observation c1c76ea2-26c3-4e6e-a6a4-a345fe1028d1 · outbound

This paper cites Root mean square layer normalization.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Root mean square layer normalization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.594954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.594954Z digest=sha256:d45dcb4bc747b2464e5234b197e58bd2dd869980cc58ab1154f7b324f452db64

Observation c919e508-7103-45e9-9caa-f3bf0e7be8d6 · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization mixup: Beyond Empirical Risk Minimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.712127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.712127Z digest=sha256:98722ebe5eaf9f494774cb03d77adab667a71962982c3ae5a2b686b602326f1f

Observation c6255f01-01fc-4fc8-804e-e9dc9e45100b · outbound

This paper cites Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.024078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.756600Z digest=sha256:61b8866c2f00a8dbddb7ee05d63482b413f4a16e747dffc3c519674a1b0b3626

Observation f8377d3a-7a91-4b77-8ca3-1083b32cc39e · outbound

This paper cites Adam can converge without any modification on update rules.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Adam can converge without any modification on update rules

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.013528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.759926Z digest=sha256:09d029ae78f402efab76246e7fb8fb4cbe9a26158e6df18f00af945d73a45f10

Observation 7a7411e0-c16f-4ab8-84f3-1b891eecdaa3 · outbound

This paper cites Why Transformers Need Adam: A Hessian Perspective.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Why Transformers Need Adam: A Hessian Perspective

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.763594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.763594Z digest=sha256:127de5f1d2487116f593886413f137ca0d3dfe857b4eab7bd1707760ee463cd6

Observation 5d7ab80c-32a2-4b23-b766-31a4adc02603 · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T12:40:42.767627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:40:42.767627Z digest=sha256:d717508fa57639c1ef8a16097281c348b8855ff1f2379930756fbb108be06e23

Observation 48ae89c5-412a-4f97-b9fe-72f49e1e412f · outbound

This paper cites Towards theoretically understanding why sgd generalizes better than adam in deep learning.

Improving Adaptive Moment Optimization via Preconditioner Diagonalization Towards theoretically understanding why sgd generalizes better than adam in deep learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:40:43.001783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T12:40:42.772228Z digest=sha256:bd332fa0bdaacb2a15176ab6cdf444ffbf00fa7b7ca5b776a041425efa9375cc

Pith citing papers

Observation e1060d59-b93e-4c5d-ab10-06d97084637f · inbound

Evolution of Optimization Methods: Algorithms, Scenarios, and Evaluations cites this paper.

Evolution of Optimization Methods: Algorithms, Scenarios, and Evaluations Improving Adaptive Moment Optimization via Preconditioner Diagonalization

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:51:21.226714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:16:58.221358Z digest=sha256:f6c9dc522284ca78cb5906a8d7c1b03094583fd09f4d3b3e11acba2d8ee5f6ef