Pith. sign in

Paper Citation Record · LEDGER

On the Variance of the Adaptive Learning Rate and Beyond

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 56 inbound Pith citation observations for arXiv:1908.03265.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.03265 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 56 of 56 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:29:22.668886Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

608
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 43527666-8bab-44bb-b011-16202f68355f · inbound

Consistency Models cites this paper.

Consistency Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:47:28.659081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T15:47:28.598205Z digest=sha256:afa85fb8abc7c208c9151f705e0b1a30dc532f5eb6a6fbd05120a8ef04cd560a

Observation 760fe633-d1de-41b2-9955-934cca7e5fef · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.197627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:1919878cfd14d5a115acc272922290c93dee9d62dacfc86400a0f3cc82bdc043

Observation 216fd231-8fdc-4802-a6af-9bb99ed6622e · inbound

Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference cites this paper.

Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference On the Variance of the Adaptive Learning Rate and Beyond

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:15:54.867502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T04:15:54.681913Z digest=sha256:553e9c812feeaacd76027713ba0068160f3e4e663cc8e76685b48a89a16f38d9

Observation 18eb65d8-34fc-45ea-a560-afdec281900c · inbound

Improved Techniques for Training Consistency Models cites this paper.

Improved Techniques for Training Consistency Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:04:10.662000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T05:04:10.600062Z digest=sha256:2cc4bd0f7c894e489ce3374a33f9843703ea487323e2c37d783feb7f019a9cd4

Observation ab8f389a-b751-480d-901d-8f36cad52591 · inbound

Learning state and proposal dynamics in state-space models using differentiable particle filters and neural networks cites this paper.

Learning state and proposal dynamics in state-space models using differentiable particle filters and neural networks On the Variance of the Adaptive Learning Rate and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:11:16.550216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:11:16.550216Z digest=sha256:ce44cb352f683ecdb470e32c4a949624bda697fa67c237dae8e6839152b78f7b

Observation e000a8d5-f4b0-4d2b-88e4-efff3ef24cd6 · inbound

Beyond adaptive gradient: Fast-Controlled Minibatch Algorithm for large-scale optimization cites this paper.

Beyond adaptive gradient: Fast-Controlled Minibatch Algorithm for large-scale optimization On the Variance of the Adaptive Learning Rate and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:39.847015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:39.847015Z digest=sha256:a4bad4c07e5c176b4fdf7a6184d633c0d3eda5c6804bae6f842dc680c7272589

Observation 4a93ff98-854e-4555-9890-62befd0919e2 · inbound

Quantity versus Diversity: Influence of Data on Detecting EEG Pathology with Advanced ML Models cites this paper.

Quantity versus Diversity: Influence of Data on Detecting EEG Pathology with Advanced ML Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T21:29:22.668886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:29:22.668886Z digest=sha256:fa292cac21e3d57d2007ee603039ad0b8e0e6bea604a7c0a49d294ed2dd488b6

Observation 7e4ce5f8-ae17-43a7-a779-efda4c0986a0 · inbound

CAdam: Confidence-Based Optimization for Online Learning cites this paper.

CAdam: Confidence-Based Optimization for Online Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T06:03:51.315324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:03:51.315324Z digest=sha256:edf461dceccc84be22eb87822045998de2095f65f52cd8e38be7b6f73d81caba

Observation f72ffc1a-e400-4bde-a3a1-d8ddbcfa11fe · inbound

Automatic Differentiation-based Full Waveform Inversion with Flexible Workflows cites this paper.

Automatic Differentiation-based Full Waveform Inversion with Flexible Workflows On the Variance of the Adaptive Learning Rate and Beyond

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:44.820508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:26:44.820508Z digest=sha256:ddf11c23c436e241bd660e971e521cbd5b63039cb846a87b247256fe2ef4fd1e

Observation 355fd820-b7cc-4138-a443-29807ff5ffbf · inbound

Some Best Practices in Operator Learning cites this paper.

Some Best Practices in Operator Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:03.709142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:03.709142Z digest=sha256:291cc7d2e8d59314e2dfdcd2e96d73cc0649faa60ecde347dc916fd32e311996

Observation ddff11e7-46f3-4f80-83da-3fa88889e118 · inbound

No More Adam: Learning Rate Scaling at Initialization is All You Need cites this paper.

No More Adam: Learning Rate Scaling at Initialization is All You Need On the Variance of the Adaptive Learning Rate and Beyond

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.366975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.366975Z digest=sha256:ef90f711198b2950566b80916d798e5f21f45fc3cc8edea661f8f9feb622f380

Observation 7aa1e192-1da6-423c-b252-6cab0852ad95 · inbound

Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data cites this paper.

Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data On the Variance of the Adaptive Learning Rate and Beyond

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T11:32:54.039495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:32:54.039495Z digest=sha256:72e22d0164d5faa24c3999d13c31e6d3034c1365cf98942d5995a11344886c5b

Observation 128c70fd-0c9b-4d1b-8255-ca8d832252b5 · inbound

From Histopathology Images to Cell Clouds: Learning Slide Representations with Hierarchical Cell Transformer cites this paper.

From Histopathology Images to Cell Clouds: Learning Slide Representations with Hierarchical Cell Transformer On the Variance of the Adaptive Learning Rate and Beyond

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:04.385407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:04.385407Z digest=sha256:46b87652e34272787a86000b110437cea0c780d632f0c290b90f8f5f836eb950

Observation 9d0f8ad0-3632-4a27-9641-b8111129220b · inbound

ReFlow6D: Refraction-Guided Transparent Object 6D Pose Estimation via Intermediate Representation Learning cites this paper.

ReFlow6D: Refraction-Guided Transparent Object 6D Pose Estimation via Intermediate Representation Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:14:17.622839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:14:17.622839Z digest=sha256:81b900f333bdb352c96853ed5e500193c93ae81a8bfbabbc7bcfbfb34bd5c2ef

Observation 101104e3-08fe-4b5d-b2a8-05102d73849c · inbound

Celo: Training Versatile Learned Optimizers on a Compute Diet cites this paper.

Celo: Training Versatile Learned Optimizers on a Compute Diet On the Variance of the Adaptive Learning Rate and Beyond

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T17:01:13.908163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:01:13.908163Z digest=sha256:40002abc5ad8caee5cb3104b35fe5e5c1fbf64821b90d65b56f434d576956329

Observation 41337159-851c-4c38-8ebb-f4896e234754 · inbound

Certified Guidance for Planning with Deep Generative Models cites this paper.

Certified Guidance for Planning with Deep Generative Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T16:51:20.102451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:51:20.102451Z digest=sha256:e00c5fc7a5bccde9137411ad74a417122b323067af7830fa44c57fd2eeefbab1

Observation 6c36cb10-3e04-4742-978e-6ea2c9d7138b · inbound

Unlearning Clients, Features and Samples in Vertical Federated Learning cites this paper.

Unlearning Clients, Features and Samples in Vertical Federated Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T15:46:20.139219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:46:20.139219Z digest=sha256:bd84dcfafec14e7b8f1df2bd8be98c8c66f008bc6566713e6aba43521df843c6

Observation f125ef0f-c08d-4f99-a0e3-9afff26d105e · inbound

CoSTI: Consistency Models for (a faster) Spatio-Temporal Imputation cites this paper.

CoSTI: Consistency Models for (a faster) Spatio-Temporal Imputation On the Variance of the Adaptive Learning Rate and Beyond

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T20:27:53.458185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:27:53.458185Z digest=sha256:9d5c5e42d594632372e6b4048d7589ead709706ad46fd30f5387501d7ddd8e11

Observation 9b8548ed-6392-44bd-aba0-c0307459498a · inbound

Rapid training of Hamiltonian graph networks using random features cites this paper.

Rapid training of Hamiltonian graph networks using random features On the Variance of the Adaptive Learning Rate and Beyond

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:17:15.702752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-19T10:17:00.343410Z digest=sha256:08921722ee9b54fed190494c378a24d61b5d7a30a4f9e90ce6762d0da5bd02e5

Observation 05fd9f67-538a-4a5b-975b-c91c9b65fd08 · inbound

Deep Potential: Recovering the gravitational potential and local pattern speed in the solar neighborhood with GDR3 using normalizing flows cites this paper.

Deep Potential: Recovering the gravitational potential and local pattern speed in the solar neighborhood with GDR3 using normalizing flows On the Variance of the Adaptive Learning Rate and Beyond

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T20:07:26.301836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:07:26.301836Z digest=sha256:863be4ab8bb0393d4e3b398d63b5eb48fed2fdb4574a7afc42f7e226c1aade81

Observation 218ef676-cdc5-4fff-9afb-a88f3cad0e79 · inbound

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam cites this paper.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the Variance of the Adaptive Learning Rate and Beyond

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.591929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.591929Z digest=sha256:7d8b632fd917c5b00d0957f36d71f40df3d8cb2634e896a575f778fb8e3ee235

Observation 9e6a5159-a688-48c3-9649-7e32726407b1 · inbound

Feature-Enhanced TResNet for Fine-Grained Food Image Classification cites this paper.

Feature-Enhanced TResNet for Fine-Grained Food Image Classification On the Variance of the Adaptive Learning Rate and Beyond

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:51.780376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:51.780376Z digest=sha256:d1b2e6a289b5769e503297f07ec9f68c3f7b03862a84f73e4f636d1d48189c1b

Observation 80f58d11-afa8-4b7e-b467-8b8a79d71466 · inbound

Criticality analysis of nuclear binding energy neural networks cites this paper.

Criticality analysis of nuclear binding energy neural networks On the Variance of the Adaptive Learning Rate and Beyond

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T06:01:11.597771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:01:11.597771Z digest=sha256:6be95ee9913ec1cfb0dbf527dae214188c830310e543ff158c22a5b54068da46

Observation d42c4cf1-b7c5-41ff-983f-823e41ef7f6f · inbound

From Next Token Prediction to (STRIPS) World Models cites this paper.

From Next Token Prediction to (STRIPS) World Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:21:36.462305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-18T16:21:22.805473Z digest=sha256:ac2215967d8326bfdf71140c82f79caf5b91d8cc59b919670674f4f16d2b8e66

Observation 37157420-cf3a-4228-9c3f-f79600f54f5b · inbound

From Next Token Prediction to (STRIPS) World Models cites this paper.

From Next Token Prediction to (STRIPS) World Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T16:41:18.775357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T16:41:18.775357Z digest=sha256:5d874177edd241777dd46efcb53148cbbe956f33f8c210928f244e882ac49ffe

Observation 53416c7b-10c1-493a-a408-472c2dd13e13 · inbound

On the Convergence of Muon and Beyond cites this paper.

On the Convergence of Muon and Beyond On the Variance of the Adaptive Learning Rate and Beyond

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:56:34.102980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-18T15:56:30.602824Z digest=sha256:fca7a6ac86be8942a9f4ee52cca4426d491936d3dd73622fb70c42514056ef6c

Observation 4a969c6c-cba2-4130-8c93-7308ff18d135 · inbound

Why Do We Need Warm-up? A Theoretical Perspective cites this paper.

Why Do We Need Warm-up? A Theoretical Perspective On the Variance of the Adaptive Learning Rate and Beyond

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.819676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:58.819676Z digest=sha256:9cc5bafb449e2c0e2be71f52e1d222c16d8613eb2d0c0ff7ec3c1f90f27fb429

Observation 8606aaef-40c8-497f-ba60-d25e8cb2dc9c · inbound

Building Deep Graph Predictors with Graph Imitation Learning cites this paper.

Building Deep Graph Predictors with Graph Imitation Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:45:18.726240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T15:44:58.327091Z digest=sha256:ba9c6ff6b6b4021e9172eb48cfabe2291f6ed7ed247af2d5b2cd514e516618e6

Observation 1cceda00-a532-4936-a3b8-d125000f1e9a · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:45:29.768850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-25T07:41:48.634351Z digest=sha256:971fce56f97b3c28ba5498b94f3036797801f8d2e673b4ba08c8729e406eddd7

Observation 6999988d-3e60-41a8-b3a6-590f5e3cf82e · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T05:37:38.449157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:37:38.449157Z digest=sha256:ec88244f8048b5d9a92ea4f2457778e00cd90af5277ec4da66088ebab3e2021c

Observation becc2260-cb50-4253-b2ff-542c3fcc3f4b · inbound

Characterizing the Instrumental Profile of LAMOST cites this paper.

Characterizing the Instrumental Profile of LAMOST On the Variance of the Adaptive Learning Rate and Beyond

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:10:03.181485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T14:09:12.016309Z digest=sha256:ea47a18899ba9a2abb4649849866382f1711f4a5c251989c7d156c62d8017d9a

Observation 6b7d67b1-d767-4ff2-851f-29266263f10f · inbound

Video-guided Machine Translation with Global Video Context cites this paper.

Video-guided Machine Translation with Global Video Context On the Variance of the Adaptive Learning Rate and Beyond

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:01.141997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T18:02:31.374359Z digest=sha256:6b3e7b89362a9810a2d32366b1f6dea9a822d975ef15c7dc13cf222a4c95a55f

Observation e23ddd4e-e4dc-4372-856d-6bf8144e9d3e · inbound

Delve into the Applicability of Advanced Optimizers for Multi-Task Learning cites this paper.

Delve into the Applicability of Advanced Optimizers for Multi-Task Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:20:59.066359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:42:48.902337Z digest=sha256:948f6212af1533b95b57fc3c26238546a40c0efbcf64b6d65926aaedf22c1a06

Observation c01c0deb-37e6-4e97-bab7-0c246942c6dc · inbound

Particle transformers for identifying Lorentz-boosted Higgs bosons decaying to a pair of W bosons cites this paper.

Particle transformers for identifying Lorentz-boosted Higgs bosons decaying to a pair of W bosons On the Variance of the Adaptive Learning Rate and Beyond

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:01:00.835334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:18:12.735038Z digest=sha256:a834c5a604e3b0b274fe39c5e24262745314731baa19c7f2c19b045cadac1dfa

Observation 8ae7bd9e-455d-4181-9297-823f6433ac9f · inbound

Neural Network Optimization Reimagined: Decoupled Techniques for Scratch and Fine-Tuning cites this paper.

Neural Network Optimization Reimagined: Decoupled Techniques for Scratch and Fine-Tuning On the Variance of the Adaptive Learning Rate and Beyond

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:22.069640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T03:26:09.751493Z digest=sha256:267f422931321c131f00e61005263ac0214df521d48d19b4362703f23ea1e1a1

Observation 59fe7434-96b1-48a5-8a05-df485960b583 · inbound

AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments cites this paper.

AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments On the Variance of the Adaptive Learning Rate and Beyond

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:31:07.615287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-09T19:50:50.653184Z digest=sha256:7cd63f614b7ae0533c66a776748bbc5fc3a7410fbe2b655b4f0125c66f12ca99

Observation e984c67a-2fbf-4c5a-a644-babd64eadc37 · inbound

Deep neural networks with Fisher vector encoding for medical image classification cites this paper.

Deep neural networks with Fisher vector encoding for medical image classification On the Variance of the Adaptive Learning Rate and Beyond

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:00:58.899036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:22:00.015928Z digest=sha256:17c5f12d217f1e1a391b5b75a46249c915a6fab803372b0ac7cece9a45176ffa

Observation 4e6ec95d-267c-4f21-8d9a-e2e9a6cd9a9c · inbound

Anon: Extrapolating Adaptivity Beyond SGD and Adam cites this paper.

Anon: Extrapolating Adaptivity Beyond SGD and Adam On the Variance of the Adaptive Learning Rate and Beyond

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:09.614347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-09T16:25:36.073417Z digest=sha256:83f5ae1d299c606d7e30b6e9ad41415a744ba0744c329ea9dd212257ae265303

Observation 00736419-3d20-4b7a-a440-5d677ac61205 · inbound

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior cites this paper.

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior On the Variance of the Adaptive Learning Rate and Beyond

Reference 297

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:16:06.693844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-08T17:47:09.591001Z digest=sha256:554eca4d8e1e9881bc3609d20dcd834e2949a4dcfcbdf4cae200733d09780db5

Observation 2918b059-9733-48f5-a3cd-47bd6101ba2e · inbound

Neural Network-Based Virtual Wheel-Speed Sensor for Enhanced Low-Velocity State Estimation cites this paper.

Neural Network-Based Virtual Wheel-Speed Sensor for Enhanced Low-Velocity State Estimation On the Variance of the Adaptive Learning Rate and Beyond

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:37:15.837211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T04:32:39.053186Z digest=sha256:24a7640fbe49575279a09b698743bfd2a341059139f0e17420c0cdd32392166c

Observation 59e925a4-f420-47fe-a40e-8169504af7f9 · inbound

Scalable Reinforcement Learning via Adaptive Batch Scaling cites this paper.

Scalable Reinforcement Learning via Adaptive Batch Scaling On the Variance of the Adaptive Learning Rate and Beyond

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:24:27.799007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T00:21:15.933174Z digest=sha256:8391757288e18526a51c60f01dbabb97b1e6a43ac4ab7789e1041157fd08131e

Observation 200df501-70ac-4759-a340-d81bee41b7fe · inbound

Scalable Reinforcement Learning via Adaptive Batch Scaling cites this paper.

Scalable Reinforcement Learning via Adaptive Batch Scaling On the Variance of the Adaptive Learning Rate and Beyond

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:57.664358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T17:17:59.698127Z digest=sha256:53952c026e4ccf49c50c1b576d121336302228847a8250bf79392bc21ca597ec

Observation 431879f9-244d-4bb3-bfd4-38146499c5ea · inbound

Why Muon Outperforms Adam: A Curvature Perspective cites this paper.

Why Muon Outperforms Adam: A Curvature Perspective On the Variance of the Adaptive Learning Rate and Beyond

Reference 169

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:06:44.922851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T07:04:21.012269Z digest=sha256:a09032a5e93bad01ec744394e23f104bb6e13413db11b6bc811427565a24a104

Observation a73ac89d-0f45-4267-a7d4-3024d2c63041 · inbound

Muon Learns More Robust and Transferable Features than Adam cites this paper.

Muon Learns More Robust and Transferable Features than Adam On the Variance of the Adaptive Learning Rate and Beyond

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:30.094776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T17:08:30.717799Z digest=sha256:9249183a177184411a3a2048bd583018d9ed27cac2f7b6e23a923c5fc92dcdd0

Observation 5b67d29d-83ba-4c55-af17-ac27193e4821 · inbound

Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning cites this paper.

Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning On the Variance of the Adaptive Learning Rate and Beyond

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:38:41.024926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T05:03:06.472027Z digest=sha256:1caccc59b2dca7021067ffb13427b0d9460f7e3255aff03679b9f155066eee74

Observation 8e1790d6-d0bb-4204-a2d9-f13af6cbcd0c · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors On the Variance of the Adaptive Learning Rate and Beyond

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:30:07.698743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:bffb1eb48713198131da6073db7f0bd17f4ecceb1272a073ca1d789dcc1c4dfc

Observation bf816f9e-848a-4f3d-a30e-92ee06128d26 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors On the Variance of the Adaptive Learning Rate and Beyond

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:08.871888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:08.871888Z digest=sha256:b207f1ba35aaa1c8e08907c50809b1ee19d5e92140d1a32c79e6d87c8554fdd0

Observation 8aa828d9-8e2a-41c4-81a9-a577794beac8 · inbound

Confidence-feedback-weighted graph matching network: online-offline laser-induced damage site matching under complex interference cites this paper.

Confidence-feedback-weighted graph matching network: online-offline laser-induced damage site matching under complex interference On the Variance of the Adaptive Learning Rate and Beyond

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:14:25.179388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T08:09:50.192952Z digest=sha256:6453a6e5e0c2687ac02daca7131c71afc14b5c64f3d792a95a0c48ef774c5ed5

Observation 3909f6fb-bb73-4b1a-8d41-6a0245ba64c8 · inbound

Point spread function wavefront recovery from in-focus stellar observations cites this paper.

Point spread function wavefront recovery from in-focus stellar observations On the Variance of the Adaptive Learning Rate and Beyond

Reference 187

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:36:40.238509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-07-02T05:34:06.325999Z digest=sha256:9a22053eb7fa2a18b18b544fd87d626c0d49103bf3f84df289c5169c4ba445ae

Observation 9acfc265-1011-43dd-a2a7-b148e9610e63 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers On the Variance of the Adaptive Learning Rate and Beyond

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:489cb4bb054531ea47658205b8a44352ca65ba99577e803297d13c9c5e315a45

Observation e6b32ccb-1afa-420f-801f-d94cfd168b29 · inbound

Unified convergence analysis for gradient descent optimization methods in the training of deep neural networks cites this paper.

Unified convergence analysis for gradient descent optimization methods in the training of deep neural networks On the Variance of the Adaptive Learning Rate and Beyond

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T20:46:05.467029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:46:05.467029Z digest=sha256:daa1ea8f957523aeaffd068491aef0fdbd350beb246acfa775e369346529d268

Observation f9e1f0fb-c62e-4366-b4a9-d290f097a92e · inbound

A machine-learned probability distribution in the phase space of turbulent channel flow for synthetic turbulence and flow reconstruction cites this paper.

A machine-learned probability distribution in the phase space of turbulent channel flow for synthetic turbulence and flow reconstruction On the Variance of the Adaptive Learning Rate and Beyond

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T16:19:27.144952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:19:27.144952Z digest=sha256:170e7075dcd52b4ee145e6536422f0ed77dda713e10092e8ee9102d32047a849

Observation a8553a71-a379-47ab-a86f-bd17140d0fb5 · inbound

GCR Spectra Reconstructed with Neutron Monitor Yield Function and Artificial Neural Networks: Comparison of Two Methods cites this paper.

GCR Spectra Reconstructed with Neutron Monitor Yield Function and Artificial Neural Networks: Comparison of Two Methods On the Variance of the Adaptive Learning Rate and Beyond

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T08:51:11.204445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:51:11.204445Z digest=sha256:245bd4e59b18dfcb79eed0c839d60dcb71e1b0280548e2aa2a3ae47c4051a635

Observation 4aa982dc-4ad3-4a8a-aace-348d101318c1 · inbound

Payne4GAIN: NLTE Corrections for Red Giants in Milky Way Mapper using H-Band Neural Network Emulators cites this paper.

Payne4GAIN: NLTE Corrections for Red Giants in Milky Way Mapper using H-Band Neural Network Emulators On the Variance of the Adaptive Learning Rate and Beyond

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-01T04:40:18.200294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:40:18.200294Z digest=sha256:6dbbdc4baaee3219f2a3f991fd49efcf86ef9bd26115329303c4ca7ebc6c6696

Observation 320d20e7-082b-4fe6-a440-833a7c068a46 · inbound

Amortized Moment Matching for Visual Generation cites this paper.

Amortized Moment Matching for Visual Generation On the Variance of the Adaptive Learning Rate and Beyond

Reference 104

Resolution
unresolved
no resolver link, observed 2026-07-30T18:58:28.223966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:58:28.223966Z digest=sha256:780921fbb69e97a4e936b91f8d2988b9dfe2e0476dc6ecbf7253bfa57a4aab5f

Observation 30ab29ce-e1d7-4d81-a48a-c360c4b99e42 · inbound

HealthCAT: An Interpretable Encoder-only Transformer Framework for Health Indicator Prediction and Temporal Interpretation of Wearable Sensor Data cites this paper.

HealthCAT: An Interpretable Encoder-only Transformer Framework for Health Indicator Prediction and Temporal Interpretation of Wearable Sensor Data On the Variance of the Adaptive Learning Rate and Beyond

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T04:27:10.763773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:27:10.763773Z digest=sha256:e8689c6065ad415cab44b5172eab82ad1653720a1c8e9cf083d331e12429013a