Pith. sign in

Paper Citation Record · LEDGER

Gradient Methods with Online Scaling Part I. Theoretical Foundations

As of 7 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 2 inbound Pith citation observations for arXiv:2505.23081.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23081 v2

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:10:18.346914Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:35:24.103094Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T18:46:45.040439Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact1
  • verified fuzzy49
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 582a86a0-fd6c-4ea8-b92a-5d0679058bde · outbound

This paper cites Disentangling Adaptive Gradient Methods from Learning Rates.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Disentangling Adaptive Gradient Methods from Learning Rates

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:06.894473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:06.894473Z digest=sha256:9eea4434a6c351ef8efbf40a62db40245307eb82f28efd1ff8527ad5f9f7b8a8

Observation 211f7c72-ff18-4c09-98d4-96bdf3d248f3 · outbound

This paper cites Parameter adaptation in stochastic optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Parameter adaptation in stochastic optimization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:33.100819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:06.990413Z digest=sha256:cb5805c27445fe1ee08943c582c9ed6f985a5f019863b285f02fb6e3a9e06ed6

Observation 599fe075-96bd-486d-81b3-a6cdd78f65b6 · outbound

This paper cites Acceleration by stepsize hedging: Silver stepsize schedule for smooth convex optimization.Mathematical Programming, pages 1–14, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Acceleration by stepsize hedging: Silver stepsize schedule for smooth convex optimization.Mathematical Programming, pages 1–14, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.891809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:07.143282Z digest=sha256:0d649aba295367941c3e6752b09ce44a939af8c8c2a25d2b8cfbbefa4a24066e

Observation 97a905fd-c4f7-41c4-8a17-67139bb6e03d · outbound

This paper cites Acceleration by stepsize hedging: Multi-step descent and the silver stepsize schedule.Journal of the ACM, 72(2):1–38, 2025.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Acceleration by stepsize hedging: Multi-step descent and the silver stepsize schedule.Journal of the ACM, 72(2):1–38, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.749701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:07.264191Z digest=sha256:d4d2f7d1cff38234014eb9631b9251a215703e3a09b3264fb12894c51ea674af

Observation c5b313fa-e5df-4908-9b14-6441dfa666ba · outbound

This paper cites Practical large-scale linear programming using primal-dual hybrid gradient.Advances in Neural Information Processing Systems, 34:20243–20257, 2021.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Practical large-scale linear programming using primal-dual hybrid gradient.Advances in Neural Information Processing Systems, 34:20243–20257, 2021

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.470056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:07.400967Z digest=sha256:8e8b9be647d33948607c13e6dd66b3ab0dab5db4b9063ccc6368b9935c19c3e4

Observation 685e8275-53dd-47c0-af85-624b89fa7cc4 · outbound

This paper cites Two-point step size gradient methods.IMA journal of numerical analysis, 8(1):141–148, 1988.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Two-point step size gradient methods.IMA journal of numerical analysis, 8(1):141–148, 1988

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.237528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:07.573874Z digest=sha256:453d8b07ae54fb3bb2e242b4b0356425814d10cfff4720dbbff2c4b518217b0a

Observation e00dd785-aae5-44e8-8e37-adb003f682b6 · outbound

This paper cites Online learning rate adaptation with hypergradient descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Online learning rate adaptation with hypergradient descent

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.987009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:07.744515Z digest=sha256:96c33e325c1033045eb822912dfd1f760a7090674e25b45582aa6d0715b9ac4c

Observation e0b9c6f3-bd1d-4b8a-9d02-fa5d750a7e55 · outbound

This paper cites Gradient descent: The ultimate optimizer.Advances in Neural Information Processing Systems, 35:8214–8225, 2022.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Gradient descent: The ultimate optimizer.Advances in Neural Information Processing Systems, 35:8214–8225, 2022

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.773172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:07.900330Z digest=sha256:94cf80f10dcacd39a8ec43617f4a823b1ac2ce4abbd1530c40cb3824428ad623

Observation 08be7905-0102-40c6-a315-5309d64fbfbe · outbound

This paper cites Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:07.974495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:07.974495Z digest=sha256:acedfc2c4e81d05316be0ed00ec01055344163e06b73805b2ab4bf3649848a67

Observation cc825b92-a85e-441e-ac2f-0b5ca3016d0c · outbound

This paper cites Non-monotonebehavioroftheheavyballmethod.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Non-monotonebehavioroftheheavyballmethod

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.467420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:08.166282Z digest=sha256:2a8a15c643f84ee8c675150cacd834e80f0290ea57ce73d50fd4aec1168721bc

Observation 698dab72-3f8c-4753-9886-28607abc2b74 · outbound

This paper cites An enhanced alternating direction method of multipliers-based interior point method for linear and conic optimization.INFORMS Journal on Computing, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations An enhanced alternating direction method of multipliers-based interior point method for linear and conic optimization.INFORMS Journal on Computing, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.295597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:08.339118Z digest=sha256:fe76fc6b6390de38ec406982cee3401ddb734f4c85f881c0ba8770cafa17f755

Observation ad823aab-24c0-474f-b3ef-d0b9e2921766 · outbound

This paper cites Uniformly Optimal and Parameter-free First-order Methods for Convex and Function-constrained Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Uniformly Optimal and Parameter-free First-order Methods for Convex and Function-constrained Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:08.488432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:08.488432Z digest=sha256:c3c1f1d38eca1877041060fc66c19768bdc30809f17019eac4ab8396545d4cac

Observation 700875c0-508d-452e-81d6-ddfa3eec3544 · outbound

This paper cites Adaptive subgradient methods for online learning and stochas- tic optimization.Journal of machine learning research, 12(7), 2011.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive subgradient methods for online learning and stochas- tic optimization.Journal of machine learning research, 12(7), 2011

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.076594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:08.656273Z digest=sha256:d5db6d739d4c980d60341c996bb64314a8cca9fc2c82fccbea1a410d3128adb9

Observation 36c42a4e-7125-48b8-9783-19a765fb6ca4 · outbound

This paper cites John Wiley & Sons, 2000.

Gradient Methods with Online Scaling Part I. Theoretical Foundations John Wiley & Sons, 2000

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.834841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:08.809382Z digest=sha256:b878c42e6e7e8a972a0b1402de6d447bad01fc5fe10601047f734a33c2676195

Observation 3ea42a4d-5fdf-4220-b05e-a03902fc660f · outbound

This paper cites Gradient Methods with Online Scaling.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Gradient Methods with Online Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:09.008540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:09.008540Z digest=sha256:adc7bdf302fb4f4a55856769efec9521a0593508e2c5ac9cf3fa34f66839188d

Observation e3cc866c-4d62-4b8a-b00e-b6e99c6b6d3f · outbound

This paper cites Scalable Approximate Optimal Diagonal Preconditioning.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Scalable Approximate Optimal Diagonal Preconditioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:09.165212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:09.165212Z digest=sha256:8d44eafb9badbdbbc76bf1c34fcd1549b8ed356512ed6d6faa4f8ed802bc16b4

Observation d7242c41-90e9-4302-98c0-6a9148e4e5ac · outbound

This paper cites Clarabel: An interior-point solver for conic programs with quadratic objectives.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Clarabel: An interior-point solver for conic programs with quadratic objectives

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:09.347627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:09.347627Z digest=sha256:feef1de84d249dc52fb3a5958b84015848852dece6bf5b84c0cb3c290b60f749

Observation 34e87645-c7c9-4b16-a7d6-cb95dde4dbfd · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Shampoo: Preconditioned stochastic tensor optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.552063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:09.577230Z digest=sha256:7a36a50e2c7d31c854e5236c504ff693ccfad5a82e04cc9b9e8dda4ef2d9f6af

Observation e82d314b-437f-4135-a23d-8f672afdbcf5 · outbound

This paper cites Introduction to online convex optimization.Foundations and Trends®in Optimization, 2(3-4):157–325, 2016.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Introduction to online convex optimization.Foundations and Trends®in Optimization, 2(3-4):157–325, 2016

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.256654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:09.773950Z digest=sha256:36bb36320fd86aa545aa304555d07457f5a6f883a3ad689d03e4e34a6291df75

Observation 830e8eff-9523-47ee-986c-3c8d79cc32c5 · outbound

This paper cites Revisiting the Polyak step size.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Revisiting the Polyak step size

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:09.962195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:09.962195Z digest=sha256:ec2493219e9657dd49b17708a9c4a07f3b5564c4cab2982c0ca879a69b92b342

Observation da8727a2-61f7-4d95-a314-16abdb8e8882 · outbound

This paper cites Adaptive online gradient descent.Advances in neural information processing systems, 20, 2007.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive online gradient descent.Advances in neural information processing systems, 20, 2007

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.949825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:10.093964Z digest=sha256:13989dfb0724ffe72f4edee19babd78c1e1a3e10502a50dbfd03ca792ec55d1e

Observation c3541b60-4e1e-4b6e-bb89-7d2e4984a221 · outbound

This paper cites Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.627509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:10.226107Z digest=sha256:166f288e4b9b5ba113b9d2b44a0a9e5f1a3333e57da9611535e55d6da715c097

Observation e6011afc-fe50-4103-97f7-d1c2d57e712e · outbound

This paper cites Restarted Primal-Dual Hybrid Conjugate Gradient Method for Large-Scale Quadratic Programming.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Restarted Primal-Dual Hybrid Conjugate Gradient Method for Large-Scale Quadratic Programming

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:10.367209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:10.367209Z digest=sha256:50bcfa2a4d99a89789622c02fa4b5d82e03ed88c697ae8bfd8344008892a9ed2

Observation 242d1643-d6d3-4f2e-a90a-ef7b5f1ef31c · outbound

This paper cites Increased rates of convergence through learning rate adaptation.Neural networks, 1(4):295–307, 1988.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Increased rates of convergence through learning rate adaptation.Neural networks, 1(4):295–307, 1988

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.321487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:10.441402Z digest=sha256:0b25d5314e62b08d9ea35381548608b087bef971948d1d81f754c8e9fe8ac048

Observation d1204d1f-fbf2-41b4-ab6e-d79bf1ea9980 · outbound

This paper cites Unconstrained online learning with unbounded losses.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Unconstrained online learning with unbounded losses

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.152637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:10.614661Z digest=sha256:5a10bde076e255f9f4c0b5b63d9af83bddb519883fddb0ceb47c33270b64783e

Observation 21918e13-066d-48bb-9e7c-8b7f63fc7b0b · outbound

This paper cites Online learning guided curvature approximation: A quasi-newton method with global non-asymptotic superlinear convergence.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Online learning guided curvature approximation: A quasi-newton method with global non-asymptotic superlinear convergence

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.917531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:10.776182Z digest=sha256:45885cca15999b4a6dec92df7f4bf14a806d0c50e0cc217d2ded5bdd381bebe9

Observation e4d11c3a-afaf-4ef0-bdf8-fd298e008b85 · outbound

This paper cites Online Learning Guided Quasi-Newton Methods with Global Non-Asymptotic Convergence.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Online Learning Guided Quasi-Newton Methods with Global Non-Asymptotic Convergence

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:10:19.028351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:10.986354Z digest=sha256:8d7e0b692115594f2cc06b32d9ad2bcf2a0f61df30e2b15a202b10872419664a

Observation 1a392729-b1cc-4389-b592-503af395f9af · outbound

This paper cites Adaptive hierarchical hyper-gradient descent.International Journal of Machine Learning and Cybernetics, 13(12):3785–3805, 2022.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive hierarchical hyper-gradient descent.International Journal of Machine Learning and Cybernetics, 13(12):3785–3805, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.781008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:11.149792Z digest=sha256:99e03064fc9e07f8e26af5174acdb8c91289ff1b8dc2ddc857983cc9d4b6633d

Observation b509697c-f221-4fff-983a-ed592bab0dd8 · outbound

This paper cites Non-asymptotic Global Convergence Analysis of BFGS with the Armijo-Wolfe Line Search.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Non-asymptotic Global Convergence Analysis of BFGS with the Armijo-Wolfe Line Search

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:11.348183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:11.348183Z digest=sha256:8caee129a7a03df7dd1bf7e532eeaf936cd22842dd6e4eb82333461dd798e077

Observation e0306c6a-5ece-4e3e-996c-773dbf121f8e · outbound

This paper cites Non-asymptotic Global Convergence Rates of BFGS with Exact Line Search.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Non-asymptotic Global Convergence Rates of BFGS with Exact Line Search

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:11.500660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:11.500660Z digest=sha256:0d09873c6f2ccca0aa69295b3b382b962c205d10721b75d35b8a7c34cfd4e0ca

Observation 2b6fc337-097e-43af-b5ff-a957bed4e696 · outbound

This paper cites Non-asymptotic superlinear convergence of standard quasi-newton methods.Mathematical Programming, 200(1):425–473, 2023.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Non-asymptotic superlinear convergence of standard quasi-newton methods.Mathematical Programming, 200(1):425–473, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.586015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:11.664806Z digest=sha256:d759460b397a7c24c8bf1101a3b47ef00d8ed4a5581d9ee83dd5ca7a4ad941c3

Observation fd73d695-ef08-4ce8-833b-5df61ebae305 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adam: A Method for Stochastic Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:11.796601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:11.796601Z digest=sha256:e55a81f1cbc893a8b1884691dedd70ea4a851349e1a27c6e834328fb68712585

Observation cda76596-1567-42f0-935a-41cdfc408c9f · outbound

This paper cites Searching for optimal per-coordinate step-sizes with multidimensional backtracking.Advances in Neural Information Processing Systems, 36, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Searching for optimal per-coordinate step-sizes with multidimensional backtracking.Advances in Neural Information Processing Systems, 36, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.347020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:11.949488Z digest=sha256:0927b37a1991e4b0dde309966b5d35a3d1db971ac97478046c4d908431c9c1be

Observation 2f51d06b-a0fb-4662-b0ab-143ebea5d09d · outbound

This paper cites Optimal and parameter-free gradient minimization methods for convex and nonconvex optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Optimal and parameter-free gradient minimization methods for convex and nonconvex optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.138443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:12.138443Z digest=sha256:e1ac0b9b4218af6b85ce2ba5f70a99cf1eff71f23e307d92737f706a3b5a1748

Observation 5f05fe23-b761-4118-9756-402baffd1343 · outbound

This paper cites A simple uniformly optimal method without line search for convex optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations A simple uniformly optimal method without line search for convex optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.329475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:12.329475Z digest=sha256:c5bad26a6d425b389fbf97ee200f8d38dd462c520e95bb47b02528ebe8e5107f

Observation 22afb333-41fb-4049-a4ec-f93bc7ff20bc · outbound

This paper cites A second look at exponential and cosine step sizes: Simplicity, adaptivity, and performance.

Gradient Methods with Online Scaling Part I. Theoretical Foundations A second look at exponential and cosine step sizes: Simplicity, adaptivity, and performance

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.024072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:12.491677Z digest=sha256:1eeeed4bf2513f9d3dfd3a53c5d713ee5d056edd742774234db597beff5547f9

Observation 58b31885-133b-402d-846b-f502f5e517c4 · outbound

This paper cites An admm-based interior-point method for large-scale linear programming.Optimization Methods and Software, 36(2-3):389–424, 2021.

Gradient Methods with Online Scaling Part I. Theoretical Foundations An admm-based interior-point method for large-scale linear programming.Optimization Methods and Software, 36(2-3):389–424, 2021

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:27.490108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:12.806062Z digest=sha256:2e2b7748b0d9064e39e2e0ca3f8c650e2cf33e89e0e5f619aa1bb2b558559440

Observation 10ffd99f-c795-4c93-9a25-2bc8a81907fd · outbound

This paper cites Pdcs: A primal-dual large-scale conic programming solver with gpu enhancements.arXiv preprint arXiv:2505.00311, 2025.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Pdcs: A primal-dual large-scale conic programming solver with gpu enhancements.arXiv preprint arXiv:2505.00311, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.017954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:13.017954Z digest=sha256:3ace142e963b0aaf8519e4d9c5d3a672fea1282bba84dc757b950effdfd2f517

Observation 13043e44-b86d-454d-b3e1-a3d1cf0a2a7e · outbound

This paper cites cuPDLP.jl: A GPU Implementation of Restarted Primal-Dual Hybrid Gradient for Linear Programming in Julia.

Gradient Methods with Online Scaling Part I. Theoretical Foundations cuPDLP.jl: A GPU Implementation of Restarted Primal-Dual Hybrid Gradient for Linear Programming in Julia

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.221669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:13.221669Z digest=sha256:0a463143abce6abb90739be57944484cfbbe8132e600748d986c1429d63ec86e

Observation 86ae7d01-ff31-4cda-8697-9c946a456fef · outbound

This paper cites cuPDLP-C: A Strengthened Implementation of cuPDLP for Linear Programming by C language.

Gradient Methods with Online Scaling Part I. Theoretical Foundations cuPDLP-C: A Strengthened Implementation of cuPDLP for Linear Programming by C language

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.373008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:13.373008Z digest=sha256:9996c1292e52ae9e6391ee5c8902f13c35e680d5b20a9285d3b81b35529083e9

Observation 8cb11ec2-df2b-4407-8d46-3c97fdc6e0e8 · outbound

This paper cites Tuning-freestep-size adaptation.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Tuning-freestep-size adaptation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:27.135518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:13.492046Z digest=sha256:21efa9605e177fbe4b27382412c4ce86eafa80b5ccbae2ddf4e1a772203de83d

Observation d537402d-9002-461e-9a01-fd926b69f4ad · outbound

This paper cites Adaptive gradient descent without descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive gradient descent without descent

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:26.619015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:13.621772Z digest=sha256:1a56e3eaddca9278c539cf15a76a10bd2966aae8c52110d9966eee74e388b64f

Observation 690cc81a-7386-4bb0-b398-2d4fd274d894 · outbound

This paper cites Adaptive proximal gradient method for convex optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive proximal gradient method for convex optimization

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:26.349459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:13.774680Z digest=sha256:50cbd0813e074bcd18d542717588eb3602b17b99c6088515ebea5485662e656e

Observation 405bcadf-473c-4c53-892d-48321954815b · outbound

This paper cites Adaptive Bound Optimization for Online Convex Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive Bound Optimization for Online Convex Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.932005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:13.932005Z digest=sha256:1da1f312c60b04e85770867b8585339dda61b47b2b86bb4138c950f0db9e3f36

Observation 3eea46ed-ab74-43c5-9a9f-6f1a7e1fb9f5 · outbound

This paper cites an unresolved cited work.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:26.037314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:14.074627Z digest=sha256:df0e4b512991e8bd7c6130309686bf303c5e0984a88d12803050dc71975154a1

Observation cda48efa-20d9-41cb-84ca-4dcc1c763e04 · outbound

This paper cites Linear convergence of first order methods for non-strongly convex optimization.Mathematical Programming, 175:69–107, 2019.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Linear convergence of first order methods for non-strongly convex optimization.Mathematical Programming, 175:69–107, 2019

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:25.725934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:14.243941Z digest=sha256:0b746855e198f045f79ac9980afb26d106570ff473d8d73bd195bc2e8dbd808f

Observation a97f9952-f43a-4b26-ab3b-4bad7f9a6d42 · outbound

This paper cites A method for solving the convex programming problem with convergence rate o (1/k2).

Gradient Methods with Online Scaling Part I. Theoretical Foundations A method for solving the convex programming problem with convergence rate o (1/k2)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:25.422538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:14.377838Z digest=sha256:8ce68dbc26994772d0b4f2b9b0206c7fa52fc9b3854ad4ec753c82858704e838

Observation af122e1a-fd8f-45d9-8581-b8e70602774c · outbound

This paper cites Springer Science & Business Media, 2013.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Springer Science & Business Media, 2013

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:25.146817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:14.540097Z digest=sha256:bcbbe08222cb82e67d199e3d4d08902024b6ed9311b9c97f1765e82c18eb31e6

Observation b6fd424a-635e-4a24-b603-d5275a7a25a8 · outbound

This paper cites Springer, 1999.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Springer, 1999

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:24.847880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:14.688071Z digest=sha256:77337af8e1fdb3605aa6427c137982847b6ccf62a431a16135708228b658251c

Observation 073b82e1-a03a-4f53-9384-b6868aad5493 · outbound

This paper cites Conic optimization via operator splitting and homogeneous self-dual embedding.Journal of Optimization Theory and Applications, 169:1042–1068,.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Conic optimization via operator splitting and homogeneous self-dual embedding.Journal of Optimization Theory and Applications, 169:1042–1068,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:24.596628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:14.887592Z digest=sha256:28ef43020cffe96381493b0648ea28467e5c4c07bb81a2abf6835b568440c91e

Observation 8fad7c9b-8518-43ef-9b02-cd5c18cbbc9c · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Online Learning: A Modern Introduction Using Convex Optimization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:15.037964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:15.037964Z digest=sha256:3482980f9be6c4d01fb5132524249a151e8cdd552fe9ea330affb6ac637f7574

Observation 11c62399-0f76-4282-9aeb-3c8a61671e7f · outbound

This paper cites Coin betting and parameter-free online learning.Advances in Neural Information Processing Systems, 29, 2016.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Coin betting and parameter-free online learning.Advances in Neural Information Processing Systems, 29, 2016

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:24.353932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:15.203695Z digest=sha256:ea5ecb2458e0aebf22eb43371a5979ac58b172580c243fbed27a2255ae7a2bf8

Observation a2d47ced-c099-4894-b9ba-99dab4fdbcd4 · outbound

This paper cites MADA: Meta-adaptive optimizers through hyper-gradient descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations MADA: Meta-adaptive optimizers through hyper-gradient descent

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:23.997951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:15.404845Z digest=sha256:685cbfaa0d06fdd5e04c729b427ad279b56cea738841a61d8827d3ac285f1d02

Observation 7d629dba-c740-40a2-b598-0f19e07ecefa · outbound

This paper cites Introduction to optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Introduction to optimization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:23.506836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:15.614007Z digest=sha256:f9bf9a345ce9056ed4e8b7f95a605ad8c8145dc8a3056568fa8210796eb8f914

Observation 0473bbe0-f075-45ac-a88c-d90e57b30df4 · outbound

This paper cites Optimal diagonal precondi- tioning.Operations Research, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Optimal diagonal precondi- tioning.Operations Research, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:23.187929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:15.729222Z digest=sha256:a19a846f873709e0037d3b32d09fba7c8574b3bea5a0eb217d3ac5c8d44380e7

Observation e227455b-df14-4f60-8c1d-8ac694052ba0 · outbound

This paper cites Lecture notes on online learning draft, 2009.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Lecture notes on online learning draft, 2009

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:22.888468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:15.885167Z digest=sha256:b2e22fef6407d1d9d16370ea0defc748b40ffde1db5b0f59a87191a3dfbf08a9

Observation deef3e93-5088-46c0-b9e4-f524ea99ca99 · outbound

This paper cites On the Convergence of Adam and Beyond.

Gradient Methods with Online Scaling Part I. Theoretical Foundations On the Convergence of Adam and Beyond

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:16.063965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:16.063965Z digest=sha256:de835b27a1d62bc8badf9adfa8142f2a73a8813781fdee849f0dfbbd8cd69f41

Observation 07821767-069b-4dad-ada2-b1eb6a819236 · outbound

This paper cites Greedy quasi-newton methods with explicit superlinear conver- gence.SIAM Journal on Optimization, 31(1):785–811, 2021.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Greedy quasi-newton methods with explicit superlinear conver- gence.SIAM Journal on Optimization, 31(1):785–811, 2021

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:22.657859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:16.251549Z digest=sha256:7e3bd52d19cc779a78247c2cfde5607a6816b45b7a050a8fa4f9d7c94adbb6d4

Observation fcd11103-fce6-4d6e-9177-588f55060dc1 · outbound

This paper cites New results on superlinear convergence of classical quasi-newton methods.Journal of optimization theory and applications, 188:744–769, 2021.

Gradient Methods with Online Scaling Part I. Theoretical Foundations New results on superlinear convergence of classical quasi-newton methods.Journal of optimization theory and applications, 188:744–769, 2021

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:22.492178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:16.373687Z digest=sha256:9e367abd5061168c9c330bf526b0f67c61c7eccaf89ce0dfac5faa0c8f3433d5

Observation 50bcfa29-c047-477c-aa49-627546b67153 · outbound

This paper cites Ratesofsuperlinearconvergenceforclassicalquasi-newtonmethods.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Ratesofsuperlinearconvergenceforclassicalquasi-newtonmethods

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:22.203383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:16.529064Z digest=sha256:370296801e57030b9ef43e726e285c62fd1d44ac4ed507103b214d0152f11391

Observation 28a70dfa-4977-4735-8bce-83f2423e0fd4 · outbound

This paper cites Convergence analysis of an adaptive method of gradient descent.University of Oxford, Oxford, M.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Convergence analysis of an adaptive method of gradient descent.University of Oxford, Oxford, M

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:21.869002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:16.744234Z digest=sha256:03f3b92d0da4ba00aaef378974a281042c63fa89db4194e26dc600b348a0774c

Observation 86f09eaa-efb1-48da-a833-c69adcb76bb9 · outbound

This paper cites Local gain adaptation in stochastic gradient descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Local gain adaptation in stochastic gradient descent

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:21.591671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:16.855472Z digest=sha256:1df10c13e21989f85b45b44767baa19c72cf7e48c75d449c38b63c2f19b212ce

Observation 6a17c6f7-b9c3-4b92-853e-850ce34d27e5 · outbound

This paper cites Adapting bias by gradient descent: An incremental version of delta-bar-delta.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adapting bias by gradient descent: An incremental version of delta-bar-delta

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:21.248493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:17.038377Z digest=sha256:4d7c63c3af76c16470746ccfa2611f30d4cd57d3de250098931f102f0c1f0016

Observation c4be4a74-e36d-4bec-be13-b0d9120ff377 · outbound

This paper cites No-regret dynamics in the fenchel game: A unified framework for algorithmic convex optimization.Mathematical Programming, 205(1):203–268, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations No-regret dynamics in the fenchel game: A unified framework for algorithmic convex optimization.Mathematical Programming, 205(1):203–268, 2024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:20.916902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:17.226881Z digest=sha256:43f4fc31dba4b461044f8e9c00346e206ed9b36713936d9cf5e2b6e60693ef33

Observation 2b5126e1-b40e-4efb-b46c-ee6866f94cb9 · outbound

This paper cites On the convergence of stochastic gradient descent with bandwidth-based step size.Journal of Machine Learning Research, 24(48):1–49, 2023.

Gradient Methods with Online Scaling Part I. Theoretical Foundations On the convergence of stochastic gradient descent with bandwidth-based step size.Journal of Machine Learning Research, 24(48):1–49, 2023

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:20.605338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:17.431702Z digest=sha256:0e89e1ff129489fdb975e0bc2ae6263b38b78e8203377b45d34fcf8c86483b8f

Observation 81a47c91-72de-40b5-8262-7b20494b7c7e · outbound

This paper cites The Role of Level-Set Geometry on the Performance of PDHG for Conic Linear Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations The Role of Level-Set Geometry on the Performance of PDHG for Conic Linear Optimization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:17.588000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:17.588000Z digest=sha256:40275e42f7143e54b9ea99ea0180831d8c52eccff351330b34db5d053796127d

Observation 09a18a02-c2b9-413e-a58c-08ea1777735d · outbound

This paper cites Adaptive powerball stochastic conjugate gradient for large-scale learning.IEEE Transactions on Big Data, 9(6):1598–1606, 2023.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive powerball stochastic conjugate gradient for large-scale learning.IEEE Transactions on Big Data, 9(6):1598–1606, 2023

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:20.317776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:17.719711Z digest=sha256:2d815c9194812caf0231ed2053008da242b2a9752cfeb2bc33c44b759bf823e8

Observation ccd1611f-fda8-426f-b2c8-293a42ac63b6 · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adam-mini: Use Fewer Learning Rates To Gain More

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:17.855272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:17.855272Z digest=sha256:057a9b44ca954530b493c4eecadc0e007ecd1472262cb9a70a92edf3ada983fa

Observation acf4447c-13d9-48d8-bdf1-9b7663236653 · outbound

This paper cites Algorithm 778: L-bfgs-b: Fortran sub- routines for large-scale bound-constrained optimization.ACM Transactions on mathematical software (TOMS), 23(4):550–560, 1997.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Algorithm 778: L-bfgs-b: Fortran sub- routines for large-scale bound-constrained optimization.ACM Transactions on mathematical software (TOMS), 23(4):550–560, 1997

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:20.001214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:18.009345Z digest=sha256:4f23a40e62c84618c3d25424f70896fc084d65cddab0c5d885eab60090ab9f8c

Observation 1caf9fe7-598d-494a-a3d9-884a50a16de7 · outbound

This paper cites Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.Advances in neural information processing systems, 33:18795–18806, 2020.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.Advances in neural information processing systems, 33:18795–18806, 2020

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:19.729131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:18.193092Z digest=sha256:dae7e48991b100068d8901ab9345230d455197bfeec4885a99a6a05865f4d2bf

Observation f8f4b0a7-8547-4732-9036-7ad4c23feadd · outbound

This paper cites Surrogate losses for online learning of stepsizes in stochastic non-convex optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Surrogate losses for online learning of stepsizes in stochastic non-convex optimization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:19.409231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:18.346914Z digest=sha256:2a81f56db6777d9a771f699eb9db77ab8ff1ec483014455064fb6fb045af1423

Observation 9400cacb-1509-4471-be06-d6ca5c05c077 · outbound

This paper cites (cited on 17).

Gradient Methods with Online Scaling Part I. Theoretical Foundations (cited on 17)

Reference 6564

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:27.789881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:10:12.619868Z digest=sha256:015e9528b42b8e389bab5f70ec078dc086af26d670a0075c4178edb961a4e618

Pith citing papers

Observation ff55bcce-22c1-47d3-8546-124f443adfee · inbound

Enhanced PDHG for Linear Programming with Online Preconditioning cites this paper.

Enhanced PDHG for Linear Programming with Online Preconditioning Gradient Methods with Online Scaling Part I. Theoretical Foundations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:24.103094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:24.103094Z digest=sha256:f5d8c6a8bbff90125dd77cd5f8762eb3c595d13c26745966dff6ce31bdfcf68f

Observation d0b4f173-62f7-41a0-89ae-6cc9a6d201e3 · inbound

Learning to accelerate distributed ADMM using graph neural networks cites this paper.

Learning to accelerate distributed ADMM using graph neural networks Gradient Methods with Online Scaling Part I. Theoretical Foundations

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:46:45.043079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T18:46:36.084583Z digest=sha256:0a35131608772f081bd28c7ad4ae707b266f55d1f6f03a4afdc5607f1e4f2a1d