Pith. sign in

Paper Citation Record · LEDGER

Stochastic Adaptive Gradient Descent Without Descent

As of 16 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2509.14969.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.14969 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:59:07.603628Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact1
  • verified fuzzy41
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28515651-884d-4a01-884c-bac83ee0ee7a · outbound

This paper cites Parameter-free FISTA by adaptive restart and backtracking.SIAM Journal on Optimization, 34(4):3259–3285, 2024.

Stochastic Adaptive Gradient Descent Without Descent Parameter-free FISTA by adaptive restart and backtracking.SIAM Journal on Optimization, 34(4):3259–3285, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.568467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.366329Z digest=sha256:e3548e9b21eb2d9aa89be184ffef8eb57292490f1bb6f871043915ff7d23636b

Observation e65ef1d3-1899-4412-b8d6-fe0c8e795fa1 · outbound

This paper cites FISTA restart using an automatic estimation of the growth parameter.Journal of Optimization Theory and Applications, 206(2):51, 2025.

Stochastic Adaptive Gradient Descent Without Descent FISTA restart using an automatic estimation of the growth parameter.Journal of Optimization Theory and Applications, 206(2):51, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.556523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.371418Z digest=sha256:0d751a7e2516ed0f178b619cbece06da6d54a224d4a887487d11cb639e103ca4

Observation 6d77bdeb-52f5-496f-a44e-62a4d89dde3d · outbound

This paper cites Complexity guarantees for Polyak steps with momentum.

Stochastic Adaptive Gradient Descent Without Descent Complexity guarantees for Polyak steps with momentum

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.546068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.375644Z digest=sha256:5dacf73b376a8211e05a3f4e9f9d9a52cd8637a407cfdb92bf4c11ae2fe41218

Observation e0313f02-1807-4d43-b1ea-fc839a17b4c0 · outbound

This paper cites Two-point step size gradient methods.IMA journal of numerical analysis, 8(1):141–148, 1988.

Stochastic Adaptive Gradient Descent Without Descent Two-point step size gradient methods.IMA journal of numerical analysis, 8(1):141–148, 1988

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:07.381309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:07.381309Z digest=sha256:ea3deed429f243f136afbc1c62620592ee2047a106bacd2af427da5ce4e18911

Observation d72975c5-2f8b-45e9-b78e-80874bfd2fdb · outbound

This paper cites an unresolved cited work.

Stochastic Adaptive Gradient Descent Without Descent Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:59:08.528569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.385336Z digest=sha256:0513d164b7f6f527612c81fec40a3e44cdac3ac3ce2ef6ddadc730a70e36e0b3

Observation a58afa3b-3cfb-4139-adb5-3cb279c76106 · outbound

This paper cites Sample size selection in optimization methods for machine learning.Mathematical programming, 134(1):127–155, 2012.

Stochastic Adaptive Gradient Descent Without Descent Sample size selection in optimization methods for machine learning.Mathematical programming, 134(1):127–155, 2012

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:07.389404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:07.389404Z digest=sha256:85fa0e9f2dbba418ca53bc3191ddf3cacf82b60dd1a8aa48cd54d3226d391318

Observation 1b59495b-ca11-45c4-8fbf-55e0dd863434 · outbound

This paper cites Making SGD parameter-free.

Stochastic Adaptive Gradient Descent Without Descent Making SGD parameter-free

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.485330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.394698Z digest=sha256:58c3dca294d4e08d435437e367799324d3f75e11cf1de5301bd64fdaff359172

Observation ee45d76b-4018-4ffa-87a4-8a087f48b818 · outbound

This paper cites Second-order step-size tuning of SGD for non-convex optimization.Neural Processing Letters, 54(3):1727–1752, 2022.

Stochastic Adaptive Gradient Descent Without Descent Second-order step-size tuning of SGD for non-convex optimization.Neural Processing Letters, 54(3):1727–1752, 2022

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.396419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.399142Z digest=sha256:99456917a4f87362069e56cee662c74210327d1c667c8d978b92e286a78163b5

Observation 0790e4b9-2d94-4595-b944-e8775f091fc1 · outbound

This paper cites Fast bundle-level methods for unconstrained and ball-constrained convex optimization.Computational Optimization and Applications, 73(1):159–199, 2019.

Stochastic Adaptive Gradient Descent Without Descent Fast bundle-level methods for unconstrained and ball-constrained convex optimization.Computational Optimization and Applications, 73(1):159–199, 2019

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.326011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.402857Z digest=sha256:cc090a960a97c894e5cdc781dbfd96f997d02a911a3fee622d87fde3274b823b

Observation c3abeb0f-0392-48a7-8dd6-8218b9f2d7f0 · outbound

This paper cites Convergence rates of gradient methods for convex optimization in the space of measures.Open J.

Stochastic Adaptive Gradient Descent Without Descent Convergence rates of gradient methods for convex optimization in the space of measures.Open J

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.315693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.406246Z digest=sha256:330ad7cc3e78b0c862e0a068b316824fac884bbb6072254928c75355245c8361

Observation 140c0664-2c14-49ef-8c24-735a3533fd20 · outbound

This paper cites New tight bounds for SGD without variance assumption: A computer-aided Lyapunov analysis.arXiv preprint arXiv:2505.17965, 2025.

Stochastic Adaptive Gradient Descent Without Descent New tight bounds for SGD without variance assumption: A computer-aided Lyapunov analysis.arXiv preprint arXiv:2505.17965, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:07.410124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:07.410124Z digest=sha256:65aaa023ce078a8551484742bb540e2da750059e985d4d11f642fbe974fdc8e3

Observation e26ba4f2-82e7-4cac-b269-8d3b3553ff1a · outbound

This paper cites Artificial constraints and hints for unbounded online learning.

Stochastic Adaptive Gradient Descent Without Descent Artificial constraints and hints for unbounded online learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.304197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.413476Z digest=sha256:815c1547623648a5d86b1d5ac35b164cf1ed4a99a2f8e0893f2fe5f9f0e4fff9

Observation 1c05c533-0bc4-49df-911f-dbf26ce0137a · outbound

This paper cites Learning-rate-free learning by d-adaptation.

Stochastic Adaptive Gradient Descent Without Descent Learning-rate-free learning by d-adaptation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:07.416714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:07.416714Z digest=sha256:5d0b7a2f9f7519bd7f3736fa6327c34719ac6a8865e4e64f65013ae3650cdd16

Observation eff84c2b-4497-480f-9b5b-b057e7cf982b · outbound

This paper cites Grad-GradaGrad? A Non-Monotone Adaptive Stochastic Gradient Method.

Stochastic Adaptive Gradient Descent Without Descent Grad-GradaGrad? A Non-Monotone Adaptive Stochastic Gradient Method

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:07.421848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:07.421848Z digest=sha256:243669b2bb1ae8f62061f35c745a81648ec9a7de15e1f93d14c35e93c46d4adc

Observation 3c24ff1d-ccba-4b4b-b9b5-52b9b9f4fada · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic opti- mization.Journal of Machine Learning Research, 12(7), 2011.

Stochastic Adaptive Gradient Descent Without Descent Adaptive subgradient methods for online learning and stochastic opti- mization.Journal of Machine Learning Research, 12(7), 2011

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.285698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.425840Z digest=sha256:9a298265b8158e057d256fc0928d329997ce395b287c7a017cdc0b6e5a2cc3cd

Observation bc8321df-3c0b-48d7-9eac-604bfc083a27 · outbound

This paper cites Duflo.Random iterative models, volume 34 ofApplications of Mathematics, New York.

Stochastic Adaptive Gradient Descent Without Descent Duflo.Random iterative models, volume 34 ofApplications of Mathematics, New York

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.272947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.429310Z digest=sha256:f48df4fff1691e2f5d27a64d0d27c12ae0280bceeb0b4e9555f47a03bc7a043c

Observation aa693b89-a4ec-43dd-a207-21f99067efa3 · outbound

This paper cites The power of adaptivity in SGD: Self-tuning step sizes with unbounded gradients and affine variance.

Stochastic Adaptive Gradient Descent Without Descent The power of adaptivity in SGD: Self-tuning step sizes with unbounded gradients and affine variance

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.261625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.433159Z digest=sha256:47882904b80f51babc2b3b200ea6d3ff35b9a3c6105f02b575f8eeac077f6c44

Observation 6e2d9cdb-3c3f-40b2-893a-7800631c9af1 · outbound

This paper cites Learning rate selection in stochastic gradient methods based on line search strategies.Applied Mathematics in Science and Engineering, 31(1):2164000, 2023.

Stochastic Adaptive Gradient Descent Without Descent Learning rate selection in stochastic gradient methods based on line search strategies.Applied Mathematics in Science and Engineering, 31(1):2164000, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.249426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.436608Z digest=sha256:60409ce456cc9606ed4e391bb519b9b9e82da8a4becf40c8fd57356879b4bc36

Observation cc2d94ea-5041-4793-9690-b9e0a3af17c5 · outbound

This paper cites Handbook of Convergence Theorems for (Stochastic) Gradient Methods.

Stochastic Adaptive Gradient Descent Without Descent Handbook of Convergence Theorems for (Stochastic) Gradient Methods

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:07.440472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:07.440472Z digest=sha256:bdd14477aaaaba2d991a466500c212040269202cf19de4a6cb092b1e617410aa

Observation cd6dab65-6b8b-468b-b38b-c5bcb06ad2bc · outbound

This paper cites A neural-network-based convex regularizer for inverse problems.IEEE Trans.

Stochastic Adaptive Gradient Descent Without Descent A neural-network-based convex regularizer for inverse problems.IEEE Trans

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.237925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.444984Z digest=sha256:a5e2ce4734b5f8b9469ff841bd70175e377aadfeef5f286209602f76e28fdee3

Observation feba319e-8c15-4fbc-a349-b7dad3d9efb4 · outbound

This paper cites On optimal universal first-order methods for minimizing heterogeneous sums.Optimization Letters, 18(2):427–445, 2024.

Stochastic Adaptive Gradient Descent Without Descent On optimal universal first-order methods for minimizing heterogeneous sums.Optimization Letters, 18(2):427–445, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.227519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.448339Z digest=sha256:5b4e0ea09fc9317eb6db5996274b826c4143b2580f7b094d8db7a4bdbcc80e90

Observation bacb869c-9fcb-4e48-a60b-22086e0761fe · outbound

This paper cites Revisiting the Polyak step size.

Stochastic Adaptive Gradient Descent Without Descent Revisiting the Polyak step size

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:07.451981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:07.451981Z digest=sha256:8345098cad26338c7925fe4b7d51007dca47cee81708eb72c49afb1efd29dbcb

Observation 7ce6b06f-3c1c-4aa6-ac22-42f9a31a9acb · outbound

This paper cites Ismailov.Ridge functions and applications in neural networks, volume 263 ofMathematical Surveys and Monographs.

Stochastic Adaptive Gradient Descent Without Descent Ismailov.Ridge functions and applications in neural networks, volume 263 ofMathematical Surveys and Monographs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.216038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.455860Z digest=sha256:530c7f5bb04c65f56ee4fc3fcdd414b141c4fd4d0e05c3e89b29def7eee8e219

Observation a175e88c-0b64-4b61-b98a-e8d77019fef2 · outbound

This paper cites DoG is SGD’s best friend: A parameter-free dynamic step size schedule.

Stochastic Adaptive Gradient Descent Without Descent DoG is SGD’s best friend: A parameter-free dynamic step size schedule

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.203786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.458997Z digest=sha256:cb78e120102017b7070bc2e499638ab75c0ccb87b059dd72c34c4868d6a66d01

Observation c50062e0-4bd4-4bdd-8fa6-fd1254a52309 · outbound

This paper cites Tuning-free stochastic optimization.

Stochastic Adaptive Gradient Descent Without Descent Tuning-free stochastic optimization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.191022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.462665Z digest=sha256:70b465e132a94f530df530fa6ea351f1429e3f9a79c082f6aa82fe1df0eef6ba

Observation 2dad2205-f1a9-4e60-835d-3f311736d5d4 · outbound

This paper cites DoWG Unleashed: An efficient universal parameter-free gradi- ent descent method.Advances in Neural Information Processing Systems, 36:6748–6769, 2023.

Stochastic Adaptive Gradient Descent Without Descent DoWG Unleashed: An efficient universal parameter-free gradi- ent descent method.Advances in Neural Information Processing Systems, 36:6748–6769, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.178752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.466111Z digest=sha256:80160c3c529dce085e6fb1b130ce1ab5a39665bb0d7ea1dbc1f776cd1357d248

Observation 303be1cc-f74d-40d8-af06-8b798d32d4e7 · outbound

This paper cites ADAM: A method for stochastic optimization.

Stochastic Adaptive Gradient Descent Without Descent ADAM: A method for stochastic optimization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.167436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.469865Z digest=sha256:c34b059f6e1c84615148f124fd695b5ef8c7824dde06f9c0e5deb8c27b4fd5f9

Observation 76274d5a-228b-4715-80d3-e5216f0852e4 · outbound

This paper cites Optimal and parameter-free gradient minimization methods for convex and nonconvex optimization.

Stochastic Adaptive Gradient Descent Without Descent Optimal and parameter-free gradient minimization methods for convex and nonconvex optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:07.473306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:07.473306Z digest=sha256:3cd85bbc1cabdfe9610e62ce2a3c5de12275b72e5f31a8a06e02651dddbc522d

Observation fb31f7df-5cd1-4727-830f-864e0dc89356 · outbound

This paper cites Adaptive proximal algorithms for convex optimization under local lipschitz continuity of the gradient.Mathematical Programming, pages 1–39, 2024.

Stochastic Adaptive Gradient Descent Without Descent Adaptive proximal algorithms for convex optimization under local lipschitz continuity of the gradient.Mathematical Programming, pages 1–39, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.150660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.476756Z digest=sha256:5078e546451939c436bf8fbafe828011f2994f375a829aa4df37aeeb51df0478

Observation 7f7e61f7-4a91-4cb1-8a2d-ae8c3823a604 · outbound

This paper cites Online to offline conversions, universality and adaptive minibatch sizes.Advances in Neural Information Processing Systems, 30, 2017.

Stochastic Adaptive Gradient Descent Without Descent Online to offline conversions, universality and adaptive minibatch sizes.Advances in Neural Information Processing Systems, 30, 2017

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.138355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.480870Z digest=sha256:87bbcc14a455bf2a77385e267495bf444cd71cb512489316bab617a7616e6968

Observation 8d01bc9b-8e29-440c-81b9-febc9acb6297 · outbound

This paper cites A simple uniformly optimal method without line search for convex optimization.

Stochastic Adaptive Gradient Descent Without Descent A simple uniformly optimal method without line search for convex optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:07.484917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:07.484917Z digest=sha256:80fc99bc39aa86426c6313cb7c0454721d225ba7add686224c5523571dcdd771

Observation 0c4f6c97-5b1e-49b4-b26f-7f3e2119ab7e · outbound

This paper cites On the convergence of stochastic gradient descent with adaptive stepsizes.

Stochastic Adaptive Gradient Descent Without Descent On the convergence of stochastic gradient descent with adaptive stepsizes

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.126222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.488891Z digest=sha256:aa88933345b069c7e50fc1805822198bfacc9310d8dd58d5c0f4c8bb5701018a

Observation 38d041ae-4ee6-477e-902f-99b01ca736b2 · outbound

This paper cites Stochastic polyak step-size for SGD: An adaptive learning rate for fast convergence.

Stochastic Adaptive Gradient Descent Without Descent Stochastic polyak step-size for SGD: An adaptive learning rate for fast convergence

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:07.493114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:07.493114Z digest=sha256:62aac56b187f04450f62eb485418fd407076c8f44c1f6265fa9c5d7637814fc5

Observation 24a15f6b-751f-4d91-894c-03ece99c5e46 · outbound

This paper cites Near-optimal Closed-loop Method via Lyapunov Damping for Convex Optimization.

Stochastic Adaptive Gradient Descent Without Descent Near-optimal Closed-loop Method via Lyapunov Damping for Convex Optimization

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:59:07.697262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.500932Z digest=sha256:7bb417ad2362167a77ff2d2eafa8f326bda33be3f77f253c3c45f1f57aa31901

Observation 4fb0c890-86e0-4d9f-9ca1-4bd35ca3b3d3 · outbound

This paper cites Adaptive gradient descent without descent.

Stochastic Adaptive Gradient Descent Without Descent Adaptive gradient descent without descent

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.107434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.505747Z digest=sha256:3ef62001d6dab3b9515e6e7a7c833a8a540a4f196d86859fcd3f71e39ed9d29f

Observation ff61b99f-f1a8-4c31-b5be-e7c1fa9d2798 · outbound

This paper cites Adaptive proximal gradient method for convex optimization.Advances in Neural Information Processing Systems, 37:100670–100697, 2024.

Stochastic Adaptive Gradient Descent Without Descent Adaptive proximal gradient method for convex optimization.Advances in Neural Information Processing Systems, 37:100670–100697, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.092416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.509441Z digest=sha256:3599453185739a19e49459da6cca13b4568d02bf3a1f2bf9a3f7a83c3ae5a33d

Observation 83ceb75f-5465-4585-a8b6-56b2eca9fbe4 · outbound

This paper cites Adaptive bound optimization for online convex optimization.

Stochastic Adaptive Gradient Descent Without Descent Adaptive bound optimization for online convex optimization

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.078631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.512636Z digest=sha256:04b64566637e904be6cdddd380cd105c0f6175480b413bfba219b267a776e641

Observation 590c4013-2aab-4b00-bd5b-7574d819ad9a · outbound

This paper cites Prodigy: An expeditiously adaptive parameter-free learner.

Stochastic Adaptive Gradient Descent Without Descent Prodigy: An expeditiously adaptive parameter-free learner

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.066315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.516726Z digest=sha256:6fc653242d8bc8cea525571893c0407c9dd7be87bfab584d9fcfa4676c3c4696

Observation 0a860d4c-3d81-4e56-97f7-2ed0db1dc669 · outbound

This paper cites Universal gradient methods for convex optimization problems.Mathematical Programming, 152(1):381– 404, 2015.

Stochastic Adaptive Gradient Descent Without Descent Universal gradient methods for convex optimization problems.Mathematical Programming, 152(1):381– 404, 2015

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.055338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.520428Z digest=sha256:7cc04c7e45449e68262b12e71dd5ececb698be42ef9479e8ee0a0b642adb5af8

Observation 9858f094-aec4-434e-97b3-b7a86d5480a1 · outbound

This paper cites A method of solving a convex programming problem with convergence rate O 1 k2.

Stochastic Adaptive Gradient Descent Without Descent A method of solving a convex programming problem with convergence rate O 1 k2

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.044167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.524703Z digest=sha256:f09e589b3303a760108ecffd83b87d428c85ef1d99b2bb423495c25700e0bf86

Observation defe4a81-6ef4-4fef-96b3-80681a1795c7 · outbound

This paper cites Springer Science & Business Media, 2013.

Stochastic Adaptive Gradient Descent Without Descent Springer Science & Business Media, 2013

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:07.530945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:07.530945Z digest=sha256:02cd7b9e6256cd5bcfab1071432ebb33169c4b060ddc19b191b294332f79712e

Observation 5ca82a62-91e6-47c9-a69e-3ffb1a9fb9f1 · outbound

This paper cites Simultaneous model selection and optimization through parameter-free stochastic learning.Ad- vances in Neural Information Processing Systems, 27, 2014.

Stochastic Adaptive Gradient Descent Without Descent Simultaneous model selection and optimization through parameter-free stochastic learning.Ad- vances in Neural Information Processing Systems, 27, 2014

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.024839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.534473Z digest=sha256:b1c513749c7ecc859257f2657877be00d6909e84cedaae85a383f97f680ada83

Observation 731ef92c-faeb-4a58-b737-16f0d94a6409 · outbound

This paper cites Normalized Gradients for All.

Stochastic Adaptive Gradient Descent Without Descent Normalized Gradients for All

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:07.539876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:07.539876Z digest=sha256:29491c4cc8469165ca3f358dc30b64f9631500ec88b1be4918902e21fd4bb41c

Observation 67d39c1b-c324-4921-aec4-45f45bce4444 · outbound

This paper cites Icml 2020 tutorial on parameter-free online optimization, 2020.

Stochastic Adaptive Gradient Descent Without Descent Icml 2020 tutorial on parameter-free online optimization, 2020

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:08.012262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.544180Z digest=sha256:3a4a5c29f4dd06ce8aca5eaec2360ac1945b16a9dded6ca291b9ba0e5f1cce87

Observation c0e0de3d-0d15-4999-a295-088e2c922ee8 · outbound

This paper cites Coin betting and parameter-free online learning.Advances in Neural Information Processing Systems, 29, 2016.

Stochastic Adaptive Gradient Descent Without Descent Coin betting and parameter-free online learning.Advances in Neural Information Processing Systems, 29, 2016

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:07.997311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.548269Z digest=sha256:7ca7992faa539ba42a610abb4beb01481dd20dd09c890591a0e00ce2fecb9978

Observation 1588ba47-5fc5-4574-81ac-439271e3e317 · outbound

This paper cites Parameter-free Stochastic Optimization of Variationally Coherent Functions.

Stochastic Adaptive Gradient Descent Without Descent Parameter-free Stochastic Optimization of Variationally Coherent Functions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:07.551940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:07.551940Z digest=sha256:bfbd1c0c9901824b684debc098050dfe2a19e6764e945d908f6261538cf577c3

Observation 141aeb97-6624-456d-a489-40cee3d94474 · outbound

This paper cites Training deep networks without learning rates through coin betting.Ad- vances in neural information processing systems, 30, 2017.

Stochastic Adaptive Gradient Descent Without Descent Training deep networks without learning rates through coin betting.Ad- vances in neural information processing systems, 30, 2017

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:07.979548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.557083Z digest=sha256:90d17eb9e854234e7cb0bdc27a5a5d6572b8cbb701c94ae06f8d130c0d8f69ce

Observation 050cc27b-5502-4d5b-bdbb-30a9a4c2d9ab · outbound

This paper cites Pedregosa, G.

Stochastic Adaptive Gradient Descent Without Descent Pedregosa, G

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:07.561094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:07.561094Z digest=sha256:5b6781c0de60e604dc0bd5a4b509e92d18c142406019194341702b01ffac5ccb

Observation d195bf5e-95d7-4a68-afab-f8d1ff788cd6 · outbound

This paper cites New York, Optimization Software,, 1987.

Stochastic Adaptive Gradient Descent Without Descent New York, Optimization Software,, 1987

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:07.953859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.564683Z digest=sha256:889f60cfbe841deb740d1678873d51a21360860f76e22233451ab86f2d5c0cb2

Observation 6bd2c4f5-c401-4783-aca7-48833feb4f82 · outbound

This paper cites Statistical complexity and optimal algorithms for nonlinear ridge bandits.The Annals of Statistics, 52(6):2557 – 2582, 2024.

Stochastic Adaptive Gradient Descent Without Descent Statistical complexity and optimal algorithms for nonlinear ridge bandits.The Annals of Statistics, 52(6):2557 – 2582, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:07.939370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.575245Z digest=sha256:6c1a641fd92bd1f748b92f568a99f0b53fdeaee7f30b74495a8dca62680ef445

Observation a0bd617d-532c-4d3c-b9d5-293dda0b98a6 · outbound

This paper cites The Barzilai and Borwein gradient method for the large scale unconstrained minimization problem.

Stochastic Adaptive Gradient Descent Without Descent The Barzilai and Borwein gradient method for the large scale unconstrained minimization problem

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:07.927117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.579375Z digest=sha256:805cb5105543d40e8cf6ad6a311c5967f94bf78d3866bb742279efcf1d4d8f69

Observation 4843c126-a99f-4cd0-8710-6d4c73e56b16 · outbound

This paper cites Robbins and D.

Stochastic Adaptive Gradient Descent Without Descent Robbins and D

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:07.912448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.583065Z digest=sha256:7c06eb040e760e25faaf94f234ae61459fe850b6bf417dbae13e05cd6dcd56ce

Observation 790950f1-63a9-41ae-aa73-4201b3363fbb · outbound

This paper cites Robles-Kelly and A.

Stochastic Adaptive Gradient Descent Without Descent Robles-Kelly and A

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:07.897216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.587995Z digest=sha256:fcfc999b99ded1e4e5270d301c06935f3c8d7b133f43628df8c9d47e4aadc6ec

Observation d30a7e49-5966-47c4-a92d-3a129da31d48 · outbound

This paper cites Optimizer benchmarking needs to account for hyperparameter tuning.

Stochastic Adaptive Gradient Descent Without Descent Optimizer benchmarking needs to account for hyperparameter tuning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:07.883611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.591732Z digest=sha256:3f7e3f4d79de148276dc2402873027972c13c60f37024a0266944a15ea003f8a

Observation 8a43bd19-7ae7-450d-9640-45793f57df6b · outbound

This paper cites Barzilai-borwein step size for stochastic gradient descent.

Stochastic Adaptive Gradient Descent Without Descent Barzilai-borwein step size for stochastic gradient descent

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:07.871122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.595793Z digest=sha256:3f2ed81357a7db7aec680072864f6bd2b2bd4ab32427645c594954c5ae907c1c

Observation 43fa126e-d9c9-45fb-8e63-d439ced752b0 · outbound

This paper cites RMSprop: Divide the gradient by a running average of its recent magnitude.

Stochastic Adaptive Gradient Descent Without Descent RMSprop: Divide the gradient by a running average of its recent magnitude

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:07.858763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.599408Z digest=sha256:c1735dc38995b601390adbf623694cc6030935083cbef33adfe238b9e79d58bb

Observation 6bd536bb-2809-4111-ae3c-ae8ee4d89d70 · outbound

This paper cites Finally, for the right-hand side of (26), from (4), θkλk = λ2 k λk−1 ≤λ k−1 1 + 1− 1 k1/2+δ θk−1 =λ k−1 1 +θk−1− θk−1 k1/2+δ =λ k−1 (1 +θk−1) 1− 1 k1/2+δ θk−1 1 +θk−1.

Stochastic Adaptive Gradient Descent Without Descent Finally, for the right-hand side of (26), from (4), θkλk = λ2 k λk−1 ≤λ k−1 1 + 1− 1 k1/2+δ θk−1 =λ k−1 1 +θk−1− θk−1 k1/2+δ =λ k−1 (1 +θk−1) 1− 1 k1/2+δ θk−1 1 +θk−1

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:59:07.848187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:59:07.603628Z digest=sha256:8acde390b4838dd1e00d99fe81c2881fd7311439814d9291c6ff209e44151edf

Pith citing papers

No inbound Pith citation observations are available.