Pith. sign in

Paper Citation Record · LEDGER

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness

As of 16 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2505.04599.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04599 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:30:25.152781Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:18:52.031537Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T16:18:52.617155Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2694726-8f38-48c1-b075-4263af3617fb · outbound

This paper cites an unresolved cited work.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:30:25.769951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.003896Z digest=sha256:c97a547c135af11e87dd09b52f6f84e8d649e5dc7bb4317e40b8e4558ebe04e5

Observation f1d6d86c-4b02-4429-b2b8-f7f311c05e42 · outbound

This paper cites Therefore, the effective learning rate of the algorithm at stept is ηt = η √ γ2 +∑ t−1 i=0 ‖F (xi,ξi)‖2 = η√ γ2 +t(ǫ2 +σ2) =αt+2.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Therefore, the effective learning rate of the algorithm at stept is ηt = η √ γ2 +∑ t−1 i=0 ‖F (xi,ξi)‖2 = η√ γ2 +t(ǫ2 +σ2) =αt+2

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.722252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.022847Z digest=sha256:0473e80e3c2763f1a96cf2d6625a0c0accc94650a80fec438b77df9468ed364c

Observation f71e23ea-10e6-4706-ac5e-30ff33ab47a8 · outbound

This paper cites We now bound the remaining constants.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness We now bound the remaining constants

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.510513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.100485Z digest=sha256:a4db12464ea7166b7f1d198146fcbb0423ba8c52e66949d9db3718e0e72d2d4a

Observation 162b8764-b904-4214-ab15-703b480c2980 · outbound

This paper cites an unresolved cited work.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:30:25.656138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.048018Z digest=sha256:887a4e8c675d2c6356180df184652466101e3091e008e65b5a522fa23a04cb48

Observation 5f2d3f58-2eb0-4417-b339-cf729753f331 · outbound

This paper cites Let algorithmADAN denote Decorrelated AdaGrad-Norm with parameters η >0 and 0<γ ≤ ∆ L1 8 log ( 1 + 48 ∆ L2 1 L0 ).

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Let algorithmADAN denote Decorrelated AdaGrad-Norm with parameters η >0 and 0<γ ≤ ∆ L1 8 log ( 1 + 48 ∆ L2 1 L0 )

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.688372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.034340Z digest=sha256:87f4dd61e22faf9ab5257a7fb01c78c45947d385344a976cc27c53f713da650a

Observation 05d82fd6-c6b7-4986-abb3-0b577140bb86 · outbound

This paper cites Near -optimal non-convex stochastic opti- mization under generalized smoothness.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Near -optimal non-convex stochastic opti- mization under generalized smoothness

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:30:24.976572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:30:24.976572Z digest=sha256:295c5a7a0cd4d1b93bcf939ccd83d651ef634295f7c1b8ecb71fcfff9c15bc8f

Observation fad08475-e6ef-4ba2-9812-fe6e7d8c2154 · outbound

This paper cites Adaptive Bound Optimization for Online Convex Optimization.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Adaptive Bound Optimization for Online Convex Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:30:24.980118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:30:24.980118Z digest=sha256:92a106c7bd26f1a9dbf3dbf9ef0c5ecd1bed1af3c9ad52515bb68da74fe3d9de

Observation 70402f0a-9d47-4096-b2f3-c8db8dc97d71 · outbound

This paper cites Variance-reduced Clipping for Non-convex Optimization.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Variance-reduced Clipping for Non-convex Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:30:24.984392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:30:24.984392Z digest=sha256:d616c01034def38a94824fb7c59a4c2baadac7cbdf595dbc965cdd4a382c63cd

Observation fc002a4f-2e9b-478e-b4d3-98062730549d · outbound

This paper cites Adagrad stepsizes: Sharp convergence over nonconvex landscapes.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Adagrad stepsizes: Sharp convergence over nonconvex landscapes

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.789540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:24.988099Z digest=sha256:c41084103ed9fcea2840b89854032cb1d43b877f14bf1a7415a580aacb2ecc84

Observation cdd59b75-7938-48fb-92d4-ce03865715d6 · outbound

This paper cites Suppose g ∈ Rd with ‖g‖ =ǫ.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Suppose g ∈ Rd with ‖g‖ =ǫ

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.555925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.085428Z digest=sha256:cbf85d5266481fee3f2519ddc1efbb4b2968f85e89d59dda5ba9f98bd06a978b

Observation 4bfd7057-a46e-485e-8197-b0daceca7bb4 · outbound

This paper cites an unresolved cited work.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:30:25.412347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.134599Z digest=sha256:dc06e373c6f4031b17a934860a173814cf49348d59bce0f7cee6903f68160945

Observation 41c979b7-8086-4eeb-a316-0491e8f52aac · outbound

This paper cites (8) 15 Published as a conference paper at ICLR 2025 The RHS of Equation 7 can be bounded as 4 L1 log ( 1 + L1gt+1 L0 ) = 4 L1 log ( 1 + ∆ L2 1 L0 ( 576(t +.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness (8) 15 Published as a conference paper at ICLR 2025 The RHS of Equation 7 can be bounded as 4 L1 log ( 1 + L1gt+1 L0 ) = 4 L1 log ( 1 + ∆ L2 1 L0 ( 576(t +

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.760537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.008158Z digest=sha256:df7e785a4faa0c005933c55375f29f931e5ffb8a450c4b54f0d1b6cfbdf5e8eb

Observation 81248594-5734-4c61-8ff6-21b98f444777 · outbound

This paper cites f is informally pictured in Figure 1b of the main text.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness f is informally pictured in Figure 1b of the main text

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.750835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.011864Z digest=sha256:e51f7bdf1a3279b5f420974c05f7cb12289d89ce8655d7bf61a244e2f815998a

Observation 03ce0f66-4c72-4d0f-b28c-2ef8ae9bc01c · outbound

This paper cites Thereforef (x0) − infxf (x) ≤ ∆.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Thereforef (x0) − infxf (x) ≤ ∆

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.741400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.015426Z digest=sha256:a87cb4ac9352d8f43d6e37cbb0ec0f985797623e727650f25fd9ef9b785261f1

Observation 51b42532-16df-4479-9363-b0ce55b2b23a · outbound

This paper cites an unresolved cited work.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:30:25.732272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.019334Z digest=sha256:04783692e9c66bcd0bf37bafc997db50d6a40bba54acf7515882b175b152046f

Observation 0dd5b158-8ee9-4016-80d8-dc3d403b0b16 · outbound

This paper cites Actually,f does not satisfy this condition becausef is not even lower bounded, due to the linear term ǫ⟨x, e1⟩.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Actually,f does not satisfy this condition becausef is not even lower bounded, due to the linear term ǫ⟨x, e1⟩

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.711410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.026471Z digest=sha256:de2afaef5ff8b8d8c4acca15f0c01581a4faf5796afb82ac525c1fed2fe67bb0

Observation 28da9ef0-af29-4dc3-839e-92775aa0863f · outbound

This paper cites Specifically, we need ˆf which is lower bounded and that satisfies: ∇ ˆf (xt) = ∇f (xt), ˆf (xt) = f (xt) for all 0 ≤t ≤T.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Specifically, we need ˆf which is lower bounded and that satisfies: ∇ ˆf (xt) = ∇f (xt), ˆf (xt) = f (xt) for all 0 ≤t ≤T

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.699836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.030305Z digest=sha256:85c24259679c9784a80c31f758056ee2056fca52f87dc9ed245ac1790aed7dde

Observation a0e8882c-eeb2-4dd2-a9ff-16b102a903de · outbound

This paper cites First, recall the definition of ψ: ˜ψ(x) = L0 L2 1 (exp (L1|x|) −L1|x| − 1).

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness First, recall the definition of ψ: ˜ψ(x) = L0 L2 1 (exp (L1|x|) −L1|x| − 1)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.677814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.038819Z digest=sha256:67148669c329c554eb7b1d5803391fbcbc6a751d785ae29c08e2d0259b7465d1

Observation e3562ffb-1431-458e-b79b-44c8185ca0ba · outbound

This paper cites an unresolved cited work.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:30:25.666623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.043439Z digest=sha256:7e6be676df01a0087d8a6d967d715484b7f8145417e7fc1a40703b4c9c6414a9

Observation 7cdca7ed-d1b5-49ec-a00b-119708470b12 · outbound

This paper cites Therefore, with the initial point x0 =m + ∆ 2ǫ , the objective satisfies f (x0) − inf x f (x) = ǫ(x0 −m) +ψ(m) =ǫ ∆ 2ǫ + ∆ 2 = ∆.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Therefore, with the initial point x0 =m + ∆ 2ǫ , the objective satisfies f (x0) − inf x f (x) = ǫ(x0 −m) +ψ(m) =ǫ ∆ 2ǫ + ∆ 2 = ∆

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.645814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.052827Z digest=sha256:6efb6373095444c06d7bae075681b268d57ae525623e7a157a2603b030346d4d

Observation b84d0707-bf2e-4b9c-9461-5f94a4638b98 · outbound

This paper cites If η ≥ √ 2γ L1σ log ( 1 + L1ǫ L0 ) , then by Lemma 3 there exists a problem instance for which Dec orrelated AdaGrad will never find an ǫ-approximate stationary point.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness If η ≥ √ 2γ L1σ log ( 1 + L1ǫ L0 ) , then by Lemma 3 there exists a problem instance for which Dec orrelated AdaGrad will never find an ǫ-approximate stationary point

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.634333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.056816Z digest=sha256:54368cb2ae56e73333d23c183c905a5486c6ab8b0ba733ba8371b2e13c9cdd9e

Observation 297f349c-c3b1-4764-b5e4-2ad024615f47 · outbound

This paper cites an unresolved cited work.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:30:25.623723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.061332Z digest=sha256:b6767e35855ddc5dfddd5df464b41fcf88e2d17d9ef207fd692527d8c91e46a3

Observation 7d83cb05-b831-4ea8-b494-0dea2b7ec6cc · outbound

This paper cites Recall the function ψ : R → R defined as ψ(x) = L0 L2 1 (exp(L1|x|) −L1|x| − 1).

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Recall the function ψ : R → R defined as ψ(x) = L0 L2 1 (exp(L1|x|) −L1|x| − 1)

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.612319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.065780Z digest=sha256:e4e65d4ad7434c3681f309dfeccde6a38769ad018b7928eba1753e9b48f3d374

Observation 8e5ee1b8-2362-45e8-9dc0-d76e99a1a971 · outbound

This paper cites In this case, the learning rate α(g) is large enough to ensure that f (xt+1) ≥ f (xt) for an exponentially increasing f.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness In this case, the learning rate α(g) is large enough to ensure that f (xt+1) ≥ f (xt) for an exponentially increasing f

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.600941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.069732Z digest=sha256:4c7f6e88d30d83115e82eb08c9926ed39b8ea26042da084d5cc1eb66b8cc4962

Observation 04a1310b-39c7-43ed-b626-2dc68e25eca2 · outbound

This paper cites This completes the induction.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness This completes the induction

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.589932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.073300Z digest=sha256:ad6416ecfeb06056c55d18933d2c3209e9c472cb09810c7ae3462f2c41f8a8d4

Observation 3cc4a260-40a6-410a-ba46-11495783b782 · outbound

This paper cites Also, ‖g1 −ℓg‖ = |c1 −ℓ|‖g‖ =ℓ −c1 = 1 −p p (c2 −ℓ) ≤ 1 −p p (σ1 +σ2ℓ) ≤σ1 +σ2ℓ, where the last inequality uses p > 1 2.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Also, ‖g1 −ℓg‖ = |c1 −ℓ|‖g‖ =ℓ −c1 = 1 −p p (c2 −ℓ) ≤ 1 −p p (σ1 +σ2ℓ) ≤σ1 +σ2ℓ, where the last inequality uses p > 1 2

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.578621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.076934Z digest=sha256:ad9a45701b60fdc300a6eecbb2b3a36d599f0b3ad64b4c3917a715d83ae691df

Observation 8edbb6e2-a660-464a-b4a4-dcee4a7d6676 · outbound

This paper cites The upper bound of ‖yi‖ in the definition of k1 ensures that Equation 31 is satisfied.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness The upper bound of ‖yi‖ in the definition of k1 ensures that Equation 31 is satisfied

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.543601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.089355Z digest=sha256:8f0a7ec65aaf3021777cfcfa28313eded9a852e92aa50e6aee41d583d23e0b9e

Observation 7251ec45-044a-4a1d-8e24-8563c5db5e59 · outbound

This paper cites an unresolved cited work.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:30:25.532573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.092804Z digest=sha256:6178f7f4682da543dbc775eb10f64d537df64e3f65793670dc045ea7636ee84a

Observation 61092573-2b38-42db-acf3-14e22a087e98 · outbound

This paper cites We can also bound β(yk1 ) using the assumed condition α(g) < 4m |g| , since we previously showed that (yk1, yk+1) satisfies Equation 29 through Equation.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness We can also bound β(yk1 ) using the assumed condition α(g) < 4m |g| , since we previously showed that (yk1, yk+1) satisfies Equation 29 through Equation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.521966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.096643Z digest=sha256:85f1430ca9eafa6ab024e3052b70150994ffbe7673475686e4a2853c9591cd07

Observation 56975a94-a671-43bf-ba50-5c8f4124e53f · outbound

This paper cites an unresolved cited work.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:30:25.500610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.105179Z digest=sha256:d603b8109b3116261132e6f66c56410bbf60e95e928826baedd7cb06b43cbab0

Observation 5fefb382-468e-4d51-b788-190b62a21471 · outbound

This paper cites For b1: b1 = 1−p1 p1 ( σ1 + ( σ2 − p1 1−p1 ) G ) ((σ2 + 1)(2p1 −.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness For b1: b1 = 1−p1 p1 ( σ1 + ( σ2 − p1 1−p1 ) G ) ((σ2 + 1)(2p1 −

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.489870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.108940Z digest=sha256:463155cac42d36f051dd43db5117d7b4a5d0d564464e6da07fd0ced605d80ade

Observation 2ea15521-7d9a-4e13-95af-0e7976cf6cb1 · outbound

This paper cites Also as in the first case, |c1 −ℓ| ≤ |c2 −ℓ|.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Also as in the first case, |c1 −ℓ| ≤ |c2 −ℓ|

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.566506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.080939Z digest=sha256:2d2f452f88f294da573f6839d85253ec53631c71f4fd0865d6162e0f52dd4c14

Observation 687c049a-651f-4bb9-8677-b92c85d43cbc · outbound

This paper cites an unresolved cited work.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:30:25.479666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.112351Z digest=sha256:108e35e3213532057690e06ac2f0a7645a6254a6f07879a8e27ecb680491e49a

Observation 5464baa6-62bf-43ff-ad3a-5ba9bc4cd8f8 · outbound

This paper cites 42 Published as a conference paper at ICLR 2025 Proof.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness 42 Published as a conference paper at ICLR 2025 Proof

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.469142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.116307Z digest=sha256:44a0afa32bffd4dc29f933123a78829e934a90743efe897d8cbec3932a70646a

Observation 7ac47581-611f-4ff7-a90c-6d649946762a · outbound

This paper cites Therefore,t ≤ ∆ 2α(ǫ)ǫ2 implies thatt<t 0 + 1, so that ˆPg(xt) ≥a by the definition of t0, and finally ‖∇f (xt)‖ =ǫ.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Therefore,t ≤ ∆ 2α(ǫ)ǫ2 implies thatt<t 0 + 1, so that ˆPg(xt) ≥a by the definition of t0, and finally ‖∇f (xt)‖ =ǫ

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.457357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.119801Z digest=sha256:0b0e7cf0f1c3ef40c5f7512db959d03b276725c27382593f510e6bb71ff33a24

Observation 8374617c-f360-421f-9471-d76aa8a90d7b · outbound

This paper cites an unresolved cited work.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:30:25.445527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.123414Z digest=sha256:2a0c937d430ec63e1b069f3ea96ece75e20b0c174e8b7e88136b82607fcce6f9

Observation 274a6c48-0ec7-4c93-9646-07fa5a89247e · outbound

This paper cites Therefore ∇f (xt) = ǫe1.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Therefore ∇f (xt) = ǫe1

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.434670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.127355Z digest=sha256:feb9c3f5f71d4efdea0aac770a9b05c6eb3fa4741ab535c8e9dcfed38ae8b7fd

Observation 34f020d3-7745-4672-90f3-a910b0f0d246 · outbound

This paper cites Together, these three equations imply that ‖∇f (xt)‖ =ǫ for allt ≤T , which is the desired conclusion.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Together, these three equations imply that ‖∇f (xt)‖ =ǫ for allt ≤T , which is the desired conclusion

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.423850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.131234Z digest=sha256:8ce865fec1235a2e73c7186c37b2aa4691cd0e7d0eedfe94a5138ef1a17688b9

Observation 1ab1cf66-ddda-4070-a7a1-1f7d0d62b3fd · outbound

This paper cites By the monotone convergence theorem, E[τ ] = limT →∞ E [Xτ ∧T ] − 1 (λ + 1)p − 1 We consider the following cases.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness By the monotone convergence theorem, E[τ ] = limT →∞ E [Xτ ∧T ] − 1 (λ + 1)p − 1 We consider the following cases

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.399789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.138041Z digest=sha256:06b608f26d2820c2dbc454923b84b3d280ce5945b4a42963a77231fccbf9e079

Observation 38ff7a7f-edb0-46fd-985b-c45f807c2ea0 · outbound

This paper cites Specifically, we need r(λ) is decreasing (59) lim λ→ 1−p p + r(λ) = 1 (60) lim λ→∞ r(λ) = 1 −p.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Specifically, we need r(λ) is decreasing (59) lim λ→ 1−p p + r(λ) = 1 (60) lim λ→∞ r(λ) = 1 −p

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.386052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.141476Z digest=sha256:0727031ac13589fc594441ab0d8e2882e85a87c6bee1adc3250d8b1383b7af66

Observation 227c42e6-5bb1-43d7-9daa-d3de12f41684 · outbound

This paper cites an unresolved cited work.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:30:25.372486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.145227Z digest=sha256:ddcd19a46b0240e7aaebac1763ee7fb2eafceb0f0f10d9470a510f66abdb0a35

Observation 2a20dbf4-4857-4e3a-bccf-e81461d8b5dd · outbound

This paper cites If ∆ L2 1 ≥L0, then T (ADAN, Fdet,ǫ ) ≥ ˜Ω (∆ 2L2 1 ǫ2 ).

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness If ∆ L2 1 ≥L0, then T (ADAN, Fdet,ǫ ) ≥ ˜Ω (∆ 2L2 1 ǫ2 )

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.360829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.148947Z digest=sha256:4a3f08c633fb0bf4ef1427aa2718f3282265376a98dacb95bcbc55fd820de238

Observation 5685fa76-27d4-4ab2-a9b7-f16a9e7e96e7 · outbound

This paper cites an unresolved cited work.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:30:25.348728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.152781Z digest=sha256:76adcb4c4e5c2e7a68bd00faf8b8e35bc5a7f79b07f9055061ccca9ef35e79de

Observation 60c35cc6-fae6-4f5b-9f98-d7a4822cc160 · outbound

This paper cites On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-15T23:30:24.853944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:30:24.853944Z digest=sha256:6f18635c1de21707aaa7276d8198a6c0943b7ebd6cb740b1db0b3f3afdfbd0fe

Observation 0fb4aa0f-b843-4aef-822b-6b94177a9ecd · outbound

This paper cites A Novel Convergence Analysis for Algorithms of the Adam Family.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness A Novel Convergence Analysis for Algorithms of the Adam Family

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-15T23:30:24.968099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:30:24.968099Z digest=sha256:5e2a830ce19dd684a137f2114cd2308b0d8af5d99362f9287ca726df33b791df

Observation a5259fda-dac4-4253-a60d-3d3f72b8f65b · outbound

This paper cites The Min-Max Complexity of Distributed Stochastic Convex Optimization with Intermittent Communication.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness The Min-Max Complexity of Distributed Stochastic Convex Optimization with Intermittent Communication

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T23:30:24.995999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:30:24.995999Z digest=sha256:814f4a3f311bea8e42912d426dd98c249b609c556c22bc50293ec15eac3f24f6

Observation 0fbac585-7ee2-41da-baac-0a9bc7e863e5 · outbound

This paper cites Generalized-Smooth Nonconvex Optimization is As Efficient As Smooth Nonconvex Optimization.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Generalized-Smooth Nonconvex Optimization is As Efficient As Smooth Nonconvex Optimization

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T23:30:24.858567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:30:24.858567Z digest=sha256:a62fb2e0ed0904f4f9cfc41603cfb136605578ebf0218768f90abec2d46debf8

Observation 169bd41b-dcaa-489f-a398-2a81204853fc · outbound

This paper cites Lower Bound for Randomized First Order Convex Optimization.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Lower Bound for Randomized First Order Convex Optimization

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T23:30:24.992058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:30:24.992058Z digest=sha256:7972dd86be024883f0e0b58e25642c387eeba15d360a59e9c892b8a606f33eaf

Observation 394bb7db-b0e1-4562-9761-40089c868c0e · outbound

This paper cites Beyond Uniform Smoothness: A Stopped Analysis of Adaptive SGD.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Beyond Uniform Smoothness: A Stopped Analysis of Adaptive SGD

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T23:30:24.863084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:30:24.863084Z digest=sha256:a8e079593f845f6f1c8942c8d42cd4597f5806fb0b7ad6891b75d4da5cd3b275

Observation e80aaad8-bef7-4932-bbbd-146d406728ea · outbound

This paper cites Convergence of Adam Under Relaxed Assumptions.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Convergence of Adam Under Relaxed Assumptions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T23:30:24.972422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:30:24.972422Z digest=sha256:0c2691d0d987a6f154370ecb4f05510df7f6ddfdb96e6008147088839c4758e9

Observation b56d61b4-c83b-44de-af95-3a3986ed3f10 · outbound

This paper cites Improved analysis of clipping algorithms for non-convex optimization.

Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness Improved analysis of clipping algorithms for non-convex optimization

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:30:25.779616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:30:25.000113Z digest=sha256:b74b79c0406eb8edbb40afffae20da15dd94b188d2697402b193659f32bb35bd

Pith citing papers

Observation 3e89b0cd-e87f-4110-b37e-38e86be44b1a · inbound

Decentralized Stochastic Nonconvex Optimization under the $(L_0,L_1)$-Smoothness cites this paper.

Decentralized Stochastic Nonconvex Optimization under the $(L_0,L_1)$-Smoothness Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:18:52.623130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T16:18:52.031537Z digest=sha256:c372b063c3f261c9f03df5b0e3a5f86511fa7ee2e5a86fd7027c72ac7862d436