Pith. sign in

Paper Citation Record · LEDGER

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping

As of 12 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2412.19529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19529 v4

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:23:05.364651Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:34:48.195468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T05:37:19.252093Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0fc28118-baf3-49a6-b1eb-9b8ff1e4a8eb · outbound

This paper cites write newline.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.160353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.160353Z digest=sha256:7847c8fd1e4fe442d1a7e9aa1760d722470c87b18916dba422dd3971bec0ca1a

Observation 24bcfccf-0945-4922-83e1-30dc3419e7ef · outbound

This paper cites Lower bounds for non-convex stochastic optimization.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Lower bounds for non-convex stochastic optimization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.165327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.165327Z digest=sha256:9a7c38b8bb1f53dc189f8e41702e7988d47c788f928d7f2e0c342801feb1d94f

Observation d3ad11f4-6b02-432f-9a2e-524cc9b75066 · outbound

This paper cites Optimization methods for large-scale machine learning.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Optimization methods for large-scale machine learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.170048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.170048Z digest=sha256:ed0658f68f1c92b7f27ed4d2cb495c54d4744e66c3a4215ea950b9632a236e3d

Observation fc9be62a-f9e0-4c6e-9034-0fac49ad27a5 · outbound

This paper cites Extrapolation and interpolation of quasi-linear operators on martingales.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Extrapolation and interpolation of quasi-linear operators on martingales

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:06.005562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.174249Z digest=sha256:b1d7315644b3b65ccb7088f36958d838f14c441a070dff278730fc8e116424d9

Observation a04d6d5b-be21-48a0-9ad1-7c5435dcad12 · outbound

This paper cites Martingale transforms.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Martingale transforms

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.994834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.178054Z digest=sha256:061a6f86ae972823ce839a725bc2b6ec7380d8fb94ab25e506e07155a562dd21

Observation ad4242fa-2a07-4bc1-be10-f72ca3ff91d1 · outbound

This paper cites Lower bounds for finding stationary points i.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Lower bounds for finding stationary points i

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.983749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.181956Z digest=sha256:c6af37bc573752db4e7b506017f0b0c106a908f15d7b5a76f9b0a7d59a3b2a5a

Observation 2b4ca97b-f132-46bd-9581-fe7ede0f56e8 · outbound

This paper cites Generalized-smooth nonconvex optimization is as efficient as smooth nonconvex optimization.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Generalized-smooth nonconvex optimization is as efficient as smooth nonconvex optimization

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.971418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.185848Z digest=sha256:0750c43f7d963575308a6f0ed045c7a323dd4352fb09a20792b6c3e9c36fb47a

Observation ad5105ba-d011-475e-91a7-7c4bba9741f7 · outbound

This paper cites Robustness to unbounded smoothness of generalized signsgd.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Robustness to unbounded smoothness of generalized signsgd

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.960483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.190382Z digest=sha256:d1e582afd36fa78a7e28348b980d20069c1985c8bb2eef45a34a003100c76324

Observation 2c750b12-965d-46cb-9e03-0e660ef42e10 · outbound

This paper cites Momentum improves normalized SGD.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Momentum improves normalized SGD

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.949463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.195061Z digest=sha256:a91d8e732ec27b0fd794b8a7eaeb641dd230005f29538708106eed31913bed9c

Observation a16ab0f0-9084-4c36-aa29-4531ae374c28 · outbound

This paper cites High-probability bounds for non-convex stochastic optimization with heavy tails.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping High-probability bounds for non-convex stochastic optimization with heavy tails

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.938843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.198729Z digest=sha256:e5737a2f88613acb5baf2730d8068ce5642297eae33ac1a7d0453ccd4fefb01a

Observation 80b17b9a-bb29-414c-9873-9c2b50e97c26 · outbound

This paper cites On the intergrability of the martingale square function.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping On the intergrability of the martingale square function

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.928488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.202455Z digest=sha256:076758d8f9f2bca6c4650d0fca9be4e410035ae3e961ba857e16658649448d50

Observation afd19ac1-36ca-4c2c-8f98-181502c26e4b · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Adaptive subgradient methods for online learning and stochastic optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.206190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.206190Z digest=sha256:2f42ab3d0ba690435adcd094f6cdd5d9b77802c0d6d00c81ebf1d2b0eb9e3bf5

Observation 58aea4d7-3ec7-4d56-8a4e-71a109f52808 · outbound

This paper cites Beyond uniform smoothness: A stopped analysis of adaptive sgd.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Beyond uniform smoothness: A stopped analysis of adaptive sgd

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.911746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.209701Z digest=sha256:3027ef6503948d0f3fdbf2afdc94b5281e3445fb7796d9621369ffc79507e33d

Observation 35ad4cc9-880d-4bf3-8011-0a57d2d20282 · outbound

This paper cites Stochastic first-and zeroth-order methods for nonconvex stochastic programming.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Stochastic first-and zeroth-order methods for nonconvex stochastic programming

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.213244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.213244Z digest=sha256:346403618b370f13a1fc141cbe7fc589246aa6b8afb3200068e0e37e628e113f

Observation f6c1cbb4-d6d2-4e45-a082-e09f526ac253 · outbound

This paper cites High-probability convergence for composite and distributed stochastic minimization and variational inequalities with heavy-tailed noise.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping High-probability convergence for composite and distributed stochastic minimization and variational inequalities with heavy-tailed noise

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.895550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.216626Z digest=sha256:c949042b8628dea2da2d5efb51de42437da998514107fc1caa8f3aa4c78e4b62

Observation 75b954ad-2c0f-4a4a-a7f2-e71bf02b0ded · outbound

This paper cites Beyond convexity: Stochastic quasi-convex optimization.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Beyond convexity: Stochastic quasi-convex optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.220660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.220660Z digest=sha256:7894d2ac3d17baa24435e69f55d34837d1d6b065999c9595c54080e3b51bb073

Observation ee49d851-ff89-406e-a56a-2afb4caed0c7 · outbound

This paper cites Revisiting Convergence of AdaGrad with Relaxed Assumptions.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Revisiting Convergence of AdaGrad with Relaxed Assumptions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.225264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.225264Z digest=sha256:4d60f3a48c525cc3497fdb187fbc28fa9f2d4ea6e097773784cba40bbb1b2c12

Observation 1a407fcb-ee45-407f-bae5-678c604665f2 · outbound

This paper cites From Gradient Clipping to Normalization for Heavy Tailed SGD.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping From Gradient Clipping to Normalization for Heavy Tailed SGD

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.229100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.229100Z digest=sha256:f751d7742c2888d912719abc21640c03612dd71ac54f52a938fe93ccbff378b6

Observation 8ee79d4f-0f3a-4261-827b-e32fd7b80bba · outbound

This paper cites Parameter-agnostic optimization under relaxed smoothness.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Parameter-agnostic optimization under relaxed smoothness

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.879086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.233151Z digest=sha256:c94618da04f1c63311ad41f375194e6079590d86f033560e15257974b97e1e69

Observation 77125654-26d1-4b33-8fa1-ae009d22de61 · outbound

This paper cites Non-convex distributionally robust optimization: Non-asymptotic analysis.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Non-convex distributionally robust optimization: Non-asymptotic analysis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.868128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.236891Z digest=sha256:6aebc38de5fab78f1769412cfeda218b0fa6a6bc11e274478bdeb05ca22066c6

Observation 46a386e5-9403-452e-a034-af1360ef99b3 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Adam: A Method for Stochastic Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.240285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.240285Z digest=sha256:4090b09e6effc4916785b85976f6dc00c164e2f0c8a97ce52237fcd0fc2e5f1b

Observation 638ce608-afd1-4407-9a37-5c6c522e82ea · outbound

This paper cites Convergence and efficiency of subgradient methods for quasiconvex minimization.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Convergence and efficiency of subgradient methods for quasiconvex minimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.857095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.243723Z digest=sha256:6a488476449dbef04e777349eaf802f026ce8ec23a10db8e7545bb200200ae9d

Observation 5fa88c4b-f563-4aae-bc01-ea316aca6af1 · outbound

This paper cites Accelerated zeroth-order method for non-smooth stochastic convex optimization problem with infinite variance.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Accelerated zeroth-order method for non-smooth stochastic convex optimization problem with infinite variance

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.845941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.247430Z digest=sha256:e3c0c98cbcafca0e404ff65ebe82ced9cb66259fbdf94ce1b367ee4f43a36b9a

Observation ebd65b1a-f0c5-4cf6-a42e-a3ce8600a10e · outbound

This paper cites First-order and stochastic optimization methods for machine learning.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping First-order and stochastic optimization methods for machine learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.250868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.250868Z digest=sha256:a214ab4351a4a3db2f4d8d96066fd690fd3590d7cf8f5a6a7fd47f5494bbbba8

Observation a0ca87c4-de89-4b4b-af29-4e7be5e90b01 · outbound

This paper cites The Power of Normalization: Faster Evasion of Saddle Points.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping The Power of Normalization: Faster Evasion of Saddle Points

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.255061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.255061Z digest=sha256:2a6a529598db9198480f29b1497b4c45d4835d53a662c8d83dc7b9a7d906755f

Observation 132f5c5f-bb67-4e2a-80ea-35306b4bf7d1 · outbound

This paper cites Convex and non-convex optimization under generalized smoothness.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Convex and non-convex optimization under generalized smoothness

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.828707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.258995Z digest=sha256:d9963dde44a21ee75d06b67654949544b48a7994a639031f9cb0e077298ce985

Observation e3645a8f-7a53-4e4c-978f-e0e5c03231de · outbound

This paper cites Convergence of adam under relaxed assumptions.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Convergence of adam under relaxed assumptions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.817103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.262381Z digest=sha256:b52a3141ec6b4126a5a147be4408cd3e218956c83c8c5a2a9ae93ed8a4c38c61

Observation 81e46ea8-f965-42e2-9c13-137eec65ac64 · outbound

This paper cites High-probability bound for non-smooth non-convex stochastic optimization with heavy tails.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping High-probability bound for non-smooth non-convex stochastic optimization with heavy tails

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.806122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.266084Z digest=sha256:e8b3f48cc885c14c63ee876899fc1d5b4b3dec17e16a4ed26c983dd7fa199937

Observation 9b9e4e8e-941d-4737-9bfa-7d0161702bc8 · outbound

This paper cites Stochastic Nonsmooth Convex Optimization with Heavy-Tailed Noises: High-Probability Bound, In-Expectation Rate and Initial Distance Adaptation.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Stochastic Nonsmooth Convex Optimization with Heavy-Tailed Noises: High-Probability Bound, In-Expectation Rate and Initial Distance Adaptation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.269460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.269460Z digest=sha256:14c2df82c6564b3f1020dea88ddc1f47ed6a8b178c601f4d22afb30529bef21c

Observation 011f68f0-915a-4656-ba06-fdcf2690a83b · outbound

This paper cites Near-Optimal Non-Convex Stochastic Optimization under Generalized Smoothness.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Near-Optimal Non-Convex Stochastic Optimization under Generalized Smoothness

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:23:05.562657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.273144Z digest=sha256:4b130d304442521e1f4017b0c0f75081d458ffa53fcaa9dc2ea66d7ab1fb7139

Observation 80e220e5-3bbd-4626-8174-1ce74974eca4 · outbound

This paper cites Breaking the lower bound with (little) structure: Acceleration in non-convex stochastic optimization with heavy-tailed noise.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Breaking the lower bound with (little) structure: Acceleration in non-convex stochastic optimization with heavy-tailed noise

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.795062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.276793Z digest=sha256:85eef79751c970ac0c41b14354d0ca683f3ae92c35bd81c4d7a31e1adbe93490

Observation 8c848301-a694-4ceb-a6c3-bd1c76924c35 · outbound

This paper cites Brendan McMahan and Matthew J.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Brendan McMahan and Matthew J

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.783834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.280377Z digest=sha256:61d90814abb1a3780ab9a6a365ccaeb66657a5f7644ca09adae3817ddc569ad4

Observation 95a200f2-69dc-4234-98d9-d6fb1b0aa4cd · outbound

This paper cites Convergence of gradient descent on separable data.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Convergence of gradient descent on separable data

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.773141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.283937Z digest=sha256:ed9ba3d1b5071f71837d820c5b9ae726a48b4af66533f9fcbb3ddf5c051fc6f5

Observation 339e4444-5c6e-4323-ac52-d31237d0a8d7 · outbound

This paper cites Lectures on convex optimization, volume 137.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Lectures on convex optimization, volume 137

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.287926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.287926Z digest=sha256:3ea5d22356e5ce72571aedc6b572db3e751448cdea0dbb89a91c803a0164bc02

Observation fd650228-1fa1-4104-994f-d706340f71d2 · outbound

This paper cites Minimization methods for nonsmooth convex and quasiconvex functions.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Minimization methods for nonsmooth convex and quasiconvex functions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.291488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.291488Z digest=sha256:edc5e808d749f74de57ae78092994aa748d122c0fd38330459a292d36dc3e9c3

Observation e66b33ec-86c6-4d0d-a181-cb0598653329 · outbound

This paper cites Improved convergence in high probability of clipped gradient methods with heavy tailed noise.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Improved convergence in high probability of clipped gradient methods with heavy tailed noise

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.747536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.296096Z digest=sha256:7f41473436abde1e7bc2dd8bccf45269bf6ae63fc56b7c8e0813c9cdfc48b7b2

Observation 62e710e2-5c40-4598-b580-5d1b39161451 · outbound

This paper cites Breaking the heavy-tailed noise barrier in stochastic optimization problems.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Breaking the heavy-tailed noise barrier in stochastic optimization problems

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.735972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.299813Z digest=sha256:d8c18bea59f3d04808667b5ec0f0e036e12192af8583347975f5dd97a5d9828f

Observation 5b44bcd0-0498-4aee-ab37-f4294002ddf1 · outbound

This paper cites On equivalence of martingale tail bounds and deterministic regret inequalities.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping On equivalence of martingale tail bounds and deterministic regret inequalities

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.725144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.303409Z digest=sha256:de3135ecb2ec6f56407f842fb7284645dae43adaf68246e6f300fd5d9390fd06

Observation 1f43cce4-4eeb-450d-bbb9-3781d43cb2a0 · outbound

This paper cites A stochastic approximation method.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping A stochastic approximation method

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.307200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.307200Z digest=sha256:ec4cf168a10c65758b9a36dcddf1a9cb65bd8f97426d96003e1edd56b288840b

Observation 262ab119-b5b0-4809-9591-85ae4bcf5d8f · outbound

This paper cites High-probability bounds for stochastic optimization and variational inequalities: the case of unbounded variance.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping High-probability bounds for stochastic optimization and variational inequalities: the case of unbounded variance

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.707260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.310737Z digest=sha256:86f4cbf12e20fccdd93f3184255d6bf913f7743b989b0af60c35709a8d14b81d

Observation 13d035c6-4109-4359-8599-e2fe27fa8e5d · outbound

This paper cites On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.314264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.314264Z digest=sha256:2c19afbf3d6920d7c0b61c91650947f8cd8dfbdf78633bb03a93763114e96070

Observation d6f61eb8-111b-476c-acda-1c1211037043 · outbound

This paper cites A tail-index analysis of stochastic gradient noise in deep neural networks.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping A tail-index analysis of stochastic gradient noise in deep neural networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.696259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.318112Z digest=sha256:aea89c067a72aaf23440f655c5f44616a3be26e88ee839629259a783a0b51946

Observation 0afb4bcc-8ae6-46d0-b042-7fc16b8c506d · outbound

This paper cites Gradient normalization provably benefits nonconvex sgd under heavy-tailed noise.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Gradient normalization provably benefits nonconvex sgd under heavy-tailed noise

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.321472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.321472Z digest=sha256:c9bf93fdb7f552a41ff2b91362892ccf0b283a34ac65c4076478ec21f613e6e1

Observation d1a17774-fafa-4920-8f77-1e7301bea51e · outbound

This paper cites Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.325034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.325034Z digest=sha256:7bf10315a02c22f037ba8ac5a054da78adbdadbcdca0ec33a69b406d3561872f

Observation 30285a24-ec49-4c23-acc5-adf26fa2ae50 · outbound

This paper cites Convergence of adagrad for non-convex objectives: Simple proofs and relaxed assumptions.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Convergence of adagrad for non-convex objectives: Simple proofs and relaxed assumptions

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.679329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.328428Z digest=sha256:e58514a3266950d6fa67d0de18eccacd8f15547b25227d92c0aac9f90b6b15e7

Observation 3ea02deb-cf8d-4a4e-acb3-1c5831536187 · outbound

This paper cites On the Convergence of Adam under Non-uniform Smoothness: Separability from SGDM and Beyond.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping On the Convergence of Adam under Non-uniform Smoothness: Separability from SGDM and Beyond

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.332066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.332066Z digest=sha256:29961076682ecdb11e8a2627a60afadb9a34766b64532281dbea607974fb5e75

Observation 807b14e1-fd80-477c-832e-25c3c34487bf · outbound

This paper cites Large Batch Training of Convolutional Networks.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Large Batch Training of Convolutional Networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.335736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.335736Z digest=sha256:9655fa1a4cf9d83f23ca0ed5f0e36eaafae323e64bdd0cec89923ff903f90e7e

Observation 3b9e8fe4-1f53-4c19-af55-00a91c1fa4b4 · outbound

This paper cites Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.339526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.339526Z digest=sha256:21d2feab38c395412f218ebf7b116606bce0ba3ccd9fdf843ead35b93c5fe6f3

Observation dc5a44f9-b838-4056-9a91-5a5472e3f179 · outbound

This paper cites Improved analysis of clipping algorithms for non-convex optimization.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Improved analysis of clipping algorithms for non-convex optimization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.343597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.343597Z digest=sha256:aa349e47e71a03d36eda302377707d3de3cec534047835c81cdfe8073b347f34

Observation 41ec6a0c-d5c0-421d-b1dd-170400ff3825 · outbound

This paper cites Why gradient clipping accelerates training: A theoretical justification for adaptivity.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Why gradient clipping accelerates training: A theoretical justification for adaptivity

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.662393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.346964Z digest=sha256:cd6e95779f8544135f8743c0d6a03a2671cbbadc1081488ed6c1d34aefce2e04

Observation 1e9764ac-5c63-4251-b2a9-30925eb9bfb2 · outbound

This paper cites Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020 c.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020 c

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.350245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.350245Z digest=sha256:6282443c8e1339ed095432b5431da528ef3931452bdfae9903b5a98571943158

Observation b318596c-0bea-45b7-889a-c199c38b5511 · outbound

This paper cites Parameter-free regret in high probability with heavy tails.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Parameter-free regret in high probability with heavy tails

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.645301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.353870Z digest=sha256:9aaa9f524800428b33f2d85ff708314d71bac60046881a4ca9833a75e768b77a

Observation 60e2cdf6-d79d-4105-89e5-052ddd30fd9d · outbound

This paper cites @esa (Ref.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping @esa (Ref

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.357279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.357279Z digest=sha256:80e52965ae9ee5cea63476e4973af6397288fc12459d55869fc724ec5c42eb67

Observation e5e25932-8001-4f9d-a7f1-8691e04d1534 · outbound

This paper cites an unresolved cited work.

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T00:23:05.360942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:23:05.360942Z digest=sha256:19eb6d38da78ec93d147ba7c149868f93226d6be54fce0a75c9aeceeb9db0a36

Observation e6d7eb5e-8a71-4dee-8029-dab6d8102179 · outbound

This paper cites denotes the set of natural numbers (excluding 0 ).

Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping denotes the set of natural numbers (excluding 0 )

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:23:05.621655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T00:23:05.364651Z digest=sha256:123f83b6396677aa1e9769d6dc7efd335e81093251e1019701904fa39c14c191

Pith citing papers

Observation 411c530c-4a62-4b55-a653-5c317d1eed7e · inbound

Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives cites this paper.

Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:37:19.253844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T05:34:48.195468Z digest=sha256:bffeaa1b77be6b406532c902a757df8d29924dc0e46260d0f2693c68182cbdc8