Pith. sign in

Paper Citation Record · LEDGER

Toward a Unified Theory of Gradient Descent under Generalized Smoothness

As of 20 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 3 inbound Pith citation observations for arXiv:2412.11773.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11773 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:47:46.392025Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:05:23.764428Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T14:46:01.413229Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28a837b3-8d0f-4e12-9adc-e796e768e00e · outbound

This paper cites For any M ≥ 0, taking ¯T (M ) such that ∥∇f (x ¯T )∥ ≤M, we get f (xT ) − f (x∗) ≤ ℓ(2M ) ∥x0 − x∗∥2 2(T − ¯T (M ) + 1).

Toward a Unified Theory of Gradient Descent under Generalized Smoothness For any M ≥ 0, taking ¯T (M ) such that ∥∇f (x ¯T )∥ ≤M, we get f (xT ) − f (x∗) ≤ ℓ(2M ) ∥x0 − x∗∥2 2(T − ¯T (M ) + 1)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.527885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:47:46.392025Z digest=sha256:130d48acdbf0a747c49c3ad378101bfb1914978957ec8e61b2184645a7022fc2

Observation 24cac4c9-6c86-493e-abd1-81a08474564f · outbound

This paper cites This function is (3.3, 1)–smooth, meaning we can run Algorithm 1 with ℓ(s) = 3.3 + s.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness This function is (3.3, 1)–smooth, meaning we can run Algorithm 1 with ℓ(s) = 3.3 + s

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T14:47:46.583136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:47:46.374599Z digest=sha256:1c19a658be8cf90b4033f566e64d50bdb7ced6ec8b5a6fb9e0acb928f4d0f491

Observation 94b02397-078e-4add-bcde-fa63625f7551 · outbound

This paper cites Large Deviations of Vector-valued Martingales in 2-Smooth Normed Spaces.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Large Deviations of Vector-valued Martingales in 2-Smooth Normed Spaces

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.340297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.340297Z digest=sha256:db1bb7d188d6ec7e18ac79c6e4e8a38f6c27bdb10d4a6ff7b313832eea63c086

Observation 798fd856-6e30-4944-bdbe-37cc476a9141 · outbound

This paper cites Federated Learning: Strategies for Improving Communication Efficiency.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Federated Learning: Strategies for Improving Communication Efficiency

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.344761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.344761Z digest=sha256:1fde01fc87d007271ea33bfd92b51d6f486a5063b34d3c444b820cffb22a2ce9

Observation e65d4d06-6939-4437-879b-52de45629bef · outbound

This paper cites Optimizing $(L_0, L_1)$-Smooth Functions by Gradient Methods.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Optimizing $(L_0, L_1)$-Smooth Functions by Gradient Methods

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.353647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.353647Z digest=sha256:1bbc99f3ab8942517dcc3d0a623c37f5df8b99c2d6f9ce13d9ddbe1c35ca182c

Observation 396e8e86-8831-423d-8707-c84d8eb55bd0 · outbound

This paper cites Gradient-Variation Online Learning under Generalized Smoothness.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Gradient-Variation Online Learning under Generalized Smoothness

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:47:46.446044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:47:46.358018Z digest=sha256:bc31d5b366d7adba2c122a46a7201ad3b70139cdbfcdf87a95c3428b457bfa2b

Observation 96fe0dda-9a44-4369-a088-4427ece666d3 · outbound

This paper cites Why gradient clipping accelerates training: A theoretical justification for adaptivity.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Why gradient clipping accelerates training: A theoretical justification for adaptivity

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.361895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.361895Z digest=sha256:66926e4812a7da95f0f2aab2b3b701cb97096f40f9d5afd0b7cbb9822c8699d7

Observation 4eb5ac19-5ad3-4592-93e8-d45fd3bd0718 · outbound

This paper cites Next, we take the step size γk = 1/(800 + 2(2f ′(x0))2) from (Li et al., 2024a) and observe that GD requires at least 20.000 iterations because f ′(x0) is huge.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Next, we take the step size γk = 1/(800 + 2(2f ′(x0))2) from (Li et al., 2024a) and observe that GD requires at least 20.000 iterations because f ′(x0) is huge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.595659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:47:46.370237Z digest=sha256:8c326faec6da0e56b1e42cb6d09cd49a2454c28b6661d08f5cb7514854723d69

Observation de1bb650-520f-4772-b690-52acced1c794 · outbound

This paper cites Using the standard differential algebra, we can solve it: dg(t) ℓ(∥∇f (x)∥ + g(t)) = dt ⇒ Z t 0 dg(v) ℓ(∥∇f (x)∥ + g(v)) = t ⇒ Z g(t) 0 dv ℓ(∥∇f (x)∥ + v) = t.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Using the standard differential algebra, we can solve it: dg(t) ℓ(∥∇f (x)∥ + g(t)) = dt ⇒ Z t 0 dg(v) ℓ(∥∇f (x)∥ + g(v)) = t ⇒ Z g(t) 0 dv ℓ(∥∇f (x)∥ + v) = t

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.569462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:47:46.378443Z digest=sha256:9f115587ff0050ec16e5d6f90f13a7d4fdc28918840ac0963aea98ec6dc88c43

Observation 43edb0f4-f171-4a37-9873-2c0ff80b75b9 · outbound

This paper cites an unresolved cited work.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:47:46.555804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:47:46.383119Z digest=sha256:3cb9232a5120f28285b5e54c70285c469f0d0c08bd74e6717d1150f02927b0fa

Observation 993fdb7d-0f8e-4e28-a9d8-7bc75035d8a6 · outbound

This paper cites Due to the strategy from Alg.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Due to the strategy from Alg

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.542435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:47:46.387503Z digest=sha256:87dc75e13cac3e5a451df19af70a64ad2ee1ab4df00f070f109288a8a263ed2f

Observation 0892a7b6-1813-4d05-ad69-5c4ba1b58b0a · outbound

This paper cites Parameter-free Clipped Gradient Descent Meets Polyak.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Parameter-free Clipped Gradient Descent Meets Polyak

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.349125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.349125Z digest=sha256:094cbc3daea7cfb5a464c7566824c2b8e11188c2c327f71161ff4b5d36f0aea2

Observation 9dd21d60-127c-4ee0-b65e-54bccbb5aa1a · outbound

This paper cites an unresolved cited work.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:47:46.608869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:47:46.366168Z digest=sha256:8a1750cb5ca22a7abde33e32aca0c3286ce67beca19b4d19e2cb88f8c0b9dee8

Observation 8dd84dcc-81f2-49d3-a939-28321be14537 · outbound

This paper cites Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.330224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.330224Z digest=sha256:d4ea28caf7cbdbf7a65ee4ece0744b0e7322827f5e86502ac59965fadce7a836

Observation d5a437c0-4c84-4287-b832-013d350911c3 · outbound

This paper cites A theoretical study of the(l 0, l1)-smoothness condition in deep learning.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness A theoretical study of the(l 0, l1)-smoothness condition in deep learning

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:47:46.621303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T14:47:46.324863Z digest=sha256:5f7da3fb44228a25cab07fbcb97a77a645c5f329ca4c80fc1f043dea10cc8a34

Observation dac0dda4-a466-4f9c-9efa-70e19b249496 · outbound

This paper cites Accelerated Objective Gap and Gradient Norm Convergence for Gradient Descent via Long Steps.

Toward a Unified Theory of Gradient Descent under Generalized Smoothness Accelerated Objective Gap and Gradient Norm Convergence for Gradient Descent via Long Steps

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:46.334977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:47:46.334977Z digest=sha256:621959d360fb6e2c14ca1af5607cf325034991ca731a08267579ece318c59d7b

Pith citing papers

Observation 596c37ec-7556-4be7-847c-315070f1c195 · inbound

Gradient-Normalized Smoothness for Optimization with Approximate Hessians cites this paper.

Gradient-Normalized Smoothness for Optimization with Approximate Hessians Toward a Unified Theory of Gradient Descent under Generalized Smoothness

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T20:05:23.764428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:05:23.764428Z digest=sha256:54f1d4f5e1243898a0146adc840426d085bd64b54b59414eb2e0baedabdb8e9b

Observation 39d2a35e-c9bc-46b6-852e-c18080516614 · inbound

Revisiting Randomized Smoothing: Nonsmooth Nonconvex Optimization Beyond Global Lipschitz Continuity cites this paper.

Revisiting Randomized Smoothing: Nonsmooth Nonconvex Optimization Beyond Global Lipschitz Continuity Toward a Unified Theory of Gradient Descent under Generalized Smoothness

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.388874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.388874Z digest=sha256:4badf631452c66862d0a4ae981d496058f6bdac3c870f2842d2b3be0fb5deb28

Observation 520b9d01-111f-4ab0-9ad5-527b2001e49e · inbound

A Few Accelerated Algorithms for Convex Optimization under $(H_0,H_1)$-Smoothness cites this paper.

A Few Accelerated Algorithms for Convex Optimization under $(H_0,H_1)$-Smoothness Toward a Unified Theory of Gradient Descent under Generalized Smoothness

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:46:01.418931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T14:46:00.979837Z digest=sha256:1506b47ecfbf894c3b98fd0c01cad2c4dc4fd9eb651ee762117a6fbc43bdde65