Pith. sign in

Paper Citation Record · LEDGER

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM

As of 23 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:1908.00700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.00700 v2

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T15:43:22.439322Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5aaf2088-de28-4c79-9948-9c36f6e0f7bb · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:23.274561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.167183Z digest=sha256:d0e0d2b4a25ad2f2780c359039f96cbf9e750d7d80dfcb16b0b837c1179c52fd

Observation 0342eb9f-8516-4047-8736-bd022bd3aeb1 · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:23.096998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.185759Z digest=sha256:2a1de6eb3145baa18465803068e436815bd5ae108194c6ec284e43ffc4bacd84

Observation fbb9d947-bd00-4788-948e-692406ba250a · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:23.175626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.174778Z digest=sha256:63481bb4bb96b95044f64fa4c0f7722116ab803a059d2785927b08e0f8295112

Observation e51f7072-900a-4fdb-8d4c-fd152e6bbfb8 · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:23.362998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.138715Z digest=sha256:4a2f3f13b92abf99c4eb72d45a1f6259035d066f51ecefbbf9d38107e1045f7b

Observation 963466fd-420e-47b8-b973-f3945f801973 · outbound

This paper cites small learning rate dilemma.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM small learning rate dilemma

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:23.342676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.146510Z digest=sha256:f78d38d88bfc4436a51cdfddbe270a22c444bb491319ce4475e022c151a404f0

Observation be65fb01-931d-40d5-b8d5-ddc2834953a3 · outbound

This paper cites Besides the figures in main text, we have repeated experiments and show results as follows.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Besides the figures in main text, we have repeated experiments and show results as follows

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:23.322470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.152883Z digest=sha256:bb1e24b3f69a861bf1c51a7a7dd5d2ea593cfb7b1050f2cb353be162f9dd7df0

Observation 3b5256ea-f896-4f98-812d-33218f1af16f · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:23.294666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.159166Z digest=sha256:5f2900214d5ad95372045e0f854ab2705946bcb18417bca9063209d4e18471ee

Observation 418484ea-8f8d-47f0-8fe3-8c66135330e2 · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:23.078499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.192714Z digest=sha256:56ada3a30549ee428b9d613133bed8bed5b526cb6ea6b9bb274185f86d1b173b

Observation 95b1d803-5af4-4a5c-bc71-3f68b94c1f48 · outbound

This paper cites With fixed L,σ,G,β 1, we have C1 =O(β2), C2 =O(dβ), C3 =O(dβ2).

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM With fixed L,σ,G,β 1, we have C1 =O(β2), C2 =O(dβ), C3 =O(dβ2)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:23.055754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.203070Z digest=sha256:503e2cba9c37c0233287169269fc9e0170df6b7d8d2aacf666bbb6acc46c7869

Observation 54de26b7-62cc-48f7-9b69-3af6a0b8e669 · outbound

This paper cites By Jensen’s inequality, 1 T T∑ t=1 (f(xt)−f∗)≥f(¯xt)−f∗, where ¯xt = 1 T ∑T t=1xt.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM By Jensen’s inequality, 1 T T∑ t=1 (f(xt)−f∗)≥f(¯xt)−f∗, where ¯xt = 1 T ∑T t=1xt

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:23.012174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.223936Z digest=sha256:1384268b1edf854a162f301fb8425fb767191356520f040ce301ce786fbbcff1

Observation e944d485-d18f-4297-82e5-1248aa07e38a · outbound

This paper cites f(¯xt)−f∗≤ D2 2µ1 √ T + β2 1d(σ2 +G2) (1−β1)2µ1T √ T (µ2 2−µ2.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM f(¯xt)−f∗≤ D2 2µ1 √ T + β2 1d(σ2 +G2) (1−β1)2µ1T √ T (µ2 2−µ2

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:22.989752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.232345Z digest=sha256:3beda05208e0b39a241da290cec4df4c467468e6e6391f0adb81dbe3966aff03

Observation 8df3b6d7-c621-4655-b47c-c9e83e8b3689 · outbound

This paper cites 42 Remark 27 The leading item of convergence order of Adam should be O( ˜C√ T ), where ˜C = D2 2µ1 + µ2 2 µ1 (σ2 +G2) + β1d(σ2+G2) 2µ1(1−β1) (µ2 2−µ2.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM 42 Remark 27 The leading item of convergence order of Adam should be O( ˜C√ T ), where ˜C = D2 2µ1 + µ2 2 µ1 (σ2 +G2) + β1d(σ2+G2) 2µ1(1−β1) (µ2 2−µ2

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:22.966435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.241813Z digest=sha256:23ee4440ad5b78df2294e092198bab24d56dd92934bf1745da7fef4fd56f5c39

Observation 608e165c-dacd-48cd-8630-623f0074dc8f · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:22.928083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.270233Z digest=sha256:df5d02f795dd3405e8e9744fcbb58d9ad57e773d3d65f63b44d80f4429b074e7

Observation 8210dcf2-c3fc-41bd-93ff-44c07021f4f7 · outbound

This paper cites f(¯xt)−f∗≤ D2 2µ3 √ T + β2 1d(σ2 +G2) (1−β1)2µ3T √ T (µ2 4−µ2.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM f(¯xt)−f∗≤ D2 2µ3 √ T + β2 1d(σ2 +G2) (1−β1)2µ3T √ T (µ2 4−µ2

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:22.868436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.314762Z digest=sha256:b4c9f6461c0202b21b06746e5f82b5e1deebf97b8538c61cc89ab256bc642991

Observation 82fd7832-5332-4e68-b8fe-fcfcae8aaa1c · outbound

This paper cites For brevity, f(¯xt)−f∗ =O( 1√ T ).

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM For brevity, f(¯xt)−f∗ =O( 1√ T )

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:22.735074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.353595Z digest=sha256:c8753c9194ee001fa54d542a8151e5c1e1dc841f4f3084af13082be9daa24dff

Observation 4fe05962-6b5d-45b7-9f3f-27ec2e0e46c6 · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:22.677674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.381079Z digest=sha256:3fb1b4168957ef9d74bdfe12b584b939029d6f4596aa4a376080e6180dd250f3

Observation 6b654323-6b9b-4c27-b712-d50fc0d93ac3 · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:22.652059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.406449Z digest=sha256:bb1024ceff7ab77c032dba336017d8b9a335fd689a45af52b1116c3db6ebc7d6

Observation 9b0abe13-6aed-49ba-a102-9d02fd4992c7 · outbound

This paper cites Set η =O( 1 T 2 ), E[f(xT +1)−f∗]≤ (1− 2λµ3 T 2 )TE[f(x1)−f∗] +O( 1 T ).

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Set η =O( 1 T 2 ), E[f(xT +1)−f∗]≤ (1− 2λµ3 T 2 )TE[f(x1)−f∗] +O( 1 T )

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:22.552564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.439322Z digest=sha256:695b22b21a4a3b6edb7dcc0445ff34b53820cb0402f34a1dd3d0017c684a623b

Observation 61af7f28-a734-4c6f-ba88-84f5487363b8 · outbound

This paper cites 38 Remark 26 The leading item from the above convergence is C1/ √ T , β plays an essential role in the complexity, and a more accurate convergence should be O(βlog (1+eβ )√ T ).

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM 38 Remark 26 The leading item from the above convergence is C1/ √ T , β plays an essential role in the complexity, and a more accurate convergence should be O(βlog (1+eβ )√ T )

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:23.036490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T15:43:22.214256Z digest=sha256:12c0cc4d0e71642f4ef3a07c319fc65bd754e74fce362e95f3bcc2f2f09b72d3

Pith citing papers

No inbound Pith citation observations are available.