Pith. sign in

Paper Citation Record · LEDGER

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM

As of 16 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:1908.00700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.00700 v2

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T15:43:22.439322Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5aaf2088-de28-4c79-9948-9c36f6e0f7bb · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:23.274561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.167183Z digest=sha256:18f8a7242293566684b210bc426b9a6c591d92ed4d4613252f0dd4e70d6dfe63

Observation 0342eb9f-8516-4047-8736-bd022bd3aeb1 · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:23.096998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.185759Z digest=sha256:b0fd3245b1afcc7213dd8c98f5312503c2e2e09b7ce722e860533188bc8f38db

Observation fbb9d947-bd00-4788-948e-692406ba250a · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:23.175626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.174778Z digest=sha256:a25bea507bd30de2a4b40c7f49d4a952ddfe9b80ccb6a0c0d4af26ab20cf65ca

Observation e51f7072-900a-4fdb-8d4c-fd152e6bbfb8 · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:23.362998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.138715Z digest=sha256:99631de69e71b73ca829142fc98d65186f7cb446b258a3b74ec0a60197dced6a

Observation 963466fd-420e-47b8-b973-f3945f801973 · outbound

This paper cites small learning rate dilemma.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM small learning rate dilemma

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:23.342676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.146510Z digest=sha256:6ac37bf616ae580e890b15ebef84feef903a09447af292ca6750f66f9e94f6c5

Observation be65fb01-931d-40d5-b8d5-ddc2834953a3 · outbound

This paper cites Besides the figures in main text, we have repeated experiments and show results as follows.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Besides the figures in main text, we have repeated experiments and show results as follows

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:23.322470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.152883Z digest=sha256:be0cf4ab2735e5b332b98509745042d86f9a638fd0ca2967685e9b440a0fe627

Observation 3b5256ea-f896-4f98-812d-33218f1af16f · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:23.294666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.159166Z digest=sha256:5529f2710de39976f76260c6c142329200c7a4dc7444951554704fe3559aee82

Observation 418484ea-8f8d-47f0-8fe3-8c66135330e2 · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:23.078499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.192714Z digest=sha256:78710f403b1327db9c4e70a8b895fa07a9d34ac88972ffcd15121cdb50a053ba

Observation 95b1d803-5af4-4a5c-bc71-3f68b94c1f48 · outbound

This paper cites With fixed L,σ,G,β 1, we have C1 =O(β2), C2 =O(dβ), C3 =O(dβ2).

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM With fixed L,σ,G,β 1, we have C1 =O(β2), C2 =O(dβ), C3 =O(dβ2)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:23.055754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.203070Z digest=sha256:6e499481d1261edf9339f9cdbc730b2927291b07bed5c93b32368bb603c474fd

Observation 54de26b7-62cc-48f7-9b69-3af6a0b8e669 · outbound

This paper cites By Jensen’s inequality, 1 T T∑ t=1 (f(xt)−f∗)≥f(¯xt)−f∗, where ¯xt = 1 T ∑T t=1xt.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM By Jensen’s inequality, 1 T T∑ t=1 (f(xt)−f∗)≥f(¯xt)−f∗, where ¯xt = 1 T ∑T t=1xt

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:23.012174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.223936Z digest=sha256:50ae4f4e03f4e4c1923ed33bd6833d13996d0b490b8a4e83f21c0e8ab597c09d

Observation e944d485-d18f-4297-82e5-1248aa07e38a · outbound

This paper cites f(¯xt)−f∗≤ D2 2µ1 √ T + β2 1d(σ2 +G2) (1−β1)2µ1T √ T (µ2 2−µ2.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM f(¯xt)−f∗≤ D2 2µ1 √ T + β2 1d(σ2 +G2) (1−β1)2µ1T √ T (µ2 2−µ2

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:22.989752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.232345Z digest=sha256:34a59cbca76041e6eccfba719b31e14d153ae093c0c4b632852b2aaa01f02733

Observation 8df3b6d7-c621-4655-b47c-c9e83e8b3689 · outbound

This paper cites 42 Remark 27 The leading item of convergence order of Adam should be O( ˜C√ T ), where ˜C = D2 2µ1 + µ2 2 µ1 (σ2 +G2) + β1d(σ2+G2) 2µ1(1−β1) (µ2 2−µ2.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM 42 Remark 27 The leading item of convergence order of Adam should be O( ˜C√ T ), where ˜C = D2 2µ1 + µ2 2 µ1 (σ2 +G2) + β1d(σ2+G2) 2µ1(1−β1) (µ2 2−µ2

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:22.966435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.241813Z digest=sha256:b978d82e0017b49b4ce5c0ec9e217527419fcf660b6236d3fb6cef73ca3470eb

Observation 608e165c-dacd-48cd-8630-623f0074dc8f · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:22.928083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.270233Z digest=sha256:7043cc2f0daf242a9e88b6e4f34ad2a4426a1c00a9a525a11f18c93d1b2644d4

Observation 8210dcf2-c3fc-41bd-93ff-44c07021f4f7 · outbound

This paper cites f(¯xt)−f∗≤ D2 2µ3 √ T + β2 1d(σ2 +G2) (1−β1)2µ3T √ T (µ2 4−µ2.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM f(¯xt)−f∗≤ D2 2µ3 √ T + β2 1d(σ2 +G2) (1−β1)2µ3T √ T (µ2 4−µ2

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:22.868436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.314762Z digest=sha256:6294750b772eb59b256d32d45d2882388c716323d6c6fb15f473460076fc959c

Observation 82fd7832-5332-4e68-b8fe-fcfcae8aaa1c · outbound

This paper cites For brevity, f(¯xt)−f∗ =O( 1√ T ).

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM For brevity, f(¯xt)−f∗ =O( 1√ T )

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:22.735074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.353595Z digest=sha256:a0d7c097c48b22327d74bcf425b3a12e7d422216ce522fc46481e4ce271bdf03

Observation 4fe05962-6b5d-45b7-9f3f-27ec2e0e46c6 · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:22.677674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.381079Z digest=sha256:1abec3619a7663075d4fef505ba18a5d7c643f23a7f3649d9bbda842f5b9f662

Observation 6b654323-6b9b-4c27-b712-d50fc0d93ac3 · outbound

This paper cites an unresolved cited work.

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:43:22.652059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.406449Z digest=sha256:99b3d02b004a9c0e46d28c4bce2826b8e55b04b2eb793459c167b686e676be4a

Observation 9b0abe13-6aed-49ba-a102-9d02fd4992c7 · outbound

This paper cites Set η =O( 1 T 2 ), E[f(xT +1)−f∗]≤ (1− 2λµ3 T 2 )TE[f(x1)−f∗] +O( 1 T ).

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM Set η =O( 1 T 2 ), E[f(xT +1)−f∗]≤ (1− 2λµ3 T 2 )TE[f(x1)−f∗] +O( 1 T )

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:22.552564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.439322Z digest=sha256:6b5316a2a5f35791af5ecbf82a44335c459aaa47e23253b2544e90b73b349bbc

Observation 61af7f28-a734-4c6f-ba88-84f5487363b8 · outbound

This paper cites 38 Remark 26 The leading item from the above convergence is C1/ √ T , β plays an essential role in the complexity, and a more accurate convergence should be O(βlog (1+eβ )√ T ).

Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM 38 Remark 26 The leading item from the above convergence is C1/ √ T , β plays an essential role in the complexity, and a more accurate convergence should be O(βlog (1+eβ )√ T )

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:43:23.036490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T15:43:22.214256Z digest=sha256:bb495bd57c5ad686260c35f37d614108a9f29977284768fc31f949fa99b8e6e0

Pith citing papers

No inbound Pith citation observations are available.