Under a new sum-of-distances cost function, gradient descent step sizes are claimed (C+ε,δ)-learnable with O~(H^3/ε^2) samples and a momentum-based two-parameter method with O~(H^4/ε^2) samples.
W.; Pfau, D.; Schaul, T.; Shillingford, B.; and de Freitas, N
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
math.OC 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Learning complexity of gradient descent and conjugate gradient algorithms
Under a new sum-of-distances cost function, gradient descent step sizes are claimed (C+ε,δ)-learnable with O~(H^3/ε^2) samples and a momentum-based two-parameter method with O~(H^4/ε^2) samples.