Foxtsage, a population-based learning-rate controller combined with SGD, reports 42% lower aggregate loss than Adam on three benchmarks, but with only about 1% accuracy gains and a 330% time increase.
A Comparison of Optimization Algorithms for Deep Learning
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In recent years, we have witnessed the rise of deep learning. Deep neural networks have proved their success in many areas. However, the optimization of these networks has become more difficult as neural networks going deeper and datasets becoming bigger. Therefore, more advanced optimization algorithms have been proposed over the past years. In this study, widely used optimization algorithms for deep learning are examined in detail. To this end, these algorithms called adaptive gradient methods are implemented for both supervised and unsupervised tasks. The behaviour of the algorithms during training and results on four image datasets, namely, MNIST, CIFAR-10, Kaggle Flowers and Labeled Faces in the Wild are compared by pointing out their differences against basic optimization algorithms.
citation-role summary
citation-polarity summary
fields
cs.NE 1years
2024 1verdicts
REJECT 1roles
baseline 1polarities
baseline 1representative citing papers
citing papers explorer
-
Foxtsage vs. Adam: Revolution or Evolution in Optimization?
Foxtsage, a population-based learning-rate controller combined with SGD, reports 42% lower aggregate loss than Adam on three benchmarks, but with only about 1% accuracy gains and a 330% time increase.