The authors derive mirror descent and mirror-less updates using the Tempesta generalized logarithm as the link function, with an approximate inverse exponential obtained via Lagrange inversion.
Exponentiated Gradient Meets Gradient Descent
1 Pith paper cite this work, alongside 10 external citations. Polarity classification is still indexing.
abstract
The (stochastic) gradient descent and the multiplicative update method are probably the most popular algorithms in machine learning. We introduce and study a new regularization which provides a unification of the additive and multiplicative updates. This regularization is derived from an hyperbolic analogue of the entropy function, which we call hypentropy. It is motivated by a natural extension of the multiplicative update to negative numbers. The hypentropy has a natural spectral counterpart which we use to derive a family of matrix-based updates that bridge gradient methods and the multiplicative method for matrices. While the latter is only applicable to positive semi-definite matrices, the spectral hypentropy method can naturally be used with general rectangular matrices. We analyze the new family of updates by deriving tight regret bounds. We study empirically the applicability of the new update for settings such as multiclass learning, in which the parameters constitute a general rectangular matrix.
citation-role summary
citation-polarity summary
fields
stat.ML 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Mirror Descent Using the Tempesta Generalized Multi-parametric Logarithms
The authors derive mirror descent and mirror-less updates using the Tempesta generalized logarithm as the link function, with an approximate inverse exponential obtained via Lagrange inversion.