REVIEW 2 cited by
Meta Mirror Descent: Optimiser Learning for Fast Convergence
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Optimisers are an essential component for training machine learning models, and their design influences learning speed and generalisation. Several studies have attempted to learn more effective gradient-descent optimisers via solving a bi-level optimisation problem where generalisation error is minimised with respect to optimiser parameters. However, most existing optimiser learning methods are intuitively motivated, without clear theoretical support. We take a different perspective starting from mirror descent rather than gradient descent, and meta-learning the corresponding Bregman divergence. Within this paradigm, we formalise a novel meta-learning objective of minimising the regret bound of learning. The resulting framework, termed Meta Mirror Descent (MetaMD), learns to accelerate optimisation speed. Unlike many meta-learned optimisers, it also supports convergence and generalisation guarantees and uniquely does so without requiring validation data. We evaluate our framework on a variety of tasks and architectures in terms of convergence rate and generalisation error and demonstrate strong performance.
Forward citations
Cited by 2 Pith papers
-
Non-Linear Strategic Classification Made Practical
A Lagrangian duality method approximates best responses for non-linear strategic classification and enables gradient-based training via the Implicit Function Theorem, yielding improved strategic accuracy on standard datasets.
-
Learnable Loss Geometries with Mirror Descent for Scalable and Convergent Meta-Learning
A meta-learning method learns a neural-network distance-generating function for mirror descent, matching or beating preconditioned baselines on few-shot image classification while providing an O(1/epsilon^2) convergence rate.
Discussion (0). Sign in to comment.