Pith. sign in

REVIEW 2 cited by

Meta Mirror Descent: Optimiser Learning for Fast Convergence

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.02711 v1 pith:Q63OFP5R submitted 2022-03-05 cs.LG math.OC

classification cs.LGmath.OC
keywords learningdescentgeneralisationconvergencemirroroptimiseroptimiserserror
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Optimisers are an essential component for training machine learning models, and their design influences learning speed and generalisation. Several studies have attempted to learn more effective gradient-descent optimisers via solving a bi-level optimisation problem where generalisation error is minimised with respect to optimiser parameters. However, most existing optimiser learning methods are intuitively motivated, without clear theoretical support. We take a different perspective starting from mirror descent rather than gradient descent, and meta-learning the corresponding Bregman divergence. Within this paradigm, we formalise a novel meta-learning objective of minimising the regret bound of learning. The resulting framework, termed Meta Mirror Descent (MetaMD), learns to accelerate optimisation speed. Unlike many meta-learned optimisers, it also supports convergence and generalisation guarantees and uniquely does so without requiring validation data. We evaluate our framework on a variety of tasks and architectures in terms of convergence rate and generalisation error and demonstrate strong performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Non-Linear Strategic Classification Made Practical

    cs.GT 2026-06 unverdicted novelty 6.0 of 10

    A Lagrangian duality method approximates best responses for non-linear strategic classification and enables gradient-based training via the Implicit Function Theorem, yielding improved strategic accuracy on standard datasets.

  2. Learnable Loss Geometries with Mirror Descent for Scalable and Convergent Meta-Learning

    cs.LG 2025-09 conditional novelty 5.0 of 10

    A meta-learning method learns a neural-network distance-generating function for mirror descent, matching or beating preconditioned baselines on few-shot image classification while providing an O(1/epsilon^2) convergence rate.

Pith tools