A reparameterization of the inner-loop update makes second-order meta-gradients computable via mixed-mode autodiff, giving over 10x memory savings with up to 25% wall-clock time savings.
Gradient-based optimization of hyperparameters
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Scalable Meta-Learning via Mixed-Mode Differentiation
A reparameterization of the inner-loop update makes second-order meta-gradients computable via mixed-mode autodiff, giving over 10x memory savings with up to 25% wall-clock time savings.