Backpropagation through the Void: Optimizing control variates for black-box gradient estimation
read the original abstract
Gradient-based optimization is the foundation of deep learning and reinforcement learning. Even when the mechanism being optimized is unknown or not differentiable, optimization using high-variance or biased gradient estimates is still often the best strategy. We introduce a general framework for learning low-variance, unbiased gradient estimators for black-box functions of random variables. Our method uses gradients of a neural network trained jointly with model parameters or policies, and is applicable in both discrete and continuous settings. We demonstrate this framework for training discrete latent-variable models. We also give an unbiased, action-conditional extension of the advantage actor-critic reinforcement learning algorithm.
This paper has not been read by Pith yet.
Forward citations
Cited by 2 Pith papers
-
Low-variance estimators overcome the phase-gradient bottleneck in complex-valued neural quantum states
Direct differentiation of the local energy at fixed samples yields an unbiased low-variance estimator for the variational Monte Carlo phase force in complex neural quantum states, with an adaptive mixture extending it...
-
Learning to Theorize the World from Observation
NEO induces compositional latent programs as world theories from observations and executes them to enable explanation-driven generalization.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.