Bayesian deep Q-learning exhibits a cold posterior effect, caused partly by misspecified Gaussian priors, and Laplace or meta-learned priors improve performance.
Exact Langevin Dynamics with Stochastic Gradients
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Stochastic gradient Markov Chain Monte Carlo algorithms are popular samplers for approximate inference, but they are generally biased. We show that many recent versions of these methods (e.g. Chen et al. (2014)) cannot be corrected using Metropolis-Hastings rejection sampling, because their acceptance probability is always zero. We can fix this by employing a sampler with realizable backwards trajectories, such as Gradient-Guided Monte Carlo (Horowitz, 1991), which generalizes stochastic gradient Langevin dynamics (Welling and Teh, 2011) and Hamiltonian Monte Carlo. We show that this sampler can be used with stochastic gradients, yielding nonzero acceptance probabilities, which can be computed even across multiple steps.
citation-role summary
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning
Bayesian deep Q-learning exhibits a cold posterior effect, caused partly by misspecified Gaussian priors, and Laplace or meta-learned priors improve performance.