DPS is a preference-based reinforcement learning algorithm that provably achieves asymptotic no-regret using posterior sampling and Bayesian linear regression credit assignment.
Then, asymptotically bound the one-sided regret rate forπi2 (Appendix A.2)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Dueling Posterior Sampling for Preference-Based Reinforcement Learning
DPS is a preference-based reinforcement learning algorithm that provably achieves asymptotic no-regret using posterior sampling and Bayesian linear regression credit assignment.