A trajectory-wise control variate estimator removes both action-level and future-trajectory variance in policy gradients, and the natural time-ordering is proven optimal under exact critic assumptions.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods
A trajectory-wise control variate estimator removes both action-level and future-trajectory variance in policy gradients, and the natural time-ordering is proven optimal under exact critic assumptions.