HyPeR is a doubly robust policy-gradient estimator that uses secondary rewards to reduce variance when target rewards are only partially observed, with data-driven tuning of the mixing weight.
Jadidinejad, Craig Macdonald, and Iadh Ounis
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A General Framework for Off-Policy Learning with Partially-Observed Reward
HyPeR is a doubly robust policy-gradient estimator that uses secondary rewards to reduce variance when target rewards are only partially observed, with data-driven tuning of the mixing weight.