Using imperfect counterfactual annotations only in the reward model part of a doubly robust estimator is the theoretically and empirically safest way to incorporate them.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation
Using imperfect counterfactual annotations only in the reward model part of a doubly robust estimator is the theoretically and empirically safest way to incorporate them.