Pith. sign in

REVIEW 1 cited by

Efficient Policy Evaluation with Offline Data Informed Behavior Policy Design

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.13734 v5 pith:YCCC3KRO submitted 2023-01-31 cs.LG

classification cs.LG
keywords policybehaviordatacarlodesignmonteofflineonline
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Most reinforcement learning practitioners evaluate their policies with online Monte Carlo estimators for either hyperparameter tuning or testing different algorithmic design choices, where the policy is repeatedly executed in the environment to get the average outcome. Such massive interactions with the environment are prohibitive in many scenarios. In this paper, we propose novel methods that improve the data efficiency of online Monte Carlo estimators while maintaining their unbiasedness. We first propose a tailored closed-form behavior policy that provably reduces the variance of an online Monte Carlo estimator. We then design efficient algorithms to learn this closed-form behavior policy from previously collected offline data. Theoretical analysis is provided to characterize how the behavior policy learning error affects the amount of reduced variance. Compared with previous works, our method achieves better empirical performance in a broader set of environments, with fewer requirements for offline data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Using imperfect counterfactual annotations only in the reward model part of a doubly robust estimator is the theoretically and empirically safest way to incorporate them.

Pith tools