Pith. sign in

REVIEW 1 cited by

Doubly Optimal Policy Evaluation for Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.02226 v2 pith:5DAHOLWH submitted 2024-10-03 cs.LG

classification cs.LG
keywords policyevaluationvariancedatamethodoptimaldata-collectingdata-processing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Policy evaluation estimates the performance of a policy by (1) collecting data from the environment and (2) processing raw data into a meaningful estimate. Due to the sequential nature of reinforcement learning, any improper data-collecting policy or data-processing method substantially deteriorates the variance of evaluation results over long time steps. Thus, policy evaluation often suffers from large variance and requires massive data to achieve the desired accuracy. In this work, we design an optimal combination of data-collecting policy and data-processing baseline. Theoretically, we prove our doubly optimal policy evaluation method is unbiased and guaranteed to have lower variance than previously best-performing methods. Empirically, compared with previous works, we show our method reduces variance substantially and achieves superior empirical performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Estimating the behavior policy from longer histories provably reduces the asymptotic variance of importance-sampling based off-policy evaluation estimators at the cost of increased finite-sample bias, with different e...

Pith tools