REVIEW 4 cited by
Off-policy evaluation beyond overlap: partial identification through smoothness
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Off-policy evaluation beyond overlap: partial identification through smoothness
read the original abstract
Off-policy evaluation (OPE) is the problem of estimating the value of a target policy using historical data collected under a different logging policy. OPE methods typically assume overlap between the target and logging policy, enabling solutions based on importance weighting and/or imputation. In this work, we approach OPE without assuming either overlap or a well-specified model by considering a strategy based on partial identification under non-parametric assumptions on the conditional mean function, focusing especially on Lipschitz smoothness. Under such smoothness assumptions, we formulate a pair of linear programs whose optimal values upper and lower bound the contributions of the no-overlap region to the off-policy value. We show that these linear programs have a concise closed form solution that can be computed efficiently and that their solutions converge, under the Lipschitz assumption, to the sharp partial identification bounds on the off-policy value. Furthermore, we show that the rate of convergence is minimax optimal, up to log factors. We deploy our methods on two semi-synthetic examples, and obtain informative and valid bounds that are tighter than those possible without smoothness assumptions.
Forward citations
Cited by 4 Pith papers
-
Triage Score: A Counterfactual Risk Assessment Instrument
Triage scores extend risk scores via additive counterfactual utilities to incorporate intervention effects in high-stakes decisions.
-
Wasserstein Policy Learning for Distributional Outcomes
Establishes finite-sample regret bounds of order sqrt(N-dim(Π)/N) for IPW and DR estimators in Wasserstein policy learning with distributional outcomes, plus a matching minimax lower bound.
-
Logging Policy Design for Off-Policy Evaluation
Derives optimal logging policies for off-policy evaluation by balancing reward concentration against action coverage in known, unknown, and partially known regimes of target policy and rewards.
-
Logging Policy Design for Off-Policy Evaluation
Derives optimal logging policies for minimizing off-policy evaluation error under known, unknown, and partially known target policies and reward distributions.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.