Pith. sign in

REVIEW 4 cited by

Off-policy evaluation beyond overlap: partial identification through smoothness

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.11812 v2 pith:WPRYZOHO submitted 2023-05-19 stat.ME math.STstat.TH

Off-policy evaluation beyond overlap: partial identification through smoothness

classification stat.ME math.STstat.TH
keywords off-policysmoothnessunderassumptionsidentificationoverlappartialpolicy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Off-policy evaluation (OPE) is the problem of estimating the value of a target policy using historical data collected under a different logging policy. OPE methods typically assume overlap between the target and logging policy, enabling solutions based on importance weighting and/or imputation. In this work, we approach OPE without assuming either overlap or a well-specified model by considering a strategy based on partial identification under non-parametric assumptions on the conditional mean function, focusing especially on Lipschitz smoothness. Under such smoothness assumptions, we formulate a pair of linear programs whose optimal values upper and lower bound the contributions of the no-overlap region to the off-policy value. We show that these linear programs have a concise closed form solution that can be computed efficiently and that their solutions converge, under the Lipschitz assumption, to the sharp partial identification bounds on the off-policy value. Furthermore, we show that the rate of convergence is minimax optimal, up to log factors. We deploy our methods on two semi-synthetic examples, and obtain informative and valid bounds that are tighter than those possible without smoothness assumptions.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Triage Score: A Counterfactual Risk Assessment Instrument

    stat.AP 2026-06 unverdicted novelty 7.0

    Triage scores extend risk scores via additive counterfactual utilities to incorporate intervention effects in high-stakes decisions.

  2. Wasserstein Policy Learning for Distributional Outcomes

    stat.ME 2026-06 unverdicted novelty 7.0

    Establishes finite-sample regret bounds of order sqrt(N-dim(Π)/N) for IPW and DR estimators in Wasserstein policy learning with distributional outcomes, plus a matching minimax lower bound.

  3. Logging Policy Design for Off-Policy Evaluation

    stat.ML 2026-05 unverdicted novelty 7.0

    Derives optimal logging policies for off-policy evaluation by balancing reward concentration against action coverage in known, unknown, and partially known regimes of target policy and rewards.

  4. Logging Policy Design for Off-Policy Evaluation

    stat.ML 2026-05 unverdicted novelty 5.0

    Derives optimal logging policies for minimizing off-policy evaluation error under known, unknown, and partially known target policies and reward distributions.