Pith. sign in

REVIEW 4 cited by

Finite Sample Analysis of Minimax Offline Reinforcement Learning: Completeness, Fast Rates and First-Order Efficiency

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.02981 v2 pith:GVY2YZ6J submitted 2021-02-05 cs.LG math.STstat.MLstat.TH

classification cs.LGmath.STstat.MLstat.TH
keywords completenessminimaxconvergenceefficiencyfastfirst-orderfunctionslearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

We offer a theoretical characterization of off-policy evaluation (OPE) in reinforcement learning using function approximation for marginal importance weights and $q$-functions when these are estimated using recent minimax methods. Under various combinations of realizability and completeness assumptions, we show that the minimax approach enables us to achieve a fast rate of convergence for weights and quality functions, characterized by the critical inequality \citep{bartlett2005}. Based on this result, we analyze convergence rates for OPE. In particular, we introduce novel alternative completeness conditions under which OPE is feasible and we present the first finite-sample result with first-order efficiency in non-tabular environments, i.e., having the minimal coefficient in the leading term.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression

    stat.ML 2025-01 conditional novelty 8.0 of 10

    DFIV achieves minimax-optimal rates in Besov spaces for nonparametric IV regression, and beats fixed-feature estimators for spatially inhomogeneous targets.

  2. Fitted Occupancy-Ratio Evaluation without Bellman Completeness

    stat.ML 2026-07 conditional novelty 7.0 of 10

    FORE estimates discounted occupancy ratios by iterating KL-projected adjoint Bellman updates, achieving convergence under ratio realizability alone without Bellman completeness.

  3. Quantile-Optimal Policy Learning under Unmeasured Confounding

    stat.ML 2025-06 conditional novelty 7.0 of 10

    Under instrumental-variable or negative-control assumptions, the authors prove a pessimism-based policy learning method achieves about 1/sqrt(n)-type regret for quantile reward objectives with unmeasured confounders.

  4. Semi-pessimistic Reinforcement Learning

    cs.LG 2025-05 reject novelty 6.0 of 10

    Semi-pessimistic pseudo labeling learns a pessimistic reward lower bound from labeled plus unlabeled data and uses it to train offline RL policies, with regret bounds under a weaker semi-coverage condition.

Pith tools