REVIEW 4 cited by
Finite Sample Analysis of Minimax Offline Reinforcement Learning: Completeness, Fast Rates and First-Order Efficiency
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
We offer a theoretical characterization of off-policy evaluation (OPE) in reinforcement learning using function approximation for marginal importance weights and $q$-functions when these are estimated using recent minimax methods. Under various combinations of realizability and completeness assumptions, we show that the minimax approach enables us to achieve a fast rate of convergence for weights and quality functions, characterized by the critical inequality \citep{bartlett2005}. Based on this result, we analyze convergence rates for OPE. In particular, we introduce novel alternative completeness conditions under which OPE is feasible and we present the first finite-sample result with first-order efficiency in non-tabular environments, i.e., having the minimal coefficient in the leading term.
Forward citations
Cited by 4 Pith papers
-
Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression
DFIV achieves minimax-optimal rates in Besov spaces for nonparametric IV regression, and beats fixed-feature estimators for spatially inhomogeneous targets.
-
Fitted Occupancy-Ratio Evaluation without Bellman Completeness
FORE estimates discounted occupancy ratios by iterating KL-projected adjoint Bellman updates, achieving convergence under ratio realizability alone without Bellman completeness.
-
Quantile-Optimal Policy Learning under Unmeasured Confounding
Under instrumental-variable or negative-control assumptions, the authors prove a pessimism-based policy learning method achieves about 1/sqrt(n)-type regret for quantile reward objectives with unmeasured confounders.
-
Semi-pessimistic Reinforcement Learning
Semi-pessimistic pseudo labeling learns a pessimistic reward lower bound from labeled plus unlabeled data and uses it to train offline RL policies, with regret bounds under a weaker semi-coverage condition.
Discussion (0). Continue with ORCID to comment.