Pith. sign in

REVIEW 2 cited by

Optimal Conservative Offline RL with General Function Approximation via Augmented Lagrangian

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.00716 v1 pith:UZN5EQEY submitted 2022-11-01 cs.LG cs.AImath.OCmath.STstat.MLstat.TH

classification cs.LGcs.AImath.OCmath.STstat.MLstat.TH
keywords offlinealgorithmsapproximationaugmentedconservativefunctionlagrangianoptimal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Offline reinforcement learning (RL), which refers to decision-making from a previously-collected dataset of interactions, has received significant attention over the past years. Much effort has focused on improving offline RL practicality by addressing the prevalent issue of partial data coverage through various forms of conservative policy learning. While the majority of algorithms do not have finite-sample guarantees, several provable conservative offline RL algorithms are designed and analyzed within the single-policy concentrability framework that handles partial coverage. Yet, in the nonlinear function approximation setting where confidence intervals are difficult to obtain, existing provable algorithms suffer from computational intractability, prohibitively strong assumptions, and suboptimal statistical rates. In this paper, we leverage the marginalized importance sampling (MIS) formulation of RL and present the first set of offline RL algorithms that are statistically optimal and practical under general function approximation and single-policy concentrability, bypassing the need for uncertainty quantification. We identify that the key to successfully solving the sample-based approximation of the MIS problem is ensuring that certain occupancy validity constraints are nearly satisfied. We enforce these constraints by a novel application of the augmented Lagrangian method and prove the following result: with the MIS formulation, augmented Lagrangian is enough for statistically optimal offline RL. In stark contrast to prior algorithms that induce additional conservatism through methods such as behavior regularization, our approach provably eliminates this need and reinterprets regularizers as "enforcers of occupancy validity" than "promoters of conservatism."

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Quantile-Optimal Policy Learning under Unmeasured Confounding

    stat.ML 2025-06 conditional novelty 7.0 of 10

    Under instrumental-variable or negative-control assumptions, the authors prove a pessimism-based policy learning method achieves about 1/sqrt(n)-type regret for quantile reward objectives with unmeasured confounders.

  2. Reachability Weighted Offline Goal-conditioned Resampling

    cs.LG 2025-06 conditional novelty 6.0 of 10

    RWS trains a positive-unlabeled reachability classifier on goal-conditioned Q-values and uses it to re-weight goal sampling, improving offline goal-conditioned RL performance on robotic manipulation benchmarks.

Pith tools