Pith. sign in

REVIEW 1 cited by

Soft-Robust Algorithms for Batch Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.14495 v2 pith:ROTBDD6W submitted 2020-11-30 cs.LG cs.AImath.OCstat.ML

classification cs.LGcs.AImath.OCstat.ML
keywords criterionalgorithmsoptimizepercentilesoft-robustconservativelearningmean
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In reinforcement learning, robust policies for high-stakes decision-making problems with limited data are usually computed by optimizing the percentile criterion, which minimizes the probability of a catastrophic failure. Unfortunately, such policies are typically overly conservative as the percentile criterion is non-convex, difficult to optimize, and ignores the mean performance. To overcome these shortcomings, we study the soft-robust criterion, which uses risk measures to balance the mean and percentile criterion better. In this paper, we establish the soft-robust criterion's fundamental properties, show that it is NP-hard to optimize, and propose and analyze two algorithms to approximately optimize it. Our theoretical analyses and empirical evaluations demonstrate that our algorithms compute much less conservative solutions than the existing approximate methods for optimizing the percentile-criterion.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    PhyB averages over the k worst dynamics models with entropy-weighted coefficients and uses Bregman-regularized policy iteration; it claims bounded pessimism, monotonic improvement, and top D4RL scores.

Pith tools