Pith. sign in

REVIEW 3 cited by

Soft-Robust Algorithms for Batch Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.14495 v2 pith:ROTBDD6W submitted 2020-11-30 cs.LG cs.AImath.OCstat.ML

classification cs.LGcs.AImath.OCstat.ML
keywords criterionalgorithmsoptimizepercentilesoft-robustconservativelearningmean
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In reinforcement learning, robust policies for high-stakes decision-making problems with limited data are usually computed by optimizing the percentile criterion, which minimizes the probability of a catastrophic failure. Unfortunately, such policies are typically overly conservative as the percentile criterion is non-convex, difficult to optimize, and ignores the mean performance. To overcome these shortcomings, we study the soft-robust criterion, which uses risk measures to balance the mean and percentile criterion better. In this paper, we establish the soft-robust criterion's fundamental properties, show that it is NP-hard to optimize, and propose and analyze two algorithms to approximately optimize it. Our theoretical analyses and empirical evaluations demonstrate that our algorithms compute much less conservative solutions than the existing approximate methods for optimizing the percentile-criterion.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty

    cs.LG 2025-06 unverdicted novelty 7.0 of 10

    DR-SAC is the first actor-critic distributionally robust RL algorithm for offline continuous control that derives a convergent robust soft policy iteration and reports up to 9.8x higher rewards than SAC under perturbations.

  2. Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    PhyB approximates Bayesian expectations in offline RL as convex combinations over dynamics model subsets with bounded discrepancy, enabling regularized policy optimization with monotonic improvement guarantees.

  3. Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief

    cs.AI 2026-05 reject novelty 6.0 of 10

    PhyB averages over the k worst dynamics models with entropy-weighted coefficients and uses Bregman-regularized policy iteration; it claims bounded pessimism, monotonic improvement, and top D4RL scores.

Pith tools