Q(s, a)−E[r]−γ·sup β≥0 −βlog Ep0s,a exp −V(s ′) β −βδ # . If using ERM method, the empirical Bellman residual is bLQ := 1 N NX i=1

Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

browse 1 citing papers

representative citing papers

DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty

cs.LG · 2025-06-14 · unverdicted · novelty 7.0

DR-SAC is the first actor-critic distributionally robust RL algorithm for offline continuous control that derives a convergent robust soft policy iteration and reports up to 9.8x higher rewards than SAC under perturbations.

citing papers explorer

Showing 1 of 1 citing paper.

DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty cs.LG · 2025-06-14 · unverdicted · none · ref 68
DR-SAC is the first actor-critic distributionally robust RL algorithm for offline continuous control that derives a convergent robust soft policy iteration and reports up to 9.8x higher rewards than SAC under perturbations.

Q(s, a)−E[r]−γ·sup β≥0 −βlog Ep0s,a exp −V(s ′) β −βδ # . If using ERM method, the empirical Bellman residual is bLQ := 1 N NX i=1

fields

years

verdicts

representative citing papers

citing papers explorer