REVIEW 4 cited by
DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present Distributional Soft Actor-Critic (DSAC), a distributional reinforcement learning (RL) algorithm that combines the strengths of distributional information of accumulated rewards and entropy-driven exploration from Soft Actor-Critic (SAC) algorithm. DSAC models the randomness in both action and rewards, surpassing baseline performances on various continuous control tasks. Unlike standard approaches that solely maximize expected rewards, we propose a unified framework for risk-sensitive learning, one that optimizes the risk-related objective while balancing entropy to encourage exploration. Extensive experiments demonstrate DSAC's effectiveness in enhancing agent performances for both risk-neutral and risk-sensitive control tasks.
Forward citations
Cited by 4 Pith papers
-
Managing Portfolios Across the Return Distribution
A quantile-targeted actor-critic (Q-A2C) produces τ-ordered portfolios in a regime example and ETF/industry tests, but the empirical evidence is in-sample and internally inconsistent.
-
Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning
ORAC combines upper-confidence-bound reward exploration with lower-confidence-bound risk-averse cost constraints and adaptive cost weighting to improve exploration in risk-averse constrained RL.
-
Value Flows
Value Flows fits the full return distribution in RL with a flow-matching critic and reweights its learning objective by estimated return variance; the central theoretical guarantee does not follow from the stated equations.
-
Uncertainty-Aware Safety-Critical Decision and Control for Autonomous Vehicles at Unsignalized Intersections
USDC combines ensemble distributional RL with a high-order control barrier function and uncertainty-based switching to reduce collisions at unsignalized intersections while preserving traffic efficiency in simulation.
Discussion (0). Continue with ORCID to comment.