Pith. sign in

REVIEW 4 cited by

DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.14547 v3 pith:OWSWQ33U submitted 2020-04-30 cs.LG cs.AI

classification cs.LGcs.AI
keywords distributionaldsacactor-criticlearningrewardsrisk-sensitivesoftalgorithm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Distributional Soft Actor-Critic (DSAC), a distributional reinforcement learning (RL) algorithm that combines the strengths of distributional information of accumulated rewards and entropy-driven exploration from Soft Actor-Critic (SAC) algorithm. DSAC models the randomness in both action and rewards, surpassing baseline performances on various continuous control tasks. Unlike standard approaches that solely maximize expected rewards, we propose a unified framework for risk-sensitive learning, one that optimizes the risk-related objective while balancing entropy to encourage exploration. Extensive experiments demonstrate DSAC's effectiveness in enhancing agent performances for both risk-neutral and risk-sensitive control tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Managing Portfolios Across the Return Distribution

    q-fin.GN 2025-10 reject novelty 6.0 of 10

    A quantile-targeted actor-critic (Q-A2C) produces τ-ordered portfolios in a regime example and ETF/industry tests, but the empirical evidence is in-sample and internally inconsistent.

  2. Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    ORAC combines upper-confidence-bound reward exploration with lower-confidence-bound risk-averse cost constraints and adaptive cost weighting to improve exploration in risk-averse constrained RL.

  3. Value Flows

    cs.LG 2025-10 reject novelty 5.0 of 10

    Value Flows fits the full return distribution in RL with a flow-matching critic and reweights its learning objective by estimated return variance; the central theoretical guarantee does not follow from the stated equations.

  4. Uncertainty-Aware Safety-Critical Decision and Control for Autonomous Vehicles at Unsignalized Intersections

    cs.RO 2025-05 conditional novelty 5.0 of 10

    USDC combines ensemble distributional RL with a high-order control barrier function and uncertainty-based switching to reduce collisions at unsignalized intersections while preserving traffic efficiency in simulation.

Pith tools