Pith. sign in

REVIEW 1 cited by

Optimal Transport-Assisted Risk-Sensitive Q-Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11774 v2 pith:A4I7PQJX submitted 2024-06-17 cs.LG cs.SYeess.SY

classification cs.LGcs.SYeess.SY
keywords optimalq-learningalgorithmpolicysafetydistributionlearningreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The primary goal of reinforcement learning is to develop decision-making policies that prioritize optimal performance without considering risk or safety. In contrast, safe reinforcement learning aims to mitigate or avoid unsafe states. This paper presents a risk-sensitive Q-learning algorithm that leverages optimal transport theory to enhance the agent safety. By integrating optimal transport into the Q-learning framework, our approach seeks to optimize the policy's expected return while minimizing the Wasserstein distance between the policy's stationary distribution and a predefined risk distribution, which encapsulates safety preferences from domain experts. We validate the proposed algorithm in a Gridworld environment. The results indicate that our method significantly reduces the frequency of visits to risky states and achieves faster convergence to a stable policy compared to the traditional Q-learning algorithm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Wasserstein Adaptive Value Estimation for Actor-Critic Reinforcement Learning

    cs.LG 2025-01 reject novelty 4.0 of 10

    WAVE adds an adaptively weighted Sinkhorn approximation of the Wasserstein distance between successive Q-value distributions to the critic loss in actor-critic reinforcement learning.

Pith tools