Pith. sign in

Conservative safety critics for exploration

6 Pith papers cite this work, alongside 32 external citations. Polarity classification is still indexing.

6 Pith papers citing it
32 external citations · external index

years

2026 6

verdicts

UNVERDICTED 6

representative citing papers

SHAPO: Sharpness-Aware Policy Optimization for Safe Exploration

cs.LG · 2026-06-08 · unverdicted · novelty 5.0

SHAPO adds a sharpness-aware adjustment to policy optimization that reweights gradients to favor conservative behavior in uncertain areas, yielding better safety-performance tradeoffs on continuous control tasks.

An Agency-Transferring Model-Free Policy Enhancement Technique

cs.LG · 2026-06-08 · unverdicted · novelty 5.0

A model-free RL method arbitrates between a functional baseline policy and a learning policy, transferring agency over time to yield a standalone policy with high goal-reaching rates and competitive returns on continuous-control tasks.

Safe-Support Q-Learning: Learning without Unsafe Exploration

cs.LG · 2026-04-28 · unverdicted · novelty 5.0

Safe-Support Q-Learning trains Q-functions and policies in reinforcement learning without ever visiting unsafe states by constraining the behavior policy to a safe set and using KL-regularized Bellman targets in a two-stage framework.

citing papers explorer

Showing 6 of 6 citing papers.