VLM-Safe-RL adds frozen VLM signals as anticipatory costs to the CMDP Lagrangian update via dual-path CLIP, VLM-Lagrange, and confidence gating, outperforming baselines on Safety-Gymnasium FormulaOne while showing partial generalization.
Penalized proximal policy optimization for safe reinforcement learning
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4verdicts
UNVERDICTED 4representative citing papers
CRAX is a new fast benchmark suite for constrained RL built on MJX, with six environment suites and tasks across difficulty levels, showing no single safe RL method dominates and benefits from curriculum learning.
PPO-EAL integrates exact augmented Lagrangian optimization into PPO for safe robotic control, with claimed theoretical guarantees and better empirical safety-performance tradeoffs on several robot benchmarks including sim-to-real gear assembly.
A single reinforcement learning policy jointly trains multiple locomotion skills for wheeled-legged robots with DC-motor constraints and learns a proprioceptive skill selector for adaptive behavior.
citing papers explorer
-
Seeing Before Colliding: Anticipatory Safe RL with Frozen Vision-Language Models
VLM-Safe-RL adds frozen VLM signals as anticipatory costs to the CMDP Lagrangian update via dual-path CLIP, VLM-Lagrange, and confidence gating, outperforming baselines on Safety-Gymnasium FormulaOne while showing partial generalization.
-
CRAX: Fast Safe Reinforcement Learning Benchmarking
CRAX is a new fast benchmark suite for constrained RL built on MJX, with six environment suites and tasks across difficulty levels, showing no single safe RL method dominates and benefits from curriculum learning.
-
PPO-EAL: Exact Augmented Lagrangian Proximal Policy Optimization for Safe Robotic Control
PPO-EAL integrates exact augmented Lagrangian optimization into PPO for safe robotic control, with claimed theoretical guarantees and better empirical safety-performance tradeoffs on several robot benchmarks including sim-to-real gear assembly.
-
MUJICA: Multi-skill Unified Joint Integration of Control Architecture for Wheeled-Legged Robots
A single reinforcement learning policy jointly trains multiple locomotion skills for wheeled-legged robots with DC-motor constraints and learns a proprioceptive skill selector for adaptive behavior.