Pith. sign in

Constrained Policy Optimization via Bayesian World Models

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Improving sample-efficiency and safety are crucial challenges when deploying reinforcement learning in high-stakes real world applications. We propose LAMBDA, a novel model-based approach for policy optimization in safety critical tasks modeled via constrained Markov decision processes. Our approach utilizes Bayesian world models, and harnesses the resulting uncertainty to maximize optimistic upper bounds on the task objective, as well as pessimistic upper bounds on the safety constraints. We demonstrate LAMBDA's state of the art performance on the Safety-Gym benchmark suite in terms of sample efficiency and constraint violation.

citation-role summary

background 1

citation-polarity summary

fields

cs.AI 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Safe Planning and Policy Optimization via World Model Learning

cs.AI · 2025-06-05 · conditional · novelty 6.0

SPOWL is a model-based safe RL method that uses a value-equivalent world model, a Lagrangian-trained safe policy, and adaptive planning thresholds to achieve low-cost, high-reward control on SafetyGymnasium tasks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Safe Planning and Policy Optimization via World Model Learning cs.AI · 2025-06-05 · conditional · none · ref 4 · internal anchor

    SPOWL is a model-based safe RL method that uses a value-equivalent world model, a Lagrangian-trained safe policy, and adaptive planning thresholds to achieve low-cost, high-reward control on SafetyGymnasium tasks.