Pith. sign in

REVIEW 21 cited by

Lyapunov-based Safe Policy Optimization for Continuous Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1901.10031 v2 pith:KZ3SOVG2 submitted 2019-01-28 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords policyalgorithmsoptimizationsafeactionagentconstrainedcontinuous
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We study continuous action reinforcement learning problems in which it is crucial that the agent interacts with the environment only through safe policies, i.e.,~policies that do not take the agent to undesirable situations. We formulate these problems as constrained Markov decision processes (CMDPs) and present safe policy optimization algorithms that are based on a Lyapunov approach to solve them. Our algorithms can use any standard policy gradient (PG) method, such as deep deterministic policy gradient (DDPG) or proximal policy optimization (PPO), to train a neural network policy, while guaranteeing near-constraint satisfaction for every policy update by projecting either the policy parameter or the action onto the set of feasible solutions induced by the state-dependent linearized Lyapunov constraints. Compared to the existing constrained PG algorithms, ours are more data efficient as they are able to utilize both on-policy and off-policy data. Moreover, our action-projection algorithm often leads to less conservative policy updates and allows for natural integration into an end-to-end PG training pipeline. We evaluate our algorithms and compare them with the state-of-the-art baselines on several simulated (MuJoCo) tasks, as well as a real-world indoor robot navigation problem, demonstrating their effectiveness in terms of balancing performance and constraint satisfaction. Videos of the experiments can be found in the following link: https://drive.google.com/file/d/1pzuzFqWIE710bE2U6DmS59AfRzqK2Kek/view?usp=sharing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 154 citations worldwide. Full citation record

  1. Self-Organizing Dual-Buffer Adaptive Clustering Experience Replay (SODACER) for Safe Reinforcement Learning in Optimal Control

    eess.SY 2026-01 unverdicted novelty 7.0 of 10

    SODACER uses fast and slow buffers with adaptive clustering for experience replay in safe RL, integrated with CBFs and Sophia optimizer to achieve faster convergence and safety on nonlinear systems like HPV transmission.

  2. Robust Shielding for Safe Reinforcement Learning

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    A sound and optimal shielding method for robust MDPs ensures LTL safety under worst-case transitions and combines with PAC sampling to produce minimally restrictive shields for learned models.

  3. RAPT: Model-Predictive Out-of-Distribution Detection and Failure Diagnosis for Sim-to-Real Humanoid Deployment

    cs.RO 2026-02 conditional novelty 6.0 of 10

    A simulation-trained recurrent model detects out-of-distribution states on a real humanoid at 50 Hz and uses gradient saliency plus an LLM to diagnose failure causes.

  4. Hierarchical Multi-Agent Reinforcement Learning with Control Barrier Functions for Safety-Critical Autonomous Systems

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A hierarchical multi-agent RL method that selects skills at a high level and enforces pointwise safety with learned CBF-QP policies achieves about 99 percent success in simulated traffic scenarios.

  5. Iteratively Learning Muscle Memory for Legged Robots to Master Adaptive and High Precision Locomotion

    cs.RO 2025-07 unverdicted novelty 6.0 of 10

    Integrates iterative learning control with a torque library to enable high-precision adaptive locomotion on bipedal and quadrupedal robots, reducing tracking errors by up to 85% and achieving over 30x faster control rates.

  6. Stochastic Neural Control Barrier Functions

    eess.SY 2025-06 reject novelty 6.0 of 10

    A framework for synthesizing and verifying neural control barrier functions for stochastic systems, including new Tanaka-formula-based safety conditions for ReLU networks.

  7. On Almost Surely Safe Alignment of Large Language Models at Inference-Time

    cs.LG 2025-02 conditional novelty 6.0 of 10

    An inference-time beam-search method with a safety-state tracker and latent critic enforces a user-supplied safety cost model, with an almost-sure guarantee only relative to that model.

  8. From Text to Trajectory: Exploring Complex Constraint Representation and Decomposition in Safe Reinforcement Learning

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A text-trajectory contrastive model with per-step cost assignment reduces safety violations in reinforcement learning agents under natural language constraints.

  9. Learning-based Model Predictive Control for Safe Exploration and Reinforcement Learning

    eess.SY 2019-06 unverdicted novelty 6.0 of 10

    Develops a learning-based MPC algorithm that uses confidence intervals on trajectories and terminal set constraints to guarantee safety throughout RL exploration and training.

  10. Lyapunov Guidance: A Unified Framework for Stabilizing Generative Flows

    cs.LG 2026-07 reject novelty 5.0 of 10

    Flow guidance is framed as Lyapunov control with a pseudo-projection for stability, but the projected flow is not shown to sample the target conditional distribution.

  11. Safe and Generalizable Hierarchical Multi-Agent RL via Constraint Manifold Control

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    Proposes hierarchical MARL framework enforcing safety via constraint manifold at low level with theoretical guarantees and stationary dynamics for stable training and generalization.

  12. LC-SAC: Lyapunov-Constrained Soft Actor-Critic via Koopman Operator Theory for Trajectory Tracking and Stabilization

    eess.SY 2026-02 reject novelty 5.0 of 10

    LC-SAC adds a Lyapunov-decrease penalty derived from an EDMD/Koopman linear model and a DARE solution to the Soft Actor-Critic update, improving simulated 2D quadrotor tracking but without a formal guarantee on the tr...

  13. Vision-based Goal-Reaching Control for Mobile Robots Using a Hierarchical Learning Framework

    cs.RO 2026-01 conditional novelty 5.0 of 10

    A hierarchical control stack coupling ORB-SLAM3, tabular Q-learning/SARSA, and robust adaptive wheel control reaches centimeter-level goal positions on a 6,000 kg off-road robot.

  14. A Fast Initialization Method for Neural Network Controllers: A Case Study of Image-based Visual Servoing Control for the multicopter Interception

    eess.SY 2025-09 conditional novelty 5.0 of 10

    A neural network for image-based visual servoing interception is initialized by fitting it to model-generated data that satisfy Lyapunov derivative constraints.

  15. Action Mapping for Reinforcement Learning in Continuous Environments with Constraints

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Decoupling feasibility from objective optimization by training the RL policy over latent actions that map to feasible actions improves sample efficiency and constraint satisfaction in continuous constrained RL.

  16. Multi-task Offline Reinforcement Learning for Online Advertising in Recommender Systems

    cs.IR 2025-06 conditional novelty 4.0 of 10

    MTORL jointly learns channel recommendation and budget allocation for online advertising from offline user journeys, and reports better accuracy and reward than prior methods on KuaiRand, Criteo, and a Taobao A/B test.

  17. Central Path Proximal Policy Optimization

    cs.LG 2025-05 conditional novelty 4.0 of 10

    C3PO augments the PPO loss with a receding ReLU penalty on the cost advantage, approximating C-TRPO's central path and improving reward-constraint trade-offs in Safety Gymnasium tasks.

  18. Effective Reward Specification in Deep Reinforcement Learning

    cs.LG 2024-12 conditional novelty 4.0 of 10

    A thesis presenting four methods (ASAF, TeamReg, CoachReg, constrained RL, goal-conditioned GFlowNets) that improve reward specification for deep RL through demonstrations, policy regularization, behavior constraints,...

  19. A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach

    cs.RO 2025-12 conditional novelty 3.0 of 10

    A position/review paper argues data-driven model predictive control is the best route to safe, adaptive, human-like autonomous-driving motion planning, but provides no new derivation or experiment.

  20. SafeRL-Lite: A Lightweight, Explainable, and Constrained Reinforcement Learning Library

    cs.LG 2025-06 conditional novelty 3.0 of 10

    SafeRL-Lite provides modular wrappers for safety-constrained and explainable DQN agents, with a CartPole demonstration showing decreasing violations and pole-angle-dominant SHAP attributions.

  21. A Review On Safe Reinforcement Learning Using Lyapunov and Barrier Functions

    eess.SY 2025-08 unverdicted novelty 2.0 of 10

    A literature review of safe RL using Lyapunov and barrier functions that identifies a shift to model-free methods since 2017, well-defined open problems per approach class, and high-dimensional scalability as the main...

Pith tools