Relearn LQR combines recursive least squares with policy gradient for on-policy data-driven LQR and proves stability of the full scheme via Lyapunov analysis with averaging and timescale separation.
A tour of reinforcement learning: The view from continuous control
2 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
An external runtime governance layer for embodied agents intercepts unauthorized actions at ~96% and recovers from runtime drift at ~91% under policy constraints in simulation, outperforming pre-execution-only baselines on continuous detection and recovery.
citing papers explorer
-
Stability-Certified On-Policy Data-Driven LQR via Recursive Learning and Policy Gradient
Relearn LQR combines recursive least squares with policy gradient for on-policy data-driven LQR and proves stability of the full scheme via Lyapunov analysis with averaging and timescale separation.
-
Harnessing Embodied Agents: Runtime Governance for Policy-Constrained Execution
An external runtime governance layer for embodied agents intercepts unauthorized actions at ~96% and recovers from runtime drift at ~91% under policy constraints in simulation, outperforming pre-execution-only baselines on continuous detection and recovery.