Pith. sign in

REVIEW 2 cited by

Policy Optimization in Control: Geometry and Algorithmic Implications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.04243 v1 pith:NXTGE3J4 submitted 2024-06-06 math.OC cs.SYeess.SYmath.DG

classification math.OCcs.SYeess.SYmath.DG
keywords controlpolicygeometrydesignoptimizationperformancepoliciesdynamic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

This survey explores the geometric perspective on policy optimization within the realm of feedback control systems, emphasizing the intrinsic relationship between control design and optimization. By adopting a geometric viewpoint, we aim to provide a nuanced understanding of how various ``complete parameterization'' -- referring to the policy parameters together with its Riemannian geometry -- of control design problems, influence stability and performance of local search algorithms. The paper is structured to address key themes such as policy parameterization, the topology and geometry of stabilizing policies, and their implications for various (non-convex) dynamic performance measures. We focus on a few iconic control design problems, including the Linear Quadratic Regulator (LQR), Linear Quadratic Gaussian (LQG) control, and $\mathcal{H}_\infty$ control. In particular, we first discuss the topology and Riemannian geometry of stabilizing policies, distinguishing between their static and dynamic realizations. Expanding on this geometric perspective, we then explore structural properties of the aforementioned performance measures and their interplay with the geometry of stabilizing policies in presence of policy constraints; along the way, we address issues such as spurious stationary points, symmetries of dynamic feedback policies, and (non-)smoothness of the corresponding performance measures. We conclude the survey with algorithmic implications of policy optimization in feedback design.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Interpretable Gradient Descent for Kalman Gain

    math.OC 2025-07 conditional novelty 7.0 of 10

    Gradient descent on the innovation loss converges to the Kalman gain under a nonstandard observability condition, with a geometric rate tied to observability and orthogonality violation.

  2. A Proximal Descent Method for Minimizing Weakly Convex Optimization

    math.OC 2025-09 conditional novelty 6.0 of 10

    A bundle-based proximal descent method achieves O(1/delta^4) for Moreau stationarity on weakly convex functions and adapts to O(1/delta^2) under smoothness and linear convergence under quadratic growth.

Pith tools