Pith. sign in

Learning Stabilizing Policies via an Unstable Subspace Representation

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We study the problem of learning to stabilize (LTS) a linear time-invariant (LTI) system. Policy gradient (PG) methods for control assume access to an initial stabilizing policy. However, designing such a policy for an unknown system is one of the most fundamental problems in control, and it may be as hard as learning the optimal policy itself. Existing work on the LTS problem requires large data as it scales quadratically with the ambient dimension. We propose a two-phase approach that first learns the left unstable subspace of the system and then solves a series of discounted linear quadratic regulator (LQR) problems on the learned unstable subspace, targeting to stabilize only the system's unstable dynamics and reduce the effective dimension of the control space. We provide non-asymptotic guarantees for both phases and demonstrate that operating on the unstable subspace reduces sample complexity. In particular, when the number of unstable modes is much smaller than the state dimension, our analysis reveals that LTS on the unstable subspace substantially speeds up the stabilization process. Numerical experiments are provided to support this sample complexity reduction achieved by our approach.

citation-role summary

extension 1

citation-polarity summary

fields

math.OC 1

years

2025 1

verdicts

CONDITIONAL 1

roles

extension 1

polarities

extend 1

representative citing papers

On the Gradient Domination of the LQG Problem

math.OC · 2025-07-11 · conditional · novelty 6.0

The LQG cost becomes gradient dominated under a history-based controller parameterization, yielding global convergence guarantees for policy gradient methods in model-based and model-free settings.

citing papers explorer

Showing 1 of 1 citing paper.

  • On the Gradient Domination of the LQG Problem math.OC · 2025-07-11 · conditional · none · ref 27 · internal anchor

    The LQG cost becomes gradient dominated under a history-based controller parameterization, yielding global convergence guarantees for policy gradient methods in model-based and model-free settings.