Pith. sign in

REVIEW 1 cited by

Learning Stabilizing Policies via an Unstable Subspace Representation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.01348 v1 pith:GSOOJ6KY submitted 2025-05-02 cs.LG math.OC

classification cs.LGmath.OC
keywords unstablesubspacepolicysystemcontroldimensionlearningapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the problem of learning to stabilize (LTS) a linear time-invariant (LTI) system. Policy gradient (PG) methods for control assume access to an initial stabilizing policy. However, designing such a policy for an unknown system is one of the most fundamental problems in control, and it may be as hard as learning the optimal policy itself. Existing work on the LTS problem requires large data as it scales quadratically with the ambient dimension. We propose a two-phase approach that first learns the left unstable subspace of the system and then solves a series of discounted linear quadratic regulator (LQR) problems on the learned unstable subspace, targeting to stabilize only the system's unstable dynamics and reduce the effective dimension of the control space. We provide non-asymptotic guarantees for both phases and demonstrate that operating on the unstable subspace reduces sample complexity. In particular, when the number of unstable modes is much smaller than the state dimension, our analysis reveals that LTS on the unstable subspace substantially speeds up the stabilization process. Numerical experiments are provided to support this sample complexity reduction achieved by our approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Gradient Domination of the LQG Problem

    math.OC 2025-07 conditional novelty 6.0 of 10

    The LQG cost becomes gradient dominated under a history-based controller parameterization, yielding global convergence guarantees for policy gradient methods in model-based and model-free settings.

Pith tools