Pith. sign in

REVIEW 1 cited by

Approximate Thompson Sampling for Learning Linear Quadratic Regulators with $O(\sqrt{T})$ Regret

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19380 v2 pith:QRYYOXRY submitted 2024-05-29 stat.ML cs.LGcs.SYeess.SY

classification stat.MLcs.LGcs.SYeess.SY
keywords approximateboundregretsamplingsqrtalgorithmexcitationlinear
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We propose a novel Thompson sampling algorithm that learns linear quadratic regulators (LQR) with a Bayesian regret bound of $O(\sqrt{T})$. Our method leverages Langevin dynamics with a carefully designed preconditioner and incorporates a simple excitation mechanism. We show that the excitation signal drives the minimum eigenvalue of the preconditioner to grow over time, thereby accelerating the approximate posterior sampling process. Furthermore, we establish nontrivial concentration properties of the approximate posteriors generated by our algorithm. These properties enable us to bound the moments of the system state and attain an $O(\sqrt{T})$ regret bound without relying on the restrictive assumptions that are often used in the literature.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Isoperimetry is All We Need: Langevin Posterior Sampling for RL with Sublinear Regret

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Posterior sampling (PSRL) and its Langevin-sampling approximation LaPSRL have sublinear regret for log-Sobolev, not only log-concave, posteriors.

Pith tools