Pith. sign in

REVIEW 1 cited by

Regret Analysis of Learning-Based MPC with Partially-Unknown Cost Function

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.02307 v2 pith:QVRSP6RY submitted 2021-08-04 math.OC cs.LG

classification math.OCcs.LG
keywords controllercostlearning-basedregretanalysischallengecontrolfinite-horizon
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The exploration/exploitation trade-off is an inherent challenge in data-driven adaptive control. Though this trade-off has been studied for multi-armed bandits (MAB's) and reinforcement learning for linear systems; it is less well-studied for learning-based control of nonlinear systems. A significant theoretical challenge in the nonlinear setting is that there is no explicit characterization of an optimal controller for a given set of cost and system parameters. We propose the use of a finite-horizon oracle controller with full knowledge of parameters as a reasonable surrogate to optimal controller. This allows us to develop policies in the context of learning-based MPC and MAB's and conduct a control-theoretic analysis using techniques from MPC- and optimization-theory to show these policies achieve low regret with respect to this finite-horizon oracle. Our simulations exhibit the low regret of our policy on a heating, ventilation, and air-conditioning model with partially-unknown cost function.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reinforcement Learning for a Discrete-Time Linear-Quadratic Control Problem with an Application

    stat.ML 2024-12 reject novelty 3.0 of 10

    The paper claims entropy regularization forces the optimal LQ feedback policy to be Gaussian and uses that to solve a mean-variance asset-liability problem, but the proof of the main theorem contains a correlation err...

Pith tools