Pith. sign in

REVIEW 2 cited by

Exploration-exploitation trade-off for continuous-time episodic reinforcement learning with linear-convex models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.10264 v1 pith:XWJC7IZX submitted 2021-12-19 cs.LG math.OCmath.PRstat.ML

classification cs.LGmath.OCmath.PRstat.ML
keywords performancelearningalgorithmarxivcasecontinuous-timecontrolepisodic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We develop a probabilistic framework for analysing model-based reinforcement learning in the episodic setting. We then apply it to study finite-time horizon stochastic control problems with linear dynamics but unknown coefficients and convex, but possibly irregular, objective function. Using probabilistic representations, we study regularity of the associated cost functions and establish precise estimates for the performance gap between applying optimal feedback control derived from estimated and true model parameters. We identify conditions under which this performance gap is quadratic, improving the linear performance gap in recent work [X. Guo, A. Hu, and Y. Zhang, arXiv preprint, arXiv:2104.09311, (2021)], which matches the results obtained for stochastic linear-quadratic problems. Next, we propose a phase-based learning algorithm for which we show how to optimise exploration-exploitation trade-off and achieve sublinear regrets in high probability and expectation. When assumptions needed for the quadratic performance gap hold, the algorithm achieves an order $\mathcal{O}(\sqrt{N} \ln N)$ high probability regret, in the general case, and an order $\mathcal{O}((\ln N)^2)$ expected regret, in self-exploration case, over $N$ episodes, matching the best possible results from the literature. The analysis requires novel concentration inequalities for correlated continuous-time observations, which we derive.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Convergence Rates of Time Discretization in Extended Mean Field Control

    math.OC 2025-08 conditional novelty 8.0 of 10

    For linear-convex extended mean field control, piecewise constant controls approximate the optimal cost at rate 1/2 and the optimal control at rate 1/4; under smoothness, the rate improves to 1.

  2. Data-Driven Exploration for a Class of Continuous-Time Indefinite Linear--Quadratic Reinforcement Learning Problems

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Data-driven adaptive exploration achieves O(N^{3/4}) regret in continuous-time linear-quadratic reinforcement learning, matching fixed-schedule methods and extending them to zero initial states.

Pith tools