REVIEW 2 cited by
Linear-Quadratic Mean-Field Reinforcement Learning: Convergence of Policy Gradient Methods
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We investigate reinforcement learning in the setting of Markov decision processes for a large number of exchangeable agents interacting in a mean field manner. Applications include, for example, the control of a large number of robots communicating through a central unit dispatching the optimal policy computed by maximizing an aggregate reward. An approximate solution is obtained by learning the optimal policy of a generic agent interacting with the statistical distribution of the states and actions of the other agents. We first provide a full analysis this discrete-time mean field control problem. We then rigorously prove the convergence of exact and model-free policy gradient methods in a mean-field linear-quadratic setting and establish bounds on the rates of convergence. We also provide graphical evidence of the convergence based on implementations of our algorithms.
Forward citations
Cited by 2 Pith papers
-
Policy Optimization for Continuous-time Linear-Quadratic Graphon Mean Field Games
A bilevel policy optimization algorithm for continuous-time linear-quadratic graphon mean field games converges linearly to best-response policies and globally to the Nash equilibrium.
-
Toward Optimal Statistical Inference in Noisy Linear Quadratic Reinforcement Learning over a Finite Horizon
In finite-horizon noisy LQ control, the policy gradient estimator and its objective cost are claimed to be asymptotically normal, and online bootstrapped confidence intervals are claimed valid with quantile error n^{-1/4}.
Discussion (0). Sign in to comment.