REVIEW 1 cited by
On the Global Convergence of Imitation Learning: A Case for Linear Quadratic Regulator
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We study the global convergence of generative adversarial imitation learning for linear quadratic regulators, which is posed as minimax optimization. To address the challenges arising from non-convex-concave geometry, we analyze the alternating gradient algorithm and establish its Q-linear rate of convergence to a unique saddle point, which simultaneously recovers the globally optimal policy and reward function. We hope our results may serve as a small step towards understanding and taming the instability in imitation learning as well as in more general non-convex-concave alternating minimax optimization that arises from reinforcement learning and generative adversarial learning.
Forward citations
Cited by 1 Pith paper
-
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration
ILDE combines VAIL imitation, GIRIL curiosity, and a state-entropy bonus to beat expert scores on 6 Atari games and match or slightly beat GIRIL on MuJoCo, using only 10% of one-life demonstrations.
Discussion (0). Continue with ORCID to comment.