REVIEW 3 cited by
Quantum Policy Iteration via Amplitude Estimation and Grover Search -- Towards Quantum Advantage for Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present a full implementation and simulation of a novel quantum reinforcement learning method. Our work is a detailed and formal proof of concept for how quantum algorithms can be used to solve reinforcement learning problems and shows that, given access to error-free, efficient quantum realizations of the agent and environment, quantum methods can yield provable improvements over classical Monte-Carlo based methods in terms of sample complexity. Our approach shows in detail how to combine amplitude estimation and Grover search into a policy evaluation and improvement scheme. We first develop quantum policy evaluation (QPE) which is quadratically more efficient compared to an analogous classical Monte Carlo estimation and is based on a quantum mechanical realization of a finite Markov decision process (MDP). Building on QPE, we derive a quantum policy iteration that repeatedly improves an initial policy using Grover search until the optimum is reached. Finally, we present an implementation of our algorithm for a two-armed bandit MDP which we then simulate.
Forward citations
Cited by 3 Pith papers
-
Model Predictive Path Integral Control as a Quantum Query Problem
The finite-ensemble MPPI control update is expressed as a ratio of bounded expectations and estimated by quantum amplitude estimation with O(m/(ε√a)) queries, a quadratic improvement over Monte Carlo.
-
HCQA: Hybrid Classical-Quantum Agent for Generating Optimal Quantum Sensor Circuits
A DQN agent with a quantum action-selection circuit generates two-qubit quantum sensor circuits that reach normalized QFI=1, the paper's claimed optimum.
-
Quantum reinforcement learning in dynamic environments
A quantum hybrid RL agent with a dissipation mechanism outlearns a classical agent in a Gridworld with a suddenly changing reward path, for suitable dissipation values.
Discussion (0). Sign in to comment.