Pith. sign in

REVIEW 1 cited by

Unentangled quantum reinforcement learning agents in the OpenAI Gym

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.14348 v1 pith:7UZJJLFB submitted 2022-03-27 quant-ph cond-mat.stat-mech

classification quant-phcond-mat.stat-mech
keywords quantumclassicalagentlearningopenaiagentsneuralreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Classical reinforcement learning (RL) has generated excellent results in different regions; however, its sample inefficiency remains a critical issue. In this paper, we provide concrete numerical evidence that the sample efficiency (the speed of convergence) of quantum RL could be better than that of classical RL, and for achieving comparable learning performance, quantum RL could use much (at least one order of magnitude) fewer trainable parameters than classical RL. Specifically, we employ the popular benchmarking environments of RL in the OpenAI Gym, and show that our quantum RL agent converges faster than classical fully-connected neural networks (FCNs) in the tasks of CartPole and Acrobot under the same optimization process. We also successfully train the first quantum RL agent that can complete the task of LunarLander in the OpenAI Gym. Our quantum RL agent only requires a single-qubit-based variational quantum circuit without entangling gates, followed by a classical neural network (NN) to post-process the measurement output. Finally, we could accomplish the aforementioned tasks on the real IBM quantum machines. To the best of our knowledge, none of the earlier quantum RL agents could do that.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CleanQRL: Lightweight Single-file Implementations of Quantum Reinforcement Learning Algorithms

    quant-ph 2025-07 conditional novelty 5.0 of 10

    The paper introduces CleanQRL, a collection of single-file implementations of quantum reinforcement learning algorithms designed to make QRL research easier to replicate and compare.

Pith tools