Pith. sign in

REVIEW 1 cited by

Parametrized quantum policies for reinforcement learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.05577 v2 pith:BVETB6NJ submitted 2021-03-09 quant-ph cs.AIcs.LGstat.ML

classification quant-phcs.AIcs.LGstat.ML
keywords learningquantumtasksclassicalparametrizedreinforcementbenchmarkinghybrid
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the advent of real-world quantum computing, the idea that parametrized quantum computations can be used as hypothesis families in a quantum-classical machine learning system is gaining increasing traction. Such hybrid systems have already shown the potential to tackle real-world tasks in supervised and generative learning, and recent works have established their provable advantages in special artificial tasks. Yet, in the case of reinforcement learning, which is arguably most challenging and where learning boosts would be extremely valuable, no proposal has been successful in solving even standard benchmarking tasks, nor in showing a theoretical learning advantage over classical algorithms. In this work, we achieve both. We propose a hybrid quantum-classical reinforcement learning model using very few qubits, which we show can be effectively trained to solve several standard benchmarking environments. Moreover, we demonstrate, and formally prove, the ability of parametrized quantum circuits to solve certain learning tasks that are intractable for classical models, including current state-of-art deep neural networks, under the widely-believed classical hardness of the discrete logarithm problem.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 10 citations worldwide. Full citation record

  1. Experimental investigation of single qubit quantum classifier with small number of samples

    quant-ph 2025-07 conditional novelty 4.0 of 10

    A silicon photonic single-qubit classifier achieves about 86 percent accuracy when trained with an average of roughly two photons per sample, matching simulation.

Pith tools