Pith. sign in

REVIEW 3 major objections 4 minor 41 references

Entangled quantum circuits outperform separable ones as feature extractors for agents learning competitive Pong.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 23:43 UTC pith:HJZSUNKE

load-bearing objection Clean first isolation of entanglement in hybrid PPO for competitive classical RL; the return gap is real, the causal story about interaction modelling is still an interpretation. the 3 major comments →

arxiv 2603.10289 v2 pith:HJZSUNKE submitted 2026-03-11 quant-ph cs.AIcs.LG

Quantum entanglement provides a competitive advantage in adversarial games

classification quant-ph cs.AIcs.LG
keywords quantum entanglementreinforcement learningparameterised quantum circuitscompetitive gamesPongrepresentation learningproximal policy optimisationquantum machine learning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Whether entanglement helps when both the environment and the rewards are fully classical remains an open question. This paper isolates that resource inside a hybrid agent that plays Pong: an eight-qubit circuit extracts features for a standard proximal-policy-optimisation actor and critic, and the only controlled change is whether the circuit is separable or uses fixed or trainable entangling gates. Entangled circuits consistently beat separable circuits with comparable parameter counts; in the low-parameter regime they also match or exceed classical multilayer-perceptron baselines. Representation-similarity analysis shows the entangled circuits learn features unlike those of classical networks, which the authors read as better capture of interactions among game-state variables. The result treats entanglement as a practical tool for representation learning under tight capacity limits rather than as a purely theoretical advantage.

Core claim

In a controlled comparison on the competitive Markov game Pong, hybrid agents whose feature extractors are eight-qubit parameterised quantum circuits with fixed CZ or trainable IsingZZ entangling gates consistently achieve higher episodic returns than identical agents whose circuits use only single-qubit gates, at matched parameter counts. In low-capacity regimes the entangled hybrids also match or exceed classical multilayer-perceptron baselines. The authors conclude that entanglement functions as a resource for representation learning in competitive reinforcement learning.

What carries the argument

An eight-qubit data-reuploading parameterised quantum circuit that maps Pong’s eight-dimensional observation vector to an eight-dimensional feature vector for classical actor–critic heads; architectures differ only by the presence or absence of fixed CZ or trainable IsingZZ entangling gates.

Load-bearing premise

The performance gap is caused by entangling gates modelling interactions among the eight state variables, rather than by other side-effects those gates introduce in the optimisation landscape or effective capacity.

What would settle it

Retrain the same hybrid agents after randomly permuting or noise-replacing the eight observation components so that genuine interactions among variables no longer exist; if entangled circuits then lose their advantage over separable ones, the interaction-modelling explanation fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Entanglement can improve feature extraction for classical reinforcement learning without quantum environments or channels.
  • Shallow entangled circuits can outperform deeper ones on this task, consistent with optimisation limits such as barren plateaus.
  • In low-parameter regimes relevant to near-term hardware, quantum feature extractors can beat classical networks of similar size.
  • Fixed entangling gates can match or exceed trainable ones on average, pointing to a trainability–expressivity trade-off.
  • Representation similarity (CKA) can separate quantum-learned features from classical ones, supporting claims of distinct function spaces.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Isolating entanglement the same way on other low-dimensional Markov games would test whether the advantage is specific to Pong’s geometry or more general.
  • If the gain mainly comes from coupling state variables, classical networks that explicitly multiply or interact features might close the gap without quantum hardware.
  • Self-play or opponent-pool evaluation, which the authors list as future work, would show whether the learned features remain useful when the opponent adapts.
  • The low-parameter edge may matter most for hybrid controllers where classical model size is limited by power, latency, or memory.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper reports a controlled empirical study of an 8-qubit data-reuploading PQC used as a feature extractor inside a PPO agent for the classical Pong Markov game. Three quantum backbones (separable, fixed-CZ entangled, trainable-IsingZZ entangled) are compared against classical MLP baselines of matched or larger parameter counts, with performance measured by final episodic return averaged over 10 random seeds (Table I, Figs. 2–3, 5). Entangled circuits consistently yield higher mean returns than separable circuits of comparable depth/parameter count; in the low-parameter regime some entangled configurations match or exceed small MLPs. Linear CKA (Fig. 4) is used to argue that the entangled representations are structurally distinct from both separable and classical ones, which the authors interpret as evidence that entanglement improves modelling of interactions among the eight state variables.

Significance. If the performance gap is robust, the work supplies the first systematic isolation of entanglement as a design variable in PQC-based agents acting in a fully classical competitive RL environment—an open question repeatedly noted in the QRL literature. The experimental design is clean on the surface (identical single-qubit U3 re-uploading layers, layer-by-layer parameter matching, multiple seeds, public CleanRL baseline), and the low-parameter regime results are practically relevant for near-term hardware. The CKA analysis and the explicit discussion of barren-plateau trade-offs further strengthen the contribution. These elements make the manuscript a useful empirical benchmark even if the causal interpretation of “interaction modelling” requires qualification.

major comments (3)
  1. [II.A–B, III.A] Section II.A–B and Discussion III.A: the central causal claim—that the return gap arises because entangling gates enable modelling of interactions among the eight state variables—is not isolated by the experimental design. Separable and entangled circuits differ only by the presence of CZ/IsingZZ gates (Eqs. 5–7, Fig. 1), yet those gates also alter the optimisation landscape, gradient variance and effective depth. The authors themselves invoke barren plateaus to explain the non-monotonic depth dependence (II.B, III.C). Without a control that preserves two-qubit correlations while removing the ability to couple distinct input coordinates (or that injects classical multiplicative interactions into a separable circuit), the performance advantage (Table I, Fig. 2) remains consistent with several non-interaction explanations.
  2. [Abstract, III.A–B, Fig. 4] Abstract, Section III.A–B and Fig. 4: linear CKA dissimilarity is presented as evidence of “improved modelling of interacting state variables.” CKA only quantifies representational divergence; it supplies no causal link to interaction modelling. The paper equates low intra-group and inter-group CKA scores with functional superiority, but this is an untested interpretive step. Either a direct probe of feature interactions (e.g., ablation of cross terms, mutual-information analysis of measured observables) or a substantial softening of the claim is required for the “functional resource for representation learning” conclusion to hold.
  3. [I, III.E, title] The evaluation is performed against a single fixed (computer-controlled) opponent. While the authors correctly flag self-play / opponent-pool evaluation as future work (III.E), the present results therefore do not yet demonstrate an advantage “in adversarial games” under the Markov-game perspective advertised in the title and introduction. At minimum the abstract and claims should be qualified to “against a fixed opponent,” or a limited multi-seed opponent evaluation should be added.
minor comments (4)
  1. [Table I] Table I caption and surrounding text: the classical 4096-parameter MLP is described as having “the same hidden dimension size as an 8-qubit quantum circuit” (2^8). This is a Hilbert-space analogy, not a parameter-count match; the wording should be clarified to avoid implying architectural equivalence.
  2. [Figs. 2, 5] Figure 2 and Figure 5 panels contain partially garbled axis labels and legends in the supplied manuscript text; ensure final production versions are clean and that the exponential-moving-average smoothing parameters are stated.
  3. [I.A] Eq. (1)–(3) restate the standard PPO objective; a brief pointer to CleanRL and the precise hyper-parameter values used (clip ε, c1, c2, γ, λ) would improve reproducibility without lengthening the main text.
  4. [Appendix A] Appendix Table II (episodic length) is correctly noted as less informative than return, yet the large standard deviations (e.g., CZ 288-layer mean 616.5 ± 745) deserve a short remark on early termination versus prolonged play.

Circularity Check

0 steps flagged

No circularity: purely empirical training comparisons on Pong with no fitted quantities renamed as predictions and no load-bearing self-citation chains.

full rationale

The paper reports controlled PPO training runs of hybrid agents whose only systematic difference is the presence/absence of CZ or IsingZZ gates inside an otherwise identical 8-qubit data-reuploading backbone (Eqs. 5–7, Fig. 1). Final episodic returns (Table I, Figs. 2–3, 5) and CKA matrices (Fig. 4) are measured quantities obtained after training; they are not algebraic consequences of any free parameter that was itself fitted to the same returns. Standard citations (CleanRL, Schulman PPO, Kornblith CKA, McClean barren plateaus) supply algorithmic scaffolding or post-hoc diagnostics and do not encode the claimed performance gap. No uniqueness theorem, ansatz, or self-definition is invoked to force the result. The causal interpretation that entanglement models state-variable interactions is an after-the-fact reading of the data, not a circular derivation. Hence the work is self-contained against external benchmarks and exhibits zero circularity of the kinds enumerated.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

The central claim rests on standard RL and quantum-circuit machinery plus a handful of modelling choices that are not independently validated outside this experiment. No new physical entities are postulated; free parameters are the usual hyper-parameters of PPO and circuit depth, none of which are fitted to produce the claimed advantage by construction.

free parameters (4)
  • PPO clip epsilon, value-loss coefficient c1, entropy coefficient c2, discount γ, GAE λ
    Standard PPO hyper-parameters taken from CleanRL defaults or left unspecified; their precise values affect absolute returns and therefore the magnitude of the reported gaps.
  • Number of PQC layers (1–6)
    Chosen by the experimenters; performance peaks at shallow depths and collapses deeper, so the reported “entanglement advantage” is conditioned on this discrete choice.
  • Hidden-layer widths of classical MLP baselines (4, 8, 16, 21, 256)
    Selected to produce parameter counts comparable to the quantum models; different widths would change the classical reference points.
  • Data-reuploading weights and biases inside each U3 gate
    Trainable parameters of the quantum feature extractor; their random initialisation produces the large seed-to-seed variance observed.
axioms (4)
  • domain assumption Pong with anti-symmetric scoring is a sufficiently faithful approximation of a competitive zero-sum Markov game for testing interaction modelling.
    Stated in the introduction and used throughout; the discussion itself later notes that a fixed opponent is inadequate and self-play is needed.
  • domain assumption Differences between circuits that share identical single-qubit gates and differ only by the presence/absence of CZ or IsingZZ gates can be attributed to entanglement.
    Core experimental design claim (Section I.B and II.A); ignores possible secondary effects on gradient variance or barren-plateau severity.
  • ad hoc to paper Linear CKA dissimilarity between feature matrices implies structurally distinct and functionally superior representations for interacting state variables.
    Interpretive step in Section II.D and Discussion III.B; CKA only measures linear similarity, not causal utility for policy optimisation.
  • domain assumption Standard PPO convergence theory and CleanRL implementation details apply unchanged to hybrid quantum-classical feature extractors.
    Implicit throughout the methods; no re-derivation or ablation of the classical optimiser is supplied.

pith-pipeline@v1.1.0-grok45 · 20711 in / 3159 out tokens · 33273 ms · 2026-07-14T23:43:03.268667+00:00 · methodology

0 comments
read the original abstract

Whether uniquely quantum resources confer advantages in fully classical, competitive environments remains an open question. Competitive zero-sum reinforcement learning is particularly challenging, as success requires modelling dynamic interactions between opposing agents rather than static state-action mappings. Here, we conduct a controlled study isolating the role of quantum entanglement in a quantum-classical hybrid agent trained on Pong, a competitive Markov game. An 8-qubit parameterised quantum circuit serves as a feature extractor within a proximal policy optimisation framework, allowing direct comparison between separable circuits and architectures incorporating fixed (CZ) or trainable (IsingZZ) entangling gates. Entangled circuits consistently outperform separable counterparts with comparable parameter counts and, in low-capacity regimes, match or exceed classical multilayer perceptron baselines. Representation similarity analysis further shows that entangled circuits learn structurally distinct features, consistent with improved modelling of interacting state variables. These findings establish entanglement as a function resource for representation learning in competitive reinforcement learning.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

41 extracted references · 2 canonical work pages · 1 internal anchor

  1. [1]

    hidden dimension size

    exp(i(λ+δ)) cos( θ 2) . (7) Hence, for separable and CZ-entangled PQCs, each layer has 6×8 = 48 trainable parameters. For IsingZZ- entangled PQCs, each layer has 6×8 + 8 = 56 trainable parameters. II. RESULTS We systematically evaluated the effect of quantum en- tanglement on the performance of a quantum–classical hybrid reinforcement learning agent in th...

  2. [2]

    J. L. Treynor, What does it take to win the trading game?, Fin. Anal. J.37, 55 (1981)

  3. [3]

    Treynor, Zero sum, Fin

    J. Treynor, Zero sum, Fin. Anal. J.55, 8 (1999)

  4. [4]

    Silver, A

    D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis, Mastering the game of go with deep neural networks and tree search, Nature529, 484 (2016)

  5. [5]

    Silver, J

    D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, Y. Chen, T. Lillicrap, F. Hui, L. Sifre, G. van den Driessche, T. Graepel, and D. Hassabis, Mas- tering the game of go without human knowledge, Nature 550, 354 (2017)

  6. [6]

    Silver, T

    D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis, Mastering chess and shogi by self-play with a general reinforcement learning algorithm, arXiv [cs.AI] (2017), arXiv:1712.01815 [cs.AI]

  7. [7]

    von Neumann, Zur theorie der gesellschaftsspiele, Math

    J. von Neumann, Zur theorie der gesellschaftsspiele, Math. Ann.100, 295 (1928)

  8. [8]

    L. S. Shapley, Stochastic games, Proc. Natl. Acad. Sci. U. S. A.39, 1095 (1953)

  9. [9]

    M. L. Littman, Markov games as a framework for multi- agent reinforcement learning, inMachine Learning Pro- ceedings 1994(Elsevier, 1994) pp. 157–163

  10. [10]

    M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, The arcade learning environment: an evaluation platform for general agents, J. Artif. Int. Res.47, 253–279 (2013)

  11. [11]

    V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, Play- ing atari with deep reinforcement learning, arXiv [cs.LG] (2013), arXiv:1312.5602 [cs.LG]

  12. [12]

    V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beat- tie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, Human-level con- trol through deep reinforcement learning, Nature518, 529 (2015)

  13. [13]

    B. J. Kagan, A. C. Kitchen, N. T. Tran, F. Habibol- lahi, M. Khajehnejad, B. J. Parker, A. Bhat, B. Rollo, A. Razi, and K. J. Friston, In vitro neurons learn and exhibit sentience when embodied in a simulated game- world, Neuron110, 3952 (2022)

  14. [14]

    L. K. Grover, A fast quantum mechanical algorithm for database search (1996), arXiv:quant-ph/9605043 [quant- ph]

  15. [15]

    E. R. Anschuetz, H.-Y. Hu, J.-L. Huang, and X. Gao, In- terpretable quantum advantage in neural sequence learn- ing, PRX Quantum4, 020338 (2023), arXiv:2209.14353 [quant-ph]

  16. [16]

    Bowles, S

    J. Bowles, S. Ahmed, and M. Schuld, Better than classical? the subtle art of benchmarking quantum machine learning models, arXiv [quant-ph] (2024), arXiv:2403.07059 [quant-ph]

  17. [17]

    P´ erez-Salinas, A

    A. P´ erez-Salinas, A. Cervera-Lierta, E. Gil-Fuster, and J. I. Latorre, Data re-uploading for a universal quantum classifier, Quantum4, 226 (2020), 1907.02085v2

  18. [18]

    Schuld, R

    M. Schuld, R. Sweke, and J. J. Meyer, Effect of data en- coding on the expressive power of variational quantum- machine-learning models, Phys. Rev. A103, 032430 (2021)

  19. [19]

    Saggio, B

    V. Saggio, B. E. Asenbeck, A. Hamann, T. Str¨ omberg, P. Schiansky, V. Dunjko, N. Friis, N. C. Harris, M. Hochberg, D. Englund, S. W¨ olk, H. J. Briegel, and P. Walther, Experimental quantum speed-up in reinforce- ment learning agents, Nature591, 229 (2021)

  20. [20]

    S. Wu, S. Jin, D. Wen, D. Han, and X. Wang, Quan- tum reinforcement learning in continuous action space, Quantum9, 1660 (2025), 2012.10711v5

  21. [21]

    Ganguly, Y

    B. Ganguly, Y. Xu, and V. Aggarwal, Quantum speedups in regret analysis of infinite horizon average-reward markov decision processes, inForty-second International Conference on Machine Learning(2025)

  22. [22]

    Sweke, J.-P

    R. Sweke, J.-P. Seifert, D. Hangleiter, and J. Eisert, On the quantum versus classical learnability of discrete dis- tributions, Quantum5, 417 (2021), 2007.14451v2

  23. [23]

    Pirnay, R

    N. Pirnay, R. Sweke, J. Eisert, and J.-P. Seifert, Su- perpolynomial quantum-classical separation for density modeling, Phys. Rev. A107, 042416 (2023)

  24. [24]

    Y. Liu, S. Arunachalam, and K. Temme, A rigorous and robust quantum speed-up in supervised machine learn- ing, Nat. Phys.17, 1013 (2021)

  25. [25]

    Wakeham and M

    D. Wakeham and M. Schuld, Inference, interference and invariance: How the quantum fourier transform can help to learn from data, arXiv [quant-ph] (2024), arXiv:2409.00172 [quant-ph]

  26. [26]

    R. S. Sutton, The bitter lesson,http://www. incompleteideas.net/IncIdeas/BitterLesson.html (2019), [Online; accessed 2025-12-10]

  27. [27]

    Chollet, On the measure of intelligence, arXiv [cs.AI] (2019), arXiv:1911.01547 [cs.AI]

    F. Chollet, On the measure of intelligence, arXiv [cs.AI] (2019), arXiv:1911.01547 [cs.AI]

  28. [28]

    Kaplan, S

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, Scaling laws for neural language models, arXiv [cs.LG] (2020), arXiv:2001.08361 [cs.LG]

  29. [29]

    Skolik, S

    A. Skolik, S. Jerbi, and V. Dunjko, Quantum agents in the gym: a variational quantum algorithm for deep Q- learning, Quantum6, 720 (2022), 2103.15084v3

  30. [30]

    Jerbi, C

    S. Jerbi, C. Gyurik, S. C. Marshall, H. J. Briegel, and V. Dunjko, Parametrized quantum policies for reinforce- ment learning, inAdvances in Neural Information Pro- cessing Systems(2021)

  31. [31]

    M. Nico, U. Christian, Y. George, K. Georgios, M. Christopher, and D. S. Daniel, Benchmarking quan- tum reinforcement learning, arXiv [quant-ph] (2025), arXiv:2501.15893 [quant-ph]

  32. [32]

    Jin, Z.-W

    Y.-X. Jin, Z.-W. Wang, H.-Z. Xu, W.-F. Zhuang, M.-J. Hu, and D. E. Liu, PPO-Q: Proximal policy optimiza- tion with parametrized quantum policies or values, arXiv [quant-ph] (2025), 2501.07085

  33. [33]

    Lazaro, J.-I

    J. Lazaro, J.-I. Vazquez, and P. Garcia-Bringas, Dis- secting quantum reinforcement learning: A system- atic evaluation of key components, arXiv [quant- ph] 10.48550/arXiv.2511.17112 (2025), arXiv:2511.17112 [quant-ph]

  34. [34]

    Y. Kubo, E. Chalmers, and A. Luczak, Combining back- propagation with equilibrium propagation to improve 13 an actor-critic reinforcement learning framework, Front. Comput. Neurosci.16, 980613 (2022)

  35. [35]

    Quantum Advantage Actor-Critic for Reinforcement Learning

    M. K¨ olle, M. Hagog, F. Ritz, P. Altmann, M. Zorn, J. Stein, and C. Linnhoff-Popien, Quantum advantage actor-critic for reinforcement learning, arXiv [quant- ph] 10.48550/arXiv.2401.07043 (2024), arXiv:2401.07043 [quant-ph]

  36. [36]

    Huang, R

    S. Huang, R. F. J. Dossa, C. Ye, J. Braga, D. Chakraborty, K. Mehta, and J. G. Ara´ ujo, Cleanrl: High-quality single-file implementations of deep re- inforcement learning algorithms, Journal of Machine Learning Research23, 1 (2022)

  37. [37]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, Proximal policy optimization algorithms, arXiv [cs.LG] (2017), arXiv:1707.06347 [cs.LG]

  38. [38]

    Bick,Towards Delivering a Coherent Self-Contained Explanation of Proximal Policy Optimization, Master’s thesis (2021)

    D. Bick,Towards Delivering a Coherent Self-Contained Explanation of Proximal Policy Optimization, Master’s thesis (2021)

  39. [39]

    J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, and H. Neven, Barren plateaus in quantum neural net- work training landscapes, Nat. Commun.9, 4812 (2018)

  40. [40]

    Kornblith, M

    S. Kornblith, M. Norouzi, H. Lee, and G. Hinton, Simi- larity of neural network representations revisited (2019), arXiv:1905.00414 [cs.LG]

  41. [41]

    Strobl, M

    M. Strobl, M. E. Sahin, L. van der Horst, E. Kuehn, A. Streit, and B. Jaderberg, Fourier fingerprints of ansatzes in quantum machine learning, arXiv [quant-ph] (2025), arXiv:2508.20868 [quant-ph]