REVIEW 1 major objections 4 minor 42 references
This paper tries to establish that RL-based quantum architecture search can be made VQE-efficient by preserving the exact, known circuit-transition dynamics and learning only the expensive post-VQE feedback, then using that learned feedback
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 05:56 UTC pith:ZGQ7YZLH
load-bearing objection A well-executed, careful model-based QAS paper whose efficiency gains mostly hold up; the unmeasured distribution shift in imagined rollouts is the main thing I'd ask for. the 1 major comments →
DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's discovery is that a feedback-only world model with exact symbolic circuit transitions can support efficient QAS policy learning. Instead of predicting ground-truth energy, the model learns a signed-log score relative to an empirical frontier, which is monotone, requires no ground-state energy, and provides improvements as rewards. The actor consumes multi-step imagined rollouts anchored on real VQE-verified prefixes; rewards are potential differences with an ensemble-pessimism term; a ranking gate switches imagination on only after pairwise accuracy suffices; and a small set of predicted-best and high-disagreement circuits is periodically verified with real VQE and returned to re
What carries the argument
The load-bearing object is a three-member recurrent ensemble over explicit gate-sequence prefixes. Each member encodes the full ordered circuit with a recurrent network and predicts a frontier-relative signed-log feedback score, combining a trainable head with a frozen randomized prior; the ensemble mean gives the predicted feedback and the standard deviation gives an epistemic uncertainty signal. That uncertainty enters as pessimism in a potential-based imagined reward, as a truncation threshold on low-confidence rollout steps, and as a risk-ranking signal for selective verification. The exact legal-transition rule is never learned, so transition error cannot compound with rollout length. A
Load-bearing premise
The learned feedback ensemble must keep ranking correctly on the imagined circuit prefixes the policy trains on, even though the policy's behavior shifts away from the real VQE-verified prefixes used to train and validate the model.
What would settle it
Measure the trained ensemble's pairwise ranking accuracy on the exact imagined prefixes generated at the end of training, evaluating those circuits with fresh real VQE: if ranking accuracy on those imagined prefixes is near chance while accuracy on real prefixes stays high, the imagined credit is not valid and the reported VQE savings are likely an artifact of reward shaping rather than genuine model-based credit assignment.
If this is right
- If correct, model-based RL-QAS can trade most real VQE evaluations for imagined rollouts anchored on verified prefixes, substantially reducing the cost of fine-error refinement.
- Because the training signal is frontier-relative and oracle-free, the approach does not require knowing the exact ground-state energy during search, widening applicability to molecules or hardware where that energy is unavailable.
- Direct greedy or beam exploitation of the same learned model does not recover the benefits; multi-step imagined policy learning is the mechanism, so predictor-assisted QAS should consume learned feedback through policy updates rather than through direct selection.
- Ensemble disagreement carries actionable information about prediction risk, giving a principled criterion for deciding which circuits deserve scarce real-VQE verification.
- The efficiency advantage grows as the target error tightens and is largest on the larger 8-qubit task, suggesting the method becomes more valuable as VQE evaluation becomes more expensive.
Where Pith is reading between the lines
- The same factorization—exact transition rules plus a learned expensive-feedback function—could apply to other sequential discrete search domains with costly evaluators, though this extension is not made in the paper.
- A direct diagnostic on the imagined prefixes the policy actually trains on would settle whether the reported savings come from valid imagined credit assignment or from benign reward shaping; the paper validates the model only on real VQE-verified prefixes.
- The emphasis on counterfactual action-ranking utility suggests that future QAS surrogate evaluation should measure decision utility rather than regression accuracy, a shift the paper's probe demonstrates but does not itself advocate as a general rule.
- A testable natural next step is to vary the selective verification rate and the pessimism strength on a new molecule to map how much of the improvement is attributable to each reliability control outside the five reported tasks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. DreamQAS is a model-based RL-QAS method that leaves circuit transitions exact and learns only the expensive post-VQE feedback via a recurrent randomized-prior ensemble. The learned signal is an oracle-free, frontier-relative signed-log score; the actor is trained on real rollouts plus multi-step imagined rollouts anchored at replay prefixes, with ranking-gated activation, pessimistic/truncated imagined returns, and selective real-VQE verification. On five HamQASBench molecular tasks under a common 15,000-episode budget and frozen evaluation, the paper reports the lowest mean frozen-policy energy error on four of five tasks, 1.6–2.0× (and 10.6× on BeH2-8q) fewer real VQE calls at matched fine-error targets, counterfactual action-ranking improvement, and same-model deployment controls showing that direct greedy/beam use of the model does not recover the imagined-policy gains.
Significance. If the claims hold, the paper makes a real contribution: it identifies a QAS-specific factorization of model-based RL in which known, exact circuit transitions are preserved and only post-VQE feedback is learned, and it argues that the model's value is decision utility rather than exact energy prediction. The experimental protocol is a strength: seed-level sustained crossings with three-point persistence and right-censoring, full reach counts, pre-specified four-level error ladders, paired bootstrap intervals, same-model deployment/horizon controls, and an unusually detailed artifact and data-lineage appendix. The paper is also candid about limitations (state-vector simulation, search-space floors, and the absence of a full calibration guarantee). The main open question is whether the mechanism conclusion about imagined rollouts is supported by the measurements actually reported.
major comments (1)
- [§4.3–§4.4, Eq. (13), Tables 2/3/18] The central mechanism claim—multi-step imagined policy learning, not direct surrogate exploitation, is what makes DreamQAS VQE-efficient—requires the ensemble to provide valid feedback on the imagined prefixes the actor actually trains on. The ranking gate (Eq. 13) and the counterfactual action probe (Table 2) are computed on real VQE-verified prefixes, and Table 18 is not split by real vs. imagined prefixes. Imagined rollouts are built by appending policy-sampled legal gates to verified prefixes; after a few steps the inputs lie on a shifted distribution that the policy itself moves while training on those rollouts. The paper never measures pairwise ranking accuracy, Spearman correlation, or calibration as a function of imagined depth. Consequently, the H=1 vs. H=5/15 contrast in Table 3(b) and the direct-greedy separation in Table 3(a) are also consistent with an inaccurate ensemble ac
minor comments (4)
- [§5.2, Fig. 2] The visual arrows in Fig. 2 (1.5×, 2.1×, 9.0×) differ from the seed-level numbers in Table 8 (1.60×, 1.97×, 10.58×). The caption says arrows are visual and numerical claims use the crossing estimator, but consider adding the exact target error to each arrow label so the reader can connect the two forms of evidence.
- [Appendix D, Eq. (22)] The FCI-referenced diagnostic r_FCI(t) = log10(E_t − E_0) − log10(E_{t+1} − E_0) is undefined if E_t equals E_0. The paper states that residual sign disagreement comes from six-digit logging precision, but it would be clearer to state the tie/zero handling explicitly.
- [Appendix D.2, Table 2] The final NReg values of 0.000 for BeH2-8q and BeH2-10q are accompanied by extensive ties in the candidate sets. The text notes this, but a per-seed tie count column would make the table self-contained and prevent over-interpretation of the zero regret.
- [Appendix H.4] The manuscript says the final repository commit 'should be exported from the frozen submission environment together with the code release.' For a reproducibility-focused appendix, provide the actual repository URL and commit hash in the manuscript rather than stating that they should be exported.
Circularity Check
No circular derivation: frontier-relative target is a disclosed label transform; decision-utility claims rest on real-VQE probes; self-citations are not load-bearing.
full rationale
The central derivation—learn a recurrent ensemble of post-VQE feedback and consume it via multi-step imagined policy gradients—is not circular. The supervised target y_t = S_k(E*(s_t)) is a deterministic, strictly monotone transform of the real optimized energy relative to an empirical frontier F_k adopted from observed runs (Eqs. 3–4). This is label normalization, not a fitted parameter: S is strictly increasing in E for fixed F_k, so the sign of the real reward r_real = y_t - y_{t+1} agrees with the sign of the true energy improvement, and the paper explicitly states that exact E_0 is used only after training for evaluation ("The exact E0 is used only after training to express frozen-policy evaluation in mHa"). The decision-utility claims are validated against matched real VQE in the counterfactual action probe (§5.3, Table 2), and the risk–coverage analysis is deliberately restricted to value-selected candidates so that disagreement is not evaluated on the same score used to select the sample (§F.1). The same-model deployment and transition-matched horizon controls are genuine ablations, not restatements of the training objective. The only quasi-self-referential elements are (i) a moving frontier label, which is disclosed and atomically relabeled in Appendix A.1, and (ii) self-citations to HamQASBench as the task source and HyRLQAS as a baseline; neither is used to justify a mathematical step or to forbid alternatives. The skeptic's concern that imagined-rollout rankings are never directly measured is a validity/limitation gap, not circularity: no equation equates the imagined reward to the final frozen-policy evaluation metric. Score 2 reflects the minor self-referential training signal and self-cited benchmark, not a circular derivation of the results.
Axiom & Free-Parameter Ledger
free parameters (9)
- score scale m =
0.1 mHa
- ensemble size K =
3
- randomized-prior scale β_rpf =
3.0
- pessimism coefficient β_pes =
1.0
- truncation threshold τσ =
0.60
- ranking gate threshold τ_pair =
0.70
- imagination horizon H =
15
- frontier refresh interval =
every 20 iterations (200 model grad steps)
- replay mixture weights =
0.25/0.35/0.20/0.20
axioms (6)
- domain assumption VQE-optimized energies E*(s) are well-defined, deterministic labels in noiseless state-vector simulation, and the simulator's per-call cost structure is the reported efficiency axis.
- standard math For d≥0 the signed-log score S(E) differs from the legacy log10(m+d) transform by a constant, so the oracle-free reward is monotone-compatible with energy ordering; for d<0 the branch is a logarithmic compression.
- domain assumption Atomic frontier adoption per refresh makes all supervised labels in one update share a single reference; historical circuits remain usable after relabeling.
- domain assumption BeH2-6q and BeH2-8q reach a search-space floor within 15,000 episodes, so zero-variance cells (0.058±<0.001, 2.17±0.00) are floor effects rather than evaluation artifacts.
- domain assumption FCI ground states E0 are externally computed (offline OpenFermion/PySCF) and used only for evaluation conversion, never for training.
- domain assumption The legality mask and stateless transition T_circ are exact and identical in real and imagined rollouts, so transition error cannot compound with imagination length.
read the original abstract
Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extending a circuit, although circuit construction and action legality are deterministic and known. We introduce DreamQAS, a model-based RL framework that preserves these exact circuit dynamics and learns only the expensive post-VQE feedback. A recurrent randomized-prior ensemble predicts an oracle-free score relative to an empirical energy frontier and supports multi-step imagined policy learning over explicit legal circuits. Ranking-based activation, uncertainty-aware pessimism and truncation, and selective real-VQE verification form a reliability-controlled learning loop. Under a common 15,000-episode budget and frozen evaluation for the RL methods, DreamQAS has the lowest mean frozen-policy energy error on four of five molecular tasks and the second-lowest on one. At fine-error targets reached by all seeds of both methods, it uses 1.6x to 2.0x fewer real VQE calls on four tasks and 10.6x fewer on BeH2-8q. Counterfactual action-ranking utility increases across all five tasks, with a mean increase of 0.346 and a 95 percent confidence interval of [0.185, 0.507], while direct greedy and beam use of the same model does not recover the gains of imagined policy learning. Ensemble disagreement also improves risk-coverage over random rejection on all three probed tasks. These results establish a world-model design for QAS whose value lies in decision-useful feedback rather than exact energy prediction.
Figures
Reference graph
Works this paper leans on
-
[1]
Nature communications , volume=
A variational eigenvalue solver on a photonic quantum processor , author=. Nature communications , volume=. 2014 , publisher=
2014
-
[2]
nature , volume=
Hardware-efficient variational quantum eigensolver for small molecules and quantum magnets , author=. nature , volume=. 2017 , publisher=
2017
-
[3]
Physics Reports , volume=
The variational quantum eigensolver: a review of methods and best practices , author=. Physics Reports , volume=. 2022 , publisher=
2022
-
[4]
Nature Reviews Physics , volume=
Variational quantum algorithms , author=. Nature Reviews Physics , volume=. 2021 , publisher=
2021
-
[5]
Quantum , volume=
Quantum computing in the NISQ era and beyond , author=. Quantum , volume=. 2018 , publisher=
2018
-
[6]
Quantum Science & Technology , volume=
Differentiable quantum architecture search , author=. Quantum Science & Technology , volume=. 2022 , publisher=
2022
-
[7]
Advanced Quantum Technologies , volume=
Expressibility and entangling capability of parameterized quantum circuits for hybrid quantum-classical algorithms , author=. Advanced Quantum Technologies , volume=. 2019 , publisher=
2019
-
[8]
arXiv preprint arXiv:2607.04845 , year=
HamQASBench: A Hamiltonian-Informed Diagnostic Benchmark for Evaluating Quantum Architecture Search , author=. arXiv preprint arXiv:2607.04845 , year=
-
[9]
npj Quantum Information , volume=
Quantum circuit architecture search for variational quantum algorithms , author=. npj Quantum Information , volume=. 2022 , publisher=
2022
-
[10]
Openreview , year=
Reinforcement Learning for Quantum Circuit Optimization: A Review , author=. Openreview , year=
-
[11]
Advances in neural information processing systems , volume=
Reinforcement learning for optimization of variational quantum circuit architectures , author=. Advances in neural information processing systems , volume=
-
[13]
arXiv preprint arXiv:2402.03500 , year=
Curriculum reinforcement learning for quantum architecture search under hardware errors , author=. arXiv preprint arXiv:2402.03500 , year=
-
[14]
International conference on machine learning , pages=
Quantumdarts: differentiable quantum architecture search for variational quantum algorithms , author=. International conference on machine learning , pages=. 2023 , organization=
2023
-
[15]
arXiv preprint arXiv:2401.09253 , year=
The generative quantum eigensolver (GQE) and its application for ground state search , author=. arXiv preprint arXiv:2401.09253 , year=
-
[16]
Machine Learning: Science and Technology , volume=
Neural predictor based quantum architecture search , author=. Machine Learning: Science and Technology , volume=. 2021 , publisher=
2021
-
[17]
EPJ Quantum Technology , volume=
Trainability maximization using estimation of distribution algorithms assisted by surrogate modelling for quantum architecture search , author=. EPJ Quantum Technology , volume=. 2024 , publisher=
2024
-
[18]
arXiv preprint arXiv:2506.06762 , year=
Benchmarking quantum architecture search with surrogate assistance , author=. arXiv preprint arXiv:2506.06762 , year=
-
[19]
arXiv preprint arXiv:1803.10122 , volume=
World models , author=. arXiv preprint arXiv:1803.10122 , volume=
-
[20]
International Conference on Learning Representations , year=
Dream to Control: Learning Behaviors by Latent Imagination , author=. International Conference on Learning Representations , year=
-
[21]
Nature , volume=
Mastering atari, go, chess and shogi by planning with a learned model , author=. Nature , volume=. 2020 , publisher=
2020
-
[22]
Advances in neural information processing systems , volume=
When to trust your model: Model-based policy optimization , author=. Advances in neural information processing systems , volume=
-
[23]
arXiv preprint arXiv:2511.04967 , year=
Hybrid Action Reinforcement Learning for Quantum Architecture Search , author=. arXiv preprint arXiv:2511.04967 , year=
-
[24]
International conference on machine learning , pages=
Qas-bench: rethinking quantum architecture search and a benchmark , author=. International conference on machine learning , pages=. 2023 , organization=
2023
-
[25]
Proceedings of the AAAI Symposium Series , volume=
BenchRL-QAS: Benchmarking reinforcement learning algorithms for quantum architecture search , author=. Proceedings of the AAAI Symposium Series , volume=
-
[26]
He et al
A GNN-based predictor for quantum architecture search: Z. He et al. , author=. Quantum Information Processing , volume=. 2023 , publisher=
2023
-
[27]
The European Physical Journal Plus , volume=
A progressive predictor-based quantum architecture search with active learning , author=. The European Physical Journal Plus , volume=. 2023 , publisher=
2023
-
[28]
Proceedings of the AAAI conference on artificial intelligence , volume=
Training-free quantum architecture search , author=. Proceedings of the AAAI conference on artificial intelligence , volume=
-
[29]
Machine Learning: Science and Technology , volume=
Reinforcement learning-based architecture search for quantum machine learning , author=. Machine Learning: Science and Technology , volume=. 2025 , publisher=
2025
-
[30]
1998 , publisher=
Introduction to reinforcement learning , author=. 1998 , publisher=
1998
-
[31]
Proceedings of the fourteenth international conference on artificial intelligence and statistics , pages=
A reduction of imitation learning and structured prediction to no-regret online learning , author=. Proceedings of the fourteenth international conference on artificial intelligence and statistics , pages=. 2011 , organization=
2011
-
[32]
Nature communications , volume=
An adaptive variational algorithm for exact molecular simulations on a quantum computer , author=. Nature communications , volume=. 2019 , publisher=
2019
-
[33]
qubit-adapt-vqe: An adaptive algorithm for constructing hardware-efficient ans
Tang, Ho Lun and Shkolnikov, VO and Barron, George S and Grimsley, Harper R and Mayhall, Nicholas J and Barnes, Edwin and Economou, Sophia E , journal=. qubit-adapt-vqe: An adaptive algorithm for constructing hardware-efficient ans. 2021 , publisher=
2021
-
[34]
Nature , volume=
Mastering diverse control tasks through world models , author=. Nature , volume=. 2025 , publisher=
2025
-
[35]
Advances in neural information processing systems , volume=
Simple and scalable predictive uncertainty estimation using deep ensembles , author=. Advances in neural information processing systems , volume=
-
[36]
Advances in neural information processing systems , volume=
Randomized prior functions for deep reinforcement learning , author=. Advances in neural information processing systems , volume=
-
[37]
International conference on machine learning , pages=
Delving into deep imbalanced regression , author=. International conference on machine learning , pages=. 2021 , organization=
2021
-
[38]
Machine learning , volume=
Simple statistical gradient-following algorithms for connectionist reinforcement learning , author=. Machine learning , volume=. 1992 , publisher=
1992
-
[39]
Icml , volume=
Policy invariance under reward transformations: Theory and application to reward shaping , author=. Icml , volume=. 1999 , organization=
1999
-
[40]
Breakthroughs in statistics: Methodology and distribution , pages=
Robust estimation of a location parameter , author=. Breakthroughs in statistics: Methodology and distribution , pages=. 1992 , publisher=
1992
-
[41]
Advances in neural information processing systems , volume=
Mopo: Model-based offline policy optimization , author=. Advances in neural information processing systems , volume=
-
[42]
ACM Sigart Bulletin , volume=
Dyna, an integrated architecture for learning, planning, and reacting , author=. ACM Sigart Bulletin , volume=. 1991 , publisher=
1991
-
[43]
Advances in neural information processing systems , volume=
Deep reinforcement learning in a handful of trials using probabilistic dynamics models , author=. Advances in neural information processing systems , volume=
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.