{"id":"6ea538d9-b3f4-4e5f-b785-864cb524a435","arxiv_id":"2607.29491","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"DreamQAS reaches fine VQE error targets with 1.6–10.6× fewer real optimizer calls than a no-imagination control by learning only the energy-feedback score and imagining policy rollouts over exact circuit rules.","lead":"This paper trains an AI that designs quantum circuits for calculating molecular energies, while avoiding the expensive quantum optimizer for most design steps by modeling only the energy score and keeping the known circuit rules exact. It shows that imagining many circuit candidates before running real checks can cut the number of real quantum optimizations by up to ten times at fine accuracy targets.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Imagined-rollout validity is never measured: the ensemble's ranking is validated on real VQE-verified prefixes (Eq. 13), not on the imagined prefixes the policy trains on (§4.3). RQ1/RQ3 gains could then be reward-shaping artifacts rather than valid model-based credit assignment.","rationale":"The paper is unusually careful: matched internal controls, common 15k budget, frozen evaluation, seed-level sustained crossings with reach counts, counterfactual probes, same-model deployment controls, and full ladders give strong internal support for the reported outcome claims. The selected fine-target savings in Table 8 are not just medians over overlapping seeds; the ranges for the two methods are separated at those rungs, so the direction of the efficiency effect is robust at the chosen targets. The weakest point is not the outcome statistics but the mechanism: the model's decision utility is measured where real labels exist, while the policy is trained on imagined continuations that are never real-VQE-labeled. This is exactly the reader's weakest_assumption. Pessimism, truncation, and DAgger are sensible mitigations but they are not measurements of imagined-prefix validity; the proposed diagnostic would close that gap. Because the reader already marks the verdict CONDITIONAL, no verdict change is needed; the condition should explicitly include imagined-prefix validation. The absence of a public code/artifact release (App. H.4 is future tense) makes this diagnostic impossible to run today, which further supports keeping the verdict conditional rather than accepting the mechanism claim outright.","tokens_in":25194,"tokens_out":16329,"duration_ms":175516,"concrete_test":"At the final checkpoint, log all imagined rollouts generated during training. Sample, per task, a matched set of imagined prefixes from each rollout step k=1,5,10,15 (e.g., 50 prefixes × up to 5 legal continuations per step). Evaluate those continuations with real VQE (diagnostic calls excluded from budget) and compute pairwise ranking accuracy and Spearman ρ_act between ensemble predictions and real energies, overall and per k. Compare against the same analysis on real replay prefixes at matched lengths and against the 0.70 gate threshold. If imagined-prefix accuracy stays ≥0.70 and does not decay with k, the premise holds; if it falls well below, the RQ1/RQ3 gains are not explained by valid imagined credit assignment.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central RQ1/RQ3 narrative is that multi-step imagined policy learning converts learned post-VQE feedback into improved search, and that this is what makes DreamQAS VQE-efficient. The ranking gate (Eq. 13) and the counterfactual action probe (§5.3, Table 2) both measure the ensemble's decision utility on real VQE-verified prefixes sampled from the replay/policy distribution. But the imagined rollouts in §4.3 are constructed by taking a verified replay prefix and appending policy-sampled legal gates under the exact transition Eq. (1). After even a few such appends, the resulting prefixes lie in a shifted distribution — one the policy itself is changing as it trains on those very rollouts. The paper never evaluates the ensemble's pairwise ranking accuracy or Spearman correlation on these imagined prefixes; the fidelity/calibration diagnostics in Table 18 are not split by imagined vs. real. This is load-bearing because the imagined reward (Eq. 11) is a potential difference: if the learned potential were accurate on imagined states, one-step imagined policy learning and direct greedy deployment should already capture most of the gain; Table 3 shows H=1 and especially direct greedy do not. That gap is consistent with the model being reliable only near the real-prefix training distribution, and the H=15 improvement could come from an inaccurate potential acting as helpful reward shaping rather than from valid credit assignment. Pessimism, truncation, and DAgger mitigate the risk empirically but do not measure imagined-prefix validity itself.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"DreamQAS is a model-based RL-QAS method that leaves circuit transitions exact and learns only the expensive post-VQE feedback via a recurrent randomized-prior ensemble. The learned signal is an oracle-free, frontier-relative signed-log score; the actor is trained on real rollouts plus multi-step imagined rollouts anchored at replay prefixes, with ranking-gated activation, pessimistic/truncated imagined returns, and selective real-VQE verification. On five HamQASBench molecular tasks under a common 15,000-episode budget and frozen evaluation, the paper reports the lowest mean frozen-policy energy error on four of five tasks, 1.6–2.0× (and 10.6× on BeH2-8q) fewer real VQE calls at matched fine-error targets, counterfactual action-ranking improvement, and same-model deployment controls showing that direct greedy/beam use of the model does not recover the imagined-policy gains.","tokens_in":25400,"tokens_out":8799,"duration_ms":112581,"significance":"If the claims hold, the paper makes a real contribution: it identifies a QAS-specific factorization of model-based RL in which known, exact circuit transitions are preserved and only post-VQE feedback is learned, and it argues that the model's value is decision utility rather than exact energy prediction. The experimental protocol is a strength: seed-level sustained crossings with three-point persistence and right-censoring, full reach counts, pre-specified four-level error ladders, paired bootstrap intervals, same-model deployment/horizon controls, and an unusually detailed artifact and data-lineage appendix. The paper is also candid about limitations (state-vector simulation, search-space floors, and the absence of a full calibration guarantee). The main open question is whether the mechanism conclusion about imagined rollouts is supported by the measurements actually reported.","major_comments":[{"comment":"The central mechanism claim—multi-step imagined policy learning, not direct surrogate exploitation, is what makes DreamQAS VQE-efficient—requires the ensemble to provide valid feedback on the imagined prefixes the actor actually trains on. The ranking gate (Eq. 13) and the counterfactual action probe (Table 2) are computed on real VQE-verified prefixes, and Table 18 is not split by real vs. imagined prefixes. Imagined rollouts are built by appending policy-sampled legal gates to verified prefixes; after a few steps the inputs lie on a shifted distribution that the policy itself moves while training on those rollouts. The paper never measures pairwise ranking accuracy, Spearman correlation, or calibration as a function of imagined depth. Consequently, the H=1 vs. H=5/15 contrast in Table 3(b) and the direct-greedy separation in Table 3(a) are also consistent with an inaccurate ensemble ac","section":"§4.3–§4.4, Eq. (13), Tables 2/3/18"}],"minor_comments":[{"comment":"The visual arrows in Fig. 2 (1.5×, 2.1×, 9.0×) differ from the seed-level numbers in Table 8 (1.60×, 1.97×, 10.58×). The caption says arrows are visual and numerical claims use the crossing estimator, but consider adding the exact target error to each arrow label so the reader can connect the two forms of evidence.","section":"§5.2, Fig. 2"},{"comment":"The FCI-referenced diagnostic r_FCI(t) = log10(E_t − E_0) − log10(E_{t+1} − E_0) is undefined if E_t equals E_0. The paper states that residual sign disagreement comes from six-digit logging precision, but it would be clearer to state the tie/zero handling explicitly.","section":"Appendix D, Eq. (22)"},{"comment":"The final NReg values of 0.000 for BeH2-8q and BeH2-10q are accompanied by extensive ties in the candidate sets. The text notes this, but a per-seed tie count column would make the table self-contained and prevent over-interpretation of the zero regret.","section":"Appendix D.2, Table 2"},{"comment":"The manuscript says the final repository commit 'should be exported from the frozen submission environment together with the code release.' For a reproducibility-focused appendix, provide the actual repository URL and commit hash in the manuscript rather than stating that they should be exported.","section":"Appendix H.4"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the stress-test concern is valid and load-bearing. The paper's empirical claims for RQ1 are carefully supported, but the mechanism narrative in RQ3 depends on an imagined-prefix validity measurement that is not present. The missing piece is a well-defined additional experiment rather than a reanalysis or correction, so I recommend major revision rather than rejection. The manuscript is otherwise unusually disciplined in its statistics and artifact reporting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth engaging: it gives a credible recipe for cutting real VQE calls in RL-QAS while keeping circuit transitions exact, and it evaluates that recipe with more care than most ML-QAS papers. The core idea—preserve the symbolic circuit dynamics and learn only the expensive post-VQE feedback—is simple but well executed. The new assembly is the frontier-relative oracle-free score, the recurrent ensemble, and imagination with pessimism/truncation/DAgger. All of that is genuinely new in the QAS literature.\n\nThe experimental protocol is the strongest part. Seed-level sustained crossings with three-point persistence, full reach counts, right-censoring of unreached seeds, pre-specified error ladders, paired bootstrap intervals for most contrasts, and counterfactual action-utility probes: that is a serious attempt to make the efficiency claim reproducible. I believe the RQ1 claim is mostly supported for the matched internal controls: at fine-error targets reached by all seeds, DreamQAS uses fewer real VQE calls than NoImag. The effect is largest at the tightest rung, which is where the model's filtering should matter, so that is not necessarily a red flag.\n\nWhere I'd push back on the paper's own story is the imagined-rollout validity premise. The ensemble's ranking is validated on real VQE-verified prefixes, not on the imagined prefixes the policy actually trains on. The policy changes the prefix distribution as it learns, and the paper never measures ranking accuracy on those shifted prefixes. Pessimism, truncation, and DAgger mitigate the risk, but they do not measure it. So the H=15 gain over H=1 could be partly reward-shaping rather than valid credit assignment. That's not fatal—the empirical gains stand on their own—but it means the mechanism story is less secure than the efficiency story.\n\nOther soft spots: the headline savings ratios are medians without confidence intervals, and the 10.6x on BeH2-8q is at a task with floor-like cells; two of the 'lowest on four' tasks have zero-variance floor cells, making those comparisons non-informative. External QAS baselines are not in the ranking, so validation is against the authors' own ablations. And there is no code, data, or commit hash—the paper explicitly says the appendix is future tense. None of the runs can be independently checked.\n\nWho should read this: anyone working on predictor-assisted QAS, model-based RL for circuit search, or amortized VQE evaluation. It deserves a serious referee: the claims are clear, the protocol is detailed, and the main weakness is a missing diagnostic rather than an obvious error. I would send it to review with a request for the imagined-prefix validity analysis and code release.","headline":"A well-executed, careful model-based QAS paper whose efficiency gains mostly hold up; the unmeasured distribution shift in imagined rollouts is the main thing I'd ask for.","tokens_in":181,"tokens_out":2251,"would_cite":true,"duration_ms":32241,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper tries to establish that RL-based quantum architecture search can be made VQE-efficient by preserving the exact, known circuit-transition dynamics and learning only the expensive post-VQE feedback, then using that learned feedback","keywords":["quantum architecture search","variational quantum eigensolver","model-based reinforcement learning","world model","uncertainty-aware ensemble","oracle-free reward","circuit optimization","sample efficiency"],"falsifier":"Measure the trained ensemble's pairwise ranking accuracy on the exact imagined prefixes generated at the end of training, evaluating those circuits with fresh real VQE: if ranking accuracy on those imagined prefixes is near chance while accuracy on real prefixes stays high, the imagined credit is not valid and the reported VQE savings are likely an artifact of reward shaping rather than genuine model-based credit assignment.","tokens_in":24897,"feed_emoji":"🧪","tokens_out":5629,"duration_ms":67203,"temperature":0.7,"pith_summary":"The paper tries to show that RL-based quantum architecture search can be made much less expensive without sacrificing final circuit quality. Its central move is to split the problem: the rule for extending a circuit is known and exact, so only the expensive energy feedback after VQE optimization needs to be learned. A recurrent ensemble predicts an oracle-free, frontier-relative score, and the policy trains on imagined rollouts that use this predicted feedback, with ensemble disagreement used to down-weight or reject unreliable predictions. On five molecular tasks under a shared 15,000-episode budget, this yields the lowest mean frozen-policy energy error on four tasks and second-lowest on one, and reaches matched fine-error targets with 1.6x to 2.0x fewer real VQE calls on four tasks and 10.6x fewer on one. The broader claim is that a world model for QAS should aim for decision-useful ranking, not exact energy prediction.","feed_headline":"A dreamed world model cuts real VQE calls up to 10.6x","feed_subtitle":"Keeping circuit rules exact while learning only VQE feedback gives top results on four of five molecules.","key_machinery":"The load-bearing object is a three-member recurrent ensemble over explicit gate-sequence prefixes. Each member encodes the full ordered circuit with a recurrent network and predicts a frontier-relative signed-log feedback score, combining a trainable head with a frozen randomized prior; the ensemble mean gives the predicted feedback and the standard deviation gives an epistemic uncertainty signal. That uncertainty enters as pessimism in a potential-based imagined reward, as a truncation threshold on low-confidence rollout steps, and as a risk-ranking signal for selective verification. The exact legal-transition rule is never learned, so transition error cannot compound with rollout length. A","core_discovery":"The paper's discovery is that a feedback-only world model with exact symbolic circuit transitions can support efficient QAS policy learning. Instead of predicting ground-truth energy, the model learns a signed-log score relative to an empirical frontier, which is monotone, requires no ground-state energy, and provides improvements as rewards. The actor consumes multi-step imagined rollouts anchored on real VQE-verified prefixes; rewards are potential differences with an ensemble-pessimism term; a ranking gate switches imagination on only after pairwise accuracy suffices; and a small set of predicted-best and high-disagreement circuits is periodically verified with real VQE and returned to re","pith_inferences":["The same factorization—exact transition rules plus a learned expensive-feedback function—could apply to other sequential discrete search domains with costly evaluators, though this extension is not made in the paper.","A direct diagnostic on the imagined prefixes the policy actually trains on would settle whether the reported savings come from valid imagined credit assignment or from benign reward shaping; the paper validates the model only on real VQE-verified prefixes.","The emphasis on counterfactual action-ranking utility suggests that future QAS surrogate evaluation should measure decision utility rather than regression accuracy, a shift the paper's probe demonstrates but does not itself advocate as a general rule.","A testable natural next step is to vary the selective verification rate and the pessimism strength on a new molecule to map how much of the improvement is attributable to each reliability control outside the five reported tasks."],"forward_implications":["If correct, model-based RL-QAS can trade most real VQE evaluations for imagined rollouts anchored on verified prefixes, substantially reducing the cost of fine-error refinement.","Because the training signal is frontier-relative and oracle-free, the approach does not require knowing the exact ground-state energy during search, widening applicability to molecules or hardware where that energy is unavailable.","Direct greedy or beam exploitation of the same learned model does not recover the benefits; multi-step imagined policy learning is the mechanism, so predictor-assisted QAS should consume learned feedback through policy updates rather than through direct selection.","Ensemble disagreement carries actionable information about prediction risk, giving a principled criterion for deciding which circuits deserve scarce real-VQE verification.","The efficiency advantage grows as the target error tightens and is largest on the larger 8-qubit task, suggesting the method becomes more valuable as VQE evaluation becomes more expensive."],"fun_headline_variants":["DreamQAS imagines VQE feedback, cuts real calls 10.6x","World model that learns only VQE feedback: 10.6x fewer real calls","RL imagines VQE outcomes to slash real calls up to 10.6x","Decision-useful world model: 10.6x fewer VQE calls"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The learned feedback ensemble must keep ranking correctly on the imagined circuit prefixes the policy trains on, even though the policy's behavior shifts away from the real VQE-verified prefixes used to train and validate the model.","fun_headline_variants_meta":{"raw":{"variants":["DreamQAS imagines VQE feedback, cuts real calls 10.6x","World model that learns only VQE feedback: 10.6x fewer real calls","RL imagines VQE outcomes to slash real calls up to 10.6x","Decision-useful world model: 10.6x fewer VQE calls"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000534,"raw_usage":{"total_tokens":2439,"prompt_tokens":810,"completion_tokens":1629,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":1540}},"tokens_in":554,"tokens_out":1629,"duration_ms":12676,"temperature":1.0,"reasoning_tokens":1540,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T05:56:38.435414+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the trained ensemble's pairwise ranking accuracy on the exact imagined prefixes generated at the end of training, evaluating those circuits with fresh real VQE: if ranking accuracy on those imagined prefixes is near chance while accuracy on real prefixes stays high, the imagined credit is not valid and the reported VQE savings are likely an artifact of reward shaping rather than genuine model-based credit assignment.","supporting_citations":[],"review_version":1}