{"id":"e7388a95-700a-4eab-8e8e-df486109b48a","arxiv_id":"2501.12130","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A hybrid ansatz that multiplies a Transformer-based neural quantum state by a circuit-generated amplitude achieves lower ground-state energy errors than the neural network alone on small spin chains and LiH.","lead":"This paper combines a parameterized quantum circuit with a neural network to approximate ground states of spin systems and small molecules, and reports lower variational energies than the neural network alone. The result is a candidate recipe for using near-term quantum processors inside variational quantum Monte Carlo simulations.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'quantum-enhanced expressivity' claim is not isolated from added parameter capacity: no equal-parameter classical control is tested, so the improvement in Figs. 5 and 7 could come from the PQC's 60-70 extra parameters rather than from quantumness.","rationale":"The reader's weakest assumption correctly identifies the central gap: no classical NQS with the same number of additional parameters as the PQC is tested, so the 'quantum-enhanced' conclusion is confounded by added capacity. My stress-test pass found no more load-bearing issue. The manuscript's stronger claims—'quantum-enhanced neural quantum states', 'notably lower relative energy', and 'enhanced expressive power provided by quantum circuits'—all rest on Fig. 5 and Fig. 7, both of which compare against a deliberately small NQS or against a much larger classical NQS, never an equal-capacity classical control. The importance-sampling derivation in Sec. II C is internally coherent, the shot-noise scaling tests in Fig. 2 are informative, and the code release is a positive independent artifact. However, those strengths do not establish the expressivity claim. A conditional verdict is therefore appropriate: the framework and numerics are worth reporting, but the central interpretive claim should be either softened to 'adding a PQC improves accuracy' or backed by the missing parameter-matched classical baseline. I see no reason to move the reader's verdict; CONDITIONAL remains the correct assessment, and the concrete test above would settle whether the concern lands.","tokens_in":18303,"tokens_out":7441,"duration_ms":77946,"concrete_test":"Reproduce Fig. 5 for LiH at 2.4 Å using the Table I hyperparameters with a purely classical NQS of the same total parameter count as the hybrid (about 495 parameters), e.g. a d=3 Transformer with the phase FFNN widened by roughly 60 parameters, or a d=4, T=1 Transformer (286 parameters) plus the same [16,8] FFNN, trained with the same Adam schedule, sample size 10^4, and the same random seeds. If this classical control reaches below chemical accuracy (1.6e-3 Ha) or approaches the hybrid's 6.0e-5 Ha, the improvement in Fig. 5 is explained by added capacity rather than by quantum expressivity. As a second check, repeat for the 7-spin AFH setup of Appendix C, matching the 70 PQC gate parameters with 70 additional classical parameters, and compare the resulting relative errors.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim (Abstract; Sec. III; Fig. 5; App. C, Fig. 7) is that a PQC 'enhances the expressive power' of an NQS beyond what a classical network alone can achieve. The evidence is a two-point comparison: NQS-d=3 (179 Transformer parameters + 256 phase-network parameters = 435 total; error 5.8e-2 Ha) versus Hybrid-d=3-Nl=4 (same classical network plus 60 PQC parameters; error 6.0e-5 Ha), with a larger NQS (6474 parameters; error 9.4e-4 Ha) as an upper bound. This design never controls for the PQC's 60 added degrees of freedom. A classical NQS with the same number of additional parameters—for example a slightly larger embedding dimension, wider phase network, or an added classical product ansatz—is not trained. Since increasing the classical parameter count from 435 to 6474 already lowers the error by roughly two orders of magnitude, parameter count is clearly a dominant factor in this regime. Whether the remaining gap to 6.0e-5 Ha is due to quantum-specific correlations is therefore untested. The sequential optimization result in Fig. 4 has the same gap: it shows that adding PQC layers after freezing the pretrained NQS improves the energy, but the proper controls—continuing to train the frozen NQS for the same number of iterations, or adding an equivalent number of classical parameters—are absent. Consequently, the 'quantum-enhanced expressivity' conclusion in the Abstract and Discussion is not established by the reported comparisons. The weaker claim that 'a PQC can augment a constrained NQS' may be true, but it is agnostic as to whether the augmentation is quantum in nature. This is a missing-control/underdetermination issue, not an internal inconsistency: the importance-sampling formalism and the numerical implementation are coherent, and the code release is valuable, but they do not resolve the control problem.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a hybrid quantum-classical variational ansatz for many-body ground-state calculations. The wavefunction is written as a product of a Parameterized Quantum Circuit (PQC) contribution and a classical Neural Quantum State (NQS), where the NQS is a Transformer autoregressive sampler supplemented by a feedforward phase network. The authors introduce a Transformer-guided importance-sampling estimator (Eqs. 8--13), address variance control through small parameter initialization and a tanh rescaling, and exploit symmetries to reduce the sampling space. Numerical tests are reported for anti-ferromagnetic Heisenberg chains with 2--8 spins and for LiH in STO-3G with an active space. The headline numerical result is that a deliberately small NQS augmented by a four-layer PQC reaches LiH energy errors near 6.0e-5 Ha, whereas the same small NQS alone has errors around 5.8e-2 Ha and a much larger NQS alone only reaches 9.4e-4 Ha. The paper concludes that a PQC enhances the expressive power of the NQS and allows higher accuracy with significantly fewer parameters.","tokens_in":18673,"tokens_out":6138,"duration_ms":68947,"significance":"If the central claim were cleanly established, this would be a useful contribution to the growing literature on hybrid quantum-classical variational methods. The importance-sampling formalism is coherent, the numerical pipeline is described in enough detail to be reproducible, and the source code is stated to be available publicly. The use of 100 independent seeds for the main error statistics is a clear strength. However, the significance of the paper currently rests on an interpretational claim that is not isolated by the reported comparisons: the hybrid ansatz is always compared against classical NQS baselines with either far fewer or far more parameters, and no equal-capacity classical control is tested. The numerical results are consistent with a more modest statement about parameter efficiency, but they do not yet establish that the PQC contributes expressivity that a classical network with the same added capacity could not provide.","major_comments":[{"comment":"The central claim that the PQC 'enhances the expressive power' of the NQS is not isolated from added parameter capacity. The comparison shown is NQS-d=3 (435 classical parameters, error 5.8e-2 Ha) versus Hybrid-d=3-Nl=4 (the same classical network plus a PQC with 60 additional parameters, error 6.0e-5 Ha), with NQS-d=8 (6474 parameters, error 9.4e-4 Ha) as an upper bound. No control of a classical NQS augmented by the same number of additional trainable degrees of freedom is reported. Since increasing the classical parameter count from 435 to 6474 already lowers the error by approximately two orders of magnitude, the observed improvement could be explained by added capacity alone. A decisive control would be a classically augmented NQS with a comparable number of extra parameters, trained under the identical protocol; without it, the 'quantum-enhanced expressivity' conclusion in the Abstract and Section III is not established.","section":"Section III, Fig. 5"},{"comment":"The sequential-optimization experiment shows that freezing the NQS and then adding PQC layers lowers the energy, but the proper classical controls are absent. One should either continue training the frozen NQS for the same number of iterations, or add an equivalent number of classical parameters after pretraining, e.g., additional Transformer blocks, wider phase-network layers, or a classical product factor. Without such controls, the improvement observed as Nl increases from 0 to 4 is compatible with the added capacity of the PQC layers rather than with any quantum-specific expressivity.","section":"Section III, Fig. 4"},{"comment":"The text states that increasing the Transformer embedding dimension from 4 to 8 for the 7-spin AFH chain 'yields comparable accuracy' to the hybrid ansatz, while increasing the classical parameter count to O(10^3). This is a potentially important comparison and should be reported quantitatively. If the d=8 NQS indeed matches the hybrid's relative error near 10^-3, then the Section III claim of 'higher accuracy with significantly fewer parameters' should be rephrased as a parameter-efficiency result; if it does not, the numerical values need to be given. As written, the caption and text leave the comparison ambiguous, and this ambiguity directly affects the main conclusion.","section":"Appendix C, Fig. 7"}],"minor_comments":[{"comment":"The word 'normarlized' should read 'normalized'.","section":"Section II.A"},{"comment":"The Heisenberg Hamiltonian as written appears to contain index typos: the last two terms should likely be sigma^Y_i sigma^Y_j and sigma^Z_i sigma^Z_j, while the printed text has sigma^Y_i sigma^Y_i and sigma^X_i sigma^X_i. The equation should be corrected to match the model actually simulated.","section":"Section III, Eq. (20)"},{"comment":"The row labeled 'Qubit size 2' is unclear for a six-qubit LiH calculation; this presumably refers to the two-dimensional qubit-state embedding, but it should be clarified so that the total qubit count is unambiguous.","section":"Appendix D, Table I"},{"comment":"It would be helpful to state explicitly which hyperparameters are shared between the NQS-only runs and the hybrid runs in Figs. 5--7, not only for the LiH run in the table but also for the AFH runs.","section":"Appendix D, Table I caption"}],"recommendation":"major_revision","confidential_remarks":"The implementation appears competent and the numerical data are plausibly correct, but the manuscript's central conceptual claim is currently an interpretation of comparisons that do not control for parameter capacity. I would recommend requesting the missing equal-capacity classical controls before publication, since the added-cost experiments are well within the scope of the paper and would substantially strengthen the conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The construction is solid: Transformer autoregressive sampler multiplied by a PQC diagonal amplitude, importance-sampling formalism in Eqs. (8)-(13), symmetry masking for electron-number conservation, code on GitHub, and new LiH/AFH benchmarks. It cites the closest prior product-ansatz work (refs. 44 and 71) fairly instead of burying it.\n\nThe soft spot is the headline. \"Quantum-enhanced expressivity\" is not isolated from added parameter capacity. The evidence is NQS-d=3 (435 parameters, 5.8e-2 Ha error) against the hybrid (435 + 60 PQC parameters, 6.0e-5 Ha), with a 6474-parameter NQS (9.4e-4 Ha) as the larger classical reference. No classical NQS with an equivalent 60-parameter augmentation is trained. So the gap could in principle be capacity rather than quantumness. That said, the picture is not as one-sided as the stress-test note suggests: the hybrid at 495 total parameters beats the 6474-parameter classical NQS by more than an order of magnitude. If raw parameter count dominated, the big NQS should win. So the PQC's functional form is doing something; what is untested is whether 60 classical parameters of the same form (say a small multiplicative correction) would do the same. The authors never run that control.\n\nAlso worth noting: the strength of the claim varies inside the paper. The Abstract and Discussion say quantum circuits provide \"enhanced expressivity\"; the Results section more cautiously says a PQC \"can effectively augment\" a constrained NQS. The numerics support the second reading, not the first. For 6-7 qubits the PQC's diagonal amplitude is a classically computable function of s and theta, so \"quantum\" in any strong sense is not established. The paper should just say \"hybrid achieves lower energy on these systems,\" which is true and useful.\n\nMechanical issues, in order of seriousness. Eq. (20) is mangled: the Heisenberg Hamiltonian as printed contains sigma_Y_i sigma_Y_i and a duplicated sigma_X_i sigma_X_i term. The 60-parameter count implies U1 and U2 share gate parameters with separate coefficients, but the text only says they have \"the same structure\"; that should be explicit. Minor: all PQC measurements are simulated, which is fine for a proposal but not flagged as a limitation.\n\nWho this is for: people working on hybrid quantum-classical variational methods or NQS who want a concrete, reproducible instance of the product ansatz with a genuinely clever sampling scheme. It deserves a serious referee. The formalism is coherent, the code is released, and the missing control is a well-defined, cheap request. I would send it to review with two demands: an equal-capacity classical baseline, and claims cut back to what the small-system numerics support.","headline":"Honest, well-cited hybrid ansatz paper with a coherent importance-sampling formalism and released code, but the 'quantum-enhanced expressivity' headline is underdetermined by a missing equal-capacity classical control.","tokens_in":19289,"tokens_out":8033,"would_cite":false,"duration_ms":79106,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A parameterized quantum circuit grafted onto an autoregressive neural network forms a variational wavefunction that reaches chemical accuracy on LiH with far fewer parameters than a standalone neural quantum state.","keywords":["neural quantum states","variational quantum Monte Carlo","parameterized quantum circuits","hybrid quantum-classical ansatz","autoregressive Transformer wavefunctions","quantum chemistry ground states","Heisenberg model","importance sampling"],"falsifier":"Repeat the LiH or 7-spin Heisenberg experiment with the small NQS expanded by the same 60 (or 70) parameters the PQC contributes, keeping the optimizer, sample size, and iteration count fixed; if the enlarged classical network reaches the same final error, the claim that the PQC adds expressivity beyond extra parameters is falsified.","tokens_in":18059,"feed_emoji":"⚛️","tokens_out":10146,"duration_ms":96286,"temperature":0.7,"pith_summary":"This paper proposes a variational ansatz that multiplies a parameterized quantum circuit (PQC) by a neural quantum state, so the circuit supplies part of the wavefunction's amplitude and phase while an autoregressive Transformer supplies normalized probabilities for direct sampling. The goal is to show that a small quantum circuit can supply the expressivity that a deliberately small classical network lacks, without scaling up the classical model. In the paper's headline result, a six-qubit LiH model with a 435-parameter neural part plus 60 circuit parameters reaches an average energy error of $6.0\\times10^{-5}$ Ha, below the chemical-accuracy threshold of $1.6\\times10^{-3}$ Ha, while a purely classical network with more than ten times the parameters reaches only $9.4\\times10^{-4}$ Ha. The same pattern appears for a seven-spin Heisenberg chain, where the hybrid ansatz ends with a relative error near $10^{-3}$ and the standalone NQS stays above $3\\times10^{-2}$. If correct, this gives a practical route to more accurate ground-state energies on near-term quantum hardware without commensurately growing the classical network.","feed_headline":"Sixty quantum parameters beat thousands of classical ones","feed_subtitle":"A hybrid PQC and Transformer wavefunction reaches sub-milli-Hartree error on LiH with a fraction of the parameters.","key_machinery":"The load-bearing object is the product ansatz of Eq. (3), in which the PQC contributes the log-amplitude $f[s;U(\\theta)] = \\sum_i c_i \\langle s| U^\\dagger(\\theta) Z_i U(\\theta) |s\\rangle$ and the classical part contributes the normalized autoregressive amplitude $\\sqrt{p(s;\\lambda_1)} e^{i\\gamma(s;\\lambda_2)}$. The Transformer's autoregressive structure makes $p(s)$ directly samplable, so the VMC step uses importance sampling with weight $\\omega(s)$ instead of Markov-chain burn-in; the PQC amplitude and phase are estimated from $Z$-basis measurements, and circuit gradients come from the parameter-shift rule. Two controls stabilize the estimator: initializing circuit parameters near zero keeps $\\omega(s)$ close to 1, while bounding the amplitude by $\\tanh$ gives a hard upper bound on the importance weight. For chemistry, masking the Transformer's output probabilities with Heaviside functions reduces the sampling space from $2^{N_e}$ to $\\binom{N_O}{N_\\uparrow}\\binom{N_O}{N_\\downarrow}$, preserving particle-number conservation.","core_discovery":"The central claim is that the hybrid wavefunction $\\langle s|\\Psi\\rangle = \\langle s|\\phi(\\theta)\\rangle\\,\\langle s|\\varphi(\\lambda)\\rangle$, with $\\phi(\\theta)$ produced by a hardware-efficient PQC and $\\varphi(\\lambda)$ by a Transformer plus a feedforward phase network, is more expressive per parameter than either component alone. Because the Transformer is autoregressive, configurations are sampled directly and independently, and importance weights $\\omega(s) = |\\langle s|\\phi\\rangle|^2 / \\mathbb{E}_{p}[|\\langle s|\\phi\\rangle|^2]$ computed from quantum measurements reweight the samples within the usual variational Monte Carlo estimator. The authors show, by pretraining the neural part and then optimizing only PQC layers, that each added quantum layer lowers the relative energy error, and they demonstrate on LiH that the hybrid beats the chemical-accuracy threshold while a much larger classical NQS does not. They interpret these results as evidence that a PQC effectively augments the wavefunction when the classical network's expressive power is constrained.","pith_inferences":["A control the paper does not run would make the quantum-expressivity claim sharper: enlarge the small NQS by the same 60 (or 70) parameters the PQC adds, using the same architecture and optimizer; matching errors would attribute the gain to capacity rather than to quantum correlations.","Because the PQC amplitude enters only through the reweighting factor $\\omega(s)$, the same sampler could be reused for other objectives, such as excited states or finite-temperature properties, by swapping the target Hamiltonian while keeping the Transformer fixed.","The paper's $O(BM)$ measurement cost, with $B$ and $M$ around $10^4$, suggests the method is feasible on noisy hardware only if shot noise and gate errors stay well below the reported energy scale; a natural extension is to test the $\\tanh$-bounded importance weight under real device noise.","A direct next benchmark suggested by the framework is the Fermi-Hubbard model, where the paper's Jordan-Wigner mapping and symmetry masking apply without modification."],"forward_implications":["If the central claim is correct, near-term quantum hardware can improve NQS ground-state calculations by adding a shallow PQC instead of enlarging the classical network.","The LiH results imply that a 60-parameter quantum circuit can push a deliberately small neural ansatz below chemical accuracy, a level the paper's much larger classical baseline does not reach.","Sequential optimization on a 7-spin Heisenberg chain shows the hybrid ending at a relative error near $10^{-3}$, versus above $3\\times10^{-2}$ for the standalone NQS.","With sample sizes and shot numbers in the $10^3$–$10^4$ range, the estimator converges and its statistical error shrinks; near $10^5$ shots the results match noise-free circuit simulations.","The symmetry-masking procedure applies to any Hamiltonian with conserved particle numbers, and the paper notes the Fermi-Hubbard Hamiltonian is a special case of the electronic Hamiltonian studied here."],"supporting_citations":[{"why":"Introduces neural quantum states as a variational ansatz, the classical baseline this work extends.","marker":"[5]"},{"why":"Supports the use of Transformer architectures as variational wavefunctions for quantum spin systems.","marker":"[15]"},{"why":"Shows autoregressive models give normalized wavefunctions with direct sampling, enabling the Transformer-guided sampler.","marker":"[53]"},{"why":"Supplies the Jordan-Wigner fermion-to-qubit mapping and local-energy construction used for molecular Hamiltonians.","marker":"[54]"},{"why":"Defines the hardware-efficient ansatz that the PQC part of the hybrid wavefunction uses.","marker":"[68]"},{"why":"Gives the parameter-shift rule used to evaluate PQC gradients in the VMC optimizer.","marker":"[73]"},{"why":"Generates the molecular Hamiltonian and one-particle density matrix used to reduce and solve the LiH problem.","marker":"[75]"},{"why":"Introduces the quantum-classical importance-sampling form that the Transformer-guided reweighting scheme is built on.","marker":"[71]"},{"why":"Provides the Adam optimizer used for training the neural and hybrid parameters.","marker":"[63]"}],"fun_headline_variants":["Quantum circuits boost neural wavefunction expressivity","Hybrid quantum-neural wavefunction outperforms classical","Fewer parameters, lower energy with quantum-enhanced NQS","Sixty quantum parameters reach chemical accuracy on LiH","Autoregressive + quantum circuit: a better wavefunction ansatz"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's quantum-expressivity conclusion rests on comparing the hybrid against a deliberately small classical network; the authors never test a classical network enlarged by the same number of extra parameters, so the improvement could in principle come from added capacity rather than from the quantum circuit.","fun_headline_variants_meta":{"raw":{"variants":["Quantum circuits boost neural wavefunction expressivity","Hybrid quantum-neural wavefunction outperforms classical","Fewer parameters, lower energy with quantum-enhanced NQS","Sixty quantum parameters reach chemical accuracy on LiH","Autoregressive + quantum circuit: a better wavefunction ansatz"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001067,"raw_usage":{"total_tokens":4440,"prompt_tokens":884,"completion_tokens":3556,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":3478}},"tokens_in":500,"tokens_out":3556,"duration_ms":28218,"temperature":1.0,"reasoning_tokens":3478,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:29:29.266041+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the LiH or 7-spin Heisenberg experiment with the small NQS expanded by the same 60 (or 70) parameters the PQC contributes, keeping the optimizer, sample size, and iteration count fixed; if the enlarged classical network reaches the same final error, the claim that the PQC adds expressivity beyond extra parameters is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the use of Transformer architectures as variational wavefunctions for quantum spin systems."},{"cited_title":"Sharir, Y","cited_arxiv_id":null,"evidence_quote":"Supplies the Jordan-Wigner fermion-to-qubit mapping and local-energy construction used for molecular Hamiltonians."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the parameter-shift rule used to evaluate PQC gradients in the VMC optimizer."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the quantum-classical importance-sampling form that the Transformer-guided reweighting scheme is built on."},{"cited_title":"Hibat-Allah, M","cited_arxiv_id":null,"evidence_quote":"Provides the Adam optimizer used for training the neural and hybrid parameters."}],"review_version":1}