{"id":"9a0f285c-497f-45dd-9f2b-c8595f053d87","arxiv_id":"2412.14466","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Bell sampling on two copies of a quantum state beats standard qubit-wise-commuting grouping for molecular energy estimates when the required accuracy is only tens of milli-Hartree or rougher, based on numerics up to 12 qubits.","lead":"This paper tests a quantum measurement trick that uses two copies of a quantum state at once, called Bell sampling, to estimate molecular energies. It finds that for rough accuracy, around tens of milli-Hartree, Bell sampling can need fewer measurements than standard Pauli-string sampling on small test molecules.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The molecular standard-deviation curves rely on a multivariate saddle-point approximation for E[b_i b_j] that is validated only on single-Pauli terms; if that approximation is inaccurate, the claimed 10–30 mHa crossover may shift or disappear.","rationale":"The reader identifies the QWC-only baseline as the weakest assumption, and that is a legitimate external-validity concern: the abstract says 'conventional sampling methods' while the numerical comparison is against QWC grouping only. I agree that a richer baseline (general-commuting or near-optimal grouping) could shrink the advantage. However, I see a more immediate internal correctness risk in the unvalidated multivariate saddle-point approximation used to produce the molecular standard-deviation curves. The paper explicitly states that the saddle-point method is validated in the single-Pauli case and then applies the same method to covariance terms without an independent check. Since the covariance terms enter Var[h] and the crossover location is a quantitative claim, an error here would directly affect the central conclusion even before considering which conventional baseline is used. The proposed Monte Carlo test would settle this by comparing the approximate variance to an essentially exact reference at the relevant N1 values. If the test passes, the numerical core of the paper stands, and the main remaining issue is the narrowness of the conventional baseline. If the test fails, the claimed advantage regime is not established. Because this concern is addressable and does not by itself prove the result wrong, the appropriate disposition remains conditional acceptance rather than rejection: the reader's verdict is unchanged, but the condition should explicitly include independent validation of the covariance calculation.","tokens_in":18404,"tokens_out":12441,"duration_ms":116871,"concrete_test":"For H4 at N1 = 10^3 and N1 = 10^4, evaluate the Bell-sampling variance of the energy estimator using the currently implemented saddle-point method (Eq. (A24) and surrounding formulas), and compare it with an independent Monte Carlo simulation of Bell-basis measurements on two copies of the exact ground state, using enough shots (e.g., 10^7) to make the reference variance tight. If the saddle-point variance differs from the Monte Carlo variance by more than 10% at either N1, recompute Fig. 3b and check whether the 10–30 mHa crossover over QWC survives. This directly tests the unvalidated multivariate saddle-point step at the systems and shot counts where the headline claim is made.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central numerical claim is that Bell sampling with sign estimation uses fewer measurements than QWC-grouped sampling when the target precision is tens of milli-Hartree. In the sign-known analysis that supports the larger molecules (Fig. 3), the Bell-sampling standard deviation is obtained from E[b_i], E[b_i^2], and E[b_i b_j]. The first two are cross-checked against exact binomial summation in Fig. 2, but E[b_i b_j] is computed only with the multivariate saddle-point formula Eq. (A24), based on a numerical saddle point solving U'_ij(m)=0 (Appendix A.5). The paper justifies this by the single-Pauli validation, but that validation does not cover the multinomial integral (A9)–(A10), whose integrand contains max(0, ·) and log terms that can push the saddle point toward a boundary. An inaccurate covariance estimate would shift the Bell standard-deviation curves in Fig. 3 and therefore change or erase the claimed 10–30 mHa advantage. This is load-bearing because it affects the sign-known results for H6 and LiH, which are the basis for the abstract's 'up to 12 qubits' claim, independent of the choice of conventional baseline. The end-to-end H4 simulation in Fig. 5 is less affected because it uses direct sampling, but it covers only one small molecule and does not by itself establish the crossover for larger systems.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper numerically assesses the performance of Bell sampling (measuring two copies of a quantum state in the Bell basis) for estimating expectation values of molecular Hamiltonians. The method simultaneously estimates the absolute values of all Pauli-string expectation values from a single circuit type; signs are estimated separately, either classically (CISD) or by conventional QWC-grouped sampling. The authors derive bias and variance formulas, introduce a saddle-point approximation for the covariance between Pauli terms, and benchmark the method against QWC-grouped conventional sampling with a doubled shot budget on H2, H4, H6, and LiH (up to 12 qubits). The central claim is that, for target precisions of tens of milli-Hartree, Bell sampling with sign estimation requires fewer measurements than conventional sampling, including end-to-end results for H4 with N2 = 5N1.","tokens_in":18675,"tokens_out":4690,"duration_ms":37580,"significance":"If the central claim holds, the paper identifies a practically relevant parameter regime where an entangled-measurement protocol reduces measurement cost for coarse energy estimates in quantum chemistry. The strengths include a careful derivation of bias and variance formulas, a transparent comparison protocol (equal number of state preparations), and explicit numerical simulations on several molecules. The paper also proposes and tests a classical (CISD) route for sign estimation, which is a useful practical addition. However, the breadth of the claim currently exceeds the evidence: the baseline is restricted to QWC grouping, the larger-molecule results rely on an unvalidated multivariate saddle-point approximation, and the end-to-end advantage is demonstrated only for H4 with a favorable N2/N1 ratio.","major_comments":[{"comment":"The standard-deviation curves for H6 and LiH in Fig. 3, which support the abstract's 'up to 12 qubits' claim, rely entirely on the multivariate saddle-point approximation for E[b_i b_j] given in Eq. (A24). This approximation is validated in Fig. 2 only for single-Pauli terms (E[b_i] and E[b_i^2]); the multinomial integral (A9)-(A10) has a different integrand with max(0, ·) and log terms that can push the saddle point toward the boundary. If Eq. (A24) is inaccurate, the 10-30 mHa crossover could shift or disappear. The authors should validate Eq. (A24) against direct Monte Carlo or exact multinomial summation for a representative set of (i,j) pairs and N1 values for at least H6 and LiH, or provide a rigorous error bound for the saddle-point approximation in the multivariate case.","section":"Appendix A.5 / Sec. III B"},{"comment":"The comparison baseline is exclusively QWC grouping, as stated in the text ('QWC grouping is chosen for comparison'), and the abstract and conclusion claim superiority over 'conventional sampling methods' without this qualifier. More efficient general-commuting and near-optimal groupings (Refs. [10, 13, 18]) are acknowledged but excluded because they require additional two-qubit gates. While the gate-count argument for QWC is reasonable, the measurement-shot comparison is incomplete. The authors should either restrict the abstract and conclusion to 'QWC-grouped conventional sampling' or include at least one additional grouping baseline (e.g., general commuting grouping with its gate overhead accounted for) to show that the advantage persists under a stronger conventional baseline.","section":"Sec. III B / abstract"},{"comment":"The end-to-end sign-estimation result uses N2 = 5N1, explicitly chosen as 'the most favorable outcome among the ratios tested empirically.' This introduces a free parameter and risks cherry-picking. The paper should report how the crossover point depends on the N2/N1 ratio and discuss how the ratio would be set in practice without prior knowledge of the optimal value. In addition, the end-to-end simulation covers only H4; the abstract's 'up to 12 qubits' claim is based on the sign-given analysis of Fig. 3, not on the full protocol with sign estimation. The authors should either add end-to-end results for a larger molecule (e.g., H6 or LiH with direct sampling) or explicitly state that the end-to-end advantage is demonstrated only for H4.","section":"Sec. III C / Fig. 5"}],"minor_comments":[{"comment":"The assumption that the sign estimators and the absolute-value estimators are uncorrelated is mentioned only in passing ('with some assumptions such as \\hat s_i and \\hat b_i being uncorrelated'). This assumption is central to the variance decomposition and should be stated more prominently, with an explicit justification that sign estimation is performed on independent measurement shots.","section":"Sec. III / Eq. (7)-(8)"},{"comment":"The caption states that the standard deviation is evaluated using only the saddle point method, while the bias is evaluated by both summation and saddle point. It would be helpful to show a direct comparison of the saddle-point standard deviation against the exact summation for at least one molecule and a few N1 values, to give the reader confidence in the approximation for the molecular case.","section":"Fig. 3 caption"},{"comment":"The paper does not discuss the possibility of multiple saddle points for Eq. (A24) or the numerical procedure for selecting the relevant one. A brief comment on uniqueness and the choice of initial guess for the SymPy nsolve would improve reproducibility.","section":"Appendix A.5"},{"comment":"There is a typo in the sentence 'the absolute values of expectation values of all the Pauli stings are measured simultaneously' — 'stings' should be 'strings'.","section":"Sec. III B"},{"comment":"The phrase 'tens of milli-Hartree' is imprecise; the numerical results indicate a crossover in the range of about 10-30 mHa, and stating this range in the abstract would make the claim more specific and falsifiable.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a relevant and timely question, and the numerical study is carefully executed within its stated scope. The main technical risk is the unvalidated multivariate saddle-point approximation, which supports the larger-molecule results; this should be addressed with additional numerical validation before the paper can be accepted. The baseline comparison to QWC only and the choice of N2 = 5N1 are additional weak points that, while not fatal, require the authors to qualify their claims or expand the benchmarks. I believe the issues are fixable within the manuscript's scope, so major revision is appropriate rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a competent, honest numerical study of an existing idea — Bell sampling for simultaneous estimation of Pauli absolute values, from Huang–Kueng–Preskill. The new content is the detailed bias and standard-deviation analysis for molecular Hamiltonians up to 12 qubits, and the identification of a regime — rough energy estimates at tens of milli-Hartree — where Bell sampling beats QWC-grouped conventional sampling even after paying for sign estimation. That regime is plausible for VQE-like use cases, and the paper does not oversell it too badly: the abstract says 'tens of milli-Hartree,' which matches the figures.\n\nWhat it does well: the statistical formalism is transparent. The estimators are defined carefully, the exact binomial/multinomial expressions for bias and variance are given in the appendix, and the single-Pauli case is cross-checked between exact summation and the saddle-point method. The end-to-end H4 simulation (Fig. 5) uses direct sampling, which is the right way to test the full protocol. The authors also acknowledge the main limitations of their comparison: QWC is not the strongest conventional baseline, and more efficient groupings cost extra two-qubit gates.\n\nWhere I'd push back:\n\n1. The multivariate saddle-point approximation for E[b_i b_j] is load-bearing for the sign-known curves on H6 and LiH, and the paper validates the saddle-point method only on single-Pauli terms. The multinomial integrand has max(0,·) and log terms, so a saddle point near the boundary could behave differently. This is not a proven error, but it is an unvalidated step. A referee should ask for a cross-check on a two-Pauli toy model or a small molecule with exact summation at modest N1.\n\n2. The end-to-end H4 claim uses N2 = 5N1, chosen post hoc as the most favorable among tested ratios. That is transparent, but a real protocol would need an a priori rule. Without that, the 'fewer measurements' statement is partly cherry-picked.\n\n3. The baseline is QWC only. The paper explains why, and the gate-count argument is fair, but the abstract's 'conventional sampling methods' is broader than what is actually tested.\n\nNone of these are fatal. The central statistical picture — Bell sampling wins at low-to-moderate precision and loses at high precision because of the 1/epsilon^4 scaling — is consistent across the molecules tested, and the bias curves back it up. The paper deserves a serious referee; it is a solid resource-estimation contribution. I would accept it with minor-to-moderate revisions, mostly to add validation for the covariance approximation and to soften the abstract's wording.","headline":"Careful numerical study of Bell sampling for coarse energy estimation; the central claims hold up, but the covariance saddle point and the N2/N1 choice need extra scrutiny.","tokens_in":19230,"tokens_out":3612,"would_cite":true,"duration_ms":29130,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that Bell sampling on two state copies estimates all Pauli expectation values at once, and that for molecular Hamiltonians up to 12 qubits it needs fewer measurements than conventional grouped sampling when target…","keywords":["Bell sampling","expectation value estimation","Pauli strings","measurement overhead","quantum chemistry","variational quantum eigensolver","qubit-wise commuting grouping","molecular Hamiltonians"],"falsifier":"Run the same H4 ground-state energy estimation at a 10 milli-Hartree target error with the conventional baseline replaced by a general-commuting or near-optimal grouping schedule; if Bell sampling no longer requires fewer total state preparations than that baseline, the central claim is false. A second check is to repeat the comparison at $\\epsilon=1$ milli-Hartree, where the $1/\\epsilon^4$ scaling of Bell sampling should make conventional sampling win.","tokens_in":18198,"feed_emoji":"⚛️","tokens_out":8173,"duration_ms":65116,"temperature":0.7,"pith_summary":"The paper asks whether Bell sampling from two copies of a quantum state can reduce the measurement cost of estimating Hamiltonian expectation values. The trick works because every doubled Pauli string $P\\otimes P$ commutes with every other, so one Bell-basis circuit yields estimates of $\\langle\\psi|P|\\psi\\rangle$'s absolute value for all Pauli strings at once; the cost is that signs must be obtained separately and the absolute-value estimator has a $1/\\epsilon^4$ shot scaling. Through bias and standard-deviation calculations for molecular Hamiltonians up to 12 qubits, the authors find that, when the target precision is no better than tens of milli-Hartree, Bell sampling beats conventional qubit-wise-commuting sampling even after sign estimation is paid for. This identifies a practical regime: coarse energy estimates in early-stage quantum chemistry algorithms could use far fewer measurements than standard approaches.","feed_headline":"Bell sampling beats conventional sampling for crude energies","feed_subtitle":"Two-copy Bell-basis measurements beat grouped sampling at tens of milli-Hartree accuracy in molecules up to 12 qubits.","key_machinery":"The carrying object is the doubled Pauli string $P\\otimes P$ acting on two copies of the state. All doubled Pauli strings commute, and the Bell measurement on each corresponding qubit pair is a joint eigenbasis of $X\\otimes X$, $Y\\otimes Y$, and $Z\\otimes Z$; reading the eigenvalues $\\Lambda_P(B)=\\prod_k \\lambda_{\\sigma_k}(B_k)$ gives an unbiased estimator $\\hat{a}(P)$ of $|\\langle\\psi|P|\\psi\\rangle|^2$. The nonlinear estimator $\\hat{b}(P)=\\sqrt{\\max(0,\\hat{a}(P))}$ introduces a state-dependent bias that the paper quantifies analytically, and the saddle-point method lets the authors evaluate that bias and the variance for molecular Hamiltonians.","core_discovery":"The central claim is that Bell sampling, meaning a joint Bell-basis measurement of two copies of $|\\psi\\rangle$, is a practically competitive way to estimate molecular Hamiltonians when the required accuracy is moderate. For any Pauli string $P$, the identity $\\langle\\psi|\\langle\\psi|P\\otimes P|\\psi\\rangle|\\psi\\rangle = |\\langle\\psi|P|\\psi\\rangle|^2$ means that squared absolute values of all Pauli expectations are encoded in a single set of Bell outcomes, and the estimator $\\hat{b}(P)=\\sqrt{\\max(0,\\hat{a}(P))}$ recovers the absolute value. The paper's numerical study estimates ground-state energies of H2, H4, H6, and LiH using exact ground states, and evaluates bias and variance through exact summation and a saddle-point approximation. With exact signs, Bell sampling matches or beats QWC grouping at 10 to 30 milli-Hartree; with signs estimated either classically by CISD or by conventional sampling with QWC grouping, it still wins for rough energy estimates, and for LiH(2o2e) the advantage extends below chemical accuracy.","pith_inferences":["Because the advantage comes from replacing an $O(M)$ circuit count with one circuit, the crossover precision should improve, moving to smaller $\\epsilon$, as the molecule grows; verifying this on a 20-plus-qubit Hamiltonian would test the paper's extrapolation.","The sign-estimation overhead is the main bottleneck; replacing the second copy by an adaptively chosen, classically simulable ancilla state could give signs and absolute values in one pass, extending the method to higher precision.","The positive bias of the max-truncated absolute-value estimator could accumulate in energy estimates when many Pauli terms have small true expectations; a testable extension is to compare signed energy estimates with and without truncation on stretched geometries where classical signs are known to be unreliable."],"forward_implications":["For molecular Hamiltonians up to 12 qubits, Bell sampling with sign estimation requires fewer total state preparations than QWC-grouped sampling when the target error is roughly 10 milli-Hartree or worse.","The number of circuits no longer grows with the number of Pauli strings $M$; the same Bell circuit yields all absolute values, so the method's overhead should scale more mildly with system size than per-Pauli projective sampling.","When signs are taken from a classical CISD calculation, the energy estimate remains competitive for rough targets, so quantum measurement shots can be traded for cheap classical approximate sign information.","The bias of $\\hat{b}(P)$ decays as $1/N_1$ for Pauli terms with large expectation values and as $1/N_1^{1/4}$ for near-zero values, so molecules with many small Pauli terms are the hard case for this estimator."],"supporting_citations":[{"why":"introduces the Bell sampling estimator $\\hat{a}(P)$ and $\\hat{b}(P)$ whose simultaneous absolute-value estimation and $1/\\epsilon^4$ guarantee the paper builds on.","marker":"[26]"},{"why":"supplies the qubit-wise-commuting grouping strategy used as the conventional sampling baseline.","marker":"[4]"},{"why":"provides the QWC grouping and minimum clique cover method used to construct the conventional baseline groups.","marker":"[9]"},{"why":"defines weighted deterministic sampling, one of the two shot-allocation rules compared for the QWC baseline.","marker":"[22]"},{"why":"defines weighted random sampling, the other shot-allocation rule used for the baseline and for the sign-estimation shots.","marker":"[46]"},{"why":"motivates the application by showing accelerated VQE with joint Bell measurement can skip sign estimation in mid-optimization.","marker":"[39]"},{"why":"suggests the adaptive ancilla-state variant that estimates signs all at once, cited as a route to improve sign handling.","marker":"[29]"},{"why":"the almost-optimal measurement schedule the paper explicitly excludes from the baseline comparison, defining the boundary of the claimed advantage.","marker":"[18]"}],"fun_headline_variants":["Bell sampling beats QWC grouping for rough molecular energies","Two-copy Bell measurement cuts measurement count for 10-30 mHa accuracy","Bell sampling wins when you only need tens of milli-Hartree","Bell measurement of two copies estimates all Pauli expectations at once","Bell-basis two-copy measurement saves measurements for moderate energies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claimed shot-count advantage is established only against qubit-wise-commuting grouping with its shot budget doubled; if a more efficient grouping that uses extra two-qubit gates is allowed as the baseline, the advantage at tens of milli-Hartree may shrink or disappear.","fun_headline_variants_meta":{"raw":{"variants":["Bell sampling beats QWC grouping for rough molecular energies","Two-copy Bell measurement cuts measurement count for 10-30 mHa accuracy","Bell sampling wins when you only need tens of milli-Hartree","Bell measurement of two copies estimates all Pauli expectations at once","Bell-basis two-copy measurement saves measurements for moderate energies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001542,"raw_usage":{"total_tokens":6183,"prompt_tokens":976,"completion_tokens":5207,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":5118}},"tokens_in":592,"tokens_out":5207,"duration_ms":28983,"temperature":1.0,"reasoning_tokens":5118,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:12:31.158579+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same H4 ground-state energy estimation at a 10 milli-Hartree target error with the conventional baseline replaced by a general-commuting or near-optimal grouping schedule; if Bell sampling no longer requires fewer total state preparations than that baseline, the central claim is false. A second check is to repeat the comparison at $\\epsilon=1$ milli-Hartree, where the $1/\\epsilon^4$ scaling of Bell sampling should make conventional sampling win.","supporting_citations":[{"cited_title":"Thus, we only consider the case of i ̸= j","cited_arxiv_id":null,"evidence_quote":"supplies the qubit-wise-commuting grouping strategy used as the conventional sampling baseline."},{"cited_title":"Yen and A","cited_arxiv_id":null,"evidence_quote":"defines weighted deterministic sampling, one of the two shot-allocation rules compared for the QWC baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines weighted random sampling, the other shot-allocation rule used for the baseline and for the sign-estimation shots."},{"cited_title":"Hamamura and T","cited_arxiv_id":null,"evidence_quote":"the almost-optimal measurement schedule the paper explicitly excludes from the baseline comparison, defining the boundary of the claimed advantage."}],"review_version":1}