{"id":"44a31652-f503-48d7-bccc-82116e2bb010","arxiv_id":"2607.08047","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":11,"one_line_summary":"NuQuLib maps realistic nuclear Hamiltonians to qubit Hamiltonians and compares T-gate costs of QPE, QKrylov, and ODMD across valence and no-core model spaces.","lead":"This paper introduces NuQuLib, a software framework that converts realistic nuclear Hamiltonians into qubit Hamiltonians and estimates the quantum resources needed by three eigenvalue algorithms. It gives nuclear quantum computing a common benchmark baseline, with explicit caveats about single-shot measurements and heuristic cost factors.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Shot-count neglect undermines the headline algorithm ordering: ODMD/QKrylov measurement overhead is disclosed but never priced, so the claimed ordering and exponents in Fig. 6 rest on an unquantified assumption.","rationale":"The reader's weakest_assumption is exactly the N_shot=1 issue, and my independent read of Sec. V B, Table II, and Fig. 6 confirms it is the most consequential unquantified input. The paper is honest about the assumption and even provides the multiplication factor, which is why this is a conditional-accept concern rather than a rejection: the framework and the per-shot T-count formulas are reproducible and the Appendix derivations look consistent. But the stated conclusions — the algorithm ordering and the scaling exponents — are headline claims, and they are not stable under a realistic shot-count insertion unless one can show the shot-count prefactor is small and roughly constant. The paper offers no such demonstration. I considered other possible concerns (Trotter step count, measurement grouping factor 3, qubitization prefactor F), but those are either explicitly parameterized or less central to the paper's main comparative claim. So the single load-bearing concern remains the shot-count neglect, matching the reader. My verdict stays CONDITIONAL: not because of a demonstrated error, but because the central comparison is presented as quantitative while depending on a disclosed but unpriced assumption that could plausibly reverse it.","tokens_in":37460,"tokens_out":1742,"duration_ms":17307,"concrete_test":"Recompute the Fig. 6(a) and 6(b) T-counts with a concrete shot-count model for QKrylov and ODMD: for each required matrix element (O(N_iter^2) for QKrylov; N_snap for ODMD), set N_shot = (est. variance)/epsilon^2 using the Hamiltonian 1-norm or the grouped-operator variance as the variance proxy and epsilon = 1 keV (or the QPE target precision). If ODMD and/or QKrylov no longer lie below Trotter-QPE at N_q = 100–1000, the headline ordering and exponents in Fig. 6 are not robust.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central quantitative claims are the Fig. 6 T-count orderings (ODMD least, QKrylov most resource-intensive; qubitization vs Trotter QPE ~10x) and the fitted exponents N_q^{3.4}–N_q^{12.1}. These numbers come from Table II cost formulas with N_shot=1 for every algorithm (Sec. V B). That is internally consistent only if all algorithms were single-shot, but they are not: QKrylov and ODMD are measurement-driven. QKrylov must estimate O(N_iter^2) overlap and Hamiltonian matrix elements to a target statistical precision; ODMD must estimate N_snap time-series overlaps. Each estimate demands N_shot ~ (variance)/epsilon^2 repetitions, and the variance scales with the number of grouped measurement circuits and with the norm of the grouped operators. Inserting realistic shot counts multiplies TODMD by N_shot(ODMD) and TQKrylov by N_shot(QKrylov), and the paper gives no formula, bound, or numerical estimate for these shot counts. The paper itself flags this in Sec. V B: 'we assumed a single-shot measurement for all algorithms ... one can simply multiply the estimated T-gate count by the number of shots required' — yet all headline comparisons, including the statement that ODMD is least resource-intensive and the quoted N_q exponents, are presented with N_shot=1. Because the prefactor for measurement-driven methods is unquantified and could be as large as 10^3–10^9 for meaningful precision (overlaps and off-diagonal matrix elements decay with overlap amplitudes), the claimed ordering and even the effective scaling exponents are not established. This is not a hidden mathematical error; the derivations in Appendix A are transparent. It is an unquantified but decisive input to the paper's primary comparative conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces NuQuLib, a software workflow that maps realistic chiral-EFT and phenomenological nuclear Hamiltonians (valence-space and no-core, with NN and selected 3N interactions) to Jordan-Wigner-encoded qubit Hamiltonians, and uses this framework to derive T-count resource estimates for three eigenvalue strategies: Trotter-based QPE, qubitization-based QPE, QKrylov, and ODMD. Central quantitative claims are the scaling exponents N_q^{3.4}–N_q^{12.1} for the model spaces in Fig. 6, the ordering that ODMD is least resource-intensive and QKrylov most resource-intensive, and the statement that qubitization-based QPE is about an order of magnitude more efficient than Trotter-based QPE. The paper also provides small-scale statevector demonstrations of QPE, QKrylov, ODMD, angular-momentum-projected state preparation, and VQE, together with an appendix of T-count derivations and block-encoding background.","tokens_in":38030,"tokens_out":10593,"duration_ms":108881,"significance":"If the resource comparisons were established, this would be a useful contribution: it provides a concrete, reproducible bridge from realistic nuclear-structure input to qubit Hamiltonians and algorithmic resource counts, and it begins to create a benchmark culture for nuclear quantum simulation analogous to quantum chemistry. The paper is honest about many of its simplifications, ships a public implementation (NuQuLib), and includes small-space exact checks for Trotter-error estimates that lend credibility to the workflow. The scaling trends, if corrected as described below, could guide early fault-tolerant algorithm selection for nuclear many-body problems. However, the headline quantitative ordering is currently not supported because the cost model sets N_shot=1 for all algorithms despite the measurement-driven nature of QKrylov and ODMD, and because the QPE controlled-evolution call count is presented inconsistently between Table II and Appendix A.","major_comments":[{"comment":"The numerical comparisons set N_shot=1 for every algorithm, as stated in Sec. V B: 'we assumed a single-shot measurement for all algorithms'. For QKrylov and ODMD this is not a harmless prefactor: the number of shots is set by the statistical precision required for O(N_iter^2) overlap/matrix elements (QKrylov) and N_snap time-series overlaps (ODMD). Overlap amplitudes can be small, and the variance depends on the grouped-measurement structure, so realistic shot counts can be orders of magnitude larger than 1. The paper suggests 'one can simply multiply the estimated T-gate count by the number of shots required,' but it provides no formula or estimate for that number. Because the headline ordering (ODMD least, QKrylov most) and the exponents in Fig. 6 are presented with N_shot=1, the central comparison is not established. Please either include realistic shot-count estimates (with their N_","section":"Sec. V B / Table II / Fig. 6"},{"comment":"The number of controlled time-evolution calls in QPE is not presented consistently. Appendix A1 derives the sum Σ_{k=0}^{N_a-1} 2^k = 2^{N_a} − 1, whereas Table II lists T_cU (2N_a − 1). For N_a=20 these differ by a factor of roughly 26,000. This factor is decisive for the QPE-vs-ODMD ordering in Fig. 6 and for the claimed qubitization advantage. Please correct the typo and state explicitly which expression was used to generate Fig. 6 and the numerical exponents.","section":"Appendix A1 vs. Table II"},{"comment":"The Trotter-based QPE estimates count one controlled Trotter step per unit time, but do not multiply by the number of Trotter steps r needed to keep the total Trotter error below the target accuracy for the full QPE evolution time. Equations (13)–(16) give error bounds, and Fig. 7 shows how the error grows with system size, but this r factor is not inserted into the plotted QPE T-counts. As a result, the comparison between Trotter-based and qubitization-based QPE in Fig. 6(c)–(d) and the 'about one order of magnitude' statement are not fixed-accuracy comparisons. The paper acknowledges this limitation in the text, but it should either be reflected in the plots or the plots should be relabeled as per-step costs.","section":"Sec. V B / Sec. V D / Fig. 6(c)-(d)"}],"minor_comments":[{"comment":"The reduction factor N_red is fixed to 2 for all model spaces, whereas the exact values in Table III range from 2.0 to 5.6. Since a smaller N_red gives a larger (more conservative) error bound, the text should state explicitly that this is an upper-bound convention, not a value extracted from the exact commutators.","section":"Table III / Sec. V D 1"},{"comment":"The QKrylov/ODMD demonstration for 20O uses a hard-core boson mapping, whereas the resource estimates in Sec. V are for Jordan-Wigner fermionic encodings. This mismatch should be stated in the demo section so readers do not treat the demonstration as validation of the resource model.","section":"Sec. VI B"},{"comment":"In monochrome print, the dashed and dashed-dotted curves in panels (b) and (d) are difficult to distinguish. Please use different marker shapes or a table of numerical values for the plotted exponents, especially for the N_q^{12.1} curve.","section":"Sec. V C / Fig. 6 caption"},{"comment":"There are several notational inconsistencies: N_q and N_Q are both used; Eq. (14) uses t both for the total evolution time and implicitly for the Trotter step; Table II writes 'N shot' while the text writes N_shot; and the superscripts in Eqs. (A1), (C17), and Table II are easy to misread in the current typesetting. A careful pass to make all exponents and subscripts explicit would improve reproducibility.","section":"General notation"}],"recommendation":"major_revision","confidential_remarks":"The authors are transparent about their assumptions, and the NuQuLib release with small-space exact checks is a real strength. The central issue is that the paper's headline ordering and exponents are presented for N_shot=1 even though the measurement-driven algorithms require shot counts that are not estimated. This is fixable within the manuscript's scope, but it is load-bearing. I also encourage the editor to have the authors confirm the QPE call-count formula, since the discrepancy between Table II and Appendix A affects the main comparison by orders of magnitude."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honestly, this is a solid infrastructure paper. The NuQuLib workflow that takes chiral-EFT or phenomenological nuclear Hamiltonians, maps them to JW-encoded qubit Hamiltonians, and produces term counts and T-count estimates across QPE, QKrylov, and ODMD — that's the real content, and it fills a gap. The counting formulas in Appendix A are transparent, the small-space demonstrations check out, and the software is public. Credit where due: the systematic term-count scalings for NN and 3N in both valence and NCSM spaces are new and will be useful to anyone planning resource estimates for nuclear simulation.\n\nThe soft spots are the measurement overhead. The paper sets N_shot=1 for all algorithms. That's disclosed in Sec. V B, and the formulas leave N_shot symbolic, but every headline figure uses N_shot=1. For QKrylov and ODMD, which are measurement-driven, the shot count to estimate overlaps and matrix elements to a fixed precision is not a constant — it depends on the variances of the estimators, which in turn can scale with the norms of the grouped operators and the size of the subspace. The paper gives no bound or estimate for this, so the claim that ODMD is the least resource-intensive and QKrylov the most is not established. It's a real limitation, but not a hidden error: the formulas are there, and the authors are upfront that shots multiply the totals. The problem is that the central comparative conclusion is presented without that multiplier.\n\nThe other heuristics — N_circ≈N_H/3, N_red=2, the SBE coefficients — are anchored to small-space exact checks and are clearly labeled. They're reasonable for a first-pass benchmark, but they shouldn't be mistaken for rigorous bounds.\n\nWho is this for? Anyone working on quantum algorithms for nuclear structure or building benchmark suites for early fault-tolerant devices. It's not a breakthrough in algorithm design; it's the kind of groundwork the field needs. I'd send it to peer review, with the request that the authors either add shot-count estimates for QKrylov/ODMD or soften the ordering claims. A conditional acceptance with revision would be appropriate.","headline":"Useful benchmark infrastructure for nuclear quantum computing, but the headline T-count ordering across algorithms is conditional on a single-shot assumption that can plausibly reverse the ranking.","tokens_in":38462,"tokens_out":2122,"would_cite":true,"duration_ms":22020,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Realistic nuclear Hamiltonians can be converted into qubit Hamiltonians and, on a shared T-count cost model, ODMD needs the fewest T-gates, QKrylov the most, and qubitization-based QPE runs an order of magnitude cheaper than Trotter-based Q","keywords":["quantum computing","nuclear many-body theory","chiral effective field theory","Jordan-Wigner encoding","T-gate resource estimation","quantum phase estimation","quantum Krylov methods","dynamic mode decomposition"],"falsifier":"Recompute the T-count comparison of Fig. 6 with N_shot set by the precision required for the measured quantities (for example N_shot ~ 1/epsilon^2 per term for QKrylov and ODMD, or calibrated shot counts), and observe whether ODMD and QKrylov remain below Trotter-QPE.","tokens_in":37335,"feed_emoji":"⚛️","tokens_out":8748,"duration_ms":77775,"temperature":0.7,"pith_summary":"Quantum computers need realistic test problems. This paper claims that atomic nuclei, described by chiral effective field theory, provide such a test bed, and it builds a workflow that maps these nuclear Hamiltonians onto qubit operators with Jordan-Wigner encoding. It then counts T-gates—the expensive resource in fault-tolerant quantum computing—for three eigenvalue algorithms. The headline findings: on a shared cost model, the measurement-based Observable Dynamic Mode Decomposition (ODMD) requires the fewest T-gates, Quantum Krylov (QKrylov) the most, and QPE built on qubitization costs about an order of magnitude less than QPE built on Trotter steps. These comparisons are meant as a consistent baseline, not final costs.","feed_headline":"ODMD needs fewest T-gates for nuclear spectra","feed_subtitle":"Qubitization beats Trotter QPE by ~10x on nuclear Hamiltonians; ODMD leads overall.","key_machinery":"The central object is NuQuLib, a workflow that maps a second-quantized nuclear Hamiltonian into a qubit Hamiltonian via Jordan-Wigner encoding with a fixed single-particle ordering. The cost model assigns each Pauli-term exponential one synthesized rotation with T-count T_epsilon about 100, and the formulas in Table II then yield T-counts from the number of Hamiltonian terms and the measurement-circuit grouping factor (about N_hatH/3). For QKrylov, measurement of overlaps and Hamiltonian matrix elements drives an N_iter^3 cost; for ODMD, Hankel-matrix snapshots drive an N_snap^2 cost; and for qubitized QPE, the cost is set by the walk operator built from block-encoding oracles, with eigenval","core_discovery":"The paper's central claim is that realistic nuclear many-body Hamiltonians—not toy models—can be used as controlled quantum-computing benchmarks, and that a shared cost model reveals clear scaling differences among candidate eigenvalue algorithms. Starting from chiral-EFT and phenomenological interactions in valence-shell and no-core model spaces, the workflow Jordan-Wigner-encodes the fermionic operators into Pauli strings, then prices each algorithm by the number of T-gates. Under the stated assumptions (one shot per circuit, T_epsilon about 100 per rotation, first-order Trotter steps), ODMD has the lowest T-count, QKrylov the highest due to its measurement circuits, and qubitized QPE beat","pith_inferences":["The resource ordering depends on the paper's explicit single-shot assumption; with realistic shot counts for the measurement-driven methods, QKrylov and ODMD costs would multiply, potentially changing the ranking.","The measurement-grouping reduction factor (about 3) was fitted on small model spaces; at the larger N_q values shown in Fig. 6, groupability may improve or worsen, which would shift QKrylov's cost relative to QPE.","The walk-operator identity E_k = lambda_H cos(theta_k) hints at Chebyshev-polynomial Krylov schemes built from powers of the walk operator rather than real-time evolution—an algorithm family the paper mentions but does not price.","The same encoding workflow could be applied to lattice nuclear interactions or scattering Hamiltonians, extending the benchmark suite beyond eigenvalue problems."],"forward_implications":["Any chiral-EFT or phenomenological shell-model Hamiltonian can be mapped to a qubit Hamiltonian and to a library of quantum circuits, so nuclear structure offers a reproducible benchmark family for quantum eigensolvers.","Under the paper's cost model, ODMD is the least T-gate-intensive of the three eigenvalue algorithms and QKrylov the most, because QKrylov must measure many overlap and Hamiltonian matrix elements.","For QPE, qubitization saves roughly an order of magnitude in T-gates over first-order Trotterization, and the gap should widen with system size since the commutator-bound-to-lambda_H ratio scales as N_q^{2.6}.","T-gate estimates for hundreds of qubits fall in the 10^{10}-10^{14} range, similar to early quantum-chemistry resource estimates, indicating that algorithmic improvements are needed before practical nuclear simulation.","Three-body interactions steepen the scaling dramatically (exponents up to N_q^{12.1} in no-core spaces), making 3N forces a particularly severe stress test for any quantum eigensolver."],"fun_headline_variants":["ODMD needs fewest T-gates for nuclear spectra","Qubitization beats Trotter ~10x for nuclear QPE","Nuclear many-body Hamiltonians rank quantum algorithms","ODMD tops T-gate cost for nuclear benchmarks"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The comparison assumes a single measurement shot per circuit for every algorithm, even though QKrylov and ODMD are measurement-driven methods that would need many repeated shots to estimate overlaps and matrix elements; if realistic shot counts are included, the claimed resource ordering may change.","fun_headline_variants_meta":{"raw":{"variants":["ODMD needs fewest T-gates for nuclear spectra","Qubitization beats Trotter ~10x for nuclear QPE","Nuclear many-body Hamiltonians rank quantum algorithms","ODMD tops T-gate cost for nuclear benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000272,"raw_usage":{"total_tokens":1445,"prompt_tokens":698,"completion_tokens":747,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":442,"completion_tokens_details":{"reasoning_tokens":693}},"tokens_in":442,"tokens_out":747,"duration_ms":6503,"temperature":1.0,"reasoning_tokens":693,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T07:56:40.699232+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the T-count comparison of Fig. 6 with N_shot set by the precision required for the measured quantities (for example N_shot ~ 1/epsilon^2 per term for QKrylov and ODMD, or calibrated shot counts), and observe whether ODMD and QKrylov remain below Trotter-QPE.","supporting_citations":[],"review_version":2}