{"id":"880d43ec-7495-431f-beba-3df120347e79","arxiv_id":"2506.09308","paper_version":3,"verdict":"REJECT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of quantum algorithm software that advertises a benchmark suite, yet the body contains no benchmarks, data, or code.","lead":"This preprint is a qualitative review of quantum computing software for condensed matter physics, covering VQE, QPE, QAOA, and QML with SDK comparisons. Its abstract promises a reproducible benchmark suite with numerical results, but the full text contains no tables, figures, code, or data, so the quantitative claims cannot be checked.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract advertises a fully reproducible benchmark suite with quantitative results, but the full text contains no benchmark data, tables, figures, or code; the central contribution is unsupported and internally contradicted by the conclusion's call for future benchmarks.","rationale":"The reader correctly identifies the weakest assumption as the existence of a concrete numerical experiment behind the advertised benchmark suite. My independent reading confirms this: the abstract promises quantitative benchmark results, but the full text contains no such results. The paper functions as a competent but non-quantitative review of known materials, and its own conclusion emphasizes the need for future benchmark development, which is incompatible with the abstract's claim that the benchmarks are already provided and released. Since the central claim of the paper is the benchmark suite itself, and that claim is entirely unsupported by the manuscript as submitted, rejection is appropriate. I agree with the reader's verdict and rationale without modification.","tokens_in":90,"tokens_out":1051,"duration_ms":23127,"concrete_test":"Check the arXiv source and any linked external files for the benchmark suite: search the full text for every occurrence of 'table', 'figure', 'simulation', 'noise', 'ZNE', 'qubit count', 'gate count', and for a code repository or supplementary data URL. If no numerical result, table, or figure exists and no repository or dataset is provided, the abstract's claim of a released, reproducible benchmark suite is falsified. Additionally, compare Section VII's statement that standardized benchmarks need to be developed with the abstract's assertion that the paper already delivers them; if both are present, the internal contradiction confirms the absence of the benchmark content.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's stated central claim, per the abstract, is a 'compact, fully reproducible benchmark suite' that 'turns qualitative claims into concrete numbers,' including qubit counts, operator weights, gate costs for Jordan-Wigner and Bravyi-Kitaev encodings, and zero-noise extrapolation results for ground-state energies. The full manuscript provides none of this. There are no tables, no figures, no numerical results, no equations defining the benchmark protocol, no hardware or simulator parameters, no noise model specifications, no lattice sizes, no ansatz descriptions, and no code repository or data link. Section IIA–D review VQE, QPE, QAOA, and QML qualitatively; Section IV surveys SDKs; Section V discusses error mitigation generally. The only mention of zero-noise extrapolation is in Section VC as a generic technique, not as a demonstrated result. Section VII explicitly states that standardized benchmarks are a 'currently underemphasized aspect' and calls for their development, directly contradicting the abstract's claim that this paper provides such benchmarks. The strongest quantitative claims therefore rest on absent experiments. This is not a matter of disagreement with consensus or subtle internal inconsistency: the advertised new content is missing entirely, so the central claim fails as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a broad review of quantum algorithm software for condensed matter physics. It surveys VQE, QPE, QA/QAOA, QML, tensor-network methods, and classical algorithms, then profiles the SDKs Qiskit, Cirq, PennyLane, and Q#, discusses hardware limitations, error mitigation, and future outlooks, and concludes with a call for standardized benchmarks. The abstract, however, makes a much stronger claim: that the paper also provides a 'compact, fully reproducible benchmark suite' with concrete results, including qubit counts, operator weights, gate costs for Jordan-Wigner and Bravyi-Kitaev encodings, and zero-noise extrapolation demonstrations. The full text contains none of these benchmark data, tables, figures, numerical equations, noise-model specifications, or code/data links; the advertised central contribution is absent from the manuscript.","tokens_in":22126,"tokens_out":3072,"duration_ms":35995,"significance":"If the benchmark suite described in the abstract existed in the paper, it would be a useful community resource for comparing quantum algorithm implementations on canonical lattice models. The survey sections are competent and will be informative to newcomers, and the paper cites a broad range of relevant literature. However, the quantitative claims in the abstract are the main advertised contribution, and they are entirely unsupported by the body of the manuscript. The paper contains no machine-checked proofs, no reproducible code artifacts, no parameter-free derivations, and no numerical results; as a result, the central claim fails as stated.","major_comments":[{"comment":"The abstract states that the paper provides a 'compact, fully reproducible benchmark suite that turns qualitative claims into concrete numbers,' including qubit counts, operator weights, and gate costs for Jordan-Wigner and Bravyi-Kitaev encodings, and zero-noise extrapolation results. I could not find any of these results in the body: there are no tables, figures, numerical data, equations defining a benchmark protocol, noise-model specifications, lattice sizes, ansatz descriptions, or code/data repository links anywhere in Sections I through VII. The advertised central deliverable is therefore absent, making the paper's central claim unsupported.","section":"Abstract and full text"},{"comment":"The abstract attributes to the paper a demonstration that zero-noise extrapolation 'restores ground-state energies and optimization quality across the noise range.' Section V.C, however, mentions zero-noise extrapolation only as one of several generic error-mitigation techniques and provides no noise model, no circuits, no extrapolation procedure, and no numerical comparison against exact-diagonalization or DMRG references. This quantitative result is claimed without any supporting evidence.","section":"Section V.C"},{"comment":"The paper says it 'tabulate[s] qubit counts, operator weights, and gate costs' for Jordan-Wigner versus Bravyi-Kitaev encodings and exposes a geometry-dependent trade-off. Section IV.A discusses fermion-to-qubit mappings only qualitatively, and Section II.E describes lattice models without any of the promised resource counts. There is no table or equation to substantiate the trade-off, so this advertised comparison is not part of the manuscript as written.","section":"Sections IV.A and II.E"},{"comment":"The conclusion states that standardized benchmarks are a 'currently underemphasized aspect' and calls for their development and adoption, saying they 'would enable more objective and rigorous comparisons of different algorithmic approaches.' This directly contradicts the abstract's claim that the present paper supplies such benchmarks. At minimum, the authors need to reconcile these statements; as it stands, the paper internally denies the existence of its own central contribution.","section":"Section VII"}],"minor_comments":[{"comment":"There is a missing space and period in 'thereby reducing latencyQiskit also includes tools' on page 8; it should read 'reducing latency. Qiskit also includes tools...'.","section":"Section IV.B"},{"comment":"The text contains 'SW AP gates' with an unintended space; it should be 'SWAP gates'.","section":"Section V.A"},{"comment":"Several headings contain stray spacing or encoding artifacts, such as 'V ariational Quantum Eigensolver', 'T wo', and 'Schr¨ odinger'; these should be corrected in the final version.","section":"Headings and typesetting"},{"comment":"Reference formatting is inconsistent: some entries are arXiv preprints without journal or DOI information, some URLs lack access dates, and [136] is a blog post. The reference list would benefit from a uniform style.","section":"References"},{"comment":"The phrase 'Q ecosystem' in the discussion of Q# is ambiguous, since Q# is the language and the development kit is the Azure Quantum Development Kit; the sentence should specify which part of the ecosystem is meant.","section":"Section VI.A"}],"recommendation":"reject","confidential_remarks":"The manuscript reads as a survey whose abstract has been expanded to promise original benchmark results that are not present. I would not recommend rejection solely on the grounds that the survey portion is unoriginal, but the mismatch is severe: the quantitative claims that should be the main contribution have no supporting content in the body. This is not a local fixable issue but a missing core deliverable, so I recommend reject. A resubmission that actually includes the benchmark section, data, and code would be a different and potentially useful paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi —\n\nThe short version: this is a competent, wide-ranging review of quantum algorithm software for condensed matter, but the abstract promises a reproducible benchmark suite with numbers, and that suite simply isn't in the text. The stress-test note is right. Read the paper and you find no tables, no figures, no code, no data, no equations for the JW/BK trade-off, no ZNE results. Section VII even says standardized benchmarks are 'currently underemphasized' and calls for their development. That directly contradicts the abstract's claim that the paper provides such benchmarks.\n\nWhat's good: the review itself is accurate and current. The descriptions of VQE, QPE, QAOA, QML, the SDKs, and classical methods like tensor networks and DMFT are all fine. The references are recent and relevant. If you need a broad survey to hand to a new student, this works.\n\nThe soft spots are proportional to how missing the advertised content is. The central claim is not just weakly supported; it's absent. The mismatch between abstract and body is a completeness problem, not a subtle disagreement. There's also no new analysis in the review sections: it's synthesis of things already published. The paper's own conclusion admits the benchmarks don't exist yet. That's a real red flag for how the abstract was written.\n\nThere's no evidence of circularity or parameter fitting — the issue is simply that the claimed new results are not in the manuscript.\n\nWho is this for? Someone who wants a broad orientation of the field, not someone looking for new results. As submitted, I'd desk-reject rather than send to referees, because the advertised contribution is missing and the remaining text is a review that could be posted as a survey without the misleading abstract. If the author later adds the actual benchmark suite with data and code, the paper becomes worth another look.\n\nSo: not a serious paper in this form, but the underlying review is honest work. My vote: reject, don't send to referees.","headline":"Abstract promises benchmarks the paper doesn't contain; competent review but the central claim is unsupported.","tokens_in":22602,"tokens_out":2953,"would_cite":false,"duration_ms":31343,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that software matters as much as hardware for quantum condensed-matter computing and claims a reproducible suite that converts algorithm comparisons into qubit counts, gate costs, and zero-noise-extrapolated energies.","keywords":["quantum algorithm software","condensed matter physics","reproducible benchmarks","Fermi-Hubbard model","Jordan-Wigner transformation","Bravyi-Kitaev transformation","zero-noise extrapolation","quantum error mitigation"],"falsifier":"Inspect the release accompanying the paper: if it contains no runnable scripts that produce Fermi-Hubbard qubit counts, gate costs, and ZNE-corrected energies for specified lattice geometries and depolarizing noise strengths, or if re-running those scripts yields numbers that do not match the abstract's claims, the benchmark-suite assertion fails. A second check is to reproduce the ZNE pipeline on a small Fermi-Hubbard instance and compare the zero-noise-extrapolated energy against exact diagonalization; disagreement beyond the claimed tolerance would falsify the noise-recovery claim.","tokens_in":21723,"feed_emoji":"⚛️","tokens_out":16260,"duration_ms":131689,"temperature":0.7,"pith_summary":"This paper tries to move quantum-algorithm software for condensed matter physics from qualitative review to quantitative benchmarking. Its central claim is that a compact, fully reproducible benchmark suite accompanies the review: VQE, QPE, quantum annealing/QAOA, and QML are each run on a canonical lattice model and checked against a classical reference such as exact diagonalization, the Bethe ansatz, or DMRG. Within that suite, the paper reports qubit counts, operator weights, and gate costs for the Fermi-Hubbard model under Jordan-Wigner and Bravyi-Kitaev encodings, exposing a geometry-dependent trade-off, and it simulates depolarizing noise to show that zero-noise extrapolation restores ground-state energies and optimization quality. If the suite is real and reusable, practitioners gain a common yardstick for comparing algorithms, software kits, and hardware, and claims of quantum advantage become testable. The manuscript itself is a narrative review; the quantitative results are promised in the abstract and in the released code rather than printed in the body.","feed_headline":"Benchmark suite turns quantum condensed-matter claims into numbers","feed_subtitle":"A review that claims reusable benchmarks for JW vs BK encodings and zero-noise extrapolation on lattice models.","key_machinery":"The machinery is a benchmark suite built around canonical lattice models. For the Fermi-Hubbard case, the central object is the fermion-to-qubit encoding: Jordan-Wigner (JW) and Bravyi-Kitaev (BK) transformations, compared by qubit count, Pauli operator weight, and gate cost, with the claimed geometry-dependent trade-off. For the noise study, the machinery is a depolarizing noise model on the simulated circuits plus zero-noise extrapolation (ZNE), a classical post-processing technique that runs circuits at several amplified noise levels and extrapolates to the zero-noise limit. Around those two quantitative studies, the suite pairs each algorithm family (VQE, QPE, QA/QAOA, QML) with a canonical model and a classical reference method, which is what makes the advertised numbers reproducible.","core_discovery":"The paper asserts that software is as decisive as hardware for realizing quantum computation in condensed matter physics, and that the field's customary qualitative reviews need to be supplemented with reproducible numbers. The discovery claim is that such a benchmark suite has been built: the Fermi-Hubbard model, mapped under Jordan-Wigner and Bravyi-Kitaev encodings, is tabulated for qubit counts, operator weights, and gate costs, revealing a trade-off between the two encodings that depends on lattice geometry; and circuits run under a depolarizing noise model show zero-noise extrapolation recovering ground-state energies and optimization quality across the noise range. Each algorithm family is demonstrated on a canonical lattice model and validated against an independent classical method, from exact diagonalization and the Bethe ansatz to matrix-product-state DMRG, so the advertised results claim to convert qualitative statements into concrete numbers. The numbers themselves do not appear in the prose of the manuscript, which instead presents the surrounding review of algorithms, software development kits, classical methods, and challenges; the numerical content is claimed to live in the released circuits, seeds, and data.","pith_inferences":["A natural next experiment is to route both encodings through a compiler pass on a fixed chip topology, since SWAP overhead on limited connectivity may reverse the raw JW-versus-BK resource ordering.","The same depolarizing-noise benchmark could be extended to compare zero-noise extrapolation against probabilistic error cancellation and readout-error correction on identical lattice models, isolating each method's contribution.","If the benchmark-template idea spreads, the field could converge on a shared set of challenge problems with fixed Hamiltonians, sizes, and classical references, making results from different quantum hardware directly comparable."],"forward_implications":["If the released suite is reproducible, researchers can directly compare Jordan-Wigner and Bravyi-Kitaev encodings on Fermi-Hubbard lattices and select the cheaper encoding for a given geometry.","If the zero-noise-extrapolation demonstration holds across the stated noise range, practitioners gain a concrete error-mitigation recipe for lattice-model simulations under depolarizing noise.","The suite's template places VQE, QPE, QAOA, and QML on the same lattice-model footing, each validated against an independent classical reference.","Widespread adoption of such standardized benchmarks would make claims of quantum advantage in condensed matter falsifiable rather than qualitative."],"supporting_citations":[{"why":"Defines the noisy intermediate-scale quantum (NISQ) context that motivates the paper's error-mitigation and benchmarking focus.","marker":"[12]"},{"why":"Introduces the variational quantum eigensolver, the first algorithm family the suite covers.","marker":"[15]"},{"why":"Introduces quantum phase estimation, the spectral algorithm the suite covers.","marker":"[16]"},{"why":"Provides the Fermi-Hubbard model as the canonical target and a quantum-hardware demonstration to benchmark against.","marker":"[20]"},{"why":"Frames resource counting for fermionic simulations, the basis for the JW-versus-BK cost comparison.","marker":"[116]"},{"why":"Supplies the Bravyi-Kitaev transformation whose resource costs the paper compares with Jordan-Wigner.","marker":"[127]"},{"why":"Establishes geometry-dependent fermion-to-qubit mapping trade-offs that the benchmark aims to expose numerically.","marker":"[128]"},{"why":"Catalogues zero-noise extrapolation and other error-mitigation techniques used in the noise study.","marker":"[140]"}],"fun_headline_variants":["Benchmark suite turns quantum condensed-matter claims into numbers","Reproducible benchmarks quantify quantum algorithms for matter","Quantum software for condensed matter gets measured, not just reviewed","Benchmark suite exposes trade-offs in quantum encodings and noise"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"That the advertised benchmark suite actually exists and was executed: the manuscript does not state lattice sizes, ansatz circuits, noise parameters, ZNE implementation details, or any numerical result, so the abstract's concrete conclusions rest on released code and data whose contents are not shown in the paper.","fun_headline_variants_meta":{"raw":{"variants":["Benchmark suite turns quantum condensed-matter claims into numbers","Reproducible benchmarks quantify quantum algorithms for matter","Quantum software for condensed matter gets measured, not just reviewed","Benchmark suite exposes trade-offs in quantum encodings and noise"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000477,"raw_usage":{"total_tokens":2429,"prompt_tokens":1071,"completion_tokens":1358,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":687,"completion_tokens_details":{"reasoning_tokens":1290}},"tokens_in":687,"tokens_out":1358,"duration_ms":11020,"temperature":1.0,"reasoning_tokens":1290,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:50:17.955832+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the release accompanying the paper: if it contains no runnable scripts that produce Fermi-Hubbard qubit counts, gate costs, and ZNE-corrected energies for specified lattice geometries and depolarizing noise strengths, or if re-running those scripts yields numbers that do not match the abstract's claims, the benchmark-suite assertion fails. A second check is to reproduce the ZNE pipeline on a small Fermi-Hubbard instance and compare the zero-noise-extrapolated energy against exact diagonalization; disagreement beyond the claimed tolerance would falsify the noise-recovery claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Bravyi-Kitaev transformation whose resource costs the paper compares with Jordan-Wigner."},{"cited_title":"Jiang, K","cited_arxiv_id":null,"evidence_quote":"Establishes geometry-dependent fermion-to-qubit mapping trade-offs that the benchmark aims to expose numerically."},{"cited_title":"Cai et al","cited_arxiv_id":null,"evidence_quote":"Catalogues zero-noise extrapolation and other error-mitigation techniques used in the noise study."}],"review_version":1}