{"id":"4a56d497-2a9d-4220-8f15-5f21ab6571ed","arxiv_id":"2508.00837","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A dataset of 55 protein fragments predicted on IBM quantum hardware via VQE, claimed to outperform AlphaFold2/3 in RMSD and docking affinity, with an underspecified and possibly unfair comparison.","lead":"This paper presents QDockBank, a dataset of 55 short protein fragments from ligand-binding pockets, with structures generated by running the Variational Quantum Eigensolver on IBM Eagle quantum hardware. The authors claim these quantum-generated fragments beat AlphaFold2 and AlphaFold3 on RMSD and docking affinity, but the comparison protocol and the biological validity of docking to isolated fragments are not established.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed docking-affinity advantage rests on an invalid metric: AutoDock Vina scores on isolated 5–14 residue fragments without the surrounding protein do not measure ligand-binding affinity, so the central claim is unsupported.","rationale":"The abstract's headline claim has two legs: structural RMSD and docking affinity. The RMSD leg is weakened by the unspecified AlphaFold protocol, but the docking-affinity leg is independently invalid because the metric does not measure what it claims. Binding affinity of a ligand to a protein is a property of the whole binding site; a 5–14 residue fragment excised from the pocket and used alone as a rigid receptor omits the majority of stabilizing contacts and presents an artificial surface. AutoDock Vina will return some score, but that score is not a binding affinity to the biological target. Since the paper explicitly uses docking affinity scores as evidence of functional superiority and as the basis for more than 90% win rates over AF2/AF3, this is the load-bearing assumption. If it fails, the central claim cannot stand regardless of the RMSD numbers. I agree with the reader's weakest_assumption. A single computational check—re-docking against full receptors—would settle whether the affinity advantage survives a realistic evaluation. The unspecified AlphaFold baseline is a separate serious flaw, but the truncated-receptor docking is the more fundamental problem because it invalidates the metric itself. The dataset may still be a useful resource, but the paper's headline conclusion is not supported as presented; the reader's REJECT verdict should stand unchanged.","tokens_in":15561,"tokens_out":6494,"duration_ms":77982,"concrete_test":"Reproduce the §6.2 affinity comparison for all 55 fragments using the full experimental protein (and, separately, the full AlphaFold-predicted protein) as the rigid receptor instead of the isolated fragment, keeping the same ligand, search box, and Vina settings. If QDockBank fragments no longer yield lower (more negative) Vina scores than the AlphaFold-derived receptors in a majority of cases, the claimed docking-affinity advantage is an artifact of truncated-receptor docking.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that QDockBank structures outperform AlphaFold2/3 in both RMSD and docking affinity depends critically on the docking-affinity metric defined in §6.1.2 and applied in §4.3.3. There, each predicted fragment is used as a rigid receptor by itself, and AutoDock Vina docks the native ligand from the full PDBbind complex against this isolated 5–14 residue peptide. This cannot measure ligand-binding affinity: the native ligand was co-crystallized against the entire protein, so most of the contacts that determine binding lie outside the fragment. Vina's score on a truncated, artificially exposed surface reflects how well the ligand packs against a small piece of the pocket, not the free energy of binding to the actual site. The conclusion that QDockBank predictions are 'more favorable' for docking is therefore not established. The AlphaFold baseline protocol is also not described (§6.2 cites ColabFold [39]), so it is unclear whether the RMSD and affinity comparisons are apples-to-apples; but even if they were, the affinity comparison would remain biologically invalid.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"QDockBank is presented as the first large-scale dataset of protein fragment structures generated on real (utility-level) quantum hardware, with 55 fragments of 5–14 residues extracted from ligand-binding pockets of PDBbind proteins. The fragments are predicted via a coarse-grained tetrahedral lattice model encoded into a Hamiltonian and optimized with VQE on IBM Eagle processors. The paper's central claim is that the quantum-generated structures outperform AlphaFold2 and AlphaFold3 in both RMSD to X-ray structures and AutoDock Vina binding-affinity scores on the corresponding native ligands. The dataset also includes quantum metadata, docking results, and a claimed coverage of amino-acid interaction types.","tokens_in":15828,"tokens_out":3165,"duration_ms":37485,"significance":"If the claims were established, the work would be significant as an engineering demonstration: it would be among the largest uses of real quantum hardware for a biomolecular modeling task, with documented execution times, qubit counts, and a publicly released dataset. The authors deserve credit for reporting extensive hardware metadata and for making the dataset available. However, the central benchmarking claim is compromised by a biologically invalid docking metric, an underspecified Hamiltonian, and an undocumented AlphaFold baseline protocol. The paper's headline conclusion—that quantum-generated fragments outperform AlphaFold in docking affinity—is not supported by the evidence as presented.","major_comments":[{"comment":"The docking-affinity comparison is not a valid measure of ligand-binding affinity. Each predicted fragment is used as a rigid receptor by itself, and AutoDock Vina docks the native ligand from the full PDBbind complex against this isolated 5–14 residue peptide. The native ligand was co-crystallized with the entire protein, so most of the contacts that determine binding lie outside the fragment. A Vina score on a truncated, artificially exposed surface reflects how well the ligand packs against a small piece of the pocket, not the free energy of binding to the actual site. The abstract's claim that QDockBank structures outperform AlphaFold2 and AlphaFold3 'in terms of ... docking affinity scores' is therefore unsupported, regardless of the numerical results.","section":"§6.1.2, §4.3.3"},{"comment":"The Hamiltonian that defines the prediction objective is never specified. Equation (1) lists four terms H_c, H_g, H_d, H_i, but the functional form of each term, the basis of the 'pairwise amino acid interaction energies' in H_i, and the parameter values are not given. The statement that λ_c = λ_g = λ_d = λ_i = 1 is insufficient without units or a definition of the energy scales. This is load-bearing because the predicted structures are the ground states of this Hamiltonian; without specifying it, the method cannot be reproduced, and the 'first-principles' characterization in Section 2.2 and the abstract is contradicted by the later reliance on a statistical potential (Miyazawa–Jernigan) in Section 6.2.","section":"§4.3.1"},{"comment":"The AlphaFold2 and AlphaFold3 baseline protocol is not described. The text says only that the comparison was made 'compared with AlphaFold2(AF2) [39] and AF3' and cites ColabFold [39] for AF2. It is not stated whether AF2/AF3 were run on the isolated fragment sequence or on the full protein with the fragment subsequently extracted, which input structures or templates were used, how the predicted structures were aligned or trimmed, or whether the same post-processing (e.g., Open Babel refinement, centering) was applied to the baselines. Without this information, the RMSD and affinity comparisons cannot be evaluated as apples-to-apples, and the central comparison is not established.","section":"§6.2"},{"comment":"The performance comparison is reported only as percentages of samples where QDock scores lower, with no confidence intervals, paired statistical tests, or analysis of the magnitude of differences. Given the small sample sizes within each group (e.g., 12 in Group L, 23 in Group M), the claim that the method 'outperforms' AlphaFold is not supported by a significance assessment. This is secondary to the invalid docking metric but further weakens the headline comparison.","section":"§6.2, Figures 2–4"}],"minor_comments":[{"comment":"The paper repeatedly calls the approach 'first-principles' while the interaction term H_i is said to encode pairwise amino acid interaction energies, and Section 6.2 explicitly invokes the Miyazawa–Jernigan statistical potential. This terminology should be revised to avoid implying the method is parameter-free.","section":"Abstract / §2.2"},{"comment":"The text states that the dataset comprises 'more than 2,000 docking tests,' but 55 fragments × 20 docking runs gives 1,100 docking simulations; if the top-10 poses are counted, the number is larger. The counting convention should be clarified.","section":"§4.2"},{"comment":"The AlphaFold3 baseline is not cited at the point of comparison; reference [39] is ColabFold. A precise citation for the AF3 version and protocol should be added.","section":"§6.2"},{"comment":"The table reports average docking metrics for a single PDB entry (4jpy). The text would benefit from error bars or per-run values, since 20 docking runs with different seeds are described but only averages are shown.","section":"§7.1, Table 4"},{"comment":"The atomic reconstruction step is described as 'applying standard amino acid templates' without specifying which templates or how side-chain conformations were chosen; this is relevant because side-chain placement can affect subsequent docking scores.","section":"§4.3.3"}],"recommendation":"reject","confidential_remarks":"The paper's engineering effort and dataset release are commendable, but the central benchmarking claim rests on a docking protocol that cannot measure ligand-binding affinity for isolated fragments, and the prediction Hamiltonian is never specified. These are load-bearing issues that cannot be fixed by local revisions; the comparison would need to be redesigned, or the claims substantially narrowed to structural RMSD only, which would change the paper's scope. I therefore recommend rejection, though the authors should be encouraged to resubmit a revised dataset description if the affinity comparison is removed or replaced with a valid evaluation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuine dataset contribution from real IBM Eagle runs, and the metadata alone is worth a look. But the paper's central claim—QDockBank structures outperform AlphaFold2/3 in docking affinity—rests on a metric that does not measure ligand binding. They dock the native ligand to an isolated 5–14 residue fragment, not the full protein. Most of the contacts that determine binding lie outside the fragment. Vina scores on a truncated piece reflect packing against a small exposed surface, not affinity. So the abstract overstates the result.\n\nWhat is new and good: the dataset itself. 55 fragments extracted from PDBbind, all generated on physical quantum hardware with VQE, including qubit counts, circuit depths, energies, runtimes. That is a first, and the public GitHub release makes it usable. The cost and runtime numbers are transparent. The interaction-coverage analysis (395/400 pairs) is a nice sanity check.\n\nSoft spots beyond the docking issue. The AlphaFold baseline is described only as 'ColabFold [39]'; there is no statement of whether AF2/AF3 were given the fragment sequence alone or the full protein, which matters enormously for RMSD comparisons. The Hamiltonian in Section 4.3.1 is given only as a weighted sum of four unnamed terms; the actual functional form of H_i (the pairwise amino acid interaction energies) is never specified, so the quantum optimization is not reproducible from the paper. And calling the approach 'first-principles' is a stretch if H_i is an empirical statistical potential, which the later Miyazawa–Jernigan reference suggests.\n\nHow soft are these? The docking flaw is load-bearing for the affinity claim. The RMSD comparison might be salvageable if the AlphaFold protocol is clarified, but as written the comparison is not apples-to-apples. The missing Hamiltonian terms are a reproducibility gap, not necessarily a conceptual error, and could be fixed by adding an appendix.\n\nBottom line: this paper deserves a serious referee because the dataset is real and novel, but the current draft should not be accepted as-is. I'd recommend peer review with major revision, and the authors should either fix the evaluation or soften the claims to structural accuracy only.","headline":"Real quantum hardware dataset, but the headline docking comparison is biologically invalid and the AlphaFold baseline is underspecified, so the central claim does not hold as written.","tokens_in":16269,"tokens_out":2557,"would_cite":false,"duration_ms":29483,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces QDockBank, a 55-fragment dataset of ligand-binding protein structures generated entirely on utility-level quantum hardware, and claims these structures beat the leading deep-learning predictors on both RMSD and…","keywords":["quantum protein structure prediction","ligand docking","variational quantum eigensolver","benchmark dataset","coarse-grained lattice model","binding affinity","RMSD","superconducting quantum processor"],"falsifier":"Run the same docking comparison on full-length protein structures (or on fragments embedded in the surrounding pocket residues) rather than on the isolated fragments: if the deep-learning models' full-length structures match or beat the quantum fragments' affinity scores, the paper's central advantage is an artifact of truncation. A simpler check is to dock the native ligand against the quantum fragment and against the corresponding full experimental pocket and compare the resulting poses.","tokens_in":15415,"feed_emoji":"🧬","tokens_out":8058,"duration_ms":79902,"temperature":0.7,"pith_summary":"QDockBank is the first large-scale dataset of protein fragment structures generated entirely on utility-level quantum computers. The paper's central claim is that these 55 fragments, drawn from ligand-binding pockets, are more accurate than the leading deep-learning predictors: in its evaluations, the quantum structures beat the older deep-learning baseline in binding affinity in 53 of 55 cases and beat the newer baseline in 50 of 55, with similar win rates on RMSD. If true, this means a first-principles, physics-based quantum pipeline can outperform data-driven models exactly where those models are weakest: short, variable peptides in functional pockets. The dataset also records docking scores, RMSD values, qubit counts, circuit depths, and execution times, making it a reproducible benchmark for quantum biomolecular modeling. The paper reports over 60 hours of processor runtime and a total computational cost above one million dollars, framing the scale at which such quantum structure prediction currently operates.","feed_headline":"Quantum-built protein fragments beat AI models on docking tests","feed_subtitle":"QDockBank's 55 fragments report lower error and stronger ligand binding than deep-learning predictions.","key_machinery":"The load-bearing machinery is a tetrahedral-lattice coarse-grained encoding combined with a four-term Hamiltonian $H_t = \\lambda_c H_c + \\lambda_g H_g + \\lambda_d H_d + \\lambda_i H_i$, where the terms enforce chirality, backbone geometry, residue-collision avoidance, and pairwise amino acid interaction energies. Each residue becomes a node with four allowed continuation directions and a fixed $\\sim$109.4° bond angle, so every conformation maps to a quantum state and the Hamiltonian's expectation value is the conformational energy. A variational quantum eigensolver—a hybrid loop in which a parameterized circuit is updated classically to lower $\\langle \\psi | U^\\dagger(\\theta) H U(\\theta) |\\psi \\rangle$—finds the low-energy state; the circuit is then measured 100,000 times and the sampled bitstrings are reconstructed into atomic coordinates. Extra ancilla qubits are allocated during compilation to shorten the circuit by reducing routing overhead, which is the strategy that makes deep circuits runnable on current hardware.","core_discovery":"On the paper's own terms, the discovery is that a coarse-grained quantum optimization—each residue mapped to a tetrahedral lattice node, the conformational energy encoded as a four-term Hamiltonian, and the ground state found with a variational quantum eigensolver on a real superconducting processor—produces fragment structures that beat the two dominant deep-learning predictors on the two metrics that matter for docking. Compared with experimentally determined X-ray structures, the quantum fragments give lower root-mean-square deviation (RMSD) of backbone carbon positions in 51 of 55 cases against the older deep-learning model and 40 of 55 against the newest one; compared with the same deep-learning models, docking against native ligands gives lower (more favorable) binding-affinity scores in 53 of 55 and 50 of 55 cases, respectively. The paper reads these results as evidence that quantum-first modeling, grounded in physical energy minimization rather than training-data statistics, can handle short ligand-binding fragments better than data-driven approaches.","pith_inferences":["A fairer head-to-head would dock ligands against full-length structures from the deep-learning models and compare those poses with poses from the quantum fragments; the paper compares all methods on isolated fragments, which likely disadvantages models that rely on global context.","If the advantage survives that embedding test, the natural next step is to use the quantum fragment as a local perturbation inside a classical pipeline: generate a full structure, then refine the pocket region with the quantum Hamiltonian.","The same Hamiltonian and variational procedure could be run with classical optimization on a simulator for the smaller fragments; a result matching the quantum hardware outputs would suggest the advantage comes from the energy model rather than from quantum noise, which is the paper's stated mechanism for escaping local minima.","Averaging the reported cost over the 55 fragments puts each structure at roughly $18,000, so extending the dataset to hundreds of fragments will require either cheaper quantum access or a hybrid screening step that selects only the most informative pockets for quantum computation."],"forward_implications":["QDockBank gives researchers a reusable, docking-ready benchmark of 55 quantum-generated fragments with metadata, so future quantum structure predictions can be compared on identical terms.","If the fragment-level accuracy holds, quantum coarse-grained modeling becomes a practical local-refinement tool for binding pockets, complementing global deep-learning predictions.","The reported win rates (96.4% affinity versus the older baseline, 90.9% versus the newer one) set concrete targets that any alternative method—classical or quantum—can be tested against.","The dataset's coverage of nearly all amino acid interaction pairs (395 of 400) makes it usable for training or validating energy functions and coarse-grained potentials, not only for docking.","The cost figures (more than 60 processor-hours, over one million dollars for 55 fragments) give a concrete baseline for judging whether quantum structure prediction is becoming economically feasible."],"supporting_citations":[{"why":"Supplies the docking engine and affinity scores used for all method comparisons.","marker":"[22]"},{"why":"Defines the variational quantum eigensolver framework that carries the energy minimization.","marker":"[24]"},{"why":"Motivates the coarse-grained tetrahedral lattice representation used for the quantum encoding.","marker":"[23]"},{"why":"Documents the 127-qubit superconducting processor on which all fragments were generated.","marker":"[27]"},{"why":"Supports the claim that noise on current quantum hardware can aid escape from local minima, justifying real-device execution.","marker":"[28]"},{"why":"Is the newest deep-learning predictor that the paper's quantum fragments claim to beat in RMSD and docking affinity.","marker":"[2]"},{"why":"Is the earlier deep-learning predictor used as a primary baseline.","marker":"[3]"},{"why":"Implements the earlier deep-learning baseline in the comparison runs.","marker":"[39]"},{"why":"Supplies the experimental structures and native ligands used for fragment extraction, RMSD, and docking evaluation.","marker":"[34]"},{"why":"Provides the reference interaction matrix used to verify amino-acid coverage.","marker":"[40]"}],"fun_headline_variants":["Quantum dataset beats AlphaFold on 55 protein fragments","QDockBank: 55 quantum fragments beat AlphaFold on docking","Quantum-built fragments outperform AlphaFold2 and 3 on docking","First million-dollar quantum dataset tops AlphaFold in docking"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes that ligand-binding quality is captured by docking a short, isolated 5-to-14-residue fragment, treated as a rigid receptor, against its native ligand; if a fragment removed from its protein context does not reproduce the real pocket's binding behavior, the reported docking-affinity advantage over deep-learning models would not carry over to full-length proteins.","fun_headline_variants_meta":{"raw":{"variants":["Quantum dataset beats AlphaFold on 55 protein fragments","QDockBank: 55 quantum fragments beat AlphaFold on docking","Quantum-built fragments outperform AlphaFold2 and 3 on docking","First million-dollar quantum dataset tops AlphaFold in docking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000628,"raw_usage":{"total_tokens":2884,"prompt_tokens":909,"completion_tokens":1975,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":1919}},"tokens_in":525,"tokens_out":1975,"duration_ms":14207,"temperature":1.0,"reasoning_tokens":1919,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:15:02.287625+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same docking comparison on full-length protein structures (or on fragments embedded in the surrounding pocket residues) rather than on the isolated fragments: if the deep-learning models' full-length structures match or beat the quantum fragments' affinity scores, the paper's central advantage is an artifact of truncation. A simpler check is to dock the native ligand against the quantum fragment and against the corresponding full experimental pocket and compare the resulting poses.","supporting_citations":[{"cited_title":"The variational quantum eigensolver: a re- view of methods and best practices.Physics Reports, 986:1–128, 2022","cited_arxiv_id":null,"evidence_quote":"Defines the variational quantum eigensolver framework that carries the energy minimization."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the coarse-grained tetrahedral lattice representation used for the quantum encoding."},{"cited_title":"Ibm quantum breaks the 100-qubit processor barrier","cited_arxiv_id":null,"evidence_quote":"Documents the 127-qubit superconducting processor on which all fragments were generated."},{"cited_title":"Evidence for the utility of quantum computing before fault tolerance.Nature, 618(7965):500–505, 2023","cited_arxiv_id":null,"evidence_quote":"Supports the claim that noise on current quantum hardware can aid escape from local minima, justifying real-device execution."},{"cited_title":"Highly accurate pro- tein structure prediction with alphafold.nature, 596(7873):583–589, 2021","cited_arxiv_id":null,"evidence_quote":"Is the earlier deep-learning predictor used as a primary baseline."},{"cited_title":"Colabfold: making protein folding ac- cessible to all.Nature methods, 19(6):679–682, 2022","cited_arxiv_id":null,"evidence_quote":"Implements the earlier deep-learning baseline in the comparison runs."},{"cited_title":"The pdbbind database: Collection of binding affinities for protein- ligand complexes with known three-dimensional structures.Journal of medicinal chemistry, 47(12):2977–2980, 2004","cited_arxiv_id":null,"evidence_quote":"Supplies the experimental structures and native ligands used for fragment extraction, RMSD, and docking evaluation."},{"cited_title":"Estima- tion of effective interresidue contact energies from protein crystal structures: quasi-chemical approxi- mation.Macromolecules, 18(3):534–552, 1985","cited_arxiv_id":null,"evidence_quote":"Provides the reference interaction matrix used to verify amino-acid coverage."}],"review_version":1}