{"id":"9da053ba-df59-46d5-8851-0c1c3444c16c","arxiv_id":"2607.09737","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Q-Score scores docking poses via GNN-predicted NBO E(2) energies solved as MWVCP with DC-QAOA, yielding rankings orthogonal to Vina and free of size bias.","lead":"Q-Score ranks protein-ligand binding by packing GNN-predicted orbital donor-acceptor energies into a graph and solving a maximum-weight clique with DC-QAOA. It is uncorrelated with classical docking scores, free of molecular-weight bias, and runs on 6-qubit NISQ hardware.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The orthogonality/enrichment claims rest on SIMG* E(2) weights that are never validated against true NBO for protein-ligand interfaces.","rationale":"The reader's weakest_assumption correctly isolates the single load-bearing condition: that SIMG* intermolecular E(2) energies are chemically reliable enough to serve as docking node weights. Everything else in the strongest claim (orthogonality, MW independence, IQEF) follows mechanically once those weights are accepted. The paper supplies no external check of that condition, only internal consistency (rho with its own E(2) and the expected-by-construction IQEF). Classical solvers already solve the 10-qubit MWVCPs exactly, three redocking targets fail under the default penalty, and no code/data are released; these are secondary. Because the concern is already flagged by the reader and the quantitative claims remain well-supported under the stated assumption, the verdict stays CONDITIONAL rather than moving to REJECT. A single DFT-NBO validation on the existing complexes would settle the issue cleanly.","tokens_in":18838,"tokens_out":643,"duration_ms":5301,"concrete_test":"On the 11 redocking complexes (or a random 50 of the 1000 generated poses), recompute true NBO second-order E(2) for every intermolecular donor-acceptor pair that SIMG* scored, using a standard DFT package (e.g., omegaB97X-D/def2-SVP). Report Spearman rho and MAE between SIMG* and true E(2) for the selected OIG anchors. If rho < 0.6 or the top-ranked anchors systematically mismatch known critical contacts (e.g., the Cys145 covalent pair in 8SKH), the chemical fidelity of the node weights is insufficient and the orthogonality/enrichment claims become GNN artifacts.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim (rho(Vina,Q)=0.05, rho(E(2),Q)=0.90, IQEF_Q=1.96, free of MW bias) treats the GNN-predicted intermolecular E(2) values as faithful orbital physics. Section IV-B states SIMG* is trained only on QM9 and GEOM (small organic molecules) and is applied zero-shot to protein-ligand complexes after a 5 A binding-site cut. No comparison of predicted vs. actual NBO E(2) (or even vs. DFT interaction energies) is reported for any of the 11 co-crystal interfaces or the 1000 generated poses. Consequently the observed orthogonality to Vina and the 2x enrichment for high-E(2) molecules could be properties of the GNN's learned representation rather than of genuine stereoelectronic complementarity. The paper itself notes that the E(2) correlation and IQEF are expected by construction; without external chemical validation of the node weights, those metrics do not establish that Q-Score captures true orbital physics.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript introduces Q-Score, a docking scoring function that replaces empirical pairwise sums with GNN-predicted NBO second-order perturbation energies E(2). These are aggregated into an Orbital Interaction Graph (OIG) whose maximum-weight vertex clique is solved by Digitized-Counterdiabatic QAOA; the clique weight is the score. Redocking on 11 co-crystal structures shows DC-QAOA recovers the exact classical optimum on 8 of 11 targets at 10 qubits (default penalty P=6). On 1000 Pocket2Mol-generated molecules across 10 targets, Q-Score is uncorrelated with Vina (mean Spearman ρ=0.05), strongly correlated with mean E(2) (ρ=0.90), free of molecular-weight bias, and yields an Interaction Quality Enrichment Factor of 1.96. Simulation mean approximation ratio is 0.94 (52% exact); 1000 circuits on IBM Eagle r3 establish a practical 6-qubit hardware boundary.","tokens_in":19115,"tokens_out":1244,"duration_ms":9548,"significance":"If the GNN-derived node weights are chemically faithful, the work supplies a compact, quantum-native scoring primitive that is demonstrably orthogonal to classical empirical scores and free of the well-known molecular-weight bias. The hybrid pipeline (SIMG* → OIG → DC-QAOA → Kabsch pose) is fully automated, the quantum solver is benchmarked at scale (1000 instances) with transparent approximation-ratio tables, and hardware execution on a 127-qubit Eagle processor is reported with match rates and circuit metrics. These elements make the paper a concrete data point for both quantum-chemistry-informed docking and NISQ combinatorial optimization, even though classical solvers remain competitive at the 10-qubit scale examined.","major_comments":[{"comment":"Section IV-B and the scoring-comparison claims (Table II, ρ(E(2),Q)=0.90, IQEF=1.96): SIMG* is trained exclusively on QM9/GEOM small-molecule data and applied zero-shot to protein–ligand interfaces after a 5 Å cut. No comparison of predicted intermolecular E(2) against true NBO (or even DFT interaction energies) is provided for any of the 11 co-crystal complexes or the 1000 generated poses. Because Q-Score is defined as the sum of those same E(2) weights (Eq. 9), the reported orbital-selectivity metrics are tautological once a clique is found; without external chemical validation of the node weights, the orthogonality and enrichment results cannot be attributed to genuine stereoelectronic physics rather than to properties of the GNN representation.","section":null},{"comment":"Section V-C, Table I and Table VI: under the default penalty P=6, three of eleven redocking targets (FXR, PRMT5, SRC) return invalid cliques whose scores exceed the classical optimum. The subsequent penalty sweep shows that the optimal P is instance-dependent (P=10 recovers two targets but degrades others). Because the Q-Score itself is the weight of the recovered clique, an unreliable constraint-enforcement mechanism directly undermines score fidelity on a non-negligible fraction of targets; an adaptive or theoretically justified penalty schedule is needed before the method can be treated as a robust scorer.","section":null},{"comment":"Section V-C (8SKH case study) and the redocking evaluation: Kabsch reconstruction yields all-atom RMSD of 3.34–3.45 Å. While the paper correctly notes that prior graph-based docking formulations often omit pose reconstruction, these RMSD values remain only modest by conventional docking standards and are reported for a single target. Without a systematic RMSD table across all 11 co-crystals (or a comparison against Vina poses on the same OIG anchors), it is difficult to judge whether the selected cliques correspond to chemically useful binding modes.","section":null}],"minor_comments":[{"comment":"Abstract and Section V-D: the phrase “enriching for strong orbital interactions at twice the random rate” should be qualified by the paper’s own acknowledgment that IQEF is expected by construction once E(2) weights are used.","section":null},{"comment":"Figure 3(b) heatmap uses “inv” for invalid cliques; a brief legend or caption note would improve readability.","section":null},{"comment":"Table VII reports wall-clock times for a single BRAF instance; adding mean ± std over the 1000-molecule set (already mentioned as 235 s) would make the scaling claim more precise.","section":null},{"comment":"The distance tolerance τ is listed as target-specific in Table VI but the selection criterion is not described; a short sentence on how τ was chosen would aid reproducibility.","section":null}],"recommendation":"major_revision","confidential_remarks":"The central novelty claim rests on the chemical fidelity of SIMG* intermolecular E(2) predictions that are never validated on biomolecular interfaces. If the authors cannot supply at least a limited DFT/NBO benchmark on a few of the co-crystal sites, the paper risks being read as a pure quantum-optimization demonstration rather than a chemistry contribution; that would affect fit for a physics.chem-ph or computational-chemistry venue. The quantum-hardware results are solid and self-contained, so a transfer to a quantum-computing journal would also be viable if the chemical-validation gap cannot be closed."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The real increment here is not the clique formulation or QAOA itself (Ding et al. already did that with pharmacophore tables). It is the swap to continuous SIMG* E(2) node weights, the clean orthogonality numbers against Vina on 1000 Pocket2Mol molecules, and the systematic 1000-circuit Eagle scaling table. That package is useful for people who care about quantum-for-chemistry applications and about size-bias-free scoring.\n\nWhat they do well: the pipeline is fully specified (binding-site cut, greedy diverse anchors, QUBO with explicit P, DC-QAOA with independent Pauli angles, Kabsch pose recovery). Simulation AR 0.94 / 52 % exact at 10q is reported cleanly; hardware match rates (65 % at 6q, 13 % at 10q) are honest and establish a concrete NISQ boundary. They themselves flag that the rho=0.90 with mean E(2) and IQEF=1.96 are expected by construction, so the non-circular claim is the near-zero correlation with Vina and with molecular weight. That orthogonality is the part worth citing.\n\nSoft spots, in proportion. Three of eleven redocking targets produce invalid cliques under the default P=6; the optimal penalty is instance-dependent and they show the trade-off table. RMSDs of 3.3–3.5 Å are only modest. Classical exact solvers already finish the 10-qubit instances instantly; the quantum-advantage argument is only asymptotic. Most importantly, SIMG* was trained on QM9/GEOM and applied zero-shot to protein–ligand interfaces with no predicted-vs-true NBO (or even DFT) check on any of the 11 co-crystals. So the “orbital quality” story could still be a property of the GNN’s learned representation rather than genuine stereoelectronics. That is the load-bearing assumption the stress-test correctly flags; it does not kill the engineering result, but it does limit how far one can push the chemical interpretation.\n\nNo code or data release, which hurts reproducibility. Citation pattern is fine; they properly credit Ding, SIMG*, and the DC-QAOA paper.\n\nWho it is for: quantum-chemistry / NISQ application people and anyone hunting for scoring functions that are not size-biased. It deserves a serious referee. I would bring it to reading group as a concrete worked example of hybrid AI–quantum docking, with the GNN-validation gap as the discussion prompt. I would cite the orthogonality and hardware tables if I were writing on quantum docking or size-bias-free scoring; I would not yet treat the E(2) numbers as validated orbital physics.","headline":"Solid engineering of GNN-NBO weights into an existing MWVCP/QAOA docking pipeline, with honest NISQ scaling data; the headline orthogonality claim is real but the orbital-physics interpretation is under-validated.","tokens_in":19774,"tokens_out":671,"would_cite":true,"duration_ms":5971,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.67.Lx","87.15.Aa","31.15.ae"],"model":"grok-4.5","headline":"Q-Score ranks protein-ligand binding by orbital donor-acceptor energies solved as a quantum clique problem, not by summing empirical atom-pair contacts.","keywords":["molecular docking","scoring function","QAOA","orbital interaction graph","NBO","NISQ","drug discovery","maximum-weight clique"],"falsifier":"Replace the GNN-predicted E(2) node weights with high-level ab initio NBO energies on a held-out set of protein-ligand complexes and check whether the Spearman correlation with classical scores remains near zero and the interaction-quality enrichment factor remains near 2.","tokens_in":19712,"feed_emoji":"⚛️","tokens_out":882,"duration_ms":9226,"temperature":0.7,"pith_summary":"Classical docking scores add up pairwise contacts and therefore favor larger molecules and miss the orbital charge-transfer effects that often decide binding specificity. This paper replaces that additive sum with Q-Score: a graph neural network predicts second-order orbital interaction energies, those energies become node weights on a compatibility graph, and a maximum-weight clique is found by digitized-counterdiabatic QAOA so that only mutually realizable anchors contribute. Across eleven co-crystal structures the quantum solver recovers the exact clique on eight targets at ten qubits. On one thousand AI-generated molecules the resulting ranks are statistically uncorrelated with a standard classical score, track orbital quality almost perfectly, show no molecular-weight bias, and enrich for strong orbital contacts at nearly twice the random rate. Hardware runs on an IBM Eagle processor confirm that six-qubit instances remain solvable on present-day noisy devices.","feed_headline":"Docking score from orbital energies, not atom-pair sums","feed_subtitle":"Quantum clique solver ranks ligands by donor-acceptor quality, free of size bias and uncorrelated with classical scores","key_machinery":"The Orbital Interaction Graph (OIG): each node is an atom-pair interaction anchor weighted by summed GNN-predicted E(2) energies, edges encode rigid-body geometric compatibility, and the Q-Score is the total weight of the maximum clique found by DC-QAOA.","core_discovery":"Encoding GNN-predicted NBO second-order energies as node weights of an Orbital Interaction Graph and solving the resulting maximum-weight vertex clique with DC-QAOA yields a docking score that is orthogonal to classical empirical scoring, driven by orbital interaction quality, free of molecular-weight bias, and capable of recovering exact optimal cliques on most redocking targets at ten qubits.","pith_inferences":["If the GNN energies prove reliable, the same OIG construction could serve as a drop-in rescoring stage for any docking engine that already produces poses.","Adaptive or instance-dependent penalty schedules would be required before the method can be applied blindly across chemically diverse targets without manual retuning.","The approach supplies a concrete test-bed for error-mitigation techniques: any method that restores ten-qubit bitstring fidelity would immediately enlarge the usable pocket size."],"forward_implications":["Docking pipelines can rank candidates by stereoelectronic complementarity instead of atom count, reducing size bias in virtual screens.","Q-Score and classical scores can be used jointly as complementary filters because they disagree on top-ranked molecules.","NISQ devices can already solve the six-qubit clique instances that arise from compact binding-site graphs.","As qubit counts grow, larger Orbital Interaction Graphs become feasible while classical exact solvers face exponential growth."],"fun_headline_variants":["Q-Score ranks docks by orbital cliques via DC-QAOA","Orbital-energy vertex cliques score binding free of size bias","Max-weight orbital graphs yield docking scores orthogonal to classics","DC-QAOA recovers exact cliques from GNN donor-acceptor weights","Quantum clique docking driven by orbital quality not pairwise sums"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The graph neural network, trained only on small-molecule quantum datasets, must produce chemically trustworthy intermolecular orbital energies for real protein-ligand interfaces; if those predicted energies are systematically wrong, both the orthogonality and the enrichment results become artifacts of the model rather than of orbital physics.","fun_headline_variants_meta":{"raw":{"variants":["Q-Score ranks docks by orbital cliques via DC-QAOA","Orbital-energy vertex cliques score binding free of size bias","Max-weight orbital graphs yield docking scores orthogonal to classics","DC-QAOA recovers exact cliques from GNN donor-acceptor weights","Quantum clique docking driven by orbital quality not pairwise sums"]},"model":"grok-4.5","effort":"low","cost_usd":0.005304,"raw_usage":{"total_tokens":1384,"prompt_tokens":750,"num_sources_used":0,"completion_tokens":93,"cost_in_usd_ticks":53040000,"prompt_tokens_details":{"text_tokens":750,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":541,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":750,"tokens_out":93,"duration_ms":4976,"temperature":1.0,"reasoning_tokens":541,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T16:37:53.997268+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Replace the GNN-predicted E(2) node weights with high-level ab initio NBO energies on a held-out set of protein-ligand complexes and check whether the Spearman correlation with classical scores remains near zero and the interaction-quality enrichment factor remains near 2.","supporting_citations":[],"review_version":1}