{"id":"44cbc401-9305-4d41-a539-d1a119e38897","arxiv_id":"2501.01811","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A TensorNet-based neural network potential (AceFF 1.0) in an NNP/MM scheme achieves RBFE accuracy close to OPLS4, exceeding GAFF2 and ANI-2x on most JACS benchmark targets, at a 2 fs timestep.","lead":"This paper tests a new neural network potential, AceFF 1.0, for computing relative protein-ligand binding free energies, and reports accuracy close to the industry-standard OPLS4 force field while beating older classical and neural force fields on most targets. It also shows the NNP can run at twice the usual timestep, roughly halving simulation cost.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline improvement over GAFF2 is within bootstrap confidence intervals; without a paired significance test the central claim of improved RBFE accuracy is not statistically established.","rationale":"The reader's conditional verdict is reasonable, and their weakest-assumption analysis of mechanical embedding (Eq. 1) correctly identifies a limitation for generalization to cases where ligand-environment polarization matters. However, the single most load-bearing issue for the central claim is more immediate: the paper's headline numerical improvements over GAFF2 are all within the reported 95% bootstrap confidence intervals. The paper never performs a paired significance test, so the observed differences could be explained by sampling variability. This matters because the abstract and conclusions make a strong comparative claim ('improved accuracy and correlation' vs GAFF2) that the data may not support at the stated precision. The concrete test proposed above would settle the issue using the already-released edge-level data. If the paired test shows significance, the central claim survives and the mechanical-embedding concern remains as a caveat about generalizability. If the paired test fails, the verdict should move toward requiring the authors to reframe the result as comparable rather than improved. Since the reader already issued a CONDITIONAL verdict with requests for clarification, my read does not change that verdict: it sharpens the condition by adding a statistical robustness requirement. Agreement with the reader is partial rather than full because the reader identified a different weakest assumption; both concerns bear on whether the practical utility claim is trustworthy, but the statistical issue is more directly tied to the reported numbers and is testable with the released artifacts.","tokens_in":19606,"tokens_out":6598,"duration_ms":73577,"concrete_test":"Re-analyze the released per-edge ΔΔG data (GitHub quantumbind_rbfe) by pairing each edge's absolute error for AceFF and GAFF2 and running a paired bootstrap or Wilcoxon signed-rank test on the difference, using the same 1000 resamples as the paper. Report the 95% confidence interval of the paired mean difference and the p-value for RMSE, MAE, and Kendall tau. If the paired CI includes 0 or p > 0.05, the headline claim of improved accuracy over GAFF2 is not statistically supported and should be softened to 'comparable'.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim is that AceFF 1.0 improves over GAFF2: for ΔG, RMSE 0.99 vs 1.17, MAE 0.79 vs 0.90, Kendall tau 0.59 vs 0.55 (Table S2). However, the reported 95% bootstrap confidence intervals overlap for every metric: RMSE [0.89,1.10] vs [1.03,1.30], MAE [0.71,0.89] vs [0.79,1.01], tau [0.52,0.65] vs [0.48,0.62]. The ΔΔG metrics in Table S4 show the same pattern: RMSE [1.10,1.33] vs [1.30,1.62], MAE [0.85,1.03] vs [0.97,1.20], tau [0.40,0.51] vs [0.34,0.49]. No significance test or paired error analysis is reported anywhere in the paper. Because the per-target differences are small and generally within the bootstrap error, the observed 'improvement' may be sampling noise rather than a genuine force-field effect. This is load-bearing: the abstract and conclusions assert 'overall improved accuracy and correlation' and frame AceFF as a practical drop-in replacement; if the paired difference is not significant, only 'comparable to GAFF2' is supported. The mechanical-embedding limitation of Eq. 1 is a real generalization risk, but it is secondary to whether the headline benchmark result is established at all.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents QuantumBind-RBFE, an alchemical relative binding free energy protocol that uses the AceFF 1.0 neural network potential for the ligand within an NNP/MM mechanical-embedding scheme and the Alchemical Transfer Method (ATM). AceFF 1.0 is benchmarked on seven targets of the JACS dataset (280 edges; PTP1B is explicitly excluded because its dianionic ligands lie outside the trained charge range, though a stress test on PTP1B is also reported). The paper claims improved accuracy and correlation over GAFF2 under the same protocol (RMSE 0.99 vs 1.17 kcal/mol, MAE 0.79 vs 0.90, Kendall tau 0.59 vs 0.55), comparable ranking but slightly worse RMSE/MAE than OPLS4/FEP+ (a cross-protocol comparison), better accuracy than ANI-2x, and stable 2 fs timestep operation. The authors conclude that AceFF 1.0 can already serve as a practical ligand force field for RBFE calculations while noting limitations for rare charge states and the mechanical-embedding approximation.","tokens_in":19882,"tokens_out":4596,"duration_ms":47054,"significance":"If the headline claims are accepted, this is a practically important result: it would demonstrate that a publicly released neural network potential can replace a classical ligand force field in an established alchemical RBFE workflow, with meaningful speed advantages over earlier NNP/MM implementations. The study has genuine strengths: triplicate 70 ns/replica ATM simulations, UWHAM analysis, 95% bootstrap confidence intervals, a same-protocol GAFF2 control, and public availability of the model, input structures, and analysis scripts. The additional PTP1B stress test, run outside the model's trained charge range, is a useful falsifiability check. However, the central quantitative claim of improved accuracy over GAFF2 is not supported by the reported statistics alone, because the bootstrap confidence intervals overlap for every headline metric and no paired significance test is provided. The mechanical-embedding assumption in Eq. (1) is also acknowledged but not stress-tested, which limits the generality of the conclusions.","major_comments":[{"comment":"The claim that AceFF 1.0 shows 'improved accuracy' over GAFF2 is not statistically established. For the all-target ΔG metrics, the 95% bootstrap intervals overlap for RMSE ([0.89,1.10] vs [1.03,1.30]), MAE ([0.71,0.89] vs [0.79,1.01]), and Kendall tau ([0.52,0.65] vs [0.48,0.62]). The same is true for the ΔΔG metrics in Table S4 (RMSE [1.10,1.33] vs [1.30,1.62], MAE [0.85,1.03] vs [0.97,1.20], tau [0.40,0.51] vs [0.34,0.49]). Since the abstract and conclusions assert 'overall improved accuracy and correlation', the authors should either report a paired significance test (e.g., a paired bootstrap over ligands or edges, or a permutation test on per-edge ΔΔG errors) or soften the claim to 'comparable to, and in some targets better than, GAFF2'. Given the data and code are public, such a test is straightforward and would resolve whether the point-estimate differences are sampling noise or a genuine force-field effect.","section":"§3.1, Table S2"},{"comment":"The bootstrap resampling unit is not specified. The reported confidence intervals are used to support the main comparisons, but the ΔG metrics are computed over ligands and the ΔΔG metrics over edges, and edges share ligands through the perturbation network. If the bootstrap resamples individual edges rather than independent units (ligands, or clusters of correlated edges), the intervals will be too narrow and the overlap assessment in the previous comment would be optimistic. Please state the resampling unit explicitly and, if edges are resampled, justify why dependencies are negligible or switch to a cluster bootstrap by ligand.","section":"§2.5"},{"comment":"The mechanical-embedding assumption that all ligand-protein and ligand-solvent nonbonded interactions are adequately described by fixed-charge classical MM (RESP charges, TIP3P, ff14SB) is the load-bearing modeling choice, but it is not stress-tested. The paper attributes AceFF's improvements over GAFF2 mainly to better treatment of ligand internal energetics, which is consistent with this choice, yet the claim that AceFF is a 'viable alternative for RBFE calculations' in general goes beyond what can be concluded from the JACS systems, where polarization and charge-transfer effects may be small for the tested ligands. I recommend either adding a limited test (e.g., comparison against a range-corrected or explicitly polarizable scheme on a subset) or explicitly restricting the conclusion to systems where the fixed-charge environment approximation is adequate.","section":"§2.1, Eq. (1)"},{"comment":"The claim that '2 fs timestep simulations demonstrate comparable accuracy to 1 fs runs' rests on point estimates whose differences are not tested for significance. For example, Table S9 shows TYK2 ΔG RMSE of 0.47 at 1 fs versus 0.74 at 2 fs, with overlapping intervals ([0.23,0.66] vs [0.38,0.81]); similar differences appear for THROMBIN (0.57 vs 0.80). The stability at 2 fs is convincing, but 'comparable accuracy' would be better supported by reporting paired per-target differences or by presenting the claim as 'within the observed bootstrap uncertainty' rather than as equivalence.","section":"§3.3, Tables S9-S10"}],"minor_comments":[{"comment":"The manuscript states that PTP1B is omitted from the evaluation, but Tables S3, S5, S6, and S7 report PTP1B rows for AceFF 1.0 and the charge-model comparison. The role of these rows should be clarified; if they are part of the 'going beyond the limits' stress test, they should be labeled as such in the table captions to avoid confusion with the main benchmark.","section":"§2.2 and Supporting Information"},{"comment":"The sentence 'We used the Amber ff14SB parameters as well as the TIP3P water model' appears without a citation in the body text; the ff14SB reference is present in the reference list but should be cited at the point of use, as done for TIP3P elsewhere.","section":"§2.4"},{"comment":"The discussion of the MCL1 decrease in Kendall tau (0.28 for AceFF vs 0.44 for GAFF2) says it is 'primarily due to the overprediction of binding affinities for five ligands', but no quantitative support for this attribution is provided; a sentence identifying those five ligands and their predicted versus experimental values would make the claim verifiable.","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The authors' affiliation with Acellera and the proprietary nature of the AceFF training set are disclosed in the manuscript, which is adequate. The strongest finding is the stable 2 fs timestep and the same-protocol GAFF2 comparison; the 'improvement over GAFF2' claim needs a paired significance test. Since the data and code are public, this is fixable within the scope of a revision. The OPLS4 comparison is explicitly cross-protocol and should be framed as indicative rather than as a head-to-head benchmark."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuine benchmark, not a hype piece. The first RBFE test of AceFF 1.0 covers 7 JACS targets, includes charged ligands, and ships structures, weights, and a tutorial. The same-protocol GAFF2 control is exactly the right comparison, and the PTP1B failure included in the paper is a sign the authors know where the model breaks. The 2 fs timestep claim is supported by NVE energy conservation checks and comparable 1 fs vs 2 fs results on a subset.\n\nThe main weak spot is statistical, and it lands on the paper's central sentence. AceFF's reported improvement over GAFF2 on ΔG has overlapping 95% bootstrap intervals for every metric: RMSE 0.99 [0.89,1.10] vs 1.17 [1.03,1.30], MAE 0.79 [0.71,0.89] vs 0.90 [0.79,1.01], Kendall tau 0.59 [0.52,0.65] vs 0.55 [0.48,0.62]. Same pattern for ΔΔG. With no paired test, the only strictly supported statement is 'comparable to GAFF2.' That doesn't kill the paper, but the abstract and conclusions should be reworded, and the authors should run a paired bootstrap over the same edges or report per-edge differences. The current confidence intervals are unpaired and not designed to test a difference.\n\nThe other caveats are secondary. The promising per-target ranking improvements are real observations but small-sample and also CI-overlapping. The mechanical-embedding coupling in Eq. 1 means the NNP only fixes ligand intramolecular energetics; if ligand–environment polarization matters, this benchmark won't capture it. The OPLS4 comparison mixes engines and protocols, and 0.99 vs 0.78 kcal/mol is more than 'slightly' worse, though the correlations are comparable. The AceFF training set is proprietary, but weights are released and the external validation is framed as a benchmark rather than a fit to binding data.\n\nBottom line: this deserves a serious referee. The experimental design is clean, the artifacts are real, and the central question—can a released NNP work as a drop-in ligand potential for ATM—is answered 'yes, at least as well as GAFF2, with a promising but statistically unproven edge.' I'd recommend peer review with a request for a paired significance analysis.","headline":"Careful, reproducible benchmark of a new NNP for RBFE; the improvement over GAFF2 is plausible but the reported bootstrap intervals overlap and the abstract overstates it.","tokens_in":20530,"tokens_out":2458,"would_cite":true,"duration_ms":25372,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a released neural network potential, AceFF 1.0, improves relative binding free energy predictions over GAFF2 and reaches ranking correlations comparable to OPLS4 when used as the ligand potential in an NNP/MM scheme.","keywords":["relative binding free energy","neural network potentials","AceFF","TensorNet","alchemical transfer method","NNP/MM","force field accuracy","drug discovery"],"falsifier":"Replace the classical coupling term in the same NNP/MM scheme with a polarizable or range-corrected quantum treatment and rerun the seven-target benchmark; if RMSE falls well below 0.99 kcal/mol, the remaining error is dominated by the mechanical-embedding coupling rather than by the ligand's intramolecular potential, weakening the paper's attribution of the improvement to AceFF. A cheaper check is to benchmark a target rich in halogen bonds or salt bridges, where fixed-charge electrostatics are known to struggle.","tokens_in":19355,"feed_emoji":"🧪","tokens_out":8701,"duration_ms":80283,"temperature":0.7,"pith_summary":"This paper tests whether a released neural network potential, AceFF 1.0, can serve as the ligand force field in relative binding free energy (RBFE) calculations. Using the Alchemical Transfer Method with an NNP/MM mechanical-embedding setup, the authors report that AceFF improves accuracy and ranking over GAFF2 on the seven-target JACS benchmark (RMSE 0.99 vs 1.17 kcal/mol, MAE 0.79 vs 0.90, Kendall tau 0.59 vs 0.55). Against OPLS4 with FEP+, AceFF shows slightly larger errors but comparable correlations (Kendall tau 0.59 vs 0.66), and the authors note the comparison mixes different simulation engines and protocols. The same simulations run stably at 2 fs, twice the usual NNP timestep, with accuracy close to 1 fs runs. If these results hold, a publicly available neural potential can be dropped into drug-discovery workflows without the torsion-scan parameterization of classical force fields.","feed_headline":"Neural net ligand force field beats GAFF2 on binding benchmark","feed_subtitle":"AceFF 1.0 matches OPLS4 in ranking and runs at twice the usual NNP timestep.","key_machinery":"The carrying object is AceFF 1.0, a one-layer TensorNet neural network potential: an equivariant message-passing architecture that represents atomic environments with Cartesian tensor features and predicts the ligand's intramolecular energy and forces. It is embedded through the NNP/MM energy expression $V = V_{\\text{NNP}}(\\mathbf{r}_{\\text{NNP}}) + V_{\\text{MM}}(\\mathbf{r}_{\\text{MM}}) + V_{\\text{NNP-MM}}(\\mathbf{r})$, in which the coupling term is the classical electrostatic and van der Waals interactions between the ligand and the surrounding protein and solvent. The free energy difference is computed with the Alchemical Transfer Method, which alchemically transfers the ligand between bound and unbound states in a dual-topology simulation box. The NNP replaces all bonded ligand terms, including torsions, so no torsion-scan fitting is needed; the paper notes that range-corrected variants that include short-range environment interactions exist but are not used here.","core_discovery":"The central claim is that AceFF 1.0, a TensorNet-based neural network potential trained on quantum-chemical energies and forces, is accurate enough to replace classical ligand force fields in RBFE calculations. In the QuantumBind-RBFE protocol, the total energy is split as $V = V_{\\text{NNP}}(\\mathbf{r}_{\\text{NNP}}) + V_{\\text{MM}}(\\mathbf{r}_{\\text{MM}}) + V_{\\text{NNP-MM}}(\\mathbf{r})$, where only the ligand's intramolecular energy is computed by the NNP and all ligand-environment interactions are classical. Across seven protein targets and 280 perturbation edges, the model outperforms GAFF2 run under the identical protocol, reduces outliers, and improves compound prioritization on most targets. The paper attributes the gain to better treatment of the ligand's internal energetics, especially strain, since the sampled conformational distributions are similar to GAFF2. AceFF 1.0 is stable at a 2 fs timestep, and the paper demonstrates that 2 fs and 1 fs runs give nearly identical errors and correlations.","pith_inferences":["Because the OPLS4 comparison uses a different molecular dynamics engine and free-energy protocol, the small accuracy gap (RMSE 0.99 vs 0.78 kcal/mol) is not a clean force-field comparison; an identical-protocol study could narrow or reverse it.","The paper switched from AM1BCC to RESP charges compared with its earlier ANI-2x work, so part of the improvement over that baseline may come from the charge model rather than the neural network itself; an ablation with identical charges would separate the two effects.","If the 2 fs stability generalizes to larger, flexible ligands, NNP/MM costs approach classical MM at 4 fs, which would make neural-potential FEP a viable screening tool rather than a retrospective benchmark.","The PTP1B failure is a natural stress test: retraining the model on -2 and +2 charged species would provide a direct test of whether the charge-range limitation, rather than the mechanical-embedding coupling, is the main barrier to broader applicability."],"forward_implications":["Publicly available neural potentials can replace classical ligand force fields in RBFE calculations, eliminating the need for torsion-scan parameterization.","AceFF's stable 2 fs timestep doubles simulation speed relative to earlier NNP models, making NNP/MM free-energy campaigns practical on existing hardware.","Ranking of the most active compounds, not just error metrics, improves over GAFF2 on most of the tested targets, which is the metric that matters in hit-to-lead and lead optimization.","Ligands outside the NNP's training distribution, such as those with a -2 charge, fail badly (RMSE above 3 kcal/mol and negative Kendall tau on PTP1B), so chemical coverage of charge states is a limiting factor."],"supporting_citations":[{"why":"Provides the seven-target benchmark with experimental affinities used as the validation test bed.","marker":"[32]"},{"why":"Supplies the OPLS4/FEP+ results used as the reference comparison for accuracy and ranking.","marker":"[33]"},{"why":"Earlier NNP/MM RBFE study with the ANI-2x potential that this work extends.","marker":"[23]"},{"why":"Introduces the NNP/MM mechanical-embedding energy expression used throughout.","marker":"[27]"},{"why":"Introduces the TensorNet architecture on which AceFF is based.","marker":"[17]"},{"why":"Defines the Alchemical Transfer Method used to compute the relative free energies.","marker":"[29]"},{"why":"Validates ATM on the benchmark and defines the protocol adopted here.","marker":"[30]"},{"why":"Releases the AceFF 1.0 model parameters used for the ligand potential.","marker":"[20]"},{"why":"Provides the molecular dynamics engine that executes the NNP/MM simulations.","marker":"[47]"},{"why":"Supplies the ANI-2x potential used as the NNP baseline for comparison.","marker":"[14]"}],"fun_headline_variants":["AceFF 1.0 NNP beats GAFF2 in binding free energy tests","QuantumBind-RBFE: neural net force field improves RBFE accuracy","TensorNet NNP outperforms GAFF2, matches OPLS4 in affinity bench","NNP ligand model doubles timestep, cuts RBFE errors vs GAFF2"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that all ligand–protein and ligand–water interactions are accurately described by fixed-charge classical molecular mechanics, with the neural network correcting only the ligand's internal energy; if polarization, charge transfer, or other quantum effects across the binding interface matter, the NNP cannot fix those errors.","fun_headline_variants_meta":{"raw":{"variants":["AceFF 1.0 NNP beats GAFF2 in binding free energy tests","QuantumBind-RBFE: neural net force field improves RBFE accuracy","TensorNet NNP outperforms GAFF2, matches OPLS4 in affinity bench","NNP ligand model doubles timestep, cuts RBFE errors vs GAFF2"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000167,"raw_usage":{"total_tokens":1272,"prompt_tokens":978,"completion_tokens":294,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":594,"completion_tokens_details":{"reasoning_tokens":203}},"tokens_in":594,"tokens_out":294,"duration_ms":3483,"temperature":1.0,"reasoning_tokens":203,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:20:32.936504+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replace the classical coupling term in the same NNP/MM scheme with a polarizable or range-corrected quantum treatment and rerun the seven-target benchmark; if RMSE falls well below 0.99 kcal/mol, the remaining error is dominated by the mechanical-embedding coupling rather than by the ligand's intramolecular potential, weakening the paper's attribution of the improvement to AceFF. A cheaper check is to benchmark a target rich in halogen bonds or salt bridges, where fixed-charge electrostatics are known to struggle.","supporting_citations":[{"cited_title":"https://huggingface.co/Acellera/AceForce-1.0, 2024; Accessed: 2024-12-30","cited_arxiv_id":null,"evidence_quote":"Releases the AceFF 1.0 model parameters used for the ligand potential."},{"cited_title":"S.; Huddleston, K","cited_arxiv_id":null,"evidence_quote":"Supplies the ANI-2x potential used as the NNP baseline for comparison."},{"cited_title":"GFN2-xTB—An accurate and broadly parametrized self-consistent tight-binding quantum chemical method with multipole electrostatics and density-dependent dispersion contributions","cited_arxiv_id":null,"evidence_quote":"Provides the seven-target benchmark with experimental affinities used as the validation test bed."},{"cited_title":"E.; Chodera, J","cited_arxiv_id":null,"evidence_quote":"Earlier NNP/MM RBFE study with the ANI-2x potential that this work extends."},{"cited_title":"K.; Greenwood, J.; Romero, D","cited_arxiv_id":null,"evidence_quote":"Introduces the NNP/MM mechanical-embedding energy expression used throughout."},{"cited_title":"TensorNet: Cartesian Tensor Representations for Efficient Learning of Molecular Potentials","cited_arxiv_id":null,"evidence_quote":"Introduces the TensorNet architecture on which AceFF is based."},{"cited_title":"https://www.acellera.com/, 2024; Accessed: 2024-12-30","cited_arxiv_id":null,"evidence_quote":"Defines the Alchemical Transfer Method used to compute the relative free energies."},{"cited_title":"D.; James, N","cited_arxiv_id":null,"evidence_quote":"Validates ATM on the benchmark and defines the protocol adopted here."}],"review_version":1}