{"id":"7fed75f3-f0c0-4d67-bbc9-9b3e82137c9b","arxiv_id":"2508.19385","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A pipeline that starts absolute binding free energy simulations from Boltz-predicted protein-ligand complexes reaches sub-1 kcal/mol mean unsigned error on four kinase targets.","lead":"This paper combines Boltz-2 structure prediction with absolute binding free energy simulations to estimate how tightly drug-like molecules bind proteins, without needing an experimental crystal structure. On four benchmark kinases the pipeline reaches average errors below 1 kcal/mol, suggesting structure prediction may replace crystallography as the starting point for affinity calculations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Manual inclusion of binding partners (e.g., cyclin for CDK2) means the 'without crystal structures' claim is not demonstrated from SMILES+sequence alone.","rationale":"The Reader's verdict is CONDITIONAL with MODERATE confidence, citing manual biological curation as the weakest assumption. This stress-test identifies the same load-bearing concern and sharpens it: the 'without crystal structures' claim is not merely weakened by the CDK2 example—it is contradicted for any target where correct modeling requires knowledge of binding partners that is not in the monomer sequence. The paper's own Figure 7C shows a large degradation in ABFE when the partner is omitted for recently deposited CDK2 structures. This is not an internal inconsistency; the paper acknowledges the need for 'careful tailoring of sequence and biological assembly.' However, it does undercut the abstract's broad feasibility claim if 'sequence information alone' is the intended input. I do not recommend moving the verdict to REJECT because the core numerical result—ABFE with predicted structures reaching MUE < 1 kcal/mol on the four benchmark targets—is plausible and the limitation is explicitly disclosed. CONDITIONAL remains the appropriate verdict, so no change is required. The concrete test proposed would determine whether the concern lands by evaluating the pipeline under the input restrictions that the central claim implies.","tokens_in":14194,"tokens_out":6298,"duration_ms":78319,"concrete_test":"Run a blind reproducibility test on a set of targets with obligate binding partners (e.g., CDK2/cyclin, CDK4/cyclinD, and another 8–12 kinases from the FEP+ benchmark). Provide the pipeline only the target chain's UniProt sequence and the ligand SMILES, with no instructions about partner chains, and require the pipeline to decide automatically whether to include additional chains (e.g., default: none). Compute ABFE with Boltz-2+POSIT for each case. If any target known to require a partner yields MUE > 1 kcal/mol or a binding-pocket expansion > 2 Å relative to the known holo structure, the 'without crystal structures' claim is not supported for that target class. As a complementary check, automate partner selection using sequence-level interaction prediction (e.g., homology to PDB complexes or AlphaFold3 multimer) and verify that the selected assembly reproduces the crystal-like pocket; t","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that Boltz-ABFE enables absolute FEP 'without experimental crystal structures'—requires that the pipeline inputs be only ligand SMILES and protein sequence(s), as stated in Section 2.1.1. Section 2.1.3 undermines this requirement: for CDK2, Boltz-2 alone predicts an artificially expanded pocket and an extended ligand conformation; including CyclinE1 restores the native pocket and the folded ligand pose, and materially improves ABFE results (Figure 7C). The decision to include the cyclin partner is biological knowledge not contained in the CDK2 sequence alone; it is normally obtained from crystal structures, homology, or prior experimental evidence. The paper does not provide an automated rule for deciding which additional chains or cofactors to include. Therefore, for target classes with obligate binding partners, the demonstrated pipeline is not 'without crystal structures' in the strong sense advertised: it silently relies on external structural/biological annotation. The benchmark in Figure 8 (four well-studied kinases) does not expose this dependency because the correct biological assembly for these systems is already known to the practitioner. This is the most load-bearing weakness because it directly targets the paper's headline contribution—the claimed removal of the need for experimental structural knowledge—rather than merely the numerical accuracy of a particular protocol.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Boltz-ABFE, a pipeline that combines Boltz-1/Boltz-2 protein-ligand co-folding predictions with a previously developed absolute binding free energy (ABFE) protocol, aiming to perform rigorous affinity calculations without experimental crystal structures. The pipeline uses ligand SMILES and protein sequence(s) as inputs, generates multiple Boltz models, corrects ligand chemistry by re-docking with OpenEye POSIT, prepares the complex with Spruce and BioSimSpace, and runs alchemical ABFE simulations with GROMACS and alchemlyb. The method is tested on four kinase targets from the FEP+ benchmark (TYK2, CDK2, JNK1, P38) with three replicates, reporting average MUE below 1 kcal/mol for Boltz-seeded simulations, comparable to crystal-structure-seeded results. The paper also explores Boltz-1 vs Boltz-2 failure modes, target deconvolution via classical scoring, truncation of low-confidence regions, and the importance of including binding partners (e.g., cyclin for CDK2) during co-folding.","tokens_in":14535,"tokens_out":4442,"duration_ms":49701,"significance":"If the central claim is sustained, the work would be a useful step toward extending FEP to early drug discovery stages where crystal structures are unavailable. The paper's strengths include a clearly specified pipeline built with open tools, a benchmark against experimental affinities, triplicate ABFE replicates, and explicit discussion of structure-quality failure modes. The paper also acknowledges important limitations, including a speculative side-chain-flip explanation and the need for future work on cofactors and apo-state offsets. However, the evidence base is narrow: four well-studied kinases, with one target (TYK2/Boltz-2) exceeding the 1 kcal/mol MUE threshold, and no demonstrated automation of biological-assembly choices. As a result, the headline claim that ABFE can be performed 'without crystal structures' is plausible but not yet convincingly demonstrated in the strong sense advertised.","major_comments":[{"comment":"The headline claim 'without experimental crystal structures' (Abstract, Sec. 2.1.1) is not established for systems that require biological-assembly knowledge. For CDK2, Boltz-2 alone predicts an artificially expanded pocket and an extended ligand pose; including CyclinE1 is necessary to restore the native pocket and improve ABFE (Figs. 7C and 7F). The decision to include the cyclin partner is not derived from the ligand SMILES and CDK2 sequence by any automated step in the pipeline; it is external biological knowledge typically obtained from crystallography, homology, or prior experiment. The four-target benchmark (Sec. 2.2, Fig. 8) does not expose this dependency because the correct assemblies for TYK2/CDK2/JNK1/P38 are already known. To support the central claim, the authors should either provide an automated, sequence-derived rule for selecting partner chains and validate it on less-c","section":"Sec. 2.1.3 / Fig. 7"},{"comment":"The 'MUE < 1 kcal/mol on average' claim is an average over four targets; the TYK2/Boltz-2 condition exceeds 1 kcal/mol, as acknowledged in the text. With only four targets, a single outlier represents 25% of the benchmark, and the three-replicate error bars do not propagate experimental affinity uncertainties. Without such propagation, the MUE/RMSE comparisons lack a statistically grounded uncertainty estimate. The proposed side-chain-flip explanation for the TYK2 Boltz-2 discrepancy is also speculative: no structural overlay, per-residue analysis, or alternative-model ABFE calculation is provided to causally link the flipped side chain to the computed error. The authors should report per-target errors with experimental error propagated and either test the side-chain hypothesis or clearly label it as a hypothesis requiring further study.","section":"Sec. 2.2 / Fig. 8"},{"comment":"The choice of POSIT over Hybrid re-docking is made based on ABFE results on TYK2 (SI Table S1), and TYK2 is then included in the benchmark set reported in Sec. 2.2. This is protocol selection on the test set, which can inflate the apparent performance for TYK2 and potentially for the average. To make the benchmark clean, the authors should either select the docking protocol using a separate validation set or an oracle not involving the reported targets, or explicitly state that TYK2 serves as a training target for the pipeline choice. The current presentation conflates tuning and evaluation, weakening the generalizability claim.","section":"Sec. 2.1.1 / Sec. 2.2"},{"comment":"The paper claims that the targets include structures the models have not seen before (Introduction), but no training-containment analysis is presented. The four FEP+ benchmark targets are public and well studied, and Boltz-2 is developed by the same group (Ref. 51), so the reader cannot rule out that these complexes or close homologs are in the training data. This matters because the central claim is about generalization to uncharacterized targets. The authors should provide evidence of training-set exclusion (e.g., sequence/structure similarity analysis to Boltz-2 training data, or validation on recently deposited structures such as 9OB2–9OB6). Without this, the benchmark may be optimistic.","section":"Sec. 4.1 / Intro"}],"minor_comments":[{"comment":"Typo: 'RSME' in the text should be 'RMSE'.","section":"Sec. 2.2"},{"comment":"The caption contains '(Figure 5G, G)' where the second reference appears to be a typo (likely 'Figure 5F, G').","section":"Fig. 5 caption"},{"comment":"The phrase 'MUE ′s > 1 kcal/mol' contains an awkward apostrophe; suggest 'MUE values > 1 kcal/mol'.","section":"Sec. 2.2"},{"comment":"The sentence describing Boltz-1+P performance relative to crystal-structure results is difficult to parse ('yielding a modest increase of 0.1 kcal/mol in mean unsigned error compared to the crystal structure results on average'). Please clarify whether this is an increase or decrease in error, and relative to which condition.","section":"Sec. 2.2"}],"recommendation":"major_revision","confidential_remarks":"The core idea is timely and the protocol is technically sound, but the headline claim is broader than the evidence. The binding-partner dependency (Sec. 2.1.3) is the most serious issue; it affects the paper's main contribution and may require either substantial additional automation/validation or a carefully moderated claim. The protocol-selection-on-TYK2 issue and the absence of training-overlap analysis also need attention before the work can be considered for publication in a leading venue. The paper may be suitable after major revision, but the current version overstates the 'without crystal structures' capability."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Jake — quick read of Boltz-ABFE. The real content is a pipeline: Boltz-2 co-folding, POSIT redocking, and the in-house ABFE protocol, tested on four FEP+ kinases. That combination is new, and comparing predicted-structure results against crystal-initiated ABFE is the right experiment. The low-confidence truncation analysis and the binding-partner inclusion study are genuinely useful observations. Credit where due: the benchmark is externally validated against experimental affinities, three replicates, and the numbers mostly land near 1 kcal/mol MUE. That is a plausible feasibility result.\n\nThe softest spot is the headline claim. Section 2.1.3 shows that for CDK2, Boltz-2 alone gives an expanded pocket and an extended ligand; including CyclinE1 fixes it and materially improves ABFE. But choosing to include that partner is biological knowledge — the kind you normally get from crystal structures, homology, or prior experiment. The paper says the pipeline starts from SMILES and protein sequence alone, but the CDK2 workflow silently uses extra structural annotation. That doesn't kill the pipeline, but it does mean 'without crystal structures' is not demonstrated in the strong sense. The accurate bound is more like 'without a crystal structure of the specific ligand–target complex, provided you already know the biological assembly.'\n\nOther concerns are minor by comparison. Four well-studied kinases is a small, likely training-overlapping set. TYK2 with Boltz-2 exceeds 1 kcal/mol. Experimental errors aren't propagated, and the sidechain-flip explanation is speculative. No code or data shipped, which makes reproducibility harder to judge.\n\nWho is this for: computational drug discovery people who want to know whether predicted structures can seed FEP. It's a useful proof-of-concept, not a definitive method. I'd send it to peer review: the question is important, the benchmark is the right kind of evidence, and the weaknesses are addressable in revision. The abstract should be softened, and the automation gap for binding-partner selection should be discussed explicitly. If a referee pushes on that, the authors can either provide an automated rule or narrow the claim.","headline":"Useful pipeline demonstration, but the 'without crystal structures' headline overreaches: for CDK2 the pipeline needs the cyclin partner, which is structural knowledge the paper doesn't automate.","tokens_in":14988,"tokens_out":2255,"would_cite":true,"duration_ms":20474,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that absolute binding free-energy calculations can run from predicted protein-ligand structures, with mean signed errors under 1 kcal/mol across four kinase targets, removing the need for experimental crystal structures in","keywords":["absolute binding free energy","Boltz-2","co-folding structure prediction","free energy perturbation","protein-ligand docking","re-docking","drug discovery","molecular dynamics"],"falsifier":"Run the pipeline on a new protein target with no experimental complex and no known binding partner, using the untailored full UniProt sequence; if the mean unsigned error against measured affinities exceeds 1 kcal/mol, or if ABFE values diverge sharply when a binding partner is later added, the claim that crystal-free predictions are generally sufficient would be refuted.","tokens_in":14166,"feed_emoji":"💊","tokens_out":6782,"duration_ms":69165,"temperature":0.7,"pith_summary":"The paper tries to establish that rigorous absolute binding free energy (ABFE) calculations—the gold-standard physics-based way to estimate how tightly a small molecule binds a protein—can start from computer-predicted protein-ligand structures instead of experimental crystal structures. It builds a pipeline that uses the Boltz-2 co-folding model to predict the complex from sequence and SMILES, repairs common prediction defects such as steric clashes, wrong bond orders, aromaticity, and stereochemistry via multi-model sampling and re-docking, then runs alchemical free energy simulations. On four kinase targets from a public benchmark, ABFE values initiated from predicted structures achieved mean unsigned errors below 1 kcal/mol, essentially matching simulations started from crystal structures. If this holds, free energy perturbation can be applied earlier in drug discovery, before crystals exist, expanding its domain of applicability.","feed_headline":"No crystal needed: predicted structures reach sub-1 kcal/mol accuracy","feed_subtitle":"Boltz-2 co-folding plus re-docking yields absolute binding free energies accurate enough for early drug discovery.","key_machinery":"The load-bearing mechanism is the coupling of a co-folding structure predictor (Boltz-2) with an alchemical absolute binding free energy (ABFE) protocol, mediated by structure repair. ABFE is a simulation in which the ligand is annihilated in the bound and unbound states and the free energy difference yields the binding affinity; the repair step—re-docking the ligand into the predicted pocket with POSIT, a shape- and pharmacophore-guided docking method—corrects ligand chemistry errors (bond orders, aromaticity, stereochemistry) that would otherwise poison the simulation. Supporting machinery includes sampling multiple Boltz models per complex to escape steric clashes, truncating low-confiden","core_discovery":"On the paper's own terms, the central claim, stated in Section 2.2, is that 'the ABFE results initiated from any of the Boltz predicted structures achieved satisfactory results with MUE < 1 kcal/mol on average.' The authors interpret this as demonstrating the feasibility of absolute FEP simulations without experimental crystal structures. The demonstration rests on a preparation step: Boltz predictions contain systematic chemical errors, and re-docking with the template-guided POSIT method fixes them; truncating low-confidence sequence regions and including binding partners (e.g., cyclin for CDK2) are needed to obtain pockets that match biology. The paper also shows that a top-down trained a","pith_inferences":["If these results generalize, the next bottleneck is not the simulation but the automation of biological context: deciding which binding partners, co-factors, or modified termini to feed into the structure predictor. The paper's CDK2 case suggests this is currently a manual choice; an automated rule would be needed for a truly crystal-free workflow.","The same pipeline could be pointed at hit identification rather than lead optimization, where absolute values across different proteins matter; the paper notes the offset problem but does not solve it.","A broader implication is a new evaluation loop: predicted complexes can be scored by how well they support ABFE convergence, turning affinity measurements into structural-validation data.","The competitive performance of the top-down affinity module on kinases hints that purely learned affinity models may dominate on well-represented protein families, while physics-based ABFE should be favored on novel targets—an easily testable trade-off across target families."],"forward_implications":["ABFE simulations can reach gold-standard accuracy (MUE < 1 kcal/mol) without an experimental complex, at least for well-behaved kinase targets.","The re-docking correction is what makes Boltz-1 predictions usable; Boltz-2's inference-time steering removes most clashes and chemistry errors but still leaves stereochemistry errors, so re-docking remains valuable.","For targets with flexible or partner-dependent binding sites, including binding partners in the co-folding input is essential; omitting them can create artificial pockets and extended ligand poses that degrade affinity predictions.","Classical scoring of co-folded poses can filter wrong targets when targets are structurally distant (AUC up to 0.91), but cannot resolve selectivity among similar off-targets (AUC about 0.62), so affinity simulations or better methods are needed for target deconvolution.","The sensitivity of ABFE to the starting structure offers an affinity-based benchmark for co-folding models that does not require crystal structures."],"supporting_citations":[{"why":"Supplies the Boltz-2 co-folding model with inference-time steering, the main generator of predicted protein-ligand complexes used in the pipeline.","marker":"[51]"},{"why":"Supplies the Boltz-1 model, the earlier co-folding predictor whose defects motivate the re-docking repair step.","marker":"[50]"},{"why":"Describes the authors' absolute binding free energy simulation protocol, the core engine for computing binding affinities from predicted structures.","marker":"[55]"},{"why":"Provides the POSIT template-guided re-docking method used to correct ligand bond-order, aromaticity, and stereochemistry errors.","marker":"[56]"},{"why":"Provides the FEP+ benchmark set and experimental affinities used to validate the pipeline and to define the sub-1 kcal/mol accuracy target.","marker":"[35]"},{"why":"Prior work combining predicted complexes with relative free energy calculations, which this paper extends to absolute FEP on a larger target set.","marker":"[53]"}],"fun_headline_variants":["Predicted structures enable sub-1 kcal/mol free energies","No crystal? Boltz-ABFE still hits 1 kcal/mol accuracy","Boltz-2 plus re-docking: FEP without crystal structures","Crystal-free FEP: predicted complexes match experiment","Boltz-ABFE: accurate binding without experimental structures"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the biological context needed for an accurate co-folding prediction—which binding partners or sequence regions to include—can be determined without consulting experimental structures; the CDK2 case shows this is currently a manual choice.","fun_headline_variants_meta":{"raw":{"variants":["Predicted structures enable sub-1 kcal/mol free energies","No crystal? Boltz-ABFE still hits 1 kcal/mol accuracy","Boltz-2 plus re-docking: FEP without crystal structures","Crystal-free FEP: predicted complexes match experiment","Boltz-ABFE: accurate binding without experimental structures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000171,"raw_usage":{"total_tokens":1113,"prompt_tokens":756,"completion_tokens":357,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":270}},"tokens_in":500,"tokens_out":357,"duration_ms":3974,"temperature":1.0,"reasoning_tokens":270,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:46:47.930487+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the pipeline on a new protein target with no experimental complex and no known binding partner, using the untailored full UniProt sequence; if the mean unsigned error against measured affinities exceeds 1 kcal/mol, or if ABFE values diverge sharply when a binding partner is later added, the claim that crystal-free predictions are generally sufficient would be refuted.","supporting_citations":[{"cited_title":"BioRxiv 2025, 2025--06","cited_arxiv_id":null,"evidence_quote":"Supplies the Boltz-2 co-folding model with inference-time steering, the main generator of predicted protein-ligand complexes used in the pipeline."},{"cited_title":"bioRxiv 2024,","cited_arxiv_id":null,"evidence_quote":"Supplies the Boltz-1 model, the earlier co-folding predictor whose defects motivate the re-docking repair step."},{"cited_title":"Optimizing Absolute Binding Free Energy Calculations for Production Usage","cited_arxiv_id":null,"evidence_quote":"Describes the authors' absolute binding free energy simulation protocol, the core engine for computing binding affinities from predicted structures."},{"cited_title":"P.; Brown, S","cited_arxiv_id":null,"evidence_quote":"Provides the POSIT template-guided re-docking method used to correct ligand bond-order, aromaticity, and stereochemistry errors."},{"cited_title":"The maximal and current accuracy of rigorous protein-ligand binding free energy calculations","cited_arxiv_id":null,"evidence_quote":"Provides the FEP+ benchmark set and experimental affinities used to validate the pipeline and to define the sub-1 kcal/mol accuracy target."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior work combining predicted complexes with relative free energy calculations, which this paper extends to absolute FEP on a larger target set."}],"review_version":1}