{"id":"ab4f35a4-f51b-4990-8e40-90a601266a0c","arxiv_id":"2411.15618","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"A two-stage equivariant neural network predicts hydration site locations and entropy/enthalpy profiles from static protein structures, trained on 4,148 explicit-water MD simulations.","lead":"This paper trains a geometric deep learning model on thousands of molecular dynamics simulations to predict where water molecules sit on protein surfaces and how tightly they bind. It targets fast, one-shot hydration site analysis for drug design, replacing costly simulation-based methods.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'near dynamics-level accuracy' claim is evaluated only against the same 20 ns restrained WATsite/Desmond labels used for training; the sole external x-ray check (Table 7) shows PHR 6-11%, so the model may be an emulator of one MD pipeline rather than a general hydration-site predictor.","rationale":"The reader's verdict is conditional largely because the ground truth is the authors' own MD pipeline. I agree that this is the most load-bearing risk. I considered two other candidates: the MUP ligand correlation and the thermodynamic profiling conditional on true coordinates. The MUP footnote about linear regression does not by itself inflate the reported Pearson r, since an affine transform leaves r unchanged, but the small sample size and the tolerance choice at 2.4 A still limit it. The thermodynamic section is evaluated only given true hydration-site coordinates, so it does not test the end-to-end pipeline; nevertheless, the location error is separately characterized in Tables 1-4. The decisive issue is external validity: the model agrees with WATsite, while the single external comparison to crystallographic waters shows a much lower PHR. The proposed 100 ns unrestrained repeat simulation directly tests whether the WATsite labels are converged and thus whether the model is a faithful surrogate or an emulator of one approximate protocol. This does not change the reader's conditional verdict; it sharpens the specific condition that must be met.","tokens_in":15191,"tokens_out":11203,"duration_ms":116104,"concrete_test":"Select 50-100 proteins at random from the test set, rerun the full WATsite analysis on 100 ns unrestrained Desmond simulations in three independent replicates, and recompute the high-occupancy (>=0.5) hydration-site list. Measure the fraction of original 20 ns restrained sites whose cluster centers move by more than 0.5 A or disappear, and compute GTRR/PHR of the trained model against the converged replicate labels. If more than ~30% of sites shift or vanish, the original labels are not a stable reference and the reported 'near dynamics-level accuracy' is an artifact of the training pipeline; if the model's PHR against the new labels stays near the Table 1 values (65.9% at 1.0 A), the concern is substantially weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 generates every training target and every test-set reference with a single protocol: 20 ns Desmond simulations with 50 kcal/mol/A^2 heavy-atom restraints, followed by WATsite clustering and WATsite's entropy estimate from Eq. 7. Tables 1-4 therefore measure how well the network imitates WATsite on sequence-distant proteins, not how close the predictions are to converged explicit-water thermodynamics. The only external check, Appendix B.1 Table 7, reports a crystallographic-water PHR of 6.26% at 0.5 A and 11.5% at 1.0 A, versus 48.3% and 65.9% against WATsite at the same cutoffs. The authors do not explain this large gap in the main text; if crystallographic water incompleteness is the reason, the PHR metric should be recomputed on ordered first-shell waters, and if the gap persists the 'high fidelity' experimental claim is unsupported. The thermodynamic profiling section inherits the same limitation because the Table 5 correlations are against WATsite enthalpy and entropy values whose absolute accuracy is not independently established; the MUP case study is a 12-complex, post hoc-toleranced demonstration rather than a substitute for that validation. Because the central claim depends on the training labels being a trustworthy ground truth, and because the one external comparison points the other way, this is the load-bearing assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a two-stage deep learning pipeline for protein hydration site analysis. In the first stage, an equivariant graph neural network is trained to predict hydration site coordinates from a static protein structure, with a loss based on a Gaussian mixture surrogate for KL divergence. In the second stage, a graph attention network predicts per-site enthalpy and entropy values, conditioned on hydration site coordinates. Training and test labels are generated by a single WATsite/Desmond explicit-water MD protocol applied to 4,148 protein structures, split by sequence similarity at 35% identity. The authors report ground-truth recovery rates and prediction hit rates on the held-out test set, layer- and occupancy-stratified results, thermodynamic correlations against WATsite references, several qualitative case studies (DsbA, TIM, HSP90, Clarin-2, MUP), and a MUP ligand-binding study correlating predicted desolvation free energies with experimental affinities. The central claim is that the model achieves near dynamics-level hydration site localization and thermodynamic profiling in a one-shot, fixed-time manner.","tokens_in":15544,"tokens_out":3273,"duration_ms":31239,"significance":"If the central claims hold, this work would be a substantial practical advance: it would replace expensive, multi-hour explicit-water MD simulations for hydration site localization and thermodynamic profiling with a seconds-scale deep learning model, and it would enable hydration analysis for AlphaFold-predicted structures and large conformational ensembles. The paper has concrete strengths: a large training dataset with a sequence-similarity-based held-out split, publicly available code and data, clearly specified evaluation metrics (GTRR/PHR), and several qualitative case studies probing generalization to mutations, conformational change, and membrane proteins. However, the load-bearing evidence for the 'near dynamics-level accuracy' and 'high fidelity' claims is incomplete. The only external experimental comparison, x-ray crystallographic waters in Appendix B.1, shows markedly lower prediction hit rates than the WATsite-based evaluation, and this discrepancy is not discussed in the main text.","major_comments":[{"comment":"The x-ray validation reports a prediction hit rate of 6.26% at 0.5 Å and 11.5% at 1.0 Å against crystallographic waters, whereas the WATsite-based test set (Table 1) reports 48.3% and 65.9% at the same cutoffs. The main text (Sections 1 and 4.1, and the Conclusion) claims that the model 'can reproduce results from molecular dynamics and experimentally resolved structures with high fidelity,' but the large gap between the WATsite and x-ray PHR values is not discussed anywhere in the main text. If incompleteness of deposited crystal waters is the explanation, the PHR should be recomputed on a curated set of ordered first-shell waters; if the gap persists, the experimental support for the central claim is much weaker than stated. This needs to be addressed explicitly.","section":"Appendix B.1, Table 7"},{"comment":"The thermodynamic profiling evaluation in Table 5 is explicitly performed 'given the coordinates of the true hydration sites' (Section 4.2). In actual use, the thermodynamic model consumes predicted coordinates, so coordinate errors can propagate into enthalpy and entropy errors. The only evaluation on predicted coordinates is the MUP case study (Section 4.3.5), which uses 12 complexes, a displacement tolerance of 2.4 Å selected post hoc, and 'transformed predictions obtained by a linear regression of the experimental values on the model predictions' (footnote to Table 6). The reported Pearson R of 0.930 is therefore not a direct measure of the untransformed model's predictive correlation, and the tolerance choice is not justified. Please report thermodynamic accuracy on the held-out test set using predicted hydration sites with a fixed, pre-specified matching protocol, and report both raw and calibrated correlations for the MUP study.","section":"Section 4.2 and Section 4.3.5"},{"comment":"All training targets and all test references are generated by a single computational protocol: 20 ns Desmond simulations with 50 kcal/mol/Å² heavy-atom restraints, WATsite clustering, and the entropy estimate in Eq. (7) based on an external-mode probability density. Because the model is trained and evaluated on labels from this same protocol, Tables 1–5 establish that the network imitates WATsite/Desmond output on sequence-distant proteins; they do not independently establish agreement with converged explicit-water thermodynamics or with experiment. This is a limitation of the evidence, not a circularity of the held-out split, but the paper should state it plainly and, if possible, include a small convergence or force-field sensitivity check (for example, longer simulations or a second water model on a subset of proteins) to bound the magnitude of protocol-dependent bias.","section":"Section 3.3 and Eq. (7)"},{"comment":"The loss L1 in Eq. (5) is described as 'a simplified surrogate for the symmetrized Kullback-Leibler divergence KL(p|q) + KL(q|p),' but as written it evaluates the mixture densities q and p only at the component centers (q(u_j) and p(x_j)) rather than integrating over the mixture components, so it is not a standard symmetrized KL between the two Gaussian mixtures. The paper does not provide a derivation of this surrogate, nor any sensitivity analysis for the Gaussian width σ = 0.5 or the weight penalty α in Eq. (6). These hyperparameters directly shape the predicted spatial distribution, and the reader cannot judge whether the reported accuracies are robust to reasonable variations. Please clarify the relationship to the true symmetrized KL and provide an ablation or sensitivity analysis.","section":"Section 3.1.1, Eq. (5)"}],"minor_comments":[{"comment":"There are several typographical errors, including 'it's' for 'its' in the Abstract and Section 1, and 'signficantly' for 'significantly' in Section 4.1. These should be corrected.","section":"Throughout"},{"comment":"The PDB ID column lists both '1IO6' and '1I06' for different ligands (SBT, PT, IPT, ET, MT); this appears to be a typographical inconsistency, and the correct identifiers should be verified.","section":"Table 6"},{"comment":"The last row of Table 13 repeats 'r = 1.5' instead of 'r = 2.0', which is presumably a typo. Please correct the cutoff labels.","section":"Appendix B.2, Table 13"},{"comment":"Reference [35] contains malformed author names ('Xiaohan Kuang, Zhaoqian Su Su, ... Jesse , Tyler Derr') and should be cleaned up. Reference [44] lacks a title; the full Bowers et al. SC2006 paper title should be provided.","section":"References"},{"comment":"The case study claims that all hydration sites were identified 'within 1.0 Å' (DsbA) or that '20/21 crystal waters predicted within 1.0 Å' (TIM), but the criteria for matching crystal waters to predictions (e.g., any-vs-unique matching, handling of absent waters) are not specified. Please state the matching rule used in these qualitative comparisons.","section":"Section 4.3.1 and Section 4.3.3"},{"comment":"The phrase 'near dynamics-level accuracy' is defined only through the WATsite-based metrics; given the x-ray results in Table 7, the wording should be qualified (e.g., 'near WATsite/Desmond-level accuracy') or the experimental validation should be strengthened.","section":"Abstract and Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely within the scope of a computational chemistry/biology journal, and the dataset and code releases are commendable. The main risk is that the central 'experimental validation' claim is not supported by the reported x-ray PHR values, and the thermodynamic profiling is validated only against the training protocol's own labels. The authors can address this with additional experiments (e.g., curated crystal-water evaluation, predicted-coordinate thermodynamic evaluation, and a raw-vs-calibrated MUP analysis) or by substantially tempering the claims. I would not reject, but the revision needs to be substantive, not just textual."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good paper to know about if you work on hydration site prediction. The genuinely new thing is scale: an SE(3)-transformer location model plus a GAT thermodynamic head, trained on WATsite/Desmond results for 4,148 proteins, with code and data released. Held-out performance on sequence-distant proteins is respectable for first-shell sites (GTRR about 63% at 0.5 A and 85% at 1.0 A), and the case studies (DsbA mutants, TIM, HSP90 conformers, an AlphaFold3 membrane protein) are believable illustrations of what the model can do. The thermodynamic correlations (0.86 entropy, 0.84 enthalpy) are decent given true coordinates.\n\nThe soft spot is not small. Every training label and every test reference comes from one pipeline: 20 ns Desmond runs with 50 kcal/mol/A^2 heavy-atom restraints, then WATsite clustering and WATsite's entropy estimate. So Tables 1-5 measure how well the network imitates WATsite, not how close it is to converged explicit-water thermodynamics or to experiment. The only external check is Appendix B.1, Table 7: against crystallographic waters, PHR drops from 48.3% to 6.26% at 0.5 A and from 65.9% to 11.5% at 1.0 A. That gap is not addressed in the main text, and for an abstract that says 'confirmed accuracy on experimental data,' that omission matters. It could be partly due to known incompleteness of crystal waters, but the paper doesn't make that argument, and the gap is an order of magnitude.\n\nThe MUP case study is a 12-ligand demonstration with a 2.4 A displacement tolerance and a footnote saying the energies are \"transformed predictions obtained by a linear regression of the experimental values on the model predictions.\" That is in-sample calibration, so the R=0.93 is weaker evidence than it appears. Also, the thermodynamic model is evaluated only given true hydration site coordinates; performance with predicted coordinates is shown only anecdotally.\n\nNone of this kills the contribution. A fast, open surrogate for the WATsite pipeline is useful in its own right, and the paper is honest enough to put the x-ray table in the appendix even if it doesn't face the implication. What I'd want before treating it as a drop-in replacement for dynamics-based profiling is a proper external validation, for example on ordered crystal waters or on independent MD from a different protocol, and some discussion of why the aggregate x-ray numbers are so much lower than the WATsite numbers.\n\nRecommendation: send it to peer review, but with a strong request to address the x-ray gap and the \"given true coordinates\" evaluation. This deserves referee time, not a desk reject.","headline":"Large-scale ML surrogate for MD-based hydration site profiling, with real code and data, but the headline accuracy claim rests on emulating one MD pipeline and the single external x-ray check is far weaker.","tokens_in":16071,"tokens_out":1840,"would_cite":true,"duration_ms":16312,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents a geometric deep network that predicts hydration site locations and thermodynamic profiles from a static protein structure in one shot, at accuracy close to explicit-water molecular dynamics.","keywords":["hydration site prediction","thermodynamic profiling","equivariant graph neural network","molecular dynamics","water networks","protein-ligand binding","desolvation free energy","structure-based drug design"],"falsifier":"Run the trained model on a set of proteins with high-resolution neutron crystallography or multi-force-field consensus hydration sites, and measure the ground-truth recovery rate. If recovery at 1.0 Å falls to the level of the appendix's x-ray prediction hit rate (11.5%) rather than the MD-based 80.2%, the claim that the model achieves near-dynamics-level accuracy on real experimental waters is not supported.","tokens_in":14993,"feed_emoji":"💧","tokens_out":4471,"duration_ms":43504,"temperature":0.7,"pith_summary":"The paper claims that a geometric deep neural network can replace expensive molecular-dynamics simulations for finding where water sits on a protein surface and how tightly it is held. Trained on hydration sites computed from thousands of explicit-water simulations, the model predicts site coordinates in a single forward pass, then assigns each site an enthalpy and entropy. On a sequence-disjoint test set it recovers 80% of true hydration sites within 1 Å and reaches correlations of 0.86 and 0.84 for entropy and enthalpy. If this holds, structure-based drug design can screen hydration thermodynamics across many protein conformations and predicted structures in seconds rather than hours.","feed_headline":"Neural network maps protein hydration sites in seconds","feed_subtitle":"Trained on thousands of MD simulations, it recovers 80% of sites within 1 Å and scores the energy cost of displacing each water.","key_machinery":"A two-stage equivariant graph neural network. The first stage seeds water nodes at every atom with solvent-accessible surface area above 0.1, then applies five SE(3)-equivariant attention layers with distance-based graph updates so that predicted water positions rotate and translate with the protein while protein atoms stay fixed; it is trained with a Gaussian-mixture Kullback-Leibler loss against WATsite hydration sites. The second stage builds a graph connecting protein atoms and hydration sites within 8 Å, applies graph attention layers, and regresses per-site enthalpy and entropy. The equivariant updates are what let the network represent water-water and water-protein interactions without imposing a fixed grid.","core_discovery":"The central claim is that hydration site localization and thermodynamic profiling, normally requiring lengthy explicit-water molecular dynamics, can be done in one shot by an equivariant graph neural network with near-dynamics-level accuracy. The model first places water nodes at solvent-exposed protein atoms and refines their positions through five equivariant attention layers, learning multi-body water networks that respect the protein surface. A second graph attention network then predicts the enthalpy and entropy of each site, and the predicted water displacement free energies correlate with measured binding affinities in the MUP ligand series (Pearson R = 0.930). The authors argue the model is fast, robust to point mutations and conformational changes, and generalizes to unseen protein sequences.","pith_inferences":["The reported accuracy is dominated by first-shell waters: second-shell sites are recovered at only 26.9% within 1 Å, so practical applications should treat second-shell predictions as candidates for refinement rather than final answers.","Because the training labels come from one MD force field and WATsite's occupancy and entropy approximations, absolute thermodynamic values may shift if labels are regenerated with a different water model, even if rankings are stable.","The same two-stage architecture could be retrained on hydration data for protein-ligand complexes or protein-protein interfaces, extending the method beyond apo protein surfaces toward biologics and co-solvent mapping.","The appendix's x-ray comparison shows much lower hit rates than the MD-based evaluation, suggesting that a sharper experimental validation against high-resolution neutron or curated crystallographic waters would be a more demanding test of the model's real-world accuracy."],"forward_implications":["Hydration site positions and thermodynamic profiles for a single protein structure are produced on the seconds timescale, making high-throughput scanning across many conformations practical.","The model can attach physically reasonable water networks to predicted protein structures from structure-prediction models, which normally omit water, opening hydration analysis for uncharacterized proteins.","Displacing predicted hydration sites yields desolvation free energies that correlate with experimental ligand binding affinities in the MUP case study, suggesting a fast scoring component for lead optimization.","The model responds to point mutations and conformational rearrangements without retraining, so it can be applied to mutant analysis and to conformational ensembles.","The architecture could be integrated into deep learning pipelines for co-folding, dynamics, or free-energy estimation that currently ignore explicit water."],"supporting_citations":[{"why":"Supplies the WATsite program whose occupancy and thermodynamic predictions are the training ground truth for the model.","marker":"[28]"},{"why":"Extends WATsite for enclosed binding sites, contributing to the protocol that generated the hydration site labels.","marker":"[39]"},{"why":"Provides the SE(3)-transformer architecture that the location prediction model is inspired by.","marker":"[36]"},{"why":"Provides the graph attention layers used in the enthalpy and entropy prediction model.","marker":"[38]"},{"why":"Used to split proteins by sequence similarity so the test set remains sequence-disjoint from training.","marker":"[42]"},{"why":"Describes the Desmond simulation protocol that produced the explicit-water trajectories from which hydration sites were calculated.","marker":"[44]"},{"why":"Source of the large set of protein structures from which the non-redundant training dataset was selected.","marker":"[40]"}],"fun_headline_variants":["One-shot hydration site mapping with graph neural nets","Deep learning predicts water positions and energetics on proteins","Graph neural network finds and scores protein hydration sites","Fast water mapping on proteins via geometric deep learning","Predicting protein hydration energy landscapes in seconds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire pipeline inherits whatever bias sits in the MD-based WATsite labels: if the simulated hydration site locations, occupancies, or entropy estimates are wrong, the model's predictions are wrong in the same way, and the appendix's lower x-ray hit rates indicate that this mismatch with experiment is real.","fun_headline_variants_meta":{"raw":{"variants":["One-shot hydration site mapping with graph neural nets","Deep learning predicts water positions and energetics on proteins","Graph neural network finds and scores protein hydration sites","Fast water mapping on proteins via geometric deep learning","Predicting protein hydration energy landscapes in seconds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000354,"raw_usage":{"total_tokens":1864,"prompt_tokens":825,"completion_tokens":1039,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":968}},"tokens_in":441,"tokens_out":1039,"duration_ms":9117,"temperature":1.0,"reasoning_tokens":968,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:06:12.991122+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained model on a set of proteins with high-resolution neutron crystallography or multi-force-field consensus hydration sites, and measure the ground-truth recovery rate. If recovery at 1.0 Å falls to the level of the appendix's x-ray prediction hit rate (11.5%) rather than the MD-based 80.2%, the claim that the model achieves near-dynamics-level accuracy on real experimental waters is not supported.","supporting_citations":[{"cited_title":"Watsite: Hydration site prediction program with pymol interface, 2014","cited_arxiv_id":null,"evidence_quote":"Supplies the WATsite program whose occupancy and thermodynamic predictions are the training ground truth for the model."},{"cited_title":"Efficient and accurate hydration site profiling for enclosed binding sites","cited_arxiv_id":null,"evidence_quote":"Extends WATsite for enclosed binding sites, contributing to the protocol that generated the hydration site labels."},{"cited_title":"Prediction of molecular field points using se(3)-transformer model","cited_arxiv_id":null,"evidence_quote":"Provides the SE(3)-transformer architecture that the location prediction model is inspired by."},{"cited_title":"Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets","cited_arxiv_id":null,"evidence_quote":"Used to split proteins by sequence similarity so the test set remains sequence-disjoint from training."},{"cited_title":"Scalable algorithms for molecular dynam- ics simulations on commodity clusters","cited_arxiv_id":null,"evidence_quote":"Describes the Desmond simulation protocol that produced the explicit-water trajectories from which hydration sites were calculated."},{"cited_title":"The pdbbind database: method- ologies and updates","cited_arxiv_id":null,"evidence_quote":"Source of the large set of protein structures from which the non-redundant training dataset was selected."}],"review_version":1}