{"id":"2f7d60e0-71cf-46ef-8970-6e6c1611b153","arxiv_id":"2505.23446","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 44-model molecular dynamics benchmark shows four-site TIP4P-type potentials, led by TIP4P/2005, reproduce experimental water structure best across 254 to 366 K.","lead":"This paper ran a common molecular dynamics protocol with 44 classical water models and compared the simulated X-ray and neutron scattering patterns against diffraction experiments from 254 to 366 K. It finds that four-site TIP4P-type models, led by TIP4P/2005, best describe liquid water structure, while several new three-site models such as OPC3 are nearly as good.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Neutron-diffraction leg of the combined ranking uses D2O reference data without isotope correction; if D2O/H2O structural differences exceed the ~10% spread that separates the top 17 models, the conclusion that TIP4P-type models are best may be an artifact of the reference choice.","rationale":"The reader identified the same weakest assumption: the neutron-diffraction leg uses D2O data as a proxy for H2O without isotope analysis or sensitivity testing. On inspection of the full text, this is indeed the most load-bearing concern. The central claim is a ranking claim: the paper concludes that four-site TIP4P-type models give the best overall agreement with experimental diffraction data, judged by Rtot, which combines XRD and ND relative R-factors. The XRD leg is internally consistent (H2O experiment compared to H2O simulation). The ND leg is not: the experimental reference is D2O, while the simulated PRDFs come from H2O models, and the weighting in Eq. (2) uses a single set of scattering lengths. The mismatch has two components: (i) the scattering-length weighting for D2O differs from H2O (b_D vs b_H, including sign changes in O-H and H-H partial weights), and (ii) the real structure of D2O is more ordered than H2O due to nuclear quantum effects. Component (ii) is physically large enough to matter: the top 17 models span less than 10% in Rtot, and D2O-vs-H2O differences in the first peak of the total structure factor are known to be on the order of several percent. The paper's own caveat that the variation among top models is 'not considered significant' makes the ranking sensitive to any systematic bias. No error bars on R-factors, no block averaging, and no jackknife over experimental data sets are provided, but the D2O proxy is a distinct and actionable bias. The concrete test proposed is feasible with data already in the literature and would settle whether the ranking survives the isotope correction. I therefore see no reason to change the reader's conditional acceptance; the concern warrants keeping the condition, not rejecting or upgrading the paper.","tokens_in":49658,"tokens_out":6652,"duration_ms":67726,"concrete_test":"At 295 K, recompute the neutron-diffraction R-factors for all 44 models using H2O experimental reference data instead of D2O data, and using hydrogen scattering lengths (b_H) in Eq. (2). Soper's 2013 paper (Ref. [65]) provides H2O total structure factors derived from H/D substitution, so the data needed for this test already exist. Then re-calculate the combined room-temperature ranking Rrel_RT (XRD, 295K + ND, 295K) and, if H2O data at other temperatures are available, recompute Rtot as well. If the identity of the top-ranked model changes, or if any of the current top-17 models moves outside the 10% spread relative to TIP4P/2005, the D2O reference is biasing the central conclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.2 states: 'Since heavy water data provides the lowest uncertainty in neutron diffraction measurements, the corresponding structure functions were used with the appropriate weights,' and Section 4.3.2 reiterates that D2O curves were compared with simulated data. In Eq. (2), the neutron weights use coherent scattering lengths; for a simulated H2O model, pairing D2O experimental data with deuterium scattering lengths (b_D) makes the comparison a D2O-weighted total structure factor of an H2O simulation. However, D2O is structurally more ordered than H2O (stronger hydrogen-bond network), so the experimental reference is systematically shifted relative to the true H2O target. The combined Rtot ranking is computed as the sum of XRD and ND relative R-factors (Section 4.3.3), and the top 17 models differ by less than 10%. A systematic isotope offset of only a few percent in the ND peak amplitudes can therefore reorder the ranking and move models in and out of the top cluster. The manuscript provides no sensitivity analysis, no isotope correction, and no estimate of how much the D2O-vs-H2O mismatch contributes to the R-factor differences. Because the central claim ('best agreement... achieved with four-site, TIP4P-type models') rests on this combined ND-based ranking, the D2O proxy is a load-bearing correctness risk, not a cosmetic issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports molecular dynamics simulations of 44 classical water models at seven temperatures (254–366 K). It computes partial radial distribution functions and total scattering structure factors, compares them against X-ray diffraction data for H2O (Skinner et al.) and neutron diffraction data for D2O (Soper; Ohtomo et al.), and introduces relative R-factors to rank the models. The central claims are that models with more than four interaction sites, and flexible or polarizable models, do not provide significant structural advantages, and that TIP4P-type four-site models, especially TIP4P/2005, give the best overall agreement, with recent three-site models (OPC3, OPTI-1T, OPTI-3T) being competitive.","tokens_in":49926,"tokens_out":6648,"duration_ms":61971,"significance":"If the central claim holds, the paper provides a useful benchmark for model selection in classical simulations of liquid water. Its strengths are a broad and uniform simulation protocol across 44 models, direct comparison with experimental TSSFs rather than derived PRDFs, cross-checks of density and self-diffusion against literature values, and full tabulation of R-factors and structural data. The work fits no new parameters: all model parameters come from prior literature, and the experimental data are external benchmarks. However, the main ranking rests on a combined XRD/ND comparison whose ND leg uses D2O reference data for H2O simulations, and the reported R-factors carry no statistical uncertainties; both issues need to be addressed before the ranking claims can be considered robust.","major_comments":[{"comment":"The neutron-diffraction leg of the comparison uses experimental D2O structure factors (Soper 283/295 K and Ohtomo et al. 298–368 K) as the reference for simulated H2O models, with deuterium scattering lengths in the weighting of Eq. (2). No isotope correction or sensitivity analysis is provided. D2O is more strongly ordered than H2O, so the reference curve is systematically shifted relative to the H2O target, and because the top 17 models in Rtot (Section 4.3.3, Table S6) are separated by less than 10%, a few-percent isotope offset could reorder the top cluster. Please either simulate D2O with the leading models to quantify the isotope effect, or add a sensitivity test that bounds the D2O/H2O contribution to the ND R-factors.","section":"Section 2.2, Eq. (2), and Section 4.3.2"},{"comment":"R-factors are reported without statistical uncertainties, yet the text states that differences among the top 17 models are 'not significant' and that any of them can be reliably used for structural analysis. Adjacent ranks in Table S6 (e.g., TIP4P/2005 with Rtot 2.50 and TIP4P/ε with 2.52) differ by less than 1%, and no estimate of the noise floor is given. The authors should add error bars (for example, block averaging over independent trajectory segments) or otherwise show that the ranking is stable under plausible simulation and experimental uncertainty. Without this, the 'no significant advantage' claim for models beyond four sites is not quantitatively supported.","section":"Section 4.3.3 and Tables 3, 4, S6"},{"comment":"The claim that 'using different ND dataset combinations from the three publications does not significantly affect model rankings' is not demonstrated. Table 4 shows that the best model changes with dataset (TIP4P-BG for the 284 K Soper data, TIP6P-Ew for the 295 K Soper data, and TIP4P/2005f for the Ohtomo data), so an explicit table of Rtot under alternative ND dataset combinations is needed to support the statement. This is load-bearing because Rtot determines the paper's central ranking.","section":"Section 4.3.3"}],"minor_comments":[{"comment":"In the first paragraph of the introduction, 'such a model that that simultaneously reflects' contains a duplicated 'that'; please correct.","section":"Section 1"},{"comment":"The TIP6P-Ew entry lists dOH = 0.98000 nm; this is an order of magnitude larger than the TIP6P value (0.09800 nm) and is almost certainly a typo for 0.098000 nm. Please verify the parameter file and correct the table, since other groups may use these parameters.","section":"Table S1"},{"comment":"The caption for the 'mix' model says it combines the intramolecular part of TIP4P/2005 with the intramolecular part of OPC, whereas Section 4.3.3 states that the mixed PRDFs use the intermolecular part of OPC and the intramolecular part of TIP4P/2005; the caption should be corrected to match the text.","section":"Figure 12 caption"},{"comment":"The term 'PRDSs' appears to be a typo for 'PRDFs'.","section":"Section 4.3.3"}],"recommendation":"major_revision","confidential_remarks":"The D2O proxy is the main risk to the central claim; a dedicated isotope sensitivity analysis is feasible and would make the ranking credible. I would also encourage the authors to deposit parameter files and analysis scripts, since the manuscript's value for future model selection depends on reproducibility of the exact simulation settings."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Takeaway: this is a competent, wide-ranging benchmark that supports the claim that for liquid water structure, TIP4P/2005 and a few other four-site models are as good as anything more complex. The comparison of 44 models over 254–366 K against both XRD and ND total structure factors is genuinely new, and the authors are careful to compare TSSFs directly rather than fitting PRDFs, which avoids a common ambiguity. The simulation protocol is sensible, densities are cross-checked, and the inclusion of recent models (OPC3, OPTI, ECCw2024, SWM4 revisions) is valuable.\n\nThe main soft spot, as the stress-test note says, is the use of D2O neutron data as the reference for simulated H2O models. They never correct for the isotope shift, and the combined Rtot ranking leans on this ND leg. I checked the text: they justify it by saying D2O data has the lowest uncertainty, but that doesn't change the fact that the reference is D2O, not H2O. D2O is slightly more ordered, so the comparison is systematically biased. That said, I think the stress-test overstates the risk. The top 17 models are separated by less than 10% in relative R-factors, and the isotope shift in the total structure factor is likely a few percent at most. It could reorder the top few, but it is unlikely to change the broad conclusion that four-site TIP4P-type models are sufficient and that polarizable or flexible models do not help. Still, an isotope correction or a sensitivity analysis is needed before the specific ranking can be trusted.\n\nOther minor issues: R-factors are reported without error bars, which matters when the spread is small; there's a likely typo in the supplementary table for TIP6P-Ew dOH (0.98 nm instead of 0.098 nm); the self-diffusion values are explicitly approximate and not corrected for finite size, which is fine but should be kept in mind. The citation pattern looks appropriate, and prior work is acknowledged.\n\nWho is this for? Anyone choosing a water model for MD who cares about structural accuracy, and model developers who want to see how recent parameterizations stack up. It deserves a serious referee; with the isotope issue addressed, it would be a solid publication. I recommend sending it for peer review, with the D2O concern raised as the main revision request.","headline":"A useful, well-run benchmark of 44 water models; the D2O-for-H2O neutron reference is the main caveat but does not sink the central conclusion.","tokens_in":50491,"tokens_out":3374,"would_cite":true,"duration_ms":33543,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that a simple rigid four-site model of water, TIP4P/2005, reproduces experimental liquid structure more accurately than more complex flexible, polarizable, or five-to-seven-site models over the full 254–366 K range.","keywords":["water models","molecular dynamics","radial distribution functions","total scattering structure factors","neutron diffraction","X-ray diffraction","TIP4P/2005","structure prediction"],"falsifier":"Recompute the combined ranking using experimental light-water neutron total scattering structure factors, such as those from H/D isotopic substitution measurements, instead of heavy-water data at 295 K; if the top-17 ordering changes by more than the current 10% spread, the paper's conclusion about four-site dominance would be overturned.","tokens_in":49444,"feed_emoji":"💧","tokens_out":5183,"duration_ms":55152,"temperature":0.7,"pith_summary":"This paper asks which of 44 classical water models best predicts the atomic-scale structure of liquid water, judged by matching measured neutron and X-ray scattering structure factors. Across temperatures from 254 K to 366 K, the best overall agreement comes from four-site TIP4P-type models, with TIP4P/2005 first in the combined ranking. Models with more interaction sites, flexibility, or polarizability did not improve structural accuracy despite higher computational cost. Recent three-site models nearly close the gap, but the simplest rigid four-site parameterization already appears sufficient for structure. A reader should care because structure prediction is foundational to molecular simulation, and the result suggests expensive model complexity is not needed for this property.","feed_headline":"Four-site water model beats costlier rivals on structure","feed_subtitle":"A 44-model simulation sweep finds TIP4P/2005 matches neutron and X-ray data best from 254 K to 366 K.","key_machinery":"The load-bearing comparison is the total scattering structure factor $S(Q)$, computed from simulated partial radial distribution functions $g_{ij}(r)$ by a weighted Fourier transform, combined with the goodness-of-fit measure $R$ normalized per data set to a relative R-factor $R_{rel} = R/R_{best}$ and summed into $R_{tot}$. This combined metric is what lifts four-site models to the top: because the X-ray and neutron weights emphasize different partials, fitting both data types simultaneously is a stricter test than fitting either alone.","core_discovery":"Forty-four classical pairwise-additive water models were simulated under an identical protocol; trajectories produced partial radial distribution functions, and X-ray and neutron weighted total scattering structure factors were compared with experimental data. The paper's central conclusion is that on the combined relative R-factor $R_{tot}$ (average X-ray relative R-factor plus average neutron relative R-factor), TIP4P/2005 ranks highest over the full temperature range, with TIP4P/ε, TIP4Q, TIP4P/2005f, TIP4P-FB, and other four-site models in close succession. The top 17 models span less than 10% in $R_{tot}$, so they are statistically comparable to one another. More complex models, including five-, six-, and seven-site, flexible, polarizable, and Buckingham-potential models, do not show a significant structural advantage; the worst performers include several polarizable models and two models with poor density. The paper further shows that the OPC model, though best for X-ray data alone, fails neutron data largely because its intramolecular geometry differs from gas-phase water geometry.","pith_inferences":["If heavy-water and light-water structures differ by more than the roughly 10% spread separating the top 17 models, the neutron leg of the ranking could shift; using light-water neutron data from H/D isotopic substitution would test this directly.","The temperature-shift behavior noted for TIP3P and OPC3 suggests that part of the apparent model error is a shifted temperature scale, so aligning models by effective temperature could change rankings at the edges of the 254–366 K window.","The same $R_{tot}$ protocol could benchmark machine-learned and other advanced water potentials against the same experimental data, quantifying whether their added cost buys structural improvement.","The paper's conclusion is property-specific: for thermodynamic, dynamic, or other non-structural properties, more complex models may still be needed."],"forward_implications":["Four-site rigid models are sufficient for pure-liquid-water structure, so simulations needing structural accuracy can use TIP4P/2005 without paying the computational cost of polarizability or flexibility.","The spread of less than 10% among the top 17 models means many cheap models give statistically indistinguishable structural fits, allowing cost to guide selection within that group.","Recent three-site models such as OPC3 and OPTI-3T are nearly as accurate as the best four-site models, providing an even cheaper alternative for large-scale simulations.","Model developers should validate against both neutron and X-ray structure factors rather than only the oxygen-oxygen distance, because the OPC case shows that intramolecular geometry can spoil neutron agreement.","For pure water at ambient pressure, adding interaction sites or polarizable terms is not justified by structural accuracy alone."],"supporting_citations":[{"why":"Defines TIP4P/2005, the model that ranks first in the combined R-factor metric.","marker":"[45]"},{"why":"Supplies the experimental X-ray total scattering structure factors used for the XRD leg from 254 to 366 K.","marker":"[101]"},{"why":"Supplies the room-temperature heavy-water neutron structure factor used as the primary ND reference.","marker":"[65]"},{"why":"Supplies the 283 K heavy-water neutron data and shows a potential fitted to neutron diffraction data.","marker":"[69]"},{"why":"Supplies the temperature-dependent heavy-water neutron structure factors used for the higher-temperature ND comparisons.","marker":"[102]"}],"fun_headline_variants":["TIP4P/2005 tops 44 water models on structure","Simple four-site water model beats complex rivals","Water structure best with TIP4P/2005, sweep of 44 models finds","Complex water models fail to beat TIP4P/2005 on structure","OPC wins X-ray but loses neutron; TIP4P/2005 overall best"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The neutron-diffraction leg of the ranking uses heavy-water (D2O) experimental data as a stand-in for simulated light-water (H2O) models, without an isotope correction or sensitivity analysis.","fun_headline_variants_meta":{"raw":{"variants":["TIP4P/2005 tops 44 water models on structure","Simple four-site water model beats complex rivals","Water structure best with TIP4P/2005, sweep of 44 models finds","Complex water models fail to beat TIP4P/2005 on structure","OPC wins X-ray but loses neutron; TIP4P/2005 overall best"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001083,"raw_usage":{"total_tokens":4557,"prompt_tokens":1002,"completion_tokens":3555,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":3457}},"tokens_in":618,"tokens_out":3555,"duration_ms":25704,"temperature":1.0,"reasoning_tokens":3457,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:45:25.938373+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the combined ranking using experimental light-water neutron total scattering structure factors, such as those from H/D isotopic substitution measurements, instead of heavy-water data at 295 K; if the top-17 ordering changes by more than the current 10% spread, the paper's conclusion about four-site dominance would be overturned.","supporting_citations":[],"review_version":1}