{"id":"4876f162-2e01-43a7-8e3b-077af6be07ab","arxiv_id":"1908.08734","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An automated active-learning procedure trains a single high-dimensional neural network potential that reproduces CCSD(T*)-F12a/aug-cc-pVTZ energies for protonated water clusters from H3O+ to H9O4+ with a fitting error of 0.06 kJ/mol per atom.","lead":"This paper builds an automated pipeline that fits neural network potentials to expensive coupled-cluster calculations for protonated water clusters, reaching fitting errors around 0.06 kJ/mol per atom. If confirmed, this makes high-accuracy quantum chemistry simulations of these clusters much faster and easier to run.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Active-learning convergence relies on two-network disagreement as an error proxy; shared architecture and training data can make disagreement vanish while both NNPs share a systematic bias, so the reported test RMSE may not certify accuracy outside the selected distribution.","rationale":"I read the paper as a methodological and applied demonstration: an automated, active-learning-driven fit of a high-dimensional NNP to CCSD(T*)-F12a/aug-cc-pVTZ energies for protonated water clusters, with extensive but not exhaustive validation. The strongest claim is that the resulting single PES represents the reference method at a fitting error of 0.06 kJ/mol per atom and is reliable for stationary points, reaction pathways, and MD/PIMD sampling. The paper provides substantial internal support: a very small training/test RMSE, consistency across cluster sizes, good harmonic frequency agreement, accurate MEP profiles, and direct CC checks along short MD/PIMD trajectory windows. The learning-curve analysis in the SI is also a meaningful convergence check. The soft spot is not the fitting quality on the collected data, but the assumption that the active-learning selection strategy, based on two-network disagreement plus symmetry-function extrapolation detection, actually covers all configurations where the NNP could deviate significantly. This is exactly the reader's weakest_assumption. The two networks are not independent models in the strong sense: they share the same descriptor space, network family, and training pipeline, so common-mode bias is plausible. The final random-split test set inherits any blind spots of the selection procedure. The concrete test I propose would settle whether the selection was representative by evaluating the NNP on configurations from the original physical ensembles that were deliberately not chosen during active learning. This is a real but nonfatal concern: the paper's validations are enough to support a conditional accept, not to reject the central claim, and not to require an immediate change in the reader's verdict. I therefore leave the verdict unchanged.","tokens_in":18109,"tokens_out":6142,"duration_ms":72260,"concrete_test":"Draw roughly 1000 configurations per cluster from the original DFT AIMD/AI-PIMD ensembles described in Sec. III that were never selected during the 80+ refinement stages, compute CCSD(T*)-F12a/aug-cc-pVTZ energies for them, and compare with the final NNP. If the per-atom RMSE on these untouched configurations is close to the Table I test-set values (0.06-0.10 kJ/mol per atom), the committee selection was representative. If the RMSE is substantially larger (for example above 0.3 kJ/mol per atom) or grows systematically with cluster size, the active-learning loop converged on a biased training distribution and the headline fitting error is not transferable. As a secondary probe, record the two-network disagreement at the final iteration for these same configurations and correlate it with the actual CCSD error, testing directly whether small committee disagreement implies small error.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is the active-learning convergence criterion described in Sec. II and used in Sec. IV A: iteration stops when the energy difference between two independently trained NNPs is small for all ensembles. This equates committee disagreement with NNP error, but the two networks share the same symmetry functions, the same architecture, the same fitting code, and largely the same training data; they differ mainly in random initialization and Kalman-filter updates. They can therefore converge to the same biased function, especially in regions that are underrepresented in the training distribution, and their disagreement can shrink without the actual CCSD error being reduced. Strategy II, which detects extrapolation by comparing symmetry-function values to the training-set range, only flags points outside the already covered range, so a smooth but systematically wrong region inside that range is invisible. The final test set is a random split of the reference data that were themselves selected by these committee/extrapolation heuristics; it is thus in-distribution and cannot by itself establish accuracy for configurations the procedure never chose. The validation on MEPs and MD trajectories is valuable, but those paths were generated by the NNP itself, so they can lie inside the learned region. This does not refute the central result for the tested structures, but it means the 'unbiased/automated' claim is only as strong as the calibration of committee disagreement to actual reference error, and that calibration is not directly demonstrated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an automated active-learning procedure for constructing a high-dimensional neural network potential (NNP) for protonated water clusters (H3O+ through H9O4+ plus H2O) trained to CCSD(T*)-F12a/aug-cc-pVTZ reference energies. The procedure iteratively selects training configurations using two strategies: (i) disagreement between two independently fitted NNPs and (ii) detection of symmetry-function values outside the range of the current training set. The final NNP achieves a training-set RMSE of 0.06 kJ/mol per atom and a test-set RMSE of 0.08 kJ/mol per atom over the full data set. The fit is validated by comparison to coupled-cluster reference calculations for binding energies of stationary points, harmonic frequencies, potential-energy scans, minimum-energy paths, and short segments of classical and path-integral molecular dynamics trajectories. The authors conclude that the automated procedure enables fast and accurate construction of coupled-cluster-quality potential energy surfaces for finite clusters.","tokens_in":18394,"tokens_out":6327,"duration_ms":62215,"significance":"If the reported accuracy and validation hold, this work is a significant methodological advance: it demonstrates that an automated, committee-based active-learning scheme can produce a single NNP that covers multiple cluster sizes at essentially coupled-cluster accuracy with a small number of reference calculations relative to grid-based approaches. The paper's strengths include a clear train/test split with the test set withheld from the fit, validation against reference data for diverse properties (binding energies, frequencies, scans, MEPs, MD/PIMD), and the public provision of the coupled-cluster reference data in the Supporting Information. The automated nature of the fitting procedure, if reliable, would make such potentials practical for systems beyond the present test case. However, the load-bearing assumption that committee disagreement is a calibrated proxy for true error is not directly tested, and the molecular dynamics validation covers only a very short trajectory window.","major_comments":[{"comment":"The claim that the automated procedure selects configurations 'in an unbiased and efficient way' rests on the assumption that committee disagreement (Strategy I) and symmetry-function range (Strategy II) are reliable proxies for the NNP error relative to the CCSD(T) reference. The manuscript does not calibrate these proxies: no comparison is shown between committee disagreement and the actual NNP error on the held-out test set, and the stopping criterion ('until the differences between the networks are converged') is only qualitative. Because the test set is a random split of the reference data that were themselves selected by these heuristics, the test RMSE does not by itself certify accuracy in configurations that the procedure might systematically ignore. The external validations (scans, MEPs, MD) are reassuring, but they sample a limited set of low-dimensional paths. Please provide a quantitative calibration of committee disagreement versus true error (e.g., a plot on the test set), define a concrete convergence threshold, and show that further iterations do not change predictions on a fixed validation set; alternatively, temper the 'unbiased' claim to 'heuristic-driven' and discuss the associated risk.","section":"Sec. II and Sec. IVA"},{"comment":"The MD/PIMD validation is limited to the last 100 steps of each 25 ps trajectory, which at the stated time steps corresponds to 50 fs of classical MD (0.5 fs step) and 25 fs of PIMD (0.25 fs step). This is a very short window and does not sample the full range of configurations explored in the trajectories, including the isomerization event described for the 600 K run (ending near the Ring isomer). Thus the statement that the NNP 'reliably describes protonated water clusters during classical MD and quantum PIMD simulations at various conditions' is only directly demonstrated for a small fraction of the simulation. Please extend the coupled-cluster recomputation to more points along the trajectories (or provide statistical measures over the full trajectory) and clarify precisely which frames were compared.","section":"Sec. IVD and Fig. 7"},{"comment":"The abstract and Sec. IVA quote a 'fitting error of 0.06 kJ/mol per atom' without specifying that this is the training-set RMSE; the independent test-set RMSE for the full data set is 0.08 kJ/mol per atom (Table I). While both values are excellent, the abstract should be precise about which error is being reported, and the difference between training and test error (about 25%) is worth a sentence because it bears on the generalization claim.","section":"Abstract and Table I"}],"minor_comments":[{"comment":"The phrase 'and recall only that' appears to be a grammatical error; it should likely read 'and recall that' or 'and recall only that' with proper punctuation.","section":"Sec. II, first paragraph"},{"comment":"The thermostat is referred to as 'Nos–Hover chain thermostat'; this should be 'Nosé–Hoover chain thermostat'.","section":"Sec. III, MD setup"},{"comment":"The phrase 'protonated water tertramer' contains a typo; it should be 'tetramer'.","section":"Sec. IVC, final paragraph"},{"comment":"The text states that 'at least 100 000 uncorrelated structures are extracted from these simulations for each cluster', but the trajectory lengths given in Sec. III (at least 100 ps AIMD and 25 ps AI-PIMD with configurations spaced 10 fs) would yield about 10,000 and 2,500 structures per cluster, respectively. Please check this number and reconcile the discrepancy.","section":"Sec. IVA, first paragraph"},{"comment":"Eq. (1) defines E_bind using n water monomers, but for H3O+ (n=1) the formula gives zero, whereas Table II reports -171.4 kcal/mol for M=1 and the table caption notes that E_bind is the energy difference between H3O+ and water. Please clarify the definition for the monomer case.","section":"Eq. (1) and Table II"},{"comment":"The text says the NNP is 'about 10^8 times faster' than the reference calculation; the stated values (7 h vs. 0.1 ms) correspond to about 2.5×10^8, so the order of magnitude is correct, but the wording could be made more precise.","section":"Sec. IVA, speed comparison"}],"recommendation":"major_revision","confidential_remarks":"The paper is well written and the central results appear sound, but the 'unbiased' claim for the active-learning procedure is stronger than the evidence supports. The committee-disagreement proxy is a common heuristic, but the authors should either calibrate it or soften the claim. The short MD validation window is a separate but important limitation that should be addressed in the revised manuscript. I am recommending major revision rather than rejection because the identified issues are fixable within the scope of the paper and the overall contribution is substantial."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a genuine, carefully validated advance—a single neural network potential fitted to CCSD(T*)-F12a/aug-cc-pVTZ energies for H3O+ through H9O4+ that hits ~0.08 kJ/mol/atom on a held-out test set. The central claim holds up for the structures they actually checked.\n\nWhat's new: they combine committee-based active learning (Strategy I) with symmetry-function-range extrapolation detection (Strategy II) into an automated loop, then join the cluster-specific training sets into one NNP. The validation is thorough for an ML potential paper: stationary-point binding energies within 0.1 kcal/mol, harmonic frequencies within ~10 cm^-1, scans along proton transfer and umbrella inversion up to 150 kJ/mol, and MEPs for isomerizations with CC single points on the NNP paths. The fact that forces weren't used in fitting but frequencies still come out that well is a good sign, and they explicitly test on H3O+ whether adding force labels changes anything (it doesn't, given enough energy points).\n\nSoft spots, in order of seriousness.\n\n1. The convergence criterion for active learning is committee disagreement. The stress-test worry is real: two NNPs sharing architecture, symmetry functions, and nearly all training data can converge to the same biased function, so the disagreement can shrink while the actual CC error in some region stays high. Strategy II catches points outside the symmetry-function range, but not smooth errors inside the range. The random test split is in-distribution. That said, the independent MEP and scan checks are genuinely outside the training distribution, and they agree. So I see this as a caveat on the \"automated/unbiased\" language, not a fatal flaw.\n\n2. The finite-temperature validation is a spot check. The \"last 100 points\" of the MD/PIMD trajectories is only 25–50 fs. That's enough to show the NNP tracks CC energies over small fluctuations, but it doesn't validate long-time stability or rare-event sampling. The earlier active-learning runs went up to 8 ns with the NNP, but those weren't re-evaluated against CC.\n\n3. Reproducibility: no code or data in the preprint; the SI apparently has the CC reference sets on the journal site. RuNNer is GPL but not in the paper. Not disqualifying, but worth a referee asking for a public data deposit.\n\nThe citation pattern is appropriate—they cite the relevant Behler/Marx/others work and competing many-body approaches like PIPs and MB-pol, and they're honest about the tradeoffs. No red flags.\n\nWho should read this: anyone building ML potentials for finite clusters, and anyone working on protonated water. It deserves serious peer review; I'd accept it with minor revisions, mainly asking for code/data and a longer CC-revalidated trajectory segment if feasible.","headline":"Solid, carefully validated NNP at CCSD(T) accuracy, though the active-learning convergence criterion is a heuristic and the MD validation is a short spot check.","tokens_in":18899,"tokens_out":3720,"would_cite":true,"duration_ms":39341,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A neural network potential fitted automatically to ~50,000 coupled cluster energies reproduces protonated water clusters from H3O+ to H9O4+ within 0.06 kJ/mol per atom.","keywords":["neural network potentials","active learning","coupled cluster","potential energy surface","protonated water clusters","path integral molecular dynamics","molecular dynamics","CCSD(T*)-F12a"],"falsifier":"Take the published final training set, run long NNP-PIMD simulations of H9O4+ at 300 K and 600 K, and compute CCSD(T*)-F12a/aug-cc-pVTZ energies for, say, one thousand uncorrelated frames that were never part of training; if the per-atom RMSE on these frames exceeds about 0.1 kJ/mol, the claim that the potential is converged for production simulations would be disproven. A cheaper check would be to compute explicit coupled cluster forces for a sample of tetramer configurations and compare them with NNP forces, since the paper only tests forces explicitly on H3O+.","tokens_in":17915,"feed_emoji":"💧","tokens_out":8113,"duration_ms":72354,"temperature":0.7,"pith_summary":"The paper attempts to show that a single neural network potential can stand in for expensive coupled cluster calculations across an entire family of molecules, not just one fixed structure. Using an automated active-learning loop, the authors fit one potential to CCSD(T*)-F12a/aug-cc-pVTZ reference energies for H2O, H3O+, H5O2+, H7O3+, and H9O4+, reaching a fitting error of 0.06 kJ/mol per atom. They then demonstrate that the same potential reproduces stationary-point energies, harmonic frequencies, reaction paths, and energies sampled by classical and path integral molecular dynamics. The practical payoff, if the claim holds, is that simulations with essentially converged electronic structure accuracy become routine for these clusters, with each energy evaluation about eight orders of magnitude cheaper than the reference calculation.","feed_headline":"Neural net matches coupled cluster accuracy for protonated water","feed_subtitle":"A single fitted potential covers five cluster sizes at 0.06 kJ/mol per atom, making coupled cluster simulations routine.","key_machinery":"The mechanism that carries the argument is committee-based active learning combined with extrapolation detection. At each iteration, two independent neural network potentials are trained on the current reference set and used to predict energies of a large pool of unlabeled configurations from DFT-based classical and path integral trajectories; the 20 configurations with largest disagreement between the two networks are selected for new coupled cluster calculations. A second selection strategy flags configurations whose atom-centered symmetry function values fall outside the range spanned by the training set, which are encountered when preliminary NNP-based simulations explore new regions. The underlying architecture is a high-dimensional neural network potential: atom-centered symmetry functions convert the structure into invariant input vectors, atomic neural networks output atomic energy contributions that sum to the total energy, and the whole function is analytically differentiable so forces are available. This machinery is what lets the training set grow automatically and keeps the number of expensive reference calculations near a practical minimum.","core_discovery":"The central claim is that the automated fitting procedure, not a hand-curated grid of configurations, produces a neural network potential whose errors sit at the intrinsic uncertainty of the coupled cluster reference. Trained on roughly 50,000 energy-only reference points and tested on 5,470 held-out points, the single NNP reaches 0.06 kJ/mol per atom RMSE on the training set and 0.08 kJ/mol per atom on the test set, with binding energies of all optimized clusters matching the reference to within 0.1 kcal/mol and harmonic frequencies within about 10 cm−1. The paper further claims the potential remains accurate far from equilibrium, along proton-transfer coordinates and minimum energy paths for isomerization, and that the same functional form supports classical molecular dynamics from 300 to 600 K and path integral simulations down to 1.67 K. In short, the authors claim to have made coupled-cluster-level force evaluation cheap and automatic for finite protonated water clusters.","pith_inferences":["The two-network disagreement test is not tied to the neural network architecture in any essential way, so the same automated loop could be applied to other machine-learned potentials such as Gaussian approximation potentials or moment tensor potentials, with the symmetry-function extrapolation test replaced by an equivalent representation-space distance.","The paper's validation re-evaluates only the last 100 frames of 25 ps trajectories; a stricter test would be to compute coupled cluster energies for thousands of uncorrelated frames from long production runs, and the reported test-set error suggests such a check would likely pass but is not explicitly demonstrated.","If the procedure transfers to larger clusters, one could imagine building a CCSD(T)-quality description of bulk proton transport from cluster-derived training data, because the symmetry functions used here already cover the full cluster with cutoff radii beyond the molecular size."],"forward_implications":["A single set of neural network parameters can describe several cluster sizes and isomers at coupled cluster accuracy, so separate fits per molecule are unnecessary.","Classical and path integral simulations of these clusters can run at essentially converged electronic structure accuracy, including quantum nuclear effects at temperatures from 1.67 K to 600 K.","Energy-only training is sufficient once the reference set is dense; adding force labels to the hydronium data did not improve the fit, so the expensive computation of coupled cluster forces can be avoided.","The automated selection protocol reduces the number of reference calculations needed compared with grid-based fitting, making larger finite systems a feasible next step."],"supporting_citations":[{"why":"Introduces the high-dimensional neural network potential architecture that maps atom-centered symmetry functions to atomic energy contributions.","marker":"[32]"},{"why":"First uses disagreement between two neural network potentials to detect configurations underrepresented in the training set; this becomes Strategy I.","marker":"[43]"},{"why":"Establishes that the flexibility of neural network potentials can be used to identify deficiencies in the training set.","marker":"[23]"},{"why":"Provides the query-by-committee learning principle and the exponential error decrease that justifies the two-network selection.","marker":"[57]"},{"why":"Defines the CCSD(T*)-F12a/aug-cc-pVTZ reference method used for all energies.","marker":"[69]"},{"why":"Supplies the symmetry function parameters optimized for water that define the NNP input representation.","marker":"[73]"},{"why":"Demonstrates NNP fitting for protonated water clusters at DFT level, the baseline this work extends to coupled cluster accuracy.","marker":"[51]"},{"why":"Earlier demonstration that machine-learned potentials can reach coupled cluster accuracy for finite clusters, motivating the current automated setup.","marker":"[47]"}],"fun_headline_variants":["Auto-fitted neural nets match coupled cluster accuracy for protonated water","One NNP for all protonated water clusters at CCSD(T) accuracy","Coupled cluster accuracy on autopilot for protonated water clusters","Automated fit: single NN potential hits CCSD(T) level for protonated water","Neural net PES at coupled cluster accuracy, no hand tuning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the disagreement between two independently trained neural networks, plus the symmetry-function range check, reliably points to the configurations where the potential most needs new coupled cluster data; if both networks share a systematic bias, the active-learning loop could stop early and the reported test error would be optimistic.","fun_headline_variants_meta":{"raw":{"variants":["Auto-fitted neural nets match coupled cluster accuracy for protonated water","One NNP for all protonated water clusters at CCSD(T) accuracy","Coupled cluster accuracy on autopilot for protonated water clusters","Automated fit: single NN potential hits CCSD(T) level for protonated water","Neural net PES at coupled cluster accuracy, no hand tuning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001367,"raw_usage":{"total_tokens":5593,"prompt_tokens":1047,"completion_tokens":4546,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":663,"completion_tokens_details":{"reasoning_tokens":4449}},"tokens_in":663,"tokens_out":4546,"duration_ms":31255,"temperature":1.0,"reasoning_tokens":4449,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:30:20.445407+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the published final training set, run long NNP-PIMD simulations of H9O4+ at 300 K and 600 K, and compute CCSD(T*)-F12a/aug-cc-pVTZ energies for, say, one thousand uncorrelated frames that were never part of training; if the per-atom RMSE on these frames exceeds about 0.1 kJ/mol, the claim that the potential is converged for production simulations would be disproven. A cheaper check would be to compute explicit coupled cluster forces for a sample of tetramer configurations and compare them with NNP forces, since the paper only tests forces explicitly on H3O+.","supporting_citations":[{"cited_title":"Artrith \\ and\\ author J","cited_arxiv_id":null,"evidence_quote":"First uses disagreement between two neural network potentials to detect configurations underrepresented in the training set; this becomes Strategy I."},{"cited_title":"Morawietz , author O","cited_arxiv_id":null,"evidence_quote":"Supplies the symmetry function parameters optimized for water that define the NNP input representation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates NNP fitting for protonated water clusters at DFT level, the baseline this work extends to coupled cluster accuracy."},{"cited_title":"Schran , author F","cited_arxiv_id":null,"evidence_quote":"Earlier demonstration that machine-learned potentials can reach coupled cluster accuracy for finite clusters, motivating the current automated setup."}],"review_version":1}