Pith. sign in

REVIEW 2 cited by

Random sampling versus active learning algorithms for machine learning potentials of quantum liquid water

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.10698 v1 pith:LUAEFIM4 submitted 2024-10-14 physics.chem-ph

classification physics.chem-ph
keywords learningactivedatapotentialsstructurestrainingliquidquantum
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Training accurate machine learning potentials requires electronic structure data comprehensively covering the configurational space of the system of interest. As the construction of this data is computationally demanding, many schemes for identifying the most important structures have been proposed. Here, we compare the performance of high-dimensional neural network potentials (HDNNPs) for quantum liquid water at ambient conditions trained to data sets constructed using random sampling as well as various flavors of active learning based on query by committee. Contrary to the common understanding of active learning, we find that for a given data set size, random sampling leads to smaller test errors for structures not included in the training process. In our analysis we show that this can be related to small energy offsets caused by a bias in structures added in active learning, which can be overcome by using instead energy correlations as an error measure that is invariant to such shifts. Still, all HDNNPs yield very similar and accurate structural properties of quantum liquid water, which demonstrates the robustness of the training procedure with respect to the training set construction algorithm even when trained to as few as 200 structures. However, we find that for active learning based on preliminary potentials, a reasonable initial data set is important to avoid an unnecessary extension of the covered configuration space to less relevant regions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Same Quality Metrics, Different Graph Drawings

    cs.CG 2025-08 unverdicted novelty 5.0 of 10

    Quality metrics for graph drawings can stay almost unchanged while a drawing is deformed into a very different target shape, so single metrics cannot certify drawing quality.

  2. The CP2K Program Package Made Simple

    physics.comp-ph 2025-08 accept novelty 2.0 of 10

    A user-oriented review of CP2K collecting input recipes for DFT, quantum chemistry, GW/BSE, spectroscopy, and embedding methods, plus a benchmark of the UZH basis and pseudopotential protocol against all-electron FP-LAPW.

Pith tools