Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

Gradient-Guided Furthest Point Sampling for Robust Training Set Selection

T0 review · 2 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read This paper claims that weighting furthest-point sampling distances by molecular force norms, with an alternating exponent that favors both high- and low-gradient structures, builds training sets up to twice as data-efficient as plain FPS wh

desk verdict GGFPS is a sensible idea with a solid MD17 story, but the main ST data-efficiency claim is undermined by an uncontrolled test-set protocol, and the SI algorithm contradicts its own sign-alternation mechanism. read the letter →

arxiv 2510.08906 v2 pith:D3KWM7F6 submitted 2025-10-10 stat.ML cs.LGphysics.chem-ph

classification stat.MLcs.LGphysics.chem-ph
keywords trainingsetselectionfurthestpointsamplinggradient-guidedkernelridgeregressionpotentialenergysurfacesmoleculardynamicsdataefficiencymodelrobustness
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces GGFPS, a training-set selection rule that multiplies furthest-point-sampling distances by a power of the molecular force norm, with the power alternating between positive and negative values across the selection run. The authors argue that because force norms track the local variance of energy labels, this weighting covers both descriptor space and label space, avoiding FPS's blind spot: FPS systematically under-samples equilibrium geometries in MD17 trajectories. On a two-dimensional multimodal test function and six MD17 molecules, GGFPS matches or beats FPS and uniform random sampling, achieving the same accuracy with up to half the training data and cutting prediction-error variance. If true, the result means simple gradient information, already a byproduct of reference calculations, can substantially lower the data cost and improve robustness of learned potential-energy surfaces.

What carries the argument

The load-bearing mechanism is the GGFPS score function s_j = g_j^{beta_k} d_j, combining the Euclidean distance from already-selected points with the L2 norm of the atomic forces. The hyperparameter beta defines an interval [-beta, beta] whose linearly spaced values are re-ordered alternately, so early selections alternate between high- and low-gradient points rather than committing to one regime; beta = 0 recovers plain FPS. The first point is also drawn with probability proportional to gradient norm. This device converts a purely geometric greedy selector into a supervised selector that trades coverage against label variance.

What would settle it

On a potential-energy surface engineered so that force norms and label variance are anticorrelated (for example, a steep-walled valley with constant energy and a flat region with a sudden energy cliff), run GGFPS and FPS at N=50; GGFPS's advantage should reverse. More directly, on an MD17 molecule, bin test configurations by force norm and check whether GGFPS's gains vanish when the kernel is changed to one that already includes gradient information.

Watch

Extended reading notes

Core claim

The central discovery is that pure descriptor-space coverage is not enough: FPS selects configurations that are mutually far apart in representation space, and on the non-uniform molecular datasets those far-apart points are systematically strained, high-energy structures. GGFPS instead scores every candidate by d_j times g_j raised to an exponent that alternates between negative and positive values across the selection. This makes the sparse early training set contain both low- and high-gradient points, then fills in medium-gradient configurations as sampling proceeds. The result is a training set that spans the entire dataset while intentionally over-weighting regions where force norms ind

Load-bearing premise

The load-bearing premise is that force norms act as a reliable proxy for local energy-label variance: if a surface has large forces in regions of nearly constant energy, or flat forces where energies vary sharply, the weighted scores will push sampling in the wrong direction.

Editorial extensions

If this is right

  • On the Styblinski-Tang function, GGFPS training sets reach the full-labeled-set mean absolute error with 50% fewer points and match FPS accuracy with up to half the training points.
  • On MD17, GGFPS training sets lower prediction errors for equilibrium and strained structures relative to both FPS and uniform random sampling, with the largest gains in the low-data regime.
  • GGFPS reduces prediction-error variance across all six MD17 molecules, by up to an order of magnitude relative to FPS and up to about 7x relative to uniform random sampling for aspirin, so models are less prone to sporadic large errors.
  • The distributional analysis implies that FPS's poor MD17 performance is systematic, not a tuning accident: it under-samples low-force-norm equilibrium geometries, and GGFPS explicitly corrects this.
  • Because forces are a standard byproduct of reference electronic-structure calculations, the method adds no new data-acquisition cost and can be applied to existing labeled datasets.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My inference: the alternating-exponent trick could improve other greedy selectors, such as active-learning queries or stratified sampling, since the pathology it fixes (early commitment to one regime) is not specific to FPS.
  • My inference: GGFPS's variance reduction suggests it should also improve uncertainty quantification; a testable extension is whether error bars from an ensemble of GGFPS-trained models are better calibrated than those from FPS or uniform-sampling ensembles.
  • My inference: the paper only tests interpolative settings where train and test come from the same potential-energy surface; the method's value for extrapolation across chemical space is an open, testable question that the authors themselves flag.
  • My inference: using force norms as a label-variance proxy could backfire on surfaces where gradients are large but energies are smooth; a cheap diagnostic on new datasets would be to compare GGFPS against FPS on a held-out validation set before committing to the beta schedule.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces Gradient-Guided Furthest Point Sampling (GGFPS), an extension of furthest point sampling that weights the FPS distance score by molecular force norms raised to an exponent β, with an alternating exponent schedule intended to balance high- and low-gradient regions. The method is evaluated with kernel ridge regression on the two-dimensional Styblinski–Tang function and on six MD17 molecular trajectories, comparing against uniform random sampling (URS), FPS, and constant-exponent variants. The authors report that GGFPS improves data efficiency and robustness, particularly for equilibrium and strained molecular configurations, and reduces prediction-error variance. The paper is empirical, with no formal derivations; the main evidence consists of learning curves, error-surface plots, and distribution analyses over bootstrap repetitions.

Significance. If the empirical claims hold, GGFPS would be a simple, practical contribution to training-set selection for molecular machine learning, where force labels are often already available from reference calculations. The benchmark is broad (six molecules, training sizes 50–1000, bootstrap repetitions), and the comparison includes both standard baselines and supervised FPS-style selectors. A notable strength is that the MD17 experiments consistently show lower mean errors and lower error variances for GGFPS across many training-set sizes, not just at a single favorable operating point. The paper also identifies a potentially important failure mode of plain FPS on Boltzmann-distributed molecular data. However, the central quantitative efficiency claim for the toy system rests on a comparison protocol with unmatched test sets, and the stated algorithmic mechanism for achieving low-gradient coverage is contradicted by the pseudocode as written. These issues need to be resolved before the contributions can be accepted at face value.

major comments (2)
  1. [Sec. IIB and Algorithm 1 (Appendix, lines 5–9)] The alternating β sequence is described as interpolating between −β and β and as 'flipping index signs during interpolation' to ensure that the sparse early training set contains both low- and high-gradient points. As written, however, the sequence does not alternate signs. With β_list ordered from −β to β, the reindexed sequence is β_N (positive), −β_2 (positive, since β_2 is negative), β_{N−1} (positive), −β_3 (positive), and so on. Since β_j is negative only for j < (N+1)/2, the even-indexed terms stay positive through roughly the first half of the iterations; for even N only the final term is negative, and for odd N no term is negative. Thus the algorithm as pseudocoded strongly favors high-gradient points and does not implement the stated low-gradient inclusion mechanism. This is load-bearing because the MD17 interpretation (Sec. IIIB, Figs. 6–8) attributes the method's success at l
  2. [Sec. IIIA and Fig. 3] The Styblinski–Tang learning curves for FPS and GGFPS are not computed on a fixed test set. The text states that labeled sets of size L are split into training sets of size N and test sets of the remaining L−N points, while URS learning curves are 'invariant to the initial labeled set size,' implying URS is evaluated on a separate fixed test set. Cross-N comparisons—such as the claim that GGFPS reaches the same accuracy as FPS with half the training points—therefore confound sampling quality with changes in test-set size and composition. The authors' own note in the Fig. 3 caption, that the apparent advantage of GGFPS at N=950 over URS at N=1,000 is an artifact of MAE, confirms that the protocol is not fully controlled. The ST efficiency claims should be recomputed with a common held-out test set for all methods and training-set sizes, and the factor-of-2 claim should be verified under t
minor comments (5)
  1. [General] Algorithm 1 is referred to as being in the SI, but it appears in the Appendix; the reference should be fixed.
  2. [Sec. IIB] The statement that molecular force norms 'indirectly describe the variance of molecular energy labels' is a heuristic motivation. A quantitative analysis relating force norms to local label variance (e.g., residuals of a descriptor-space model) on the MD17 sets would substantially strengthen the conceptual basis of the method.
  3. [Fig. 3 and SI Fig. 9] The y-axis label 'MAE ST Function [arb. u.]' and the caption use inconsistent notation for β and β′; please unify notation across the figure and text.
  4. [Abstract and Sec. I] The abstract contains a typo ('Styblinksi-Tang'); the correct spelling is Styblinski–Tang.
  5. [Reproducibility] No data or code availability statement is provided. For an empirical benchmark paper, making the implementation and bootstrap scripts available would be important for reproducibility and for resolving the algorithmic ambiguity above.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: GGFPS is an empirical sampling method, β is cross-validated, and the reported gains do not reduce to the method's inputs by construction.

full rationale

The paper makes no formal derivation claim; it reports an empirical comparison of sampling strategies. GGFPS's only fitted parameter, β, is selected by grid-search cross-validation on each training fold, not by fitting to the held-out test errors, and β=0 explicitly recovers FPS, providing a non-circular control. The supervised use of force norms is the method's design premise, not a renamed prediction of the outcome. Self-citations (FCHL19, sGDML, and related descriptor work) are used as standard representations and tools, not as load-bearing support for the central superiority claim. The ST learning-curve comparison does contain a potential evaluation mismatch—GGFPS/FPS test on the L−N leftover points while URS curves are described as invariant to the labeled set size, and the paper itself admits the N=950 MAE advantage 'is an artifact of using the MAE metric, and disappears when RMSE is used instead'—but this is a validity concern about experimental protocol, not circularity, because the reported advantage is not equivalent to the inputs by construction. No circular step can be quoted and exhibited from the paper's equations or protocol. Therefore the circularity score is 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on three classes of assumptions: (i) force norms proxy label variance, (ii) the MD17 setup and FCHL19/KRR model are representative, and (iii) the printed β schedule behaves as described. The last is contradicted by the paper's own algorithm—a key finding of this review.

free parameters (3)
  • β (gradient bias exponent) = 0–2, selected per fold by grid-search CV (SI Fig. 18 shows MD17 optimal β distributions)
    Controls the strength of force-norm weighting in GGFPS; tuned on the labeled pool for each training-set fold, so the reported gains partly reflect hyperparameter selection not available to FPS/URS.
  • KRR kernel width σ = Optimized per fold; SI Fig. 9 shows GGFPS optimal widths 1.5–2x smaller than FPS/URS on ST
    Standard KRR hyperparameter tuned by 5-fold CV; affects all methods but interacts with sampling density.
  • KRR regularization λ = Not reported numerically, tuned by CV
    Standard ridge penalty in Eq. (7), tuned per fold for all methods.
assumptions (4)
  • domain assumption Molecular force norms (‖F_i‖₂) are a valid proxy for the local variance of energy labels, so weighting sampling by g_j^β covers label-space variance.
    Sec. IIB: 'Because molecular force norms indirectly describe the variance of molecular energy labels...' This is the key modeling premise; if false for a given PES, GGFPS could underperform FPS.
  • domain assumption MD17 trajectory samples are approximately Boltzmann-distributed, and test configurations come from the same PES as the labeled pool.
    Sec. IIIB: 'samples from the chemical potential energy surface are approximately Boltzmann distributed'; the train/test split (25k labeled, 50k test) assumes the pooled selection and evaluation distribution match.
  • ad hoc to paper The alternating β sequence in Algorithm 1 (lines 5–9) produces a mixture of positive and negative β values in early iterations.
    The reindexing rule gives β_k = β_{N−⌊(k−1)/2⌋} for odd k and −β_{⌊k/2⌋+1} for even k with β_1 = −β, which yields non-negative β for the first ~N/2 iterations, contradicting the prose claim of sign flipping. The paper's experimental results implicitly assume the implementation differs from the printed pseudo-code.
  • domain assumption FCHL19 descriptors and a local Gaussian kernel yield a KRR model accurate enough for the comparisons.
    Sec. IID; standard in the field, but the conclusions are specific to this representation/kernel; other models might not show the same ordering.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Gradient-Guided Furthest Point Sampling for Robust Training Set Selection." pith.science (2026). https://pith.science/paper/D3KWM7F6

@misc{pith2026251008906,
  author       = {Pith},
  title        = {Pith review of: Gradient-Guided Furthest Point Sampling for Robust Training Set Selection},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/D3KWM7F6}},
  note         = {Machine review of arXiv:2510.08906}
}
read the original abstract

Training set sampling methods are used to improve model performance and lower data costs in machine learning problems relevant to chemistry. We introduce Gradient Guided Furthest Point Sampling (GGFPS), a simple extension of Furthest Point Sampling (FPS) that leverages molecular force norms to guide efficient sampling of configurational spaces of molecules. Numerical evidence is presented for a toy system (the Styblinski-Tang function) as well as for molecular dynamics trajectories from the MD17 dataset. Our numerical results indicate superior data efficiency and model robustness when using GGFPS compared to FPS and uniform random sampling (URS), as well as established supervised FPS-style selectors, PCov-FPS and PCov-CUR. Distribution analysis of the MD17 data suggests that FPS systematically under-samples equilibrium geometries, resulting in large test errors for relaxed structures. GGFPS cures this artifact and (i) enables up to twofold reductions in training cost without sacrificing predictive accuracy compared to FPS in the 2-dimensional Styblinski-Tang system, (ii) systematically lowers prediction errors for equilibrium as well as strained structures in MD17, and (iii) systematically decreases prediction error variances across all of the MD17 configuration spaces. These results suggest that gradient-aware sampling methods hold great promise as effective training set selection tools, and that naive use of FPS may result in imbalanced training and inconsistent prediction outcomes.

Figures

Figures reproduced from arXiv: 2510.08906 by the authors.

Figure 1
Figure 1. Left: A subset (orange) of labeled data (grey) from the Lennard-Jones potential, which represent a well performing [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Left: A contour plot of the ST function surface in two dimensions. Center-left: The ST function gradient norm [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. The MAE predictions of the 2D ST function [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (15 more)
Figure 5
Figure 5. Figure 5: Scatter plots of the 2D ST function absolute test [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 3
Figure 3. Figure 3: However, the FPS models are less robust, with [PITH_FULL_IMAGE:figures/full_fig_p006_3.png]
Figure 6
Figure 6. Figure 6: MD17 FCHL19 force norm (top row) and energy (bottom row) distributions of training sets, each with 100 [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Top row: The MAE predictions of the MD17 trajectories with respect to training set size, [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: MD17 trajectory absolute test errors versus force norms for FPS, GGPFS, and URS training sets (blue, orange, [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 9
Figure 9. Figure 9: Left: The RMSE predictions of the 2D ST function surface with respect to training set size for URS, FPS, and [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Scatter plots of the 2D ST function test errors corresponding to models trained on FPS (top row) and GGFPS [PITH_FULL_IMAGE:figures/full_fig_p014_10.png]
Figure 11
Figure 11. Figure 11: Left: A contour plot of the toy function surface in two dimensions. Center-left: The toy function gradient norm [PITH_FULL_IMAGE:figures/full_fig_p015_11.png]
Figure 12
Figure 12. Figure 12: Top row: The RMSE predictions of the MD17 trajectories with respect to training set size, [PITH_FULL_IMAGE:figures/full_fig_p015_12.png]
Figure 13
Figure 13. Figure 13: MD17 trajectory test errors versus force norms for FPS, GGPFS, and URS training sets (blue, orange, green), [PITH_FULL_IMAGE:figures/full_fig_p016_13.png]
Figure 14
Figure 14. Figure 14: MD17 trajectory test errors versus force norms for FPS, GGPFS, and URS training sets (blue, orange, green), [PITH_FULL_IMAGE:figures/full_fig_p016_14.png]
Figure 15
Figure 15. Figure 15: MD17 trajectory test errors versus force norms for FPS, GGPFS, and URS training sets (blue, orange, green), [PITH_FULL_IMAGE:figures/full_fig_p017_15.png]
Figure 16
Figure 16. Figure 16: MD17 trajectory test errors versus force norms for FPS, GGPFS, and URS training sets (blue, orange, green), [PITH_FULL_IMAGE:figures/full_fig_p017_16.png]
Figure 17
Figure 17. Figure 17: MD17 trajectory test errors versus force norms for FPS, GGPFS, and URS training sets (blue, orange, green), [PITH_FULL_IMAGE:figures/full_fig_p018_17.png]
Figure 18
Figure 18. Figure 18: MD17 FCHL19 optimal cross-validated GGFPS [PITH_FULL_IMAGE:figures/full_fig_p018_18.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Multi-Scale Machine Learning Framework for Coupled Chemical, Spin, and Structural Disorder in Alloys

    cond-mat.mtrl-sci 2026-07 conditional novelty 6.0 of 10

    A coupled MC+MD framework using GNNs and MLIPs reproduces Fe-Co phase transition (1,000 K) and melting (1,690 K) temperatures while capturing structural transitions in Fe-Co-C alloys.

Reference graph

Works this paper leans on

76 extracted references · 8 canonical work pages · cited by 1 Pith paper

  1. [1]

    The global minimum of the function is located atx∗ = [−2.903534,−2.903534], with a function value off(x ∗) =−78.33198

    The Styblinski-Tang Function The Styblinski-Tang (ST) function is a multi-modal, d-dimensional benchmark function used to test opti- mization algorithms.[70, 71] It contains a mixture of wells surrounded by steep walls and is defined ford= 2 as f(x) = 1 2 2X i=1 x4 i −16x 2 i + 5xi (8) wherex= [x 1, x2]T ∈R 2. The global minimum of the function is located...

  2. [2]

    backwards

    The MD17 Trajectories We represent the configurations of the MD17 as- pirin, toluene, malonaldehyde, naphthalene, paraceta- mol, and uracil trajectories[27] with FCHL19 [72], a faster albeit slightly less accurate version of the original Faber–Christensen–Huang–Lilienfeld (FCHL) representation.[73] FCHL19 is a smooth local repre- sentation that contains r...

  3. [3]

    Simon-Gabriel and B

    C.-J. Simon-Gabriel and B. Schölkopf, Journal of Ma- chine Learning Research19, 1 (2018)

  4. [4]

    Hofmann, B

    T. Hofmann, B. Schölkopf, and A. J. Smola, The An- nals of Statistics36, 1171 (2008)

  5. [5]

    Herbrich,Learning Kernel Classifiers: Theory and Algorithms(The MIT Press, 2001)

    R. Herbrich,Learning Kernel Classifiers: Theory and Algorithms(The MIT Press, 2001)

  6. [6]

    Vapnik,The nature of statistical learning theory (Springer science & business media, 2013)

    V. Vapnik,The nature of statistical learning theory (Springer science & business media, 2013)

  7. [7]

    M. Rupp, A. Tkatchenko, K.-R. Müller, and O. A. von Lilienfeld, Phys. Rev. Lett.108, 058301 (2012)

  8. [8]

    J.BehlerandM.Parrinello,Phys.Rev.Lett.98,146401 (2007)

Show all 76 references
  1. [10]

    O. A. von Lilienfeld, Angew. Chem. Int. Ed Engl.57, 4164 (2018)

  2. [11]

    Hansen, G

    K. Hansen, G. Montavon, F. Biegler, S. Fazli, M. Rupp, M. Scheffler, O. A. von Lilienfeld, A. Tkatchenko, and K.-R. Müller, J. Chem. Theory Comput.9, 3404 (2013), http://pubs.acs.org/doi/pdf/10.1021/ct400195d

  3. [12]

    Hansen, F

    K. Hansen, F. Biegler, O. A. von Lilienfeld, K.-R. Müller, and A. Tkatchenko, J. Phys. Chem. Lett.6, 2326 (2015)

  4. [13]

    F. A. Faber, L. Hutchison, B. Huang, J. Gilmer, S. S. Schoenholz, G. E. Dahl, O. Vinyals, S. Kearnes, P. F. Riley, and O. A. von Lilienfeld, J. Chem. Theory Com- put.13, 5255 (2017)

  5. [14]

    G. N. Simm and M. Reiher, J. Chem. Theory Comput. 14, 5238 (2018)

  6. [15]

    Proppe, S

    J. Proppe, S. Gugler, and M. Reiher, J. Chem. Theory Comput.15, 6046 (2019), 1906.09342

  7. [16]

    Gugler and M

    S. Gugler and M. Reiher, J. Chem. Theory Comput. 18, 6670 (2022)

  8. [17]

    Molecular similarity in ma- chine learning of energies in chemical reaction net- works,

    S. Gugler and M. Reiher, “Molecular similarity in ma- chine learning of energies in chemical reaction net- works,” (2025)

  9. [18]

    Zhang and C

    Y. Zhang and C. Ling, npj Comput Mater4, 25 (2018)

  10. [19]

    Chmiela, H

    S. Chmiela, H. E. Sauceda, I. Poltavsky, K.-R. Müller, and A. Tkatchenko, Comput. Phys. Commun.240, 38 (2019)

  11. [20]

    A. S. Christensen, L. A. Bratholm, F. A. Faber, and O. Anatole von Lilienfeld, J. Chem. Phys.152, 044107 (2020)

  12. [21]

    K. T. Butler, D. W. Davies, H. Cartwright, O. Isayev, and A. Walsh, Nature559, 547 (2018)

  13. [22]

    Himanen, A

    L. Himanen, A. Geurts, A. S. Foster, and P. Rinke, Adv. Sci.6, 1900808 (2019)

  14. [23]

    Batra, L

    R. Batra, L. Song, and R. Ramprasad, Nat Rev Mater 6, 655 (2021)

  15. [24]

    Ramprasad, R

    R. Ramprasad, R. Batra, G. Pilania, A. Mannodi- Kanakkithodi, and C. Kim, npj Comput Mater3, 54 (2017). 11

  16. [25]

    Schmidt, M

    J. Schmidt, M. R. G. Marques, S. Botti, and M. A. L. Marques, npj Comput Mater5, 1 (2019)

  17. [26]

    G. R. Schleder, A. C. M. Padilha, C. M. Acosta, M. Costa, and A. Fazzio, J. Phys. Mater.2, 032001 (2019)

  18. [27]

    P. O. Dral, J. Phys. Chem. Lett.11, 2336 (2020)

  19. [28]

    Carleo, I

    G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby, L. Vogt-Maranto, and L. Zdeborová, Rev. Mod. Phys.91, 045002 (2019), 1903.10563

  20. [29]

    Chmiela, A

    S. Chmiela, A. Tkatchenko, H. E. Sauceda, I. Poltavsky, K. T. Schütt, and K.-R. Müller, Science Advances3, e1603015 (2017)

  21. [30]

    Schütt, P.-J

    K. Schütt, P.-J. Kindermans, H. E. Sauceda Felix, S. Chmiela, A. Tkatchenko, and K.-R. Müller, in Advances in Neural Information Processing Systems, Vol. 30, edited by I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Gar- nett (Curran Associat...

  22. [31]

    K. T. Schütt, F. Arbabzadah, S. Chmiela, K. R. Müller, and A. Tkatchenko, Nature Communications8(2017), 10.1038/ncomms13890

  23. [32]

    Ramakrishnan, P

    R. Ramakrishnan, P. Dral, M. Rupp, and O. A. von Lilienfeld, Scientific Data1, 140022 (2014)

  24. [33]

    van der Oord, M

    C. van der Oord, M. Sachs, D. P. Kovács, C. Ortner, and G. Csányi, Npj Comput. Mater.9(2023)

  25. [34]

    Dragoni, T

    D. Dragoni, T. D. Daff, G. Csányi, and N. Marzari, Physical Review Materials2(2018), 10.1103/physrev- materials.2.013808

  26. [35]

    R. K. Cersonsky, B. A. Helfrecht, E. A. Engel, S. Kli- avinek, and M. Ceriotti, Machine Learning: Science and Technology2, 035038 (2021)

  27. [37]

    Wengert, G

    S. Wengert, G. Csányi, K. Reuter, and J. T. Margraf, Chemical Science12, 4536–4546 (2021)

  28. [38]

    Célerse, M

    F. Célerse, M. D. Wodrich, S. Vela, S. Gallarati, R. Fab- regat, V. Juraskova, and C. Corminboeuf, Journal of Chemical Information and Modeling64, 1201–1212 (2024)

  29. [39]

    T. T. Nguyen, E. Székely, G. Imbalzano, J. Behler, G. Csányi, M. Ceriotti, A. W. Götz, and F. Pae- sani, The Journal of Chemical Physics148(2018), 10.1063/1.5024577

  30. [40]

    J. Qi, T. W. Ko, B. C. Wood, T. A. Pham, and S. P. Ong, npj Computational Materials10(2024), 10.1038/s41524-024-01227-4

  31. [41]

    V. L. Deringer, A. P. Bartók, N. Bernstein, D. M. Wilkins, M. Ceriotti, and G. Csányi, Chemical Reviews 121, 10073–10141 (2021)

  32. [42]

    Riquelme-Granada, K

    N. Riquelme-Granada, K. A. Nguyen, and Z. Luo, inCommunications in Computer and Information Sci- ence,Communicationsincomputerandinformationsci- ence (Springer International Publishing, Cham, 2021) pp. 195–222

  33. [43]

    Y. Wen, Z. Li, Y. Xiang, and D. Reker, Digit. Discov. (2023)

  34. [44]

    He and E

    H. He and E. A. Garcia, IEEE Trans. Knowl. Data Eng. 21, 1263 (2009)

  35. [45]

    Mannodi-Kanakkithodi, G

    A. Mannodi-Kanakkithodi, G. Pilania, T. D. Huan, T. Lookman, and R. Ramprasad, Scientific reports6, 1 (2016)

  36. [46]

    V. Botu, R. Batra, J. Chapman, and R. Ramprasad, The Journal of Physical Chemistry C121, 511–522 (2016)

  37. [47]

    A. P. Bartók, S. De, C. Poelking, N. Bernstein, J. R. Kermode, G. Csányi, and M. Ceriotti, Science Ad- vances3(2017), 10.1126/sciadv.1701816

  38. [48]

    D. J. Rosenkrantz, R. E. Stearns, and P. M. Lewis, II, SIAM J. Comput.6, 563 (1977)

  39. [49]

    Eldar, M

    Y. Eldar, M. Lindenbaum, M. Porat, and Y. Zeevi, in Proceedings of the 12th IAPR International Conference on Pattern Recognition, Vol. 2 - Conference B: Com- puter Vision & Image Processing. (Cat. No.94CH3440- 5)(1994) pp. 93–97 vol.3

  40. [50]

    Imbalzano, A

    G. Imbalzano, A. Anelli, D. Giofré, S. Klees, J. Behler, and M. Ceriotti, The Journal of Chemical Physics148 (2018), 10.1063/1.5024611

  41. [51]

    P. Rowe, V. L. Deringer, P. Gasparotto, G. Csányi, and A. Michaelides, The Journal of Chemical Physics153 (2020), 10.1063/5.0005084

  42. [52]

    Wengert, G

    S. Wengert, G. Csányi, K. Reuter, and J. T. Mar- graf, Journal of Chemical Theory and Computation18, 4586–4593 (2022)

  43. [53]

    Boulangeot, F

    N. Boulangeot, F. Brix, F. Sur, and E. Gaudry, Jour- nal of Chemical Theory and Computation (2024), 10.1021/acs.jctc.4c00367

  44. [54]

    R. Li, C. Zhou, A. Singh, Y. Pei, G. Henkelman, and L. Li, The Journal of Chemical Physics160(2024), 10.1063/5.0187892

  45. [55]

    Montes de Oca Zapiain, M

    D. Montes de Oca Zapiain, M. A. Wood, N. Lub- bers, C. Z. Pereyra, A. P. Thompson, and D. Perez, npj Computational Materials8(2022), 10.1038/s41524- 022-00872-x

  46. [56]

    Karabin and D

    M. Karabin and D. Perez, The Journal of Chemical Physics153(2020), 10.1063/5.0013059

  47. [57]

    V. L. Deringer, M. A. Caro, and G. Csányi, Nature Communications11(2020), 10.1038/s41467-020-19168- z

  48. [58]

    Y. Zuo, C. Chen, X. Li, Z. Deng, Y. Chen, J. Behler, G. Csányi, A. V. Shapeev, A. P. Thompson, M. A. Wood, and S. P. Ong, The Journal of Physical Chem- istry A124, 731–745 (2020)

  49. [59]

    X. Jia, A. Lynch, Y. Huang, M. Danielson, I. Lang’at, A. Milder, A. E. Ruby, H. Wang, S. A. Friedler, A. J. Norquist, and J. Schrier, Nature573, 251–255 (2019)

  50. [60]

    Musil, S

    F. Musil, S. De, J. Yang, J. E. Campbell, G. M. Day, and M. Ceriotti, Chem. Sci.9, 1289 (2018)

  51. [61]

    J. S. Smith, B. Nebgen, N. Lubbers, O. Isayev, and A. E. Roitberg, J. Chem. Phys.148, 241733 (2018)

  52. [62]

    Gould, B

    T. Gould, B. Chan, S. G. Dale, and S. Vuckovic, Chem- ical Science15, 11122–11133 (2024)

  53. [63]

    Korth and S

    M. Korth and S. Grimme, Journal of Chemical Theory and Computation5, 993–1003 (2009)

  54. [64]

    R. P. Feynman, Physical Review56, 340–343 (1939)

  55. [65]

    G. C. Schatz, inLecture Notes in Chemistry(Springer Berlin Heidelberg, Berlin, Heidelberg, 2000) pp. 15–32

  56. [66]

    Chmiela, A

    S. Chmiela, A. Tkatchenko, H. E. Sauceda, I. Poltavsky, K. T. Schütt, and K.-R. Müller, Sci Adv3, e1603015 (2017)

  57. [67]

    F. A. Faber, A. S. Christensen, B. Huang, and O. A. von Lilienfeld, J. Chem. Phys.148, 241717 (2018). 12

  58. [68]

    T. D. Huan, R. Batra, J. Chapman, S. Krishnan, L. Chen, and R. Ramprasad, npj Computational Ma- terials3(2017), 10.1038/s41524-017-0042-y

  59. [69]

    Panknin, Danny, Stefan Chmiela, Klaus Robert Muller, and Shinichi Nakajima., Transactions on Machine Learning Research. (2023)

  60. [70]

    Schölkopf and A

    B. Schölkopf and A. J. Smola,Learning with kernels: support vector machines, regularization, optimization, and beyond(MIT press, 2002)

  61. [71]

    D. G. Krige, Journal of the Southern African Institute of Mining and Metallurgy52, 119 (1951)

  62. [72]

    Jamil and X.-S

    M. Jamil and X.-S. Yang, International Journal of Mathematical Modelling and Numerical Optimisation 4, 150 (2013)

  63. [73]

    Styblinski and T.-S

    M. Styblinski and T.-S. Tang, Neural Networks3, 467 (1990)

  64. [74]

    A. S. Christensen, L. A. Bratholm, F. A. Faber, and O. Anatole von Lilienfeld, The Journal of Chemical Physics152(2020), 10.1063/1.5126701

  65. [75]

    F. A. Faber, A. S. Christensen, B. Huang, and O. A. von Lilienfeld, The Journal of Chemical Physics148, 241717 (2018)

  66. [76]

    A. P. Bartók, M. C. Payne, R. Kondor, and G. Csányi, Phys. Rev. Lett.104, 136403 (2010)

  67. [77]

    Altman and M

    N. Altman and M. Krzywinski, Nat Methods15, 399 (2018)

  68. [78]

    Bellman, Press, NJ (1961)

    R. Bellman, Press, NJ (1961). 13 V. APPENDIX Algorithm 1:Gradient-Guided Furthest Point Sampling (GGFPS). For a constantβ′ value, lines 5–9 can be left out and line 11 will change toβ←β′. Input:g∈R N (gradient norms),D∈R N×N (distance matrix),N(target size),β≥0(gradient expone...

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.