REVIEW 3 major objections 4 minor 75 references
Efficient dataset generation for machine learning perovskite alloys
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Cluster-based selection and uncertainty-driven active learning reduce the number of DFT training calculations needed for perovskite alloy ML models by up to half, matching accuracy on binary CsPb(Cl/Br)$_3$ and extending to ternary…
desk verdict Useful data-efficiency workflow with a solid controlled test on CsPb, but the headline CsSn accuracy number is an in-sample model-selection metric, not an unbiased estimate. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The method rests on three components. The many-body tensor representation (MBTR) turns each atomic structure into a fixed-length vector, so both clustering and kernel-based regression work in the same geometric feature space. Constrained k-means clustering over those vectors selects a spread of structures, favoring the low-symmetry phases that are hardest to predict. For active learning, a Gaussian process regression (GPR) model shares the kernel of the kernel ridge regression (KRR) predictor, so its predictive mean equals the KRR prediction while its standard deviation $\sigma$ estimates the uncertainty; relaxation trajectories run with the fast KRR model, and the GPR $\sigma$ along the trajectory decides which structures get DFT labels. The two-stage acquisition—first threshold-based uncertainty during the relaxation path, then highest-uncertainty near-equilibrium endpoints—is what turns a model with 88% relaxation convergence into one with full convergence.
What would settle it
Run the CsSn(Cl/Br/I)$_3$ active learning loop again with the same data but with the uncertainty scores replaced by random draws from the same distribution; if the final relaxation error and convergence rate match the uncertainty-driven run, the claimed data savings are not attributable to the acquisition rule. A more direct check is to record, along each ML relaxation trajectory, the actual KRR force error at the step where GPR $\sigma$ exceeds the threshold; if high-$\sigma$ steps are not statistically associated with high force errors, the mechanism is broken.
Extended reading notes
Core claim
The paper's central claim is that the bottleneck of ML for perovskite alloys—the DFT cost of generating training data—can be reduced by replacing random data collection with informed selection. Its validation on a fixed CsPb(Cl/Br)$_3$ dataset shows two things: k-means clustering over the many-body tensor representation selects single-point structures that reach the same energy and force accuracy as randomly chosen data with roughly 20% fewer calculations, and selecting relaxation trajectories according to maximum Gaussian-process uncertainty reaches the same accuracy with roughly half the trajectories. The paper then demonstrates the scheme on a new problem, generating a CsSn(Cl/Br/I)$_3$ dataset from scratch with a two-stage active learning protocol. The first stage chases relaxation convergence by sending trajectories whose uncertainty exceeds a threshold to DFT; the second stage refines accuracy near equilibrium by labeling the most uncertain relaxed endpoints. The final ML model relaxes all 100 test structures, with an average relaxation error of 0.45 meV/atom, comparable to the error previously achieved for the simpler binary alloy.
Load-bearing premise
The load-bearing premise is that the uncertainty estimate produced by the Gaussian process model reliably identifies where the kernel ridge regression's structure relaxations will go wrong; if that calibration fails, choosing trajectories by uncertainty would not outperform random sampling.
Editorial extensions
If this is right
- For binary CsPb(Cl/Br)$_3$, the combination of cluster-selected initial data, active relaxation data, and pruning reaches the same energy, force, and relaxation accuracy as random sampling while using about 20% less initial data and about half the relaxation data.
- For the ternary alloy, an ML relaxation model can be trained from a 16,000-structure initial set plus actively selected relaxation snapshots, with all test relaxations converging and an average relaxation error of 0.45 meV/atom.
- The final pruning step makes the model smaller and faster at prediction time without sacrificing accuracy, because full relaxation trajectories contain many nearly identical snapshots.
- The workflow transfers to other ML models that provide energies, forces, and uncertainties, though cluster counts and acquisition schedules should be adjusted to the model's learning rate.
- Extending the scheme to molecular dynamics potentials would require swapping structure relaxations for MD simulations inside the active learning loop, but the clustering-based initial dataset would remain valid and avoids wasting DFT data from MD trajectories.
Reading between the lines
- An extension the authors leave implicit: the 20% and 50% savings are measured on fixed datasets and emulated relaxations, so a live deployment on a fresh binary alloy would be the direct test of whether the savings persist when the pool of unlabelled structures is not precomputed.
- Because the acquisition rule leans entirely on GPR uncertainty, a testable prediction is that corrupting or shuffling those uncertainty scores should erase the active learning advantage and bring performance back to random sampling; that experiment is cheap to run on the same CsSn data.
- The paper stops single-point data generation for CsSn after 16,000 structures because energy errors had converged, but force errors were still falling; training on energies and forces jointly would likely make the initial cluster-selected set substantially smaller, aligning with the authors' own remark about future work.
- If the scheme generalizes, the same cluster-and-acquire loop could be applied to other structurally complex materials where DFT is the bottleneck, such as molecular crystals or battery interphases, with the caveat that the optimal cluster count and uncertainty threshold will need retuning per material family.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a three-step data-generation workflow for machine-learning surrogate models of perovskite alloys: k-means clustering to select a diverse initial set of single-point DFT structures, a two-stage active-learning protocol that adds DFT relaxation trajectories selected by Gaussian-process-regression uncertainty, and a final clustering-based pruning of the relaxation data. The scheme is validated on the precomputed CsPb(Cl/Br)3 dataset, where the authors report that clustering saves roughly 20% of data in the initial single-point stage and that active learning and pruning save roughly 50% relative to random selection. The same workflow is then applied to generate a new dataset for the ternary alloy CsSn(Cl/Br/I)3, for which the final ML model is reported to relax all test structures and to achieve an average relaxation error of 0.45 meV/atom.
Significance. The CsPb(Cl/Br)3 half of the paper is a genuinely controlled study: the active-versus-random comparisons in Section III B and the clustering-versus-random comparisons in Sections III A and III C provide concrete, quantitative evidence of data savings. The authors also make the generated data and code available through NOMAD, Zenodo, and GitLab, which is a practical strength for reproducibility. If the ternary claims were established, the workflow would be a useful template for building ML potentials for compositionally complex alloys. However, the ternary evaluation as reported is not yet convincing: the same 100 relaxation structures are used for monitoring, for deciding the stage switch and stopping point, and for the final accuracy numbers, so the reported 0.45 meV/atom is not an unbiased estimate of generalization error. In addition, the ternary active-learning experiment has no random-acquisition baseline, so the headline data-efficiency claim for the ternary system is not yet supported.
major comments (3)
- [Section IV B, Fig. 7] The 100 relaxation structures described at the start of Section IV B are used simultaneously as the monitoring set in each active-learning iteration, as the basis for switching from stage 1 to stage 2 'after approximately 2000 structures', as the stopping criterion, and as the test set for the final numbers in Fig. 7b,c. The reported average relaxation error of 0.45 meV/atom and the statement that all test relaxations converge are therefore model-selection metrics, not unbiased estimates for unseen structures. Please re-evaluate the final model on a disjoint held-out set that was not used for monitoring or early stopping, or use nested evaluation with a separate monitoring set, and report the resulting relaxation accuracy and convergence rate.
- [Section IV B, Fig. 7c] The ternary active-learning experiment has no random-acquisition baseline and no repeated runs. The learning curve in Fig. 7c therefore cannot demonstrate that the GPR-uncertainty acquisition rule is what improves the model, rather than simply the addition of more relaxation data. Please add a comparison against random trajectory selection at the same data budget, and preferably also against a simpler uncertainty-threshold rule, with multiple repetitions or bootstrap confidence intervals. Without such a baseline, the data-efficiency claim for the ternary alloy is not established.
- [Section II C and Section IV B] The acquisition rule in the active-learning loop assumes that the GPR predictive uncertainty sigma reliably identifies structures where the KRR model will make large errors during relaxation. The paper does not provide a calibration check or an ablation for this assumption, and the CsSn monitoring curve cannot validate it because of the test-set reuse identified above. I recommend reporting a calibration plot (binned predicted GPR sigma versus actual KRR/DFT energy or force error on a held-out set), or at least comparing acquisitions selected by GPR sigma with acquisitions selected by random sampling or by a different uncertainty proxy.
minor comments (4)
- [Section III B] The text states force prediction errors of 17.6 meV/Å for active learning and '19.6 meV/atom for random selection'; the second unit should be meV/Å, not meV/atom.
- [Section IV A and Section IV B] The total number of relaxation snapshots added by the active-learning protocol for CsSn is not reported explicitly; the text only states that stage switching occurred after approximately 2000 added structures. Please report the final training-set size so that the 'about 20% more than the single point structures' statement in Section V can be checked.
- [Figures 2, 4, 5, 6, 7] The learning curves are shown without error bars even though the binary results are stated to be means of three or five randomized splits. Reporting standard deviations or confidence bands would make the apparent differences between methods easier to assess and would strengthen the quantitative claims.
- [Section IV A] The decision to stop single-point data generation at 16,000 structures is based on the energy learning curve while the force error is still decreasing (Fig. 6b). A sentence explaining why the force error was not considered a stopping criterion would clarify the workflow.
Circularity Check
CsSn ternary validation reuses the same 100-structure test set for active-learning monitoring, stage switching, early stopping, and the reported 0.45 meV/atom relaxation MAE, so that headline metric is a model-selection quantity rather than an independent prediction.
-
fitted input called prediction
[Section IV B, FIG. 7c]
"We monitored the performance of the ML model by repeating the relaxation test described above in each iteration of the active learning loop. The resulting learning curves are shown in FIG. 7c. ... At this point, we changed to the second stage of the active learning protocol. ... With the final ML model, all the test relaxations converge and the MAE is only 0.45 meV/atom."
The 'relaxation test described above' is the same 100-structure test set (25 structures per phase) whose final errors are displayed in FIG. 7b and summarized as 0.45 meV/atom. The learning curve in FIG. 7c is computed on this same set and is the explicit decision criterion for switching from stage 1 to stage 2 and for stopping the active-learning loop. Therefore the reported final relaxation MAE is the value of the objective used to select the model and its stopping point, not an unbiased estimate on unseen structures. The abstract's 'robust and highly accurate' characterization of the CsSn model rests on this in-sample metric.
-
fitted input called prediction
[Section IV A, FIG. 6]
"After each batch finished computing, we refitted our MBTR-KRR model to monitor the convergence of the energy and force predictions on the test set. ... After adding 16 000 training structures the energy MAE has converged to approximately 0.5 meV/atom."
The 2,600 atomic structures generated 'for model testing' are the same set on which the MBTR-KRR model is monitored after every DFT batch, and the reported convergence to approximately 0.5 meV/atom is read from the learning curve that also drives the decision to stop single-point DFT data generation. Thus this reported accuracy is a monitoring/selection metric rather than an independent held-out evaluation. This affects the single-point model accuracy statement, though the paper's headline relaxation-error claim and the CsPb controlled comparison are the more important pieces of evidence.
full rationale
The core data-efficiency claim of the paper is supported by the CsPb(Cl/Br)3 experiments in Section III, where active learning is compared against random trajectory selection on disjoint training and test trajectory sets, and where clustering is compared against random selection using held-out single-point structures. Those controlled comparisons give independent grounding to the up-to-20% and 50% data-reduction numbers. The CsSn(Cl/Br/I)3 extension, however, is not circularly derived from a self-citation chain; its problem is that the same 100 relaxation test structures are used simultaneously as the monitoring set for stage switching and early stopping and as the final evaluation set. This makes the reported 0.45 meV/atom relaxation MAE and the 'all test relaxations converge' statement model-selection outcomes, not unbiased predictions on unseen data. The same reuse pattern appears in the single-point monitoring in Section IV A, where the 2,600-structure test set drives both the stopping decision and the reported convergence value. No random-acquisition baseline is reported for CsSn, so the active-learning-specific data savings are not independently demonstrated for the ternary system; that is a lack of control rather than circularity. Self-citations to prior MBTR-KRR, DScribe, and BOSS work refer to established tools and do not carry the load of the central derivation. Overall the circularity is partial: the binary-alloy validation is sound, but the ternary claim reduces to a selection metric on the test set that was used to choose the model.
Assumptions & free parameters
free parameters (14)
- Initial dataset cluster count (binary) =
500
- Initial dataset cluster count (ternary) =
2000
- MBTR xmin =
-0.1 /angstrom
- MBTR xmax =
0.6 /angstrom
- MBTR N =
50
- MBTR sigma =
1.752e-2 /angstrom (CsPb), 4.279e-2 /angstrom (CsSn)
- MBTR rcut =
6.27 angstrom (CsPb), 7.04 angstrom (CsSn)
- MBTR wcut =
1.0e-3
- KRR alpha =
1.0e-5
- KRR gamma =
1.280e-4 (CsPb), 4.279e-2 (CsSn)
- GPR uncertainty threshold =
0.5 meV/atom (initial)
- DFT force convergence limit (stage 1) =
0.1 eV/angstrom
- Stage 2 force threshold =
0.005 eV/angstrom
- Structures per phase per active learning iteration =
25 generated; 2 selected
assumptions (5)
- domain assumption PBEsol DFT with ZORA and tier-2 basis sets provides accurate reference energies and forces for halide perovskites.
- domain assumption The MBTR descriptor (k=2 term) is a sufficient representation of perovskite structures for both KRR energy prediction and k-means clustering.
- domain assumption The Gaussian process regression posterior standard deviation sigma is a meaningful measure of kernel ridge regression prediction error.
- domain assumption The four space groups (Pm-3m, P4/mbm, I4/mcm, Pnma) span the structurally relevant configuration space of the studied perovskite alloys.
- domain assumption BFGS relaxation with the ML model converges to the same local minima as DFT relaxation when the ML energy surface is accurate.
Cite this review
Pith. "Pith review of Efficient dataset generation for machine learning perovskite alloys." pith.science (2026). https://pith.science/paper/YNE7DLC4
@misc{pith2026250605777,
author = {Pith},
title = {Pith review of: Efficient dataset generation for machine learning perovskite alloys},
year = {2026},
howpublished = {\url{https://pith.science/paper/YNE7DLC4}},
note = {Machine review of arXiv:2506.05777}
}
abstract
Lead-based perovskite solar cells have reached high efficiencies, but toxicity and lack of stability hinder their wide-scale adoption. These issues have been partially addressed through compositional engineering of perovskite materials, but the vast complexity of the perovskite materials space poses a significant obstacle to exploration. We previously demonstrated how machine learning (ML) can accelerate property predictions for the CsPb(Cl/Br)$_3$ perovskite alloy. However, the substantial computational demand of density functional theory (DFT) calculations required for model training prevents applications to more complex materials. Here, we introduce a data-efficient scheme to facilitate model training, validated initially on CsPb(Cl/Br)$_3$ data and extended to the ternary alloy CsSn(Cl/Br/I)$_3$. Our approach employs clustering to construct a compact yet diverse initial dataset of atomic structures. We then apply a two-stage active learning approach to first improve the reliability of the ML-based structure relaxations and then refine accuracy near equilibrium structures. Tests for CsPb(Cl/Br)$_3$ demonstrate that our scheme reduces the number of required DFT calculations during the different parts of our proposed model training method by up to 20% and 50%. The fitted model for CsSn(Cl/Br/I)$_3$ is robust and highly accurate, evidenced by the convergence of all ML-based structure relaxations in our tests and an average relaxation error of only 0.5 meV/atom.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
C. Liu, Y. Yang, H. Chen, J. Xu, A. Liu, A. S. Bati, H. Zhu, L. Grater, S. S. Hadke, C. Huang, et al., Bimolec- ularly passivated interface enables efficient and stable in- verted perovskite solar cells, Science 382, 810 (2023)
work page 2023
-
[2]
45 meV/ atom. V. DISCUSSION For the fixed sized CsPb(Cl/Br) 3 dataset, we observed that clustering with a larger number of clusters produced better results. This is to be expected, because we chose the training structures randomly from each cluster. For a fixed dataset size, a large number of smaller clusters in- creases diversity, whereas a smaller number ...
work page 2000
-
[3]
National Renewable Energy Laboratory, Best research- cell efficiency chart, https://www.nrel.gov/pv/ cell-efficiency.html (2024), accessed: 6 Nov 2024
work page 2024
- [4]
-
[5]
M. Lu, Y. Zhang, S. Wang, J. Guo, W. W. Yu, and A. L. Rogach, Metal halide perovskite light-emitting devices: promising technology for next-generation displays, Adv. Funct. Mater. 29, 1902008 (2019)
work page 2019
-
[6]
F. Ma, Y. Zhao, Z. Qu, and J. You, Developments of highly efficient perovskite solar cells, Acc. Mater. Res. 4, 716 (2023)
work page 2023
-
[7]
X.-K. Liu, W. Xu, S. Bai, Y. Jin, J. Wang, R. H. Friend, and F. Gao, Metal halide perovskites for light-emitting diodes, Nat. Mater. 20, 10 (2021)
work page 2021
-
[8]
A. Fakharuddin, M. K. Gangishetty, M. Abdi-Jalebi, S.-H. Chin, A. R. bin Mohd Yusoff, D. N. Congreve, W. Tress, F. Deschler, M. Vasilopoulou, and H. J. Bolink, Perovskite light-emitting diodes, Nat. Electron. 5, 203 (2022)
work page 2022
Show all 75 references
-
[9]
Park and S
B.-W. Park and S. I. Seok, Intrinsic instability of inorganic–organic hybrid halide perovskite materials, Adv. Mater. 31, 1805337 (2019). 10
2019
-
[10]
Zhou and Y
Y. Zhou and Y. Zhao, Chemical stability and instability of inorganic halide perovskites, Energy Environ. Sci. 12, 1495 (2019)
2019
-
[11]
T. Hu, D. Li, Q. Shan, Y. Dong, H. Xiang, W. C. Choy, and H. Zeng, Defect behaviors in perovskite light- emitting diodes, ACS Mater. Lett. 3, 1702 (2021)
2021
-
[12]
B. Chen, S. Wang, Y. Song, C. Li, and F. Hao, A critical review on the moisture stability of halide perovskite films and solar cells, Chem. Eng. J. 430, 132701 (2022)
2022
-
[13]
Ke and M
W. Ke and M. G. Kanatzidis, Prospects for low-toxicity lead-free perovskite solar cells, Nat. Commun. 10, 965 (2019)
2019
-
[14]
Giustino and H
F. Giustino and H. J. Snaith, Toward lead-free perovskite solar cells, ACS Energy lett. 1, 1233 (2016)
2016
-
[15]
Zhang, Y
Y. Zhang, Y. Ma, Y. Wang, X. Zhang, C. Zuo, L. Shen, and L. Ding, Lead-free perovskite photode- tectors: progress, challenges, and opportunities, Adv. Mater. 33, 2006691 (2021)
2021
-
[16]
Konstantakou and T
M. Konstantakou and T. Stergiopoulos, A critical review on tin halide perovskite solar cells, J. Mater. Chem. A 5, 11518 (2017)
2017
-
[17]
Zhang, Y
Y. Zhang, Y. Liu, and S. F. Liu, Composition engineering of perovskite single crystals for high-performance opto- electronics, Adv. Funct. Mater. 33, 2210335 (2023)
2023
-
[18]
Saliba, Polyelemental, multicomponent perovskite semiconductor libraries through combinatorial screening, Adv
M. Saliba, Polyelemental, multicomponent perovskite semiconductor libraries through combinatorial screening, Adv. Energy Mater. 9, 1803754 (2019)
2019
-
[19]
K. Ji, M. Anaya, A. Abfalterer, and S. D. Stranks, Halide perovskite light-emitting diode technologies, Adv. Opt. Mater. 9, 2002128 (2021)
2021
-
[20]
S. Sun, A. Tiihonen, F. Oviedo, Z. Liu, J. Thapa, Y. Zhao, N. T. P. Hartono, A. Goyal, T. Heumueller, C. Batali, et al., A data fusion approach to optimize com- positional stability of halide perovskites, Matter 4, 1305 (2021)
2021
-
[21]
F. Xu, M. Zhang, Z. Li, X. Yang, and R. Zhu, Challenges and perspectives toward future wide-bandgap mixed- halide perovskite photovoltaics, Adv. Energy Mater. 13, 2203911 (2023)
2023
-
[22]
Karlsson, Z
M. Karlsson, Z. Yi, S. Reichert, X. Luo, W. Lin, Z. Zhang, C. Bao, R. Zhang, S. Bai, G. Zheng, et al. , Mixed halide perovskites for spectrally stable and high- efficiency blue light-emitting diodes, Nat. Commun. 12, 361 (2021)
2021
-
[23]
Y. Li, L. G. McKinney, Y. He, S.-Y. Liu, and S. Wang, First-principles investigation of stable lead-free halide perovskite materials CsSnCl x Bry I3 – x – y for solar cell ap- plications, J. Phys.: Condens. Matter 35, 435501 (2023)
2023
-
[24]
Lyu, J.-H
M. Lyu, J.-H. Yun, P. Chen, M. Hao, and L. Wang, Ad- dressing toxicity of lead: Progress and applications of low-toxic metal halide perovskites and their derivatives, Adv. Energy Mater. 7, 1602512 (2017)
2017
-
[25]
Himanen, A
L. Himanen, A. Geurts, A. S. Foster, and P. Rinke, Data- driven materials science: Status, challenges, and perspec- tives, Adv. Sci. 6, 1900808 (2019)
2019
-
[26]
J. S. Bechtel and A. Van der Ven, First-principles ther- modynamics study of phase stability in inorganic halide perovskite solid solutions, Phys. Rev. Mater. 2, 045401 (2018)
2018
-
[27]
S. S. Chong, Y. S. Ng, H.-Q. Wang, and J.-C. Zheng, Advances of machine learning in materials science: Ideas and techniques, Front. Phys. 19, 13501 (2023)
2023
-
[28]
Schmidt, M
J. Schmidt, M. R. G. Marques, S. Botti, and M. A. L. Marques, Recent advances and applications of machine learning in solid-state materials science, npj Comput. Mater. 5, 83 (2019)
2019
-
[29]
D. P. Kov´ acs, I. Batatia, E. S. Arany, and G. Cs´ anyi, Evaluation of the mace force field architecture: From medicinal chemistry to materials science, J. Chem. Phys. 159 (2023)
2023
-
[30]
W. Ye, C. Chen, Z. Wang, I.-H. Chu, and S. P. Ong, Deep neural networks for accurate predictions of crystal stability, Nat. Commun. 9, 3800 (2018)
2018
-
[31]
Batzner, A
S. Batzner, A. Musaelian, L. Sun, M. Geiger, J. P. Mailoa, M. Kornbluth, N. Molinari, T. E. Smidt, and B. Kozinsky, E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials, Nat. Commun. 13, 2453 (2022)
2022
-
[32]
Stanev, C
V. Stanev, C. Oses, A. G. Kusne, E. Rodriguez, J. Paglione, S. Curtarolo, and I. Takeuchi, Machine learn- ing modeling of superconducting critical temperature, npj Comput. Mater. 4, 29 (2018)
2018
-
[33]
A. Jain, S. P. Ong, G. Hautier, W. Chen, W. D. Richards, S. Dacek, S. Cholia, D. Gunter, D. Skinner, G. Ceder, et al. , Commentary: The materials project: A materials genome approach to accelerating materials innovation, APL mater. 1 (2013)
2013
-
[34]
K. T. Sch¨ utt, H. E. Sauceda, P.-J. Kindermans, A. Tkatchenko, and K.-R. M¨ uller, SchNet – A deep learn- ing architecture for molecules and materials, J. Chem. Phys. 148, 241722 (2018)
2018
-
[35]
Laakso, M
J. Laakso, M. Todorovi´ c, J. Li, G.-X. Zhang, and P. Rinke, Compositional engineering of perovskites with machine learning, Phys. Rev. Mater. 6, 113801 (2022)
2022
-
[36]
Curtarolo, W
S. Curtarolo, W. Setyawan, G. L. Hart, M. Jahnatek, R. V. Chepulskii, R. H. Taylor, S. Wang, J. Xue, K. Yang, O. Levy, et al. , Aflow: An automatic framework for high- throughput materials discovery, Comput. Mater. Sci. 58, 218 (2012)
2012
-
[37]
Marion, A
M. Marion, A. ¨Ust¨ un, L. Pozzobon, A. Wang, M. Fadaee, and S. Hooker, When less is more: Investigating data pruning for pretraining llms at scale, arXiv preprint arXiv:2309.04564 (2023)
2023 arXiv
-
[38]
S. Yang, Z. Xie, H. Peng, M. Xu, M. Sun, and P. Li, Dataset pruning: Reducing training data by examining generalization influence, in The Eleventh International Conference on Learning Representations (2023)
2023
-
[39]
Wengert, G
S. Wengert, G. Cs´ anyi, K. Reuter, and J. T. Mar- graf, Data-efficient machine learning for molecular crystal structure prediction, Chem. Sci. 12, 4536 (2021)
2021
-
[40]
Killamsetty, X
K. Killamsetty, X. Zhao, F. Chen, and R. Iyer, Retrieve: Coreset selection for efficient and robust semi-supervised learning, Adv. neural inf. process. syst. 34, 14488 (2021)
2021
-
[41]
Sivaraman, A
G. Sivaraman, A. N. Krishnamoorthy, M. Baur, C. Holm, M. Stan, G. Cs´ anyi, C. Benmore, and ´A. V´ azquez- Mayagoitia, Machine-learned interatomic potentials by active learning: amorphous and liquid hafnium dioxide, npj Comput. Mater. 6, 104 (2020)
2020
-
[42]
J. Qi, T. W. Ko, B. C. Wood, T. A. Pham, and S. P. Ong, Robust training of machine learning interatomic potentials with dimensionality reduction and stratified sampling, npj Comput. Mater. 10, 43 (2024)
2024
-
[43]
Zhang, G
L. Zhang, G. Cs´ anyi, E. Van Der Giessen, and F. Maresca, Atomistic fracture in bcc iron revealed by active learning of gaussian approximation potential, npj Comput. Mater. 9, 217 (2023). 11
2023
-
[44]
Settles, Active Learning Literature Survey , Com- puter Sciences Technical Report 1648 (University of Wisconsin–Madison, 2009)
B. Settles, Active Learning Literature Survey , Com- puter Sciences Technical Report 1648 (University of Wisconsin–Madison, 2009)
2009
-
[45]
Zhang, D.-Y
L. Zhang, D.-Y. Lin, H. Wang, R. Car, and W. E, Active learning of uniformly accurate interatomic potentials for materials simulation, Phys. Rev. Mater. 3, 023804 (2019)
2019
-
[46]
Gubaev, E
K. Gubaev, E. V. Podryabinkin, G. L. Hart, and A. V. Shapeev, Accelerating high-throughput searches for new alloys with active learning of interatomic potentials, Comput. Mater. Sci. 156, 148 (2019)
2019
-
[47]
Todorovi´ c, M
M. Todorovi´ c, M. U. Gutmann, J. Corander, and P. Rinke, Bayesian inference of atomistic structure in functional materials, npj Comput. Mater. 5, 35 (2019)
2019
-
[48]
Schran, K
C. Schran, K. Brezina, and O. Marsalek, Committee neu- ral network potentials control generalization errors and enable active learning, J. Chem. Phys. 153 (2020)
2020
-
[49]
J. Li, F. Pan, G.-X. Zhang, Z. Liu, H. Dong, D. Wang, Z. Jiang, W. Ren, Z.-G. Ye, M. Todorovi´ c, et al. , Struc- tural disorder by octahedral tilting in inorganic halide perovskites: New insight with bayesian optimization, Small Struc. , 2400268 (2023)
2023
-
[50]
Jinnouchi, J
R. Jinnouchi, J. Lahnsteiner, F. Karsai, G. Kresse, and M. Bokdam, Phase transitions of hybrid perovskites sim- ulated by machine-learning force fields trained on the fly with bayesian inference, Phys. Rev. Lett. 122, 225701 (2019)
2019
-
[51]
Himanen, M
L. Himanen, M. O. J. J¨ ager, E. V. Morooka, F. F. Canova, Y. S. Ranawat, D. Z. Gao, P. Rinke, and A. S. Foster, Dscribe: Library of descriptors for machine learn- ing in materials science, Comp. Phys. Commun. 247, 106949 (2020)
2020
-
[52]
Huo and M
H. Huo and M. Rupp, Unified representation of molecules and crystals for machine learning, Mach. Learn.: Sci. Technol. 3, 045017 (2022)
2022
-
[53]
Stuke, M
A. Stuke, M. Todorovi´ c, M. Rupp, C. Kunkel, K. Ghosh, L. Himanen, and P. Rinke, Chemical diversity in molecu- lar orbital energy predictions with kernel ridge regression, J. Chem. Phys. 150 (2019)
2019
-
[54]
Laakso, L
J. Laakso, L. Himanen, H. Homm, E. V. Morooka, M. O. J. J¨ ager, M. Todorovi´ c, and P. Rinke, Updates to the DScribe library: New descriptors and derivatives, J. Chem. Phys. 158, 234802 (2023)
2023
-
[55]
P. S. Bradley, K. P. Bennett, and A. Demiriz, Con- strained k-means clustering, Microsoft Research, Red- mond 20, 0 (2000)
2000
-
[56]
L. Fang, J. Laakso, P. Rinke, and X. Chen, Machine- learning accelerated structure search for ligand-protected clusters, J. Chem. Phys. 160 (2024)
2024
-
[57]
See Supplemental Material at [URL will be inserted by publisher] for additional details on the machine learning model hyperparameters and analysis of the coreset selec- tion tests
-
[58]
Levy-Kramer, k-means-constrained (2018)
J. Levy-Kramer, k-means-constrained (2018)
2018
-
[59]
V. Blum, R. Gehrke, F. Hanke, P. Havu, V. Havu, X. Ren, K. Reuter, and M. Scheffler, Ab initio molecular simulations with numeric atom-centered orbitals, Com- put. Phys. Commun. 180, 2175 (2009)
2009
-
[60]
Fletcher, Practical methods of optimization (John Wi- ley & Sons, 2000)
R. Fletcher, Practical methods of optimization (John Wi- ley & Sons, 2000)
2000
-
[61]
X. Ren, P. Rinke, V. Blum, J. Wieferink, A. Tkatchenko, S. Andrea, K. Reuter, V. Blum, and M. Scheffler, Resolution-of-identity approach to Hartree-Fock, hybrid density functionals, RPA, MP2, and GW with numeric atom-centered orbital basis functions, New J. Phys. 14, 053020 (2012)
2012
-
[62]
V. Havu, V. Blum, P. Havu, and M. Scheffler, Efficient O(N ) integration for all-electron electronic structure cal- culation using numeric basis functions, J. Comput. Phys. 228, 8367 (2009)
2009
-
[63]
J. P. Perdew, A. Ruzsinszky, G. I. Csonka, O. A. Vydrov, G. E. Scuseria, L. A. Constantin, X. Zhou, and K. Burke, Restoring the density-gradient expansion for exchange in solids and surfaces, Phys. Rev. Lett. 100, 136406 (2008)
2008
-
[64]
S. V. Levchenko, X. Ren, J. Wieferink, R. Johanni, P. Rinke, V. Blum, and M. Scheffler, Hybrid function- als for large periodic systems in an all-electron, numeric atom-centered basis framework, Comput. Phys. Com- mun. 192, 60 (2015)
2015
-
[65]
All the codes used for the CsSn(Cl/Br/I)3 dataset generation are available in a Git- Lab repository [67]
and Zenodo [66]. All the codes used for the CsSn(Cl/Br/I)3 dataset generation are available in a Git- Lab repository [67]. A. Initial dataset At the single point data generation stage, we followed the same steps as for the CsPb(Cl/Br) 3 dataset. We gen- erated 100 structures f...
2000
-
[66]
van Lenthe, E
E. van Lenthe, E. J. Baerends, and J. G. Snijders, Rel- ativistic regular two-component hamiltonians, J. Chem. Phys. 99, 4597 (1993)
1993
-
[67]
NOMAD Repository, https://doi.org/10.17172/NOMAD/2024.11.08-1
2024 doi
-
[68]
Zenodo Repository, https://doi.org/10.5281/zenodo.14056015
-
[69]
GitLab Repository, https://gitlab.com/cest-group/learnsolar-cssnclbri
-
[70]
Stuke, P
A. Stuke, P. Rinke, and M. Todorovi´ c, Efficient hyperpa- rameter tuning for kernel ridge regression with Bayesian optimization, Mach. Learn.: Sci. Technol. 2, 035022 (2021). Supplementary Material: Efficient dataset generation for machine learning perovskite alloys Henrietta Hom...
2021
-
[71]
and HDBSCAN (Hierarchical Density-Based Spatial Clustering of Appli cations with Noise) [2] are presented in FIG. S1. These tests were performed, similarly to the validation re sults presented in the article, with the set of 7500 single point CsPb(Cl/Br) 3 structures used as t...
-
[72]
279 × 10− 2 rcut (˚ A) 6.27 7.04 wcut 1
752 × 10− 2 4. 279 × 10− 2 rcut (˚ A) 6.27 7.04 wcut 1. 0 × 10− 3 1. 0 × 10− 3 α 1. 0 × 10− 5 1. 0 × 10− 5 γ 1. 280 × 10− 4 4. 279 × 10− 2 TABLE S1: Hyperparameters of the ML models for CsPb(Cl/Br) 3 and CsSn(Cl/Br/I) 3. S3. MACHINE LEARNING MODEL HYPERP ARAMETERS The machine ...
-
[73]
Zhang, R
T. Zhang, R. Ramakrishnan, and M. Livny, Birch: an efficient d ata clustering method for very large databases, ACM sigmod record 25, 103 (1996)
1996
-
[74]
R. J. Campello, D. Moulavi, and J. Sander, Density-based clu stering based on hierarchical density estimates, in Pacific-Asia conference on knowledge discovery and data mining (Springer, 2013) pp. 160–172
2013
-
[75]
Laakso, M
J. Laakso, M. Todorovi´ c, J. Li, G.-X. Zhang, and P. Rinke, Compositional engineering of perovskites with machine learni ng, Physical Review Materials 6, 113801 (2022)
2022
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.