REVIEW 3 major objections 4 minor 1 cited by
Materials-Discovery Workflows Guided by Symbolic Regression: Identifying Acid-Stable Oxides for Electrocatalysis
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read An active-learning workflow using ensembles of SISSO symbolic-regression models finds 12 acid-stable oxides out of 1470 candidates in only 30 hybrid-DFT evaluations, versus 2 for random selection.
desk verdict A genuinely useful SISSO-based active-learning workflow with a real screening application, but the headline efficiency claim rests on a threshold mismatch that needs fixing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is an ensemble construction the paper calls bagging with Monte-Carlo dropout of primary features: each ensemble member is a SISSO model trained on a bootstrap sample of the data with a random 20% subset of the 14 primary features retained. Averaging these models gives the mean prediction $\Delta G_{\text{pbx,ESISSO}}$, and their spread gives the uncertainty estimate $\sigma_{\text{ESISSO}}$ used in the acquisition function $\text{POF} = F\!\left(\frac{\tau - \Delta G_{\text{pbx,ESISSO}}}{\sigma_{\text{ESISSO}}}\right)$, the Gaussian cumulative probability that a candidate is acid-stable. Together they turn SISSO into a closed-loop discovery engine that selects one oxide per iteration for hybrid-DFT evaluation and retrains.
What would settle it
A decisive test would be to rerun the same 30-iteration campaign with an acquisition function that only exploits the ensemble mean, ignoring the uncertainty term; if the mean-only strategy also finds 12 acid-stable oxides, then the calibrated uncertainty estimates are not essential to the claimed gain.
Extended reading notes
Core claim
The central claim is that bagging with Monte-Carlo dropout of primary features creates SISSO ensembles whose mean prediction and standard deviation are good enough to steer active learning. On a held-out test set the ensemble reaches a mean absolute error of 0.26 eV/atom, down from 0.34 for a single SISSO model and 0.29 for plain bagging, and its miscalibration score drops from 2.76 for bagging to 1.76, still overconfident but markedly improved. The acquisition function, probability of feasibility, is the Gaussian cumulative probability that a candidate's Pourbaix decomposition free energy lies below the stability threshold. In 30 iterations this strategy identifies 12 acid-stable oxides, several of which are missed by PBE-level screening, and the SISSO descriptor map shows the selected materials clustering in the low-energy region.
Load-bearing premise
The ranking of candidates by probability of feasibility assumes that each prediction error is Gaussian with standard deviation equal to the ensemble's spread, and that this spread reliably orders which candidates the model knows least about; the paper's own calibration analysis shows the ensemble is still overconfident, so the search's efficiency rests on this unproven ranking property.
Editorial extensions
If this is right
- If correct, hybrid-DFT screening of oxide stability becomes practical for thousands of materials instead of dozens.
- The same workflow can be applied to other properties whose governing parameters are unknown, as long as primary features and an ensemble of symbolic-regression models can be defined.
- The descriptor maps produced by SISSO give interpretable axes for the materials space, showing which chemical motifs, such as Mo, Ta, and W, favor acid stability.
- Because five of the twelve discovered oxides are misclassified by PBE, the result implies that cheaper exchange-correlation functionals can mislead high-throughput screens and that ensemble uncertainty flags where higher-level calculations are needed.
- The workflow reduces the number of expensive calculations by roughly a factor of six relative to random selection in this pool, a direct measure of its efficiency.
Reading between the lines
- I infer that Monte-Carlo feature dropout works because it prevents the ensemble from over-agreeing on a single descriptor set; a similar effect could be obtained by sampling over the SISSO hyperparameters q and D, which the paper does not test.
- I infer that the POF criterion will be most valuable in early iterations when the surrogate is least certain; after the descriptor space is well mapped, pure exploitation may become nearly as efficient, an effect visible in the shrinking error bars in the paper's Figure 2.
- I infer that the 12 oxides, which are rich in Mo, Ta, and W, suggest a design rule worth testing in the laboratory: acid stability at pH 0 and 1.23 V correlates with these oxophilic early-transition-metal frameworks, and the descriptor map could be used to propose new ternary compositions.
- I infer that the overconfidence remaining, with a mean z of 1.76, means the workflow's hit rate could be improved further by replacing the Gaussian cumulative-distribution assumption in POF with a heavier-tailed or nonparametric calibration, a modification the paper leaves open.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops an active-learning (AL) workflow based on ensembles of SISSO symbolic-regression models. To quantify prediction uncertainty, the authors compare three ensemble strategies (bagging, model-complexity bagging, and bagging with Monte-Carlo dropout of primary features) and find that the latter yields the lowest MAE and the most balanced miscalibration scores on held-out data. They then use this ensemble in an AL campaign with a probability-of-feasibility (POF) acquisition function to discover acid-stable oxides from a pool of 1470 candidates, using DFT-HSE06 as the ground-truth evaluator. In 30 AL iterations, they report identifying 12 acid-stable oxides, whereas random selection finds only 2. They also provide SISSO-derived descriptor maps of oxide stability and compare PBE and HSE06 predictions for the discovered materials.
Significance. If the findings are robust, this is a valuable demonstration of symbolic regression in closed-loop materials discovery, going beyond interpolation-based AL by offering interpretable descriptors and uncertainty estimates. The ensemble comparison is carefully done with 30 independent trials, and the public availability of code, data, and a tutorial strengthens reproducibility. The practical result—discovering acid-stable oxides that are missed by GGA-based screening—is of interest to the electrocatalysis community. However, the central efficiency claim rests on several choices that are not fully justified, notably the threshold τ used in the acquisition function and the lack of repeated AL campaigns; these need to be addressed to establish the quantitative advantage claimed.
major comments (3)
- [Results (POF definition)] Equation (1) defines the acquisition probability with τ = 0.00, but the paper labels an oxide as acid-stable only if ΔG_pbx^OER ≤ 0.1 eV/atom (stated in the paragraph after Eq. (1) and in Figure 2). Since σ_ESISSO varies across candidates, this threshold mismatch does not merely shift all POF values by a constant; it changes the ranking of candidates and therefore the set of materials acquired in each AL iteration. The reported efficiency gain (12 acid-stable oxides found by POF vs. 2 by random selection in 30 iterations) is thus contingent on an arbitrary threshold choice. The authors should either justify τ = 0, rerun the AL campaign with τ = 0.1, or at least show that the conclusions are insensitive to τ over a plausible range.
- [Results (AL campaigns)] The central claim that POF-guided AL identifies 12 acid-stable oxides in 30 iterations is based on a single AL campaign. The ensemble comparison in Figure 1 is repeated over 30 independent trials, but the AL results in Figure 2 are not; given the stochasticity of bootstrap sampling and Monte-Carlo feature dropout, a different random seed could yield a different trajectory and a different number of discoveries. The authors should report the distribution of outcomes over multiple campaigns (e.g., 5–10 seeds) or at least provide a sensitivity analysis with respect to the initial training set and the random seed.
- [Results (uncertainty calibration)] The ensemble method used for acquisition has a mean miscalibration score z = 1.76, indicating overconfident uncertainty estimates. Equation (1) treats σ_ESISSO as a Gaussian standard deviation, but the calibration analysis shows that POF values are not true probabilities. The paper acknowledges this limitation but does not test whether the AL efficiency actually derives from the uncertainty term. A useful control would be an acquisition strategy based only on the ensemble mean prediction (e.g., selecting the lowest predicted ΔG_pbx^OER) to separate the exploitation and exploration contributions; if that baseline performs comparably to POF, the uncertainty-driven exploration claim would need to be moderated.
minor comments (4)
- [Figure 2] The filled and open square markers are difficult to distinguish in grayscale; consider using different shapes or colors with accessible palettes.
- [Results (POF definition)] The acronym 'POF' is defined as 'probability of feasibility', but the term 'feasibility' is not standard for a stability threshold; consider 'probability of stability' or define the target event more explicitly.
- [Methods (SISSO)] In the operator set in Eq. (2), the division φ1/φ2 is included but there is no statement about how zero denominators are handled; please clarify this implementation detail.
- [Results (initial dataset)] The composition and selection criteria of the initial 250-oxide training set are deferred to the Supplementary Material; a brief statement in the main text would help the reader judge potential bias in the training distribution.
Circularity Check
No significant circularity: the SISSO-guided discovery claim is validated by independent DFT-HSE06 calculations, not by the fitted surrogate.
full rationale
The paper's derivation chain is self-contained against an external benchmark. SISSO is a surrogate fitted to DFT-HSE06 values of ΔG_pbx, and ensembles provide mean predictions and uncertainties used by the POF acquisition strategy. Crucially, the acid-stable materials reported are not taken from SISSO predictions; they are confirmed by direct DFT-HSE06 Pourbaix analysis, as stated: 'the selected oxides along with computed acid stability are added to the dataset' and Figure 2 compares SISSO predictions against DFT-HSE06 values. The central claim of 12 acid-stable oxides in 30 iterations is therefore an empirical outcome of the closed-loop workflow, not a consequence of fitting or self-citation. The SISSO-related citations are to published methods and code (e.g., refs. [13], [15], [39]) that are externally validated and not target-specific; they do not smuggle in the result. No equation in the paper reduces to its own inputs by construction. The noted mismatch between the POF threshold τ=0.00 and the acid-stability threshold ΔG≤0.1 eV/atom is a potential acquisition-objective alignment concern, but it is not circularity, because the final discovery count is determined by DFT calculations, not by the POF score. Overall, the paper's logic does not exhibit any of the seven circularity patterns, so the score is 0.
Assumptions & free parameters
free parameters (6)
- SISSO rung q =
2
- SISSO descriptor dimension D =
2
- Ensemble size k =
10
- Feature dropout fraction =
0.2 (20% retained)
- POF threshold tau =
0.00 eV/atom
- Acid-stability labeling threshold =
0.1 eV/atom
assumptions (4)
- domain assumption DFT-HSE06 Pourbaix decomposition free energy is an accurate proxy for acid stability under OER conditions (pH=0, 1.23 V).
- domain assumption The candidate space of 1470 oxides and the initial 250 training oxides are representative enough for the AL comparison.
- ad hoc to paper Prediction errors of the SISSO ensemble are Gaussian with mean and variance from the ensemble, as assumed in Eq. (1).
- domain assumption SISSO's selected descriptors generalize across the materials space (extrapolation via physical parameters).
Cite this review
Pith. "Pith review of Materials-Discovery Workflows Guided by Symbolic Regression: Identifying Acid-Stable Oxides for Electrocatalysis." pith.science (2026). https://pith.science/paper/TY7WHNPX
@misc{pith2026241205947,
author = {Pith},
title = {Pith review of: Materials-Discovery Workflows Guided by Symbolic Regression: Identifying Acid-Stable Oxides for Electrocatalysis},
year = {2026},
howpublished = {\url{https://pith.science/paper/TY7WHNPX}},
note = {Machine review of arXiv:2412.05947}
}
read the original abstract
The efficiency of active learning (AL) approaches to identify materials with desired properties relies on the knowledge of a few parameters describing the property. However, these parameters are unknown if the property is governed by a high intricacy of many atomistic processes. Here, we develop an AL workflow based on the sure-independence screening and sparsifying operator (SISSO) symbolic-regression approach. SISSO identifies the few, key parameters correlated with a given materials property via analytical expressions, out of many offered primary features. Crucially, we train ensembles of SISSO models in order to quantify mean predictions and their uncertainty, enabling the use of SISSO in AL. By combining bootstrap sampling to obtain training datasets with Monte-Carlo feature dropout, the high prediction errors observed by a single SISSO model are improved. Besides, the feature dropout procedure alleviates the overconfidence issues observed in the widely used bagging approach. We demonstrate the SISSO-guided AL workflow by identifying acid-stable oxides for water splitting using high-quality DFT-HSE06 calculations. From a pool of 1470 materials, 12 acid-stable materials are identified in only 30 AL iterations. The materials property maps provided by SISSO along with the uncertainty estimates reduce the risk of missing promising portions of the materials space that were overlooked in the initial, possibly biased dataset.
Figures
Forward citations
Cited by 1 Pith paper
-
Materials Database from All-electron Hybrid Functional DFT Calculations
An open database of 7,024 inorganic materials computed with all-electron HSE06 hybrid DFT, including stability metrics and a SISSO model for HSE06 band gaps.
Reference graph
Works this paper leans on
-
[1]
D. P. Tabor, L. M. Roch, S. K. Saikin, C. Kreisbeck, D. Sheberla, J. H. Montoya, S. Dwaraknath, M. Aykol, C. Ortiz, H. Tribukait, et al., Accelerating the discovery 7 of materials for clean energy in the era of smart automa- tion, Nature reviews materials 3, 5 (2018)
work page 2018
-
[2]
A. Ludwig, Discovery of new materials using combi- natorial synthesis and high-throughput characterization of thin-film materials libraries combined with compu- tational methods, npj Computational Materials 5, 70 (2019)
work page 2019
-
[3]
E. O. Pyzer-Knapp, J. W. Pitera, P. W. Staar, S. Takeda, T. Laino, D. P. Sanders, J. Sexton, J. R. Smith, and A. Curioni, Accelerating materials discovery using ar- tificial intelligence, high performance computing and robotics, npj Computational Materials 8, 84 (2022)
work page 2022
-
[4]
A. Merchant, S. Batzner, S. S. Schoenholz, M. Aykol, G. Cheon, and E. D. Cubuk, Scaling deep learning for materials discovery, Nature , 1 (2023)
work page 2023
-
[5]
J. H. Montoya, K. T. Winther, R. A. Flores, T. Bligaard, J. S. Hummelshøj, and M. Aykol, Autonomous intelligent agents for accelerated materials discovery, Chemical Sci- ence 11, 8517 (2020)
work page 2020
-
[6]
T. Lookman, P. V. Balachandran, D. Xue, and R. Yuan, Active learning in materials science with emphasis on adaptive sampling using uncertainties for targeted de- sign, npj Computational Materials 5, 21 (2019)
work page 2019
- [7]
- [8]
Show all 39 references
-
[9]
Qian, B.-J
X. Qian, B.-J. Yoon, R. Arr´ oyave, X. Qian, and E. R. Dougherty, Knowledge-driven learning, optimiza- tion, and experimental design under uncertainty for ma- terials discovery, Patterns 4 (2023)
2023
-
[10]
W. Ye, X. Lei, M. Aykol, and J. H. Montoya, Novel inor- ganic crystal structures predicted using autonomous sim- ulation agents, Scientific Data 9, 302 (2022)
2022
-
[11]
A. G. Kusne, H. Yu, C. Wu, H. Zhang, J. Hattrick- Simpers, B. DeCost, S. Sarker, C. Oses, C. Toher, S. Cur- tarolo, et al., On-the-fly closed-loop materials discovery via bayesian active learning, Nature communications 11, 5966 (2020)
2020
-
[12]
J. H. Montoya, C. Grimley, M. Aykol, C. Ophus, H. Sternlicht, B. H. Savitzky, A. M. Minor, S. B. Tor- risi, J. Goedjen, C.-C. Chung, et al., How the ai-assisted discovery and synthesis of a ternary oxide highlights ca- pability gaps in materials science, Chemical Science 15, 5...
2024
-
[13]
Ouyang, S
R. Ouyang, S. Curtarolo, E. Ahmetcik, M. Scheffler, and L. M. Ghiringhelli, Sisso: A compressed-sensing method for identifying the best low-dimensional descriptor in an immensity of offered candidates, Physical Review Mate- rials 2, 083802 (2018)
2018
-
[14]
Foppa, L
L. Foppa, L. M. Ghiringhelli, F. Girgsdies, M. Hasha- gen, P. Kube, M. H¨ avecker, S. J. Carey, A. Tarasov, P. Kraus, F. Rosowski, et al., Materials genes of hetero- geneous catalysis from clean experiments and artificial intelligence, MRS bulletin , 1 (2021)
2021
-
[15]
T. A. Purcell, M. Scheffler, and L. M. Ghiringhelli, Re- cent advances in the sisso method and their implemen- tation in the sisso++ code, The Journal of Chemical Physics 159 (2023)
2023
-
[16]
Y. Wang, N. Wagner, and J. M. Rondinelli, Symbolic regression in materials science, MRS Communications 9, 793 (2019)
2019
-
[17]
E. S. Muckley, J. E. Saal, B. Meredig, C. S. Roper, and J. H. Martin, Interpretable models for extrapolation in scientific machine learning, Digital Discovery 2, 1425 (2023)
2023
-
[18]
Pourbaix, Atlas of electrochemical equilibria in aque- ous solutions, NACE (1966)
M. Pourbaix, Atlas of electrochemical equilibria in aque- ous solutions, NACE (1966)
1966
-
[19]
J. Heyd, G. E. Scuseria, and M. Ernzerhof, Hybrid func- tionals based on a screened coulomb potential, The Jour- nal of chemical physics 118, 8207 (2003)
2003
-
[20]
Kokott, F
S. Kokott, F. Merz, Y. Yao, C. Carbogno, M. Rossi, V. Havu, M. Rampp, M. Scheffler, and V. Blum, Ef- ficient all-electron hybrid density functionals for atom- istic simulations beyond 10,000 atoms, arXiv preprint arXiv:2403.10343 (2024)
2024 arXiv
-
[21]
H. Li, Y. Lin, J. Duan, Q. Wen, Y. Liu, and T. Zhai, Stability of electrocatalytic oer: from principle to appli- cation, Chemical Society Reviews (2024)
2024
-
[22]
M. Ju, Y. Zhou, F. Dong, Z. Guo, J. Wang, and S. Yang, Two-dimensional oer catalysts: Is there a win-win so- lution for their activity and stability?, ACS Materials Letters 6, 3602 (2024)
2024
-
[23]
I. C. Man, H.-Y. Su, F. Calle-Vallejo, H. A. Hansen, J. I. Mart ´ ınez, N. G. Inoglu, J. Kitchin, T. F. Jaramillo, J. K. Nørskov, and J. Rossmeisl, Universality in oxygen evolu- tion electrocatalysis on oxide surfaces, ChemCatChem 3, 1159 (2011)
2011
-
[24]
Y. Xu, K. Fan, Y. Zou, H. Fu, M. Dong, Y. Dou, Y. Wang, S. Chen, H. Yin, M. Al-Mamun, et al., Ra- tional design of metal oxide catalysts for electrocatalytic water splitting, Nanoscale 13, 20324 (2021)
2021
-
[25]
S. Lu, L. M. Ghiringhelli, C. Carbogno, J. Wang, and M. Scheffler, On the uncertainty estimates of equivariant- neural-network-ensembles interatomic potentials, arXiv preprint arXiv:2309.00195 (2023)
2023 arXiv
-
[26]
Bauer, P
S. Bauer, P. Benner, T. Bereau, V. Blum, M. Boley, C. Carbogno, R. Catlow, G. Dehm, S. Eibl, R. Ernstor- fer, et al., Roadmap on data-centric materials science, Modelling and Simulation in Materials Science and En- gineering (2024)
2024
-
[27]
Pernot, Calibration in machine learning uncertainty quantification: beyond consistency to target adaptivity, APL Machine Learning 1 (2023)
P. Pernot, Calibration in machine learning uncertainty quantification: beyond consistency to target adaptivity, APL Machine Learning 1 (2023)
2023
-
[28]
Palmer, S
G. Palmer, S. Du, A. Politowicz, J. P. Emory, X. Yang, A. Gautam, G. Gupta, Z. Li, R. Jacobs, and D. Mor- gan, Calibration after bootstrap for accurate uncertainty quantification in regression models, npj Computational Materials 8, 115 (2022)
2022
-
[29]
M. Liu, A. Gopakumar, V. I. Hegde, J. He, and C. Wolverton, High-throughput hybrid-functional dft cal- culations of bandgaps and formation energies and multi- fidelity learning with uncertainty quantification, Physical Review Materials 8, 043803 (2024)
2024
-
[30]
Huang and J
L.-F. Huang and J. M. Rondinelli, Electrochemical phase diagrams for ti oxides from density functional calcula- tions, Physical Review B 92, 245126 (2015)
2015
-
[31]
A. K. Singh, L. Zhou, A. Shinde, S. K. Suram, J. H. Mon- toya, D. Winston, J. M. Gregoire, and K. A. Persson, Electrochemical stability of metastable materials, Chem- istry of Materials 29, 10159 (2017)
2017
-
[32]
H. Park, Y. Kim, S. Choi, and H. J. Kim, Data driven computational design of stable oxygen evolution catalysts by dft and machine learning: Promising electrocatalysts, Journal of Energy Chemistry 91, 645 (2024). 8
2024
-
[33]
Wang, Y.-R
Z. Wang, Y.-R. Zheng, I. Chorkendorff, and J. K. Nørskov, Acid-stable oxides for oxygen electrocatalysis, ACS Energy Letters 5, 2905 (2020)
2020
-
[34]
S. Back, K. Tran, and Z. W. Ulissi, Discovery of acid- stable oxygen evolution catalysts: high-throughput com- putational screening of equimolar bimetallic oxides, ACS applied materials & interfaces 12, 38256 (2020)
2020
-
[35]
G. K. K. Gunasooriya and J. K. Nørskov, Analysis of acid-stable and active oxides for the oxygen evolution reaction, ACS Energy Letters 5, 3778 (2020)
2020
-
[36]
V. Blum, R. Gehrke, F. Hanke, P. Havu, V. Havu, X. Ren, K. Reuter, and M. Scheffler, Ab initio molecular simulations with numeric atom-centered orbitals, Com- puter Physics Communications 180, 2175 (2009)
2009
-
[37]
Gjerding, T
M. Gjerding, T. Skovhus, A. Rasmussen, F. Bertoldo, A. H. Larsen, J. J. Mortensen, and K. S. Thygesen, Atomic simulation recipes: A python framework and li- brary for automated workflows, Computational Materials Science 199, 110731 (2021)
2021
-
[38]
J. J. Mortensen, M. Gjerding, and K. S. Thygesen, Myqueue: Task and workflow scheduling system, Jour- nal of Open Source Software 5, 1844 (2020)
2020
-
[39]
Scheffler, C
M. Scheffler, C. Carbogno, L. M. Ghiringhelli, et al., Sisso++: A c++ implementation of the sure- independence screening and sparsifying operator ap- proach, Journal of Open Source Software 7, 3960 (2022)
2022
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.