Pith. sign in

REVIEW 4 major objections 4 minor 72 references

Guided Synthesis of EMT Zeolites by Machine Learning

T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Trained on 174 in-house synthesis runs, the paper's machine-learning pipeline proposes six recipes for the zeolite framework EMT; five crystallize as pure EMT, including two with silicon-to-aluminum ratios below every value the model was tr

desk verdict Solid experimental core with five validated EMT recipes, but the headline extrapolation claim is contradicted by the paper's own Table S1. read the letter →

arxiv 2608.03760 v1 pith:IAVMGE6D submitted 2026-08-04 cond-mat.mtrl-sci

classification cond-mat.mtrl-sci
keywords EMTzeolitesynthesismachinelearningconditiondiscoveryOSDA-freefeatureimportancePCAlatentspaceexperimentalvalidation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether machine learning trained on a modest set of recorded synthesis attempts can guess which untried recipes will crystallize the zeolite framework EMT rather than its close relatives FAU or SOD. Using 174 in-house syntheses described by 13 parameters, the authors train classifiers — tree ensembles and a pretrained tabular foundation model perform about equally — to separate pure-EMT outcomes from mixed or failed ones. They then generate candidate recipes in a dimension-reduced (principal-component) version of the parameter space, screen them by predicted probability, physical plausibility, and novelty, and test the six most promising. Five of the six crystallize as pure EMT, an 83% hit rate against the 32% baseline of the original data, and two of the five use Si/Al stoichiometries below every value in the training set. The paper's claim is that this closed loop — predict, screen, synthesize, validate — transfers beyond the training domain and can be reused for other frameworks.

What carries the argument

The load-bearing mechanism is a two-stage pipeline. A proposer generates candidate synthesis conditions in a three-dimensional latent space from principal component analysis of the 13 recorded features, sampling 19 evenly spaced values per component from 30% below to 30% above the observed range, then inverts the transform; generating in latent space respects strong feature correlations, so candidates remain physically plausible. An evaluator screens them with three filters: physical soundness (no negative ratios), predicted pure-EMT probability above 0.85 from an XGBoost (gradient-boosted tree) binary classifier, and a dissimilarity score — minimum Euclidean distance to any training sample

What would settle it

Rerun each of the five successful recipes while deliberately varying only factors outside the 13 descriptors — stirring rate, gel-aging time, cooling rate, vessel geometry — and check whether pure EMT persists; any flip to hybrid or non-EMT shows the descriptor set does not determine the outcome. Alternatively, fix the pipeline's other features and synthesize at Si/Al = 1.5 and 2.0: if neither yields pure EMT, the two out-of-range successes are a narrow accident rather than evidence of extrapolation.

Watch

Extended reading notes

Core claim

A classifier trained on 174 organic-template-free EMT synthesis attempts, each described by 13 ordinary recipe features, proposes new crystallization conditions that yield pure EMT zeolite with 83% success (five of six tested recipes), a 2.6-fold improvement over the 32% experience-guided baseline. The discovery pipeline compresses the 13 features into three principal components, samples 6,859 candidates stretching 30% beyond the observed ranges, maps them back, discards physically invalid ones, keeps the 374 with predicted pure-EMT probability above 0.85, and retains the 28 whose minimum Euclidean distance to the training data exceeds 0.5. The six highest-probability candidates were synthes

Load-bearing premise

The load-bearing premise is that the 13 recorded descriptors — amounts, ratios, precursor sources, pre-dissolution flags, time, and temperature — determine the crystallization outcome, so any unrecorded factor such as mixing, gel aging, cooling rate, or thermal history that also shifts the phase would make the model's confident predictions fail to transfer, a fragility the paper itself acknowledges in noting that minor temperature or time variations can flip EMT to FAU or SOD

Editorial extensions

If this is right

  • EMT recipes can be proposed and screened computationally, and the top candidates validate at an 83% success rate — more than 2.5 times the 32% rate of the experience-guided runs that built the dataset.
  • The pipeline reaches compositions the training data never contained: pure EMT formed at Si/Al = 1.77 and 2.30, both below the dataset's lower bound of 2.5.
  • The model transfers to other laboratories' syntheses, correctly predicting 8 of 12 literature EMT cases inside the training envelope and 3 of 4 beyond it, with missing descriptors imputed from the training set.
  • The workflow — propose in a reduced latent space, screen by predicted probability and novelty, validate in the lab — is put forward as a general route for other zeolite frameworks and other crystalline materials with many coupled synthesis variables.
  • Because the trained classifier is cheap to run, the 22 high-probability candidates already listed in the supplement are immediately available for further experimental testing.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The single hybrid-EMT failure among the six tested candidates is itself informative: feeding that outcome back into the training set would test whether the 0.85 probability cutoff is well calibrated near the pure/hybrid boundary, something the paper does not do.
  • The two out-of-range successes suggest the pure-EMT window in Si/Al extends lower than the training data implied; probing Si/Al near 1.5–2.0 with the pipeline's other features fixed would map the true phase boundary and show whether the extrapolation is reliable or fortunate.
  • The misclassified literature case (72 °C, 24 h) hints that time–temperature interactions are the least transferable part of the descriptor set; a fixed-composition study sweeping time and temperature would reveal whether a missing descriptor (aging, heating rate, or mixing) is needed for cross-laboratory transfer.
  • The 0.5 dissimilarity threshold discarded 346 of 374 high-probability candidates, so the reported 83% is conditional on that specific novelty cut; relaxing it would trade hit rate for wider exploration, and the paper does not map that trade-off.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper trains machine-learning classifiers (XGBoost, Random Forest, CatBoost, TabPFN) on 174 in-house OSDA-free EMT synthesis attempts to predict pure-EMT, hybrid-EMT, or non-EMT outcomes. A binary classifier is trained on a 101-sample subset restricted to Si/Al in [2,10] and NaOH in [30,50]. The authors propose new synthesis conditions by sampling a PCA latent space, filtering by physical soundness, predicted EMT probability, and dissimilarity to the training set. Six candidates are selected, and five are experimentally confirmed as pure EMT by PXRD and SEM; one is hybrid-EMT. The paper claims that two successful conditions have Si/Al ratios (1.77 and 2.3) below the training data's lower bound of 2.5, thus demonstrating extrapolation. An external evaluation on 16 literature-reported EMT syntheses yields 11/16 correct predictions. The central experimental result—5/6 pure EMT—is encouraging, but the extrapolation claim is contradicted by the paper's own Supplemental Table S1.

Significance. If the extrapolation claim were supported, this would be a useful demonstration of model-guided zeolite synthesis beyond the training domain, with practical value for EMT and potentially other frameworks. The paper's strengths include direct experimental validation with PXRD and SEM, reproducible code/data availability, comparison of four ML methods, and feature-importance analysis that aligns with known synthesis trends. However, the central claim of out-of-range Si/Al prediction is undermined by an internal inconsistency in the reported training range, and the reported 83% success rate is a conditional precision under the model's own high-confidence threshold rather than an unbiased generalization estimate. The paper's significance is therefore currently overstated; the underlying experimental result remains valuable if the claims are corrected.

major comments (4)
  1. [Section III.C and Table S1] The paper's key extrapolation claim is contradicted by its own Supplemental Table S1. Section III.C states that two successful conditions have Si/Al ratios 1.77 and 2.3, falling below the dataset's lower bound of 2.5, and the Introduction repeats this claim. However, Table S1 reports the in-house dataset's Si/Al range as 0.5–25.8. If this table is correct, both 1.77 and 2.3 are inside the training range, and the extrapolation claim is false. Moreover, the binary classifier used for screening was trained on the Section III.B filtered subset with 2≤Si/Al≤10, so its training lower bound is 2, not 2.5; under that definition, Si/Al=2.3 is also in-domain. The authors must either reconcile Table S1 with the text or remove the extrapolation claim.
  2. [Section III.C, Table IV, and Fig. 7] The 83% success rate is not an unbiased estimate of the pipeline's generalization. The six candidates were selected because the model assigned them probabilities of 0.86–0.93; the success rate is therefore conditional on high model confidence. Comparing this to the 32% prevalence of pure-EMT in the full dataset is not a valid baseline, because the six candidates are not a random sample from the same distribution. The paper should report precision at the chosen threshold with confidence intervals, or compare against candidates drawn randomly from the same PCA/search region but not selected by the classifier, to substantiate the claimed 2.6-fold improvement.
  3. [Section III.C and Table S1] All six experimentally validated candidates are nearly identical to each other and closely resemble the canonical OSDA-free EMT literature route summarized in Table S1: Na2SiO3 and NaAlO2 as precursors, pre-dissolved Si and Al, H2O=450, temperature 40–42°C, time ~72 h. This is a narrow local region of the synthesis space, not broad exploration. The claim that the pipeline 'explores the synthesis space and identifies six promising new conditions' should be tempered to acknowledge that the successful conditions are minor variations around a known EMT-favorable recipe, and the dissimilarity scores in Table IV (0.67–4.98) are modest in absolute terms.
  4. [Section II.A and Introduction] The model assumes that the 13 recorded descriptors are sufficient to determine crystallization outcome. The paper itself notes in the Introduction that minor variations in temperature or reaction time can cause EMT-to-FAU or EMT-to-SOD transitions. Since each of the six proposed conditions was validated in only one synthesis run, with no replicates and no report of uncontrolled variables such as mixing, gel aging, cooling rate, or operator-specific details, the five successes could be partly due to luck or unrecorded factors. Reporting repeated syntheses for at least the claimed extrapolation cases would strengthen the conclusion.
minor comments (4)
  1. [Section III.C] Typographical issue: 'Altogether, 6859 new candidate synthesis conditions were created... Finally, each candidate... via the inverse PCA transform' is clear, but later 'the PCA-generated populate' should be 'the PCA-generated pool' or 'population.'
  2. [Supplemental Material, Section S1] The text reads 'Vmicore (cm3/g)'—this appears to be a typo for 'Vmicro' (micropore volume).
  3. [Supplemental Material, Section S9] In the introductory sentence, 'Section IIID of the main manuscript' should be 'Section III.D'.
  4. [Table S1 and Section III.B] The in-house dataset ranges in Table S1 (e.g., Si/Al 0.5–25.8, H2O 173–650) are not consistently referenced in the main text, which contributes to the confusion about the training domain. A concise statement of which ranges were used for the binary model's training set would improve clarity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: experimental validation and external literature checks are independent of model outputs; the extrapolation concern is an internal factual inconsistency, not a circular derivation.

full rationale

The paper's derivation chain is: assemble a 174-sample in-house dataset (Sec. II.A), train ML classifiers (Sec. II.B), inspect feature importance (Sec. III.A), build a focused binary model on a 101-sample filtered subset (Sec. III.B), generate and screen candidates via PCA latent-space sampling plus XGBoost probability and dissimilarity filters (Sec. III.C), then experimentally validate six top-scoring candidates, and finally benchmark on 16 literature-reported EMT syntheses (Sec. III.D). The experimentally observed outcomes are new data not used in training, so the 5/6 success rate at high model probability is not forced by construction. The literature validation is explicitly disclosed by the authors as partially imputed (SM S9: 'Because the imputed values are drawn from the in-house training dataset... rather than a fully independent test across all model inputs'), which is an honest limitation, not circular reasoning. Feature-importance statements are descriptive summaries of the trained model, not derivations. No load-bearing self-citation chain is present: the Wen/Adhikari references are background on ML uncertainty, and the Rimer-related synthesis citations are external literature anchors, not arguments that assume the paper's conclusion. The main concern raised by the reader—that the claimed Si/Al extrapolation (1.77 and 2.3) conflicts with Table S1's stated in-house Si/Al range of 0.5–25.8—is a real internal inconsistency affecting the strength of the extrapolation claim, but it is a factual correctness/methodological reporting issue, not a case of a prediction reducing to its input by definition or by fitted parameters. Likewise, the 83% success rate is a conditional precision estimate at a high-confidence screening threshold rather than an unbiased generalization estimate; this is a selection-bias caveat, not circularity. The paper is therefore not circular in its reasoning, though its headline extrapolation claim warrants closer factual scrutiny.

Assumptions & free parameters 8 free parameters · 8 assumptions · 0 invented entities

The central claim rests on the representativeness of 174 in-house experiments and on a series of hand-chosen thresholds (Si/Al and NaOH bounds, 0.85 probability, 0.5 dissimilarity, 30% PCA expansion). No new physical entities are postulated. The ML model's hyperparameters are underreported, and the literature-validation imputation is acknowledged to bias inputs toward the in-house domain.

free parameters (8)
  • Si/Al filter bounds = [2, 10]
    Used to define the pure-EMT region and construct the binary training set (Section III.B, Figure 4).
  • NaOH filter bounds = [30, 50]
    Same filter used to exclude 73 samples from the binary task (Section III.B, Figure 4).
  • PCA expansion factor = 30% beyond observed min/max; 19 samples per principal component; 3 components
    The proposer's sampling parameters are chosen by hand (Section III.C).
  • Probability threshold = 0.85
    Candidates with predicted pure-EMT probability above 0.85 retained (Section III.C).
  • Dissimilarity threshold = 0.5
    Minimum Euclidean distance to the training set for candidate retention (Section III.C).
  • Categorical rounding thresholds = 0.6 for Si-source one-hot difference; 0.9 for binary features
    Post-processing of inverse-PCA candidate feature vectors (Section S6).
  • XGBoost hyperparameters = not reported
    Tuned on the validation split but exact values are absent from the paper and SM.
  • Missing-value imputation means = in-house feature means (e.g., Si size 3.25 nm)
    Mean imputation for the cleaned dataset and for literature descriptors (Sections S1, S9).
assumptions (8)
  • domain assumption The 13 recorded synthesis descriptors are sufficient to determine the crystallization outcome (pure, hybrid, or non-EMT).
    The whole model maps only these features to labels; unmeasured factors such as mixing, gel aging, and thermal history are assumed irrelevant (Section II.A, Table I).
  • domain assumption The 174 in-house experiments are representative of the OSDA-free EMT synthesis space.
    The dataset was built around canonical EMT routes with deliberate broadening (Section II.A); representativeness is not independently established.
  • domain assumption PXRD-based labeling of pure-EMT, hybrid-EMT, and non-EMT outcomes is correct for every training sample.
    Labels are taken from characterization (Section S8); no independent re-verification of the full 174-sample label set is provided.
  • domain assumption Mean imputation of missing literature descriptors does not systematically bias the external validation.
    Section S9 acknowledges that imputation from in-house means may place literature cases closer to the in-house domain.
  • domain assumption Inverse PCA mapping produces chemically realistic precursor combinations after filtering.
    The proposer relies on the 3-D PCA manifold of the training data; fidelity of inverse transforms for extrapolated points is unverified (Section III.C).
  • ad hoc to paper The hand-chosen bounds Si/Al in [2,10] and NaOH in [30,50] correctly delineate the pure-EMT region.
    These bounds are derived from the in-house data (Figure 4) and then used to exclude samples and define the binary task.
  • domain assumption Model probability scores are sufficiently calibrated for thresholding at 0.85.
    No calibration analysis is reported; the threshold is used to select 374 of 3375 candidates (Section III.C).
  • standard math Principal component analysis and tree-ensemble methods behave as standard tools.
    Background methods are used without proofs (Sections II.B and III.C).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Guided Synthesis of EMT Zeolites by Machine Learning." pith.science (2026). https://pith.science/paper/IAVMGE6D

@misc{pith2026260803760,
  author       = {Pith},
  title        = {Pith review of: Guided Synthesis of EMT Zeolites by Machine Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IAVMGE6D}},
  note         = {Machine review of arXiv:2608.03760}
}
read the original abstract

Zeolites are microporous crystalline materials with diverse frameworks, widely used in industrial applications such as petroleum refining and molecular separation. Unlike most zeolites, EMT can be synthesized under mild conditions (at low temperatures and without the use of organic structure-directing agents), making it attractive for cost-effective and environmentally sustainable production. However, the specific synthesis conditions that selectively produce EMT rather than similar frameworks like FAU are not yet well established. In this work, we develop machine learning (ML) models to guide the discovery of synthesis conditions for EMT zeolites. Our dataset comprises 174 experimental synthesis attempts, recording reaction time, temperature, silica and alumina sources, Si/Al stoichiometric ratio, and other synthesis parameters. We apply both classical ML methods and pretrained foundation models to predict zeolite framework outcomes from these synthesis parameters. Feature importance analysis identifies critical parameters for EMT formation, validating known synthesis principles. Leveraging the ML models, we explore the synthesis space and identify six promising new conditions for EMT formation. Experimental validation confirms EMT crystallization in five cases, including two with Si/Al stoichiometric ratios outside the training dataset's range. Evaluation on independent literature-reported synthesis conditions further demonstrates the generalizability of the model. This work demonstrates a data-driven approach to accelerating zeolite synthesis, closing the loop between ML prediction and experimental validation.

Figures

Figures reproduced from arXiv: 2608.03760 by the authors.

Figure 1
Figure 1. FIG. 1. Feature importance analysis using SHAP for the [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 3
Figure 3. FIG. 3. Confusion matrix for the XGBoost ternary classifica [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 2
Figure 2. FIG. 2. Distribution of synthesis outcome labels with respect [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: FIG. 4. Distribution of binary synthesis outcome labels with respect to Si/Al stoichiometric ratio, NaOH molar amount, and [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: FIG. 5. Feature importance analysis using SHAP for the bi [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: FIG. 6. A two-stage pipeline for discovering new synthesis conditions for EMT zeolites. (a) New candidate synthesis conditions [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: FIG. 7. Distribution of the six experimentally verified synthe [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: FIG. 8. Model predictions for literature-reported EMT syn [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

72 extracted references · 69 canonical work pages

  1. [1]

    M. E. Davis, Ordered porous materials for emerging ap- plications, Nature417, 813 (2002)

  2. [2]

    Primo and H

    A. Primo and H. Garcia, Zeolites as catalysts in oil re- fining, Chemical Society Reviews43, 7548 (2014)

  3. [3]

    Mintova, M

    S. Mintova, M. Jaber, and V. Valtchev, Nanosized micro- porous crystals: emerging applications, Chemical society reviews44, 7207 (2015)

  4. [4]

    Y. Li, L. Li, and J. Yu, Applications of zeolites in sus- tainable chemistry, Chem3, 928 (2017)

  5. [5]

    J. Li, M. Gao, W. Yan, and J. Yu, Regulation of the Si/Al ratios and Al distributions of zeolites and their impact on properties, Chemical Science14, 1935 (2023)

  6. [6]

    Ma and Z.-P

    S. Ma and Z.-P. Liu, Machine learning potential era of zeolite simulation, Chemical Science13, 5055 (2022)

  7. [7]

    R. E. Fletcher, S. A. Wells, K. M. Leung, P. P. Edwards, and A. Sartbaeva, Intrinsic flexibility of porous materials; theory, modelling and the flexibility window of the EMT zeolite framework, Structural Science71, 641 (2015)

  8. [8]

    Ghojavand, E

    S. Ghojavand, E. Dib, and S. Mintova, Flexibility in zeo- lites: origin, limits, and evaluation, Chemical Science14, 12430 (2023)

Show all 72 references
  1. [9]

    Ohsuna, B

    T. Ohsuna, B. Slater, F. Gao, J. Yu, Y. Sakamoto, G. Zhu, O. Terasaki, D. E. Vaughan, S. Qiu, and C. R. A. Catlow, Fine structures of zeolite-linde-l (ltl): surface structures, growth unit and defects, Chemistry–A Eu- ropean Journal10, 5031 (2004)

  2. [10]

    The model cor- rectly predicts 3 of 4 EMT cases (75%) in this region

    This is considered an extrapolation regime, as the median dissimilarity score of these cases is ˜d= 39.19, which is approximately 14 times larger than that of the literature cases in the in-house region. The model cor- rectly predicts 3 of 4 EMT cases (75%) in this region. The...

  3. [11]

    S. Yang, M. Lach-hab, E. Blaisten-Barojas, X. Li, and V. L. Karen, Machine learning study of the heulandite family of zeolites, Microporous and Mesoporous Materi- als130, 309 (2010)

  4. [12]

    E. Dib, E. B. Clatworthy, A. A. Paecklar, J. Grand, N. Barrier, and S. Mintova, Role of al distribution in co2 adsorption capacity in rho zeolites, The Journal of Physical Chemistry C127, 3 (2022)

  5. [14]

    Maldonado, M

    M. Maldonado, M. D. Oleksiak, S. Chinta, and J. D. Rimer, Controlling crystal polymorphism in organic-free synthesis of na-zeolites, Journal of the American Chemi- cal Society135, 2641 (2013)

  6. [15]

    Morin, A

    S. Morin, A. Berreghis, P. Ayrault, N. Gnep, and M. Guisnet, Dealumination of zeolites part viiiacidity and catalytic properties of hemt zeolites dealuminated by steaming, Journal of the Chemical Society, Faraday Transactions93, 3269 (1997)

  7. [16]

    Moliner, F

    M. Moliner, F. Rey, and A. Corma, Towards the ratio- nal design of efficient organic structure-directing agents for zeolite synthesis, Angewandte Chemie International Edition52, 13880 (2013)

  8. [17]

    M. D. Oleksiak and J. D. Rimer, Synthesis of zeolites in the absence of organic structure-directing agents: factors governing crystal selection and polymorphism, Reviews in Chemical Engineering30, 1 (2014)

  9. [18]

    Muraoka, Y

    K. Muraoka, Y. Sada, D. Miyazaki, W. Chaikittisilp, and T. Okubo, Linking synthesis and structure descriptors from a large collection of synthetic records of zeolite ma- terials, Nature Communications10, 4459 (2019)

  10. [19]

    Slater, J

    B. Slater, J. Titiloye, F. Higgins, and S. Parker, Atomistic simulation of zeolite surfaces, Current Opinion in Solid State and Materials Science5, 417 (2001)

  11. [20]

    G. M. Psofogiannakis, J. F. McCleerey, E. Jaramillo, and A. C. Van Duin, Reaxff reactive molecular dynamics sim- ulation of the hydration of cu-ssz-13 zeolite and the for- mation of cu dimers, The Journal of Physical Chemistry C119, 6678 (2015)

  12. [21]

    Tesson, W

    S. Tesson, W. Louisfrema, M. Salanne, A. Boutin, E. Fer- rage, B. Rotenberg, and V. Marry, Classical polarizable force field to study hydrated charged clays and zeolites, The Journal of Physical Chemistry C122, 24690 (2018)

  13. [22]

    Misturini, F

    A. Misturini, F. Rey, and G. Sastre, Molecular sim- ulation of biobutanol recovery using lta and cha zeo- lite nanosheets with an external surface, The Journal of Physical Chemistry C126, 17680 (2022)

  14. [23]

    Filippone, F

    F. Filippone, F. Buda, S. Iarlori, G. Moretti, and P. Porta, Structural and electronic properties of sodalite: An ab initio molecular dynamics study, The Journal of Physical Chemistry99, 12883 (1995)

  15. [24]

    Astala, S

    R. Astala, S. M. Auerbach, and P. Monson, Density func- tional theory study of silica zeolite structures: Stabilities and mechanical properties of sod, lta, cha, mor, and mfi, The Journal of Physical Chemistry B108, 9208 (2004)

  16. [25]

    H. Shi, A. N. Migues, and S. M. Auerbach, Ab initio and classical simulations of the temperature dependence of zeolite pore sizes, Green Chemistry16, 875 (2014). 10

  17. [26]

    J. Rey, P. Raybaud, and C. Chizallet, Ab initio simu- lation of the acid sites at the external surface of zeolite beta, ChemCatChem9, 2176 (2017)

  18. [27]

    Kim and K

    N. Kim and K. Min, Accelerated discovery of zeolite structures with superior mechanical properties via active learning, The Journal of Physical Chemistry Letters12, 2334 (2021)

  19. [28]

    Ducamp and F.-X

    M. Ducamp and F.-X. Coudert, Prediction of Thermal Properties of Zeolites through Machine Learning, The Journal of Physical Chemistry C126, 1651 (2022)

  20. [29]

    Gaillac, S

    R. Gaillac, S. Chibani, and F.-X. Coudert, Speeding Up Discovery of Auxetic Zeolite Frameworks by Machine Learning, Chemistry of Materials32, 2653 (2020)

  21. [30]

    M. Wu, S. Zhang, and J. Ren, Ai-empowered digital de- sign of zeolites: Progress, challenges, and perspectives, APL Materials13, 10.1063/5.0253847 (2025)

  22. [31]

    Y. Yan, J. Li, M. Qi, X. Zhang, J. Yu, and R. Xu, Database of open-framework aluminophosphate synthe- ses: introduction and application (i), Science in China Series B: Chemistry52, 1734 (2009)

  23. [32]

    Jensen, E

    Z. Jensen, E. Kim, S. Kwon, T. Z. Gani, Y. Rom´ an- Leshkov, M. Moliner, A. Corma, and E. Olivetti, A ma- chine learning approach to zeolite synthesis enabled by automatic literature data extraction, ACS Central Sci- ence5, 892 (2019)

  24. [33]

    E. Pan, S. Kwon, Z. Jensen, M. Xie, R. G´ omez- Bombarelli, M. Moliner, Y. Rom´ an-Leshkov, and E. Olivetti, ZeoSyn: A Comprehensive Zeolite Synthe- sis Dataset Enabling Machine-Learning Rationalization of Hydrothermal Parameters, ACS Central Science10, 729 (2024)

  25. [34]

    J. D. Evans and F.-X. Coudert, Predicting the mechani- cal properties of zeolite frameworks by machine learning, Chemistry of Materials29, 7833 (2017)

  26. [35]

    Moliner, Y

    M. Moliner, Y. Rom´ an-Leshkov, and A. Corma, Machine learning applied to zeolite synthesis: The missing link for realizing high-throughput discovery, Accounts of Chemi- cal Research52, 2971 (2019)

  27. [36]

    Conroy, R

    B. Conroy, R. Nayak, A. L. R. Hidalgo, and G. J. Millar, Evaluation and application of machine learning princi- ples to zeolite lta synthesis, Microporous and Mesoporous Materials335, 111802 (2022)

  28. [37]

    Schwalbe-Koda, D

    D. Schwalbe-Koda, D. E. Widdowson, T. A. Pham, and V. A. Kurlin, Inorganic synthesis-structure maps in ze- olites with machine learning and crystallographic dis- tances, Digital Discovery2, 1911 (2023)

  29. [38]

    D. A. Carr, M. Lach-hab, S. Yang, I. I. Vaisman, and E. Blaisten-Barojas, Machine learning approach for structure-based zeolite classification, Microporous and Mesoporous Materials117, 339 (2009)

  30. [39]

    S. Yang, M. Lach-hab, I. I. Vaisman, and E. Blaisten- Barojas, Identifying zeolite frameworks with a machine learning approach, The Journal of Physical Chemistry C 113, 21721 (2009)

  31. [40]

    Raman, Forecasting low framework density zeolites from synthesis descriptors using machine learning, Jour- nal of Solid State Chemistry327, 124290 (2023)

    G. Raman, Forecasting low framework density zeolites from synthesis descriptors using machine learning, Jour- nal of Solid State Chemistry327, 124290 (2023)

  32. [41]

    X. Li, H. Han, N. Evangelou, N. J. Wichrowski, P. Lu, W. Xu, S.-J. Hwang, W. Zhao, C. Song, X. Guo,et al., Machine learning-assisted crystal engineering of a zeolite, Nature communications14, 3152 (2023)

  33. [42]

    X. Peng, R. Pan, X. Li, W. Zhong, and F. Qian, Molecu- lar descriptor-assisted interpretable machine learning: A scheme for guiding the synthesis of zeolites with target structures, Chemical Engineering Science308, 121378 (2025)

  34. [43]

    A. J. Cohen, P. Mori-S´ anchez, and W. Yang, Insights into current limitations of density functional theory, Science 321, 792 (2008)

  35. [44]

    M. Wen, S. M. Blau, X. Xie, S. Dwaraknath, and K. A. Persson, Improving machine learning performance on small chemical reaction data with unsupervised con- trastive pretraining, Chemical Science13, 1446–1458 (2022)

  36. [45]

    J. Dai, S. Adhikari, and M. Wen, Uncertainty quantifica- tion and propagation in atomistic machine learning, Re- views in Chemical Engineering 10.1515/revce-2024-0028 (2024)

  37. [46]

    Gandhi and M

    A. Gandhi and M. F. Hasan, Machine learning for the design and discovery of zeolites and porous crystalline materials, Current Opinion in Chemical Engineering35, 100739 (2022)

  38. [47]

    Hastie, R

    T. Hastie, R. Tibshirani, and J. Friedman,The elements of statistical learning: data mining, inference, and pre- diction(Springer, New York, NY, USA, 2009)

  39. [48]

    Hollmann, S

    N. Hollmann, S. M¨ uller, L. Purucker, A. Krishnakumar, M. K¨ orfer, S. B. Hoo, R. T. Schirrmeister, and F. Hut- ter, Accurate predictions on small data with a tabular foundation model, Nature637, 319 (2025)

  40. [49]

    E.-P. Ng, H. Awala, K.-H. Tan, F. Adam, R. Retoux, and S. Mintova, EMT-type zeolite nanocrystals synthesized from rice husk, Microporous and Mesoporous Materials 204, 204 (2015)

  41. [50]

    See supplemental material at [url will be inserted by publisher] for data processing procedures for the ma- chine learning models, additional results, and experimen- tal synthesis protocols (2026)

  42. [51]

    Breiman, Random forests, Machine Learning45, 5 (2001)

    L. Breiman, Random forests, Machine Learning45, 5 (2001)

  43. [52]

    Chen and C

    T. Chen and C. Guestrin, Xgboost: A scalable tree boost- ing system, inProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining(2016) pp. 785–794

  44. [53]

    Prokhorenkova, G

    L. Prokhorenkova, G. Gusev, A. Vorobev, A. V. Doro- gush, and A. Gulin, Catboost: unbiased boosting with categorical features, Advances in neural information pro- cessing systems31, 10.48550/arXiv.1706.09516 (2018)

  45. [54]

    S. M. Lundberg and S.-I. Lee, A unified approach to inter- preting model predictions, Advances in neural informa- tion processing systems30, 10.48550/arXiv.1705.07874 (2017)

  46. [55]

    Wendelbo, M

    R. Wendelbo, M. St¨ ocker, H. Junggreen, H. B. Mostad, and D. E. Akporiaye, Tumbling approach towards template-free synthesis of emt zeolite, Microporous and Mesoporous Materials28, 361 (1999)

  47. [56]

    Wu and K.-j

    C.-n. Wu and K.-j. Chao, Synthesis of faujasite zeolites with crown-ether templates, Journal of the Chemical So- ciety, Faraday Transactions91, 167 (1995)

  48. [58]

    C” in the “Type

    V. Georgieva, A. Vicente, C. Fernandez, R. Retoux, A. Palˇ ci´ c, V. Valtchev, and S. Mintova, Control of Na-EMT zeolite synthesis by organic additives, Crystal Growth & Design15, 1898 (2015). Supplemental Material: Guided Synthesis of EMT Zeolites by Machine Learning Emmanuel...

  49. [59]

    A similar rounding procedure was applied to all remaining binary categorical features with values close to 1 (e.g., 0.9) assigned to 1 and all other values set to 0

    Candidates that did not meet this criterion (e.g., cases with values of 0.6 and 0.4 for two silicon sources) were discarded. A similar rounding procedure was applied to all remaining binary categorical features with values close to 1 (e.g., 0.9) assigned to 1 and all other val...

  50. [60]

    dissimilarity index

    Finally, after applying these manual adjustments, we recomputed the predicted proba- bilities and dissimilarity scores following the procedure described in the main text. Finally, candidates with a minimum “dissimilarity index” of 0.5 were retained. From the resulting 28 candi...

  51. [61]

    Solution A was prepared by dissolving sodium aluminate and sodium hydroxide in deionized (DI) water in a polypropylene bottle, with stirring in an ice bath until the mixture was homogeneous

  52. [62]

    Sodium silicate was used as the nominal silica source unless otherwise stated

    Solution B was prepared by adding one of the silica sources (sodium silicate, fumed silica, or tetraethyl orthosilicate (TEOS)) to DI water in a separate polypropylene bottle. Sodium silicate was used as the nominal silica source unless otherwise stated

  53. [63]

    The resulting growth mixture was stirred in the ice bath for an additional 10 minutes

    Solution B was stirred in an ice bath until homogeneous, then added to Solution A. The resulting growth mixture was stirred in the ice bath for an additional 10 minutes

  54. [64]

    The mixture was heated under static conditions in a Thermo Fisher Precision oven at a temperature between 35 and 45 ◦C

  55. [65]

    The bottles were removed from the oven at various time intervals between 3 and 72 hr 9 and cooled to 25 ◦C in a water bath

  56. [66]

    The solid product was recovered by three cycles of centrifugation (6000 rpm for 5 min- utes using a Corning LSE centrifuge) and washing with DI water

  57. [67]

    Prediction

    The final product was dried in air at room temperature prior to characterization. S8. CHARACTERIZA TION OF ML-GUIDED EMT SYNTHESES a. X-ray diffraction.Powder X-ray diffraction (PXRD) patterns of dried solids were collected on a Rigaku SmartLab diffractometer with a Cu Kαsourc...

  58. [68]

    E.-P. Ng, D. Chateigner, T. Bein, V. Valtchev, and S. Mintova, Capturing ultrasmall emt zeolite from template-free systems, Science335, 70 (2012)

  59. [69]

    E.-P. Ng, H. Awala, K.-H. Tan, F. Adam, R. Retoux, and S. Mintova, EMT-type zeolite nanocrystals synthesized from rice husk, Microporous and Mesoporous Materials204, 204 (2015)

  60. [70]

    Maldonado, M

    M. Maldonado, M. D. Oleksiak, S. Chinta, and J. D. Rimer, Controlling crystal polymorphism in organic-free synthesis of na-zeolites, Journal of the American Chemical Society135, 2641 (2013)

  61. [71]

    Wendelbo, M

    R. Wendelbo, M. St¨ ocker, H. Junggreen, H. B. Mostad, and D. E. Akporiaye, Tumbling ap- proach towards template-free synthesis of emt zeolite, Microporous and Mesoporous Materials 28, 361 (1999)

  62. [72]

    Wu and K.-j

    C.-n. Wu and K.-j. Chao, Synthesis of faujasite zeolites with crown-ether templates, Journal of the Chemical Society, Faraday Transactions91, 167 (1995)

  63. [73]

    Dougnier and J.-L

    F. Dougnier and J.-L. Guth, EMT zeolite synthesis: Na+ vs. OH–effect, Journal of the Chemical Society, Chemical Communications , 951 (1995)

  64. [74]

    Georgieva, A

    V. Georgieva, A. Vicente, C. Fernandez, R. Retoux, A. Palˇ ci´ c, V. Valtchev, and S. Mintova, Control of Na-EMT zeolite synthesis by organic additives, Crystal Growth & Design15, 1898 (2015). 14

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.