{"id":"c4ec380d-57f0-4b06-884c-c21c223c2627","arxiv_id":"2608.03760","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Machine learning proposed six new EMT zeolite synthesis conditions; five produced pure EMT in the lab, including two extrapolations below the training Si/Al range.","lead":"An ML pipeline trained on 174 experimental EMT zeolite synthesis attempts proposed six new recipes, five of which produced pure EMT zeolite in follow-up experiments. The study closes the loop between machine learning prediction and lab validation for zeolite synthesis, including two recipes that extrapolate beyond the Si/Al ratios seen in the training data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Training Si/Al range contradicts extrapolation claim: Table S1 lists 0.5–25.8, so Si/Al 1.77 and 2.3 may be in-domain, not extrapolation.","rationale":"The reader's verdict is CONDITIONAL, and I agree that the paper requires corrections before the headline claims are taken at face value. However, the most load-bearing concern is not the descriptor-sufficiency issue that the reader identified as the weakest assumption. The descriptor-sufficiency argument is plausible but somewhat speculative—it is not directly contradicted by evidence in the paper and applies to any ML-driven synthesis study. The Si/Al range inconsistency, in contrast, is an internal contradiction that directly undermines the paper's central novelty: the claim of extrapolation beyond the training stoichiometry. The abstract, Introduction, and Section III.C all assert that all training Si/Al ratios exceeded 2.5, yet Table S1 reports an in-house range starting at 0.5. If Table S1 is correct, the two 'extrapolated' successes at 1.77 and 2.3 are within the training distribution, and the extrapolation claim is unsupported. If the text is correct, Table S1 is wrong, which still signals a data-integrity problem. This can be settled by inspecting the publicly available dataset. The five experimental confirmations are valuable evidence, but they do not by themselves establish extrapolation; they establish that the model can identify near-literature conditions with high probability. Therefore the verdict should remain CONDITIONAL, with the specific condition that the Si/Al range inconsistency be resolved and the extrapolation claim be either verified against the actual data or retracted.","tokens_in":20012,"tokens_out":3252,"duration_ms":40424,"concrete_test":"Download the cleaned dataset from the provided GitHub repository and compute the minimum and full distribution of the Si/Al feature across the 174 samples. If any training sample has Si/Al ≤ 2.3, the claim 'all synthesis mixture Si/Al ratios in the training data exceeded 2.5' is false, and the extrapolation claim must be retracted. Also verify the Si/Al values of the two claimed extrapolated candidates (1.77 and 2.3) against the actual training distribution. If the training range indeed includes these values, the paper should be revised to remove or substantially weaken the extrapolation claim.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central novelty claim is that the ML pipeline extrapolates beyond the training data, with two successful conditions at Si/Al 1.77 and 2.3 lying below the asserted training lower bound of 2.5 (Section III.C; also Abstract and Introduction). However, Table S1 in the Supplemental Material states the in-house dataset's Si/Al range as 0.5–25.8. If any training sample has Si/Al below 2.5—or below 2.3—the extrapolation claim is falsified by the paper's own data table. This is not a cosmetic discrepancy: the 'nontrivial and noteworthy achievement' hinges on these two points being genuinely out-of-distribution. The six proposed conditions are also tightly clustered near canonical literature recipes (all use Na2SiO3/NaAlO2, Pd:Si/Pd:Al = yes, H2O = 450, T ≈ 40–42°C, t ≈ 72 h), so even the in-domain successes may reflect interpolation in a known EMT-favorable regime rather than discovery. The reader's descriptor-sufficiency concern is real but secondary; the Si/Al inconsistency is concrete, checkable, and directly attacks the strongest claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains machine-learning classifiers (XGBoost, Random Forest, CatBoost, TabPFN) on 174 in-house OSDA-free EMT synthesis attempts to predict pure-EMT, hybrid-EMT, or non-EMT outcomes. A binary classifier is trained on a 101-sample subset restricted to Si/Al in [2,10] and NaOH in [30,50]. The authors propose new synthesis conditions by sampling a PCA latent space, filtering by physical soundness, predicted EMT probability, and dissimilarity to the training set. Six candidates are selected, and five are experimentally confirmed as pure EMT by PXRD and SEM; one is hybrid-EMT. The paper claims that two successful conditions have Si/Al ratios (1.77 and 2.3) below the training data's lower bound of 2.5, thus demonstrating extrapolation. An external evaluation on 16 literature-reported EMT syntheses yields 11/16 correct predictions. The central experimental result—5/6 pure EMT—is encouraging, but the extrapolation claim is contradicted by the paper's own Supplemental Table S1.","tokens_in":20414,"tokens_out":4791,"duration_ms":61003,"significance":"If the extrapolation claim were supported, this would be a useful demonstration of model-guided zeolite synthesis beyond the training domain, with practical value for EMT and potentially other frameworks. The paper's strengths include direct experimental validation with PXRD and SEM, reproducible code/data availability, comparison of four ML methods, and feature-importance analysis that aligns with known synthesis trends. However, the central claim of out-of-range Si/Al prediction is undermined by an internal inconsistency in the reported training range, and the reported 83% success rate is a conditional precision under the model's own high-confidence threshold rather than an unbiased generalization estimate. The paper's significance is therefore currently overstated; the underlying experimental result remains valuable if the claims are corrected.","major_comments":[{"comment":"The paper's key extrapolation claim is contradicted by its own Supplemental Table S1. Section III.C states that two successful conditions have Si/Al ratios 1.77 and 2.3, falling below the dataset's lower bound of 2.5, and the Introduction repeats this claim. However, Table S1 reports the in-house dataset's Si/Al range as 0.5–25.8. If this table is correct, both 1.77 and 2.3 are inside the training range, and the extrapolation claim is false. Moreover, the binary classifier used for screening was trained on the Section III.B filtered subset with 2≤Si/Al≤10, so its training lower bound is 2, not 2.5; under that definition, Si/Al=2.3 is also in-domain. The authors must either reconcile Table S1 with the text or remove the extrapolation claim.","section":"Section III.C and Table S1"},{"comment":"The 83% success rate is not an unbiased estimate of the pipeline's generalization. The six candidates were selected because the model assigned them probabilities of 0.86–0.93; the success rate is therefore conditional on high model confidence. Comparing this to the 32% prevalence of pure-EMT in the full dataset is not a valid baseline, because the six candidates are not a random sample from the same distribution. The paper should report precision at the chosen threshold with confidence intervals, or compare against candidates drawn randomly from the same PCA/search region but not selected by the classifier, to substantiate the claimed 2.6-fold improvement.","section":"Section III.C, Table IV, and Fig. 7"},{"comment":"All six experimentally validated candidates are nearly identical to each other and closely resemble the canonical OSDA-free EMT literature route summarized in Table S1: Na2SiO3 and NaAlO2 as precursors, pre-dissolved Si and Al, H2O=450, temperature 40–42°C, time ~72 h. This is a narrow local region of the synthesis space, not broad exploration. The claim that the pipeline 'explores the synthesis space and identifies six promising new conditions' should be tempered to acknowledge that the successful conditions are minor variations around a known EMT-favorable recipe, and the dissimilarity scores in Table IV (0.67–4.98) are modest in absolute terms.","section":"Section III.C and Table S1"},{"comment":"The model assumes that the 13 recorded descriptors are sufficient to determine crystallization outcome. The paper itself notes in the Introduction that minor variations in temperature or reaction time can cause EMT-to-FAU or EMT-to-SOD transitions. Since each of the six proposed conditions was validated in only one synthesis run, with no replicates and no report of uncontrolled variables such as mixing, gel aging, cooling rate, or operator-specific details, the five successes could be partly due to luck or unrecorded factors. Reporting repeated syntheses for at least the claimed extrapolation cases would strengthen the conclusion.","section":"Section II.A and Introduction"}],"minor_comments":[{"comment":"Typographical issue: 'Altogether, 6859 new candidate synthesis conditions were created... Finally, each candidate... via the inverse PCA transform' is clear, but later 'the PCA-generated populate' should be 'the PCA-generated pool' or 'population.'","section":"Section III.C"},{"comment":"The text reads 'Vmicore (cm3/g)'—this appears to be a typo for 'Vmicro' (micropore volume).","section":"Supplemental Material, Section S1"},{"comment":"In the introductory sentence, 'Section IIID of the main manuscript' should be 'Section III.D'.","section":"Supplemental Material, Section S9"},{"comment":"The in-house dataset ranges in Table S1 (e.g., Si/Al 0.5–25.8, H2O 173–650) are not consistently referenced in the main text, which contributes to the confusion about the training domain. A concise statement of which ranges were used for the binary model's training set would improve clarity.","section":"Table S1 and Section III.B"}],"recommendation":"major_revision","confidential_remarks":"The internal inconsistency between Table S1 and the Si/Al extrapolation claim is the central issue. If the correct in-house lower bound is 0.5, then the paper's headline result—successful extrapolation to Si/Al 1.77 and 2.3—collapses, and the contribution reduces to a modest model-guided interpolation study. The 5/6 experimental success rate is still worth reporting, but the manuscript needs substantial revision of the claims and a more careful statistical framing. Please ask the authors to provide the exact training ranges used by the screening model and to re-benchmark the success rate against an appropriate random or low-confidence baseline. If the extrapolation claim cannot be substantiated, it should be removed from the abstract and conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Sanjay—\n\nThe short version: this is a genuinely useful experimental paper with a framing problem. The five newly validated EMT recipes are real, and the closed-loop ML-to-experiment pipeline is a good template. But the paper's most marketable claim—that the model extrapolates to Si/Al ratios below the training range—is contradicted by the paper's own supplemental table. Table S1 lists the in-house Si/Al range as 0.5–25.8; the abstract and Section III.C claim the training data all had Si/Al above 2.5. Those can't both be true. Since two of the five successes are at Si/Al 1.77 and 2.3, the extrapolation story as told doesn't hold up. The 2.3 case isn't even out of the filtered training region (2≤Si/Al≤10) they define in Section III.B. Only 1.77 is actually below that bound.\n\nThe paper does several things well. The dataset of 174 in-house attempts is a real contribution, and the PXRD/SEM confirmation of the five new recipes is solid evidence. The feature-importance analysis matches known chemistry (temperature, Si/OH, Si/Al). The literature benchmark is presented with explicit caveats about imputation—that's the right kind of honesty. Code and data are public.\n\nThe other soft spots are real but more minor. The 83% success rate is conditional on selecting the six candidates with the highest predicted probabilities, so it's not an unbiased estimate of generalization; the baseline comparison to 32% is apples-to-oranges. The screening thresholds (probability >0.85, dissimilarity <0.5) appear post hoc. And the six chosen conditions cluster tightly around the canonical recipes—all use Na2SiO3/NaAlO2, H2O=450, T~40–42°C, t~72h—so the 'discovery' is really within a known EMT-favorable regime, with the possible exception of the Si/Al=1.77 case. The descriptor-sufficiency worry (mixing, aging, thermal history not captured) is legitimate but not unique to this paper; it's a general limitation of synthesis datasets.\n\nNet: the experimental contribution is worth refereeing and eventually citing. But the extrapolation claim needs to be rewritten or dropped before publication. If they fix the Si/Al inconsistency, the paper is a solid application piece. If they don't, a careful referee should push hard on it.\n\nI'd send it to peer review—the data and validation deserve referee time—but I'd want the authors to reconcile the numbers.","headline":"Solid experimental core with five validated EMT recipes, but the headline extrapolation claim is contradicted by the paper's own Table S1.","tokens_in":20835,"tokens_out":3577,"would_cite":true,"duration_ms":40274,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Trained on 174 in-house synthesis runs, the paper's machine-learning pipeline proposes six recipes for the zeolite framework EMT; five crystallize as pure EMT, including two with silicon-to-aluminum ratios below every value the model was tr","keywords":["EMT zeolite","zeolite synthesis","machine learning","synthesis condition discovery","OSDA-free synthesis","feature importance","PCA latent space","experimental validation"],"falsifier":"Rerun each of the five successful recipes while deliberately varying only factors outside the 13 descriptors — stirring rate, gel-aging time, cooling rate, vessel geometry — and check whether pure EMT persists; any flip to hybrid or non-EMT shows the descriptor set does not determine the outcome. Alternatively, fix the pipeline's other features and synthesize at Si/Al = 1.5 and 2.0: if neither yields pure EMT, the two out-of-range successes are a narrow accident rather than evidence of extrapolation.","tokens_in":19967,"feed_emoji":"🧪","tokens_out":13415,"duration_ms":138131,"temperature":0.7,"pith_summary":"The paper asks whether machine learning trained on a modest set of recorded synthesis attempts can guess which untried recipes will crystallize the zeolite framework EMT rather than its close relatives FAU or SOD. Using 174 in-house syntheses described by 13 parameters, the authors train classifiers — tree ensembles and a pretrained tabular foundation model perform about equally — to separate pure-EMT outcomes from mixed or failed ones. They then generate candidate recipes in a dimension-reduced (principal-component) version of the parameter space, screen them by predicted probability, physical plausibility, and novelty, and test the six most promising. Five of the six crystallize as pure EMT, an 83% hit rate against the 32% baseline of the original data, and two of the five use Si/Al stoichiometries below every value in the training set. The paper's claim is that this closed loop — predict, screen, synthesize, validate — transfers beyond the training domain and can be reused for other frameworks.","feed_headline":"Five of six machine-learning recipes yield pure EMT zeolite","feed_subtitle":"Trained on 174 runs, the model beats the 32% baseline and finds Si/Al ratios never tried before.","key_machinery":"The load-bearing mechanism is a two-stage pipeline. A proposer generates candidate synthesis conditions in a three-dimensional latent space from principal component analysis of the 13 recorded features, sampling 19 evenly spaced values per component from 30% below to 30% above the observed range, then inverts the transform; generating in latent space respects strong feature correlations, so candidates remain physically plausible. An evaluator screens them with three filters: physical soundness (no negative ratios), predicted pure-EMT probability above 0.85 from an XGBoost (gradient-boosted tree) binary classifier, and a dissimilarity score — minimum Euclidean distance to any training sample","core_discovery":"A classifier trained on 174 organic-template-free EMT synthesis attempts, each described by 13 ordinary recipe features, proposes new crystallization conditions that yield pure EMT zeolite with 83% success (five of six tested recipes), a 2.6-fold improvement over the 32% experience-guided baseline. The discovery pipeline compresses the 13 features into three principal components, samples 6,859 candidates stretching 30% beyond the observed ranges, maps them back, discards physically invalid ones, keeps the 374 with predicted pure-EMT probability above 0.85, and retains the 28 whose minimum Euclidean distance to the training data exceeds 0.5. The six highest-probability candidates were synthes","pith_inferences":["The single hybrid-EMT failure among the six tested candidates is itself informative: feeding that outcome back into the training set would test whether the 0.85 probability cutoff is well calibrated near the pure/hybrid boundary, something the paper does not do.","The two out-of-range successes suggest the pure-EMT window in Si/Al extends lower than the training data implied; probing Si/Al near 1.5–2.0 with the pipeline's other features fixed would map the true phase boundary and show whether the extrapolation is reliable or fortunate.","The misclassified literature case (72 °C, 24 h) hints that time–temperature interactions are the least transferable part of the descriptor set; a fixed-composition study sweeping time and temperature would reveal whether a missing descriptor (aging, heating rate, or mixing) is needed for cross-laboratory transfer.","The 0.5 dissimilarity threshold discarded 346 of 374 high-probability candidates, so the reported 83% is conditional on that specific novelty cut; relaxing it would trade hit rate for wider exploration, and the paper does not map that trade-off."],"forward_implications":["EMT recipes can be proposed and screened computationally, and the top candidates validate at an 83% success rate — more than 2.5 times the 32% rate of the experience-guided runs that built the dataset.","The pipeline reaches compositions the training data never contained: pure EMT formed at Si/Al = 1.77 and 2.30, both below the dataset's lower bound of 2.5.","The model transfers to other laboratories' syntheses, correctly predicting 8 of 12 literature EMT cases inside the training envelope and 3 of 4 beyond it, with missing descriptors imputed from the training set.","The workflow — propose in a reduced latent space, screen by predicted probability and novelty, validate in the lab — is put forward as a general route for other zeolite frameworks and other crystalline materials with many coupled synthesis variables.","Because the trained classifier is cheap to run, the 22 high-probability candidates already listed in the supplement are immediately available for further experimental testing."],"supporting_citations":[{"why":"Supplies the canonical organic-template-free EMT route whose gel composition anchors the in-house dataset.","marker":"[12]"},{"why":"Supplies the other canonical low-temperature EMT route and the evidence that small time or temperature shifts switch EMT to FAU or SOD.","marker":"[13]"},{"why":"Provides TabPFN, the pretrained foundation model benchmarked against the classical classifiers on the small dataset.","marker":"[47]"},{"why":"Provides Random Forest, one of the baseline tree ensembles.","marker":"[50]"},{"why":"Provides XGBoost, the gradient-boosted tree algorithm selected for the screening classifier.","marker":"[51]"},{"why":"Provides the SHAP feature-importance method used to identify the controlling synthesis parameters.","marker":"[53]"},{"why":"Provides the mean-decrease-in-impurity importance method used alongside SHAP.","marker":"[46]"},{"why":"Supplies an external literature EMT synthesis (rice-husk route) used in the validation set.","marker":"[48]"},{"why":"Supplies low-NaOH literature EMT syntheses that define the extrapolation region of the external validation.","marker":"[54]"},{"why":"Supplies the short, low-temperature literature EMT synthesis the model misclassifies, used to bound the model's transferability.","marker":"[56]"}],"fun_headline_variants":["ML-guided synthesis finds EMT zeolites in 5 of 6 tries","Machine learning boosts pure EMT zeolite yield to 83%","AI finds new Si/Al ratios for pure EMT zeolite","Data-driven recipe discovery yields five novel EMT syntheses","ML beats baseline 2.6x in EMT zeolite synthesis"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the 13 recorded descriptors — amounts, ratios, precursor sources, pre-dissolution flags, time, and temperature — determine the crystallization outcome, so any unrecorded factor such as mixing, gel aging, cooling rate, or thermal history that also shifts the phase would make the model's confident predictions fail to transfer, a fragility the paper itself acknowledges in noting that minor temperature or time variations can flip EMT to FAU or SOD","fun_headline_variants_meta":{"raw":{"variants":["ML-guided synthesis finds EMT zeolites in 5 of 6 tries","Machine learning boosts pure EMT zeolite yield to 83%","AI finds new Si/Al ratios for pure EMT zeolite","Data-driven recipe discovery yields five novel EMT syntheses","ML beats baseline 2.6x in EMT zeolite synthesis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000161,"raw_usage":{"total_tokens":1084,"prompt_tokens":766,"completion_tokens":318,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":231}},"tokens_in":510,"tokens_out":318,"duration_ms":4498,"temperature":1.0,"reasoning_tokens":231,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T13:11:43.155750+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun each of the five successful recipes while deliberately varying only factors outside the 13 descriptors — stirring rate, gel-aging time, cooling rate, vessel geometry — and check whether pure EMT persists; any flip to hybrid or non-EMT shows the descriptor set does not determine the outcome. Alternatively, fix the pipeline's other features and synthesize at Si/Al = 1.5 and 2.0: if neither yields pure EMT, the two out-of-range successes are a narrow accident rather than evidence of extrapolation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the canonical organic-template-free EMT route whose gel composition anchors the in-house dataset."},{"cited_title":"Hastie, R","cited_arxiv_id":null,"evidence_quote":"Provides TabPFN, the pretrained foundation model benchmarked against the classical classifiers on the small dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides Random Forest, one of the baseline tree ensembles."},{"cited_title":"Breiman, Random forests, Machine Learning45, 5 (2001)","cited_arxiv_id":null,"evidence_quote":"Provides XGBoost, the gradient-boosted tree algorithm selected for the screening classifier."},{"cited_title":"Gandhi and M","cited_arxiv_id":null,"evidence_quote":"Provides the mean-decrease-in-impurity importance method used alongside SHAP."},{"cited_title":"Hollmann, S","cited_arxiv_id":null,"evidence_quote":"Supplies an external literature EMT synthesis (rice-husk route) used in the validation set."},{"cited_title":"Wu and K.-j","cited_arxiv_id":null,"evidence_quote":"Supplies the short, low-temperature literature EMT synthesis the model misclassifies, used to bound the model's transferability."}],"review_version":1}