{"id":"4dd0500e-0458-4602-bfca-63881003a0c0","arxiv_id":"2608.11062","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An uncertainty-aware random forest combined with M3GNet relaxation and static DFT screening of 5.5 million compounds yields 209 ultra-low and 227 ultra-high work function surfaces.","lead":"This paper screens about 5.5 million compounds with a machine-learning model that reports calibrated uncertainty, then relaxes and recalculates the most promising surfaces to identify hundreds of materials with extremely low or high work functions. If the computed values are confirmed, the resulting candidate list could accelerate work on electron emitters, solar-cell contacts, and catalysts.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"M3GNet-relaxed geometries are used for final static DFT work functions without validation against DFT relaxation; Fig. 6's 0.5-0.6 eV relaxation shifts make this the decisive uncertainty in the extreme-surface counts.","rationale":"The reader's verdict is CONDITIONAL and identifies M3GNet relaxation as the weakest assumption. I agree. The final claim is the extreme-work-function counts, and those counts are determined by static SCF values on M3GNet-relaxed geometries. The manuscript's own Fig. 6 demonstrates that relaxation changes predicted work functions by hundreds of meV, and no comparison of M3GNet to DFT-relaxed geometry is given. Since many surfaces sit near the 2.0/6.0 eV thresholds by design, geometry errors of this magnitude can plausibly change the final counts. A secondary inconsistency, abstract reports 209 low-WF surfaces while the introduction reports 219, reinforces that the counts should be checked, but it is not the main correctness risk. The calibrated uncertainty and domain-of-applicability analyses are internally coherent and useful, but they address RF prediction error rather than the M3GNet-to-DFT geometry gap. Given the data and code availability statement, the proposed subset DFT-relaxation check is feasible and would settle the concern. Therefore the CONDITIONAL verdict remains appropriate; this is the same load-bearing concern as the reader's weakest assumption, so no verdict change is needed.","tokens_in":16209,"tokens_out":3687,"duration_ms":83527,"concrete_test":"Select 30-50 candidate surfaces spanning the low and high regimes, including representative lanthanide-, metalloid-, and P-terminated surfaces from Table 1. For each, perform a full DFT relaxation with the same VASP-PBE settings (520 eV cutoff, 0.04 A^-1 k-points, dipole correction) starting from the M3GNet-relaxed geometry, and recompute the work function from the DFT-relaxed slab. Compare the static-SCF values on M3GNet geometry with the values on DFT-relaxed geometry; report the distribution of signed differences and the number of surfaces that move across the 2.0 eV and 6.0 eV thresholds. If threshold-crossing counts change by more than 10-20%, the screening conclusion is not stable under geometry relaxation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the set of 209 low- and 227 high-work-function surfaces produced by static PBE SCF on M3GNet-relaxed slabs (Sec. 2.3). For this claim to be robust, M3GNet must yield surface geometries close to DFT relaxation across the screened chemistry, including lanthanide-, metalloid-, and phosphorus-terminated surfaces. The paper provides no such validation: Fig. 6 compares RF predictions on unrelaxed vs M3GNet-relaxed structures, not M3GNet vs DFT geometries, and Fig. 7 compares ML to DFT on the same M3GNet geometry, so neither checks the geometry itself. Figure 6 shows relaxation changes predicted work functions by up to 0.5-0.6 eV for high-work-function surfaces, and the final counts use thresholds at 2.0 and 6.0 eV. A systematic M3GNet error of this size could change which surfaces cross these thresholds and thus the headline count. The protocol also relaxes only the top 1/3 of the slab with a 0.05 eV/A force cutoff, and M3GNet was trained mostly on bulk Materials Project data; its fidelity for low-coordination surface atoms, especially f-electron and metalloid systems, is not established. Because the DFT step is single-shot SCF, any M3GNet geometry error propagates directly into the final reported work functions. A secondary count inconsistency (abstract says 209 low-WF surfaces, introduction says 219) is not the main correctness risk, but it reinforces that the final candidate set should be treated as conditionally supported pending verification.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a multi-fidelity screening workflow for discovering materials with extreme work functions. The authors augment a previously published random forest model for work function prediction with calibrated uncertainty estimates and a domain-of-applicability (DoA) assessment, then screen about 5.5 million compounds from the GNoME and Alexandria databases. After filtering for stability and metallic character, they generate roughly 11 million surfaces, apply the ML model, relax selected slabs with the M3GNet universal interatomic potential, and perform static PBE DFT calculations on about 1.5k relaxed surfaces. They report 209 surfaces (136 materials) with DFT work functions below 2.0 eV and 227 surfaces (172 materials) above 6.0 eV, and they identify chemical trends such as lanthanide-rich terminations favoring low work functions and metalloid or phosphorus terminations favoring high work functions.","tokens_in":16536,"tokens_out":3895,"duration_ms":37612,"significance":"If the reported candidate sets are reliable, they represent a substantial expansion of computationally predicted extreme-work-function materials, with potential impact on thermionic emission, catalysis, and electronic device contacts. The workflow itself is a useful demonstration of combining uncertainty quantification and DoA screening with MLIP-based relaxation and targeted DFT. The paper also makes its data and code publicly available, which supports reproducibility. However, the central quantitative claims rest on the adequacy of M3GNet-relaxed geometries as proxies for DFT-relaxed surfaces, and on the transferability of an RF model trained on unrelaxed surfaces to relaxed geometries; these points are not directly validated in the manuscript.","major_comments":[{"comment":"The final DFT work functions are computed as single-shot static SCF calculations on M3GNet-relaxed slab geometries, but the manuscript provides no direct validation that M3GNet geometries are accurate proxies for DFT-relaxed surface geometries across the screened chemistry space, including lanthanide-, metalloid-, and phosphorus-terminated surfaces. Figure 6 compares ML predictions on unrelaxed versus M3GNet-relaxed surfaces, and Figure 7 compares ML predictions with DFT on the same M3GNet-relaxed geometry; neither checks the geometry itself against DFT relaxation. Since Figure 6 shows that relaxation can shift predicted work functions by 0.5-0.6 eV for high-work-function surfaces, a systematic M3GNet geometry error of comparable size would directly change which surfaces cross the 2.0 eV and 6.0 eV thresholds, and hence alter the headline counts of 209 and 227. The authors should add a validation set in which a stratified random subset of final candidates is fully relaxed with DFT and the resulting work functions are compared with the static-DFT-on-M3GNet values, or alternatively relax the final candidate slabs with DFT before reporting final work functions.","section":"Sec. 2.3, Figs. 6 and 7"},{"comment":"The random forest model was trained on DFT work functions of unrelaxed surfaces, yet it is applied to M3GNet-relaxed surfaces for re-prediction after relaxation. The model has never seen relaxed geometries in training, and the features used (interlayer distances, packing fractions, layer angles) change upon relaxation. The manuscript does not demonstrate that the unrelaxed-trained model remains valid for relaxed geometries, nor does it retrain or fine-tune on relaxed-surface data. Because the relaxed-surface ML predictions are used to select which surfaces proceed to DFT, any systematic bias in this transfer could distort the final candidate set. Please either retrain the model on relaxed surfaces, provide evidence that the model transfers reliably (e.g., by benchmarking on a set of DFT-relaxed surfaces with known work functions), or clearly restrict the ML-based selection to unrelaxed features and justify the use of relaxed-geometry predictions.","section":"Secs. 2.1 and 2.3"},{"comment":"The DoA analysis in Figure 5e shows that all training surfaces with work functions above 7 eV are flagged as out-of-domain, yet the workflow explicitly targets high-work-function surfaces above 6.0 eV. The manuscript does not explain how the final high-work-function candidates avoid being flagged as out-of-domain, or how the DoA cutoff interacts with the extreme high-work-function regime. This is not necessarily a fatal flaw, but it needs discussion because the DoA assessment is presented as a guardrail for trustworthy predictions, and the reader needs to know whether the high-work-function candidates are in-domain according to the calibrated DoA metric.","section":"Sec. 3.3 and Fig. 5e"}],"minor_comments":[{"comment":"The abstract states 209 low-work-function surfaces below 2.0 eV, while the introduction states 219 such surfaces; Sec. 3.4 and the conclusion both report 209. This is likely a typo in the introduction, but it should be corrected for consistency because the count is a headline result.","section":"Abstract vs. Introduction"},{"comment":"The DFT uncertainty is fixed at ±0.1 eV based on an \"empirical estimate,\" but no justification or reference is provided for this value. Since these error bars are used to state that 79% and 81% of points fall within uncertainty bounds, the authors should either justify the ±0.1 eV value with data or present the comparison without fixed error bars.","section":"Sec. 3.5 and Fig. 7"},{"comment":"There are minor language issues: \"from the the bulk Fermi level\" in the first sentence of the introduction, \"We augmented the previously random forest machine learning model\" in the conclusion, and a duplicated sentence in the conclusion beginning \"We augmented...\". These should be cleaned up.","section":"Introduction and Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The two major comments are coupled: if M3GNet geometry is validated against DFT relaxation and the model-transfer issue is addressed, the paper's central claims would be substantially strengthened. The count inconsistency and fixed uncertainty are minor but should be fixed. I do not see a circularity problem in the use of the RF model: the DFT validation is an external benchmark, and no fitted parameter is used as a target. However, the reliance on self-cited prior methods is not itself a concern. The paper is within scope for a computational materials science journal, and the data/code availability is a positive feature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a solid multi-fidelity screening paper that produces genuinely new candidate lists for extreme work functions, but I would not treat the headline counts as settled. The workflow—RF model plus calibrated uncertainties plus domain-of-applicability filter, then M3GNet relaxation, then static PBE SCF—is coherent and the scale is impressive: roughly 11 million surfaces screened down to a few hundred DFT-verified candidates. The chemical trends, especially lanthanide-rich low-work-function terminations and metalloid/phosphorus high-work-function terminations, are interesting and not in the prior literature.\n\nThe paper does several things well. The uncertainty calibration and DoA analysis are carefully done and build on prior methods without overselling them. The feature importance analysis aligns with physical intuition. The candidate tables and data availability are useful.\n\nThe soft spots are real. The final DFT values are static SCF on M3GNet-relaxed structures, and there is no validation of M3GNet geometries against DFT relaxation for any subset. Figure 6 shows that relaxation changes predicted work functions by up to 0.5–0.6 eV, which is the same magnitude as the margins around the 2.0 and 6.0 eV thresholds. A systematic geometry error that size could shift which surfaces cross the thresholds, so the 209/227 counts should be viewed as provisional. The abstract/introduction mismatch (209 vs 219 low-WF surfaces) is a minor but annoying count error. The fixed ±0.1 eV DFT uncertainty is empirical, not a rigorous estimate.\n\nThat said, the central screening logic holds up. The role of the ML model is to prioritize; the final values come from PBE, so this is not circular. I would send a serious referee, but the authors should be asked to either validate M3GNet relaxation against DFT for a representative subset of the extreme candidates or soften the quantitative claims. The paper is a good reading-group example of how to think about relaxation accuracy in ML-driven screening.","headline":"A solid, useful multi-fidelity screening paper whose headline counts are provisional until the M3GNet-relaxed geometries are validated against DFT relaxation.","tokens_in":17078,"tokens_out":2733,"would_cite":true,"duration_ms":30372,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An uncertainty-aware machine-learning screen turns 5.5 million compounds into 436 extreme-work-function surfaces, ready for DFT-level follow-up.","keywords":["work function","machine learning screening","uncertainty calibration","domain of applicability","multi-fidelity screening","density functional theory","surface termination","electron emission"],"falsifier":"Take the 436 reported surfaces and re-relax each slab with DFT rather than the machine-learned potential, then recompute the work function with the same DFT-PBE settings; if a substantial fraction of surfaces no longer fall below 2.0 eV or above 6.0 eV, the extreme-work-function counts are an artifact of the relaxation proxy.","tokens_in":15995,"feed_emoji":"⚡","tokens_out":6619,"duration_ms":55373,"temperature":0.7,"pith_summary":"The paper tries to show that extreme work functions—the energy needed to pull an electron out of a surface—can be found at scale by layering cheap machine-learning filters under expensive density-functional-theory checks. It augments a published random-forest work-function predictor with calibrated error bars and an in-domain/out-of-domain guardrail, relaxes the most promising slab surfaces with a universal machine-learned interatomic potential, and confirms the survivors with DFT-PBE. Out of roughly 5.5 million compounds screened, the workflow reports 209 surfaces with DFT work functions below 2.0 eV and 227 surfaces above 6.0 eV, corresponding to 136 and 172 unique materials. If these numbers hold, they give device designers a concrete candidate pool for electron-emitting and hole-injecting materials, and they expose chemical motifs—lanthanide-rich surfaces for low work function, metalloid- or phosphorus-terminated surfaces for high—that extend familiar design rules.","feed_headline":"Machine-learning screen finds 436 extreme work-function surfaces","feed_subtitle":"Calibrated ML filters plus DFT checks turn 5.5 million compounds into concrete emitter and hole-injector candidates.","key_machinery":"The load-bearing mechanism is the multi-fidelity funnel: a calibrated random-forest regression model supplies fast per-surface work-function estimates; a kernel-density estimate of the training data supplies a dissimilarity score that rejects out-of-domain predictions at a cutoff of $d=0.97$; a universal machine-learned interatomic potential relaxes the top slab region of candidate surfaces; and targeted static DFT-PBE calculations on the relaxed slabs set the final work function. The funnel's role is to concentrate the expensive DFT budget on surfaces whose ML predictions are both extreme and trustworthy, rather than to replace DFT.","core_discovery":"The paper's central claim is that a trust-aware multi-fidelity funnel can replace brute-force DFT screening for surface properties. Starting from roughly 5.5 million bulk compounds in two open computational databases, the authors keep only non-radioactive, thermodynamically stable, metallic compounds, then use an uncertainty-calibrated random-forest model to predict work functions for about 11 million low-index surfaces while flagging predictions outside its training domain. Surfaces passing the domain filter and ranked in the extreme tails are relaxed with a universal machine-learned interatomic potential and re-ranked; finalists receive a static DFT-PBE calculation. The endpoint is a set of 209 surfaces with $\\Phi_{\\mathrm{DFT}} < 2.0$ eV and 227 surfaces with $\\Phi_{\\mathrm{DFT}} > 6.0$ eV, with lanthanide-rich terminations over-represented at the low end and metalloid or phosphorus terminations at the high end.","pith_inferences":["Because the final DFT values are single-shot static calculations on machine-learned-relaxed slabs, a full DFT relaxation of a random subset of the 436 surfaces would be the most direct stress test; the paper's own relaxation analysis suggests some fraction of the extreme counts could shift across the thresholds.","The lanthanide-rich low-work-function pattern suggests a targeted expansion into lanthanide sulfides, selenides, and pnictides—a chemical neighborhood the current candidates sample sparsely.","The same calibrated-ML-plus-domain-filter funnel could be turned on other surface-sensitive quantities, such as cleavage energy or surface energy, where interpolation reliability matters as much as raw accuracy."],"forward_implications":["The published list of 209 low- and 227 high-work-function surfaces, mapped to 136 and 172 unique materials, is a directly usable candidate pool for thermionic emitters, field emitters, hole-transport layers, and electron-blocking layers.","Surface relaxation moves predicted work functions by up to 0.5–0.6 eV, so any screening pipeline that scores unrelaxed slabs will mis-rank extreme candidates; relaxation must be an explicit stage.","Uncertainty calibration plus domain-of-applicability filtering lets a screening campaign deprioritize predictions that are merely extrapolations, which is what makes the 99.8% reduction in search space defensible.","Lanthanide-rich surface terminations are a newly emphasized low-work-function motif, and metalloid- or phosphorus-terminated surfaces a high-work-function motif, giving synthetic chemists compositional targets beyond the familiar alkali and alkaline-earth rules."],"supporting_citations":[{"why":"Supplies the random-forest model, the 58,332-surface training set, and the surface-enumeration toolkit that the whole funnel builds on.","marker":"[4]"},{"why":"Graph-network band-gap model used to exclude compounds whose PBE band gap underestimates how metallic they really are.","marker":"[31]"},{"why":"One of the two large computed-materials databases screened, contributing the stable GNoME compound set.","marker":"[32]"},{"why":"The other screened database, contributing the broader Alexandria compound set that most low-work-function finalists come from.","marker":"[33]"},{"why":"Calibrated bootstrap method that rescales raw random-forest uncertainty into statistically consistent error bars.","marker":"[36]"},{"why":"Kernel-density applicability-domain method whose dissimilarity cutoff filters out unreliable extrapolated predictions.","marker":"[38]"},{"why":"Universal machine-learned interatomic potential used to relax candidate surface slabs before DFT.","marker":"[39]"},{"why":"PBE exchange-correlation functional used in the final static DFT work-function calculations.","marker":"[43]"}],"fun_headline_variants":["Uncertainty-aware ML finds 436 extreme work-function surfaces","Multi-fidelity screen nets 436 extreme work-function surfaces","Calibrated ML sieve reveals 436 extreme work-function surfaces","Trust-aware screening uncovers 436 extreme work-function surfaces","From 5.5M compounds to 436 extreme work-function surfaces"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The screening counts stand or fall on the premise that the machine-learned interatomic potential used for relaxation gives geometries close enough to true DFT-relaxed geometries that single-shot DFT work functions on those geometries remain reliable; the paper's own Figure 6 shows relaxation changes predicted work functions by up to 0.5–0.6 eV.","fun_headline_variants_meta":{"raw":{"variants":["Uncertainty-aware ML finds 436 extreme work-function surfaces","Multi-fidelity screen nets 436 extreme work-function surfaces","Calibrated ML sieve reveals 436 extreme work-function surfaces","Trust-aware screening uncovers 436 extreme work-function surfaces","From 5.5M compounds to 436 extreme work-function surfaces"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000551,"raw_usage":{"total_tokens":2650,"prompt_tokens":989,"completion_tokens":1661,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":605,"completion_tokens_details":{"reasoning_tokens":1575}},"tokens_in":605,"tokens_out":1661,"duration_ms":12223,"temperature":1.0,"reasoning_tokens":1575,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:01:18.779343+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the 436 reported surfaces and re-relax each slab with DFT rather than the machine-learned potential, then recompute the work function with the same DFT-PBE settings; if a substantial fraction of surfaces no longer fall below 2.0 eV or above 6.0 eV, the extreme-work-function counts are an artifact of the relaxation proxy.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Graph-network band-gap model used to exclude compounds whose PBE band gap underestimates how metallic they really are."},{"cited_title":"Ghahremanpour, P.J","cited_arxiv_id":null,"evidence_quote":"The other screened database, contributing the broader Alexandria compound set that most low-work-function finalists come from."},{"cited_title":"Palmer, S","cited_arxiv_id":null,"evidence_quote":"Calibrated bootstrap method that rescales raw random-forest uncertainty into statistically consistent error bars."},{"cited_title":"Jacobs, L.E","cited_arxiv_id":null,"evidence_quote":"Universal machine-learned interatomic potential used to relax candidate surface slabs before DFT."}],"review_version":1}