{"id":"25467220-2126-4803-a482-3260c06455d8","arxiv_id":"2411.16770","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A regression model trained on interatomic-potential data predicts the same grain boundary energy scaling from fundamental properties that density functional theory finds, supporting the use of potential ensembles as synthetic materials.","lead":"Using many approximate interatomic potentials as a stand-in ensemble for real metals, the authors build a regression model that predicts a grain boundary energy scaling factor from basic material properties, then show it also works for density functional theory data. The approach offers a recipe for predicting expensive large-scale material quantities from cheap first-principles properties and for choosing what to include when training new interatomic potentials.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The same-pool validation is confounded if the DFT E0 reference is fitted to a different, smaller set of grain boundaries than the IP E0; because the LM model is approximate, E0 depends on the fitted boundary set, so the reported Pd, Pt, Ni errors may reflect a definitional mismatch.","rationale":"The reader correctly identifies the same-statistical-pool assumption as the load-bearing premise, but frames the weakness as breadth (one QoI, seven metals). My concern is more specific: the validation may not even measure a single well-defined E0 for both IPs and DFT. If the DFT E0 is fitted to different grain boundaries than the IP E0, the comparison in Table 4 is not apples-to-apples, and the reported errors could arise from this mismatch rather than from a failure of the pool assumption. This is testable and would substantially change the confidence in the paper's 'confirming' language. The paper has genuine strengths: an out-of-sample DFT test, recovery of known correlations, use of OpenKIM-curated data, and a clear data pipeline. These support a conditional acceptance, but the condition should include demonstrating that the DFT E0 is computed from a comparable set of boundaries as the IP E0, or reporting the sensitivity of the errors to the boundary subset.","tokens_in":66511,"tokens_out":7838,"duration_ms":77245,"concrete_test":"Obtain the exact list of grain boundaries in the Zheng et al. [24] DFT database for the seven metals. For each IP in the training set, refit E0 using only those same tilt axes and angles with the same wield code and LM parameters, then re-train the three-factor model on the restricted E0 values and compare predictions with the DFT E0 values. If the per-metal percent errors change by more than 5 percentage points for any of Ni, Pd, or Pt, or if the restricted IP E0 distributions differ systematically from the full-curve fits, the reported validation is confounded by the choice of GB subset.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim is that IPs and DFT belong to the same statistical pool, validated by agreement between the regression-predicted E0 and the E0 fitted directly to DFT GB energies (Table 4, Fig. 6). This validation is meaningful only if the two E0 values are the same quantity. Section 2.2 states that IP E0 is obtained by least-squares fitting of the LM model to GB energy curves across four tilt axes ([001], [111], [110], [112]) per IP. Section 3.3 does not state how the DFT E0 was fitted beyond 'directly from first-principles GB energy data' from Ref. [24]. If the DFT dataset contains only a subset of GB planes and angles (e.g., a few special boundaries per metal), then—because the LM model is approximate (the fits in Fig. 1 deviate from the atomistic data)—E0 becomes dependent on the specific boundaries included in the fit. The comparison would then conflate a genuine same-pool test with a definitional mismatch. The systematic negative errors for Ni (-14.8%), Pd (-21.9%), and Pt (-16.5%) are consistent with such a mismatch. This is the most load-bearing threat to the 'confirming' conclusion because it attacks the operationalization of the validation, not just its breadth.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a statistical strategy for discovering correlations between fundamental microscopic properties (\"canonical properties\") and large-scale quantities of interest (QoIs), using an ensemble of classical interatomic potentials (IPs) as \"synthetic materials.\" As a proof of principle, the QoI is the scaling factor E0 in the lattice-matching (LM) model of symmetric tilt grain boundary (GB) energies in FCC metals. The authors extract GB energy curves and canonical properties for 235 IPs from OpenKIM, fit support vector and multilinear regression models, and identify unstable stacking fault energy, vacancy migration energy, vacancy formation energy, and C44 as the most informative predictors. They then use DFT-computed canonical properties to predict E0 and compare these predictions with E0 values fitted directly to DFT GB energies from the literature (Table 4, Fig. 6). The paper reports adjusted R-squared of 0.899 in nested k-fold cross-validation and interprets the DFT agreement as confirmation that IPs and DFT belong to the same statistical pool.","tokens_in":66800,"tokens_out":6525,"duration_ms":64389,"significance":"The proposed idea is genuinely useful if the same-pool assumption holds: it would allow regression models built from cheap classical IP simulations to be used with first-principles canonical properties to predict QoIs that are beyond DFT reach, and to guide the selection of properties for training new IPs. The paper has notable strengths: the use of the standardized OpenKIM test driver, the public availability of data and code on GitHub, the nested k-fold cross-validation protocol, and a genuine out-of-sample test in which a model trained on IP data is evaluated against independent DFT results. The work also recovers known GB-energy correlations and proposes new predictive relationships. The main limitations are the narrow scope of the validation (seven FCC metals and one scalar QoI) and, more importantly, the underspecified operationalization of the DFT E0 reference, which makes the central same-pool claim currently less secure than the abstract's wording suggests.","major_comments":[{"comment":"The central validation is not fully operationalized. In Section 2.2 the IP E0 is defined as a least-squares fit of the LM model to GB energy curves across four tilt axes ([001], [111], [110], [112]), but Section 3.3 does not state which tilt axes, GB planes, or tilt angles from Ref. [24] were used to obtain the \"Coefficient, exact fit using DFT GB energies\" in Table 4. Because the LM model is approximate (as shown by the deviations in Fig. 1), E0 is not a unique material constant but depends on the set of boundaries included in the fit. If the DFT data from Ref. [24] contain only a subset of the boundaries used in the IP fits, the comparison in Table 4 would conflate a genuine same-pool test with a definitional mismatch between two differently fitted E0 values. This concern is reinforced by the systematic negative errors for Ni (-14.8%), Pd (-21.9%), and Pt (-16.5%). Please specify the exact DFT boundary set, fit the DFT E0 with the same least-squares protocol used for the IP data, and report the sensitivity of E0 to the choice of boundary set. In addition, the statement in Section 4.3 that all DFT predicted values fall within the error bars of the regression prediction is not verifiable from Fig. 6 as printed, because no error bars are shown in that figure; please provide the prediction intervals used for this claim.","section":"Section 3.3, Table 4, Fig. 6"},{"comment":"The conclusion that IPs and DFT \"belong to the same statistical pool\" is stronger than the evidence presented. The validation covers only seven FCC metals, a single QoI, and one scalar descriptor (E0) of the full GB energy curves, with per-species percent errors up to 21.9% (Pd). These errors are not necessarily disqualifying if they are random and within well-calibrated error bars, but the systematic sign of the errors for Ni, Pd, and Pt is the pattern one would expect if the DFT E0 were fitted to a different and less complete set of boundaries. Please either temper the same-pool claim to reflect this limited validation, or add a robustness test, such as refitting the DFT E0 with the identical boundary set used for the IP data, a leave-one-species-out analysis, or a comparison on additional QoIs and boundary types.","section":"Section 4.3 and Conclusions"}],"minor_comments":[{"comment":"The text states that \"This calculation across tilt axes leads to the final dataset of 300 scaling coefficients\" and also mentions 235 unique IP models and 1040 GB energy curves; the relationship between these three numbers is not explained and should be clarified.","section":"Section 2.2"},{"comment":"The regression coefficients in Eq. (6) are reported with five significant digits, which is disproportionate to the accuracy of the model; fewer digits would be appropriate.","section":"Equation (6)"},{"comment":"The sign convention for \"Percent Error\" is not defined; a short sentence stating whether negative values correspond to underprediction of the DFT-fitted E0 would remove ambiguity.","section":"Table 4"},{"comment":"The caption of Fig. 6 describes gray circles as \"calculated scaling coefficient results that are outside of the quartiles,\" but the caption of Fig. S20 uses the same wording for canonical property values; please make clear in each caption whether the outlying points refer to E0 or to the property shown.","section":"Fig. 6 caption"},{"comment":"The caption of Fig. S19 states \"Strong correlation between rVFPE and uVFPE can be seen,\" but the figure shows stacking fault and twinning energies; this is evidently a copy-paste error and should be corrected.","section":"SI Section S5, Fig. S19"}],"recommendation":"major_revision","confidential_remarks":"The manuscript presents a promising and potentially influential approach, and the empirical material is valuable. The main issue is that the validation claimed in the abstract and Section 4.3 depends on a DFT reference E0 whose fitting procedure is not specified; this is fixable with additional analysis but is load-bearing for the paper's central claim. I would also encourage the authors to check that their GitHub repository matches the data described in the manuscript, since the availability statement is one of the paper's strengths."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Jason, here's my read on the Jasperson et al. paper. The core idea is good: use a collection of interatomic potentials as 'synthetic materials' to regress expensive QoIs against cheap first-principles properties, then test the regression on real DFT data. The transferability test is the genuine novelty—training on IP data, predicting DFT-derived grain boundary energy scaling factors—and the top three-factor model (uSFE, rVFPE, VME) with adjusted R² = 0.899 is a solid proof of concept. The curation is careful: 1040 GB curves across four tilt axes, nested k-fold CV, error bars. They recover known correlations (C44, stacking fault, vacancy energies) and add a few new ones; that's fine—the method is the contribution.\n\nThe soft spots are real but not fatal. First, the validation covers one QoI and seven FCC metals, with per-species errors up to 21.9% for Pd. Calling this 'confirming' the same-statistical-pool assumption is too strong; 'consistent with' is about right. Second, and more concerning, is the DFT E0 comparison. Section 2.2 explains that IP E0 is fit by least squares to GB energy curves across four tilt axes. Section 3.3 doesn't say how the DFT E0 was fit—whether it used the same tilt axes and angles or a smaller subset. Since the LM model is approximate (the fits in Fig. 1 deviate from the atomistic data), E0 is sensitive to which boundaries are included. If the DFT data from Ref. [24] covers only a few special boundaries, the comparison could be measuring a definitional mismatch, not a statistical-pool agreement. The systematic negative errors for Ni, Pd, Pt are consistent with this worry. The authors need to clarify this before the 'confirming' conclusion stands.\n\nThe data exclusions (153 curves, 76 pair potentials, property cutoffs) are documented, so not a red flag, but they do shrink the pool and should be motivated more strongly. Also, the code isn't out yet—the GitHub link says 'will be made available'—so reproducibility is pending.\n\nIn short: a worthwhile proof-of-principle, honestly presented, but with an overreaching abstract and a validation step that needs tightening. This is for computational materials scientists and IP developers. I'd send it to peer review and ask for (a) a precise description of the DFT E0 fitting procedure, (b) a softened claim in the abstract, and (c) a versioned code release. I'd bring it to the reading group and would cite it if the DFT fitting gets clarified.","headline":"A credible proof-of-principle for using IP ensembles to discover property correlations, but the 'same statistical pool' claim is validated on one QoI and the DFT E0 comparison may have a definitional mismatch—worth peer review with revisions.","tokens_in":67306,"tokens_out":2834,"would_cite":true,"duration_ms":27759,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper attempts to show that regression models trained on classical interatomic potentials can predict first-principles grain boundary energies, because approximate potentials and density functional theory draw from the same…","keywords":["grain boundary energy","interatomic potentials","canonical properties","lattice matching model","statistical inference","synthetic materials","support vector regression","density functional theory"],"falsifier":"One concrete test: compute DFT canonical properties and the directly fitted $E_0$ scaling factor from DFT grain boundary energies for a metal outside the seven studied here, such as Ir or Th, and compare the regression prediction interval trained only on potential data. If the DFT $E_0$ falls outside that interval, or if the sign of a top predictor's correlation with $E_0$ reverses, the same-statistical-pool assumption would be refuted.","tokens_in":66332,"feed_emoji":"⚛️","tokens_out":6971,"duration_ms":62858,"temperature":0.7,"pith_summary":"The paper tries to establish that classical interatomic potentials, treated as an ensemble of \"synthetic materials,\" can reveal correlations between fundamental microscopic properties and large-scale quantities of interest (QoIs) that hold for first-principles data and therefore for real materials. As a proof of principle, it builds a regression model that predicts the scaling factor $E_0$ of the lattice-matching model for symmetric tilt grain boundary energies in face-centered cubic (FCC) metals from canonical properties such as stacking fault energies, vacancy formation and migration energies, and elastic constants. Trained only on potential-derived data from a large curated repository, the model predicts density functional theory (DFT) $E_0$ values for seven FCC metals with per-metal errors mostly under 17 percent and a worst case near 22 percent. The authors take this agreement as evidence that potential data and DFT data belong to the same statistical pool, which is the key assumption enabling the whole approach. If the claim holds, regression models of this kind give a general route to predict large-scale QoIs from first-principles properties and to choose which properties should enter the training set of a new interatomic potential.","feed_headline":"IP-trained model predicts DFT grain boundary energies","feed_subtitle":"A regression built on hundreds of interatomic potentials matches first-principles results for seven FCC metals, validating a new…","key_machinery":"The machinery has three parts. First, the lattice-matching (LM) model [22] collapses an entire orientation-dependent symmetric tilt grain boundary energy curve into a single material-dependent scaling factor $E_0$ by writing $\\gamma = E_0\\left(1 - c/c_0\\right)$ in terms of a covariance between Gaussian-broadened lattice densities; this $E_0$ is the quantity to be predicted. Second, a large archived collection of classical interatomic potentials is treated as a pool of synthetic materials: each potential-plus-species combination gives one $E_0$ value and one set of canonical properties, yielding 300 scaling coefficients assembled from 1040 grain boundary energy curves. Third, support vector regression with a radial basis function kernel and a simpler three-factor multilinear regression are fit with nested k-fold cross-validation; the linear model $E_0 = 18.52125\\,\\text{uSFE-FCC} + 0.79610\\,\\text{rVFPE-FCC} + 0.67369\\,\\text{VME-FCC} - 0.15919$ is then applied to DFT-computed indicator properties and compared against $E_0$ obtained by directly fitting DFT grain boundary energies.","core_discovery":"On the paper's own terms, the central discovery is that correlations between canonical properties and a large-scale QoI discovered using approximate interatomic potentials survive transfer to DFT data. For the grain boundary energy scaling factor $E_0$, a three-factor linear regression trained on potential-derived results, using FCC unstable stacking fault energy, relaxed vacancy formation potential energy, and vacancy migration energy, achieves an adjusted $R^2$ of about 0.899 in nested cross-validation and predicts DFT $E_0$ values for Ag, Al, Au, Cu, Ni, Pd, and Pt with per-metal errors between about $-2.7\\%$ and $-21.9\\%$. Because all DFT predictions fall within the regression error bars, the authors conclude that potentials and DFT belong to the same statistical pool, which validates the proposed statistical inference approach and identifies the most relevant canonical properties to target when fitting potentials for grain-boundary-related simulations.","pith_inferences":["Extension: the validation covers only one QoI, the symmetric tilt grain boundary energy scaling in seven FCC metals, so the strongest testable version of the same-statistical-pool claim would require checking other QoIs, other crystal classes, and boundary types such as twist and asymmetric grain boundaries.","Extension: if the pooling assumption holds generally, the approach offers a transfer check for machine-learned potentials: a potential that breaks the discovered correlations with DFT canonical properties could be identified as outside the pool before being used in large-scale simulation.","Extension: the discovery that vacancy migration and formation energies outperform elastic constants as predictors suggests a concrete, falsifiable recommendation for future potential fitting, namely that defect energetics should be prioritized in training sets for grain-boundary-focused potentials."],"forward_implications":["A regression model of this kind can estimate a large-scale QoI for a real material using only DFT values of a few small canonical properties, bypassing expensive DFT simulations of the large-scale quantity itself.","The properties flagged as strong predictors, such as unstable stacking fault energy, vacancy formation energy, vacancy migration energy, and the shear elastic constant $C_{44}$, are the properties that should be included in the training set of a potential designed to reproduce grain boundary energetics.","The same workflow can be applied to other QoIs beyond DFT reach, such as plastic flow strength, by running an ensemble of existing potentials, regressing the QoI against canonical properties, then feeding first-principles values of the top predictors into the regression.","The consistency between potential-based and DFT results for $E_0$ implies that past atomistic grain boundary simulations using archived potentials can be treated as samples from the same statistical distribution as first-principles results, despite the approximations in each individual potential."],"supporting_citations":[{"why":"Supplies the lattice-matching model whose material scaling factor E0 is the quantity of interest.","marker":"[22]"},{"why":"Provides the automated grain boundary energy calculations used to generate the potential-based training data.","marker":"[23]"},{"why":"Supplies the DFT grain boundary energy database used as validation ground truth.","marker":"[24]"},{"why":"Describes the open repository of interatomic potentials that provides the ensemble of synthetic materials.","marker":"[17]"},{"why":"Demonstrates the workflow on plastic flow strength, a QoI beyond DFT reach.","marker":"[45]"},{"why":"Provides the regression and cross-validation implementations used to fit the predictive models.","marker":"[43]"}],"fun_headline_variants":["IP ensembles predict DFT grain boundary energies","Hundreds of potentials match DFT grain boundary data","Same statistical pool: IPs and DFT agree on GB energies","Statistical bridge from synthetic potentials to real QoIs","Canonical property model survives DFT transfer test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the correlations between properties found using classical interatomic potentials are the same correlations that hold for density functional theory results, and hence for real materials.","fun_headline_variants_meta":{"raw":{"variants":["IP ensembles predict DFT grain boundary energies","Hundreds of potentials match DFT grain boundary data","Same statistical pool: IPs and DFT agree on GB energies","Statistical bridge from synthetic potentials to real QoIs","Canonical property model survives DFT transfer test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1325,"prompt_tokens":949,"completion_tokens":376,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":565,"completion_tokens_details":{"reasoning_tokens":302}},"tokens_in":565,"tokens_out":376,"duration_ms":4003,"temperature":1.0,"reasoning_tokens":302,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:33:09.954556+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One concrete test: compute DFT canonical properties and the directly fitted $E_0$ scaling factor from DFT grain boundary energies for a metal outside the seven studied here, such as Ir or Th, and compare the regression prediction interval trained only on potential data. If the DFT $E_0$ falls outside that interval, or if the sign of a top predictor's correlation with $E_0$ reverses, the same-statistical-pool assumption would be refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the regression and cross-validation implementations used to fit the predictive models."},{"cited_title":"J., Conti, S","cited_arxiv_id":null,"evidence_quote":"Supplies the lattice-matching model whose material scaling factor E0 is the quantity of interest."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the automated grain boundary energy calculations used to generate the potential-based training data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the DFT grain boundary energy database used as validation ground truth."},{"cited_title":"B., Elliott, R","cited_arxiv_id":null,"evidence_quote":"Describes the open repository of interatomic potentials that provides the ensemble of synthetic materials."},{"cited_title":"Cross-scale covariance for material property prediction","cited_arxiv_id":"2406.05146","evidence_quote":"Demonstrates the workflow on plastic flow strength, a QoI beyond DFT reach."}],"review_version":1}