{"id":"e0be54e5-1e2f-4592-8324-fad68efb492c","arxiv_id":"1908.03665","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A Bayesian regularized regression with Bayesian information criterion feature selection predicts configurational energies of refractory high entropy alloys with held-out RMSE around 0.6 meV, and ensemble sampling across supercell sizes improves accuracy and stability.","lead":"This paper fits a machine learning model to density functional theory data to predict the energy of different atomic arrangements (configurations) in high entropy alloys, and uses Bayesian statistics to choose how complex the model should be. A smart generalist might read it because it shows a practical way to handle very small datasets in computational materials science and to know how much to trust the prediction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The model is validated only on random configurations, while the intended Monte Carlo use requires accuracy on ordered/SRO states; the pair-only EPI assumption is untested there.","rationale":"The reader's weakest assumption is the pair-only, size-independent EPI model plus absence of a supercell-size offset. The size-offset part is substantially tested by Fig. 5, where training on 16/32/64-atom cells and testing on 128/160-atom cells yields mean RMSE below 1 meV; I therefore do not see that as the most load-bearing unresolved issue. The more consequential gap is distributional: every DFT point in the paper is a random configuration, so the excellent RMSE demonstrates interpolation within a narrow random-disorder window. The stated application (abstract, Section 1, Step 6) is Monte Carlo simulation of ordering, which samples configurations far from random. For such states, omitted triplet/many-body interactions and possibly longer-ranged pair interactions can matter, and the paper provides no evidence either way. This is a missing-validation concern rather than an internal inconsistency, so it supports the existing CONDITIONAL verdict rather than rejection. The ordered-configuration test is a single, decisive check.","tokens_in":15216,"tokens_out":12627,"duration_ms":158247,"concrete_test":"Generate about 200 ordered/SRO configurations for NbMoTaW in 128-atom supercells (e.g., B2-like layering and SQS with Warren-Cowley SRO parameters spanning roughly -0.2 to +0.2), compute LSMS energies, and predict them with the EPI model fitted only to the 600 random configurations from 16/32/64-atom cells. If the RMSE on these ordered configurations exceeds about 1-2 meV, the pair-only, random-trained surrogate is not robust for Monte Carlo; if it stays near the 0.6 meV random-configuration value, the extrapolation concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All training and testing configurations are described as 'randomly drawn' (Section 3.1), so the reported RMSEs (Table 1) characterize predictions only near random disorder. The EPI Hamiltonian in Eq. (8) contains pair terms only; triplet and higher-order interactions are omitted by construction. On random configurations, omitted many-body terms may be small or partly absorbed into effective pair coefficients, but the stated purpose is to feed the surrogate into Monte Carlo simulations of order-disorder transitions, where sampled configurations develop strong short-range order. Sections 3.1-3.3 contain no validation on ordered or strongly SRO configurations, and Step 6 explicitly defers the Monte Carlo application to future work. The central claim of robustness for thermodynamics therefore rests on an untested extrapolation from the random-configuration training distribution to the ordered states Monte Carlo will visit. BIC-based shell selection, being fit to the same random configurations, cannot by itself correct this gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Bayesian regularized regression framework, combined with an effective pair interaction (EPI) model and Bayesian information criterion (BIC) based feature selection, to predict the configurational energy of refractory high entropy alloys from sparse first-principles data. The method is demonstrated on NbMoTaW, NbMoTaWV, and NbMoTaWTi, using DFT energies from supercells of 16--128 or 20--160 atoms, with training data drawn from the three smaller supercells and testing on the largest supercell. The reported held-out test RMSE is around 0.6 meV for all three alloys. The paper also introduces an ensemble sampling strategy that combines configurations from different supercell sizes, analyzes the uncertainty and correlations of the fitted pair interaction parameters, and shows that BIC-based selection of the number of coordination shells reduces the risk of overfitting and underfitting when data are limited.","tokens_in":15415,"tokens_out":4903,"duration_ms":54427,"significance":"If the claims hold, the paper offers a practical recipe for constructing surrogate Hamiltonians for multicomponent alloys with quantified parameter uncertainty, which is valuable because conventional cluster expansion becomes combinatorially intractable for high entropy alloys. The work is concrete: the algorithm steps are clearly enumerated, the DFT data generation is described, and the held-out test evaluation gives a quantitative accuracy statement. The explicit use of BIC for shell truncation and the comparison of Bayesian regression against ordinary least squares are useful methodological contributions. The main limitation is that the validation is restricted to random configurations, so the stated suitability for Monte Carlo simulations of order-disorder transitions is not directly demonstrated.","major_comments":[{"comment":"The regression model in Eq. (8) assumes that a single set of pair interaction coefficients J_m^{X,Y} describes the configurational energy across all supercell sizes, but the training data combine DFT energies from supercells of 16, 32, 64, and 128 atoms (or 20, 40, 80, and 160 atoms). If the per-atom DFT energies carry any supercell-size-dependent reference offset, for example from different Brillouin-zone sampling or from the definition of the configurational energy itself, then the fitted coefficients and the BIC-selected number of shells would be biased. The manuscript does not discuss this possibility or include a size-dependent offset term in the model. The authors should either add a supercell-size indicator or intercept in the regression, or demonstrate explicitly that no such offset is present for these materials and this DFT setup.","section":"§2.2, Eq. (8) and §3.1, ensemble sampling"},{"comment":"All training and testing configurations are described as randomly drawn, and the reported RMSE values therefore characterize prediction accuracy only for configurations near random disorder. The paper's stated purpose, however, is to feed the fitted Hamiltonian into Monte Carlo simulations for modeling thermodynamics and order-disorder transitions, where the sampled configurations develop short-range order and may include ordered states. The EPI model of Eq. (8) contains pair interactions only, and the paper does not validate the model on ordered or strongly short-range-ordered configurations. The central claim of robustness for thermodynamic applications is thus an extrapolation beyond the tested distribution. I recommend adding validation on ordered supercells or on configurations with strong SRO, or explicitly limiting the accuracy claim to random configurations.","section":"§3.1 and Step 6 (Section 2.5)"}],"minor_comments":[{"comment":"Equation (9) writes E = JP + ε with no constant term, while Eq. (5) contains a concentration-dependent constant J0 that is said to be discarded. For absolute DFT energies, the regression must include an intercept or the energies must be centered; the manuscript should state which convention is used.","section":"§2.2, Eq. (9)"},{"comment":"The text refers to 'RMSE results' for different shell numbers but does not explicitly state whether these RMSEs are computed on the held-out test supercell or on training data. Please clarify that the comparison in Figures 12--14 uses the same held-out test set as Table 1, and that BIC selection is performed using only training data.","section":"§3.3, Figures 12--14"},{"comment":"The caption of Figure 8 reads 'NbMoTaTi' but the text and Table 1 refer to NbMoTaWTi; please correct the caption.","section":"Figure 8 caption"},{"comment":"In Eq. (14), the notation N(E|JP, λ2) implies λ2 is the variance, but later in Eq. (16) λ2 is treated as a gamma-distributed precision parameter. Please make the variance/precision convention consistent.","section":"§2.3, Eq. (14)"},{"comment":"The column headers 'TrainingεR (meV)' and 'TestingεR (meV)' are missing spaces; consider formatting as 'Training εR (meV)' for readability.","section":"§3.1, Table 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript relies extensively on the authors' own prior work (refs [49]--[54]) for the EPI model, Bayesian regression, and Latin hypercube sampling. This is not a problem in itself, but the novelty relative to ref. [50] should be made clear, and the editor may wish to ask the authors to specify which aspects of the current paper are new beyond that earlier study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a useful incremental paper on fitting configurational energies of refractory HEAs from sparse DFT data. The new pieces are the BIC-based shell selection as a function of dataset size and the ensemble sampling across supercell sizes; both are demonstrated to lower held-out RMSE compared to single-supercell training. The Bayesian ridge machinery is standard and appears correctly applied, and the held-out test RMSE of ~0.6 meV on 128/160-atom supercells gives real support to the accuracy claim for the tested distribution of configurations.\n\nThe soft spots are real but not fatal. First, the paper mixes DFT total energies from different supercell sizes in one regression without discussing a per-cell-size reference offset. If the per-atom energies carry a size-dependent constant, the fitted pair coefficients and the BIC shell choice could be biased. This is addressable by adding size offsets or demeaning per supercell. Second, and more conceptually: every training and test configuration is random. The stated downstream purpose is Monte Carlo simulation of order-disorder transitions, which will visit strongly SRO and ordered configurations. The pair-only EPI model omits many-body interactions by construction, and the paper gives no evidence that it transfers to those configurations. The authors do say Monte Carlo is future work (Step 6), so the claim \"robust for thermodynamics\" is overstated; the paper is really about robust prediction on random configurations with sparse data. That is still a legitimate contribution, but the title and abstract should be scoped accordingly. Minor issues: no error bars on the headline RMSE values, no code/data release, and DFT settings are not detailed enough for reproduction.\n\nThe citation pattern is fine; the self-citations are to the EPI model and the LHS tools, which are relevant. I did not find the circularity that sometimes comes with data-driven papers: the test set is held out and from a larger supercell.\n\nRecommendation: worth a serious referee. It needs major revision to temper the thermodynamic claim and to address the size-offset and validation-on-SRO questions, but the core method and experiments are solid enough that a good referee could turn it into a citable paper.","headline":"Solid incremental methods paper for sparse-data HEAs configurational energy; accuracy holds on random configs but thermodynamic claims outrun the validation.","tokens_in":15933,"tokens_out":2222,"would_cite":true,"duration_ms":23078,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A pair-interaction model fitted with Bayesian regression and BIC shell selection predicts the configurational energy of refractory high-entropy alloys with a held-out RMSE around 0.6 meV.","keywords":["high entropy alloys","configurational energy","effective pair interaction model","Bayesian regularized regression","Bayesian information criterion","uncertainty quantification","refractory alloys","first-principles calculations"],"falsifier":"Refit Eq. (8) with an additive supercell-size offset and with triplet correlation features on the same DFT data; if either change materially lowers the held-out RMSE or shifts the BIC-selected shell count for these alloys, the pair-only, size-independent assumption behind the 0.6 meV claim is falsified.","tokens_in":15026,"feed_emoji":"⚛️","tokens_out":8715,"duration_ms":84415,"temperature":0.7,"pith_summary":"This paper argues that a deliberately simple Hamiltonian—only pair interactions between atoms, truncated at a coordination shell chosen by Bayesian model selection—is enough to predict the configurational energy of refractory high-entropy alloys from sparse DFT data. The authors fit the effective pair interaction (EPI) model to randomly drawn configurations of NbMoTaW, NbMoTaWV, and NbMoTaWTi using Bayesian regularized regression, and report held-out testing RMSE values around 0.6 meV for all three alloys. They further show that choosing the number of shells with the Bayesian information criterion avoids both overfitting on small datasets and underfitting on larger ones, and that pooling configurations from several supercell sizes gives more stable predictions than training on any single small supercell. If correct, the payoff is a cheap substitute for direct DFT that can be fed into Monte Carlo simulations of order-disorder transitions.","feed_headline":"Sparse Bayesian model predicts high-entropy alloy energies to 0.6 meV","feed_subtitle":"Bayesian ridge regression and shell selection turn sparse DFT data into Monte Carlo-ready Hamiltonians.","key_machinery":"The load-bearing object is the effective pair interaction (EPI) model, an Ising-like Hamiltonian that maps a configuration to an energy through the probabilities $P^{X|Y}_m$ of finding element $X$ in the $m$-th coordination shell around element $Y$, for each independent pair $X\\neq Y$. Three pieces carry the argument: Bayesian regularized regression with conjugate gamma priors on the noise and coefficient precisions, which supplies stable point estimates and uncertainty quantification for the pair coefficients; the BIC shell-selection criterion, which decides how many of the shells to keep; and ensemble sampling, which draws configurations from several supercell sizes so the training data include both short-range and long-range order.","core_discovery":"The central claim is that the configurational energy of a multicomponent alloy can be written, to good accuracy, as a sum of chemically distinct pair probabilities in the first few coordination shells, $E = N\\sum_{X\\neq Y,m} J^{X,Y}_m P^{X|Y}_m$, and that the coefficients $J^{X,Y}_m$ can be estimated reliably by Bayesian $\\ell^2$-regularized regression even when the number of DFT configurations is small. To set model complexity, the paper minimizes the Bayesian information criterion expressed through the residual sum of squares, $\\mathrm{BIC}_{\\mathrm{RSS}} = n_d \\log(\\mathrm{RSS}/n_d) + k\\log(n_d)$, and shows that the selected shell count grows sensibly with dataset size. With the BIC-selected shells and an ensemble sampling strategy that mixes supercells of 16–160 atoms, the fitted models reach testing RMSE about 0.6 meV on the three refractory alloys, and the uncertainty in the pair interactions drops sharply once a few hundred configurations are available.","pith_inferences":["Because the EPI model is pair-only, a direct extension would be to add triplet correlation features and compare BIC; if triplets systematically lower held-out RMSE on these alloys, the 0.6 meV claim would need to be weakened.","The BIC-selected shell count can be read as a measured interaction range, which suggests a testable prediction: independent electronic-structure calculations should find longer-ranged or more frustrated effective interactions in NbMoTaWV and NbMoTaWTi than in NbMoTaW.","A stricter stress test than the paper reports is to train the model on one refractory alloy and predict another; success would indicate the pair coefficients capture transferable ordering physics, while failure would mean they encode chemistry-specific fits.","The paper stops short of propagating the Bayesian parameter uncertainties into Monte Carlo free energies; doing so would reveal whether a 0.6 meV energy error is small enough for accurate order-disorder transition temperatures."],"forward_implications":["A Monte Carlo simulation of order-disorder transitions can replace direct DFT calls with the fitted EPI Hamiltonian, because the surrogate predicts large-supercell energies from small-supercell training data to within about 1 meV.","The BIC shell count supplies a data-size-dependent recipe—2–3 shells for small datasets, 5–6 for medium ones, 6–9 for large ones—so model complexity no longer has to be set arbitrarily.","Ensemble sampling across supercell sizes should be part of similar data-driven alloy models, since it lowers both the mean and the standard deviation of the prediction error compared with training on a single small supercell.","Increasing the number of training configurations from 100 to 400 reduces the variance of the fitted pair interactions by more than an order of magnitude, so the framework can tell users how much confidence to place in each energy estimate.","For NbMoTaWV and NbMoTaWTi the best model retains more coordination shells than for NbMoTaW, indicating that nearest-neighbor-only pair models would systematically underfit those alloys."],"supporting_citations":[{"why":"Supplies the EPI model and the pair-probability features used for the configurational-energy regression.","marker":"[50]"},{"why":"Provides the locally self-consistent multiple-scattering DFT method generating the training and test energies.","marker":"[53]"},{"why":"Precedent for Bayesian fitting of cluster expansions that the regularized regression extends.","marker":"[22]"},{"why":"Gives the sparsity argument for physical interactions that motivates feature selection and regularization.","marker":"[21]"},{"why":"Basis for the Bayesian evidence approximation behind the BIC shell-selection criterion.","marker":"[52]"},{"why":"Latin hypercube sampling procedure used in the ensemble sampling strategy across supercell sizes.","marker":"[54]"},{"why":"General cluster expansion formalism whose many-body terms the EPI model truncates to pair interactions.","marker":"[19]"}],"fun_headline_variants":["Sparse data, precise energies: Bayesian model hits 0.6 meV","Bayesian shell selection: accurate alloy energies from tiny datasets","Data-efficient Bayesian model predicts alloy energies within 0.6 meV","Tiny DFT datasets, big accuracy: Bayesian model for alloy energies","Shell selection + Bayesian regression = 0.6 meV alloy energies from few data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model only works if the configurational energy is fully determined by two-body pair probabilities with coefficients that do not depend on supercell size, so significant triplet or many-body interactions, or a size-dependent offset in the DFT energies, would bias the fitted coefficients and the reported 0.6 meV accuracy.","fun_headline_variants_meta":{"raw":{"variants":["Sparse data, precise energies: Bayesian model hits 0.6 meV","Bayesian shell selection: accurate alloy energies from tiny datasets","Data-efficient Bayesian model predicts alloy energies within 0.6 meV","Tiny DFT datasets, big accuracy: Bayesian model for alloy energies","Shell selection + Bayesian regression = 0.6 meV alloy energies from few data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000503,"raw_usage":{"total_tokens":2479,"prompt_tokens":990,"completion_tokens":1489,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":1392}},"tokens_in":606,"tokens_out":1489,"duration_ms":12003,"temperature":1.0,"reasoning_tokens":1392,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:06:08.847602+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Refit Eq. (8) with an additive supercell-size offset and with triplet correlation features on the same DFT data; if either change materially lowers the held-out RMSE or shifts the BIC-selected shell count for these alloys, the pair-only, size-independent assumption behind the 0.6 meV claim is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the locally self-consistent multiple-scattering DFT method generating the training and test energies."},{"cited_title":"Mueller, G","cited_arxiv_id":null,"evidence_quote":"Precedent for Bayesian fitting of cluster expansions that the regularized regression extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the sparsity argument for physical interactions that motivates feature selection and regularization."},{"cited_title":"Zhang, M","cited_arxiv_id":null,"evidence_quote":"Basis for the Bayesian evidence approximation behind the BIC shell-selection criterion."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Latin hypercube sampling procedure used in the ensemble sampling strategy across supercell sizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"General cluster expansion formalism whose many-body terms the EPI model truncates to pair interactions."}],"review_version":1}