{"id":"58749219-af40-4e94-8e5c-6cc587bee770","arxiv_id":"2606.01305","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Hybrid ML models learn Redlich-Kister coefficients from elemental descriptors to enable zero-shot extrapolation of CALPHAD interaction parameters for unseen elements in FCC alloys.","lead":"This paper shows machine learning can predict Redlich-Kister interaction coefficients for alloy free energies using elemental descriptors, tested on formation energies from a machine-learning interatomic potential for 14 FCC elements. A smart generalist might read it to see how data-driven methods could reduce the experimental burden for building thermodynamic databases in new material systems.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Universal MLIP energies used as ground truth lack reported DFT/experimental validation for the 14-element FCC set","rationale":"The reader's weakest_assumption matches the load-bearing point exactly. Because the abstract (and therefore the central claim) positions the MLIP energies as the sole source of ground truth for benchmarking all three model classes, any unquantified surrogate error directly undermines the unification and extrapolation assertions. No other internal inconsistency is visible from the given material; the concern is therefore isolated to this unverified data-generation step.","tokens_in":1726,"tokens_out":367,"duration_ms":12275,"concrete_test":"Pick 8–10 binary/ternary FCC compositions spanning the 14 elements (including at least two leave-one-element-out cases); recompute formation energies with DFT using the same lattice and k-point settings as the MLIP training; if MAE > 8 meV/atom or if errors correlate with descriptor distance from training elements, retrain the RK predictors on the DFT subset and re-evaluate the hybrid advantage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The hybrid ML4RK claim requires that the surrogate formation energies accurately reflect true thermodynamics so that learned RK coefficients and zero-shot extrapolation performance are meaningful. The paper generates all training/benchmark data from one universal MLIP and treats those values as ground truth for both composition-based and descriptor-based models. No section quantifies MLIP error against DFT (or experiment) across the composition space or for out-of-distribution elements; universal MLIPs are known to exhibit composition-dependent biases in multi-principal-element alloys. If those biases are present, the reported complementarity between RK, pure ML, and ML4RK, as well as the extrapolation advantage, could be artifacts of the surrogate rather than physical behavior.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that embedding elemental descriptors into the Redlich-Kister (RK) framework yields a hybrid ML4RK model that unifies data-efficient RK fitting (when binaries are known) with zero-shot extrapolation to unseen elements, demonstrated via leave-one-element-out benchmarking on formation energies of 14-element FCC alloys generated by a universal MLIP. Three model classes are compared: composition-based RK/ML, descriptor-based ML, and the hybrid.","tokens_in":1873,"tokens_out":364,"duration_ms":7845,"significance":"If the surrogate energies are faithful, the hybrid approach would offer a physically interpretable route to extend CALPHAD interaction parameters to data-scarce or unknown binaries while retaining the formalism's robustness. The leave-one-element-out design directly tests transferability, which is a genuine strength for the claimed unification of regimes.","major_comments":[{"comment":"Abstract and benchmarking description: all training, validation, and test data are generated from a single universal MLIP and treated as ground truth, yet no section quantifies MLIP error against DFT or experiment for the 14-element FCC set or for out-of-distribution compositions. Because the reported complementarity between RK, pure ML, and ML4RK (and the extrapolation advantage) rests entirely on these surrogate values, any composition-dependent bias in the MLIP would render the benchmarking results non-physical.","section":"Abstract / benchmarking description"}],"minor_comments":[{"comment":"The abstract states that RK models are 'most data-efficient when binary information is available' but supplies no quantitative metrics, error bars, or tables comparing RMSE or parameter counts across the three classes.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive comment on our benchmarking approach. We respond to the major comment point by point below.","responses":[{"response":"We agree that the study relies exclusively on formation energies generated by a single universal MLIP, treated as ground truth for benchmarking purposes. This choice was deliberate to enable a large, internally consistent dataset across 14 elements, which supports the leave-one-element-out protocol necessary to evaluate zero-shot extrapolation to unseen elements. We acknowledge that the absence of direct quantification of MLIP errors versus DFT or experiment for the specific 14-element FCC set (including out-of-distribution compositions) means that any composition-dependent biases in the surrogate could influence the observed complementarity and extrapolation performance, rendering the results relative rather than absolute. In the revised manuscript we will add a dedicated paragraph in the Methods section summarizing the published validation metrics of the universal MLIP against DFT for relevant alloy systems, and we will revise both the abstract and the benchmarking description to explicitly note that all comparisons are performed within this surrogate landscape. We will also add a limitations subsection discussing the implications of potential MLIP biases for physical transferability. These textual revisions directly address the concern while preserving the demonstration of the hybrid ML4RK framework.","revision_made":"yes","referee_comment":"[Abstract / benchmarking description] Abstract and benchmarking description: all training, validation, and test data are generated from a single universal MLIP and treated as ground truth, yet no section quantifies MLIP error against DFT or experiment for the 14-element FCC set or for out-of-distribution compositions. Because the reported complementarity between RK, pure ML, and ML4RK (and the extrapolation advantage) rests entirely on these surrogate values, any composition-dependent bias in the MLIP would render the benchmarking results non-physical."}],"tokens_in":1329,"tokens_out":385,"duration_ms":19995,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's core move is to train models that predict Redlich-Kister interaction parameters from elemental descriptors rather than fitting them only to composition data. This lets the hybrid ML4RK model do zero-shot extrapolation to binaries involving elements left out of training, while still keeping the RK functional form. That unification of the data-efficient RK regime with the extrapolative strength of descriptor ML is the part that is not already in the cited prior work.\n\nWhat works is the clean leave-one-element-out setup across the three model classes on the 14-element FCC set. It shows RK staying strong when some binary data exists and pure descriptor ML handling the no-data case, with the hybrid sitting in between. The framing around extending CALPHAD to data-scarce systems is reasonable.\n\nThe soft spot is exactly the one flagged in the stress test. All training and test formation energies come from a single universal MLIP with no reported comparison to DFT or experiment for these compositions or for the held-out elements. Universal MLIPs can carry composition-dependent biases, especially in multi-principal-element space. If those biases are present, the reported complementarity and the extrapolation gains could be artifacts of the surrogate rather than physical behavior. The abstract also gives no error bars, no quantitative metrics, and no external validation, so the strength of the claims is hard to judge from what is shown.\n\nThis is for people building or extending CALPHAD databases who already work with MLIPs or descriptors. It is worth sending to peer review because the modeling idea is concrete and the benchmarking design is straightforward; the main requirement would be adding the missing validation against first-principles or experimental data before the results can be taken as reliable.","headline":"The hybrid ML4RK approach that folds elemental descriptors into Redlich-Kister coefficients is the actual new piece, but the whole evaluation sits on unvalidated universal MLIP energies treated as ground truth.","tokens_in":2369,"tokens_out":425,"would_cite":false,"duration_ms":11350,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Machine learning predicts Redlich-Kister coefficients for unknown alloy binaries by using elemental descriptors.","keywords":["CALPHAD","Redlich-Kister","machine learning","alloy thermodynamics","interaction parameters","elemental descriptors","FCC alloys"],"falsifier":"Direct comparison of ML4RK-predicted interaction parameters against measured thermodynamic data for at least one binary system whose elements were held out during training.","tokens_in":2635,"feed_emoji":"","tokens_out":545,"duration_ms":13458,"temperature":0.7,"pith_summary":"The paper establishes that a hybrid ML-augmented Redlich-Kister model learns interaction coefficients directly from elemental descriptors, allowing prediction of parameters for binaries lacking data. This hybrid unifies the data efficiency of standard RK models when binary information exists with the zero-shot extrapolation of pure descriptor-based ML models when elements are absent from training. A sympathetic reader would care because CALPHAD thermodynamic modeling has long been limited by sparse experimental data and composition-only functional forms, restricting predictions for new alloy chemistries.","feed_headline":"ML predicts Redlich-Kister parameters for unknown alloy binaries","feed_subtitle":"Elemental descriptors enable extrapolation to data-scarce systems while retaining thermodynamic structure.","key_machinery":"The ML-augmented Redlich-Kister (ML4RK) model that learns RK interaction coefficients from physically informed elemental descriptors.","core_discovery":"Using formation energies of 14-element FCC alloys from a universal MLIP, the hybrid ML4RK approach embeds elemental descriptors into the RK framework to predict interaction parameters for otherwise unknown or data-scarce binaries, as shown in leave-one-element-out tests that demonstrate complementary strengths across model classes.","pith_inferences":["The same descriptor embedding could be tested on other thermodynamic models beyond RK polynomials.","If the MLIP accuracy assumption holds, the method could accelerate screening of high-entropy alloys before any experimental binary measurements.","A natural extension would be to incorporate temperature dependence or multi-component interactions using the same descriptor foundation."],"forward_implications":["RK models remain most data-efficient when binary data is available.","Descriptor-based ML enables genuine extrapolation to entirely new elements.","The hybrid unifies both regimes for broader coverage of alloy systems.","This route combines ML transferability with the interpretability and robustness of thermodynamic formalisms."],"fun_headline_variants":["ML4RK embeds descriptors into RK for unknown alloy binaries","Hybrid ML augments Redlich-Kister to predict data-scarce binaries","Elemental descriptors extend RK models beyond training elements","ML learns RK coefficients from physical descriptors in CALPHAD","Descriptor-based ML unifies with thermodynamic formalisms for alloys"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Formation energies generated by the universal MLIP are accurate enough to serve as ground truth for training and benchmarking the RK coefficient predictors.","fun_headline_variants_meta":{"raw":{"variants":["ML4RK embeds descriptors into RK for unknown alloy binaries","Hybrid ML augments Redlich-Kister to predict data-scarce binaries","Elemental descriptors extend RK models beyond training elements","ML learns RK coefficients from physical descriptors in CALPHAD","Descriptor-based ML unifies with thermodynamic formalisms for alloys"]},"model":"grok-4.3","cost_usd":0.003941,"raw_usage":{"total_tokens":2014,"prompt_tokens":660,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":39412000,"prompt_tokens_details":{"text_tokens":660,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1274,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":660,"tokens_out":80,"duration_ms":9018,"temperature":1.0,"reasoning_tokens":1274,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T16:51:50.040057+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Direct comparison of ML4RK-predicted interaction parameters against measured thermodynamic data for at least one binary system whose elements were held out during training.","supporting_citations":[],"review_version":1}