{"id":"fb5908bc-4592-48c4-8479-360d6a0abc97","arxiv_id":"2504.18366","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A neural network trained on 91 binary steel systems predicts mixing enthalpies for unassessed binary liquid alloys and outputs Redlich-Kister parameters for CALPHAD databases.","lead":"This paper trains a neural network to predict how much heat is released or absorbed when pairs of liquid metals are mixed, using a steel thermodynamics database. It converts those predictions into the standard Redlich-Kister format, which lets CALPHAD databases be extended without new experiments.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The sub-1 kJ/mol uncertainty bound for extrapolated binaries is asserted, not demonstrated: Sec. 3.3 checks consistency with CALPHAD data already in the database, while the actual missing Al-Sn/Al-Sb predictions have no ground-truth comparison.","rationale":"The reader's CONDITIONAL verdict is appropriate. The paper is honest about the Fe-Sn and Fe-Sb LOOCV failures and provides open data and code, which are real strengths. My stress test does not reject the method outright; it identifies that the headline uncertainty number is not supported for the actual use case, namely extrapolation to binaries absent from the database. The LOOCV failure for Fe-Sn/Fe-Sb is closely related but not identical to the Al-Sn prediction, because the final model used for Al-Sn includes Fe-Sn in training. The central gap is that no test compares extrapolated predictions to physical reality: the uncertainty quantification compares against CALPHAD values for systems already in the database, and the paper's own Sec. 3.4 limitation statement concedes that sparse-element systems perform worse while asserting a belief rather than a measured bound. A single independent experimental check on one missing sparse-element binary would settle whether the <1 kJ/mol claim holds. Until then, the database amendments should be treated as provisional, which is exactly what the CONDITIONAL verdict conveys.","tokens_in":9579,"tokens_out":5563,"duration_ms":60833,"concrete_test":"Extract the published RK parameters for an originally missing binary with a sparse element, e.g., Al-Sn from Table S3, and compare the resulting Hmix curve at 1873 K against independent calorimetric data not used in training. If the maximum absolute deviation exceeds 1 kJ/mol at any composition, the abstract's uncertainty claim fails for the amended database; if no such experimental data is available, retrain the model with all Sn-containing systems removed and quantify the cold-start error, which the paper already shows is large for Fe-Sn.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the abstract claim that the model predicts Hmix of systems not in training with uncertainty below 1 kJ/mol, the paper needs a validated error estimate for genuinely extrapolated binaries. The evidence offered is weaker than that claim. The 10-fold CV in Sec. 3.1 splits points randomly, so every binary system appears in the training set; it cannot test unseen systems. The LOOCV in Sec. 3.2 is the right test, but it covers only Fe-X systems already in the database and fails for Fe-Sn and Fe-Sb when Sn/Sb disappear from the training data. The uncertainty quantification in Sec. 3.3 removes a primary and one correlated secondary system, then compares predictions with the CALPHAD values of the primary, which are derived from the same database family used to train the model. This measures internal consistency, not physical accuracy, and it is performed for Cu-based systems with many Fe-Z training systems. The final paragraph of Sec. 3.4 concedes that sparse As/Sn/Sb systems 'seem to perform relatively worse' and then states 'we believe' the extrapolated error is below 1 kJ/mol. The actual deliverable includes missing binaries such as Al-Sn and Al-Sb, where the only training information about Sn or Sb is the single Fe-Sn or Fe-Sb system; no held-out data exists to test that transfer. If the true error in those predictions exceeds 1 kJ/mol, reintegrating the RK parameters into a CALPHAD database would introduce biased enthalpies while the stated uncertainty understates the risk.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a neural-network model trained on CALPHAD-derived and experimental liquid-phase mixing-enthalpy data for 91 binary systems of 22 elements at 1873 K, using Matminer elemental and composition features. The authors report a 10-fold cross-validated MAE of 16.52 J/mol, a leave-one-system-out test on Fe-X binaries, and an uncertainty analysis that removes a primary and one correlated secondary system. They then use the model to predict Hmix for all missing binaries, extract Redlich-Kister parameters, and claim that the uncertainty of the predicted mixing enthalpy is below 1 kJ/mol, so that the parameters can be reintegrated into a CALPHAD database.","tokens_in":9882,"tokens_out":2408,"duration_ms":25206,"significance":"If the extrapolation claim were fully supported, this would be a valuable tool for augmenting CALPHAD databases with first-pass descriptions of unassessed binary systems, especially for tramp elements relevant to steel recycling. The paper has genuine strengths: the LOOCV protocol is the right kind of test for unseen systems; the authors openly report the Fe-Sn and Fe-Sb failures; the data and code are shared through a Zenodo repository; and the Redlich-Kister extraction addresses practical use by the CALPHAD community. However, the evidence as presented supports a more limited claim than the abstract makes: reliable extrapolation is demonstrated only for systems whose elements appear elsewhere in the training set, and the sub-1 kJ/mol uncertainty bound is asserted for systems that were not tested against any held-out data.","major_comments":[{"comment":"The headline MAE of 16.52 J/mol is reported for the best of 10 folds in a point-level k-fold split, so every binary system appears in the training set. This metric cannot support the abstract's claim that the model predicts mixing enthalpy for systems not present in the training dataset. The paper should either report metrics from the LOOCV as the primary out-of-system measure or clearly label the k-fold result as an interpolation test.","section":"Sec. 3.1 and Fig. 3"},{"comment":"The LOOCV is performed only on Fe-X systems, and it fails outright for Fe-Sn and Fe-Sb because no other Sn- or Sb-containing system remains in training. Yet the final deliverable includes missing binaries such as Al-Sn and Al-Sb, where the only training information about Sn or Sb is the single Fe-Sn or Fe-Sb system. The paper does not provide a held-out test for this regime, so the transferability that is load-bearing for those predictions is not demonstrated. At minimum, the abstract and Sec. 3.4 should be reworded to state that reliable extrapolation is conditional on the element being represented elsewhere in the training set.","section":"Sec. 3.2 and Sec. 3.4, Fig. 7"},{"comment":"The uncertainty quantification removes a primary and one correlated secondary system and compares the resulting predictions with CALPHAD values from the same database family used for training. This measures internal consistency of the database, not physical accuracy, and it is shown only for Cu-based systems with many Fe-Z training systems. The abstract's statement that 'the estimated uncertainty of the model is below 1 kJ/mol for the predicted mixing enthalpy' is therefore not supported for the genuinely extrapolated predictions. The paper should either provide a validated error estimate for missing systems (e.g., against experimental data for at least a few held-out systems) or replace the uncertainty claim with a clearly labeled consistency measure.","section":"Sec. 3.3 and abstract"},{"comment":"The authors concede that As-X, Sn-X, and Sb-X systems 'seem to perform relatively worse' and then state 'we believe' that extrapolated errors will be below 1 kJ/mol. This is an explicit admission that the central uncertainty claim is not demonstrated for the systems that matter most for the proposed application. The manuscript should either provide quantitative evidence for the error on such systems or remove the unqualified sub-1 kJ/mol claim from the abstract and summary.","section":"Sec. 3.4, final paragraph"}],"minor_comments":[{"comment":"The phrase 'amended with several direct experimental reports' is vague; the paper should specify which systems and how many experimental data points were added.","section":"Abstract"},{"comment":"The Redlich-Kister expansion is written with indices k=1...n, whereas the CALPHAD convention typically starts at k=0; this may confuse readers comparing with standard database files.","section":"Eq. (1)"},{"comment":"The caption abbreviates RMSE as 'RSME'; please correct the typo.","section":"Fig. 6 caption"},{"comment":"The heatmap uses green and red fields, but the caption does not state the color convention explicitly; please add a legend or explicit sentence.","section":"Sec. 2.1 and Fig. 1"},{"comment":"The precision of the tabulated RK parameters is much higher than the claimed physical accuracy; the footnote is helpful but could be strengthened by adding an explicit statement of the implied uncertainty in the parameter values.","section":"Sec. 2.4 and Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper has a useful core and honest reporting of limitations, but the abstract and summary overstate what is demonstrated. The key fix is to align the central claims with the actual evidence: point-level k-fold, LOOCV on Fe-X only, and a consistency-based uncertainty estimate do not together support a sub-1 kJ/mol guarantee for unseen binaries, especially for Al-Sn and Al-Sb. This is a correctable framing issue rather than a fatal flaw, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid engineering paper, not a breakthrough. The genuinely new bits are the experiment-amended MatCalc steel database, the leave-one-system-out test, the double-quarantine UQ scheme, and the Redlich-Kister parameter extraction for 140 missing binaries. That last deliverable is useful: practitioners can import RK coefficients and see what the model suggests, even if they spot-check before using.\n\nThe LOOCV is the right test and they run it properly on Fe-X systems. They report failures honestly: Fe-Sn and Fe-Sb collapse when Sn and Sb vanish from training. The double-quarantine experiment gives a real sense of sensitivity for Cu-based systems. Data and code are on Zenodo. That counts.\n\nNow the soft spots, in proportion. The 16.52 J/mol MAE is from the best of 10 point-level folds, so it says nothing about unseen systems. The LOOCV only covers Fe-X, and only when the other element is represented elsewhere. The uncertainty section checks consistency with CALPHAD values from the same database family, not physical ground truth; the 0.5 kJ/mol standard deviation measures training-data sensitivity, not prediction error. The abstract's \"below 1 kJ/mol\" for truly extrapolated binaries is not supported. The final paragraph admits sparse As/Sn/Sb systems \"seem to perform relatively worse\" and then says \"we believe\" the error is below 1 kJ/mol. That is belief, not evidence. For the actual deliverable, including Al-Sn and Al-Sb, no held-out test exists. The risk is real: putting RK parameters into CALPHAD with an uncertainty that understates the error could quietly bias downstream phase-equilibria predictions.\n\nThe citation pattern is fine; they credit Deffrennes et al. and other prior ML-CALPHAD work. The paper also states its own limitations, which I appreciate. It works at one temperature and for binaries only, and they say so.\n\nWho is this for? CALPHAD practitioners who want a quick way to generate provisional binary liquid descriptions, especially for tramp-element systems in steel recycling. It is not a general-purpose thermodynamic model. A serious referee should engage with it, but the review should ask for system-level error metrics (LOOCV MAE/RMSE), an explicit per-element support analysis, and either experimental spot-checks or a downgraded uncertainty claim. I would send it to peer review and would not desk-reject.","headline":"Useful ML pipeline for filling missing binary liquid Hmix in steel CALPHAD databases, with honest LOOCV, but the <1 kJ/mol uncertainty claim for extrapolated systems is asserted rather than demonstrated.","tokens_in":10447,"tokens_out":1680,"would_cite":true,"duration_ms":17314,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A neural network trained on 91 known binary liquid alloys can predict the mixing enthalpy of unseen binary systems with estimated uncertainty below 1 kJ/mol, and the resulting Redlich-Kister parameters can be written directly into CALPHAD…","keywords":["CALPHAD","neural network","mixing enthalpy","liquid phase","Redlich-Kister parameters","leave-one-out cross-validation","thermodynamic databases","materials informatics"],"falsifier":"Measure the mixing enthalpy of a predicted-but-unseen binary system such as Al-Sn or Al-Sb by high-temperature calorimetry at 1873 K and compare the curve to the model's prediction; if the deviation exceeds the claimed 1 kJ/mol uncertainty at any composition, the central uncertainty claim is falsified for that system.","tokens_in":9382,"feed_emoji":"🧪","tokens_out":6589,"duration_ms":60373,"temperature":0.7,"pith_summary":"The paper tries to establish that a neural network can predict the mixing enthalpy of liquid binary alloy systems that were entirely absent from its training data, with an estimated uncertainty below 1 kJ/mol. If true, the approach turns an existing thermodynamic database into a generator of Redlich-Kister parameters for the 140 missing binary combinations among 22 elements, so those parameters can be plugged directly back into CALPHAD databases. The authors argue this can spare substantial experimental calorimetric work on systems such as Fe-Sn, Fe-Sb, and Al-Sb that are becoming relevant due to scrap recycling in steel production. The claim is demonstrated through k-fold cross-validation, leave-one-out cross-validation, and an uncertainty quantification that removes two correlated systems at a time.","feed_headline":"Neural net predicts unseen binary alloy enthalpies under 1 kJ/mol","feed_subtitle":"Redlich-Kister parameters for all 140 missing binary systems are ready for CALPHAD databases.","key_machinery":"The central object is a neural network regression model mapping a descriptor vector of a binary liquid at a given composition to its mixing enthalpy $H_{\\mathrm{mix}}$. The descriptors combine 151 atomic features, 181 composition-based features, plus Miedema and Yang alloy features extracted with Matminer, chosen to encode electronegativity differences, orbital overlaps, and semi-empirical enthalpic trends. The model is trained at a fixed temperature, 1873 K, on 2000 equidistant composition points per binary system extracted from the MatCalc open database with pycalphad, augmented by experimental enthalpy data. Validation uses 10-fold cross-validation and leave-one-system-out cross-validation to test genuine extrapolation; the final predictions are fitted to a fourth-order Redlich-Kister polynomial $H_{\\mathrm{mix}} = x_1 x_2 \\sum_{k=1}^{4} L_k (x_1 - x_2)^{k-1}$, whose parameters are tabulated for direct insertion into a CALPHAD database file.","core_discovery":"On its own terms, the paper claims that a feedforward neural network, trained on 91 binary liquid-phase systems from an open steel database supplemented with experimental mixing enthalpies, learns composition-dependent interaction behavior well enough to predict mixing enthalpy curves for binary systems it has never seen. The key demonstration is leave-one-out validation: when a whole binary system such as Fe-Cu is quarantined, the network reproduces the CALPHAD curve, including demixing, as long as some other binary containing each element (Fe and Cu) remains in training. The two failures, Fe-Sn and Fe-Sb, occur precisely when no other Sn- or Sb-containing system is available. The paper further reports that a model trained on all data achieves a mean absolute error of 16.52 J/mol on held-out points and R²=0.99, and that uncertainty from removing a second correlated system stays below about 0.5 kJ/mol, leading to the claim that extrapolated predictions for the 140 missing binaries carry errors under 1 kJ/mol.","pith_inferences":["The paper's own LOOCV failures for Fe-Sn and Fe-Sb imply that the 140 'missing' predictions should be stratified by how many correlated binaries exist for the involved elements; binaries pairing two sparse elements (e.g., As-Sn, As-Sb, Sn-Sb) are the least supported, and a calorimetric spot-check on one of them would be the sharpest test of the 1 kJ/mol claim.","Because the descriptor set includes Miedema and Yang features that already encode semi-empirical estimates of mixing enthalpy, part of the apparent predictive accuracy may be inherited from those physics-informed features rather than learned from the thermodynamic database; ablating those features and retraining would separate the two contributions.","The uncertainty-quantification procedure suggests a natural active-learning loop: remove the binary whose removal most degrades predictions for a target system, run an experiment there, add it to the training set, and repeat to build a complete 22-element matrix with minimal calorimetric effort."],"forward_implications":["All 140 missing binary liquid systems among the 22 considered elements become available as fitted Redlich-Kister parameters, so a CALPHAD database can be amended without new calorimetric measurements.","Systems relevant to scrap recycling, such as Fe-Sn, Fe-Sb, Fe-As, Al-Sn, and Al-Sb, receive a first thermodynamic description at 1873 K, useful for initial phase-stability screening.","The same workflow can be extended to temperature-dependent training data, since the constant-temperature restriction is presented as a proof of concept rather than a fundamental limit.","The predicted parameters are formatted to be read directly into existing thermodynamic database files, making the output usable by standard CALPHAD software.","If the uncertainty claim holds, the model provides quantitatively reliable interpolation across the binary composition matrix, complementing prior work on ternary extrapolation from known binaries."],"supporting_citations":[{"why":"Supplies the 91 binary liquid-phase systems and their thermodynamic descriptions used as the training basis.","marker":"[10]"},{"why":"Extracts 2000 equidistant Hmix data points per binary system from the thermodynamic database.","marker":"[11]"},{"why":"Justifies the fixed temperature 1873 K at which all predictions are made.","marker":"[12]"},{"why":"Provides the atomic and composition-based descriptors, including Miedema and Yang features.","marker":"[17]"},{"why":"The data-driven comparison study for liquid-phase mixing enthalpy that this work extends.","marker":"[5]"},{"why":"One of the experimental sources used to amend the database with direct mixing enthalpy data.","marker":"[13]"},{"why":"Supplies experimental Cu-Fe mixing enthalpy values used to supplement the training set.","marker":"[14]"},{"why":"Provides calorimetric mixing enthalpy measurements for liquid iron alloys used as additional training data.","marker":"[15]"}],"fun_headline_variants":["AI learns binary enthalpy trends, predicts unseen alloys to <1 kJ/mol","Neural net extends CALPHAD to 140 binary systems","Deep learning predicts mixing enthalpy for unknown binary alloys","AI fills CALPHAD gaps: predicts enthalpies for 140 missing binaries"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the premise that interaction knowledge from binaries containing an element transfers to unseen binaries with that element—a premise the paper shows fails when a quarantined element disappears from training (Fe-Sn and Fe-Sb), so the blanket 1 kJ/mol claim for all 140 missing binaries assumes this transfer works for every element, including sparse As, Sn, and Sb.","fun_headline_variants_meta":{"raw":{"variants":["AI learns binary enthalpy trends, predicts unseen alloys to <1 kJ/mol","Neural net extends CALPHAD to 140 binary systems","Deep learning predicts mixing enthalpy for unknown binary alloys","AI fills CALPHAD gaps: predicts enthalpies for 140 missing binaries"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000818,"raw_usage":{"total_tokens":3605,"prompt_tokens":990,"completion_tokens":2615,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":2542}},"tokens_in":606,"tokens_out":2615,"duration_ms":18997,"temperature":1.0,"reasoning_tokens":2542,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:18:27.835071+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the mixing enthalpy of a predicted-but-unseen binary system such as Al-Sn or Al-Sb by high-temperature calorimetry at 1873 K and compare the curve to the model's prediction; if the deviation exceeds the claimed 1 kJ/mol uncertainty at any composition, the central uncertainty claim is falsified for that system.","supporting_citations":[{"cited_title":"Povoden-Karadeniz, Open databases for MatCalc, https://www.matcalc.at/index.php/ databases/open-databases, [Online; accessed 11-Dec-2024] (2024)","cited_arxiv_id":null,"evidence_quote":"Supplies the 91 binary liquid-phase systems and their thermodynamic descriptions used as the training basis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extracts 2000 equidistant Hmix data points per binary system from the thermodynamic database."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies the fixed temperature 1873 K at which all predictions are made."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the atomic and composition-based descriptors, including Miedema and Yang features."},{"cited_title":"Tanaka, N","cited_arxiv_id":null,"evidence_quote":"One of the experimental sources used to amend the database with direct mixing enthalpy data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies experimental Cu-Fe mixing enthalpy values used to supplement the training set."},{"cited_title":"Iguchi, Y","cited_arxiv_id":null,"evidence_quote":"Provides calorimetric mixing enthalpy measurements for liquid iron alloys used as additional training data."}],"review_version":1}