{"id":"8bbf75ae-6c2b-4b00-9352-0c8074323632","arxiv_id":"2412.19353","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"SuperSalt is a MACE-based machine learning potential that predicts thermophysical properties of 11-cation chloride melts with near-DFT accuracy and enables Bayesian optimization of salt compositions.","lead":"Researchers built a single machine-learning force field, SuperSalt, that models molten chloride salts made from 11 different metal cations with accuracy close to density functional theory but at a fraction of the cost. The model predicts density, heat capacity, expansion, and structure across thousands of possible salt mixtures, and a Bayesian search finds target compositions in a few iterations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The transferability claim rests on MACE-MP0-generated training and test trajectories; if MACE-MP0 under-samples Zr/Zn-rich liquid configurations, DFT labels and the shared-generator Test1/Test2 sets cannot reveal the gap.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the workflow relies on MACE-MP0-generated configurations to span the liquid configurational space, and the test sets are generated by the same pipeline, so the validation can inherit the generator's blind spots. The paper's own Methods acknowledge MACE-MP0 is not accurate for molten salts, and the element-resolved error analysis shows Zr, the highest-valence cation, has the largest force errors—consistent with incomplete coverage of difficult chemical environments. The AIMD comparisons on nine compositions are encouraging but too sparse to certify the full 11-cation composition space, especially because the BO demonstration targets only density and does not independently validate the discovered compositions. These points reinforce the reader's CONDITIONAL verdict rather than overturning it. The paper has genuine strengths: the parity plots show low RMSE on the held-out Test1/Test2 sets, the property comparisons against AIMD and experiment are reasonable, and the active-learning dataset of ~70,000 DFT-labeled structures is substantial. The missing artifacts (code, data, weights) and the lack of error bars on MLIP-predicted properties further support the conditional stance, but the generator-bias issue is the most technically decisive because it directly challenges the scope of the central transferability claim. A single targeted experiment—building an AIMD-sampled test set for Zr/Zn-rich compositions—would provide a clear pass/fail signal. No ad hominem is intended; this is a standard and addressable MLIP validation gap.","tokens_in":15496,"tokens_out":2893,"duration_ms":31370,"concrete_test":"Generate an independent test set from unbiased AIMD, not from MACE-MP0 MD: choose ~20 random 11-cation compositions enriched in ZrCl4 and ZnCl2 (e.g., 10-30 mol% each), run 50 ps AIMD at 1200 K using the same VASP PBE-D3 settings as the paper, subsample ~500 decorrelated configurations, and compute SuperSalt force/energy RMSE against DFT labels. Compare to the reported Test2 RMSEs (24.4 meV/Å for forces, 1.3 meV/atom for energies). If the AIMD-sampled RMSE is within roughly 2x of the Test2 values and SuperSalt-MD densities/RDFs for these compositions match AIMD within the paper's stated tolerances (2% density, overlapping RDF peaks), the generator-bias concern is resolved. If the RMSE degrades substantially or the structural properties deviate, the transferability claim must be restricted to compositions and coordination environments represented in MACE-MP0-sampled trajectories.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that one MACE potential trained on unary, binary, and 11-cation chloride data transfers to arbitrary compositions across the 11-cation space with near-DFT accuracy. The load-bearing condition is that the raw configurations fed to DFT labeling adequately cover the liquid configurational space of all 11 cations, especially the high-valence Zr4+ and Zn2+ environments that already show the largest force errors (Fig. 2, third column). The Methods state that MACE-MP0 'may not be accurate enough for predicting thermophysical properties of molten systems,' yet all raw training configurations are generated by MACE-MP0-driven MD. Crucially, the Test1 and Test2 sets are generated the same way: the Methods say test configurations were produced by PACKMOL followed by MD 'similar to what was done for generating training data.' Thus the held-out sets share the generator's configurational biases. If MACE-MP0 explores only a subset of relevant liquid structures—for example, incorrect Zr-Cl or Zn-Cl coordination statistics—active learning can only select diverse samples from that subset, and DFT labels cannot recover the unsampled regions. The property validations against AIMD (density, RDFs, heat capacity, thermal expansion) cover only nine compositions and do not systematically probe Zr/Zn-rich corners of the composition space. The concern is therefore not that the reported RMSEs are wrong, but that they may not be representative of the claimed full compositional space. A positive control using configurations sampled from unbiased AIMD trajectories would settle whether the generator coverage assumption actually holds.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents SuperSalt, a MACE-based equivariant neural network interatomic potential for 11-cation chloride melts (Li, Na, K, Rb, Cs, Mg, Ca, Sr, Ba, Zn, Zr). Training configurations are generated by MACE-MP0-driven MD for unary, binary (at three compositions per pair), and 11-component systems, then selected via HDBSCAN active learning and labeled with PBE-D3 DFT; the final database contains roughly 70,000 structures. The authors report low energy/force RMSEs on training, validation, and two held-out sets (a 3300-structure ternary set and an 800-structure random multicomponent set), and they compare density, RDFs, heat capacity, and thermal expansion against AIMD and some experimental data. A Bayesian optimization workflow is demonstrated for targeting compositions with desired densities. The central claim is that a single potential trained on 1-, 2-, and 11-component data transfers to all intermediate compositions across the 11-cation space with near-DFT accuracy.","tokens_in":15795,"tokens_out":6372,"duration_ms":56429,"significance":"If the transferability claim holds, this would be a practically useful advance: a single MLIP covering a broad molten-salt family could replace many system-specific potentials and accelerate composition screening. The paper's strengths include a large and documented training set, a reproducible active-learning workflow, explicit property comparisons against AIMD, and a concrete application to Bayesian optimization. The reported energy and force errors are low, and the property comparisons show reasonable agreement. However, the central transferability claim rests on an assumption that the MACE-MP0-generated configurations adequately sample the liquid configurational space of all 11 cations, especially the high-valence Zr and Zn environments that already show the largest force errors. This assumption is not demonstrated, and because the held-out test sets are generated with the same MACE-MP0 workflow, the reported RMSEs may not be representative of the full compositional space. The paper is well within the scope of the journal and the community will likely find the dataset and workflow useful, but the strength of the claims currently exceeds the evidence.","major_comments":[{"comment":"The raw training and test configurations are all generated by MACE-MP0-driven MD, a model the paper itself states 'may not be accurate enough for predicting thermophysical properties of molten systems.' The Test1 and Test2 sets are generated the same way ('similar to what was done for generating training data'), so the held-out errors share the generator's configurational biases. If MACE-MP0 under-samples Zr-rich or Zn-rich liquid structures—the elements with the largest force errors in Figure 2 (third column)—active learning can only select from that biased distribution, and DFT labels cannot recover the unsampled regions. To support the transferability claim, please provide evidence that the MACE-MP0 sampling covers the relevant liquid configurational space, for example by comparing MACE-MP0-generated coordination statistics or RDFs with AIMD for ZrCl4- and ZnCl2-rich compositions, or by adding DFT-labeled configurations from independent sampling (AIMD, classical force fields, or other generators) to both training and test sets.","section":"Methods §MD workflows; Results §Training and testing results"},{"comment":"The property comparisons against AIMD use the same PBE-D3 functional and similar ~100-atom cells as the training labels, so they demonstrate consistency with the training electronic-structure method rather than independent physical accuracy. The experimental comparison in Figure 4b covers only six systems and shows an average density deviation near 5%. More importantly, the AIMD validation set includes only one Zr-containing composition (the 11-component Li2Na4Mg3K5Ca9ZnRb3Sr2Zr6Cs2Ba2Cl74) and one Zn-containing ternary (0.58NaCl-0.12CaCl2-0.3ZnCl2), both at low concentrations of the high-valence cations. The claim of transferability to untested compositions, especially 4- to 10-component and Zr/Zn-rich systems, is therefore not fully supported. Please validate SuperSalt on a broader set of compositions that systematically vary the Zr and Zn fractions, or temper the transferability claim to the compositions actually tested.","section":"Results §Density, Heat capacity, and Thermal expansion; Figure 4"},{"comment":"The training set contains only 1-, 2-, and 11-component systems, and the ternary test set uses a single composition A0.33B0.33C0.34 per ternary system. The statement that a potential trained on 11 elements 'inherently describes all 2048 suballoys' and 'exhibits excellent transferability' across all intermediate compositions is an extrapolation from 1-, 2-, and 11-component data. The paper does not report how many distinct compositions appear in Test2, nor how the 800 random configurations are distributed across 4-, 5-, ..., 10-component systems. Please provide error statistics broken down by number of components and by composition region (e.g., Zr-rich, Zn-rich, alkali-rich) for both Test1 and Test2, so that the transferability claim is quantified rather than asserted.","section":"Results §Comprehensive workflow; Results §Training and testing results"}],"minor_comments":[{"comment":"There is a typo 'Tes 2' in the sentence describing the Test2 dataset; it should read 'Test 2.'","section":"Results §Training and testing results"},{"comment":"The caption contains 'T arget one' and 'T arget two'; these should be 'Target one' and 'Target two.'","section":"Figure 5 caption"},{"comment":"The names 'V ASP 6.4.2' and 'PA W-PBE' have unintended spaces; they should read 'VASP 6.4.2' and 'PAW-PBE.'","section":"Methods §DFT computations"},{"comment":"The statement 'the entire melt-quench region was mapped to the SuperSalt model with only 2% of configuration space (initial structures)' is unclear and appears inconsistent with the later statement that ~70,000 configurations were selected from more than 20,000,000 raw configurations (about 0.35%). Please clarify what the 2% refers to.","section":"Results §Comprehensive workflow"},{"comment":"The comparison of SuperSalt to MACE-MP0 in Figure S1 should specify exactly how the DFT-D3 correction was applied to MACE-MP0 predictions, since MACE-MP0 was trained on PBE data without D3. Without this detail, readers cannot assess whether the comparison is a fair test of the universal potential's accuracy for molten salts.","section":"Methods §MD workflows; Figure S1"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely problem and the dataset plus workflow are likely to be valuable to the molten-salt community. My main reservation is that the central transferability claim currently rests on an unvalidated assumption about the coverage of the MACE-MP0 generator, and the test and validation sets share the generator's biases. This is fixable in a revision by adding independent DFT-labeled configurations or systematic composition-stratified error reporting, and by appropriately softening the claims where the evidence is incomplete. I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Candidly: this is one of the more useful MLIP papers I've read this year. It delivers a single MACE potential for 11 chloride cations, trained on unary/binary/11-component melts with an HDBSCAN active-learning scheme, and shows held-out errors on ternary and random multicomponent sets that are an order of magnitude better than MACE-MP0. The property validation (density, Cp, thermal expansion, RDFs) against AIMD over nine compositions is real evidence, and the Bayesian optimization demo is a nice practical payoff. I'd call the central claim -- near-DFT accuracy across the 11-cation space -- largely supported for the compositions tested.\n\nThe soft spots are real but not disqualifying. The biggest one is exactly what the stress-test note flags: the raw configurations for both training and the held-out Test1/Test2 sets come from MACE-MP0-driven MD. The paper admits MACE-MP0 is not accurate for molten salt properties, and if it under-samples Zr/Zn-rich liquid structures, active learning can only pick from that subset. The AIMD comparisons cover nine compositions, including one 11-cation melt, but they don't systematically probe the high-valence corners where force errors are largest. I'd like a positive control: a handful of AIMD-generated liquid configurations outside the MACE-MP0 distribution, used as an additional test set. That would settle the generator-coverage question.\n\nAlso, no code, data, or model weights are released. For a field that increasingly expects artifacts, that's a reproducibility gap. It's not fatal, but it limits how much independent verification the claims get.\n\nThe BO section is fine as a proof of concept, though the max-density result (pure BaCl2) is somewhat obvious from atomic weights. The second BO target is a better test of precision.\n\nOverall: this deserves a serious referee. I'd send it out with a request for the artifacts and the extra AIMD-configuration test. The paper is honest, the numbers are good, and the 1/2/11 training insight is worth disseminating even if the universal-coverage phrasing is a bit strong.","headline":"Strong, useful MLIP paper; transferability claim is solid for tested compositions but rests on a shared MACE-MP0 generator that deserves an independent AIMD-configuration check.","tokens_in":16399,"tokens_out":2392,"would_cite":true,"duration_ms":22909,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SuperSalt: one neural-network potential covers all 11-cation chloride melts at near-DFT accuracy.","keywords":["molten salts","machine learning interatomic potential","equivariant neural network","MACE","transferability","active learning","Bayesian optimization","chloride melts"],"falsifier":"Take a ZrCl4-rich multicomponent melt (or add a 12th cation not in the training set), generate configurations from long molecular dynamics runs, and compare SuperSalt forces and predicted densities and heat capacities against fresh AIMD (PBE-D3) calculations; if force RMSEs on Zr environments exceed the roughly 25 meV/Å range seen for random multicomponent tests, the claimed transferability is limited. A simpler check is to compute SuperSalt's force error on configurations from a melt-quench trajectory of pure ZrCl4 at 1200 K, since Zr4+ already shows the largest per-element force errors in the paper's parity plots.","tokens_in":15255,"feed_emoji":"🧂","tokens_out":5508,"duration_ms":43577,"temperature":0.7,"pith_summary":"The paper sets out to show that a single machine-learned interatomic potential can cover the entire liquid-phase composition space of an 11-cation chloride salt family, not just one or two specific melts. By training a MACE equivariant neural network on configurations drawn only from unary, binary, and 11-component systems, the authors claim near-DFT accuracy for density, bulk modulus, radial distribution functions, heat capacity, and thermal expansion, with force errors roughly an order of magnitude below the universal MACE-MP0 potential on the same tests. If correct, this would replace the current practice of fitting a new potential for every salt composition with one reusable model, and it would make large-scale screening of salt compositions practical. The paper also demonstrates that coupling this potential with Bayesian optimization can locate compositions with target densities after only a handful of molecular dynamics runs.","feed_headline":"Single model nails molten salt properties across 11 cations","feed_subtitle":"SuperSalt reaches near-DFT accuracy on densities, heat capacity, and structure for all chloride salt compositions.","key_machinery":"The central object is the MACE (multilayer atomic cluster expansion) architecture, an equivariant message-passing neural network that builds many-body atomic descriptors from tensor products of spherical basis functions; with two layers, lmax=3, and 4-body messages it maps local atomic environments to energies and forces. Around it sits a data-generation workflow: MACE-MP0-driven melt-quench MD generates raw configurations for 1-, 2-, and 11-component salts; HDBSCAN active learning selects a few hundred diverse structures per subsystem; and PBE-D3 DFT labels roughly 70,000 structures (about 7 million atoms). The paper's efficiency claim rests on the assumption that salt physics is dominated by pairwise electrostatics, so 1- and 2-component data capture the essential interactions and a small amount of 11-component data teaches the potential to handle many-element environments.","core_discovery":"SuperSalt, a MACE potential fitted to DFT (PBE-D3) data for 11 chloride salts (LiCl, NaCl, KCl, RbCl, CsCl, MgCl2, CaCl2, SrCl2, BaCl2, ZnCl2, ZrCl4), achieves near-DFT accuracy for energies and forces across the $2^{11}$ composition space, including ternary and random multicomponent mixtures never seen in training. Energy RMSEs are 0.5 meV/atom on training and validation sets, 0.6 meV/atom on 3300 ternary configurations, and 1.3 meV/atom on 800 random multicomponent configurations; force RMSEs range from 13.7 to 24.4 meV/Å. Predicted densities deviate from AIMD by less than 2% and from experiment by about 5%, heat capacities by up to 4.8%, and thermal expansion coefficients by about 6.7%. Relative to the general-purpose MACE-MP0 foundation model, SuperSalt cuts energy errors by roughly 40–90 times and force errors by roughly 7–10 times on molten salt tests.","pith_inferences":["The same unary-plus-binary-plus-high-order training strategy may transfer to other chemically similar liquid families, such as fluoride or bromide melts, if pairwise electrostatics dominate there as well.","A sharper test of transferability would be to hold out entire binary subsystems from training and check whether SuperSalt still predicts them; the reported Test1 and Test2 sets do not fully isolate this.","The 5% density deviation from experiment likely reflects systematic DFT-D3 error rather than potential error, so coupling SuperSalt with empirical density corrections could improve absolute property predictions.","The HDBSCAN active-learning pipeline could be reused to build compositional foundation potentials for other liquid electrolytes or oxide melts where enumerating all compositions is infeasible."],"forward_implications":["A single SuperSalt potential replaces dozens of system-specific MLIP fittings for chloride melts, since it reliably predicts all 165 ternary and random multicomponent compositions from training on only 1-, 2-, and 11-component systems.","Compositions never present in training, including all ternary mixtures, are predicted at near-DFT accuracy with force RMSEs below about 25 meV/Å.","Bayesian optimization on top of SuperSalt-MD finds target-density compositions in as few as six iterations, which the paper argues is impractical with empirical or ab initio methods at this scale.","Extending the approach to a 12th element requires only adding one unary, its 12 binary systems, and active-learned 12-component configurations, per the authors' stated plan.","The potential enables nanosecond-scale molecular dynamics of multicomponent melts with near-DFT force accuracy, opening the way to screening properties like viscosity that are too expensive for direct AIMD."],"supporting_citations":[{"why":"Supplies the MACE-MP0 universal potential used for raw configuration generation and as the baseline comparator that SuperSalt outperforms by an order of magnitude.","marker":"[15]"},{"why":"The Dirichlet distribution method that controls and generates the 11-component compositions in the training set.","marker":"[26]"},{"why":"The MACE model chosen as the neural-network architecture for SuperSalt because of its higher-order many-body accuracy.","marker":"[27]"},{"why":"HDBSCAN hierarchical density-based clustering, the method used in active learning to partition the trajectory data into uncorrelated clusters.","marker":"[31]"},{"why":"A prior neural-network interatomic potential for NaCl melts used as a reference for training RMSEs and RDF comparisons.","marker":"[35]"},{"why":"The MACE architecture reference that defines the equivariant message-passing framework and training methodology used in the Methods.","marker":"[42]"},{"why":"The DFT-D3 dispersion correction applied in all DFT reference calculations, which is essential for matching experimental melt densities.","marker":"[51]"},{"why":"Prior work showing that unary and binary data at intermediate compositions suffice for compositionally transferable potentials, motivating the training-set design.","marker":"[53]"}],"fun_headline_variants":["SuperSalt: near-DFT accuracy for 11-cation molten salts","One MLIP predicts all 11 chloride salt properties near DFT","SuperSalt reaches near-DFT accuracy across 11 cation melts","Molten salt MLIP hits near-DFT precision for 11 cations","SuperSalt: universal molten salt force field with DFT accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The raw configurations fed to the DFT labels are generated by the MACE-MP0 universal potential, which the paper itself says may be inaccurate for molten salts; if that generator never visits some chemically important liquid structure, especially around high-valence Zr4+ and Zn2+ ions, the training labels cannot teach the potential about it.","fun_headline_variants_meta":{"raw":{"variants":["SuperSalt: near-DFT accuracy for 11-cation molten salts","One MLIP predicts all 11 chloride salt properties near DFT","SuperSalt reaches near-DFT accuracy across 11 cation melts","Molten salt MLIP hits near-DFT precision for 11 cations","SuperSalt: universal molten salt force field with DFT accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000455,"raw_usage":{"total_tokens":2289,"prompt_tokens":951,"completion_tokens":1338,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":567,"completion_tokens_details":{"reasoning_tokens":1246}},"tokens_in":567,"tokens_out":1338,"duration_ms":9887,"temperature":1.0,"reasoning_tokens":1246,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T00:40:43.540626+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a ZrCl4-rich multicomponent melt (or add a 12th cation not in the training set), generate configurations from long molecular dynamics runs, and compare SuperSalt forces and predicted densities and heat capacities against fresh AIMD (PBE-D3) calculations; if force RMSEs on Zr environments exceed the roughly 25 meV/Å range seen for random multicomponent tests, the claimed transferability is limited. A simpler check is to compute SuperSalt's force error on configurations from a melt-quench trajectory of pure ZrCl4 at 1200 K, since Zr4+ already shows the largest per-element force errors in the paper's parity plots.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Dirichlet distribution method that controls and generates the 11-component compositions in the training set."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The MACE model chosen as the neural-network architecture for SuperSalt because of its higher-order many-body accuracy."},{"cited_title":"In: Pacific-Asia Conference on Knowledge Discovery and Data Mining, pp","cited_arxiv_id":null,"evidence_quote":"HDBSCAN hierarchical density-based clustering, the method used in active learning to partition the trajectory data into uncorrelated clusters."},{"cited_title":"Cell Rep","cited_arxiv_id":null,"evidence_quote":"A prior neural-network interatomic potential for NaCl melts used as a reference for training RMSEs and RDF comparisons."},{"cited_title":"Advances in Neural Information Processing Systems 35, 11423–11436 (2022)","cited_arxiv_id":null,"evidence_quote":"The MACE architecture reference that defines the equivariant message-passing framework and training methodology used in the Methods."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The DFT-D3 dispersion correction applied in all DFT reference calculations, which is essential for matching experimental melt densities."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior work showing that unary and binary data at intermediate compositions suffice for compositionally transferable potentials, motivating the training-set design."}],"review_version":1}