{"id":"c46ea2e1-3b78-4e5f-b3a5-3b005d3f8228","arxiv_id":"2506.12592","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Motif-based sampling of training configurations produces machine learning potentials that accurately predict alloy properties across compositions, as validated against experiments for phase diagrams, melting, short-range order, thermal expansion, and heat capacity.","lead":"Researchers introduce motif-based sampling (MBS), a way to build machine learning potential training sets that more uniformly cover local chemical environments in alloys. The method produces alloy models that match experimental phase diagrams, melting points, short-range order, and heat capacities across compositions, from simple binaries to high-entropy alloys.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MBS-vs-RS comparison does not isolate motif diversity from correlated changes in energy distribution and higher-order correlations; the causal claim needs a matched control.","rationale":"The paper's strongest and most distinctive contribution is a new training-set construction method with a specific mechanistic claim: uniformly sampling local chemical motifs is what makes MLPs accurate across alloy compositions. The direct evidence for this mechanism is the MBS-versus-RS comparison in Fig. 3b, which is exactly the load-bearing point the reader identified. The empirical validations in Figs. 5-9 are extensive and support the practical value of the resulting potentials, but they compare MBS-trained models against experimental data rather than against RS-trained models, so they do not test the mechanism. The causal claim could fail if MBS changes other properties of the training distribution. MBS is not a minimal intervention on the motif histogram: atomic swaps necessarily affect pair correlations, energy ordering, and possibly local strain. The paper reports comparable total SRO and thermal noise, but does not report matching the energy distribution or higher-order correlations. The high-SRO test sets are generated by reverse Monte Carlo, and MBS's swap-based motif flattening may inadvertently produce configurations more similar to those test sets than RS does, making MBS look better for reasons unrelated to motif diversity. A matched control experiment using RS configurations selected to reproduce MBS's energy and structural statistics would settle this. If the control matches MBS, the central mechanistic claim is weakened, although the practical method may still be valuable; if MBS still wins, the claim is strengthened. The reader's conditional verdict remains appropriate: the method is promising and well validated empirically, but the central mechanism is not yet fully isolated and the lack of released code and data further supports a conditional rather than unconditional accept.","tokens_in":16552,"tokens_out":4842,"duration_ms":66299,"concrete_test":"Construct a matched control from the RS configuration pool: select a subset whose DFT energy histogram, radial and angular distribution functions, and Warren-Cowley parameters through at least the third shell match the MBS training set, while keeping motif packing density at the RS level (select configurations rather than performing atomic swaps). Retrain the same 10-MLP ensemble on this control and evaluate on the five SRO test sets used in Fig. 3b. If the control matches MBS accuracy, motif diversity is not the operative variable; if MBS remains more accurate, the causal claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on Fig. 3b and the Methods section 'Sampling of chemical motifs': MBS training sets outperform RS training sets because they increase motif packing density. But MBS is implemented via intracell atomic swaps that flatten the motif distribution, and this operation changes more than the motif histogram. The paper states that composition, thermal noise, lattice parameters, and total SRO (eq. 1) are matched, but it does not report matching the DFT energy distribution, the distribution of interatomic distances or angular environments, or higher-order correlations beyond the first shell. If MBS systematically shifts training energies toward the energy range of the reverse-Monte-Carlo test sets, or biases pair and angular statistics toward the test distribution, the apparent error reduction would reflect dataset alignment rather than motif diversity. The widening MBS advantage at high SRO in Fig. 3b is exactly the pattern expected if the swap procedure generates configurations that resemble the high-SRO test set. Because the methodology is framed as 'motif-based,' this causal attribution is load-bearing; the extensive experimental benchmarks in Figs. 5-9 validate the resulting potentials but do not distinguish this mechanism from correlated confounds.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes motif-based sampling (MBS) to build machine-learning-potential training sets for metallic alloys. Using a local chemical motif decomposition from the authors' earlier work, MBS promotes a more uniform distribution of chemical motifs in the training data via intra-cell atomic swaps. The authors compare MBS with random sampling on CrCoNi and report improved energy accuracy on reverse-Monte-Carlo test sets with controlled short-range order. They then train PACE potentials for CrCoNi, AuPt, CuAu, and TaTiVW, and validate the models against experiments for phase diagrams, melting temperatures, Warren–Cowley short-range-order parameters, thermal expansion, heat capacity, and stacking-fault energies. The central claim is that motif diversity in the training set is a critical determinant of MLP accuracy for solid solutions across composition space.","tokens_in":16805,"tokens_out":4834,"duration_ms":61970,"significance":"The paper has clear strengths: it targets an important practical problem (accurate MLPs for disordered alloys across compositions), it evaluates the method on external held-out DFT data and on extensive experimental comparisons, and it proposes a low-cost modification to existing training pipelines. The experimental validations in Figs. 5–9 are broad and, if the mechanism is sound, demonstrate a practically useful method. The main risk is that the causal attribution to motif diversity rests on one controlled comparison (Fig. 3b), and that comparison does not currently exclude correlated confounds. The Cu3Au short-range-order validation also relies on an underspecified standardization step. These issues are fixable, but they are load-bearing for the paper's central claim.","major_comments":[{"comment":"The statement that 'any differences in model performance can be attributed specifically to motif representation' is too strong. MBS is implemented by intra-cell atomic swaps that flatten the motif histogram, and those swaps also change the distribution of DFT energies, pair and angular statistics, and higher-order correlations in the training set. Since the test sets are reverse-Monte-Carlo configurations with controlled SRO, a swap-induced resemblance to high-SRO test configurations could explain the widening MBS advantage in Fig. 3b even if motif diversity were not the operative variable. The paper should report matched diagnostics (energy histograms, pair-correlation functions beyond the first shell, and angular environment distributions) for the MBS and RS training sets, and should add a control experiment in which motif diversity is varied while the energy distribution is held fixed or otherwise decorrelated from the test SRO. Without such a control, the central causal claim that uniform motif sampling drives the improved generalization is not established.","section":"Motif-based sampling for solid solutions; Fig. 3b"},{"comment":"The Cu3Au comparison uses 'standardized' Warren–Cowley parameters, but the standardization procedure is not described and appears to be a post-hoc normalization that removes a systematic bias attributed to the DFT functional. As presented, the reader cannot tell whether the agreement in Fig. 7b reflects a fitted scaling parameter or a parameter-free prediction. The paper should report the raw α values, the exact normalization formula, any fitted parameter values and their uncertainties, and the sensitivity of the conclusion to the normalization. This is important because the quantitative SRO claim for CuAu rests on this panel; the CrNi agreement in Fig. 7a is quantitative and does not rely on such standardization.","section":"Chemical short-range order; Fig. 7b"},{"comment":"The key methodological details on which the paper's claims rest are referenced only as Methods sections: 'Sampling of chemical motifs', 'Training and testing datasets', and 'Warren-Cowley parameters'. The manuscript should include the full MBS acceptance criterion, the exact definitions of motif packing density and Jensen–Shannon divergence used, and the complete SRO standardization procedure, so that the central comparison in Fig. 3b and the Cu3Au validation can be reproduced and independently assessed.","section":"Throughout; methods completeness"}],"minor_comments":[{"comment":"The text refers to 'fig. 3d', but the figure caption lists only panels (a)–(c); either add the panel or correct the reference.","section":"Fig. 3"},{"comment":"The reported MatterSim error of 'up to 4,500 meV/atom' and '10,861% variation across compositions' should be checked; the numbers are surprisingly large and the baseline for the percentage variation is not defined.","section":"Fig. 2a"},{"comment":"The 10-model ensemble results would be more informative with error bars or shaded intervals; as plotted, the reader cannot assess the ensemble spread behind the MBS and RS curves.","section":"Fig. 3b"},{"comment":"The maximum deviations of 0.1% and 0.3% for lattice parameters should be accompanied by the temperature range over which they are evaluated and by the experimental uncertainty, so that the reader can interpret the agreement.","section":"Fig. 8a"},{"comment":"The paper would be strengthened by a statement on data and code availability, including whether the trained PACE potentials, the MBS and RS training sets, and the test sets are released.","section":"Code and data availability"}],"recommendation":"major_revision","confidential_remarks":"The paper builds directly on the authors' prior motif framework (refs 8 and 14), and the present novelty is the sampling strategy plus the broad experimental validation. The main technical issue for the editor to weigh is the control in Fig. 3b: the causal claim requires either matched diagnostics showing that only the motif histogram differs, or a control experiment that separates motif diversity from energy-distribution alignment. The experimental benchmarks are strong, and I would view a suitably qualified version of the central claim as publishable after the authors address this point."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"MBS is a genuinely useful training-set construction method, and the paper backs it with a lot of careful validation. The information-theoretic sampling objective is new, and the authors demonstrate across five alloy systems that it produces potentials with lower energy errors on random solid solutions than random sampling or SQS, and then go on to reproduce experimental phase diagrams, melting temperatures, short-range order, thermal expansion, heat capacities, and stacking-fault energies. The phase diagram and melting point results in particular are impressive, and the comparison with universal potentials (MatterSim, Orb, MACE) is fair and informative. If you work on MLPs for alloys, this is worth a read and probably a cite.\n\nThe soft spot is the causal claim. The paper argues that because composition, thermal noise, lattice parameters, and total SRO are matched, the performance difference between MBS and random sampling must be due to motif representation. That inference is too strong. The swap operation that flattens the motif histogram can also shift the energy distribution, alter higher-order correlations beyond the first shell, and change the distribution of interatomic distances and angles. The widening MBS advantage at high SRO in Fig. 3b is exactly what you would expect if the swap procedure makes training configurations more similar to the high-SRO test set. So the mechanistic attribution—motif diversity drives the gain—is not cleanly established. This is fixable: run a control where you match the energy distribution and the relevant correlation functions between MBS and RS datasets, or train on RS configurations selected to have the same motif distribution. Without that, the practical message survives but the conceptual one is weaker.\n\nMinor issues: the Cu3Au SRO comparison uses a post-hoc standardization of the Warren-Cowley parameters; the authors disclose it and cite known DFT biases, but it is not a clean validation. There is also no released code or data, which makes reproducibility harder. The DFT reference quality is acknowledged in the discussion, so that is fair.\n\nOverall, this paper deserves a serious referee. I'd send it to review, but I'd ask the authors to address the matched-control issue and either release data or provide a very detailed dataset description before I would call the causal mechanism established.","headline":"A genuinely useful training-set construction method with strong experimental benchmarking, but the central comparison doesn't cleanly isolate motif diversity from correlated structural changes.","tokens_in":17269,"tokens_out":2483,"would_cite":true,"duration_ms":26228,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A machine-learning potential trained on a uniform spread of local chemical motifs can model an alloy across its full composition range without retraining.","keywords":["machine learning potentials","solid solutions","high-entropy alloys","chemical short-range order","motif-based sampling","phase diagrams","alloy thermodynamics","compositional transferability"],"falsifier":"Train MBS and random-sampling datasets that are matched in composition, thermal noise, short-range order, and interatomic-distance distribution, then compare energy errors on a held-out solid-solution set. If the MBS advantage vanishes once distance statistics are controlled, motif diversity is not the active variable; if it survives, the mechanism is confirmed.","tokens_in":16365,"feed_emoji":"⚛️","tokens_out":7368,"duration_ms":82217,"temperature":0.7,"pith_summary":"The paper claims that the accuracy of machine-learning potentials—models that predict atomic energies from local structure—is set less by model architecture or dataset size than by how well the training set represents local chemical environments. It introduces motif-based sampling, which reshuffles atoms within training configurations so that local coordination motifs, the polyhedral arrangements of neighbors around each atom, appear with roughly uniform frequency. Potentials trained this way keep energy errors low across the full composition range of solid solutions, and across four alloy families—CrCoNi, AuPt, CuAu, and TiTaVW—they reproduce experimental phase diagrams, melting temperatures, short-range order, thermal expansion, heat capacity, and stacking-fault energies without retraining. If true, this makes it practical to build one potential that works across an entire alloy phase space instead of retraining for each composition.","feed_headline":"Train on motif diversity, model every alloy composition","feed_subtitle":"A balanced spread of local chemical motifs lets one MLP match experiments across four alloy systems.","key_machinery":"The load-bearing object is the local chemical motif—the coordination polyhedron around an atom that encodes its chemical neighborhood—and the motif-based sampling (MBS) procedure built on it. MBS starts from chemically random configurations and performs intracell atomic swaps to drive the distribution of motifs toward uniform, quantified by Jensen–Shannon divergence to a uniform target and by motif packing density, the percentage of distinct motifs sampled. Two auxiliary components complete the training set: phase sampling adds metastable lattices and ordered intermetallics, and thermal perturbations add synthetic vibrations and thermal expansion. The argument is that these three components together cover the chemical, structural, and vibrational subspace an alloy actually occupies.","core_discovery":"The central claim is that chemical motif diversity, not raw data volume, is the key determinant of whether an machine-learning potential can resolve the small energy differences that govern disordered alloys. On the CrCoNi solid solution, universal potentials show errors as large as thousands of meV/atom varying wildly with composition; a motif-based-sampling-trained model reduces and flattens those errors across the ternary triangle. The authors show the improvement is physically meaningful: the typical energy differences associated with chemical short-range order are about 10 meV/atom, the scale that motif-based sampling makes resolvable. Using one model per system, the paper reproduces experimental phase boundaries for CrNi, CrCo, and AuPt, melting temperatures within a few percent for CrCoNi and TaTiVW-derived alloys, Warren–Cowley short-range-order parameters, thermophysical properties, and composition-dependent stacking-fault energies for CrCoNi.","pith_inferences":["By extension, the same motif-diversity principle should apply to other disordered materials—glasses, liquids, irradiated or defective crystals—where the relevant degrees of freedom are local environments rather than lattice periodicity; the paper does not test these cases.","A causal test that goes beyond the paper: randomize the pairing between motif identity and atomic coordinates in training data. If energy errors track motif diversity rather than pair-distance statistics, the proposed mechanism is confirmed; if not, the improvement has a different source.","The paper's phase-diagram workflow still assumes knowledge of which competing phases to include; combining MBS with generative structure search could close that gap, a direction the authors note but do not pursue.","The fixed-dataset-size gains suggest MBS could be folded into active-learning loops to cut first-principles data costs further, but that combination is not demonstrated here."],"forward_implications":["A single MBS-trained potential can replace per-composition retraining, so phase diagrams, melting curves, and ordering tendencies across a composition space become computable from one model.","The accuracy gain is concentrated where it matters thermodynamically: near the 10 meV/atom energy scale that separates competing ordered and disordered states.","Because MBS works by dataset construction rather than architecture change, it can be added to existing machine-learning-potential training pipelines at negligible cost.","The resulting models resolve the energetic biases behind chemical short-range order, making Warren–Cowley parameters and stacking-fault energies accessible from simulation.","MBS training sets can serve as a benchmark for testing whether universal potentials actually handle disordered phases."],"supporting_citations":[{"why":"Defines local chemical motifs and their frequencies and correlations in metallic alloys; MBS builds directly on this decomposition.","marker":"[8]"},{"why":"Shows how to characterize chemical motifs with equivariant graph neural networks, supplying the machinery for quantifying motif diversity.","marker":"[14]"},{"why":"The special quasi-random structure method that serves as the standard baseline; the paper compares MBS against it.","marker":"[24]"},{"why":"The atomic cluster expansion model architecture used to train the evaluated potentials.","marker":"[26]"},{"why":"Multi-cell Monte Carlo method used to compute solid-state phase boundaries.","marker":"[54]"},{"why":"Solid–liquid coexistence method used to estimate melting temperatures.","marker":"[55]"},{"why":"Experimental melting temperature of CrCoNi used as the benchmark for the coexistence simulation.","marker":"[57]"},{"why":"Experimental solidus and liquidus temperatures for refractory high-entropy alloys used to validate melting predictions.","marker":"[58]"},{"why":"Experimental Warren–Cowley parameters for Cu3Au used to validate short-range-order predictions.","marker":"[62]"}],"fun_headline_variants":["Motif diversity, not data volume, makes alloy MLPs accurate","Information-theoretic sampling trains alloy MLPs across compositions","One MLP per alloy system, trained on chemical motif diversity","Sampling chemical motifs unlocks MLPs for whole alloy landscapes","Key to alloy MLPs: balanced sampling of local chemical motifs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim would collapse if the accuracy gain attributed to greater chemical-motif diversity actually comes from a correlated change—such as an altered distribution of interatomic distances or a bias toward the specific test configurations—rather than from motif diversity itself.","fun_headline_variants_meta":{"raw":{"variants":["Motif diversity, not data volume, makes alloy MLPs accurate","Information-theoretic sampling trains alloy MLPs across compositions","One MLP per alloy system, trained on chemical motif diversity","Sampling chemical motifs unlocks MLPs for whole alloy landscapes","Key to alloy MLPs: balanced sampling of local chemical motifs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000725,"raw_usage":{"total_tokens":3221,"prompt_tokens":885,"completion_tokens":2336,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":501,"completion_tokens_details":{"reasoning_tokens":2252}},"tokens_in":501,"tokens_out":2336,"duration_ms":19084,"temperature":1.0,"reasoning_tokens":2252,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:45:03.305766+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train MBS and random-sampling datasets that are matched in composition, thermal noise, short-range order, and interatomic-distance distribution, then compare energy errors on a held-out solid-solution set. If the MBS advantage vanishes once distance statistics are controlled, motif diversity is not the active variable; if it survives, the mechanism is confirmed.","supporting_citations":[{"cited_title":"Quantifying chemical short-range order in metallic alloys","cited_arxiv_id":null,"evidence_quote":"Defines local chemical motifs and their frequencies and correlations in metallic alloys; MBS builds directly on this decomposition."},{"cited_title":"Multi-cell Monte Carlo method for phase prediction","cited_arxiv_id":null,"evidence_quote":"Multi-cell Monte Carlo method used to compute solid-state phase boundaries."},{"cited_title":"Melting line of aluminum from simulations of coexist- ing phases","cited_arxiv_id":null,"evidence_quote":"Solid–liquid coexistence method used to estimate melting temperatures."},{"cited_title":"Optimization of conflicting properties via en- gineering compositional complexity in refractory high en- tropy alloys","cited_arxiv_id":null,"evidence_quote":"Experimental solidus and liquidus temperatures for refractory high-entropy alloys used to validate melting predictions."},{"cited_title":"Analysis of short-range order in Cu3Au using X-ray pair distribution functions","cited_arxiv_id":null,"evidence_quote":"Experimental Warren–Cowley parameters for Cu3Au used to validate short-range-order predictions."}],"review_version":1}