{"id":"94da4cb3-892b-4304-96bf-d45fe710576e","arxiv_id":"2411.13689","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Reusing FAIR data from a prior active learning campaign reduced the cost of a new alloy melting-temperature search by roughly ten times.","lead":"This paper shows that reusing stored simulation data from a public workflow database cut the number of simulations needed to find an optimal alloy from about sixty to about six. It is a concrete case study of how FAIR data sharing can speed up materials discovery.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 10x speedup claim assumes the reported minimum compositions were not already present in the 265-point FAIR training set; the paper never states whether they were, so 'discovery with three compositions' may actually be retrieval.","rationale":"The reader's weakest assumption was that the prior maximization campaign's data may be biased away from the low-melting Cu-rich region, causing extrapolation error. My concern is adjacent but more specific and more directly load-bearing: the reported minima may already be present in the reused training data. This would not merely bias the model; it would undermine the claim that the AL workflow discovered the minimum with only three new simulations. The paper's own text enables this concern: Section 2.3 states that all 265 compositions from Work 1 were used, and Section 3.2 reports the final minima without stating whether those compositions were previously characterized. Because the central claim is a speed comparison between a workflow that starts with 265 labeled points and one that starts with 40, the provenance of the specific target points is essential. If the minima are in the training data, the comparison is not between two active-learning searches but between a lookup/verification and a genuine search. This is fixable by a data query, so the verdict remains conditional rather than outright rejection: the authors can clarify whether the minima were previously stored and, if so, reframe the contribution as demonstrating FAIR-data reuse for known-optimum retrieval, or rerun the minimization on a target not in the training set to support the discovery claim. I partially agree with the reader because the distribution-bias concern is real, but the exact-composition-presence question is more decisive for the headline result.","tokens_in":7707,"tokens_out":8993,"duration_ms":83390,"concrete_test":"Query nanoHUB ResultsDB / the 265-composition training set for the three reported optimal compositions (Cr40Cu50Ni10, Cr40Cu50Co10, Cr50Cu50) and for near-duplicates within 5 at% of each. If any of these compositions are present, the paper must explicitly state that the minima were already known, and the 'testing only three compositions' discovery claim is unsupported; the 10x speedup then needs recomputation against a target genuinely absent from the training set. If they are absent, the concern is reduced but not eliminated, and a follow-up ablation excluding all compositions within 5 at% of each selected candidate would test whether the three-iteration success survives without near-neighbor memorization.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central comparison in Section 3.2 is 65 simulations (original workflow, minimization task) versus about 6 simulations (FAIR-data workflow, same task), leading to the 10x speedup claim. This comparison is only a fair test of active-learning acceleration if the minimization targets were not already labeled in the reused FAIR data. The new workflow starts from all 265 compositions collected during Work 1's maximization campaign (Section 2.3), and the final reported minima are Cr40Cu50Ni10 (1566 K), Cr40Cu50Co10 (1564 K), and Cr50Cu50 (1535 K). The paper says these are 'consistent with the low values predicted for Cu-rich systems in Work 1' (Section 3.2), but it never states whether the exact optimal compositions were previously simulated and stored in ResultsDB. If they were, then the AL loop did not discover them; it selected candidates whose melting temperatures were already in the training set, and 'testing only three compositions' is re-verification, not discovery. In that case the claimed speedup reduces to a data-lookup advantage over Work 1's 65 simulations, which had to label unknown compositions. Even if the exact compositions are absent, the paper does not report how close the selected candidates are to the training set, so the risk of near-neighbor memorization rather than generalization remains unresolved. This is the most load-bearing gap because the entire headline result—10x faster active-learning discovery—depends on the answer.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper demonstrates the reuse of FAIR simulation workflows and automatically stored data to accelerate a second active-learning optimization. Using 265 compositions previously characterized by the meltheas Sim2L in Farache et al. (Work 1), the authors train a random forest to predict alloy melting temperature and linear models for the solidus/liquidus input temperatures of the coexistence MD method. They report that the number of simulations needed to converge a melting temperature for a given composition drops from 4.4 to 1.3 on a 54-composition held-out test set, and that an active-learning minimization over the same 555-composition design space finds low-melting Cu-rich alloys in 3 compositions and about 6 simulations, compared with 16 compositions and 65 simulations when the original Work 1 workflow is applied to the same minimization task. The paper claims this constitutes a 10x speedup enabled by FAIR data reuse.","tokens_in":8029,"tokens_out":7058,"duration_ms":76283,"significance":"If the claims hold, this is a valuable concrete demonstration that FAIR data infrastructure can accelerate later materials optimizations, and the paper has real strengths: it uses a sequestered test set for the per-composition simulation reduction, it runs the original workflow as a control on the same minimization task, and it ships reproducible nanoHUB workflows with DOIs. However, the central quantitative claim currently rests on an unstated relationship between the reused training set and the final AL-selected minima, as well as on the representativeness of data collected in a maximization campaign for a minimization objective. These are answerable with additional analysis, and the conceptual contribution is likely sound, but the reported 10x speedup needs verification before the result can be fully credited.","major_comments":[{"comment":"The headline claim that the workflow finds the minimum-melting alloy by testing only three compositions is not fully supported because the manuscript never states whether the three final compositions (Cr40Cu50Ni10, Cr40Cu50Co10, Cr50Cu50) were already present in the 265-composition ResultsDB set used to initialize the random forest. If any of these exact compositions were previously simulated and stored, the AL loop is selecting an already-labeled candidate, and the reported 3-composition/6-simulation count is re-verification rather than discovery; the 10x speedup against the 65-simulation control would then conflate data lookup with active-learning acceleration. Please report the overlap of every acquisition-function-selected candidate with the training set, and for candidates not present, give their nearest-neighbor distances in composition space to the training data.","section":"§3.2, Fig. 4, and Data availability"},{"comment":"The 265 prior compositions were collected during Work 1's maximization campaign, and the paper does not characterize how that campaign sampled the low-melting-temperature, Cu-rich region in which the reported minima lie. If the training distribution is sparse or absent there, the random forest extrapolates rather than interpolates, and the apparent speedup could be dominated by model extrapolation error rather than by FAIR data reuse. Please provide coverage statistics for the training set (for example, the number of points with at least 40 at.% Cu and the range of melting temperatures sampled) and report the model's predictive uncertainty for the selected candidates.","section":"§2.3 and §3.2"},{"comment":"The quantitative claims are inconsistent across the paper: the Abstract says the workflow reduced the number of simulations per composition to one; §3.1 reports an average of 1.3 simulations per composition on the 54-composition test set; §3.2 reports approximately 2 simulations per alloy in the minimization run; and §4 concludes that the minimum was found with only two simulations. Since the headline speedup depends on these numbers, the paper should state a single, well-defined metric (simulations per composition and total simulations for the optimization) and correct the abstract and conclusions accordingly.","section":"Abstract, §3.1, §3.2, and §4"}],"minor_comments":[{"comment":"The parity plot should distinguish training and test points and report the test-set R² or MAE; as presented, the undifferentiated plot makes it difficult to judge generalization from the random forest.","section":"Fig. 2"},{"comment":"The text says 'most acquisition functions' succeeded but does not list which ones or how many iterations each required; a table with per-acquisition-function totals (compositions, simulations, and whether the minimum was found) would make the robustness claim checkable, especially because MLI failed within 40 compositions in the control run.","section":"§3.2 and Figs. S1–S3"},{"comment":"The rendering 'F AIR' with a space appears in the title and abstract; if this is intentional, it should be explained, otherwise it should be corrected to 'FAIR'.","section":"Title and text"},{"comment":"There are several typographical and formatting errors, including 'estimate theTsol' in §3.1, inconsistent capitalization of 'Sim2L'/'sim2l', and 'is the demonstrate' in §2.2; these should be corrected in a final pass.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the journal's scope and the FAIR-data-reuse concept is timely. The main decision hinges on whether the authors can show that the AL-selected minima were not already present in the reused training set and that the training data provide adequate coverage of the target region. If they can, the corrected quantitative claims and an explicit statement of the comparison metric would make the paper publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a useful case study, not a branch-level advance. The genuinely new pieces are the linear Tsol/Tliq estimators, the decision to start the AL loop from all 265 prior results in ResultsDB, and the minimization variant of the same acquisition functions. The held-out comparison is the right design: on 54 held-out compositions, the data-driven simulation-parameter estimates cut the average number of simulations per composition from 4.4 to 1.3. And the controlled minimization run using the original Work 1 workflow took 16 compositions and 65 simulations, versus 3 compositions and about 6 simulations with the FAIR-data workflow. That comparison on the same task makes the 10x speedup claim credible.\n\nThe paper also ships the workflow and data through nanoHUB, so the results are reproducible and the FAIR argument is concrete rather than rhetorical. Credit is due for that.\n\nSoft spots are mostly clarity. The abstract says \"one\" simulation per composition, but the held-out average is 1.3 and the minimization runs used about 2 per alloy; minor overstatement. The speedup bundles the effect of more training data and better simulation parameters, which is the point of FAIR reuse, but an ablation separating the two would strengthen the attribution.\n\nThe more load-bearing question, which the stress-test raises, is whether the reported minima—Cr40Cu50Ni10, Cr40Cu50Co10, Cr50Cu50—were already present in the 265-composition training set. The paper never says. If they were, the AL loop selected previously labeled compositions and the story changes from discovery to retrieval. It is plausible they weren't, because Work 1 was a maximization campaign and these are low-melting, Cu-rich boundary compositions. But the authors should confirm this explicitly and report how close the selected candidates sit to the training distribution. That is an easy fix.\n\nWho is this for? The materials-informatics and cyberinfrastructure community, especially people arguing for public data repositories. It is a single alloy family and one MD workflow, so impact is subfield-level, but the demonstration is clean. I would accept it for peer review, conditional on the authors addressing the training-set overlap and moderating the abstract's \"one simulation\" claim.","headline":"A solid FAIR-data reuse case study with a convincing controlled comparison; the main fix is to verify the discovered minima weren't already in the reused training set.","tokens_in":8559,"tokens_out":2820,"would_cite":true,"duration_ms":922206,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Reusing FAIR simulation data cuts an alloy search from about 65 simulations to about 6.","keywords":["active learning","FAIR data","molecular dynamics","melting temperature","multi-principal component alloys","random forest","workflow reuse","nanoHUB"],"falsifier":"Hold out every training composition with copper content at or above 50 at.%, retrain the random forest, and run the active-learning minimization for the lowest melting temperature; if the model then fails to find Cr50Cu50 within the original budget—or if its predicted melting temperature for that composition differs from the MD coexistence result by more than the observed ~10 K run-to-run spread—the speedup depended on the prior campaign having already sampled the answer region.","tokens_in":7506,"feed_emoji":"🔬","tokens_out":6459,"duration_ms":58990,"temperature":0.7,"pith_summary":"This paper claims that reusing FAIR-compliant simulation workflows and their automatically stored results can accelerate active learning in materials discovery. In a case study on multi-principal component alloys, the authors take the stored data from a previous campaign that maximized melting temperature and use it to train a random forest model, derive better starting temperatures for molecular dynamics simulations, and seed a new active learning search for the alloy with the lowest melting temperature. The result is a reported 10x speedup: about 6 simulations and 3 compositions to find the minimum, versus 65 simulations and 16 compositions with the original workflow. The paper uses this example to argue that FAIR data infrastructure turns each optimization into a reusable asset for future ones.","feed_headline":"FAIR data reuse cuts alloy search from 65 simulations to 6","feed_subtitle":"Reusing stored results, active learning finds the lowest-melting alloy in three compositions","key_machinery":"The load-bearing mechanism is the FAIR data reuse pipeline built on nanoHUB's Sim2L (a published, containerized simulation workflow with declared inputs and outputs) and ResultsDB (the database that automatically stores every workflow run's input-output pairs). This stored history serves two purposes: it provides training data for a random forest that predicts alloy melting temperature from composition, and it reveals that the melting-temperature simulation's starting temperatures Tsol and Tliq follow linear trends against the MD-determined melting temperature, so they can be set from the model prediction instead of by trial-and-error adjustment. These two pieces together let active learning start from an accurate model and finish each composition with roughly one simulation.","core_discovery":"The central discovery claim is that the automatically indexed input-output records of a nanoHUB Sim2L workflow constitute enough training data to both calibrate the physics simulation parameters and initialize a predictive model, so that a subsequent optimization runs in a fraction of the original cost. Specifically, the stored melting temperatures from Work 1's 265 compositions make it possible to express the coexistence-method input temperatures Tsol and Tliq as linear functions of a random-forest predicted melting temperature, with R2 of about 0.99 compared to 0.28 for the rule-of-mixtures estimate. That reduces the average number of simulations needed for a converged melting temperature from 4.4 to 1.3 on a held-out test set of 54 compositions. The same pre-trained model, with acquisition functions negated to minimize rather than maximize, identifies the lowest-melting alloy in the design space within about 3 compositions and about 6 simulations, a tenfold reduction relative to the original workflow applied to the same minimization task.","pith_inferences":["The reported 10x speedup conflates two improvements—a better initial model and better per-composition simulation parameters; the paper does not ablate the two, so the marginal contribution of each remains untested.","The stored dataset was collected during a maximization campaign, so its coverage of the low-melting, Cu-rich region is incidental; applying the same recipe to a design space that the prior campaign did not explore could produce a much smaller speedup or fail.","Because all acquisition functions selected compositions with exactly 50% Cu, the search effectively reduced to a one-dimensional problem; a broader design space with multiple competing optima would be a stiffer test of the claimed acceleration.","If the community standardizes on FAIR workflow tools like Sim2L, the first campaign in a new materials space remains expensive, but every subsequent campaign—even one targeting a different property—inherits its data; this could change how research groups value data publication."],"forward_implications":["The optimization cost for a materials property can drop by an order of magnitude when a prior, differently-directed optimization has already populated a queryable database.","Stored simulation data can be used to tune simulation parameters (starting temperatures) as well as to train the surrogate model, removing a labor-intensive manual step.","The same active-learning acquisition functions work for minimization once their target values are negated, so a model trained for one optimization direction transfers to the opposite direction.","If workflows and results are FAIR by default, each new optimization contributes to a growing dataset that makes later optimizations in the same or neighboring design spaces cheaper.","Averaging 1.3 simulations per composition (rather than 4.4) on the test set means the data-driven temperature estimator alone reduces total MD cost before the search even starts."],"supporting_citations":[{"why":"Work 1; supplies the original active-learning workflow, the 555-composition design space, the 265-characterization dataset, and the baseline of 15 compositions / ~4 simulations per composition that this paper improves upon.","marker":"[21]"},{"why":"The meltheas Sim2L on nanoHUB, whose runs generate the melting-temperature data that the paper reuses.","marker":"[23]"},{"why":"The solid-liquid coexistence method used to compute melting temperatures in the simulations.","marker":"[24]"},{"why":"The embedded-atom-method interatomic potential for Cr-Co-Cu-Fe-Ni alloys used in all MD runs.","marker":"[25]"},{"why":"Describes Sim2Ls as FAIR simulation workflows, the infrastructure that makes the data reusable.","marker":"[11]"},{"why":"Defines the FAIR guiding principles that the workflow and database follow.","marker":"[1]"}],"fun_headline_variants":["FAIR data reuse cuts alloy search from 65 to 6 simulations","Active learning with FAIR data finds alloys 10x faster","Reusing data speeds alloy discovery: 3 compositions, 6 simulations","Data-driven shortcut: lowest-melting alloy found in 3 tests","Accelerated alloy search via FAIR data: 10x speedup"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The random forest and the linear Tsol/Tliq estimators are trained on 265 compositions gathered while searching for the highest melting temperature; the paper assumes this stored dataset covers the low-melting-temperature region well enough to guide the minimization search, so the apparent 10x speedup could largely disappear if that prior data were biased toward high-melting compositions.","fun_headline_variants_meta":{"raw":{"variants":["FAIR data reuse cuts alloy search from 65 to 6 simulations","Active learning with FAIR data finds alloys 10x faster","Reusing data speeds alloy discovery: 3 compositions, 6 simulations","Data-driven shortcut: lowest-melting alloy found in 3 tests","Accelerated alloy search via FAIR data: 10x speedup"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000186,"raw_usage":{"total_tokens":1351,"prompt_tokens":996,"completion_tokens":355,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":612,"completion_tokens_details":{"reasoning_tokens":261}},"tokens_in":612,"tokens_out":355,"duration_ms":3469,"temperature":1.0,"reasoning_tokens":261,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:59:26.292543+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Hold out every training composition with copper content at or above 50 at.%, retrain the random forest, and run the active-learning minimization for the lowest melting temperature; if the model then fails to find Cr50Cu50 within the original budget—or if its predicted melting temperature for that composition differs from the MD coexistence result by more than the observed ~10 K run-to-run spread—the speedup depended on the prior campaign having already sampled the answer region.","supporting_citations":[{"cited_title":"High entropy alloy melting point calculation, Mar 2020","cited_arxiv_id":null,"evidence_quote":"The meltheas Sim2L on nanoHUB, whose runs generate the melting-temperature data that the paper reuses."},{"cited_title":"Melting line of aluminum from simulations of coexisting phases","cited_arxiv_id":null,"evidence_quote":"The solid-liquid coexistence method used to compute melting temperatures in the simulations."},{"cited_title":"Model interatomic potentials and lattice strain in a high- entropy alloy","cited_arxiv_id":null,"evidence_quote":"The embedded-atom-method interatomic potential for Cr-Co-Cu-Fe-Ni alloys used in all MD runs."},{"cited_title":"The fair guiding principles for scientific data management and stewardship","cited_arxiv_id":null,"evidence_quote":"Defines the FAIR guiding principles that the workflow and database follow."}],"review_version":1}