{"id":"9551ff6e-ac8e-473c-9da9-9302701981c0","arxiv_id":"2509.10293","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"OpenCSP is an open pressure-diverse dataset and model suite that matches or beats larger universal atomistic models on high-pressure crystal structure prediction with far fewer training data.","lead":"This paper presents OpenCSP, an open dataset of about 1.5 million pressure-labeled crystal configurations and three neural network potentials trained on it for high-pressure structure prediction. The models report higher stress accuracy and better recovery of high-pressure crystal structures than much larger universal models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-model benchmarks use different DFT labels/evaluation (ABACUS for OpenCSP vs native/VASP for baselines), so the reported virial and pressure advantages may be code artifacts rather than model quality.","rationale":"The reader's verdict is CONDITIONAL, and the weakest assumption points to the same issue. After reading the manuscript, I find the most load-bearing weakness is the inconsistent DFT reference between OpenCSP and the baselines. This is not a minor detail: virial outputs are known to vary across DFT codes and pseudopotential sets, and the paper itself acknowledges relabeling with ABACUS to match its training settings, while leaving baselines on their native labels. This asymmetry affects all three quantitative comparisons that support the central claim: MPTrj virial errors (III.B), GNoME formation energies (III.C), and pressure-controlled relaxation (III.D). In the CSP success-rate benchmark (III.E), the comparison is less affected by the code mismatch because success is defined by structural matching, not energy values; however, the distributional overlap with the training pipeline remains a secondary concern. The proposed test is feasible and decisive: if OpenCSP's advantage disappears under unified labeling, then the abstract's claim of superior high-pressure enthalpy ranking and stability prediction is unsupported; if it persists, the concern is resolved. I therefore recommend keeping the CONDITIONAL verdict, pending this re-evaluation.","tokens_in":19096,"tokens_out":5794,"duration_ms":66132,"concrete_test":"Recompute the comparisons in Table III (MPTrj-S/L) and Table IV (pressure relaxation) using a single DFT code for all models: either evaluate all structures with ABACUS or all with VASP (e.g., relabel MPTrj structures for all baselines with ABACUS, or use original VASP labels for OpenCSP). The key check is whether OpenCSP's virial MAE advantage and its pressure-reproduction advantage persist when all models are judged against identical reference data. If the advantage shrinks or reverses, the reported superiority is a label artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim of superior high-pressure performance depends on quantitative comparisons in which the reference labels are not code-consistent. In Sec. III.B (MPTrj cross-dataset), OpenCSP errors are computed against ABACUS relabels of the test structures, while MACE, MatterSim, and GRACE errors use the original MPTrj labels. In Sec. III.D (pressure relaxation), final structures are evaluated with ABACUS for OpenCSP and with VASP for the baselines, so the reported final pressures are not on the same footing. Virial/stress predictions are sensitive to DFT code, pseudopotential, and basis-set choices; the paper provides no quantification of ABACUS vs VASP virial differences for these test structures. The reported differences are large (virial MAE approximately 20-30 meV/atom for OpenCSP vs 110-170 meV/atom for baselines on MPTrj; pressure deviations at 50 GPa of 0-2 GPa for OpenCSP vs 40-58 GPa for MACE/GRACE). If the inter-code virial/pressure disagreement is of this magnitude, the benchmark advantage is an artifact. Because every headline comparison (virial accuracy, pressure reproduction, GNoME formation energy) carries this asymmetry, the evidence base for 'comparable or superior performance' is not yet secure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces OpenCSP, an open-source dataset of ~1.5 million DFT-labeled configurations generated by random CALYPSO structure searches under a pressure-aware DP-GEN concurrent-learning workflow, together with three DPA3-based models (OpenCSP-L6/L12/L24) trained for joint energy, force, and virial prediction. The central claim is that, despite using one to two orders of magnitude fewer training data than large universal models, OpenCSP achieves comparable or superior performance in high-pressure crystal structure prediction, particularly in virial accuracy, pressure-controlled relaxation, and CSP success rates at elevated pressures. Benchmarks are presented against MACE-MPA-0, MatterSim v1 5M, and GRACE-2L-OAM on in-distribution accuracy, MPTrj cross-dataset transfer, GNoME formation energies, pressure-constrained relaxations, and pressure-resolved CSP tasks.","tokens_in":19351,"tokens_out":4380,"duration_ms":52457,"significance":"If the central claim is established, OpenCSP would be a valuable community resource: it is an open, pressure-resolved dataset with explicit stress labels, and it demonstrates that targeted pressure-aware data generation can be more efficient than indiscriminate large-scale data collection. The paper's strengths include the public release of dataset and models, a realistic active-learning pipeline, held-out compositional splits for the in-distribution test, and honesty about the in-distribution versus zero-shot status of the baselines on MPTrj. However, the headline quantitative comparisons that support the high-pressure superiority claim are built on an evaluation asymmetry in DFT references, and the high-pressure CSP benchmark does not control for compositional overlap with the training pipeline. These issues must be addressed before the data-efficiency and superiority claims can be considered secure.","major_comments":[{"comment":"The abstract and conclusion claim 'comparable or superior performance in high-pressure enthalpy ranking,' but no direct enthalpy-ranking benchmark is presented. Table IV measures pressure reproduction and §III.E measures structural matching to literature structures; neither evaluates whether the models rank competing candidate structures correctly by enthalpy at a given pressure. A direct benchmark — e.g., generating multiple candidate structures per composition, ranking them by model enthalpy, and comparing the ranking/energy ordering to DFT enthalpies computed with a single code — is needed to substantiate the enthalpy-ranking claim. Without it, the superiority claim rests on an indirect proxy.","section":"General"}],"minor_comments":[{"comment":"The abbreviation 'sAlex' is used without definition; please introduce it (presumably a subset of the Alexandria dataset) at first use.","section":"§II"},{"comment":"Typo: 'labled' should be 'labeled' in the description of the initial dataset.","section":"§V.A"},{"comment":"The claim that MatterSim results are 'true zero-shot' would benefit from a citation or verification that MPTrj structures are not included in MatterSim's 17M training set, since the paper otherwise emphasizes this distinction.","section":"§III.B"},{"comment":"The phrase 'consistent, model-independent assessment' is misleading when different DFT codes are used for different models; please rephrase or, preferably, use a single common evaluator.","section":"§III.D / Table IV"},{"comment":"The data availability statement lists only generic AIS Square URLs. Please provide direct and permanent identifiers (e.g., DOI) for the specific dataset and model versions used in the paper.","section":"§X"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's main experimental design issue is the inconsistent DFT label reference across models, which is fixable but essential. If the authors' conclusions survive a common-label re-evaluation, the paper could be a strong contribution. Please also consider whether the emphasis on 'open source' is fully consistent with 'other scripts are available from the authors upon request' and with the reliance on a proprietary CALYPSO code, though this is secondary. The high-pressure reference set is drawn mostly from the same community's CALYPSO work; that is not a flaw per se, but it raises the bar for demonstrating that the advantage is not distributional overlap."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Alright, here's my read. The useful thing in this paper is the dataset: 1.5 million pressure-resolved DFT configurations, generated via CALYPSO sampling inside a DP-GEN active learning loop. That plus the three trained DPA3 models is a real contribution to the community. The pressure-aware sampling and the data-efficiency claim are interesting, and the reported virial improvements are large.\n\nWhat's done well: the methods are transparent (pseudopotentials, training schedules, selection criteria all given), the in-distribution and MPTrj zero-shot results look internally consistent, and the pressure-controlled relaxation table is impressive taken by itself. The data release is genuinely useful.\n\nThe soft spot, and it's the big one: the cross-model benchmarks are not on equal footing. In the MPTrj cross-dataset, OpenCSP errors are measured against ABACUS relabels while MACE, MatterSim, and GRACE are measured against the original MPTrj labels. Same story in the GNoME formation-energy test and in the pressure-relaxation table: ABACUS for OpenCSP, VASP for the baselines. The paper states this openly and the rationale is coherent—evaluate each model on labels from its own training distribution—but the comparison then conflates model quality with label source. Virial is sensitive to code, pseudopotential, and basis-set choices, and the gaps are so large that a code-consistent re-evaluation is essential before claiming the models are actually better.\n\nA couple of smaller things. The high-pressure CSP benchmark sets are small (8, 12, 11, 23 tasks at each pressure) and most reference structures come from CALYPSO-based papers, many from the same groups, so distributional overlap is a legitimate worry. The abstract promises 'enthalpy ranking' but I don't see a direct enthalpy-ranking benchmark; the CSP success rate is structure matching, not enthalpy ordering. And the benchmark scripts are only on request, which slows reproducibility even though the data and models are public.\n\nBottom line: this is a solid engineering contribution and the dataset deserves to be used. But the headline comparative claim is not yet established. Send it to a serious referee with instructions to require a common-label comparison (for instance, all evaluated on the same DFT code) and larger high-pressure test sets. I'd accept it for review, not desk reject.","headline":"Useful open dataset and models for high-pressure CSP, but the benchmark advantages over large models rest on a code-inconsistent evaluation that needs fixing before the comparative claims are credible.","tokens_in":19875,"tokens_out":7168,"would_cite":true,"duration_ms":66752,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"OpenCSP claims that targeted, pressure-aware data collection—not raw model scale—is what makes machine-learned interatomic potentials reliable for crystal structure prediction under tens to hundreds of gigapascals.","keywords":["crystal structure prediction","high-pressure materials","machine learning interatomic potentials","uncertainty-guided active learning","enthalpy ranking","virial accuracy","data-efficient training","concurrent learning"],"falsifier":"Evaluate OpenCSP and the baselines on a held-out set of high-pressure phases produced by an independent search algorithm, using identical DFT settings for all models; if OpenCSP's crystal-structure-prediction success-rate and virial advantages shrink to noise, the claim that targeted pressure sampling is the cause is not supported.","tokens_in":18940,"feed_emoji":"💎","tokens_out":6220,"duration_ms":67117,"temperature":0.7,"pith_summary":"OpenCSP argues that the reason large atomistic models struggle at high pressure is not model size but training data: their corpora cluster near ambient equilibrium and lack explicit pressure labels. To fix this, the authors built a 1.5-million-structure dataset by random high-pressure structure search with uncertainty-guided concurrent learning, then trained three deep graph-network potentials of modest scale. They show these models match or beat much larger universal potentials in energy, force, and especially virial accuracy, and in crystal structure prediction at 50–200 GPa, despite using one to two orders of magnitude less training data. If right, this means pressure-aware data selection can substitute for brute-force scaling in extreme-condition materials discovery.","feed_headline":"1.5M structures outperform far bigger models under high pressure","feed_subtitle":"Targeted pressure sampling, not data volume, decides if machine-learned potentials can rank compressed phases.","key_machinery":"The carrying mechanism is the dataset-construction loop, not a new architecture. Random structure proposals are relaxed at random target pressures drawn from 0–100 GPa; an ensemble of graph-network potentials estimates force or enthalpy uncertainty along the relaxation; only the most uncertain configurations go to DFT labeling; and the labeled set is folded back into training. The models are deep graph neural networks with 6–24 message-passing layers, jointly trained on energy, force, and virial, with virial treated as a first-class target because it controls the PV term in enthalpy.","core_discovery":"The paper claims that a curated, pressure-resolved dataset of about 1.5 million DFT-labeled configurations—built by proposing random compressed structures, relaxing them under randomly sampled target pressures, and relabeling only the most uncertain with DFT—is sufficient to train machine-learned potentials that outperform far larger universal models in high-pressure crystal structure prediction. The strongest evidence is in the virial (pressure–volume) term: on cross-dataset tests, OpenCSP models have several times smaller virial error and a much lighter error tail, which directly benefits enthalpy ranking. In pressure-controlled relaxation, the OpenCSP models reproduce target pressures of","pith_inferences":["This suggests that for extreme-condition modeling, targeted uncertainty-guided acquisition may beat brute-force scaling; a direct test would compare cost-per-accurate-enthalpy against large generic datasets on a fixed budget.","The same strategy could transfer to other structure-search generators, such as evolutionary or diffusion-based methods, or to other target properties like temperature and defect equilibria; the paper's benchmark protocol ties the result to one search pipeline, but the data-selection logic is general.","The virial advantage over baselines may partly reflect differences in DFT codes and pseudopotentials rather than physical accuracy alone; re-labeling all structures with one common DFT setting would test this.","Pressure-resolved open datasets like this one could make 'pressure MAE' a standard reporting metric for atomistic models, shifting evaluation from aggregate energy accuracy toward the enthalpy-relevant quantities that matter for high-pressure prediction."],"forward_implications":["A 1.5-million-configuration, pressure-diverse dataset yields virial error several times smaller than 12M–113M-scale baselines on a held-out trajectory test, with a lighter tail of large PV errors.","Models trained this way can relax unseen ternary structures to target pressures of 0, 50, 100, 150, and 200 GPa with mean deviations within about 3 GPa, whereas baselines drift to errors above 80% at 50 GPa.","In CSP recovery tests at 50–150 GPa, OpenCSP models reach success rates roughly 20–30 percentage points higher than the baselines; at 0 GPa all models are comparable at about 60%.","Increasing network depth from 6 to 24 layers steadily improves energy and force accuracy and cross-dataset generalization, with diminishing returns for force and virial beyond 12 layers."],"fun_headline_variants":["1.5M curated structures beat larger models under pressure","Pressure-aware data beats model scale in crystal prediction","Small pressure-tuned dataset outranks big universal models","Targeted pressure sampling wins over raw data volume"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The comparisons assume that literature-reported high-pressure structures recovered within 2,000 generated candidates are the right measure of prediction quality, and that DFT labels from different electronic-structure codes are comparable enough for direct error comparison.","fun_headline_variants_meta":{"raw":{"variants":["1.5M curated structures beat larger models under pressure","Pressure-aware data beats model scale in crystal prediction","Small pressure-tuned dataset outranks big universal models","Targeted pressure sampling wins over raw data volume"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000452,"raw_usage":{"total_tokens":2122,"prompt_tokens":768,"completion_tokens":1354,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":1300}},"tokens_in":512,"tokens_out":1354,"duration_ms":11612,"temperature":1.0,"reasoning_tokens":1300,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T17:58:11.441717+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate OpenCSP and the baselines on a held-out set of high-pressure phases produced by an independent search algorithm, using identical DFT settings for all models; if OpenCSP's crystal-structure-prediction success-rate and virial advantages shrink to noise, the claim that targeted pressure sampling is the cause is not supported.","supporting_citations":[],"review_version":1}