{"id":"89013db8-53d4-4128-95f6-9e680ba75665","arxiv_id":"2607.08518","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A multi-task graph model scores crystals for dynamical stability and enriches stable generative candidates ~2.7× on PhononBench and up to ~5–6× under hard DFT screening.","lead":"PhononScore ranks crystal structures by dynamical stability in seconds, without full phonon calculations. It can raise the share of stable candidates from generative models from about 31% to 84%, making large-scale materials screening far more practical.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Headline PhononBench enrichment is same-label MatterSim train/eval; transfer to independent DFT is only partly stress-tested.","rationale":"The paper’s strongest evidence is large-scale ranking enrichment under an explicit stability definition, with useful multi-task design, formula-level splits, ALIGNN ablation, and a DFT fine-tune story. That is enough for a methods contribution as a fast ranker under stated label regimes, with top candidates still needing high-fidelity phonon checks—exactly the reader’s CONDITIONAL framing. The single most load-bearing concern is not architecture or Top-K math; it is that the main 2.72× / 83.7% numbers are same-supervision-family (MatterSim train and MatterSim eval). Independent DFT transfer is shown on a different distribution and hard-screening is resampled, so the leap to general closed-loop readiness remains conditional. I agree with the reader’s weakest_assumption and do not raise a stronger internal inconsistency; no change of verdict is warranted beyond reinforcing CONDITIONAL until generator-pool DFT revalidation is reported.","tokens_in":24469,"tokens_out":625,"duration_ms":7501,"concrete_test":"Take the Top-100 PhononScore picks from each of the nine PhononBench generators (or a fixed 200–300 stratified subset), recompute harmonic ω_min with DFT-PBE (or a second independent uMLIP not used in training) under the paper’s −0.1 THz criterion; report true DFT Top-100 stable rates and enrichment vs the original pool. If average DFT Top-100 stability falls well below ~70–80% or enrichment collapses toward ~1×, the headline PhononBench claim does not transfer beyond MatterSim labels.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central operational claim (30.7%→83.7% Top-100, 2.72× enrichment on PhononBench) is evaluated with the same MatterSim+phonopy harmonic ω_min labels used for pretraining (Fig. 1a–b; Results “PhononScore enables efficient enrichment…”; Methods). Formula-level exclusion blocks compositional leakage but not shared-label-family bias: the model can learn MatterSim-specific soft-mode signatures and still look excellent when ranked against MatterSim ground truth. ALIGNN ω_min ablation (Table 1) and multi-threshold tables (Appendix B) show multi-task ranking helps under that label family, but do not prove the ranking matches DFT/experimental dynamical stability for generated candidates. PhononScore-DFT transfer (balanced Top-100 93%, hard-screen 5–6×) is real but on Materials Project-like DFT structures and synthetic resampling of a balanced 1k set—not on the nine-generator PhononBench pools with independent DFT phonons. Discussion already notes harmonic/anharmonic limits; that is the softest load-bearing hinge for “closed-loop inverse design” language.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript introduces PhononScore, a multi-task graph neural scoring function that ranks crystal structures by dynamical stability without explicit phonon calculations. Structures are encoded as periodic atom and line graphs; a shared encoder jointly optimizes minimum-phonon-frequency regression, multi-threshold stability classification, local geometry likelihood via a mixture density network, and a threshold-aware pairwise ranking loss, then combines standardized head outputs into a unified score. Training uses a multi-fidelity set of 157,463 structures (MatterSim-labeled generated and MP40 data for pretraining; DFT-PBE for fine-tuning to PhononScore-DFT), with formula-level train/test exclusion. On PhononBench pools from nine generators, PhononScore raises average Top-100 stability (ω_min > −0.1 THz) from 30.7% to 83.7% (2.72× enrichment; Top-10 ~97.5%), outperforming an ALIGNN ω_min baseline. PhononScore-DFT reaches 93% Top-100 on a balanced DFT-PBE set and ~5–6× enrichment under resampled hard-screening, with supporting calibration, case studies (e.g., K–I–O), and an online platform.","tokens_in":24811,"tokens_out":1486,"duration_ms":13631,"significance":"If the ranking claims hold under independent high-fidelity labels, PhononScore would fill a genuine gap between crystal generative models and expensive phonon validation, analogous to docking or confidence scores in drug discovery. Strengths include a large multi-fidelity dataset with formula-level exclusion, explicit multi-task and ranking objectives, a clear ALIGNN ablation (Table 1), multi-threshold robustness (Appendix Table 2), DFT transfer and hard-screening analyses, calibration-by-bin plots, nontrivial case orderings, and public code, data, and a web service. These make the work practically useful for high-throughput reranking and potentially for RL/active-learning feedback, even if full closed-loop inverse design remains aspirational.","major_comments":[{"comment":"Results “PhononScore enables efficient enrichment…” and Fig. 1a–b / Methods: the headline PhononBench claim (30.7%→83.7% Top-100, 2.72×) uses MatterSim+phonopy ω_min both as pretraining labels and as evaluation ground truth. Formula-level exclusion blocks compositional leakage but not shared-label-family bias. Table 1 shows multi-task ranking beats ALIGNN under that same label family; it does not establish that the ranking matches independent DFT (or experiment) on generator pools. The central claim for generated candidates should be restated as MatterSim-label enrichment unless a subset of PhononBench structures is revalidated with DFT phonons, or the DFT transfer section is elevated as the primary high-fidelity claim.","section":"Results; Fig. 1; Table 1"},{"comment":"Results “Transferability…” and hard-screening (Fig. 3d–e): PhononScore-DFT transfer is demonstrated on Materials Project-like DFT structures and on synthetic resampling of a balanced 1k set, not on the nine-generator PhononBench pools with independent DFT labels. Hard-screening enrichment (5–6×) is therefore conditional on that distribution. For the closed-loop / generator-reranking narrative, either report DFT phonons on a stratified sample of high- vs low-PhononScore generated candidates, or clearly limit claims about generator pools to MatterSim-level ranking and treat DFT results as transfer on MP-like chemistry.","section":"Results (DFT transfer / hard-screening); Discussion"},{"comment":"Methods (unified scoring) and Appendix C: evaluation uses within-pool z-score combination with fixed α=0.25, β=2.0 chosen on the MP20 generated validation benchmark for mean Top-100 stable rate. That is appropriate for pool reranking but makes absolute scores pool-dependent and couples weight selection to the same MatterSim-labeled generator distribution used in the headline metric. Report sensitivity of PhononBench and DFT metrics to α,β (beyond the heatmap) and state explicitly that single-CIF scores require inspecting head components rather than Seval.","section":"Methods; Appendix C"}],"minor_comments":[{"comment":"Abstract and Introduction: “nine” generators vs PhononBench “7” models and Appendix figures using eight sources—align counts and naming throughout.","section":"Abstract; Introduction; Fig. 2"},{"comment":"Eqs. (12)–(17) vs Appendix I: training-time mix (0.6 S_thr + 0.1 S_geom) differs from evaluation weights; state both clearly in one place to avoid confusion.","section":"Methods; Appendix I"},{"comment":"Fig. 2f / Fig. 3f: add space-group and composition labels consistently; ensure ω_min units and stability threshold are readable in all panels.","section":"Figures 2–3"},{"comment":"Discussion already notes harmonic/anharmonic limits; a short quantitative caveat on near-threshold (±0.1 THz) sensitivity would help readers use the score in practice.","section":"Discussion"},{"comment":"Minor typos and notation: “Candinates”, “Traning”, inconsistent ω_min vs ωmin, and arXiv-style citation formatting for journal submission.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The work is timely and the engineering contribution is real; the main risk for a high-impact materials journal is overclaiming “dynamical stability” for generated crystals when the flagship metric is same-pipeline MatterSim enrichment. If the authors reframe the PhononBench result as MatterSim-label ranking and keep DFT transfer as a separate, carefully scoped claim—or add even a modest independent DFT check on generated Top-K—the paper is close to publishable. Scope fits cond-mat.mtrl-sci / computational materials well."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is operational: they built a multi-task ranking score (ω_min regression + multi-threshold classification + geometry MDN + pairwise enrichment) that reranks generator pools in seconds and ships code, data, and a web tool. On PhononBench the Top-100 stability rate goes from 30.7% to 83.7% (2.72×), Top-10 ~97.5%; PhononScore-DFT hits 93% Top-100 on a balanced DFT set and ~5–6× under hard resampling. That is the right problem framing for generative materials work.\n\nWhat is actually new is not “another GNN on crystals,” but treating dynamical stability as a drug-discovery-style scoring problem with a ranking objective and a multi-fidelity stack (MatterSim pretrain → DFT fine-tune). Formula-level exclusion is the right leakage control. Table 1 vs ALIGNN ω_min, multi-threshold tables, bin-calibration plots, and the K–I–O / polymorph case series are honest evidence that multi-task ranking beats pure frequency regression under their labels. Citations and methods look standard and reproducible enough for a methods paper.\n\nSoft spot, in proportion: the headline PhononBench enrichment is train/eval on the same MatterSim+phonopy harmonic ω_min family. Formula exclusion blocks composition leakage, not shared-label bias. DFT transfer is real but on MP-like structures and synthetic hard pools, not independent DFT on the nine-generator pools. Discussion already flags harmonic/anharmonic limits; that is the hinge for “closed-loop inverse design” language, not a reason to dismiss the ranking results. Free weights (α, β, λ’s) are fixed after a validation sweep—fine if users treat the score as a ranker and still phonon-validate tops.\n\nThis is for people running crystal generators, active learning, or RL who need a cheap stability prior before MLIP/DFT phonons. I would bring it to reading group, cite the enrichment protocol and dataset, and send it to peer review. Engage as a calibrated ranker under stated label regimes, not as a substitute for high-fidelity phonons.","headline":"Solid, usable scoring-function paper for a real bottleneck; headline PhononBench numbers are same-label MatterSim, but DFT transfer and ablations still make it worth engaging.","tokens_in":25460,"tokens_out":549,"would_cite":true,"duration_ms":6121,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"PhononScore ranks crystal candidates for dynamical stability in seconds and raises the average stable share of generator pools from 30.7% to 83.7% in the top 100.","keywords":["crystal generation","scoring function","dynamical stability","phonons","materials screening","graph neural networks","multi-fidelity learning","reranking"],"falsifier":"Hold out a large set of generator candidates never seen in training, rank them with PhononScore, then recompute harmonic phonons with independent high-fidelity DFT-PBE for the top-100 versus a random sample of the same pool; if the true DFT stable fraction in the top-100 does not substantially exceed the pool baseline under ω_min > −0.1 THz, the central enrichment claim fails.","tokens_in":25302,"feed_emoji":"⚛️","tokens_out":724,"duration_ms":12344,"temperature":0.7,"pith_summary":"Crystal generators can invent huge numbers of candidate materials, but most of them are dynamically unstable and would never form real solids. This paper argues that you do not need a full phonon calculation for every candidate: a learned scoring function, PhononScore, can read the crystal structure and output a single stability score fast enough for large-scale reranking. Trained on a multi-fidelity set of more than 150,000 phonon-labeled crystals, it lifts the average dynamical-stability rate of pools from nine generators from 30.7% to 83.7% among the top 100, with top-10 rates near 97.5%. A DFT-finetuned version further enriches rare stable structures under hard, imbalanced screening. If that ranking signal holds, generators, active learning, and closed-loop design can get a second-level stability feedback loop without paying for explicit phonons on every structure.","feed_headline":"Score crystals for phonon stability in seconds, not minutes","feed_subtitle":"A learned ranker lifts stable candidates from ~31% to ~84% in top-100 generator pools.","key_machinery":"PhononScore: a graph-neural scoring function on periodic atom and line graphs that jointly learns minimum phonon frequency, multi-threshold stability classification, local geometry likelihood, and a threshold-aware ranking loss, then combines the standardized heads into one unified reranking score.","core_discovery":"The authors claim that dynamical-stability screening for generated crystals can be cast as learning a calibrated ranking score rather than predicting full phonon spectra, and that their multi-task PhononScore—pretrained on large MatterSim-labeled data and optionally fine-tuned on DFT-PBE phonons—recovers true stability order well enough to deliver roughly 2.7× enrichment on PhononBench (30.7% → 83.7% top-100 average) and about 5–6× enrichment under scarce-stable DFT hard-screening, at second-level cost.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["PhononScore ranks crystal dynamical stability in seconds","Learned score lifts generator pools 31% to 84% stable","2.7x enrichment of stable crystals without full phonons","Second-level phonon-aware ranking for crystal generation","DFT-tuned PhononScore hits 93% top-100 dynamical stability"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The main enrichment numbers rest on treating harmonic minimum phonon frequencies from the same machine-learning potential workflow as faithful labels for both training and judging stability, so if those labels mis-order near-threshold or anharmonic crystals, the reported gains can look large without matching real first-principles or experimental stability.","fun_headline_variants_meta":{"raw":{"variants":["PhononScore ranks crystal dynamical stability in seconds","Learned score lifts generator pools 31% to 84% stable","2.7x enrichment of stable crystals without full phonons","Second-level phonon-aware ranking for crystal generation","DFT-tuned PhononScore hits 93% top-100 dynamical stability"]},"model":"grok-4.5","effort":"low","cost_usd":0.006368,"raw_usage":{"total_tokens":1696,"prompt_tokens":904,"num_sources_used":0,"completion_tokens":69,"cost_in_usd_ticks":63680000,"prompt_tokens_details":{"text_tokens":904,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":723,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":904,"tokens_out":69,"duration_ms":6347,"temperature":1.0,"reasoning_tokens":723,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T06:05:06.149248+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Hold out a large set of generator candidates never seen in training, rank them with PhononScore, then recompute harmonic phonons with independent high-fidelity DFT-PBE for the top-100 versus a random sample of the same pool; if the true DFT stable fraction in the top-100 does not substantially exceed the pool baseline under ω_min > −0.1 THz, the central enrichment claim fails.","supporting_citations":[],"review_version":1}