{"id":"9941caa0-8c39-4fa9-a6e2-7184988ab18b","arxiv_id":"2501.05341","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"An equivariant graph neural network predicts spin-crossover energies for thousands of metal complexes and enriches candidate detection about four-fold over random screening.","lead":"A machine-learning model trained on 1,439 density-functional-theory spin-switching energies was used to screen more than 11,000 transition-metal complexes from the Cambridge Structural Database for spin-crossover candidates. The authors report that their relevance-based classifier roughly quadruples the chance of finding a candidate compared with random selection, aiming to speed up materials discovery for quantum information devices.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported 4-fold enrichment and 17-fold reduction are computed from a Bayes table (81/167, 1065/1129) that does not match the stated 80% precision and 1.7% FPR operating point; the two metric sets imply different confusion matrices.","rationale":"The paper's contribution includes a new 1,439-entry r2SCAN spin-switching-energy database, a compact 915-parameter equivariant GNN, and a large-scale CSD screen. The regression MAE of 360 meV and the coordination-shell analysis are credible and independently useful; the paper also concedes the debatable nature of the −200 to 500 meV label window in its conclusions. However, the headline screening benefit — the nearly four-fold enrichment and 17-fold reduction in redundant calculations — is quantitative and is repeated in the abstract and conclusion. That claim is not reproducible from the text because the reported precision-recall operating point (80% precision, 35% recall, 1.7% FPR) and the Bayes posterior (56% precision, 94% rejection) cannot arise from the same confusion matrix. The counts 81/167 and 1065/1129 are inconsistent with a 144-molecule held-out test set and are never sourced to a specific split or threshold. This directly changes the practical yield of the screen: applying 56% instead of 80% precision to the 1,076 predicted candidates lowers the expected number of true candidates from ~861 to ~603. The reader's weakest-assumption choice (DFT label reliability) is a valid external-validity caveat, but the more immediate, checkable threat is the internal inconsistency of the headline metrics. Conditional acceptance is appropriate: the authors should report the held-out test contingency table at the deployed relevance threshold and recompute the enrichment and rejection factors from that table. If the corrected statistics still support the main conclusions, the paper stands; if not, the quantitative claims need to be revised.","tokens_in":14023,"tokens_out":8809,"duration_ms":83685,"concrete_test":"Obtain or reconstruct the model's predictions on the exact held-out test set and build the 2x2 contingency table at relevance 0.9 and at relevance 0.5 (TP, FP, FN, TN, plus test-set class prevalence). Check whether the table at 0.9 yields TP≈58/FP≈15 (precision 80%, recall 35%, FPR 1.3%) or TP≈81/FP≈64 (precision 56%, recall 48%, FPR 5.7%). Then recompute enrichment as precision divided by test prevalence and redundant-calculation reduction as 1/FPR, using the same table that justifies the 1,076-molecule screen. If the corrected numbers differ from the published '56%' and '94%' / 17-fold claims, revise the headline statistics and the expected number of true candidates among the 1,076 predicted molecules accordingly.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"In the Results section (Figure 5 discussion), the chosen relevance threshold 0.9 is reported with precision ≈ 80%, recall ≈ 35%, and a false-positive rate of just 1.7%. Immediately afterward, Bayes' theorem is invoked with P(ΔÊ)=167/1296, P(ΔE|ΔÊ)=81/167, and P(ΔE|ΔÊc)=1−1065/1129, yielding 0.5586, which is quoted as a four-fold enrichment over 13% random picking, and the 94% refusal / 17-fold reduction is derived from the same table. These two descriptions are numerically incompatible. If the 1296-set contains 167 true positives and 1129 true negatives, then recall 35% implies TP ≈ 58; precision 80% then implies FP ≈ 15 and FPR ≈ 1.3%, giving a 1/FPR ≈ 77-fold reduction, not 17-fold. Conversely, TP=81 and FP=64 (the Bayes table) gives precision 55.9%, recall 48.5%, and FPR 5.7% — a different operating point from the selected threshold. Moreover, 81 and 1065 cannot be held-out test counts for a ~144-molecule test set; they appear to be training/CV aggregates or a different threshold, but no provenance is given. The large-scale screen applies the model at relevance 0.9 to 1,076 molecules and infers ~861 good candidates from the 80% test precision; if the true precision at that threshold were the Bayes value of 56%, the expected number would be ~603. The central enrichment, rejection, and downstream candidate-count claims therefore rest on an unreported and internally inconsistent contingency table.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper constructs a data set of 1,439 mononuclear transition-metal complexes from the Cambridge Structural Database, computes r2SCAN spin-switching energies ΔEHL for each, and trains an equivariant graph neural network with 915 trainable parameters to predict ΔEHL. On a 10% holdout set the model achieves MAE 360 meV and R² 0.88. The regression is converted into a classifier through a relevance function built from sigmoids on the interval −200 ≤ ΔEHL ≤ 500 meV; at relevance 0.9 the authors report precision ≈80%, recall ≈35%, and false-positive rate 1.7%. They then use Bayes' theorem with the numbers 167/1296, 81/167, and 1065/1129 to claim a 56% posterior probability, a four-fold improvement over random picking, a 94% refusal probability, and a 17-fold reduction in redundant DFT calculations. The model is applied to 11,356 CSD molecules, yielding 1,076 predicted candidates, of which about 861 are expected to be good based on the 80% test precision. The paper concludes by acknowledging that the chosen energy window is debatable and depends on the functional and on omitted zero-point corrections.","tokens_in":14365,"tokens_out":5763,"duration_ms":58517,"significance":"The screening pipeline is potentially useful: the data set of DFT-computed spin-switching energies, the compact equivariant model, and the relevance-based classification scheme are all sensible components, and the coordination-shell analysis agrees qualitatively with the spectrochemical series. The use of a standard holdout split, the small parameter count, and the explicit statement of the label definition are strengths. However, the headline enrichment and reduction claims are currently not internally consistent, and the numerical basis for them is not reported with sufficient provenance. The practical significance of the work therefore cannot be assessed until the statistical claims are reconciled.","major_comments":[{"comment":"The operating point at relevance 0.9 is incompatible with the contingency table used in the Bayes calculation. The text reports precision ≈80%, recall ≈35%, and FPR ≈1.7% at relevance 0.9. The Bayes calculation uses P(ΔÊ)=167/1296, P(ΔE|ΔÊ)=81/167, and P(ΔE|ΔÊ^c)=1−1065/1129. This table implies recall = 81/167 = 48.5%, FPR = 64/1129 = 5.7%, and precision = 81/(81+64) = 55.9%, none of which matches the stated operating point. In addition, counts such as 81, 167, 1065, and 1129 cannot be held-out test counts for a test set of roughly 144 molecules; their provenance (training set, cross-validation aggregate, or a different threshold) is not given. Because the four-fold enrichment, the 94% refusal probability, and the 17-fold reduction are all derived from this Bayes table, the central screening claims must be recomputed from one explicitly sourced confusion matrix at the selected relevance threshold.","section":"Results and discussion, Figure 5(b) and the Bayes paragraph"},{"comment":"The random-picking baseline used for the four-fold claim is the training-set prevalence 167/1296 ≈ 12.9%. The large-scale screen is performed on a different pool of 11,356 CSD molecules, whose composition and prevalence of candidates need not match the training set. The enrichment factor is not well-defined for the actual discovery scenario unless the baseline prevalence is stated for the screened pool or an explicit assumption is made that the training prevalence transfers. Please clarify which baseline is being used and recompute the enrichment factor accordingly.","section":"Results and discussion, Bayes paragraph and large-scale screen"},{"comment":"The statement that approximately 861 of the 1,076 predicted materials are expected to have ΔEHL within the range of interest is obtained by multiplying 1,076 by the 80% test precision. Since the 80% precision and the precision implied by the Bayes table (≈56%) are mutually inconsistent, this expectation is not currently supported. Recomputing with the Bayes-table precision gives roughly 603 instead of 861, a material difference. The expected number of true candidates must be derived from the same consistent confusion matrix and threshold used for the other screening claims.","section":"Results and discussion, large-scale screen paragraph"}],"minor_comments":[{"comment":"The sentence beginning 'The choice of ligands The attractive feature...' is incomplete and should be rewritten.","section":"Introduction, p. 2"},{"comment":"The same sigmoid form is used for the lower and upper relevance thresholds, but the parameters c and s are not specified separately for each boundary; please give the explicit values used.","section":"Equation (3) and Figure 5(a)"},{"comment":"Because the holdout set is only about 10% of 1,439 molecules, the reported precision and recall values are subject to substantial sampling uncertainty; reporting exact test-set counts and confidence intervals would make the operating point more interpretable.","section":"Results and discussion, Figure 5(b)"},{"comment":"There is a typo: 'minimizing the the number of redundant calculations' contains a duplicated 'the'.","section":"Results and discussion, p. 15"},{"comment":"The abstract and conclusions state the improvement as a property of the model, but the improvement is relative to the computational label −200 ≤ ΔEHL ≤ 500 meV; this qualification should appear wherever the four-fold and 17-fold numbers are quoted.","section":"Concluding remarks"}],"recommendation":"major_revision","confidential_remarks":"The inconsistency between the stated precision/recall/FPR operating point and the Bayes contingency table is the central issue: the abstract's headline numbers depend on it. The issue appears fixable by reporting a single, sourced confusion matrix and recomputing the derived quantities, so I recommend major revision rather than rejection. The regression results and the data set itself seem sound and useful."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this paper gives you a genuinely useful dataset and a plausible regression pipeline, but the headline screening claims don't hold together. The 1,439-entry r2SCAN spin-switching database for Cr/Mn/Fe/Co complexes is a real resource, and the EGNN with 915 parameters reaches a test MAE of 360 meV. The coordination-shell analysis shows the model is actually learning metal-centered ligand-field chemistry, not just memorizing sizes—that's the most interesting part of the paper.\n\nThe trouble is the classification numbers. The text quotes precision ≈ 80%, recall ≈ 35%, and FPR ≈ 1.7% at relevance 0.9, then immediately uses Bayes' theorem with 167/1296, 81/167, and 1−1065/1129 to get a 56% posterior and a four-fold enrichment. Those two descriptions are numerically incompatible. The Bayes table implies precision ≈ 56% and FPR ≈ 5.7%, not 80% and 1.7%. And 81/167 can't be held-out test counts, since the test set is only ~144 molecules. You could rescue the calculation by saying the Bayes numbers come from cross-validation or a different threshold, but the paper doesn't say that. As written, the central 'four-fold improvement' and '17-fold reduction' claims rest on an unexplained, inconsistent contingency table. The downstream count of ~861 expected good candidates among 1,076 should be ~603 if the true precision is 56%.\n\nOther soft spots: the −200 to 500 meV window is admittedly debatable and excludes zero-point corrections; the screen applies the model to molecules up to 280 atoms, far beyond the training distribution, without an applicability-domain analysis; and code/weights aren't public, only descriptors on request. These are addressable.\n\nI'd send this to peer review — the dataset and the interpretability analysis deserve referee time — but the revision must pin down a single, consistent confusion matrix for the operating threshold that the authors actually recommend. As written, the screening claims are overstated, even though the underlying model and data are worth engaging with.","headline":"Useful dataset, credible regression, but the four-fold enrichment claim is internally inconsistent.","tokens_in":14988,"tokens_out":4625,"would_cite":true,"duration_ms":42812,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a small equivariant graph neural network, trained on density functional theory spin-switching energies for 1,439 transition-metal complexes, can identify spin-crossover candidates with roughly four times the hit…","keywords":["spin crossover","equivariant graph neural networks","materials screening","density functional theory","r2SCAN","relevance-based classification","transition metal complexes"],"falsifier":"Take a random sample of 50 to 100 of the 1,076 predicted candidates and recompute $\\Delta E_{\\mathrm{HL}}$ with a higher-level method such as CASPT2 or CCSD(T), or add zero-point corrections at the DFT level; if the fraction that still falls inside the $-200$ to $500$ meV window is far below the claimed about 80% precision, the screening gain is inflated. A smaller set of 20 to 30 synthesized and magnetically measured complexes would provide a direct test of the predictions.","tokens_in":13773,"feed_emoji":"🧲","tokens_out":7180,"duration_ms":65485,"temperature":0.7,"pith_summary":"This paper claims that a small equivariant graph neural network can screen for spin-crossover candidates far more efficiently than random high-throughput searching. The authors assemble 1,439 mononuclear Cr, Mn, Fe, and Co complexes, compute their low-spin to high-spin energy differences with r2SCAN density functional theory, and train a 915-parameter graph network that predicts those energies with a test mean absolute error of 360 meV. They then convert the regression into a classifier using a relevance function centered on the candidate window $-200 \\le \\Delta E_{\\mathrm{HL}} \\le 500$ meV, obtaining a 56% posterior probability of finding a true candidate versus 13% for random picking, about 80% precision at 35% recall, and a 94% refusal probability that reduces redundant DFT calculations by roughly 17-fold. Applied to 11,356 additional complexes, the model predicts 1,076 potential spin-crossover materials, most of them iron-based.","feed_headline":"Graph network quadruples spin-crossover candidate hit rate","feed_subtitle":"Trained on 1,439 DFT energies, it rejects unsuitable molecules at 94% confidence, cutting redundant runs.","key_machinery":"The load-bearing object is an equivariant graph neural network, a message-passing architecture equivariant to rotations, translations, and permutations, modified with an attention layer that feeds atom oxidation states and bond orders into the edges. A single convolutional layer keeps the model to 915 trainable parameters. The regression output is converted into a candidate classifier through a relevance function built from two sigmoids centered at the interval boundaries, following the precision-recall-for-regression method; Bayes' theorem then turns precision and recall into the reported posterior and refusal probabilities. A coordination-shell ablation, in which graphs are built from progressively larger shells centered on the transition metal versus on a random atom, serves as the diagnostic showing that the model learns the local dominance of the metal coordination environment.","core_discovery":"The paper's central claim is that a graph neural network with only 915 trainable parameters can learn the relationship between molecular structure and spin-switching energy well enough to act as a high-throughput pre-screen for spin crossover. Trained on r2SCAN DFT values of $\\Delta E_{\\mathrm{HL}}$ for 1,439 mononuclear complexes drawn from a crystallographic structure database, the network reaches a test mean absolute error of 360 meV and $R^2 = 0.88$. Converting the regression into a classification with a relevance function whose sigmoid thresholds sit at the candidate-window boundaries yields about 80% precision at 35% recall on the held-out set, a 56% posterior probability of finding a candidate versus 13% for random picking, and a refusal probability near 94% that corresponds to an approximate 17-fold reduction in redundant Kohn-Sham DFT calculations. On a broader set of 11,356 additional complexes, the model predicts 1,076 plausible spin-crossover candidates, with Fe-based species dominating the list.","pith_inferences":["The 56% posterior is computed using the training-set prevalence as the prior; in a different chemical library with a different base rate of candidates, the posterior would shift even if precision and recall stayed the same.","Because the label window excludes zero-point and thermal contributions, the 1,076 candidates should be read as \"worth a higher-level or experimental check\" rather than confirmed spin-crossover materials; a prospective test on a subset would be the natural next step.","The fourfold gain is measured against random picking from the same database; in a real pipeline where chemists already bias their searches toward known spin-crossover ligand families, the practical advantage over that stronger baseline could be smaller.","The same equivariant architecture combined with relevance-based classification could transfer to other narrow-window material properties, such as singlet-triplet gaps or redox potentials, wherever DFT labels define a target interval."],"forward_implications":["If the claimed enrichment holds, screening a large structure database can be reduced from tens of thousands of DFT calculations to a few hundred follow-up calculations on the predicted candidates.","The coordination-shell result implies that only the inner coordination environment, roughly five shells around the metal, carries the essential information, so future screening models could be built from smaller molecular fragments.","The dominance of Fe among the 1,076 predicted candidates matches the known prevalence of Fe(II) spin-crossover chemistry and suggests the model can rank ligand families for targeted synthesis.","The relevance threshold gives an adjustable operating point: tightening it trades recall for precision, so experimentalists can choose a stricter cutoff that yields a smaller, higher-confidence candidate list."],"supporting_citations":[{"why":"Supplies the equivariant message-passing architecture that the model is built on.","marker":"[56]"},{"why":"Provides the crystallographic structures of the molecular complexes used to build the training and screening datasets.","marker":"[58]"},{"why":"Defines the r2SCAN exchange-correlation functional used to compute the spin-switching energy labels.","marker":"[59]"},{"why":"Shows that r2SCAN gives reasonable spin-crossover energy differences for metal complexes, justifying the label method.","marker":"[60]"},{"why":"Documents the accuracy barriers in high-throughput spin-crossover screening that motivate the generous candidate window.","marker":"[62]"},{"why":"Supplies the relevance-function method that converts the regression output into a precision-recall classifier.","marker":"[88]"},{"why":"Provides the gradient-boosting decision-tree baseline whose inferior precision-recall performance supports the graph network's advantage.","marker":"[67]"}],"fun_headline_variants":["Graph net with 915 parameters quadruples spin-crossover hits","Tiny graph network boosts spin-crossover candidate finding 4x","Equivariant GNN pre-screens spin-crossover materials 4x better","Spin-crossover screening: GNN refuses 94%, 17x fewer DFT runs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the r2SCAN density functional energies define ground truth: a complex counts as a spin-crossover candidate exactly when its $\\Delta E_{\\mathrm{HL}}$ falls between $-200$ and $500$ meV, and the paper itself concedes this window is debatable and that zero-point corrections were omitted. Any systematic bias in these labels is inherited by the 56% enrichment figure and the 1,076 predicted candidates.","fun_headline_variants_meta":{"raw":{"variants":["Graph net with 915 parameters quadruples spin-crossover hits","Tiny graph network boosts spin-crossover candidate finding 4x","Equivariant GNN pre-screens spin-crossover materials 4x better","Spin-crossover screening: GNN refuses 94%, 17x fewer DFT runs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000548,"raw_usage":{"total_tokens":2579,"prompt_tokens":869,"completion_tokens":1710,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":1629}},"tokens_in":485,"tokens_out":1710,"duration_ms":15393,"temperature":1.0,"reasoning_tokens":1629,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:15:04.731577+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of 50 to 100 of the 1,076 predicted candidates and recompute $\\Delta E_{\\mathrm{HL}}$ with a higher-level method such as CASPT2 or CCSD(T), or add zero-point corrections at the DFT level; if the fraction that still falls inside the $-200$ to $500$ meV window is far below the claimed about 80% precision, the screening gain is inflated. A smaller set of 20 to 30 synthesized and magnetically measured complexes would provide a direct test of the predictions.","supporting_citations":[{"cited_title":"G.; Welling, E","cited_arxiv_id":null,"evidence_quote":"Supplies the equivariant message-passing architecture that the model is built on."},{"cited_title":"R.; Bruno, I","cited_arxiv_id":null,"evidence_quote":"Provides the crystallographic structures of the molecular complexes used to build the training and screening datasets."},{"cited_title":"W.; Kaplan, A","cited_arxiv_id":null,"evidence_quote":"Defines the r2SCAN exchange-correlation functional used to compute the spin-switching energy labels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that r2SCAN gives reasonable spin-crossover energy differences for metal complexes, justifying the label method."},{"cited_title":"G.; Trickey, S","cited_arxiv_id":null,"evidence_quote":"Documents the accuracy barriers in high-throughput spin-crossover screening that motivate the generous candidate window."},{"cited_title":"?N i Cs79 ]4] > MĜ8 W= hڕ","cited_arxiv_id":null,"evidence_quote":"Supplies the relevance-function method that converts the regression output into a precision-recall classifier."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the gradient-boosting decision-tree baseline whose inferior precision-recall performance supports the graph network's advantage."}],"review_version":1}