{"id":"3253de8c-58af-448e-bf2f-25b7839e1a40","arxiv_id":"2603.18876","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new bond-centric database and an element-specific covalent-bonding descriptor (BA) are fit from DFT ICOHP data, with only in-sample parity validation and no ML benchmark.","lead":"This paper introduces a database of chemical-bond-level electronic structure data for 36,377 inorganic crystals, plus a new element-by-element 'bonding attractivity' scale fit to 3.6 million bond records. The pitch is that giving machine-learning models these pre-computed bonding features will make materials-property predictions more accurate when training data is scarce.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"BA is validated only in-sample; no out-of-sample split or ML benchmark supports the abstract's claim that BA enables accurate predictions with limited data.","rationale":"The reader's weakest_assumption focuses on the product ansatz's neglect of orbital character and crystal-field effects. That is a real concern, and the paper itself concedes it for multi-orbital elements. However, the more load-bearing issue is that the central claim about ML utility is never tested: all parity plots are in-sample, and no ML experiment is reported. Even an imperfect product ansatz could still be a useful feature if it transfers to unseen structures; conversely, a perfect in-sample fit would not justify the abstract's 'limited data' promise. The reader's rationale does mention the absence of out-of-sample validation and ML benchmarks, so my stress test partially overlaps with the stated weakest_assumption. I recommend keeping the original CONDITIONAL verdict: the manuscript presents a substantial dataset and a compact descriptor, but the headline claim needs an explicit transferability test and an ML benchmark before acceptance. This is an addressable gap rather than a fundamental flaw, so no verdict change is needed.","tokens_in":16447,"tokens_out":3906,"duration_ms":45638,"concrete_test":"Release MattKeyBond and perform a crystal-level 80/20 holdout split: fit (η_A^0, L_A, M_A) on the training crystals, predict ICOHP for all bonds in held-out crystals, and report per-element MAE/R² against both the in-sample fit and a Pauling-EN + bond-length baseline. Then train a small-data ML model for one target property (e.g., formation energy or band gap) with and without BA as input features at training sizes such as 100, 500, and 2,000; if BA does not reduce held-out test error in the low-data regime, the limited-data claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section IV fits the three BA parameters of Eqs. 7-8 by least squares to all 3,665,789 ICOHP values and then validates with parity plots (Appendix Figs. 7-10). Because the same records are used to fit and to evaluate, the parity plots are in-sample diagnostics; they do not establish that BA predicts ICOHP for unseen bonds or crystal environments. The abstract's stronger claim—that BA 'enables accurate predictions even with limited data'—is never tested: no ML model is trained with or without BA, no small-data learning curve is shown, and no comparison against Pauling electronegativity or bond-valence baselines is provided. The paper itself concedes that for elements with multiple bonding orbitals (e.g., boron), environmental and orbital dependencies introduce additional complexity, meaning the fitted η_A may absorb environment-specific effects rather than represent an intrinsic, transferable element property. If so, BA is at best a compact interpolation of the database, not a physically predictive descriptor. This is the load-bearing gap: the headline utility claim rests on transferability, and transferability is unmeasured.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MattKeyBond, a bond-centric database built from high-throughput DFT calculations on 36,377 materials, using Closest Wannier Function (CWF) downfolding and crystal-orbital Hamilton population (ICOHP) analysis to produce bond-resolved electronic-structure data. Building on this database, the authors propose a new elemental descriptor, Bonding Attractivity (BA), defined by a multiplicative ansatz (Eq. 7) in which the ICOHP of an A–B bond is the product of element-specific BA values η_A(R,x_A) η_B(R,x_B), with each η obeying an exponential bond-length dependence and linear valence-state modulation (Eq. 8). The three parameters (η0_A, L_A, M_A) per element are obtained by least-squares fitting to 3,665,789 ICOHP records, and the resulting parameters are presented in a periodic-table map. The manuscript claims that BA is a physically interpretable, transferable descriptor that will relieve machine-learning models from relearning quantum mechanics from geometry and enable accurate predictions with limited data. The paper also emphasizes limitations, including the neglect of orbital character and crystal-field effects and the special complexity of elements with multiple bonding orbitals.","tokens_in":16756,"tokens_out":2945,"duration_ms":33899,"significance":"If the central claims are substantiated, MattKeyBond would be a valuable community resource: a large, bond-resolved database with explicit energy-dimensional descriptors, and BA could provide a compact, chemically interpretable scale of covalent bonding strength complementary to electronegativity. The paper has clear strengths: the high-throughput workflow is described in detail, the use of CWF for automated and parameter-light Wannier downfolding is timely, and the scale of the dataset (36,377 materials, 3.6M bond records) is substantial. The central scientific contribution, however, is the transferability of BA as a descriptor for predicting ICOHP and downstream material properties. That claim currently rests on in-sample parity plots and is not yet supported by out-of-sample tests, quantitative accuracy metrics, or ML benchmarks. The manuscript openly acknowledges the main limitations, which is commendable, but the abstract and summary currently overstate what has been demonstrated.","major_comments":[{"comment":"The validation of BA is entirely in-sample. The three parameters per element are obtained by least-squares fitting to all 3,665,789 ICOHP values, and the same records are then used to construct parity plots. This does not establish that BA can predict ICOHP for unseen bonds, elements, or crystal environments. The paper should provide an out-of-sample test, e.g., a random holdout split, a leave-one-structure-out split, or a split by chemical composition, and report quantitative error metrics (MAE, RMSE, R²) per element or per bond type. Without such evidence, the claim that BA 'enables accurate predictions' remains unsupported.","section":"Section IV, Eqs. (7)–(8); Appendix Figs. 7–10"},{"comment":"The headline utility claim—that BA enables accurate machine-learning predictions even with limited data—is never tested. No ML model is trained with or without BA features, no small-data learning curves are shown, and no comparison against baseline descriptors (e.g., Pauling electronegativity, bond valence, or plain geometric features) is provided. The abstract and summary should either be tempered to what is demonstrated (a compact parametrization of the database) or the manuscript should add the missing experiments that support the broader claim.","section":"Abstract; Section V; Section IV"},{"comment":"The product ansatz of Eq. (7) neglects orbital character, crystal-field effects, and magnetic state, and the paper itself concedes that 'for elements with multiple bonding orbitals (e.g., boron), stronger environmental and orbital dependencies introduce additional complexity.' This is load-bearing because the fitted η_A may absorb environment-specific effects rather than represent an intrinsic, transferable element property. The paper should quantify how much of the variance in the parity plots is attributable to such effects, and should state the domain of applicability more precisely. At a minimum, a per-element error analysis would show which elements fail the ansatz; currently the parity plots are too dense to judge this.","section":"Section IV, last paragraph; Eq. (7)"}],"minor_comments":[{"comment":"The phrase 'The fitting availability can also be found in the Fig3. 5∼8 of the Appendix' is unclear. It should say 'Parity plots for the fit are shown in Appendix Figs. 7–10.'","section":"Section IV, text near 'Fitting availability'"},{"comment":"The valence configuration for He is listed as '1s1', which is almost certainly a typo for '1s2'. Please check the orbital configurations for other elements as well, since they appear inconsistent in places (e.g., several entries list the same orbital count for different configurations).","section":"Table I"},{"comment":"In the periodic table layout, the element sequence appears to read '... 77 78 78 80 81 82 83', which duplicates atomic number 78 and omits 79 (Au). Please correct the figure.","section":"Figure 5"},{"comment":"The composite projection strategy involving f_n, D_an, and C_an is introduced without a clear explanation of why it avoids the influence of 'out shelled orbitals.' Please provide a more explicit justification or a reference.","section":"Appendix B, Eq. (B10) and surrounding text"},{"comment":"The notation ⟨RAa| D |R′Bb⟩⟨R′Bb| H |RAa⟩ is ambiguous: the second matrix element appears to lack a bra or ket. Please clean up the notation and define all operators consistently.","section":"Appendix C, Eq. (C5)"}],"recommendation":"major_revision","confidential_remarks":"The database and workflow are promising, and the authors are transparent about limitations. However, the central BA descriptor is currently validated only in-sample, and the abstract overclaims its predictive utility. The requested revisions—out-of-sample validation, quantitative error metrics, and a moderation of the ML claim—are feasible within the manuscript's scope. I do not see grounds for rejection, but the current form is not yet acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague —\n\nThe genuinely useful thing here is MattKeyBond: 36,377 materials with bond-resolved ICOHP and Wannier-based orbital-resolved data. That is a real resource, and the CWF downfolding pipeline is a sensible way to do it at scale. The paper also deserves credit for the BA idea as a compact, element-specific bond-strength scale analogous to electronegativity — that's a reasonable construct, and the periodic table trends (H above F, etc.) are physically interesting.\n\nBut the validation is not there. The three parameters per element are least-squares fit to all 3.6 million ICOHP records, and the parity plots in Figures 7–10 are just the fit reproducing the fitted data. That tells you about the flexibility of the functional form, not about whether BA transfers to bonds or structures outside the training set. The abstract says BA 'enables accurate predictions even with limited data,' but no ML model is trained, no small-data learning curve is shown, and no baseline comparison to Pauling electronegativity or bond-valence is given. That is a load-bearing gap.\n\nThe paper itself concedes the product ansatz of Eq. 7 is only accurate for elements with a single dominant valence orbital, and that boron and other multi-orbital cases introduce complexity the descriptor can't capture. That's not fatal — it's a simplification — but it means the fitted η_A values likely absorb environment-specific effects, so calling them intrinsic element properties is a stretch.\n\nAlso, no code or data link is provided, which makes the database unverifiable in practice. The authors say they will release it, but for a paper whose main product is data, that shouldn't be a promise.\n\nAll of this is fixable. Release the data, add out-of-sample splits (e.g., hold out entire structures or elements), report errors with spread, and run a simple ML comparison with and without BA as a feature. That would take maybe a month of work.\n\nVerdict: worth a serious referee, because the database is valuable and the descriptor idea is testable. But I would not accept it in current form. It's a conditional at best.\n\nWould I cite it? If the database is released and the out-of-sample numbers come in, yes. As it stands, no.","headline":"Substantial new database, but the Bonding Attractivity descriptor is only validated on the data it was fit to; the abstract's prediction claim is unsupported.","tokens_in":17278,"tokens_out":2039,"would_cite":false,"duration_ms":20192,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single element-specific quantity, Bonding Attractivity, captures the covalent bond strength of any atom pair across the periodic table.","keywords":["Bonding Attractivity","ICOHP","chemical bonding descriptor","materials database","machine learning for materials","covalency","band structure","Wannier functions"],"falsifier":"Perform a held-out test: fit BA on a subset of crystal structures, then predict ICOHP for bonds in crystal structures excluded from the fit; if prediction errors for elements with multiple bonding orbitals (e.g., boron) markedly exceed the fit's internal scatter, the product ansatz is not portable.","tokens_in":16288,"feed_emoji":"⚛️","tokens_out":5799,"duration_ms":52679,"temperature":0.7,"pith_summary":"This paper argues that chemical bonding, usually hidden inside machine-learning black boxes, can be extracted into a simple element-specific descriptor: Bonding Attractivity (BA). BA assigns each element a baseline covalent-bonding strength, a length-decay scale, and a valence-state modulation factor, and it predicts the bond energy of any atom pair as the product of the two elements' BA values. The authors fit BA using 3.6 million bond-resolved energies computed from first principles for 36,377 crystals, covering hydrogen through bismuth. If BA holds up, machine-learning models for materials would no longer have to rediscover quantum mechanics from atomic coordinates; they could consume pre-computed bonding physics as features, making accurate predictions from limited data feasible.","feed_headline":"Bonding Attractivity scale predicts covalency from H to Bi","feed_subtitle":"Element-specific covalency descriptor from 3.6M bonds helps models learn from small datasets.","key_machinery":"The load-bearing mechanism is the multiplicative ansatz of Eq. (7): -ICOHP_AB = η_A η_B, with η_A(R,x_A) = η^0_A exp(-(R-2r_A)/L_A + M_A x_A). This ansatz turns bond strength into a separable product of two element-specific scalars, each described by a baseline attractivity, a characteristic decay length, and a valence-state modulation factor, so that the full periodic table is captured by 249 parameters rather than by per-pair bond data.","core_discovery":"The central claim is that the covalent part of any A–B bond's energy, measured by the integrated crystal orbital Hamilton population (ICOHP), equals the product of the Bonding Attractivities of the two atoms: -ICOHP_AB = η_A(R,x_A)·η_B(R,x_B), where each η is an element-specific exponential function of bond length R and valence state x. Fitted to over 3.6 million DFT-derived bond records, the three parameters (baseline η^0, decay length L, valence modulation M) form a compact periodic table of covalency. A key difference from electronegativity: hydrogen tops the BA scale while fluorine tops the electronegativity scale, reflecting BA's focus on orbital-hybridization energy rather than charge-","pith_inferences":["One testable extension: BA parameters fitted on crystals could be transferred to molecular or surface-adsorption systems, where similar hybridization physics governs bond strength, but the paper only validates on crystalline compounds.","The product ansatz deliberately drops orbital character; a natural refinement would be to assign separate BA values for σ and π channels, which may improve accuracy for elements with multiple active orbitals.","If the descriptor's transferability is confirmed on held-out crystal structures, it could replace learned latent atom embeddings in graph neural networks, shrinking training-set requirements for new chemistries.","The observed sign oscillations of the valence modulation factor suggest that BA could be used to test classical assumptions about how oxidation state strengthens or weakens covalent bonding."],"forward_implications":["Machine-learning models can consume BA as a pre-computed physical feature, reducing the need to infer bond physics from geometry and improving accuracy when training data are scarce.","BA provides a human-readable covalency scale across 83 elements, complementing electronegativity by separating hybridization energy from charge-transfer energy.","The database of 3.6 million bond-resolved ICOHP values enables systematic search, comparison, and classification of bonding at atom-pair resolution across thousands of compounds.","Bond-strength descriptors of this kind may serve as direct indicators for properties tied to interatomic force constants and electron-phonon coupling, such as hardness or superconductivity."],"fun_headline_variants":["Hydrogen tops new covalency scale, not fluorine","Bond energy = product of two element-specific parameters","3.6M bonds yield a periodic table of covalency","Black-box ML replaced by physical bond features"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire construction rests on the assumption that a single element-specific scalar, multiplied across a bond, can reproduce the bond's covalent energy regardless of orbital character and local crystal environment.","fun_headline_variants_meta":{"raw":{"variants":["Hydrogen tops new covalency scale, not fluorine","Bond energy = product of two element-specific parameters","3.6M bonds yield a periodic table of covalency","Black-box ML replaced by physical bond features"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00096,"raw_usage":{"total_tokens":3918,"prompt_tokens":725,"completion_tokens":3193,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":3128}},"tokens_in":469,"tokens_out":3193,"duration_ms":21610,"temperature":1.0,"reasoning_tokens":3128,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T17:50:27.089160+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Perform a held-out test: fit BA on a subset of crystal structures, then predict ICOHP for bonds in crystal structures excluded from the fit; if prediction errors for elements with multiple bonding orbitals (e.g., boron) markedly exceed the fit's internal scatter, the product ansatz is not portable.","supporting_citations":[],"review_version":1}