{"id":"ae99334c-efe0-4fa9-b456-c28c712f4460","arxiv_id":"2505.10750","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"CatBoost corrections bring six Skyrme-HFB nuclear mass models to roughly 0.2 MeV test accuracy while preserving generalization to newly measured nuclei.","lead":"This paper uses the CatBoost machine learning algorithm to correct nuclear mass predictions from six Skyrme Hartree-Fock-Bogoliubov models. It reports that the corrections reduce test errors to about 0.2 MeV and also improve predictions for 21 nuclei measured after AME2020.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The ~0.2 MeV test accuracy is measured on random 80/20 splits that are effectively interpolation in a strongly correlated Z,N residual surface; the external generalization test uses only 21 nuclei and the paper's own extrapolation test degrades with distance, so the 'good generalization' claim…","rationale":"The paper's empirical work is internally consistent: bare HFB residuals are in line with previous mass tables, the 500-run random splits give stable rms values, and the 21 newly measured nuclei are a useful external check. I credit these results. However, the load-bearing assumption for the generalization claim is that the residual is smooth in the selected features, and the paper's evidence for this is weaker than the headline suggests. The random 80/20 split is an interpolation test, the external set is only 21 nuclei with no distance control, and the paper's own drip-line extrapolation test shows accuracy eroding with distance. The reader's weakest assumption correctly identifies the smooth-residual assumption, but the sharper issue is that the validation strategy does not separate interpolation from extrapolation. A spatial-block cross-validation would settle this directly. The reader's conditional verdict is therefore appropriate; I would keep it conditional and explicitly require this additional validation before the generalization claim is accepted.","tokens_in":24559,"tokens_out":7564,"duration_ms":83621,"concrete_test":"Run a spatial-block cross-validation: for each held-out nucleus, exclude from training all nuclei within |ΔZ| ≤ 2 and |ΔN| ≤ 2 (or, more simply, all nuclei in the same isotopic chain), using the same M8 features and the already-tuned hyperparameters, and report the held-out rms for each Skyrme set. If this blocked rms stays below ~0.3 MeV, the generalization claim is supported; if it rises to ≳0.5 MeV, the ~0.2 MeV test result is an interpolation artifact rather than evidence of learned missing physics.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central empirical claim is that CatBoost-refined HFB models reach ~0.2 MeV rms on held-out nuclei and generalize to newly measured nuclei. This requires the residual delta(Z,N) = B_HFB - B_exp to be a smooth, learnable function of the seven M8 features so that training on random 80% of AME2020 transfers to unseen points. The paper's main validation, in Figs. 3 and 5 and Table V, is a random 4:1 split. Because HFB residuals are strongly correlated in Z,N, a random split is mostly interpolation between close neighbours, so it does not strongly test extrapolative generalization. The only out-of-distribution checks are the 21 nuclei outside AME2020 and the drip-line extrapolation test in Fig. 16. The 21-nucleus set is small and no analysis is given of how far these nuclei are, in feature space, from the training region; Fig. 16 shows rms increasing with extrapolation distance and, at the largest distance (pi4), the refined model is no better than the bare HFB model. The Summary itself concedes that 'in the very extreme nuclear regions... the extrapolation performance of the machine learning will be unreliable.' Thus the claim of 'good generalization abilities' is not yet quantitatively bounded; the 0.2 MeV numbers may be interpolation performance rather than evidence that CatBoost has learned the underlying physics. A secondary, related issue is that the M8 feature set and hyperparameters are selected using the same data that later defines test sets, with no nested cross-validation, so the reported test errors are mildly optimistic.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies the CatBoost gradient-boosting algorithm to refine six Skyrme-HFB mass tables (SkM*, SkP, SLy4, SV-min, UNEDF0, UNEDF1) by learning the residual delta(Z,N) = B_HFB - B_exp from AME2020 data. Seven nuclear features are used (Table I), eight feature combinations are compared (Table II), and the M8 combination is selected using 10-fold cross-validation. After a grid search over hyperparameters (Table III), the reported test-set rms values are 0.189-0.259 MeV, with model-repair coefficients RMR = (sigma_HFB - sigma_test)/sigma_HFB between 84.5% and 96.3% (Table V). The authors also test 21 nuclei outside AME2020, reporting post-refinement rms values around 0.14-0.29 MeV, and examine extrapolation to drip-line nuclei in Fig. 16. They conclude that CatBoost can 'pick up the missing physics' and that the refined predictions are nearly independent of the bare model, while conceding in the Summary that extrapolation in very extreme regions is unreliable.","tokens_in":24979,"tokens_out":4565,"duration_ms":45819,"significance":"If the 0.2 MeV refined accuracy and the generalization claim held, this would be a practically useful contribution to nuclear mass prediction, with potential impact on r-process abundance calculations and the interpretation of new mass measurements. The paper's strengths are the systematic six-force comparison, the repeated random-split evaluation with 500 runs and uncertainty estimates, the use of external AME2020 data, and the out-of-sample check with 21 newly measured nuclei. The machine-checked style of the numerical tables is a useful feature. However, the central generalization claim is currently overstated: the evaluation protocol does not provide a fully independent test set for feature and hyperparameter selection, and the extrapolation evidence is either small in sample size (21 nuclei) or shows performance degradation with distance (Fig. 16). The interpolation performance is solid, but the extrapolation claim needs additional support before the paper can be accepted.","major_comments":[{"comment":"The evaluation protocol does not separate model selection from final testing. The M8 feature set is chosen using 10-fold cross-validation on the full dataset (Fig. 3), and the optimal hyperparameters in Table IV are selected by a further 10-fold CV on the same full dataset. The reported test-set rms values in Table V are then obtained from random 4:1 splits of that same dataset. This means the hyperparameters and feature set may have been partially tuned on nuclei that later appear in the test folds, so the 0.189-0.259 MeV test values are not fully independent of model selection. Please use a nested cross-validation scheme or hold out a completely untouched test set for the final evaluation, and report how the test rms changes when feature/hyperparameter selection is performed only on training folds.","section":"Section III, Figs. 3, 5, 8 and Table IV"},{"comment":"The abstract and the Summary claim 'good generalization abilities', but the paper's own extrapolation test in Fig. 16 shows that the refined-model rms increases with extrapolation distance, and at the largest proton-rich distance (pi4) the CatBoost-refined model is not better than the bare HFB model. The Summary itself concedes that 'in the very extreme nuclear regions ... the extrapolation performance of the machine learning will be unreliable'. Additionally, the 21-nucleus test in Fig. 15 is not accompanied by an analysis of how far those nuclei are, in the chosen feature space, from the training region; most of them lie near the stability line and are plausibly close neighbors of training nuclei. The claim of 'good generalization' should therefore be restricted to interpolation or supported by a quantitative feature-distance analysis and a clear statement of the distance at which performance degrades.","section":"Section III, Fig. 16 and Section IV"},{"comment":"The claim that the refined predictions 'hardly depend' on the bare HFB model is contradicted by the construction B_ml = B_th - delta, where delta is trained on B_th - B_exp. The residual is a function of the bare model by construction, so the weak Pearson correlation between the bare and refined rms values across 500 splits does not establish independence of the bare model. The correct statement is that the residual fit removes most of the large model bias, not that the refined predictions are nearly independent of the underlying physical model. Please rephrase the text around Fig. 14 accordingly.","section":"Section III, Fig. 14"},{"comment":"The model-repair coefficient RMR = (sigma_HFB - sigma_test)/sigma_HFB is presented as a measure of the algorithm's 'model-repair ability'. However, sigma_HFB is computed over all 2457 nuclei while sigma_test is computed on a random 20% subset, and the test set is also used for early stopping during training. The reported RMR values of 84.5-96.3% are therefore optimistic in a way that should be acknowledged. If nested cross-validation is adopted as suggested above, the RMR values should be recomputed on the held-out folds only.","section":"Table V"}],"minor_comments":[{"comment":"There are several typos: 'the the Hartree-Fock-Bogoliubov', 'paraterer', 'accurancies', and 'Intrestingly'. These should be corrected.","section":"Abstract"},{"comment":"The scaling factor F is defined as 1/(delta_max - delta_min), but it is not stated over which dataset (training set, testing set, or all data) the maximum and minimum residuals are taken.","section":"Fig. 11 caption"},{"comment":"The iteration domain [1000, 5000] with increment 2000 yields only 1000, 3000, and 5000. If 2000 and 4000 were also scanned, this should be stated; otherwise the domain/increment combination should be corrected.","section":"Table III"},{"comment":"The vertical axis label in Fig. 16(b) is missing. The text refers to 'rms derivations', but the axis should explicitly state the quantity and its units.","section":"Fig. 16"},{"comment":"The loss function L(y, F(x)) is introduced abstractly, but the actual metric used throughout the paper is the rms residual. For self-containedness, please state explicitly that the squared-error loss is used in the CatBoost training.","section":"Eq. (1)"},{"comment":"Reference [45] (Mass Explorer) is an online resource; please include an access date and, if possible, a version or a permanent identifier.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a competent empirical comparison within the journal's scope. The main issue is the gap between the strong generalization claim in the abstract and the actual validation protocol: model selection is not nested inside the final test splits, and the only out-of-sample generalization evidence is either small (21 nuclei) or shows degradation with extrapolation distance. These points are fixable with additional analysis and a more careful wording of the claims. No concerns about scientific integrity are raised by the manuscript content."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does what it says: CatBoost refinement of six Skyrme-HFB mass tables gives test rms around 0.2 MeV, stable over 500 random splits, with a check on 21 post-AME2020 nuclei. The new bits are the systematic six-force comparison, the SHAP/feature analysis, and the observation that the best bare force (UNEDF0) is not the best refined one. That last ranking reversal is interesting, though probably just regression to the mean: larger bare rms leaves more for the ML to learn.\n\nCredit where due: the authors use default hyperparameters for the initial comparison, run 500 repetitions, show the distributions, and honestly report that extrapolation degrades with distance and becomes unreliable in extreme regions. That is more honest than many ML papers. The tables and figures are clear, and the citation pattern is normal for this field.\n\nThe soft spots are real but not fatal. First, the feature set M8 and the hyperparameters are selected on the same data that later defines the test sets; without nested cross-validation, the 0.2 MeV numbers are mildly optimistic. Second, the construction Bml = Bth - delta means the refined prediction obviously depends on the bare model, so the near-zero correlations in Fig. 14 are an artifact of comparing rms across splits, not evidence of independence. The paper's wording in Section III overclaims there. Third, the abstract still says 'good generalization abilities' even though the Summary walks it back; the 21-nuclei test is too small and lacks any feature-space distance analysis, and Fig. 16 shows the refined model stops beating the bare one at the largest extrapolation distance. Those are moderate issues, not load-bearing flaws. The RMR comparison across forces is also a bit apples-to-oranges, but the authors themselves flag that.\n\nWho is this for: practitioners who want a quick CatBoost recipe for mass tables and a sense of which Skyrme interaction benefits most. It is a solid empirical benchmark, not a conceptual breakthrough. I would send it to peer review rather than desk reject. I would ask for nested CV, a correction to the independence claim, and a quantitative statement about interpolation versus extrapolation. With those revisions, the central result—~0.2 MeV test accuracy after CatBoost refinement—should hold.","headline":"A competent, incremental CatBoost application to six Skyrme-HFB mass tables with solid ~0.2 MeV interpolation accuracy, but the 'generalization' claim needs tightening and the independence claim is wrong.","tokens_in":25496,"tokens_out":2130,"would_cite":true,"duration_ms":24312,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports that a CatBoost machine-learning refinement of six Skyrme-HFB nuclear mass tables cuts test-set errors to roughly 0.2 MeV and generalizes to nuclei measured after the training data.","keywords":["nuclear mass","machine learning","CatBoost","Skyrme force","Hartree-Fock-Bogoliubov","binding energy","mass model refinement","generalization"],"falsifier":"Train the CatBoost-refined model on nuclei in one region, say $Z<50$, and test on heavier nuclei with a similar split, or evaluate the 21 newly measured nuclei after excluding all nuclei in their isotopic chains from training; if the rms error jumps far above 0.2 MeV or fails to beat the bare HFB table, the smooth-residual assumption is falsified.","tokens_in":24366,"feed_emoji":"⚛️","tokens_out":9560,"duration_ms":92335,"temperature":0.7,"pith_summary":"The paper sets out to show that a gradient-boosting algorithm called CatBoost can repair the binding-energy predictions of six Skyrme-Hartree-Fock-Bogoliubov mass tables, bringing test-set root-mean-square errors from 1.4-7.0 MeV down to roughly 0.2 MeV. The authors train CatBoost on the difference between each theoretical mass and the experimental mass in the 2020 Atomic Mass Evaluation, using seven features based on proton number, neutron number, parity, and distance to magic numbers. They report that every refined model reaches about 0.2 MeV test error, reduces large model bias, and still agrees with 21 masses measured after the training data. The practical interest is that microscopic mean-field mass tables, often too inaccurate for nuclear-structure or astrophysical applications, could be upgraded as a package rather than re-fitted.","feed_headline":"CatBoost trims Skyrme-HFB mass errors to 0.2 MeV","feed_subtitle":"The same trained models also match 21 nuclei measured after the training data, so the correction is not just curve fitting.","key_machinery":"The load-bearing object is the mass residual $\\delta(Z,N)=B_{\\rm th}(Z,N)-B_{\\rm exp}(Z,N)$ between a Skyrme-HFB table and experiment, together with the CatBoost gradient-boosting algorithm that learns $\\delta$ from seven features: $Z$, $N$, $N/Z$, even-odd flags ${\\rm Zeo},{\\rm Neo}$, and distances $|Z-m|$, $|N-m|$ to magic numbers. CatBoost's ordered boosting replaces standard gradient estimates to reduce prediction shift; the refined mass is then $B_{\\rm ml}=B_{\\rm th}-\\delta_{\\rm learned}$. The claimed repair is quantified by a model-repair coefficient $R_{\\rm MR}=(\\sigma_{\\rm HFB}-\\sigma_{\\rm test})/\\sigma_{\\rm HFB}$, which exceeds 84% for all six forces.","core_discovery":"For each of six Skyrme parameter sets (SkM*, SkP, SLy4, SV-min, UNEDF0, UNEDF1), the paper's central claim is that the residual $B_{\\rm HFB}(Z,N)-B_{\\rm exp}(Z,N)$ is a learnable function of the seven-feature set and that CatBoost can learn it from the measured nuclei with $Z,N\\ge 8$. After hyperparameter tuning with ten-fold cross-validation, the refined models reach testing-set rms deviations of 0.189-0.259 MeV, model-repair coefficients above 84%, and residuals that become visibly more random across the nuclear chart. On 21 nuclei measured after AME2020, the refined models give rms deviations of 0.137-0.293 MeV. The authors also claim that the best bare Skyrme force (UNEDF0, 1.43 MeV) is not the best refined force, indicating that the algorithm captures different missing physics for different interactions.","pith_inferences":["A testable extension is to treat CatBoost as a universal repair layer: if the residual remains smooth for other mean-field mass tables, the same seven-feature pipeline should produce comparable roughly 0.2 MeV test errors without changing the physics model.","The reversal of the best-force ranking after refinement suggests model selection should be done on refined, not bare, predictions; retraining the six models when the next atomic mass evaluation appears could check which Skyrme force then yields the best extrapolation.","Because the refined masses are smooth in $Z$ and $N$, they can be used to refit the coefficients of simple liquid-drop-type mass formulas; the paper already does this, and the same trick could generate improved constraints on Skyrme energy-density functionals.","The near-0.2 MeV test floor, larger than the roughly 25 keV experimental uncertainty, may reflect irreducible model randomness; one can probe this by testing whether the error saturates as training data grow, which would indicate how much missing physics is recoverable from present features."],"forward_implications":["All six adopted Skyrme-HFB tables can be repaired to roughly 0.2 MeV test accuracy, bringing microscopic mass models closer to the level needed for astrophysical applications.","The refined models generalize to 21 nuclei measured after AME2020, with rms deviations of 0.137-0.293 MeV, so the correction is not limited to nuclei already in the training set.","The residual distributions become more random and the large model bias decreases, indicating that the algorithm is capturing part of the missing physics rather than only memorizing the training data.","The ranking of Skyrme forces by predictive power changes after refinement: UNEDF0 is best in the bare HFB calculations, but UNEDF1 gives the best refined test error, so model selection should be based on refined rather than bare performance.","Extrapolation to drip-line nuclei worsens with distance from the training region, but the refined predictions often remain better than the bare HFB predictions except for the most distant proton-rich cases."],"supporting_citations":[{"why":"Supplies the experimental binding energies from AME2020 that define the residual target and the training/testing data.","marker":"[4]"},{"why":"Supplies the CatBoost algorithm with ordered boosting, the method used to learn the mass residuals.","marker":"[32]"},{"why":"Supplies the HFB solver whose mass calculations underlie the six theoretical tables being refined.","marker":"[42]"},{"why":"Supplies the six Skyrme-HFB mass tables used as the theoretical baselines for the residual learning.","marker":"[45]"},{"why":"Provides the feature-selection scheme and the roughly 0.2 MeV accuracy baseline from earlier machine-learning mass refinement against which CatBoost results are compared.","marker":"[34]"},{"why":"Supplies the earlier tree-based model-repair study and the extrapolation protocol adapted in this paper.","marker":"[37]"},{"why":"Supply the 21 newly measured nuclear masses outside AME2020 used to test the generalization ability of the refined models.","marker":"[78–88]"}],"fun_headline_variants":["CatBoost cuts nuclear mass errors to ~0.2 MeV","Best Skyrme force loses crown after CatBoost refinement","CatBoost refines Skyrme-HFB masses to 0.2 MeV","CatBoost mass model generalizes to 21 new nuclei"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the difference between a Skyrme-HFB binding energy and the measured value varies smoothly with proton number, neutron number, parity, and distance to magic numbers, so patterns learned on measured nuclei carry over to unmeasured and extrapolated nuclei; if that difference has random or local structure the features cannot capture, the claimed accuracy would not persist.","fun_headline_variants_meta":{"raw":{"variants":["CatBoost cuts nuclear mass errors to ~0.2 MeV","Best Skyrme force loses crown after CatBoost refinement","CatBoost refines Skyrme-HFB masses to 0.2 MeV","CatBoost mass model generalizes to 21 new nuclei"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000989,"raw_usage":{"total_tokens":4291,"prompt_tokens":1140,"completion_tokens":3151,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":756,"completion_tokens_details":{"reasoning_tokens":3078}},"tokens_in":756,"tokens_out":3151,"duration_ms":21629,"temperature":1.0,"reasoning_tokens":3078,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:04:49.906773+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the CatBoost-refined model on nuclei in one region, say $Z<50$, and test on heavier nuclei with a similar split, or evaluate the 21 newly measured nuclei after excluding all nuclei in their isotopic chains from training; if the rms error jumps far above 0.2 MeV or fails to beat the bare HFB table, the smooth-residual assumption is falsified.","supporting_citations":[{"cited_title":"Prokhorenkova, G","cited_arxiv_id":null,"evidence_quote":"Supplies the CatBoost algorithm with ordered boosting, the method used to learn the mass residuals."},{"cited_title":"Stoitsov, J","cited_arxiv_id":null,"evidence_quote":"Supplies the HFB solver whose mass calculations underlie the six theoretical tables being refined."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the six Skyrme-HFB mass tables used as the theoretical baselines for the residual learning."},{"cited_title":"Gao, Y.-J","cited_arxiv_id":null,"evidence_quote":"Provides the feature-selection scheme and the roughly 0.2 MeV accuracy baseline from earlier machine-learning mass refinement against which CatBoost results are compared."},{"cited_title":"Liu, H.-L","cited_arxiv_id":null,"evidence_quote":"Supplies the earlier tree-based model-repair study and the extrapolation protocol adapted in this paper."}],"review_version":1}