{"id":"e55b9d3d-009f-4dad-8626-f7dd373c8e32","arxiv_id":"2505.15363","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Replacing neural-network activation functions with physics-based ones, plus soft Garvey-Kelson and bound constraints, cuts extrapolation error for nuclear masses from 1173 keV to 396 keV on the outermost measured nuclei.","lead":"A neural network that swaps its usual math for physics-based functions predicts the masses of rare atomic nuclei much better than standard neural networks, including in regions never used for training. The approach works by baking in simple physics rules, and it could give astrophysicists better numbers for the exotic nuclei forged in stellar explosions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported extrapolation RMS is computed on validation+test combined, so the headline 396 keV is not a clean out-of-sample number; test-only performance is never reported.","rationale":"The reader's stated weakest assumption is the Garvey-Kelson penalty's validity far from stability. That is a real concern for how the method will perform on genuinely unknown nuclei, and it is worth testing by evaluating GK residuals on the measured outermost nuclei. However, I think the most load-bearing issue for the central claim as stated is more basic: the reported extrapolation error is computed on the union of validation and test sets, while model selection was performed on the validation extrapolation subset. The test subset was explicitly created to avoid selection bias and then never evaluated separately. This is evident from Sec. III, Sec. IV, and the Table II caption. The concern is concrete and easily settled by re-running the evaluation on the held-out test subset. It is also independent of the GK physics assumption. Since the reader's rationale already flags this evaluation issue but did not make it the weakest assumption, I mark partial agreement and keep the CONDITIONAL verdict.","tokens_in":8355,"tokens_out":3554,"duration_ms":32031,"concrete_test":"Split Table II's extrapolation RMS by subset: using the exact model checkpoint selected by validation extrapolation RMS (Sec. IV), compute RMS on only the held-out extrapolation test subset defined in Sec. III, separately for PAF and DenseNet, and report the ensemble spread across the 16 models. If the PAF test-only RMS is not clearly below DenseNet's (or exceeds roughly 600 keV), the headline '396 keV vs 1173 keV' should be replaced by the test-only comparison.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The numerical anchor of the central claim is Table II's extrapolation RMS (396 keV vs 1173 keV). The table caption states it is the 'combined validation and test datasets,' and Sec. IV says the model was selected using the lowest RMS on the extrapolation validation set, with activation functions and hyperparameters tuned on that same set. Hence the reported 396 keV includes nuclei whose errors influenced model selection and early stopping; it is not a fully out-of-sample estimate. The paper itself (Sec. III) motivates separate test sets 'to avoid selecting a model biased on the validation datasets,' but no test-only extrapolation RMS is given anywhere. Without that number, the magnitude of the improvement over DenseNet is unverified and could be substantially inflated by selection bias. The GK and bound penalties (Eqs. 3-4) are a distinct concern for generalization to unmeasured drip-line nuclei, but the evaluation gap directly undermines the claimed measured-region extrapolation accuracy.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a neural-network model for nuclear masses that uses only N and Z as inputs, with physics-related activation functions (PAF) replacing conventional activations, L0 regularization for sparsity, and two training penalties: a bound penalty and a soft Garvey-Kelson constraint. Using AME2020 data divided into training, validation, and test sets for both interpolation and extrapolation, a 16-member ensemble is trained and evaluated. The central claim is that the PAF network extrapolates far better than conventional DenseNet and ConvNet architectures, with reported extrapolation RMS errors of 396 keV versus 1173 keV for DenseNet (N, Z > 0), and that it reproduces proton and neutron drip lines.","tokens_in":8436,"tokens_out":5292,"duration_ms":50451,"significance":"The physics-informed activation-function idea is a genuinely interesting and credible direction for extrapolation in nuclear mass tables. The authors are careful to avoid using global mass models or magic-number knowledge, and they provide a comparison with conventional networks and several mass models, along with useful interpretability checks. If the reported extrapolation performance survives an unbiased evaluation, the paper would be a valuable contribution to the machine-learning nuclear-physics literature. However, the headline extrapolation error combines the validation set used for model selection with the held-out test set, so the central quantitative claim is not yet convincingly established; a test-only evaluation is needed before the comparison can be trusted.","major_comments":[{"comment":"The headline extrapolation RMS in Table II is computed on the combined validation and test datasets, while Sec. IV states that model selection, early stopping, activation-function choices, and hyperparameters were tuned using the lowest RMS on the extrapolation validation set. The reported 396 keV therefore includes nuclei whose errors influenced the selected model, and it is not a fully out-of-sample estimate. Sec. III explicitly motivates separate test sets to avoid selecting a model biased on the validation datasets, but no test-only extrapolation RMS is reported anywhere. Please report test-only RMS for PAF, DenseNet, and ConvNet, and give validation and test RMS separately for both interpolation and extrapolation; if the test-only numbers differ materially from Table II, the main quantitative comparison must be revised.","section":"Sec. IV and Table II"},{"comment":"The alternative evaluation using newly updated AME2020 data is incomplete. The text says 'We trained the same DNN and PAF networks,' but the numbers that follow are only for ConvNet (426 and 446 keV) and PAF (251 and 273 keV). Since the central comparison of the paper is PAF versus DenseNet, the absence of the DenseNet result makes the concluding sentence 'Clearly, the RMS errors on such data configuration do not fully reflect the poor extrapolation performance of DNNs' unsupported. Please include DenseNet (and ideally all compared models) in this evaluation or temper the conclusion.","section":"Sec. V (AME2020 vs AME2012 evaluation)"},{"comment":"The Garvey-Kelson penalty is applied to random points in the extrapolation region with a fixed threshold G = 0.5 u, which implicitly assumes that the GK relation remains approximately valid far from stability, including near the drip lines. The paper gives no evidence for this assumption in the unmeasured region and no sensitivity analysis on G or on the bound range in Eq. (3). A concrete robustness check—for example, varying G and the bound limits, or measuring GK residuals on held-out outermost nuclei as a function of distance from the training region—would substantially strengthen the extrapolation claim. This is particularly relevant because the paper claims coverage up to the drip lines, where experimental constraints are sparse.","section":"Sec. II, Eqs. (3)-(5)"}],"minor_comments":[{"comment":"Please state the number of nuclei in the interpolation and extrapolation sets, separately for the validation and test splits and for the N, Z ≥ 8 and N, Z > 0 selections; the current table reports only total counts.","section":"Table II"},{"comment":"The agreement of the predicted drip lines with experiment is presented qualitatively; a quantitative metric, such as the mean distance in N and Z between the predicted and experimental drip lines, would make the claim more precise.","section":"Fig. 4"},{"comment":"The statement that 'DNNs with different activation functions, structure, and regularization also gave no significant improvements in the extrapolation regions' is not backed by a table or figure; please list the tested variants and their RMS values, or remove the claim.","section":"Sec. V"},{"comment":"The Iverson-bracket notation in the L0 norm is not defined; a one-sentence definition would help readers outside the machine-learning community.","section":"Eq. (2)"},{"comment":"The paper appropriately notes that the deep-ensemble uncertainties may be less well calibrated, but the caption of Fig. 3 should state explicitly that the outer-region trimming is based on an approximate ensemble-standard-deviation cutoff.","section":"Sec. IV and Fig. 3"}],"recommendation":"major_revision","confidential_remarks":"The work is within the scope of the journal and I see no signs of questionable research practices. The main issue is an evaluation-transparency gap that the authors can likely fix by re-running their saved ensembles and reporting test-only RMS values; the AME2020-vs-AME2012 comparison should also include the DenseNet result. I would not require code release, but more architectural detail and a sensitivity analysis of the GK and bound penalties would improve reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, useful paper for the ML-for-nuclear-masses niche, and the central idea—replacing generic activations with physics-motivated functions, plus sparse regularization and soft physics penalties—genuinely changes extrapolation behavior. The drip-line reproduction is the most convincing visual evidence, and the paper is honest about its own limitations, including poorly calibrated uncertainties and the need for future work.\n\nThe main soft spot is evaluation. Table II reports combined validation+test extrapolation RMS, and the model was selected on the validation subset. No test-only extrapolation RMS appears anywhere. That means the 396 vs 1173 keV comparison could be inflated by selection bias. This is load-bearing for the quantitative claim, so it matters more than a minor quibble. The second evaluation using newly measured AME2020 nuclei absent from AME2012 gives PAF 273 keV extrapolation RMS, but the corresponding DenseNet number is not reported, so it does not replace the missing test-only comparison.\n\nThe physics penalties are transparently stated, but the Garvey-Kelson relation is assumed to remain roughly valid near the drip lines, and the bound range is hand-tuned. These are secondary concerns because the penalties are soft and the paper is upfront about them. A bigger practical issue is that no code or complete architecture details are provided; reproduction would be painful. The ablation showing that N*x and Z*x matter is useful, and the citation pattern is fine.\n\nWho is this for? Nuclear physicists using machine learning for mass predictions, particularly for r-process astrophysics. The paper deserves a serious referee. I would ask for a clean test-only extrapolation RMS, a DenseNet comparison on the AME2020-vs-AME2012 split, and at least a commitment to release code before acceptance.","headline":"The core result is plausible, but the reported 396 keV extrapolation RMS is not a clean out-of-sample number because it combines the validation set used for model selection; a test-only number is needed before trusting the magnitude of the improvement.","tokens_in":9102,"tokens_out":2113,"would_cite":true,"duration_ms":21445,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["21.10.Dr","07.05.Mh"],"model":"deepseek-v4-flash","headline":"A neural network that swaps generic computer-science activation functions for physics-related ones extrapolates nuclear masses to 396 keV RMS beyond the measured region, a threefold improvement over a conventional DenseNet, and reproduces…","keywords":["physics-related activation functions","nuclear mass extrapolation","neural networks","L0 regularization","Garvey-Kelson relation","drip lines","AME2020","sparse learning"],"falsifier":"Train a conventional ReLU DenseNet with exactly the same bound penalty, Garvey-Kelson penalty, and outer-ring validation protocol as the PAF network; if its extrapolation RMS falls from 1173 keV to near 396 keV, the extrapolation gain is caused by the physics penalties and not by the activation-function replacement, which would refute the paper's central attribution.","tokens_in":8024,"feed_emoji":"⚛️","tokens_out":13038,"duration_ms":103556,"temperature":0.7,"pith_summary":"The paper claims that the poor extrapolation of neural networks for nuclear masses is not intrinsic to deep learning but follows from the generic, computer-science-motivated activation functions used inside the network. Replacing those functions with elementary physics-related operations, such as multiplication, division, logarithm, sine, and combinations built from neutron number N and proton number Z, and pruning the network with L0 sparsity, gives a model that uses only N and Z as inputs and reaches 396 keV RMS on the outermost AME2020 nuclei not used in training. This is roughly a threefold improvement over a conventional DenseNet (1173 keV) and is achieved without any global mass-model prediction or knowledge of magic numbers. The same network also forms proton and neutron drip lines close to the experimental ones, so the result matters because it suggests neural mass models can be made to extrapolate reliably into the unknown regions where nuclear astrophysics and other applications need masses.","feed_headline":"Physics activations cut nuclear mass extrapolation error to 396 keV","feed_subtitle":"Replacing standard activations with physics-based ones triples extrapolation accuracy and recovers drip lines.","key_machinery":"The key object is the PAF (physics-related activation function) network, a dense network whose hidden nodes apply a menu of physics-motivated functions: $1/N$, $1/Z$, $1/A$, $\\log(x)$, $\\sin(x)$, parity terms, $N/Z$, $Z/N$, $N^x$, $Z^x$, $A^{1/3}$, $A^{2/3}$, and $((N-Z)/A)x$, with inputs derived only from neutron number $N$ and proton number $Z$. L0 sparsity, implemented with hard-concrete stochastic gates, prunes weights and keeps the model far from over-parametrization. Two extra soft penalties applied during training carry the extrapolation: a bound penalty that keeps mass-excess predictions inside $[-0.1, 2]$ u in the unmeasured region, and a Garvey-Kelson penalty that discourages the six-mass combination $\\Delta(N,Z)$ from deviating by more than 0.5 u. These components, rather than a larger network, are what let the model extrapolate.","core_discovery":"The central claim, stated on the authors' own terms, is that a feed-forward network whose hidden nodes apply physics-related activation functions, rather than ReLU-like computer-science functions, can extrapolate nuclear masses beyond the measured landscape in a physically consistent way. Trained only on 1,736 of the 2,548 AME2020 masses and guided by a bound penalty plus a soft Garvey-Kelson penalty on randomly sampled extrapolation points, an ensemble of sixteen such networks predicts the outermost nuclei with 396 keV RMS error (N, Z > 0), compared with 1173 keV for a conventional DenseNet. The model is not given any existing mass-model prediction or magic-number input, yet its activation outputs show structure at magic numbers and its predicted separation-energy landscape terminates at two-nucleon drip lines close to the experimental ones. The authors therefore conclude that the extrapolation weakness of neural networks in this problem is dominated by the choice of nonlinear functions and by regularization of the unmeasured region, not by network architecture or data size.","pith_inferences":["Beyond the paper, the same PAF recipe is a testable strategy for other nuclear observables such as charge radii, beta-decay half-lives, or fission barriers, where generic networks extrapolate poorly; if the gain transfers, activation-function engineering is a general route to physics extrapolation.","We infer that an ablation separating the Garvey-Kelson penalty from the PAF activation functions would clarify whether the extrapolation gain comes mostly from the physics-guided regularization or from the functions themselves; the paper does not isolate these two contributions.","Because the bound penalty's lower limit was chosen from the measured mass excess of $^{118}$Sn, the model's extrapolation is normalized to known extremes; predictions outside that range should be treated with caution until the bound is re-derived from a wider set of measured nuclei."],"forward_implications":["A neural network with no magic-number input produces the known shell structure: activation outputs display kinks at magic numbers, so the model offers an interpretable path to see where physics effects enter.","The drip lines emerge even though most near-drip-line nuclei were excluded from training, implying the PAF network captures the physics needed to separate bound from unbound nuclei.","Extrapolation at 396 keV on the outer ring of AME2020, without using any global mass model, suggests the method can supply mass predictions for nuclei beyond current measurements.","The spread across the sixteen-model ensemble gives a per-nucleus uncertainty estimate, so the PAF approach can provide calibrated ranges for applications, not only central values.","Omitting or altering the $N^x$ and $Z^x$ functions degrades extrapolation by 10 to 20 percent, indicating that these terms are what the nuclear mass system is most sensitive to."],"supporting_citations":[{"why":"introduces the idea that replacing activation functions with scientific functions helps extrapolation, which the PAF approach extends","marker":"[14]"},{"why":"provides the L0 regularization with hard-concrete gates used to prune network parameters","marker":"[17]"},{"why":"supplies the bound penalty that keeps extrapolation predictions inside a physically reasonable mass-excess range","marker":"[22]"},{"why":"gives the Garvey-Kelson relation used as a soft penalty on randomly sampled extrapolation points","marker":"[28]"},{"why":"supplies the AME2020 experimental masses from which the training, validation, and test splits are drawn","marker":"[29]"},{"why":"supports the deep-ensemble strategy used to aggregate predictions and estimate uncertainties","marker":"[31]"}],"fun_headline_variants":["Physics activations triple nuclear mass extrapolation accuracy","Physics-based activations lift neural net extrapolation 3x","Magic numbers emerge: physics activations improve mass extrapolation","Neural net with physics activations cuts error to 396 keV","Physics activations recover nuclear drip lines in extrapolation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Garvey-Kelson relation, an algebraic connection among six neighboring nuclear masses, stays approximately true for nuclei far outside the measured region, including near the drip lines; if it breaks down there, the training penalty pushes predictions toward wrong values and the reported extrapolation accuracy would not survive.","fun_headline_variants_meta":{"raw":{"variants":["Physics activations triple nuclear mass extrapolation accuracy","Physics-based activations lift neural net extrapolation 3x","Magic numbers emerge: physics activations improve mass extrapolation","Neural net with physics activations cuts error to 396 keV","Physics activations recover nuclear drip lines in extrapolation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000533,"raw_usage":{"total_tokens":2560,"prompt_tokens":936,"completion_tokens":1624,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":1543}},"tokens_in":552,"tokens_out":1624,"duration_ms":11099,"temperature":1.0,"reasoning_tokens":1543,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:19:15.038312+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a conventional ReLU DenseNet with exactly the same bound penalty, Garvey-Kelson penalty, and outer-ring validation protocol as the PAF network; if its extrapolation RMS falls from 1173 keV to near 396 keV, the extrapolation gain is caused by the physics penalties and not by the activation-function replacement, which would refute the paper's central attribution.","supporting_citations":[{"cited_title":"Speciﬁcally, the bound range of mass excesses was set to [−0.1, 2] u (atomic mass unit) by deﬁning B1, B2=0.95, 1.05 u","cited_arxiv_id":null,"evidence_quote":"supplies the bound penalty that keeps extrapolation predictions inside a physically reasonable mass-excess range"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"gives the Garvey-Kelson relation used as a soft penalty on randomly sampled extrapolation points"},{"cited_title":"Lakshminarayanan, A","cited_arxiv_id":null,"evidence_quote":"supports the deep-ensemble strategy used to aggregate predictions and estimate uncertainties"}],"review_version":1}