{"id":"303b30ca-f41c-4793-b22f-ce8d469e92e5","arxiv_id":"2607.14039","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A heteroscedastic Bayesian Model Combination, with uncertainty growing with model disagreement and distance in principal-component space, gives better-calibrated extrapolation errors for nuclear separation energies in the Z=46-52 region.","lead":"This paper extends a Bayesian method for combining nuclear models so that the predicted uncertainty grows as models extrapolate away from known data, instead of staying fixed. In tests on heavy nuclei near the proton and neutron drip lines, it reports better-calibrated error bars than the constant-uncertainty version.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"d_i, the key distance predictor in Eq. (13b), is never defined; as written, PC coordinates are measured from the model-mean origin, not from the training-data centroid, so the HeEM's error growth may not reflect true extrapolation distance.","rationale":"The reader's weakest assumption identifies a concrete, verifiable flaw: the key predictor d_i in the heteroscedastic error model is not defined, and the natural interpretation from Eq. (12) is inconsistent with the stated meaning. This is the most load-bearing concern because the central claim is that the error model is physically motivated by extrapolation distance. If d_i is not actually a distance from training data, the error model's superior calibration could be an artifact of the predictor choice rather than a genuine reflection of extrapolation uncertainty. The synthetic data experiments are the paper's strongest evidence, but they inherit the same ambiguity. The issue is addressable by a precise definition and re-analysis, so a CONDITIONAL verdict is appropriate. I agree with the reader's assessment and see no reason to change the verdict.","tokens_in":16278,"tokens_out":10997,"duration_ms":106668,"concrete_test":"Recompute all HeEM results with d_i explicitly defined as the Euclidean distance from the training-set centroid in PC space (d_i = ||zeta_i - zeta_train_mean||), and also run a variant without the d_i term (sigma_i^2 = alpha + gamma v_i). If the MACE and reduced chi^2 improvements over HoEM persist in both cases, the ambiguous d_i is not load-bearing; if they disappear, the claimed calibration superiority is an artifact of the undefined distance predictor.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that HeEM improves calibration rests on the error model in Eq. (13b) using a meaningful distance from training data. The paper states d_i is 'distance from the center of the training data in PC space' but never defines it. Eq. (12) defines zeta_i = x_c^i V, where x_c^i is the row of centered model predictions (centered by subtracting the model mean phi_0(x_i)). The origin of this PC space is the ensemble-mean model, not the centroid of training-set projections. A nucleus far from known data but with model consensus can have small ||zeta_i|| and thus small d_i, while a well-known nucleus with large model spread can have large d_i. If d_i does not measure extrapolation distance, the physically motivated rationale for the growing error bands is weakened, and the calibration gain may be driven primarily by the inter-model variance v_i or by the validation-set selection. The synthetic tests do not resolve this because they use the same ambiguous d_i.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces a heteroscedastic extension of Bayesian Model Combination (BMC) for nuclear mass predictions. In place of a constant error scale (HoEM), the error variance is modeled as σ_i^2 = α + β_1 d_i + γ_1 v_i, where d_i is described as a distance in principal-component space and v_i is the variance among the six EDF predictions (Eq. 13b). The parameters are inferred jointly with the BMC weights on AME2020 training data. The paper validates the method with leave-one-model-out synthetic experiments (Tables I–II) and with held-out validation data for Z = 46–52 (Tables III–V), reporting improved reduced χ² and MACE for the linear HeEM. It then applies HeEM to separation energies, Q_α, and drip-line existence probabilities.","tokens_in":16574,"tokens_out":4552,"duration_ms":45787,"significance":"The proposal is timely: the constant-variance assumption is a recognized limitation in multi-model nuclear UQ, and the heteroscedastic form is simple and physically motivated. The synthetic validation is a genuine strength: six independent leave-one-out experiments with different surrogate models provide evidence that the calibration gain is not an artifact of one benchmark. The real-data validation against AME2020 and the two new experimental comparisons (Cd masses, 104Te Q_α) are also valuable. If the ambiguities around the distance predictor and validation protocol are resolved, the paper would make a solid contribution to uncertainty quantification in nuclear theory.","major_comments":[{"comment":"The key predictor d_i is never defined. The text calls it 'distance from the center of the training data in PC space,' but Eq. (12) defines coordinates ζ_i by projecting the centered deviation vector x_c^i = x_0^i − φ_0(x_i) onto the SVD axes. The origin of this PC space is the ensemble-mean model, not the centroid of the training-set projections. A nucleus far from known data with model consensus can have small ||ζ_i||, so d_i as written need not measure extrapolation distance. Since d_i is the first driver in Eq. (13b), the physical rationale for the expanding HeEM bands is not established. Please define d_i explicitly (e.g., distance from the training-data centroid in PC space), report its empirical correlation with validation residuals, and state whether the reported results are robust to that definition.","section":"§III.A, Eqs. (12)–(13b)"},{"comment":"The validation set is used both to select hyperparameters (the number of retained components p) and to choose between linear and quadratic HeEM, and the same set is then used to report the headline calibration metrics in Tables I–V. This double use can bias the reported improvements. Please either split off a separate test set, perform nested cross-validation, or at minimum quantify how much of the HeEM advantage survives when p and model form are fixed a priori.","section":"§III (p selection) and §IV–V (validation)"},{"comment":"The claim that validation-set MACE is a strong predictor of extrapolation performance is supported by the synthetic experiment, but the comparison between validation and prediction-set MACE is presented without uncertainty estimates. Given that only six surrogate models are used, the average improvements (e.g., validation MACE 7.31 vs. 10.90) could be sensitive to one or two leave-one-out folds. Reporting per-fold MACE or a standard error would strengthen this inference.","section":"Table II and §IV"}],"minor_comments":[{"comment":"The text says M is 'the number of samples in the validation or test dataset,' but MACE is averaged over credible interval levels. M should be the number of p_j values used in the calibration curve.","section":"Eq. (15)"},{"comment":"The caption says 'MACE score averaged across error models,' but the table appears to average across synthetic surrogate models, not error models. Please clarify.","section":"Table II caption"},{"comment":"The caption reads 'proton and separation energies'; this should be 'proton and neutron separation energies.'","section":"Table V caption"},{"comment":"The figure caption refers to 'distance (11),' but Eq. (11) defines Δ(N,Z). Please clarify whether the plotted quantity is exactly Δ(N,Z) or a scaled version.","section":"Fig. 2"},{"comment":"The posterior values of the error-model parameters α, β_1, γ_1 are not reported. Reporting their posterior means and credible intervals would help readers judge whether the inferred distance and variance dependencies are physically reasonable.","section":"§III.A"}],"recommendation":"major_revision","confidential_remarks":"I am sympathetic to the approach and think the synthetic leave-one-out experiment is a meaningful strength. The main issues are the undefined distance predictor and the double use of the validation set for model selection and evaluation. Both are fixable within the manuscript's scope; if addressed, I would likely support publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. The paper's real contribution is letting the BMC error scale grow with two predictors—distance in PC space and inter-model variance. That's a sensible extension of the homoscedastic version in [13,14], and the synthetic leave-one-model-out tests are the best evidence: across six surrogates, both HeEM variants pull reduced chi-square toward unity and lower MACE on the validation and prediction sets. The application to Z=46–52 drip lines and the confrontation with new Cd masses and 104Te alpha decay are also well chosen and fairly discussed.\n\nThe soft spots are real but not fatal. Most important: d_i is never defined. The text calls it 'distance from the center of the training data in PC space,' but Eq. (12) defines zeta_i as a projection of the centered prediction vector, and the origin of that PC space is the ensemble-mean model, not the training-data centroid. So as written, a nucleus far from known data but with model consensus can have small d_i. If that's what the code does, the distance predictor isn't measuring extrapolation distance, and the physical motivation for the expanding error band weakens. I suspect the authors meant d_i = ||zeta_i - zeta_train_centroid||, but it's absent. They need to write it down. Second, the same validation set is used to pick p and the error-model form and then to report the final calibration metrics. That is leakage, and it makes the real-data advantage look a bit sharper than it should. The synthetic test mitigates this because the holdout model is genuinely outside the training set, but the AME2020 numbers should be treated as model selection output. Minor: no code or data release, which would help.\n\nOverall the central idea holds up. The improvement in calibration is consistent, and the drip-line probabilities are a nice demonstration. With a precise definition of d_i and a cleaner separation between model selection and evaluation, this is a solid paper. I'd send it to a competent referee and ask for those fixes.","headline":"Useful heteroscedastic extension of BMC with solid synthetic validation, but the key distance predictor is undefined and validation is used for selection—fixable before acceptance.","tokens_in":17066,"tokens_out":2519,"would_cite":true,"duration_ms":24948,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["21.10.Dr","21.60.Jz","02.50.Tt"],"model":"deepseek-v4-flash","headline":"A Bayesian ensemble method with error bars that grow with model disagreement and extrapolation distance calibrates nuclear uncertainties more honestly than constant-variance error bars, especially near the drip lines.","keywords":["heteroscedastic error model","Bayesian model combination","nuclear drip lines","uncertainty quantification","energy density functionals","separation energies","mean absolute calibration error","principal component analysis"],"falsifier":"Recompute the heteroscedastic model with d_i defined as the Euclidean distance from the training-data centroid in principal-component space (or in the original (N,Z) grid) and compare calibration metrics; if the improvement over homoscedastic error disappears or inverts, the current d_i is not measuring extrapolation distance and the physical motivation for the growing error bands collapses.","tokens_in":16177,"feed_emoji":"⚛️","tokens_out":3065,"duration_ms":33503,"temperature":0.7,"pith_summary":"The paper argues that the standard constant-variance error assumption in Bayesian Model Combination is unrealistic for nuclear predictions, because model disagreement and prediction error grow as one extrapolates toward the drip lines. It introduces heteroscedastic error scales that depend on a principal-component distance from the training data and on the variance across the model ensemble. Using synthetic data and holdout experimental data from the 2020 Atomic Mass Evaluation, the paper finds that the heteroscedastic model yields reduced chi-squared values closer to unity and lower mean absolute calibration error than the homoscedastic baseline. The practical payoff is that predicted probabilities of nuclear existence become more graded and statistically honest in regions where models genuinely diverge.","feed_headline":"Uncertainty that grows with extrapolation beats flat error bars","feed_subtitle":"A Bayesian ensemble scales error by model spread and distance from known data, sharpening drip-line existence probabilities.","key_machinery":"The central object is the heteroscedastic error scale σ_i² = α + β₁ d_i + γ₁ v_i, built from two predictors: d_i, a distance in the principal-component space of the model ensemble, and v_i, the variance of individual model predictions at each nucleus. This scale replaces the single global σ₀ of the homoscedastic model, letting the likelihood and posterior allocate more uncertainty to nuclei where models extrapolate far from the data or disagree strongly. The BMC weights are still calibrated on principal components of the centered model matrix, but now every posterior sample carries its own error scale, so reported credible intervals reflect both weight uncertainty and error-model uncertainty","core_discovery":"The paper establishes that a heteroscedastic error model, with variance σ_i² = α + β₁ d_i + γ₁ v_i where d_i is a distance in principal-component space and v_i is the inter-model variance, outperforms a constant-variance homoscedastic error model for calibrating uncertainty in extrapolated separation energies. On six synthetic benchmarks, the heteroscedastic model brings the average reduced chi-squared from 1.32 down to about 1.08, and on real data it consistently lowers the MACE relative to the homoscedastic baseline. The resulting drip-line existence probabilities inherit the growing uncertainty, producing smoother transitions from bound to unbound regions rather than artificially sharp cu","pith_inferences":["The paper never explicitly defines d_i — it is introduced as a 'distance from the center of the training data in PC space,' but the PC-space origin is the ensemble-mean model. Recomputing results with d_i defined as a genuine Euclidean distance to the training-data centroid would test whether the claimed calibration gains are driven by true extrapolation distance or by the variance term alone.","Because the homoscedastic model is nested within the heteroscedastic one, a posterior comparison of α, β₁, γ₁ could yield a formal Bayes factor for whether heteroscedasticity is statistically required for a given observable, extending beyond the paper's metric-based argument.","The same variance-function approach should transfer to any ensemble of physical models with two available diagnostics — an extrapolation distance and an inter-model spread — such as climate or materials-science ensembles, where constant-error assumptions are equally common and equally suspect.","A natural next step would be to use the heteroscedastic error model's Q_alpha predictions as calibration data to refit energy density functionals in the ¹⁰⁰Sn region, a direction the paper hints at but does not pursue."],"forward_implications":["Drip-line existence probabilities become more conservative and graded: nuclei in regions of strong model disagreement receive non-trivial existence probabilities even when the central prediction is negative.","Validation-set calibration metrics (MACE) become a reliable proxy for deep-extrapolation performance, allowing practitioners to select error models without direct access to unknown territory.","The method distinguishes regions where extrapolation is genuinely unconstrained from regions where models agree despite distance, such as the N=126 shell gap, where uncertainty growth is moderated.","New mass measurements near the limits of stability, such as the recent Cd and Te data, provide direct tests of whether the heteroscedastic bands are honestly sized; the paper shows such comparisons for ⁹⁶⁻⁹⁸Cd and ¹⁰⁴Te."],"fun_headline_variants":["Dynamic errors improve nuclear model forecasts","Growing uncertainty sharpens drip-line predictions","Bayesian ensemble adapts error to extrapolation","Heteroscedastic uncertainty beats flat error bars","Extrapolation-aware error model for nuclei"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that d_i genuinely measures how far a nucleus is from the training data in principal-component space, but the paper never explicitly defines d_i; if d_i does not track extrapolation distance, the claimed calibration improvement may be an artifact of the predictor choice rather than a physical effect.","fun_headline_variants_meta":{"raw":{"variants":["Dynamic errors improve nuclear model forecasts","Growing uncertainty sharpens drip-line predictions","Bayesian ensemble adapts error to extrapolation","Heteroscedastic uncertainty beats flat error bars","Extrapolation-aware error model for nuclei"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000147,"raw_usage":{"total_tokens":1026,"prompt_tokens":754,"completion_tokens":272,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":206}},"tokens_in":498,"tokens_out":272,"duration_ms":2909,"temperature":1.0,"reasoning_tokens":206,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T02:55:15.892299+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the heteroscedastic model with d_i defined as the Euclidean distance from the training-data centroid in principal-component space (or in the original (N,Z) grid) and compare calibration metrics; if the improvement over homoscedastic error disappears or inverts, the current d_i is not measuring extrapolation distance and the physical motivation for the growing error bands collapses.","supporting_citations":[],"review_version":1}