{"id":"899fddc8-5c06-4060-9440-8baa0548de86","arxiv_id":"2512.17744","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A simulation-based neural-likelihood analysis of Planck 2018 and DESI DR2 data reports a weak preference (tilde_Delta = 0.12, 68% CL interval spanning both signs) for the normal neutrino mass hierarchy.","lead":"This paper uses a machine-learning inference pipeline to combine Planck CMB spectra with DESI galaxy-distance data and asks whether the three neutrino masses are ordered 'normally' or 'inverted'. The result mildly favors the normal order, but the uncertainty still permits the inverted order.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Forward model's fixed-theta0 covariance (Eq. 10) and missing low-ell/lensing may bias tilde_Delta; the paper's own As/tau bias suggests misspecification, so the reported preference for the normal hierarchy is not yet robust.","rationale":"The paper is a straightforward, well-scoped application of established SBI methods to a real question. The new reparameterization tilde_Delta is physically motivated and the mapping to total neutrino mass is correctly implemented. The authors perform useful validation (rank statistics, P-P plot, true-vs-predicted) that demonstrates the neural likelihood is internally consistent within the simulator. However, these tests cannot validate the forward model itself. The paper's own observation that A_s and tau are biased relative to Planck is direct evidence that the forward model does not faithfully represent the data-generating process; this is the weakest link in the chain from simulated data to the reported tilde_Delta. The high-ell CMB power spectra contain lensing information, and the lensing amplitude depends on the neutrino mass fraction; if the covariance or the missing low-ell data distort the amplitude calibration, the hierarchy parameter can be biased. The abstract/body inconsistency (6x10000 vs 2x5000) further complicates reproduction but is secondary. The conditional verdict is appropriate: the methodological template is sound and worth publishing as a demonstration, but the specific numerical preference for the normal hierarchy should be treated as provisional until the fixed-theta0 approximation and low-ell/lensing omissions are shown not to shift the result. I therefore recommend no change to the reader's CONDITIONAL verdict.","tokens_in":8893,"tokens_out":8094,"duration_ms":86529,"concrete_test":"Rerun the SNLE inference (same 2-round/5000-simulation setup) with the covariance in Eq. (10) evaluated at each sampled theta (or at least at a representative massive-neutrino cosmology, e.g., theta with sum m_nu=0.06 eV and A_s/tau at the values inferred in Fig. 6) instead of at the fixed Planck best-fit theta0. Compare the resulting marginal posterior of tilde_Delta to the published one. If the central value or the 68% interval shifts by more than ~0.1 (about one-third of the quoted interval width), the fixed-theta0 approximation is a significant source of systematic error and the normal-hierarchy preference is not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the assumption that the approximate Planck simulator in Section II (Gaussian white noise plus beam, cosmic-variance covariance evaluated at a fixed theta0 via Eq. 10, no low-ell Commander/SimAll and no lensing likelihood) yields an unbiased likelihood for the actual Planck and DESI data. This is the load-bearing assumption, and it is demonstrably imperfect: Section III states that the inferred ln(10^10 A_s) and tau are 'a little larger' than Planck 2018. Since A_s, tau, and the lensing-induced smoothing of high-ell C_ell are degenerate with the neutrino mass hierarchy parameter tilde_Delta, the same misspecification could shift the tilde_Delta posterior. The rank-statistic and P-P validations in Figs. 3-5 only check self-consistency of the learned posterior under the same simulator; they cannot detect a systematic offset between the simulator and the real data-generation process. Without a calibration or robustness check of tilde_Delta itself, the reported value 0.12216^{+0.26193}_{-0.29243} is not a secure measurement of the hierarchy preference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies the LtU-ILI pipeline with Sequential Neural Likelihood Estimation (SNLE) to infer the neutrino mass hierarchy parameter \\tilde{\\Delta} from Planck 2018 TT/TE/EE power spectra and DESI DR2 BAO distance ratios. The forward model embeds CLASS in a simulator that adds Planck-like noise, a Gaussian beam, and cosmic-variance noise with the covariance evaluated at a fixed fiducial cosmology (Eq. 10). After sequential training and validation with rank statistics, P-P plots, and true-vs-predicted comparisons, the posterior for the observed data yields \\tilde{\\Delta}=0.12216^{+0.26193}_{-0.29243} (68% CL), which the authors interpret as a slight preference for the normal hierarchy.","tokens_in":9279,"tokens_out":4756,"duration_ms":52533,"significance":"If robust, this would be a useful demonstration of simulation-based inference for a cosmological parameter that is difficult to sample with explicit likelihoods, and it would add a mild, independent data preference for the normal neutrino mass ordering. The paper uses a standard, widely used ILI pipeline and includes appropriate internal validation checks (rank statistics, P-P coverage) that are often missing in applications of SBI. The approximate Planck simulator is clearly described, and the authors acknowledge its limitations. However, the central numerical claim is not yet supported to the standard required for a journal publication: the abstract and main text disagree on training configuration and on the central value/uncertainty, and the approximate forward model is shown to bias at least some parameters, with no calibration of the hierarchy parameter itself.","major_comments":[{"comment":"The abstract states that the analysis uses '6 rounds of 10000 simulations' and finds \\tilde{\\Delta}=0.12^{+0.21}_{-0.23} (68% CL), while Section III describes '2 rounds of training' with '5000 simulated data-parameter pairs' each and reports \\tilde{\\Delta}=0.12216^{+0.26193}_{-0.29243} (68% CL). These are materially different training configurations and different posterior summaries. The reader cannot tell which result is the one being claimed. This is a load-bearing reproducibility issue; please correct the inconsistency and report the actual configuration and result.","section":"Abstract vs. Section III"},{"comment":"The covariance of the simulated high-\\ell spectra is evaluated at a fixed Planck best-fit \\theta_0 that does not include the neutrino hierarchy parameter, while the mean spectrum is evaluated at the sampled \\theta. This discards the parameter dependence of cosmic variance, including the effect of massive neutrinos on C_\\ell. The simulator also omits the low-\\ell Commander/SimAll likelihoods and the CMB lensing likelihood. The authors themselves note (Section III) that the inferred \\ln(10^{10}A_s) and \\tau are 'a little larger' than in Planck 2018. Since A_s, \\tau, and lensing are degenerate with neutrino mass and hence with \\tilde{\\Delta}, this misspecification could bias the reported \\tilde{\\Delta} posterior. The rank-statistic and P-P tests in Figs. 3-5 only check self-consistency of the learned posterior under the same simulator; they do not validate the simulator's fidelity. Please a","section":"Section II, Eq. (10) and Section III"},{"comment":"The central claim is a 'slight preference' for \\tilde{\\Delta}>0, but the reported 68% interval, [-0.29243, +0.26193] around 0.12216, includes both positive and negative values. The posterior probability P(\\tilde{\\Delta}>0 | x_o) is not stated. A credible interval that crosses zero does not by itself establish a preference; the preference claim should be quantified by the posterior probability of the normal-hierarchy side, or by a Bayes factor with a specified prior. Please report this quantity and state whether it changes under the two training configurations.","section":"Section III, Fig. 6"}],"minor_comments":[{"comment":"The covariance-matrix expression in Eq. (10) is difficult to parse because of the layout and missing parentheses. Please rewrite it in explicit symmetric-matrix form with each element clearly labeled (e.g., Cov(TT,EE) = ...).","section":"Section II, Eq. (10)"},{"comment":"The parameter list uses '100\\theta_s' but the text refers to '100\\theta_s' and Figure 2 caption mentions '\\theta_0 does not include the neutrino mass hierarchy parameter'. Please make the parameter naming consistent and specify the prior range for \\tilde{\\Delta} in the main text (currently only in Eq. (5)).","section":"Section II, Eq. (5)"},{"comment":"The observed data label is given both as 'x_o' and 'x_0'. Please use one symbol consistently.","section":"Section III"},{"comment":"The validation figures are described qualitatively ('almost the case', 'good agreement'). Please add quantitative summary statistics, such as p-values for the uniformity of the rank statistics or the maximum deviation in the P-P plot, especially for \\tilde{\\Delta}.","section":"Section III, Fig. 3-5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's main scientific claim is a weak preference for the normal hierarchy. The abstract/full-text inconsistency and the uncalibrated approximate simulator are the main blockers. The paper would be much stronger if the author reruns with a single stated configuration and provides a coverage or bias calibration for \\tilde{\\Delta} against a more complete data model. The scope is appropriate for the journal, and the ILI demonstration is useful, but the present version is not yet ready for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a legitimate but routine application of an existing SBI pipeline (LtU-ILI/SNLE) to the neutrino mass hierarchy question, with one genuinely new piece: the reparameterization tilde_Delta = -sign(Delta) lg(|Delta|), which turns the U-shaped posterior that an upper limit on sum m_nu implies into a T-shaped one that neural density estimators handle better. That is a clever idea, and the paper shows it works on Planck+DESI data. The validation tests (rank statistics, P-P plots, true-versus-predicted) are the right checks for self-consistency, and the posterior on the base parameters is broadly consistent with Planck. The authors are also honest that the constraints would improve with lensing.\n\nThe main problems are two. First, the abstract says 6 rounds of 10000 simulations and quotes tilde_Delta = 0.12^{+0.21}_{-0.23}, while the body says 2 rounds of 5000 and quotes 0.12216^{+0.26193}_{-0.29243}. That inconsistency needs to be caught before publication. Second, the Planck forward model is simplified: the high-ell cosmic-variance covariance is evaluated at a fixed Planck best-fit theta0, and there is no low-ell or lensing likelihood. The authors themselves note that ln(10^10 A_s) and tau come out a little higher than Planck. Since A_s, tau, and lensing smoothing are all degenerate with neutrino mass, that misspecification could shift tilde_Delta. The rank-statistic and P-P validations only check the simulator against itself, not against the real data. So the numerical value 0.12216 is not a secure measurement; the qualitative weak preference for normal hierarchy is plausible and consistent with earlier work, but it is not demonstrated robustly.\n\nI also think the stress-test note is fair here. The fixed-theta0 covariance is a real approximation and the paper's own As/tau bias is evidence the simulator is imperfect. The authors don't offer a calibration for tilde_Delta.\n\nWho gets value: readers interested in SBI methodology for CMB, and people tracking the neutrino hierarchy question. It's a methods demonstration, not a new cosmological constraint.\n\nRecommendation: send it to a referee, but only after the authors fix the abstract/body discrepancy and add a robustness check on tilde_Delta against the forward-model choices. The core idea is sound enough to deserve referee time. I would not cite it as a measurement, though.","headline":"Reasonable SBI demonstration with a useful new parameterization, but internal inconsistencies and an imperfect forward model keep the reported hierarchy preference from being secure.","tokens_in":9708,"tokens_out":2829,"would_cite":false,"duration_ms":28201,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A simulation-based analysis of Planck and DESI data finds a slight preference for the normal neutrino mass hierarchy.","keywords":["neutrino mass hierarchy","simulation-based inference","implicit likelihood inference","sequential neural likelihood estimation","CMB power spectra","DESI DR2 BAO","cosmological parameters","Planck 2018"],"falsifier":"Recompute the Δ̃ posterior using the official Planck 2018 likelihoods (low-ℓ Commander/SimAll, Plik, and lensing) together with DESI DR2 distance ratios, within the same simulation-based framework or with an explicit likelihood. If the resulting posterior is centered at or near zero, or if it excludes positive values at high probability, the mild preference for the normal hierarchy reported here would be an artifact of the simplified forward model. Alternatively, a direct comparison of the simulated spectra to the real Planck spectra at the inferred best fit could reveal whether the covariance","tokens_in":8791,"feed_emoji":"🌌","tokens_out":8619,"duration_ms":84421,"temperature":0.7,"pith_summary":"This paper asks whether current cosmological data can tell the normal and inverted neutrino mass orderings apart, and answers with a qualified yes. Instead of writing down an explicit likelihood, the authors fit a neural-network model of the likelihood to simulations from a forward model of CMB spectra and BAO distance ratios, and they estimate a reparametrized hierarchy parameter Δ̃ = −sign(Δ)·log₁₀|Δ|. The resulting 68% interval, Δ̃ = 0.12216^{+0.26193}_{−0.29243}, sits mostly above zero, which favors the normal hierarchy (Δ > 0), but the interval easily contains zero and negative values, so the preference is not decisive. If the result holds up to more careful checks, it demonstrates that simulation-based inference can recover neutrino-mass information from current data without an explicit likelihood, and it gives a practical parametrization for future analyses.","feed_headline":"Cosmic data mildly favor normal neutrino mass hierarchy","feed_subtitle":"A simulation-based inference with a reparametrized hierarchy parameter yields a positive but not decisive signal from Planck and DESI data.","key_machinery":"The central object is the reparametrized hierarchy parameter Δ̃ = −sign(Δ)·log₁₀|Δ|, with Δ = (m₃−m₁)/(m₁+m₃). Because an upper limit on the total neutrino mass ∑mᵥ produces a U-shaped posterior in Δ, the paper maps it to a T-shaped parameter that is easier for a neural density estimator to capture near the origin. The inference engine is sequential neural likelihood estimation (SNLE): a masked autoregressive flow learns an approximate likelihood q_w(x|θ) from simulated data–parameter pairs, and the posterior for the observed data is then sampled by MCMC. The forward model combines a Boltzmann solver for CMB spectra with Planck-like white noise and a Gaussian beam, cosmic-variance realizatio","core_discovery":"The paper claims the neutrino mass hierarchy can be constrained by simulation-based inference from Planck 2018 CMB spectra and DESI DR2 BAO distance ratios, using a forward model that simulates these observables from cosmological parameters. The key step replaces Δ = (m3−m1)/(m1+m3) with Δ̃ = −sign(Δ)·lg|Δ|, turning the U-shaped total-mass posterior into a T-shaped target that a neural density estimator learns more easily. After sequential training, the posterior gives Δ̃ = 0.12216^{+0.26193}_{−0.29243} (68% CL), a positive central value that slightly favors the normal hierarchy (m1 < m3). Validation on a held-out test set shows the learned posterior is globally consistent.","pith_inferences":["The paper's own validation notes that ln(10^10 A_s) and τ are biased relative to Planck 2018; if a similar small systematic affects Δ̃, the positive central value could be a mild artifact, so the normal-hierarchy preference should be checked with the full official likelihoods.","The abstract states 6 rounds of 10,000 simulations while the main text describes 2 rounds of 5,000; this inconsistency should be clarified because the simulation budget affects the fidelity of the learned likelihood.","The same reparametrization could be applied directly inside explicit likelihood analyses (e.g., nested sampling) to make the hierarchy posterior easier to sample and visualize.","If future data keep pushing the total neutrino mass toward lower values, Δ̃ is expected to concentrate near zero and the sign may become decisive one way or the other; this parametrization offers a direct way to track that trend."],"forward_implications":["The reported positive Δ̃ gives a non-decisive but real statistical preference for the normal neutrino mass ordering, aligning with the direction favored by terrestrial oscillation data.","The Δ̃ parametrization converts the two-mode structure of the mass-sum posterior into a near-Gaussian target, making it useful for any future neutrino-mass analysis, likelihood-free or not.","Because the learned posterior is amortized, new observations can be incorporated by re-weighting or fine-tuning without re-simulating millions of spectra.","Extending the same pipeline to CMB lensing, bispectra, or galaxy clustering should shrink the uncertainty on Δ̃ and could make the hierarchy preference decisive.","The method reproduces Planck's constraints on the base ΛCDM parameters, supporting simulation-based inference as a viable alternative to explicit likelihoods for cosmological parameter estimation."],"fun_headline_variants":["Simulation-based inference: normal neutrino hierarchy slightly ahead","Cosmic data lean toward normal neutrino ordering","Neural likelihood mildly boosts normal neutrino hierarchy","Planck and DESI data hint at normal neutrino mass hierarchy"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing assumption is that the approximate Planck forward model used here — Gaussian white noise plus a Gaussian beam, with the high-multipole cosmic-variance covariance evaluated at a fixed Planck best-fit cosmology and with no low-multipole or lensing likelihoods — is accurate enough that its imperfections do not bias the neutrino-hierarchy parameter beyond the quoted uncertainty. The paper's own finding of slightly biased A_s and τ values shows this approximation","fun_headline_variants_meta":{"raw":{"variants":["Simulation-based inference: normal neutrino hierarchy slightly ahead","Cosmic data lean toward normal neutrino ordering","Neural likelihood mildly boosts normal neutrino hierarchy","Planck and DESI data hint at normal neutrino mass hierarchy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000254,"raw_usage":{"total_tokens":1383,"prompt_tokens":702,"completion_tokens":681,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":446,"completion_tokens_details":{"reasoning_tokens":620}},"tokens_in":446,"tokens_out":681,"duration_ms":8091,"temperature":1.0,"reasoning_tokens":620,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T15:10:49.952144+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the Δ̃ posterior using the official Planck 2018 likelihoods (low-ℓ Commander/SimAll, Plik, and lensing) together with DESI DR2 distance ratios, within the same simulation-based framework or with an explicit likelihood. If the resulting posterior is centered at or near zero, or if it excludes positive values at high probability, the mild preference for the normal hierarchy reported here would be an artifact of the simplified forward model. Alternatively, a direct comparison of the simulated spectra to the real Planck spectra at the inferred best fit could reveal whether the covariance","supporting_citations":[],"review_version":1}