{"id":"cd12d899-b18c-47f3-969f-9770c0117428","arxiv_id":"2502.01452","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A hierarchical Bayesian stacking framework improves sensitivity for detecting faint high-energy neutrino source populations when per-source flux predictions are uncertain.","lead":"This paper develops a Bayesian statistical method for combining many faint neutrino source candidates to detect a population, and tests it on simulations. It reports that the Bayesian approach finds weaker sources than standard maximum likelihood stacking when source brightness estimates are uncertain.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The §5 sensitivity claim relies on a Bayesian prior normalized by the true simulated signal count n̂_sig; an analyst would not know this value, so the reported 2σ advantage may be an artifact of the truth-informed prior.","rationale":"Good-faith reading: the paper is a methods demonstration; the framework is clearly explained and the PyTorch-optimized implementation is a real asset. The strongest claim is the sensitivity comparison. The reader's weakest assumption is spot on: the Bayesian detection prior uses the true total signal count as its normalization. This is the single most consequential issue because it directly affects the headline conclusion. I do not see an internal inconsistency in the likelihood derivation; the simplified example's 'uniform' prior expression (π(ξvar)=2exp(L0+2σL)) is odd and likely miswritten, but it does not drive the main result. The central concern stands, and the proposed test would settle it. Since the reader already assigned CONDITIONAL on this basis, my read leaves the verdict unchanged.","tokens_in":12395,"tokens_out":4280,"duration_ms":39900,"concrete_test":"Recompute the Section 4.4 odds ratios for σ_S = 1.0 (Fig. 2, bottom right) using a scale prior that does not involve the simulated n̂_sig: e.g., let the total signal expectation be S = Σ_j n̂*_j and marginalize over log10 S uniform on [log10(N_obs/100), log10(100·N_obs)] (or over [0, diffuse-flux upper bound]), keeping all per-source log-normal priors. Generate the same number of background and signal realizations and compare the median Bayesian p-value with the sample-normalization ML p-value. If the Bayesian p-value rises to within a factor of ~3 of the frequentist value, the claimed 2σ advantage is an artifact of the truth-informed prior; if it remains ≲10^-4, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the prior in Section 4.4: π(Σ_j n̂*_j / n̂_sig) = Uniform(0,2), with n̂_sig defined in Section 4.2 as the expected number of signal events in the simulated population. This is not a vague prior; it is an informative prior tightly centered on the truth, with mean exactly n̂_sig and width only 2× the true value. In a real search the total signal normalization is unknown a priori (at best bounded by the observed diffuse flux), so this prior leaks the simulation answer into the analysis. For the 1.0 dex scatter case the Bayesian method yields p ≈ 8.9×10^-5 whereas sample-normalization ML gives p ≈ 1.4×10^-2; the factor-50 advantage and the claimed 2σ gain are plausibly driven by the prior's knowledge of the total signal rate rather than by Bayesian marginalization. The log-normal per-source priors are centered on the modeled n̂_j, which is a legitimate use of catalog estimates, but the overall scale prior should also be derived from observables, not from the truth.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a Bayesian framework for 'deep-stacking' searches for high-energy neutrinos from large catalogs of faint candidate sources. The method combines a Poisson likelihood over reconstructed events with priors on per-source expected event counts, marginalizes over nuisance parameters, and uses a Bayes factor / odds ratio as a test statistic calibrated with background-only simulations. The authors demonstrate the approach in two settings: a simplified 1000-source example reconstructing the average source luminosity, and a more realistic simulation comparing Bayesian deep-stacking with two maximum-likelihood stacking variants under increasing scatter between the injected source fluxes and the fluxes assumed in the analysis. They report that the Bayesian method outperforms the frequentist approaches, with a factor-50 improvement in p-value and roughly 2σ added sensitivity at 1 dex scatter.","tokens_in":12635,"tokens_out":3602,"duration_ms":34939,"significance":"If the reported sensitivity advantage survives a fully fair comparison, the paper would make a useful contribution to neutrino multi-messenger searches and population inference. The hierarchical Bayesian formulation is natural for this problem, the use of calibrated test-statistic distributions to compare Bayesian and frequentist methods is commendable, and the authors are transparent about computational limitations. The main methodological ideas — incorporating source-flux uncertainties as priors and marginalizing rather than profiling — are clearly presented and potentially valuable. However, the central quantitative claim in Section 4 is currently supported by a prior that uses the true simulated signal count as its normalization anchor, and the simplified example in Section 3 contains a prior expression that is not a uniform distribution as claimed. These issues are load-bearing for the paper's main demonstrations, so the significance of the reported sensitivity gain is not yet established.","major_comments":[{"comment":"The prior on the total signal normalization is informative in a way that leaks the simulation truth into the analysis. The text states: 'A uniform prior was used for 0 ≤ Σ_j n̂*_j / n̂_sig ≤ 2', where n̂_sig is defined in Section 4.2 as the expected number of signal events from the source population in the simulated data. An analyst performing a real search would not know n̂_sig; at best the total normalization could be bounded by the observed diffuse flux, which would give a far wider and less informative prior. Because the claimed advantage at 1.0 dex scatter (p ≈ 8.9×10^-5 for Bayesian versus p ≈ 1.4×10^-2 for sample-normalization maximum likelihood in Fig. 2) occurs precisely where the total-scale prior matters most, the reported sensitivity gain may be driven by this truth-centered prior rather than by Bayesian marginalization. The authors should either re-run the comparison with a prior anchored to observables (e.g., a broad prior derived from the diffuse astrophysical neutrino flux) or demonstrate that the reported gains are robust to the choice of prior scale.","section":"Section 4.4"},{"comment":"The prior for the average luminosity is stated as 'We adopted a non-informative uniform distribution for the luminosity: π(ξvar) = 2 exp(L0 + 2σL).' This expression is not uniform in L0: it grows exponentially with L0 and depends on the parameter being estimated and on σL. As written it is not a valid probability density over L0 (nor a uniform density over log L0). This is not a minor typographical issue, because the simplified example in Section 3 is one of the paper's two demonstrations of Bayesian parameter reconstruction and of superiority over maximum likelihood. The authors should correct the prior specification, state the actual support of the prior, and re-run the Monte Carlo results in Fig. 1 to verify that the claimed unbiasedness and improved performance still hold.","section":"Section 3.2"},{"comment":"The scatter width σS used in the log-normal per-source priors is set to the same value as the scatter injected into the simulated data. The text says, 'we computed various scenarios of scattering between the injected source flux and the n̂_j used in the analysis model using the same set of widths, i.e., σS = (0, 0.3, 0.5, 1.0)'. In a real analysis σS is not known a priori; it would need to be estimated from data, marginalized over, or treated with a hyperprior. Since this parameter controls how much the likelihood can adapt to source-by-source deviations, the comparison as presented is optimistic for both the Bayesian and the source-normalization methods. The authors should show that the reported ranking of methods is robust to a mismatch between the assumed and true σS, for example by evaluating the Bayesian analysis with a moderately mis-specified σS.","section":"Section 4.4"}],"minor_comments":[{"comment":"There are several typographical errors that should be corrected: 'a called deep-stacking' in the abstract, 'frequentiest' in Section 4, 'nuissance' in Section 4.3, 'Incoporating' and 'strenght' in the Conclusion, and 'Jeffries prior' in Section 2.3 (should be Jeffreys).","section":"Throughout"},{"comment":"The sentence 'To match this challenge, the likelihood calculations for this manuscript have been optimized using the PyTorch package' appears without a preceding full stop in the submitted text; please fix the punctuation.","section":"Section 4.3"},{"comment":"In Eq. (2.6), the notation p(x|Ltot > 0, ξ) is used, but the right-hand side should be a marginal likelihood with the source luminosities integrated out; this is presumably implied by the context but should be stated explicitly for clarity.","section":"Section 2.6"}],"recommendation":"major_revision","confidential_remarks":"The core issue is the truth-informed prior in Section 4.4; this is a correctness concern that can in principle be fixed by re-running the simulation with an observable-based prior. The simplified example in Section 3.2 also needs a corrected prior specification. If the authors address these points, the paper could become acceptable; as it stands, the central sensitivity claim is not yet supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe core of this paper is a hierarchical Bayesian framework for stacking faint neutrino sources, with a pedagogical intro, a realistic simulation, and a direct comparison against two frequentist likelihood approaches. That is genuinely useful. It builds on Capel et al. (2020, 2024), and the authors say so plainly. The presentation is clear, and the use of PyTorch for the high-dimensional likelihood is sensible.\n\nThe soft spot is the one the stress-test flags, and it is load-bearing. In Section 4.4 the Bayesian prior on total signal normalization is Uniform(0, 2 n̂_sig), where n̂_sig is the expected number of signal events in the simulated population. That is an informative prior centered on the truth; it leaks the simulated answer into the analysis. In a real search you would have to set that scale from an observable like the diffuse flux, not from the injected signal count. With a 1-dex scatter, the Bayesian p-value is ~9e-5 versus ~1.4e-2 for sample-normalization ML, a factor of 50, but most of that looks like the prior, not marginalization. The authors should rerun with a scale chosen from the diffuse flux (they even mention that option in Section 2.3), and show how the result changes when the prior is deliberately off-scale.\n\nTwo more things. The simplified example in Section 3.2 writes the prior on average luminosity as π(ξvar)=2 exp(L0+2σL), which is not a uniform prior despite the text saying so. That needs correcting. And the 'source normalization' maximum-likelihood approach in Section 4.3 seems to overfit by construction—its background TS distribution is inflated relative to the null—so the comparison in that column is not as clean as the text implies.\n\nNone of this kills the framework. The paper is honest about its caveats, including the computational limits for dense catalogs. But the headline claim of 2σ added sensitivity is not currently supported. The right path is a revision that fixes the prior, re-estimates the p-values, and then reassesses the claimed gain.\n\nI'd send this to a serious referee: the methodological contribution is real, the flaw is identified clearly in the text so an editor can hand it to a statistician, and the fix is straightforward. I would not cite the current version's sensitivity numbers, but I would cite the framework once the normalization prior is fixed.\n\nBest,\n[Your name]","headline":"Clear, useful framework for Bayesian deep-stacking, but the headline sensitivity gain is inflated by a prior that uses the true signal count; the main result needs a redo with a data-derived prior.","tokens_in":13153,"tokens_out":2102,"would_cite":false,"duration_ms":19547,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Bayesian deep-stacking beats maximum likelihood on faint neutrino sources.","keywords":["Bayesian inference","deep-stacking","high-energy neutrinos","source catalog stacking","hierarchical Bayesian model","maximum likelihood comparison","source luminosity uncertainty","neutrino astronomy"],"falsifier":"Rerun the realistic simulation exactly as described but with a prior on the total expected signal that is not centered on the truth, for example uniform over a decade-wide range whose midpoint is off by a factor of three, and compare Bayesian deep-stacking with the source-normalization maximum-likelihood method at 1 dex scatter; if the Bayesian advantage falls below the claimed roughly 2σ, the truth-centered prior is the load-bearing element.","tokens_in":12173,"feed_emoji":"🔭","tokens_out":6624,"duration_ms":57276,"temperature":0.7,"pith_summary":"The paper tries to establish that deep-stacking, combining large numbers of individually undetectable neutrino source candidates into one statistical search, is best carried out with Bayesian inference because the dominant uncertainty is the poorly known mapping from a source's observable properties to its expected neutrino flux. In a realistic simulation of roughly 120 cataloged sources, it finds that when per-source flux expectations are scattered by 1 dex relative to the true luminosities, Bayesian deep-stacking keeps the median detection significance above 3σ, while the standard maximum-likelihood stack drops below it; the authors describe the gain as about 2σ of added sensitivity, equivalent to doubling the effective signal strength. The same framework reconstructs population properties more accurately than maximum likelihood in a simplified 1000-source example, because marginalizing over individual source luminosities uses all information without overfitting. If the claim holds, future catalog-driven neutrino searches would gain sensitivity by adopting Bayesian odds-ratio tests rather than frequentist likelihood ratios.","feed_headline":"Bayesian deep-stacking doubles sensitivity to faint neutrino sources","feed_subtitle":"When per-source flux predictions are uncertain, the Bayesian test keeps detection above 3σ where the standard search fails.","key_machinery":"The machinery is the hierarchical Bayesian odds-ratio test built on the Poisson product likelihood $p(x|\\xi)\\propto e^{-\\hat n_{\\rm tot}}\\prod_i [\\sum_j \\hat n_j S_j(x_i) + \\hat n_{\\rm bg} B(x_i)]$, where $\\hat n_j$ is the expected event count of catalog source $j$, $S_j$ is the signal spatial probability, and $B$ is the background probability. Deep-stacking is the strategy of including all such sources, even very distant and faint ones, so that the search targets population detection rather than individual detections. The key move is to put a log-normal prior on each $\\hat n_j$ with width $\\sigma_S$ characterizing flux-model scatter, marginalize over the $\\hat n_j$ inside the odds ratio, and then reinterpret that odds ratio as a frequentist test statistic by calibrating it on background-only simulations. A practical simplification makes the marginalization tractable: when the average source spacing is much larger than the detector angular resolution, each event contributes only to the nearest source's signal term, reducing the computation to a series of one-dimensional integrals. The prior on the total signal scale, taken uniform over 0 to 2 times the true expected signal, is what carries the sensitivity claim.","core_discovery":"The central claim is that a hierarchical Bayesian treatment of source-stacking searches for high-energy neutrinos is more sensitive and more accurate than the standard frequentist likelihood-ratio approach whenever individual source flux expectations are uncertain. The paper writes the unbinned Poisson likelihood for a catalog search and treats each source's expected event count as a nuisance parameter with a log-normal prior whose width encodes the uncertainty of the flux model. It then forms the odds ratio between the signal and background hypotheses, marginalizes over the nuisance parameters, and uses the odds ratio itself as a test statistic calibrated on background-only pseudo-data. In the realistic simulation, the Bayesian method outperforms both a standard sample-normalization maximum-likelihood search and a source-normalization search with per-source nuisance parameters, and it is the only method that reaches average significance above 3σ for 1 dex flux scatter. The paper also shows, in a simplified scenario with 1000 cataloged sources, that the Bayesian posterior on the average source luminosity is less biased and has smaller error bars than maximum-likelihood estimates.","pith_inferences":["A natural extension, not explored in the paper, is to apply the same odds-ratio construction to other astrophysical transient searches or to gravitational-wave source catalogs, where per-source distance and mass uncertainties play a role similar to neutrino flux-model uncertainty.","A practical prescription suggested but not tested is to set the total-signal prior from an independent measurement such as the diffuse neutrino flux, which would remove the truth-centered-prior dependence; testing this would show how much of the reported gain survives in a real search.","Treating the spectral index as a hyperparameter with its own prior, rather than fixing it to the injected value, is a natural extension that could change the energy-bin weights and hence the sensitivity comparison."],"forward_implications":["Catalog-based neutrino stacking searches should prefer Bayesian odds-ratio tests over sample-normalization maximum likelihood when per-source flux predictions are unreliable, because the standard method's significance degrades as flux scatter grows.","A real search can report posterior credible intervals on population parameters such as average luminosity, giving physical interpretation of the stacked population rather than only a detection p-value.","The gain is expected to grow with the number of sources in the catalog, since marginalization avoids overfitting the many per-source nuisance parameters that maximum-likelihood searches must fit.","If source density becomes high enough that point-spread functions overlap, the nearest-source approximation fails and full multi-dimensional marginalization would be required, which the paper identifies as the current computational bottleneck."],"supporting_citations":[{"why":"Supplies the deep-stacking motivation: adding sources out to redshift about 0.3 greatly increases search sensitivity, defining the regime the paper targets.","marker":"[11]"},{"why":"Provides the earlier Bayesian hierarchical model constraining neutrino source populations from reconstructed fluxes, which the present work extends to event-level data.","marker":"[16]"},{"why":"Gives the hierarchical Bayesian point-source likelihood that the paper's Eq. 2.3 closely follows, serving as the methodological starting point.","marker":"[17]"},{"why":"One of the standard stacking searches whose likelihood framework the realistic simulation implements as the frequentist baseline.","marker":"[18]"},{"why":"Another standard stacking search used as a template for the unbinned maximum-likelihood ratio test.","marker":"[19]"},{"why":"Supplies the numerical quadrature routine used to carry out the one-dimensional marginalization integrals in the odds ratio.","marker":"[20]"},{"why":"Provides the observed diffuse neutrino spectral index used to set the simulated source spectrum.","marker":"[21]"},{"why":"Normalizes the simulated background to observed track event counts, fixing the signal-to-background ratio of the test.","marker":"[24]"}],"fun_headline_variants":["Bayesian deep-stacking beats standard neutrino searches","Bayesian stacking finds faint neutrino sources at 3σ","Uncertain flux? Bayesian neutrino stacking wins","Bayesian method sharpens neutrino source detection","Deep-stacking with Bayesian stats outperforms likelihood"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The sensitivity demonstration assumes the analysis uses a prior on the total expected signal that is centered on the true simulated value (uniform from zero to twice that value), information a real search would not have; if that prior is misspecified, the reported 2σ advantage may shrink.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian deep-stacking beats standard neutrino searches","Bayesian stacking finds faint neutrino sources at 3σ","Uncertain flux? Bayesian neutrino stacking wins","Bayesian method sharpens neutrino source detection","Deep-stacking with Bayesian stats outperforms likelihood"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1281,"prompt_tokens":891,"completion_tokens":390,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":507,"completion_tokens_details":{"reasoning_tokens":319}},"tokens_in":507,"tokens_out":390,"duration_ms":4544,"temperature":1.0,"reasoning_tokens":319,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T15:14:42.267193+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the realistic simulation exactly as described but with a prior on the total expected signal that is not centered on the truth, for example uniform over a decade-wide range whose midpoint is off by a factor of three, and compare Bayesian deep-stacking with the source-normalization maximum-likelihood method at 1 dex scatter; if the Bayesian advantage falls below the claimed roughly 2σ, the truth-centered prior is the load-bearing element.","supporting_citations":[{"cited_title":"A hierarchical Bayesian approach to point source analysis in high-energy neutrino telescopes","cited_arxiv_id":"2406.14268","evidence_quote":"Gives the hierarchical Bayesian point-source likelihood that the paper's Eq. 2.3 closely follows, serving as the methodological starting point."}],"review_version":1}