{"id":"0792441a-feb7-42e4-86c3-c3961108e8e0","arxiv_id":"1908.05637","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A proposed IceCube analysis method sums the likelihood ratios of many candidate neutrino flares per source and claims better discovery potential for multi-flare signals than single-flare or time-integrated stacking.","lead":"This paper presents a new statistical search method, called multiflare stacking, that looks for multiple neutrino flares from the same source by summing evidence over all candidate flares instead of only the brightest one. The authors use simulations to argue that this method can beat existing single-flare and time-integrated searches when sources flare many times at moderate strength.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claim of improved discovery potential rests on analytic published baselines, not on running the competing methods on the same pseudo-experiments; a head-to-head comparison is needed.","rationale":"The reader correctly identifies the injection-model slice as a weak point, and the paper itself is honest that the power-law parameters are descriptive and that alpha = 3 is illustrative. My concern is closely related but slightly more load-bearing: even within the chosen slice, the comparison against 'existing' methods is not a direct comparison. The single-flare baseline is an analytic binomial estimate rather than the actual single-flare time-dependent fit run on the same pseudo-experiments, and the time-integrated baseline is a published limit from a different catalog and dataset. This makes the quantitative claim of superiority vulnerable to differences in event selection, livetime, catalog composition, and trial factors. The paper's central claim would be on solid ground only if all three analyses were applied to identical simulations. That said, this is a methods paper with a plausible construction, and the authors do not overclaim beyond the illustrative role of the injection model. The appropriate verdict remains CONDITIONAL: the method is promising but needs validation on identical pseudo-experiments across the parameter space before the discovery-potential comparison can be taken at face value. My read therefore does not change the reader's verdict.","tokens_in":6011,"tokens_out":7241,"duration_ms":82779,"concrete_test":"Regenerate (or reuse) the pseudo-experiments used for Figure 4 and run the actual single-flare time-dependent search of [3] and the time-integrated stacking search of [6] on exactly the same events, source catalog, and livetime. Compute the 3-sigma discovery-potential surfaces for all three methods over a grid varying alpha in {2, 3, 4}, I0 across the plotted range, flare duration in {10, 100, 300} days, and spectral index E^-2 and E^-2.5. If the multiflare method does not outperform both baselines over a substantial fraction of this grid, the headline claim in the abstract must be restricted to the demonstrated region.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that multiflare stacking yields a significant increase in discovery potential over both the existing single-flare fit and time-integrated stacking is supported in Figure 4 by comparisons to analytic limits, not by running those methods on the same simulated data. In Section 4, the 'single flare limits' are 'calculated analytically by integrating equation 3.1' with a binomial test, and the 'time integrated limits' are taken from the published flux limits of [6], which were derived for Fermi 2LAC blazars, while the catalog under test is the top 500 Fermi 3LAC blazars. This conflates method performance with catalog, dataset, and analysis differences. In addition, the discovery potential is shown only for alpha = 3, 100-day flares, and an E^-2 spectrum; the paper itself calls this choice 'purely for the purposes of demonstration.' Because the injection power law of Eq. 3.1 is explicitly descriptive rather than predictive, the quantitative size of the claimed improvement, and even whether it exists, may depend on the unconstrained parameters. Without identical pseudo-experiments and the actual single-flare and time-integrated algorithms applied to the same events, the abstract's unqualified claim is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes \"multiflare stacking,\" an untriggered, time-dependent, source-stacking search for multiple neutrino flares from the same source. Test windows are formed from all event pairs passing an S/B threshold, each window is assigned a per-flare test statistic TS_j based on a box-shaped signal hypothesis (Eqs. 2.1–2.6), overlapping windows are removed by a greedy decorrelation procedure, and the remaining TS_j values are summed into a global test statistic ~TS (Eq. 2.7). The method is extended to source catalogs via a binomial stacking procedure that optimizes the number of sources stacked (Eqs. 2.8–2.9). The performance is studied with injected signals drawn from a descriptive power-law intensity distribution (Eq. 3.1), and 3σ discovery potentials are presented for two catalogs: the top 500 Fermi 3LAC blazars and a self-triggered catalog of the 32 highest-energy IceCube events. The central claim, stated in the abstract and Section 5, is that for signals consisting of many small flares, this method gives a significant increase in discovery potential over both the existing single-flare fit and time-integrated stacking.","tokens_in":6298,"tokens_out":4794,"duration_ms":48734,"significance":"If the headline claim were robustly established, the method would fill a genuine gap: existing untriggered searches are optimized for single dominant flares or time-integrated emission, while the TXS 0506+056 results motivate a search for repeated moderate flares. The paper is clearly written and honest about its limitations; it explicitly labels the injection power law as descriptive rather than predictive and calls the chosen parameter slice (α=3, 100-day flares, E⁻²) purely demonstrative. The algorithm is well defined, and the simulation framework is self-consistent, including a useful sanity check in Figure 3 showing that multiflare stacking recovers the injected number of signal events more accurately than the single-flare fit. However, the central quantitative claim is currently supported only by comparisons to analytic or published limits rather than by direct algorithm-to-algorithm comparisons on identical pseudo-experiments, and the exploration of the injection parameter space is very narrow. These issues make the unqualified abstract claim stronger than the evidence presented.","major_comments":[{"comment":"The central claim of a \"significant increase in discovery potential\" is not established by a head-to-head comparison. The single-flare limits are calculated analytically by integrating Eq. 3.1 and applying a binomial test, not by running the existing single-flare likelihood fit on the same simulated events. The time-integrated limits are taken from the published flux limits of Ref. [6], which were derived for the Fermi 2LAC blazar catalog, whereas the test catalog here is the top 500 Fermi 3LAC blazars. This conflates algorithm performance with differences in catalog composition, data sample, and analysis details. Please re-run the competing methods (or representative implementations thereof) on the same pseudo-experiments used for the multiflare discovery-potential curves, or explicitly restrict the claim to \"improvement relative to analytic baselines derived from Ref. [6]\" and adjust the abstract accordingly.","section":"Section 4, Figures 4a and 4b"},{"comment":"The discovery potential is demonstrated only for a single slice of the injection parameter space: α=3, flare durations fixed to 100 days, and an E⁻² spectrum. The paper itself states that this choice is \"purely for the purposes of demonstration.\" Because Eq. 3.1 is descriptive and the true neutrino-flare intensity distribution is unconstrained, the size of the claimed improvement, and even whether it exists in a given region, may depend on α and on the flare duration distribution. Please show at least a few additional slices (e.g., α=2 and α=4, and a second flare duration) or, if computation time is a barrier, state explicitly as a caveat in the abstract and conclusions that the improvement is demonstrated only for this illustrative slice.","section":"Section 4, Figure 4 and Section 3, Eq. 3.1"},{"comment":"The self-triggered catalog is constructed from the 32 highest-energy events in the same data sample used for the search, and these source events are then removed before computing test statistics. The paper does not explain how the background-only simulations used for the p-value calibration and discovery potential account for the look-elsewhere effect inherent in choosing the catalog on the same data. Unlike a fixed external catalog, the source locations here are data-dependent; if the simulations do not replicate the catalog-selection step, the reported p-values and the Figure 4b discovery potential may be over-optimistic. Please clarify the simulation procedure for this catalog or add a trials factor for the catalog selection.","section":"Section 4.2, self-triggered catalog"}],"minor_comments":[{"comment":"The text says the likelihood is \"minimized\" with respect to ns_j and γ_j, but Eq. 2.5 is a product of probabilities and the quoted TS is the standard −2 log-likelihood ratio; please clarify that the quantity minimized is the negative log-likelihood (equivalently, the likelihood is maximized).","section":"Section 2.2, Eq. 2.5"},{"comment":"The motivation for the ∆Tdata/∆Tj prefactor as a trials correction is stated only briefly. Since the final p-value calibration is performed with background Monte Carlo, this term is not incorrect, but a sentence explaining why this particular functional form is used (rather than, e.g., a factor depending on the number of test windows formed) would help the reader.","section":"Section 2.2, Eq. 2.6"},{"comment":"The caption states that \"the solid blue line corresponds to the observation of a single, 100 day flare with 13 events,\" but it is not clear how this line relates to the single-flare analytic limit or to the plotted discovery-potential curves. Please clarify the meaning of this line in the legend or caption.","section":"Section 4, caption of Figure 4"},{"comment":"The normalization constant Am is described as the \"overall power law normalization,\" but the units of N(I) are not specified; defining N(I) as the expected number of flares per interval dI would make the interpretation of Am and Io clearer.","section":"Section 3, Eq. 3.1"},{"comment":"There are a few typographical and style issues: \"gaussian\" should be \"Gaussian\" in Section 2.2; the stray period after \"gaussian\" in the same sentence should be removed; and references [4] and [8] are cited as \"these proceedings\" without page numbers, which is acceptable for ICRC but should be completed in the final version.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"This is an ICRC proceedings paper, so the depth of the numerical study is understandably limited. The method itself is interesting and clearly described, and the simulation framework is self-consistent. The main issue is that the headline claim of significant improvement over existing methods rests on comparisons to analytic or published baselines that are not directly comparable to the new method. That is fixable by either running the competitor algorithms on the same pseudo-experiments or substantially qualifying the claim. I would not reject the paper; the science is likely sound, but the quantitative support should be strengthened before the abstract's unqualified claim can stand."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Multi-flare stacking is a genuinely new construction: sum per-window test statistics over decorrelated non-overlapping flares, with a binomial stacking step for catalogs. That directly targets the repeated-flare hypothesis motivated by TXS 0506+056, and the paper shows with injections that the global TS recovers the total number of signal events much better than the single-flare fit. The method section is clear, and the simulations are self-consistent. I buy the qualitative improvement for many moderate flares.\n\nThe soft spots are real but not fatal. The quantitative claim in the abstract—\"significant increase in discovery potential\"—rests on comparisons to analytic limits and to the published time-integrated limits of [6], not to the same algorithms run on the same pseudo-experiments. That matters because [6] was derived for Fermi 2LAC, while the catalog under test is top-500 3LAC; the improvement could be partly a catalog effect. Also, the discovery potential is shown only for alpha=3, 100-day flares, E^-2, a slice the authors themselves call demonstrative. Since Eq. 3.1 is descriptive, the size of the improvement could shift elsewhere in parameter space. Those concerns don't undermine the method's logic; they just mean the headline number is not yet established.\n\nMinor: no code or data released, but that's normal for an ICRC proceedings. The self-triggered catalog is a reasonable idea and handled carefully (removing source events from the TS). The citation pattern is appropriate; [3] is the right base, [4] the right borrowing.\n\nWho is this for? Anyone working on time-dependent neutrino searches or multimessenger flares. It deserves a serious referee: the construction is worthwhile and the limitations are fixable. My recommendation: send to peer review, but ask for a head-to-head comparison on identical pseudo-experiments and a scan over more of the parameter space before accepting the quantitative claim.","headline":"A useful new test statistic for repeated neutrino flares, with a plausible qualitative advantage that the quantitative comparison doesn't yet fully pin down.","tokens_in":6761,"tokens_out":2091,"would_cite":true,"duration_ms":20360,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper constructs an untriggered, time-dependent, source-stacking search for repeated neutrino flares, and shows by simulation that summing per-flare test statistics after decorrelation boosts discovery power for many moderate flares.","keywords":["IceCube","neutrino flares","source stacking","time-dependent search","multi-flare stacking","blazars","discovery potential","untriggered search"],"falsifier":"Recalculate the 3-$\\sigma$ discovery potentials in Figure 4 with injected flare intensities drawn from $\\alpha=2$ and $\\alpha=4$, with flare durations of 10 and 300 days, and with $E^{-1.5}$ and $E^{-2.5}$ spectra; if multi-flare stacking no longer beats the single-flare and time-integrated methods across a substantial middle range of $I_o$, the paper's central claim is not robust to the unknown flare distribution.","tokens_in":5829,"feed_emoji":"🧊","tokens_out":10504,"duration_ms":82249,"temperature":0.7,"pith_summary":"This paper proposes a way to search IceCube data for astrophysical sources that emit many neutrino flares, without requiring a gamma-ray trigger. The method builds candidate flare time windows from pairs of events, assigns each window a test statistic for a flare hypothesis, discards background-like windows, removes overlapping windows, and sums the remaining per-flare statistics into a global test statistic per source. Using simulated injections drawn from a power-law distribution of flare intensities, the paper shows that for signals made of many moderate-to-small flares, this 'multi-flare stacking' has higher discovery potential than both the single-flare fit and the time-integrated stacking search. The authors also describe a binomial stacking step for source catalogs and apply the method to the top 500 Fermi 3LAC blazars and to a self-triggered catalog built from high-energy IceCube events, while noting that the injected flare-intensity power law is descriptive rather than predictive.","feed_headline":"Summing small neutrino flares sharpens IceCube discovery power","feed_subtitle":"A multi-flare stacking statistic finds repeated moderate flares better than single-flare or time-integrated searches.","key_machinery":"The load-bearing object is the global test statistic of Eq. 2.7, obtained by summing the per-window test statistics $TS_j$ after a decorrelation procedure removes windows that overlap with a higher-test-statistic window. Each per-window statistic is a likelihood ratio comparing a box-shaped flare hypothesis, with a source signal PDF built from spatial, energy, and uniform-in-window time terms, to the background, with a $\\Delta T_{\\rm data}/\\Delta T_j$ trials penalty that compensates for the many short windows. The decorrelation step turns the list of candidate windows into a 'neutrino flare curve' for a source, and the sum over the surviving windows is the discovery statistic. For catalogs, a binomial stacking step selects how many top sources to co-add, at the cost of a trial factor.","core_discovery":"The paper's central claim is that an untriggered, time-dependent, source-stacking search can be built by summing individual flare test statistics, and that this sum substantially improves discovery potential when a source produces many comparable flares. The global test statistic is $\\widetilde{TS} = \\sum_{j: TS_j>0} TS_j$ after a decorrelation step that keeps only non-overlapping windows with the largest test statistics. Over a source catalog, the per-source values are summed, and a binomial procedure chooses the number of top sources to stack. In simulations with flare intensities drawn from $N(I)=A_m(\\alpha-1)/I_o\\,(I/I_o)^{-\\alpha}$, with flare durations fixed at 100 days and $E^{-2}$ spectra, the method recovers injected signal events more accurately than the single-flare fit and reaches 3-$\\sigma$ discovery at a lower normalization $A_m$ than either existing method over a middle range of flare intensities. For very small flares the time-integrated method remains superior, and for very large flares the single-flare fit is competitive.","pith_inferences":["The same summed-statistic construction could be run with windows seeded by external gamma-ray flare catalogs instead of the neutrino S/B cut, which would test whether the multi-flare gain persists when window selection is independent of the neutrino data.","If the true flare population has a steeper or shallower intensity index than $\\alpha=3$, the region where multi-flare stacking wins will shift, so measuring $\\alpha$ from future multi-flare candidates is the key observational step.","Letting flare durations float rather than fixing them at 100 days would map how much of the gain comes from the box-window construction and how much from the assumed flare length."],"forward_implications":["A source that flares many times at moderate strength, like the behavior suggested for TXS 0506+056, becomes testable on the full IceCube sample with one coherent statistic rather than by combining separately chosen flares.","A detection with this method would directly reveal the number, timing, and fitted spectra of individual neutrino flares through the per-source flare curve.","For the Fermi 3LAC blazar catalog, multi-flare stacking improves the 3-sigma discovery potential over the existing limits in the moderate-flare region.","The self-triggered catalog of high-energy IceCube events offers a way to look for neutrino flares without relying on electromagnetic associations."],"supporting_citations":[{"why":"Motivates the multiple-flare hypothesis with the TXS 0506+056 result and defines the separate data-period approach that this method generalizes.","marker":"[1]"},{"why":"Provides the evidence for a neutrino flare from TXS 0506+056 that motivates an untriggered multi-flare search.","marker":"[2]"},{"why":"Supplies the per-window signal and background PDFs, the likelihood, and the single-flare fitting machinery that the multi-flare test statistic builds on.","marker":"[3]"},{"why":"Describes the binomial stacking and trial-factor treatment used to optimize how many catalog sources to stack.","marker":"[4]"},{"why":"Motivates the power-law model of flare intensities used for injected signals and discovery-potential calculations.","marker":"[5]"},{"why":"Provides the Fermi 3LAC catalog and the existing time-integrated limits that define the comparison baseline for the first discovery-potential plot.","marker":"[6]"},{"why":"Provides the diffuse astrophysical spectrum fit used to choose the 200 TeV threshold for the self-triggered high-energy event catalog.","marker":"[7]"},{"why":"Motivates the self-triggered catalog idea with a similar catalog of high-energy IceCube events.","marker":"[8]"}],"fun_headline_variants":["Summing flare test statistics boosts neutrino discovery","Untriggered stacked flare search finds repeated neutrino bursts","For many small neutrino flares, summed stacking wins","Stacked neutrino flare signals increase discovery potential"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the real neutrino flare population is reasonably described by the power law $N(I)=A_m(\\alpha-1)/I_o\\,(I/I_o)^{-\\alpha}$ with parameters near the illustrated $\\alpha=3$ slices, since the paper's quantitative claim of improved discovery potential is computed only for that assumed population, with flare durations fixed at 100 days and $E^{-2}$ spectra.","fun_headline_variants_meta":{"raw":{"variants":["Summing flare test statistics boosts neutrino discovery","Untriggered stacked flare search finds repeated neutrino bursts","For many small neutrino flares, summed stacking wins","Stacked neutrino flare signals increase discovery potential"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000736,"raw_usage":{"total_tokens":3297,"prompt_tokens":964,"completion_tokens":2333,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":580,"completion_tokens_details":{"reasoning_tokens":2274}},"tokens_in":580,"tokens_out":2333,"duration_ms":19744,"temperature":1.0,"reasoning_tokens":2274,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:07:45.431076+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recalculate the 3-$\\sigma$ discovery potentials in Figure 4 with injected flare intensities drawn from $\\alpha=2$ and $\\alpha=4$, with flare durations of 10 and 300 days, and with $E^{-1.5}$ and $E^{-2.5}$ spectra; if multi-flare stacking no longer beats the single-flare and time-integrated methods across a substantial middle range of $I_o$, the paper's central claim is not robust to the unknown flare distribution.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the per-window signal and background PDFs, the likelihood, and the single-flare fitting machinery that the multi-flare test statistic builds on."},{"cited_title":"Braun, M","cited_arxiv_id":null,"evidence_quote":"Describes the binomial stacking and trial-factor treatment used to optimize how many catalog sources to stack."},{"cited_title":"O'Sullivan, PoS(ICRC2019)973 (these proceedings)","cited_arxiv_id":null,"evidence_quote":"Motivates the power-law model of flare intensities used for injected signals and discovery-potential calculations."},{"cited_title":"Giomi, et al","cited_arxiv_id":null,"evidence_quote":"Provides the Fermi 3LAC catalog and the existing time-integrated limits that define the comparison baseline for the first discovery-potential plot."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the diffuse astrophysical spectrum fit used to choose the 200 TeV threshold for the self-triggered high-energy event catalog."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the self-triggered catalog idea with a similar catalog of high-energy IceCube events."}],"review_version":1}