{"id":"931de563-60c3-49e3-a6b1-83bd2490206b","arxiv_id":"2608.11862","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"SPORE samples realistic reconstructed neutrino events for point and extended sources from tabulated instrument response functions, with native multi-detector support.","lead":"SPORE is an open-source Python package that turns tables of detector response into simulated neutrino events, letting researchers test how several telescopes would observe the same source together. It was checked against public IceCube data and matches the observed sky distribution well, with a low-energy mismatch the authors trace to the coarse public response tables.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Low-energy overprediction is attributed to the public smearing matrix, but the forward-fold control does not rule out SPORE's default effective-area smoothing/trimming (App. B); a controlled re-run is needed.","rationale":"The paper is transparent and the package appears well designed: the HDF5 IRF format is clearly specified, the sampling algorithms are described in enough detail to be checked, and the repository is said to include figure-generation scripts. The HESE and psi2 validations are reasonable sanity checks, and the paper correctly notes that the HESE comparison is not an independent test of the flux model. The strongest support for the central claim is the round-trip consistency test, and the most load-bearing point in that test is the attribution of the sub-600 GeV energy-proxy overprediction. The paper offers a forward fold as evidence that the discrepancy is intrinsic to the released smearing matrix rather than the sampler; that is real support, and it plausibly rules out a bug in the event-sampling loop. However, the forward fold as described does not isolate the smearing matrix from the default effective-area smoothing and trimming that App. B applies before sampling. If those defaults are responsible, the discrepancy would appear even with a fine-grained smearing matrix, and the paper's explanation would need revision even though the package itself might still be usable with different settings. This is an empirical attribution question, not an internal inconsistency or a derivation error, so it does not justify rejection. The reader's conditional verdict already captures this uncertainty; my concern narrows it to the effective-area preprocessing rather than the sampler, but it does not change the recommended outcome. A single controlled re-run with preprocessing disabled would settle whether the concern lands.","tokens_in":15920,"tokens_out":3901,"duration_ms":42514,"concrete_test":"Re-run the Sec 6.2 northern-track energy-proxy comparison and the direct forward fold with the same released smearing matrix but with effective-area preprocessing disabled (smoothing_sigma=0, trim_isolated=False), using the repository's figure-generation scripts. If the sub-600 GeV overprediction remains at the reported factor ~2-3, the attribution to the smearing matrix is supported; if the ratio drops toward unity, the App. B defaults are the cause and the paper's explanation and default settings must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Sec 6.2 energy round-trip claims the sub-600 GeV overprediction is caused by the coarse true-energy binning of the released smearing matrix, supported by the statement that a direct forward fold 'with no event sampling at all' reproduces a comparable factor. But the forward fold is not, as reported, a control on the effective-area processing: App. B applies default trim_isolated and Gaussian smoothing in (ln E, ln Aeff) with sigma=1.5 energy bins before any sampling, and the text does not say the forward fold used an unprocessed effective area. If the smoothing/trimming broadens the low-energy acceptance, the same forward fold would overpredict even with a fine-grained smearing matrix. The validation therefore does not distinguish between 'IRF binning is the cause' and 'default Aeff preprocessing is the cause'; the latter would be a property of SPORE's defaults, not of the public IRFs, and would weaken the claim that any published response table can be plugged in without code modification. The paper's own Sec 7 calls the coarse smearing matrix the dominant systematic, but that is an assertion, not a demonstrated isolation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents SPORE, an open-source Python package that samples individual neutrino events from tabulated instrument response functions (IRFs) in a standardized HDF5 format. The package supports point and extended sources, multi-detector joint simulations, and provides reconstructed directions, energies, times, and morphologies. The authors validate SPORE with three tests: a comparison to the IceCube HESE 7.5-year public data release, a round-trip consistency test against the IceCube 10-year point-source data release, and an angular-distribution test for a point-source analysis near NGC-1068. The central claim is that any neutrino telescope with published response tables can be plugged into SPORE without code modification, and that the sampler reproduces observed data to the stated accuracy. The paper is accompanied by public code, IRF files, and scripts that reproduce all figures.","tokens_in":16205,"tokens_out":4274,"duration_ms":45385,"significance":"If the central claim holds, SPORE fills a genuine gap in the neutrino simulation ecosystem: it produces final-level reconstructed events directly from public IRFs, complementing Prometheus's detector-level simulation and toise's aggregate sensitivity calculations. The inverse-CDF sampling approach is fast and avoids MCMC burn-in, and the joint smearing-matrix support is a useful feature for realistic event generation. The open-source release with IRF files and figure scripts is a notable strength. However, the validation is largely round-trip in nature, and the energy-proxy comparison shows a factor-of-three excess at low energies whose attribution to the public IRFs rather than to SPORE's own preprocessing is not fully demonstrated. The quantitative validation claims therefore need strengthening before the plug-and-play assertion is fully supported.","major_comments":[{"comment":"The claim that the sub-600 GeV overprediction is 'traced to the coarse true-energy binning of the public smearing matrix rather than to the sampling' (Sec. 6.2) is not established by the forward-fold control as described, because the forward fold uses the same default effective-area preprocessing (trim_isolated and Gaussian smoothing with smoothing_sigma=1.5 energy bins, App. B) that is applied before any sampling. If that preprocessing broadens the low-energy acceptance, the forward fold would overpredict regardless of the smearing-matrix binning, and the discrepancy would be a property of SPORE's defaults rather than of the public IRFs. Please re-run the forward fold with smoothing_sigma=0 and trim_isolated disabled, or otherwise demonstrate that the excess persists when the effective-area table is unmodified. Without this, the abstract's statement that the discrepancy is 'traced to the coarse true-energy binning' is not supported.","section":"§6.2, App. B"},{"comment":"The ratio panels in Fig. 5 show no uncertainty bands, so the quantitative agreement claims ('10–15% level' in declination, 'within about a third' in energy proxy) cannot be assessed against statistical or systematic scatter. Given that the atmospheric normalization is fit to the same data and only 40 pseudo-experiments are averaged, adding Poisson/ensemble bands or per-bin error bars to the Obs./Pred. panels is necessary to support these claims.","section":"§6.2, Fig. 5"},{"comment":"The HESE goodness-of-fit test is circular in the sense stated in the text: the same best-fit Monte Carlo defines both the reference expectation and the test-statistic distribution, so the agreement primarily demonstrates internal consistency of the sampler. The paper acknowledges this, but the sentence in Sec. 6.1 that the test 'indicates that spore reproduces both the mean expectation and the size of the Poisson scatter' overstates what can be concluded. I recommend either rephrasing to emphasize the consistency-check nature, or adding an independent validation against an analytic expectation for a simple case.","section":"§6.1, Fig. 4"}],"minor_comments":[{"comment":"The definition PSF(ψ|E) = d/dψ P(ψ'≤ψ|E) is a probability density, not a cumulative distribution; the text immediately says the PSF is stored as an inverse CDF. Please clarify the notation to avoid ambiguity between the density and the stored quantile function.","section":"§3.2, Eq. (3.3)"},{"comment":"The declination labels in Fig. 2 appear as boxes (δ = □45◦, δ = □90◦), likely a rendering or encoding issue. Please ensure the minus signs are displayed correctly.","section":"Fig. 2"},{"comment":"The sentence about the spurious loader warning when the smearing subgroup is present is confusing; consider suppressing the warning automatically when the joint smearing subgroup is detected, or rewording the sentence to state clearly that the warning is expected and harmless in that case.","section":"App. B"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern is valid: the forward-fold control in Sec. 6.2 does not isolate the smearing-matrix binning from SPORE's default effective-area preprocessing, and the paper's own Sec. 7 labels the coarse smearing matrix as the dominant systematic without demonstrating that the defaults are not responsible. This is fixable with a controlled re-run. The validation methodology is otherwise sound for a round-trip consistency test, but the 'plug-and-play' claim would be strengthened by an example with a non-IceCube IRF or by explicit tests of the default preprocessing choices."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take on SPORE (2608.11862). This is a genuinely useful piece of infrastructure. What's new isn't the inverse-CDF sampling—that's standard—but the combination of a standardized HDF5 IRF schema, an extended-source sampler that handles full RA/dec-dependent templates, and native multi-detector joint simulation. If it works as advertised, it fills a real gap between toise's aggregate sensitivity curves and Prometheus's full detector simulation. The authors wrote up the round-trip validations honestly, including the sub-600 GeV overprediction, which they attribute to the coarse public smearing matrix.\n\nWhat it does well: the software design is sensible, the paper ships code and figure-generation scripts, and the HESE test includes the right caveat that the comparison is against a best-fit expectation, not an independent flux test. The Sec. 7 limitations list is genuinely candid: they state the smoothing/trimming is applied by default, that grid error is linear in cell size, that morphology labels are approximate, and that some detector fields are recorded but unused. That transparency earns credit.\n\nThe soft spots, in proportion. The low-energy discrepancy: the stress-test note is right that the forward-fold control does not fully rule out the default effective-area smoothing/trimming as the cause. The paper says the forward fold reproduces the factor, but doesn't state whether that fold used the processed or unprocessed effective area. So the attribution to 'intrinsic to the released IRFs' is asserted, not demonstrated. This is a concrete, fixable ambiguity: run the forward fold with smoothing disabled and with a fine-grained smearing matrix, and report both. It doesn't invalidate the package, but it does mean the 'energy proxy is a prediction' claim is softer than it looks.\n\nSecond, the declination validation has a fitted atmospheric normalization, no uncertainty bands on the ratio plots, and a residual tilt the authors say they haven't identified. Minor, but it should be acknowledged more fully in the text. Third, reproducibility is not fully pinned: no commit hash or version tag in the paper; the repo exists, but a referee should ask for the exact release.\n\nOverall, the central argument holds up. The sampler inverts the IRFs correctly as far as the round-trip shows, and the multi-detector design is real. The paper deserves peer review—a serious referee can ask for the control re-run and get it. I'd take it to a reading group.","headline":"A useful and honest infrastructure paper; the low-energy validation gap is an unambiguous ambiguity in the control, not a fatal flaw, and it deserves refereeing.","tokens_in":16678,"tokens_out":2933,"would_cite":true,"duration_ms":27200,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"spore reads tabulated telescope responses and outputs simulated neutrino events, reproducing a public ten-year track sample's declination distribution to 10-15%.","keywords":["neutrino astronomy","instrument response functions","event-level Monte Carlo","inverse-CDF sampling","multi-detector joint analysis","HDF5","extended source simulation","point source analysis"],"falsifier":"Take the same detector response but replace the smearing matrix with a fine-grained version, e.g. true-energy bins of width 0.1 in $\\log_{10}(E/\\mathrm{GeV})$ in ten declination bands, and re-run the energy round-trip: if the sampled energies still over-predict the observed rate below 600 GeV by about a factor of three, then the paper's attribution of the discrepancy to the released binning would be wrong and the sampler or the preprocessing would be implicated.","tokens_in":15743,"feed_emoji":"🔭","tokens_out":9131,"duration_ms":84858,"temperature":0.7,"pith_summary":"spore is a new open-source Python package that generates event-level Monte Carlo samples of neutrinos from point and extended sources using only tabulated instrument response functions (effective area, point spread function, energy resolution). The package encodes these responses in a selection-agnostic HDF5 format, so any neutrino telescope with published response tables—public or private—can be plugged in without modifying code. The central validation is a round-trip test: ingesting the public ten-year track data release of a South Pole telescope and sampling synthetic events reproduces the observed northern-sky declination distribution to 10-15%. The reconstructed-energy distribution matches to within about a third over the bulk of the sample, with a factor-of-three over-prediction below 600 GeV that the paper attributes to the coarse true-energy binning of the public smearing matrix rather than to the sampler. If correct, the package makes joint multi-telescope analyses—combining events from different detectors under a shared source flux—straightforward and reproducible.","feed_headline":"Response tables alone can now simulate neutrino events","feed_subtitle":"Open-source sampler reproduces a public sky distribution to 10-15%, enabling joint multi-telescope studies.","key_machinery":"The load-bearing object is the standardized HDF5 instrument-response file: for each event morphology it stores the effective area $A_{\\mathrm{eff}}(\\zeta,E)$, the PSF inverse CDF $\\psi(E,u)$ mapping a uniform quantile to an angular deflection, and the energy-resolution inverse CDF $\\Delta(u)$ for the log-ratio $\\ln(E_{\\mathrm{reco}}/E_{\\mathrm{true}})$, or optionally a joint five-dimensional smearing histogram $P(E_{\\mathrm{reco}}, \\psi, \\sigma_{\\psi} | E_{\\mathrm{true}}, \\delta)$. The sampling engine is a hierarchical inverse-CDF algorithm that precomputes a weight grid $w(i,j,k) = A_{\\mathrm{eff}}(\\zeta_{ij}, E_k) \\times \\Phi(\\delta_i, \\alpha_j, E_k)$, factors it as $p(\\sin\\delta)\\, p(\\ln E|\\sin\\delta)\\, p(\\alpha|\\sin\\delta,\\ln E)$, draws three uniform variates per event, and jitters each coordinate within its cell. This replaces rejection or MCMC sampling with direct inverse-CDF draws, gives errors linear in the cell size, and lets each detector independently Poisson-sample its count for joint analyses.","core_discovery":"The paper's central claim is that the full event-level response of a neutrino telescope—direction smearing, energy resolution, and event morphology—can be captured by three tabulated functions, stored in a standardized HDF5 format, and used to Monte Carlo sample individual reconstructed events without any proprietary simulation chain. The sampler draws true directions and energies from the source flux weighted by the effective area, then applies the point-spread-function and energy-resolution inverse CDFs to produce reconstructed directions and energies, including per-event angular uncertainties when the response supplies them. A hierarchical inverse-CDF factorization over $\\sin\\delta$, right ascension, and $\\ln E$ handles extended sources with arbitrary RA- and declination-dependent flux maps, and multiple detectors can be passed to the same sampler to produce a joint, detector-tagged event list. The validation demonstrates that the sampled declination distribution of northern-sky tracks matches the observed public data to 10-15%, that the high-energy starting-event sample's deposited-energy spectrum tracks the published best-fit expectation, and that the point-source $\\psi^2$ background shape is recovered. The reported energy-proxy discrepancy below 600 GeV is presented as an input-resolution artifact, not a sampler failure, because a direct forward fold of the same IRFs reproduces it.","pith_inferences":["A natural stress test would be to apply the package to a second telescope with published response tables and public event data; if its sky and energy distributions are reproduced at similar accuracy, the selection-agnostic claim would be considerably stronger.","The identified low-energy bias is an actionable message for data releases: publishing smearing matrices with finer true-energy bins, or with kernels unweighted by the release simulation spectrum, would remove the dominant systematic this paper had to work around.","The same inverse-CDF factorisation could be extended to time-dependent effective areas or to detectors with non-separable responses (such as radio arrays) by generalising the HDF5 schema, opening event-level joint fitting across different technologies.","If adopted widely, the package would let sensitivity projections for proposed detectors be produced from response tables alone, making detector-design comparisons and multi-telescope forecasts reproducible by people outside the collaborations."],"forward_implications":["Any neutrino telescope that publishes effective area, point spread function, and energy resolution in the standardized HDF5 format can be simulated at event level with no code changes, so future detectors can enter joint studies as soon as their response tables exist.","Joint multi-detector analyses reduce to passing a list of detectors to one sampler: each detector Poisson-samples its own count, and the combined event list is tagged by detector, sharing a common source flux normalization.","Extended sources with full RA- and declination-dependent flux maps are sampled natively by the same inverse-CDF engine, avoiding MCMC burn-in and multimodal convergence problems.","The validation sets expectations for public-response-based simulations: declination distributions can be trusted at the 10-15% level, and reconstructed-energy predictions are reliable to about a third over the bulk of the sample, with a known low-energy bias traceable to coarsely binned public smearing matrices.","Both Poisson and fixed-N sampling modes, plus transient and steady-state effective-area averaging, cover the standard pseudo-experiment workflows used in point-source searches."],"supporting_citations":[{"why":"Supplies the ten-year point-source IRFs and observed data used in the round-trip validation.","marker":"[18]"},{"why":"Provides the high-energy starting event sample, its best-fit flux parameters, and the Monte Carlo used as the IRF source for the second validation.","marker":"[3]"},{"why":"Gives the nearby active galaxy spectral parameters and livetime used in the point-source angular-distribution test.","marker":"[6]"},{"why":"Provide the atmospheric flux model used in the round-trip prediction.","marker":"[23, 24]"},{"why":"Supplies the cosmic-ray flux model feeding the atmospheric flux calculation.","marker":"[25]"},{"why":"Provides the hadronic interaction model used in that calculation.","marker":"[26]"},{"why":"Establishes the normalization-uncertainty context used to interpret the fitted atmospheric normalization.","marker":"[27]"},{"why":"Defines the right-ascension scrambling technique used to build the background template in the angular-distribution validation.","marker":"[28]"},{"why":"Describes the point-source search method whose workflow the angular-distribution test follows.","marker":"[29]"}],"fun_headline_variants":["From IRFs to events: SPORE simulates neutrino skies","Neutrino events from public response tables alone","SPORE sampler reproduces IceCube sky to 10-15%","Tabulated IRFs suffice to sample neutrino events","SPORE: multi-telescope neutrino sampling from tables"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The energy round-trip conclusion assumes that the factor-of-three over-prediction below 600 GeV comes from the coarse true-energy binning of the public smearing matrix (0.5 in $\\log_{10}(E/\\mathrm{GeV})$, three declination bands, kernels averaged over the release simulation spectrum) rather than from the sampler itself or from the default effective-area smoothing and trimming.","fun_headline_variants_meta":{"raw":{"variants":["From IRFs to events: SPORE simulates neutrino skies","Neutrino events from public response tables alone","SPORE sampler reproduces IceCube sky to 10-15%","Tabulated IRFs suffice to sample neutrino events","SPORE: multi-telescope neutrino sampling from tables"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000331,"raw_usage":{"total_tokens":1923,"prompt_tokens":1102,"completion_tokens":821,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":718,"completion_tokens_details":{"reasoning_tokens":740}},"tokens_in":718,"tokens_out":821,"duration_ms":8713,"temperature":1.0,"reasoning_tokens":740,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:24:20.138927+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same detector response but replace the smearing matrix with a fine-grained version, e.g. true-energy bins of width 0.1 in $\\log_{10}(E/\\mathrm{GeV})$ in ten declination bands, and re-run the energy round-trip: if the sampled energies still over-predict the observed rate below 600 GeV by about a factor of three, then the paper's attribution of the discrepancy to the released binning would be wrong and the sampler or the preprocessing would be implicated.","supporting_citations":[],"review_version":1}