{"id":"5b43bbaf-87e2-4e77-900a-a5bb33a3f4aa","arxiv_id":"1908.05228","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A data-driven library of 67 individual core collapse supernova spectral templates, extended to about 1600 Angstroms, is constructed and demonstrated in simulations of photometric supernova surveys.","lead":"This paper builds a library of 67 spectral time-series templates for core collapse supernovae from publicly available photometry and spectroscopy, using Gaussian processes to interpolate and extend the data into the near-ultraviolet. The templates, together with revised luminosity functions, are designed for simulating realistic supernova populations in time-domain surveys such as LSST and for estimating contamination in cosmological supernova samples.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Near-UV template diversity is regularized by a same-sample prior; in-sample UV residuals leave the z~1 suitability claim untested.","rationale":"The reader's weakest assumption identifies the same-sample prior as the mechanism that can suppress UV diversity, and I agree that is the most suspicious component. I would frame the load-bearing issue slightly more broadly: the quoted 0.02/0.1 mag reproduction in Figure 7 is entirely in-sample, because the templates are remangled and re-calibrated against the same photometry that is later compared (Sections 2.2 and 3.2). That means neither the optical nor the UV residual is independent evidence that the templates generalize to new events or faithfully represent population diversity. The near-UV prior is the clearest concrete route by which in-sample behavior can hide poor out-of-sample performance, especially for events with sparse UV coverage, which is the majority of the 67-event sample. The paper's own demonstration at z>0.4 shows that the near-UV extension materially changes contamination predictions, so this is exactly where the claim is load-bearing. The proposed jackknife test directly separates the effect of the prior from the event's own UV data and would show whether the prior is simply a mild regulariser or a dominant constraint that under-represents diversity. I do not see grounds to reject the paper: the method is well documented, the code is open-source, and the library is a real resource. But the suitability claim should remain conditional on an out-of-sample UV check. The reader's CONDITIONAL verdict therefore stands unchanged, and the requested conditions (soften the fully data-driven wording, add out-of-sample validation, quantify template uncertainties) are appropriate.","tokens_in":34043,"tokens_out":8613,"duration_ms":101812,"concrete_test":"Jackknife the near-UV prior: for each of the 67 SNe, rebuild the sub-class colour prior from the other 66 SNe only (excluding the target), keep the target's own UV photometry unchanged, and rebuild the target template using the published code. Then compare synthetic UV photometry from the jackknife template to the observed UV photometry and to the standard template built with the target included in the prior. If the jackknife RMS in the Swift/UVOT filters exceeds the quoted 0.1 mag for events with fewer than about three UV epochs, or if the prior-included versus prior-excluded UV fluxes differ by more than 0.1 mag, the same-sample prior dominates the UV extension and the library's UV diversity is under-represented; the z~1 suitability claim would need qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline near-UV extension (Section 2.2.1) is the least secure part of the central claim. For each SN, the 2D Gaussian-process prior is a sub-class average colour surface built from the same 67 SNe, multiplied by that SN's V-band light curve, with fixed GP length scales (sigma_lambda = 100 A, sigma_t = 30 days). When UV photometry is sparse, the posterior mean is pulled toward this average, so an event whose true UV evolution deviates from the sub-class mean is rendered more average in the library. The only test of the prior (Section 2.2.1) uses two of the best-observed SNe (SN 2013by and iPTF13bvn) and does not quantify shrinkage for the sparse-UV majority. Moreover, Figure 7 is an in-sample check: the photometry used to build each template (after the remangling in Section 2.2.2 that forces synthetic fluxes to match observed photometry) is the same photometry later compared, so the quoted 0.1 mag UV residuals do not demonstrate predictive accuracy. This matters because the stated suitability up to z~1 depends on rest-frame UV fidelity, and the paper's own simulation (Section 4.3, Fig. 9) shows the near-UV extension changes the predicted contamination by a factor of two at z>0.4. The library remains useful, but the suitability claim should be qualified until UV diversity is tested out-of-sample.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper constructs a library of 67 rest-frame, daily-sampled spectral time-series templates for core-collapse supernovae, using literature photometry and spectroscopy. Light curves are interpolated with Gaussian processes, observed spectra are flux-calibrated by mangling to the photometry, and the wavelength coverage is extended into the near-UV with a two-dimensional GP interpolation that uses a sub-class average colour prior. Templates are optionally corrected for host-galaxy extinction and are provided in a form suitable for use with snana. The library is applied to simulate core-collapse contamination in a PS1-like photometric SN Ia sample, comparing three luminosity-function choices, and the results are compared with the observed Hubble-residual distribution.","tokens_in":34381,"tokens_out":4918,"duration_ms":52469,"significance":"If the validation caveats are addressed, this is a valuable community resource: it provides a larger, event-by-event template library than existing alternatives, retains object-to-object diversity, extends into the near-UV, and is released as open-source code with an example snana integration. The simulation application is an independent use of the templates in the sense that they are not tuned to the PS1 contamination data, and the paper honestly reports that the standard luminosity functions underproduce the observed contamination. The main scientific value lies in enabling more realistic simulations for photometric classification and SN Ia cosmology, but the headline claims about near-UV fidelity and usefulness up to z~1 currently rest on in-sample validation.","major_comments":[{"comment":"The near-UV extension is validated only in-sample. The prior for the two-dimensional GP is a sub-class average colour surface built from the same 67 SNe, with fixed length scales sigma_lambda=100 Å and sigma_t=30 days; for events with sparse UV photometry the posterior is pulled toward this sub-class average, so the library is likely to under-represent true UV diversity. The test described in §2.2.1 uses only SN 2013by and iPTF13bvn and does not quantify shrinkage for the remaining majority of events. Since §4.3 and Fig. 9 show that using the near-UV-extended templates changes the predicted contamination by a factor of two at z>0.4, the suitability claim up to z~1 should be qualified until the UV extension is tested on held-out events or through a cross-validation scheme.","section":"§2.2.1"},{"comment":"The quoted photometric recovery of 0.02 mag in the optical and 0.1 mag in the near-UV is an in-sample test: the same extinction-corrected photometry used to construct each template, and subsequently forced to match through the remangling step in §2.2.2, is the photometry compared with the final synthetic template. This demonstrates internal consistency but not predictive accuracy on new events or on data deliberately withheld during construction. An out-of-sample test, such as leaving out a filter or an epoch during template construction and checking the residuals, would support the claimed accuracy; without it, the 0.1 mag near-UV figure should be described as a fitting residual rather than as an expected template error.","section":"§3.2 / Fig. 7"},{"comment":"The host-extinction corrections carry large, acknowledged systematic uncertainties—Na i D equivalent widths with large scatter, a fixed RV=3.1, and average reddening assumed for four SNe—yet the de-reddened templates are a central product used in the §4.2.1 simulations with the R14 luminosity functions. After corrections, the uvw1−V scatter of the templates remains about 0.5–0.7 mag for stripped-envelope SNe, and Fig. A4 shows that the median host reddening in the sample is factors of two to four below the Prentice et al. (2016) values for stripped-envelope types. The paper should quantify how these uncertainties propagate into the simulated contamination rates, or it should restrict the claims made for the de-reddened-template application.","section":"Appendix A / Fig. A2"},{"comment":"The abstract and Section 5 state that the templates are built with no assumption of any parametric form or model for the light curves, but Section 2.1.2 explicitly uses the parametric power law f(t)=alpha(t-t0)^n for early phases, with n fixed to 1.5 for stripped-envelope SNe and 0.935 for hydrogen-rich SNe, and Eq. (3) adds a parametric shock-breakout component. The claim should be qualified to state that no parametric form is assumed over most of the light curve, or the early-rise parametrization should be integrated into the GP model; as written, the abstract overstates the data-driven nature of the method.","section":"Abstract and §2.1.2"}],"minor_comments":[{"comment":"The sentence describing Fig. 10 contains a typo: 'SNe Ib ans SNe Ic' should read 'SNe Ib and SNe Ic'.","section":"§4.3"},{"comment":"The exponential term in Eq. (3) is ambiguous; writing exp[-(t-t0)/tau] or otherwise clarifying the functional form of the shock-breakout component would improve reproducibility.","section":"Eq. (3)"},{"comment":"The column header 'Numb. of Ref. Spectra Host' appears to combine the spectrum count and the reddening reference into one heading; splitting these into two clearly labelled columns would prevent confusion.","section":"Table 2"},{"comment":"The classification scheme lists six sub-types but SN 1987A is also included as a template; the text should clarify whether SN 1987A is treated as a seventh class in simulations or is assigned to one of the six classes.","section":"§3.1.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the scope of MNRAS and represents a useful contribution. The main concern is that the near-UV validation and the photometric-recovery test are both in-sample; I would ask the authors to add a hold-out or cross-validation experiment, or to clearly reword the accuracy and z~1 suitability claims. The 'fully data-driven' wording also needs adjustment given the parametric early-rise model."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper gives the community a genuinely useful thing: 67 individual-event spectral templates for core-collapse SNe, daily sampled, near-UV extended, with open-source code. Building each template from its own photometry and spectroscopy, rather than averaging events into class templates, preserves event-to-event diversity, which matters for contamination studies. The 2D GP interpolation of the flux surface is a sensible way to extend into the near-UV and fill sparse spectral coverage. The simulation demonstration with SNANA is a real application, not a toy; it shows the contamination predictions are sensitive to both template library and luminosity functions, which is itself a useful result.\n\nWhat I'd flag, in proportion:\n\n1. The \"fully data-driven\" claim is overstated. The abstract and summary say no parametric form is assumed, but Section 2.1.2 fits the early rise with a power law f(t) ~ (t-t0)^n and sometimes a shock-breakout bump. It's a small fraction of the light curve, but the claim as written is wrong. Easy fix: say \"fully data-driven except for the earliest phases\" or similar.\n\n2. The near-UV validation is in-sample. Figure 7 compares synthetic photometry from the final templates to the same observed photometry that went into building them, after remangling forces those fluxes to match. The 0.02/0.1 mag residuals therefore don't demonstrate predictive accuracy. The prior for the 2D GP is built from the same 67 SNe, so if an event's UV evolution deviates from the sub-class average, the template gets pulled toward that average. The paper tests the prior on two well-observed events, but doesn't quantify shrinkage for the sparse-UV majority. Since the z~1 suitability claim rests on rest-frame UV fidelity, that claim needs qualification. This is the weakest part of the paper, but it is not fatal: the templates still preserve a lot of diversity, and the simulation example is an independent use of the library.\n\n3. Host-extinction corrections carry large systematics; the authors acknowledge this in Appendix A and in the main text. The four SNe without reddening estimates are given a median value, which is a reasonable stopgap. Fine.\n\n4. The simulation results are conditional on assumed luminosity functions and rates; the paper is appropriately careful, noting the Hubble residual analysis alone can't disentangle LF, color, and stretch effects.\n\nOverall: the central deliverable—the template library and code—is solid and deserves to be in the literature. The paper would benefit from softening the data-driven language, labeling the residual test as in-sample, and either adding an out-of-sample check or explicitly stating the UV diversity limitation. I'd send this to a serious referee. The authors know the limitations; the issue is that some claims in the abstract and summary outrun the evidence.","headline":"Useful library of 67 core-collapse SN templates with real code, but the 'fully data-driven' label and the in-sample UV validation are weaker than the abstract implies.","tokens_in":34930,"tokens_out":2910,"would_cite":true,"duration_ms":30667,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper builds a fully data-driven library of 67 core-collapse supernova spectral templates, extended into the near-UV, that reproduces observed photometry to about 0.02 mag in the optical and 0.1 mag in the near-UV, and uses it to…","keywords":["core-collapse supernovae","spectral templates","Gaussian processes","near-ultraviolet photometry","time-domain surveys","supernova contamination","luminosity functions","photometric classification"],"falsifier":"Take a well-observed stripped-envelope event, rebuild its template with the near-UV photometry withheld, and compare the near-UV light curve predicted by the GP surface with the actual satellite UV measurements; if the withheld fluxes deviate systematically by more than about 0.1 mag across several events, the claim that the UV extension is faithful fails.","tokens_in":33833,"feed_emoji":"💥","tokens_out":7944,"duration_ms":71679,"temperature":0.7,"pith_summary":"The paper sets out to supply the missing ingredient for simulating realistic populations of core-collapse supernovae in time-domain surveys: a library of spectral time-series templates built event by event from real data, with no assumed light-curve model. Using 67 published core-collapse supernovae, the authors construct daily-sampled, rest-frame templates extended into the near-UV to about 1600 \\AA{}, with an optional correction for host-galaxy dust. They claim the templates recover the observed photometry to within 0.02 mag in the optical and about 0.1 mag in the near-UV, and that the library, combined with luminosity functions and relative rates, can simulate core-collapse contamination in photometric type Ia samples out to $z\\sim1$. They demonstrate this by repeating a published photometric SN Ia contamination analysis with the new templates and three different luminosity-function choices, finding that the predicted contamination level depends strongly on the assumed luminosity functions and extinction treatment. The template-building code is open-source and generalisable to any transient with well-sampled photometry and multiple spectra.","feed_headline":"67 supernova templates reproduce survey photometry to 0.02 mag","feed_subtitle":"Data-driven, near-UV-extended core-collapse templates let simulations test contamination in type Ia samples out to z~1.","key_machinery":"The load-bearing object is the two-dimensional Gaussian-process flux surface $f(t,\\lambda)$, built by combining flux-calibrated spectra and broad-band photometry (including satellite UV filters) on a 60 \\AA{} wavelength grid. A Matern 3/2 kernel with fixed length scales $\\sigma_\\lambda = 100\\,$\\AA{} and $\\sigma_t = 30$ days interpolates the surface, and the GP mean is a sub-class average colour surface multiplied by the event's own V-band light curve. This one surface performs three jobs: it extends the optical spectra into the near-UV, it interpolates between sparsely sampled spectra, and it lets the code extrapolate additional daily spectra, which are then remangled so that synthetic and observed photometry agree in every filter.","core_discovery":"The central claim is that a spectral template for a core-collapse supernova can be built entirely from its own multi-band photometry and sparse spectroscopy, with no parametric light-curve or SED model. Each event becomes its own template: Gaussian processes interpolate each light curve, a mangling step flux-calibrates each observed spectrum against the interpolated photometry, and a two-dimensional Gaussian process over time and wavelength combines spectra with near-UV photometry into a continuous flux surface $f(t,\\lambda)$. Re-sampling that surface daily and remangling against photometry yields a spectral time series that reproduces the observed colours. The paper claims this near-UV-extended, event-by-event library preserves the diversity of core-collapse supernovae and is accurate enough for simulations of photometric surveys to $z\\approx1$.","pith_inferences":["A natural stress test not performed in the paper is a true out-of-sample UV validation: withhold the near-UV photometry from the GP surface for several well-observed events and compare the predicted UV light curves with the measured ones; the fixed prior would be expected to mask genuine UV outliers.","The fixed GP length scales encode an assumption that the flux surface is smooth on 100 \\AA{} and 30-day scales; transients with fast UV spectral evolution (early shock breakout or flash ionisation) may be smoothed over, which is a testable limitation when applying the code to other transients.","The selection criteria (UV photometry, at least five spectra, pre-peak coverage) bias the library against faint or highly reddened events, and the paper notes its stripped-envelope templates have lower median host reddening than published samples; the claimed diversity is therefore the diversity of well-observed, relatively unobscured events.","The UV extension is anchored by photometry rather than by a physical UV line-blanketing model, so at high redshift the simulated rest-frame UV colours inherit whatever the GP surface does between anchors; comparing simulated high-z colours with observed rest-frame UV colours of the same events would quantify that systematic."],"forward_implications":["The library plugs directly into existing supernova simulation pipelines, so photometric SN Ia analyses can generate core-collapse contamination with event-level spectral diversity rather than a handful of averaged templates.","At redshifts above about 0.4, the near-UV extension roughly doubles the predicted core-collapse contamination compared with optical-only templates, making the UV coverage essential for high-redshift contamination estimates.","Luminosity functions tuned against one template library do not transfer cleanly to another; the paper's repeats give contamination fractions of 3.7, 9.5, and 7.5 per cent depending on the luminosity function and extinction treatment.","Host-extinction-corrected templates combined with simulated dust overproduce bright contaminants by about a factor of three, indicating that the extinction model or luminosity function needs revisiting.","Because the code is open-source and data-driven, the library can be extended to any future transient with well-sampled photometry and multiple spectra, growing the diversity coverage of the sample."],"supporting_citations":[{"why":"Defines the predecessor 41-template library and the SNPhotCC simulation set that this work upgrades.","marker":"Kessler et al. 2010a,b"},{"why":"Provides the spectral repository from which most template input spectra were drawn.","marker":"Yaron & Gal-Yam 2012"},{"why":"Supplies the photometric catalogue from which most light curves were downloaded.","marker":"Guillochon et al. 2017"},{"why":"Provides the satellite UV photometry that drives the near-UV extension to about 1600 Angstroms.","marker":"Brown et al. 2014"},{"why":"Supplies the R-band core-collapse luminosity functions adopted in one simulation family.","marker":"Li et al. 2011"},{"why":"Supplies the B-band luminosity functions adopted in the extinction-corrected simulation family.","marker":"Richardson et al. 2014"},{"why":"Defines the published photometric SN Ia sample and contamination simulations that the demonstration repeats.","marker":"Jones et al. 2017"},{"why":"Provides the revised subtype classifications and relative rates used to weight simulated events.","marker":"Shivvers et al. 2017"},{"why":"Provides colour-evolution-based host reddening estimates for stripped-envelope events.","marker":"Stritzinger et al. 2018b"},{"why":"Gives the dust extinction law used to correct photometry and templates.","marker":"Cardelli et al. 1989"}],"fun_headline_variants":["No parametric SN models: each supernova builds its own template","67 SNe yield data-driven templates for survey simulations","Near-UV extended templates from photometry, no light-curve fits","Spectral templates that mimic real core-collapse diversity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The near-UV extension of every template leans on a sub-class average colour prior with fixed smoothing scales, so an event whose ultraviolet behaviour is genuinely unusual is pulled toward the average and its true UV diversity may be lost.","fun_headline_variants_meta":{"raw":{"variants":["No parametric SN models: each supernova builds its own template","67 SNe yield data-driven templates for survey simulations","Near-UV extended templates from photometry, no light-curve fits","Spectral templates that mimic real core-collapse diversity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00042,"raw_usage":{"total_tokens":2172,"prompt_tokens":969,"completion_tokens":1203,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":585,"completion_tokens_details":{"reasoning_tokens":1134}},"tokens_in":585,"tokens_out":1203,"duration_ms":18283,"temperature":1.0,"reasoning_tokens":1134,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:37:25.715427+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a well-observed stripped-envelope event, rebuild its template with the near-UV photometry withheld, and compare the near-UV light curve predicted by the GP surface with the actual satellite UV measurements; if the withheld fluxes deviate systematically by more than about 0.1 mag across several events, the claim that the UV extension is faithful fails.","supporting_citations":[],"review_version":1}