{"id":"e9d17b1c-0f45-4970-b3d7-cd5728d7d5bf","arxiv_id":"1908.07762","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A grid of 5,346 supernova models fitted to eleven type IIP supernovae recovers progenitor and nickel masses with precision comparable to progenitor imaging and yields a typical explosion energy of 10^50.52 erg.","lead":"This paper tests a computer model that predicts supernova brightness over time against eleven real exploding stars with known progenitor details. The model recovers the stars' starting masses and nickel amounts from lightcurves alone, matching the precision of direct pre-explosion images.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Parameter uncertainties rest on an uncalibrated 0.25 mag model-error floor; the precision claim is not yet established.","rationale":"The reader's weakest assumption correctly identifies the model-error term in Eq. (1) as the hinge on which the precision claim turns. My only refinement is that the danger is not simply that 0.25 mag is too small: several fits have chi2_min/N below 0.3, so the floor is generous as an uncorrelated random error. The real problem is that it is the wrong kind of error. The grid's systematic omissions, which the paper itself lists in Sections 2 and 6, produce correlated residuals and biases that Delta(chi2) thresholds cannot capture. The injection-recovery test is the cleanest way to settle this because it measures actual coverage of the quoted 68% intervals under the model variations the authors identify as their main limitations. If coverage is near 68%, the conditional verdict can be upgraded; if not, the central claim should be downgraded or narrowed to the 'relative ranking' language the authors use in Section 6. Since the paper already presents a conditional verdict and candidly acknowledges its limitations, my read does not change the reader's verdict.","tokens_in":47714,"tokens_out":10044,"duration_ms":107865,"concrete_test":"Injection-recovery test: generate synthetic V-band lightcurves with the same pipeline but from models deliberately outside the fitting grid, such as the MESA 15.6 Msun progenitor from Appendix C, CSM density multiplied by 0.5 and 2, beta=2 and 6, and Z=0.008 and 0.020, plus a few binary-stripped progenitors if available. Add realistic photometric noise and sampling matching the 11 SNe, then run the published fitting code and measure the fraction of input parameters recovered within the quoted Delta(chi2)=5.89 ranges. If that fraction is substantially below 68%, or if the injected value sits outside the recovered range even when the fit is visually good, the 0.25 mag floor and the precision claim fail.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that lightcurve fitting alone yields progenitor and nickel masses with precision comparable to progenitor imaging. This statement is only as strong as the parameter uncertainties in Tables 2-3, and those uncertainties are computed from chi2 in Eq. (1) with a fixed model-error sigma_tot containing a 0.25 mag floor. That floor was estimated as half the magnitude difference between models of adjacent explosion energies; it does not account for the physical assumptions of the grid: single-star only BPASS progenitors, fixed Z=0.014, fixed wind beta=5 and CSM density tied to the de Jager mass-loss prescription, simplified bolometric corrections, and the absence of binary progenitors. Because the synthetic lightcurves are deterministic and the residuals from the best fits are strongly correlated in time, the standard Delta(chi2)=5.89 confidence intervals do not measure systematic parameter uncertainty. The fact that several reduced chi2_min/N values are well below unity (e.g., 21.6/93 for SN2003gd, 7.7/39 for SN2004A) does not validate the error budget: it indicates the 0.25 mag term dominates and is large as a random-error model, while the correlated model error it is meant to represent is neither quantified nor propagated. A larger true systematic error would widen the quoted ranges; a differently-shaped correlated error could also shift best-fit values. Either way, the claimed equivalence to imaging precision is not demonstrated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a supernova lightcurve population synthesis (CURVEPOPS) validation study. The authors compute a grid of 5,346 SNEC supernova lightcurve models from BPASS single-star progenitors, varying initial mass (6-26 M_sun), explosion energy (log E = 50-52 in 0.25 dex steps), nickel mass (10^-3 to 10^-1 M_sun in 0.25 dex steps), and nickel mixing (three boundary fractions). They fit V-band lightcurves of 11 type IIP supernovae that have pre-explosion progenitor detections, using a chi-square statistic with a fixed 0.25 mag model error floor. They compare derived initial masses and nickel masses to progenitor-imaging and literature values, report a mean explosion energy of log(E/erg) = 50.52 +/- 0.10, and suggest correlations of nickel mass and mixing with initial mass. Two objects (SN2005cs, SN2006my) with poor (class C) fits are excluded from the quantitative analysis. The central claim is that, where a good fit is obtained, lightcurve fitting alone yields progenitor and nickel mass estimates with precision comparable to progenitor imaging.","tokens_in":47949,"tokens_out":10863,"duration_ms":94274,"significance":"If the central claim is established, this would be a valuable result: it would show that, for a subset of type IIP supernovae, lightcurve fitting can recover progenitor masses with precision comparable to pre-explosion imaging, and it would constrain the typical explosion energy of these events. The paper has clear strengths: the initial-mass validation against the independent Smartt (2015) progenitor-imaging sample is genuine and generally consistent; the public release of the SNEC input files and model lightcurves is a community resource; and the authors are transparent about many limitations (single-star only, fixed metallicity, coarse nickel-mixing grid, simplified bolometric corrections). However, the quantitative uncertainty estimates are built on an ad hoc and uncalibrated model error floor, the nickel-mass comparison is partly circular, and two of eleven objects are excluded from the population conclusions. These issues are load-bearing for the headline precision claim and must be addressed before the paper can be accepted.","major_comments":[{"comment":"The model error floor of 0.25 mag is introduced as 'half of the magnitude difference between models of successive explosion energies,' but no justification is given for why this represents a valid statistical error model for the residuals, nor why it should be added in quadrature as an independent Gaussian error. The resulting chi2_min values are far below N_obs for most objects (e.g., SN2003gd chi2_min=21.6 vs N_obs=93; SN2004A chi2_min=7.7 vs N_obs=39; SN2012ec chi2_min=10.0 vs N_obs=46), showing that the adopted sigma_tot dominates as a random-error term while the actual model deficiencies (end-of-plateau shape, early-time CSM features) are systematic and time-correlated. Because the quoted parameter uncertainties in Tables 2 and 3 are all derived from Delta(chi2)=5.89 contours using this sigma_tot, the claimed precision comparable to progenitor imaging is not underpinned by a validated error model. The authors should either calibrate the model error against the residuals of the best-fitting models or explicitly propagate the known systematic uncertainties (single-star assumption, fixed Z and beta, simplified bolometric corrections) into the parameter ranges. This is the central issue for the paper's headline claim.","section":"Sec. 4.2, Eq. (1)"},{"comment":"The quantitative conclusions, including the mean explosion energy log(E/erg)=50.52+/-0.10, exclude SN2005cs and SN2006my because their fits are classified as poor (C). SN2005cs, however, is one of the best-observed low-mass IIP supernovae and a canonical object for progenitor studies. The paper states that the poor fits indicate missing model physics, but it does not test how the mean explosion energy or the nickel-mass relation would change if these objects were included with plausible parameters (e.g., from the literature or from the mass-constrained fits). The abstract's phrase 'most of the type IIP supernovae' is based on at most 9 of 11 objects, and the sensitivity of the population conclusions to the excluded objects should be quantified.","section":"Sec. 5 and Sec. 5.2.1"},{"comment":"The validation of the recovered nickel mass is partly circular. Many of the 'literature' nickel masses listed in Tables 2-3 are derived from lightcurve modeling with similar one-dimensional explosion-plus-radiation-transport assumptions, sometimes with the same kind of code (e.g., Hendry et al. 2005a, 2006; Smartt et al. 2009; Fraser et al. 2011; Tomasella et al. 2013; Dall'Ora et al. 2014; Bose et al. 2015). The agreement between 'This Work' and 'Literature' in Figure 5 is therefore not a fully independent check on the accuracy of the nickel mass recovery. The authors should explicitly separate which literature nickel masses are independent of lightcurve fitting (e.g., from late-time nebular spectroscopy, as in Jerkstrand et al. 2015b) and base the validation claim on that subset.","section":"Sec. 5.1 and Tables 2-3"},{"comment":"The population statement 'most of the type IIP supernovae have an explosion energy of the order of log(E_exp/ergs)=50.52+/-0.10' does not specify the exact sample, weighting, or error propagation used. From Table 2, class A fits alone (SN2003gd, SN2004A, SN2012A, SN2012ec) give a mean near 50.56 with a small scatter, while including class B fits gives a mean near 50.53 but with a larger spread; the individual uncertainty bars are asymmetric and often comparable to the grid spacing. Because the per-object uncertainties from the Delta(chi2) method are unreliable (see the first major comment), the quoted +/-0.10 is not a robust measure of the uncertainty on the typical explosion energy. The authors should state the sample size, the sample standard deviation, and the standard error explicitly, and show how the result changes if class C objects are included.","section":"Sec. 5.2.1"}],"minor_comments":[{"comment":"There are several typographical issues, including 'codeB PA S S' (missing spaces) and '1050.5erg s-1' in Figure 1 captions, which should read 10^50.5 erg (energy, not power); the same unit error appears in the grid definition in Section 2.","section":"Sec. 2"},{"comment":"The summation notation is ambiguous; the expression should be written as sum over observed data points i of [(y_model(t_i) - y_obs,i)/sigma_tot]^2, with sigma_tot explicitly defined as the quadrature sum of the photometric error and the 0.25 mag model error.","section":"Eq. (1)"},{"comment":"Several table entries are garbled or hard to read, e.g., SN2006my's explosion energy appears as '50.751.13 -0.63' in Table 2, and SN2012aw's initial mass appears as '140.5-2' in Table 3; these need careful reformatting, especially for the asymmetric uncertainties.","section":"Tables 2 and 3"},{"comment":"The sentence 'The first magnitude measurement of the SNe is not necessarily the explosion date' should refer to a single supernova ('the SN'), and the phrase 'the difference between minimum and maximum is larger than 10 days' could be restated as 'if the allowed explosion-date range exceeds 10 days.'","section":"Sec. 4.2"},{"comment":"The phrase 'directed in pre-explosion imaging' should be 'detected in pre-explosion imaging.'","section":"Sec. 5.1"},{"comment":"The sentence 'we plot the best fitting model (black line) along with the lightcurves at are within the 1-sigma uncertainty in grey' contains a grammatical error and should read '...along with the lightcurves that are within the 1-sigma uncertainty in grey.'","section":"Appendix A"},{"comment":"The model identifier 'MESAv10398' should be written as 'MESA r10398.'","section":"Appendix C"},{"comment":"The sentence about constrained and unconstrained fits appears to have the comparison reversed: the text says 'the scatter is less in the latter case' (the unconstrained case), but the constrained fits should have less scatter in mass; the wording should be clarified.","section":"Sec. 6"}],"recommendation":"major_revision","confidential_remarks":"This is an honest and useful pilot study, and the public release of the model grid is a strength. The main obstacle is the unvalidated 0.25 mag model error floor, from which all the headline precision claims are derived; this is fixable in revision by calibrating the error model or by softening the claim to 'consistent within uncertainties.' The nickel-mass circularity and the sensitivity to the two excluded objects also need attention. I would not recommend rejection, because the initial-mass comparison to Smartt (2015) is genuinely informative and the limitations are openly discussed, but the central quantitative claim needs substantial strengthening."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it is a genuinely useful validation exercise: it takes a pre-computed grid of 5,346 BPASS+SNEC lightcurve models, fits them to 11 well-observed type IIP SNe with known progenitors, and shows that the inferred initial masses broadly agree with pre-explosion imaging estimates. That is the right test and the results are encouraging. Second, the abstract's claim that lightcurve fitting alone gives progenitor and nickel masses 'comparable in precision to those obtained from progenitor imaging' is not fully supported, because the quoted uncertainties come from a chi-square fit with an ad hoc 0.25 mag model-error floor that does not capture the systematic uncertainties of the grid.\n\nThe paper does several things well. The grid itself is a community resource and they make the SNEC inputs and synthetic lightcurves available. They are transparent about the two poor fits (2005cs, 2006my) and exclude them from the quantitative summary, though they still show them. The appendix test of how late-stage stellar evolution changes the structure is a nice, honest check. The observation that their masses run slightly high relative to Smartt (2015) is also reported openly and matches what Morozova et al. (2018) and Davies & Beasor (2018) found.\n\nThe soft spots are in the error budget. The 0.25 mag floor is exactly what it sounds like: half the magnitude difference between adjacent explosion-energy models. It is not calibrated against the other grid assumptions (single-star only, fixed Z=0.014, fixed wind beta, wind density tied to de Jager rates, simple bolometric corrections). The residuals are strongly correlated in time, so the Delta chi^2 = 5.89 intervals in Tables 2-3 are not a measure of systematic parameter uncertainty. Several reduced chi^2 values are well below unity (e.g., 21.6/93 for SN2003gd), which suggests the floor is doing a lot of work. A larger or differently-shaped systematic error would widen those error bars, and it could shift the best-fit values. The nickel-mass comparison is also partly circular, since many of the 'literature' nickel values in Table 2 come from lightcurve fits using the same type of explosion codes. And the mean explosion energy of log(E)=50.52±0.10 is based on only four class A fits, so the 'typical IIP energy' conclusion is preliminary.\n\nNone of this sinks the paper. The validation against independent imaging is real, and the approach will be useful for surveys where pre-explosion images do not exist. But the precision claim needs to be restated with a proper treatment of systematics, or at least with the error bars labeled as internal scatter rather than total uncertainty. I'd suggest the authors calibrate the model error using the full grid, e.g. by comparing fits across a denser set of models or by adding a correlated-error term, and report both statistical and systematic uncertainties. I'd also like to see the fitting code shipped with a commit hash, not just the grid outputs.\n\nThis is a solid, honest paper and deserves a serious referee. I would send it to review, but I would ask for a careful revision of the error analysis and a more cautious abstract.","headline":"Useful validation of a large-grid lightcurve fitting method for IIP SNe, but the precision claim rests on an uncalibrated 0.25 mag model-error floor.","tokens_in":48540,"tokens_out":3082,"would_cite":true,"duration_ms":31321,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Light-curve fitting alone recovers type IIP supernova progenitor masses with precision rivaling progenitor imaging.","keywords":["type IIP supernovae","supernova light-curve fitting","progenitor mass","population synthesis","explosion energy","nickel-56 mass","circumstellar medium","BPASS models"],"falsifier":"Apply the same fitting grid to a larger sample of type IIP supernovae that have independent mass estimates from late-time nebular spectra or host stellar populations, avoiding any use of pre-explosion imaging. If the fitted masses systematically disagree with those independent estimates by more than the quoted uncertainties, the fixed 0.25 magnitude error is too small and the precision claim fails.","tokens_in":47486,"feed_emoji":"💥","tokens_out":8484,"duration_ms":164424,"temperature":0.7,"pith_summary":"This paper asks whether a type IIP supernova's light curve alone can reveal the mass of the star that exploded, without ever imaging the star before it died. The authors fit a grid of more than five thousand synthetic light curves, built from single-star progenitor models with varying initial mass, explosion energy, nickel mass and nickel mixing, to eleven supernovae whose progenitors have been directly detected in archival images. For nine of the eleven supernovae the fit is good, and the fitted initial and nickel masses agree with the values inferred from progenitor imaging with comparable uncertainties. The two poor fits suggest the grid is missing some progenitor variety, such as binary interactions or metallicity differences. If this validation is correct, light-curve fitting can substitute for progenitor imaging when the star itself is not resolvable.","feed_headline":"Fitting light curves recovers progenitor masses, no imaging needed","feed_subtitle":"Nine of eleven nearby type IIP supernovae are matched with precision equal to direct progenitor imaging.","key_machinery":"The load-bearing machinery is a grid of 5,346 synthetic V-band light curves computed by the SuperNova Explosion Code (SNEC) from single-star stellar models of the BPASS population synthesis, covering initial masses from 5 to 26 solar masses, explosion energies log(E_exp/ergs) from 50 to 52, nickel masses log(M_Ni/M_sun) from -3 to -1, and three nickel mixing prescriptions. Each observed light curve is fitted to every model by minimising $chi^{2}$, with the explosion epoch also searched and a fixed 0.25 magnitude model error added in quadrature to the photometric uncertainty. The best-fitting model assigns initial mass, explosion energy, nickel mass, mixing parameter and explosion date, with parameter uncertainties read off from the range of the $chi^{2}$ surface in each parameter.","core_discovery":"The paper's central claim is that light-curve fitting over a large grid of synthetic models recovers the initial progenitor mass and nickel-56 mass of type IIP supernovae with precision comparable to what is achieved by detecting and modelling the progenitor star in pre-explosion images. The validation is against eleven nearby supernovae with directly observed progenitors; nine of the eleven are well fitted and their derived parameters are consistent with the progenitor-imaging values within the quoted uncertainties. The paper also finds a typical explosion energy of log(E_exp/ergs)=50.52±0.10 for the group, no strong dependence of explosion energy on initial mass, and tentative indications that both nickel mass and nickel mixing depend on the initial progenitor mass. In short, the light curve itself is positioned as a usable probe of the progenitor, not just of the explosion.","pith_inferences":["If the precision holds in a larger sample, light-curve fitting can measure progenitor masses for type IIP supernovae at distances where photometry is possible but the progenitor is not resolvable, effectively enlarging the statistical sample of core-collapse progenitors by an order of magnitude.","The mass-energy degeneracy the paper notes in the chi^2 surfaces might be broken by fitting the same grid to multi-band light curves, since the paper uses only V-band data; this is a testable extension of the same machinery.","The 0.25 magnitude model error is a placeholder; re-deriving it by calibrating against supernovae with independent mass anchors would directly test whether the precision claim is robust to the model's known simplifications."],"forward_implications":["A well fitted light curve by itself gives an initial progenitor mass that can stand in for a detected progenitor, extending mass measurements to supernovae beyond the local volume where the progenitor is resolvable.","The narrow typical explosion energy, log(E_exp/ergs)=50.52±0.10, provides a quantitative prior for core-collapse explosion models.","The tentative dependence of nickel mass on initial mass suggests a link from explosion nucleosynthesis to the pre-collapse carbon-oxygen core mass.","Including a circumstellar medium tied to the red supergiant wind reproduces the early light-curve brightening without tuning wind parameters supernova by supernova.","Two of the eleven supernovae are not well reproduced by any single-star model, marking a clear target for adding binary interactions or metallicity variations to the grid."],"supporting_citations":[{"why":"Supplies the list of eleven type IIP supernovae with directly detected progenitors and the reference progenitor masses used for validation.","marker":"Smartt (2015)"},{"why":"Provides the BPASS single-star stellar evolution models that are exploded to build the synthetic light-curve grid.","marker":"Eldridge et al. (2017)"},{"why":"Presents the SuperNova Explosion Code (SNEC) used to compute the synthetic light curves from the progenitor structures.","marker":"Morozova et al. (2015)"},{"why":"Introduces the CURVEPOPS population-synthesis project and the single-star model set that this paper extends to varying explosion parameters.","marker":"Eldridge et al. (2018)"},{"why":"Supplies the wind-acceleration and circumstellar-medium prescription that shapes the early light curves in the grid.","marker":"Moriya et al. (2018)"},{"why":"Provides the earlier SNEC-based light-curve fitting of the same supernovae, the comparison baseline for the claimed precision.","marker":"Morozova et al. (2018)"},{"why":"Re-evaluates progenitor masses from pre-explosion imaging, an independent set of mass estimates against which the fitted masses are checked.","marker":"Davies & Beasor (2018)"}],"fun_headline_variants":["Light curves rival direct imaging for nine type IIP progenitor masses","Synthetic light-curve grid recovers progenitor and nickel-56 masses","Explosion energy for nine II-P SNe constrained by light-curve synthesis","Light-curve population synthesis validates against nine direct progenitors","Progenitor mass from light curve matches imaging in nine of eleven"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the synthetic light curves never deviate from the true light curve by more than the adopted 0.25 magnitude model error; if the real systematic errors from binary progenitors, differing metallicities, uncertain wind mass loss and the unmodelled final stellar evolution stages are larger, the quoted parameter uncertainties and the claimed precision are not justified.","fun_headline_variants_meta":{"raw":{"variants":["Light curves rival direct imaging for nine type IIP progenitor masses","Synthetic light-curve grid recovers progenitor and nickel-56 masses","Explosion energy for nine II-P SNe constrained by light-curve synthesis","Light-curve population synthesis validates against nine direct progenitors","Progenitor mass from light curve matches imaging in nine of eleven"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001356,"raw_usage":{"total_tokens":5501,"prompt_tokens":937,"completion_tokens":4564,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":4473}},"tokens_in":553,"tokens_out":4564,"duration_ms":443461,"temperature":1.0,"reasoning_tokens":4473,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:57:08.833198+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the same fitting grid to a larger sample of type IIP supernovae that have independent mass estimates from late-time nebular spectra or host stellar populations, avoiding any use of pre-explosion imaging. If the fitted masses systematically disagree with those independent estimates by more than the quoted uncertainties, the fixed 0.25 magnitude error is too small and the precision claim fails.","supporting_citations":[],"review_version":1}