{"id":"25d1335e-05b8-45b8-987e-dcf9b2a5a7ba","arxiv_id":"2507.08731","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Early spectra statistically separate type II from IIb supernovae, and applying this result to low-resolution surveys suggests many IIb are mislabeled as II, revising the inferred IIb fraction upward.","lead":"Analyzing 866 early spectra of type II and IIb supernovae, this study finds that hydrogen and helium line strength and width can separate the two types, with IIb showing stronger absorption. Applying a machine-learning classifier to low-resolution spectra, the authors reclassify 34 objects, suggesting the true fraction of IIb supernovae may be nearly double previous estimates.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Rate revision from 4.0% to 7.26% is unsupported: the SEDM sample is not the magnitude-limited BTS sample, and the reported before/after fractions are inconsistent with the stated sample counts.","rationale":"The spectral distinction between SNe II and IIb appears robust: the median differences in pEW and FWHM are large (e.g., pEW Halpha at 10-20 days is 40.9 vs 9.6 A), the sample is large, and several independent ML methods agree. I therefore do not object to the qualitative claim. My concern targets the quantitative rate revision, which is the most consequential and least supported part of the strongest claim. The training-label circularity identified by the reader is real but secondary: only 16 of 393 objects were reclassified, and removing them would likely not erase the separation. The KS-test pooling issue is also real but does not affect the large effect sizes. The rate claim, however, is built on a non-representative sample and appears internally inconsistent in its arithmetic. Section 2.1 disclaims rate recovery, yet the abstract and conclusions assert a rate revision. The SEDM sample is not the BTS sample; it is a TNS convenience sample. And the 4.0% and 7.26% figures do not match the sample counts in the text. This is a concrete, checkable flaw that should be settled before the rate claim is used. The manuscript otherwise provides a valuable classification tool and a plausible demonstration that early Halpha and He I strengths separate SNe II and IIb.","tokens_in":71488,"tokens_out":7942,"duration_ms":84286,"concrete_test":"Recompute the IIb fractions directly from the stated sample counts in Section 5.2 and Tables B.1-B.3: before reclassification, 22 IIb/145 SNe = 15.2%; after reclassification, count the revised classifications (26 II to IIb, 1 IIb to II, 1 II to 87A-like, 6 inconclusive) and verify whether 4.0% and 7.26% appear in any subset. If the reported percentages cannot be reproduced, the rate revision is arithmetically unsupported. Separately, retrain the RFC after removing the 16 Section 3.1 reclassified objects to test whether the label circularity changes the SEDM reclassifications.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline rate revision (4.0% to 7.26% IIb fraction) is the most load-bearing part of the central claim, but it is not supported by the data as presented. Section 2.1 explicitly states the main sample is balanced and 'not to recover the intrinsic population rates,' yet Section 5.3 uses the SEDM/TNS sample to derive a rate shift. That sample is a convenience collection of SEDM spectra from TNS, not the ZTF Bright Transient Survey magnitude-limited sample claimed in the abstract. More concretely, the arithmetic in Figure 13 does not reproduce the stated counts: Section 5.2 gives 92 II and 14 IIb initially among 106 SNe, and after adding new classifications, 123 II and 22 IIb among 145 SNe, so the initial IIb fraction is 22/145 = 15.2%, not 4%. The 4% appears to come from the full TNS SEDM population (25/634), but the reclassification was only performed on the 145-SNe subsample, and no weighting or selection model is provided to extrapolate. The 7.26% after reclassification likewise cannot be recovered from any transparent combination of Tables B.1-B.3. If the rate claim is retained, it requires a clear denominator, an unbiased sample definition, and a propagation of the reclassification fraction.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper compiles 866 early-time spectra of 393 SNe II and IIb from WISeREP, measures the pseudo-equivalent width (pEW) and FWHM of Hα and He I λ5876, and finds that SNe IIb show systematically stronger and broader features, with the largest differences in the 10–20 day interval. The authors support this with KS tests, density contours, QDA, t-SNE+LDA, and a Random Forest Classifier, and then apply the classifier to low-resolution SEDM spectra drawn from TNS, claiming 34 misclassifications and an increase in the estimated SNe IIb fraction from 4.0% to 7.26%.","tokens_in":71732,"tokens_out":9157,"duration_ms":89353,"significance":"If the spectral separation is real, the paper provides a large, carefully measured sample and a practical, publicly released classification tool; the comparison of measurement methods against IRAF and the release of median spectra and code are concrete strengths. However, the headline rate revision is not supported by the sample definition and internal arithmetic, and the statistical tests ignore the clustering of multiple spectra within individual SNe, so the central quantitative claims need substantial revision.","major_comments":[{"comment":"The claimed revision of the SNe IIb fraction from 4.0% to 7.26% is not reproducible from the stated sample counts. Section 5.2 gives an initial comparison sample of 106 SNe (92 II and 14 IIb) and a final sample of 145 SNe (123 II and 22 IIb) after adding 39 new objects; the IIb fraction of the final sample is therefore 22/145 = 15.2%, not 7.26%, and the 'before' 4.0% appears to be 25/634 from the full TNS SEDM collection rather than from the 145-SNe sample that was actually reclassified. Moreover, Tables B.2 and B.3 list 25 and 39 objects, respectively, and the number of definite II-to-IIb reclassifications in those tables does not add up to the 26 reported in the text; the 34 'misclassified' count and the before/after percentages therefore require an explicit denominator, a selection model for the SEDM/TNS sample, and a reconciliation of the table entries.","section":"Section 5.3, Figure 13"},{"comment":"The KS tests treat all 866 spectra as independent, but the sample contains 393 SNe with 271 single-spectrum objects, 42 with two spectra, 22 with three, and 55 with four or more; repeated spectra from the same SN are correlated measurements of a single object. This clustering inflates the significance of the reported p-values (e.g., 7.9e-51 for the 0–40 d pEW comparison) and can bias the identification of the 10–20 d interval as the most significant. The authors should repeat the KS analysis with one randomly selected spectrum per SN, or use a mixed-effects model or a bootstrap resampled by SN, to verify that the time-interval ranking is robust.","section":"Section 4.1, Table 2"},{"comment":"The reclassification of 16 objects before the analysis was based on visual inspection of the same Hα and He I features that later serve as the pEW/FWHM inputs to the classifier in Section 4.2.4. This is a mild circularity because the training labels are partly derived from the discriminating features, so the measured separation and the classifier accuracy partly encode the authors' visual judgments. The authors should quantify the impact by rerunning the classification with the original TNS/WISeREP labels or by excluding the 16 reclassified objects from the training set.","section":"Section 3.1, Table 1"},{"comment":"The low-resolution sample is described in the abstract as from the 'Zwicky Transient Facility Bright Transient, a magnitude-limited survey,' but Section 5.2 states that the sample was gathered as 'all available SEDM spectra for SNe II and IIb from the TNS' (with objects lacking light curves excluded). These are different selections; the TNS collection is not necessarily the magnitude-limited BTS sample, and no selection function is given. Extrapolating the reclassification fraction from the 106/145-SNe subsample to the full population therefore requires a model relating the subsample to the underlying rate, which is not provided.","section":"Section 5.2 and abstract"}],"minor_comments":[{"comment":"The phrase 'Zwicky Transient Facility Bright Transient' should be 'Zwicky Transient Facility Bright Transient Survey (BTS)', and the sample used in Section 5.2 should be explicitly identified as a TNS SEDM collection, not as BTS itself.","section":"Abstract"},{"comment":"The text states that the initial sample contains 449 low-redshift SNe, but the final dataset after cuts is 393 SNe; please state the number of objects removed by each cut (resolution, explosion-date quality, and the SEDM exclusion).","section":"Section 2.1"},{"comment":"The table caption 'Comparison of precision and recall scores for the different classification methods' does not match the column layout; the columns labelled 'pEW' and 'FWHM' appear to be separate sub-tables, and the caption should clarify which features each block uses.","section":"Table 3"},{"comment":"The caption 'FWHM HeI vs pEW Halpha' is likely a typo; the panel shows FWHM of He I versus FWHM of Hα, and the axis labels should match.","section":"Figure A.7 caption"},{"comment":"Please report the number of trees, maximum depth, and other hyperparameters of the Random Forest Classifier, and use repeated cross-validation rather than a single 60/40 split to avoid optimistic precision/recall estimates.","section":"Section 4.2.4"}],"recommendation":"major_revision","confidential_remarks":"The paper's spectral comparison is potentially useful, but the headline rate revision is internally inconsistent and the abstract overstates the connection to the BTS magnitude-limited sample. The clustering issue in the KS tests and the label circularity are also substantive. These are fixable within the manuscript's scope, but they require a significant rewrite of Section 5 and the abstract."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, the core of this paper is worth your time: a large, quantitative comparison of early-time Halpha and He I lines in SNe II vs IIb, with a classifier that actually seems to work. The 866-spectrum sample and the public code/data on Zenodo are real assets. The most interesting number in the abstract - the IIb fraction jumping from 4.0% to 7.26% - does not survive arithmetic. Section 5.2 describes a 145-SNe SEDM/TNS sample with 22 IIb; that is 15.2%, not 4%. The 4% appears to come from the full TNS SEDM population, but the reclassification was only performed on the subsample, and no selection model bridges the gap. Calling this a magnitude-limited survey is also generous; it is a convenience sample of SEDM spectra from TNS, not the ZTF BTS sample. This is a load-bearing flaw, and the rate claim should be withdrawn or completely redone.\n\nThe rest is in better shape. The pEW and FWHM differences are consistent across time bins, cleanest in the 10-20 day interval, and replicated by several independent ML methods. That is a useful, citable result for anyone doing early spectroscopic classification. Two smaller caveats: the 16 reclassifications in Section 3.1 were done using the same Halpha/He I features that later feed the classifier, so the labels are not fully independent of the features - mild circularity, not fatal. And the KS tests treat multiple spectra from one SN as independent observations, which inflates significance; a per-SN bootstrap or mixed-effects version would be more honest.\n\nBottom line: the spectral separation and the classification tool deserve a serious referee. The rate revision does not deserve to be in the abstract until the denominator is fixed.","headline":"The spectral separation result and the classification tool are solid and citable, but the headline IIb rate revision (4.0% to 7.26%) does not survive arithmetic and should not be used until the denominator and sample selection are fixed.","tokens_in":72319,"tokens_out":2595,"would_cite":true,"duration_ms":31224,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Type IIb supernovae show stronger H-alpha and He I absorption than type II in the first 40 days, a difference strong enough to reclassify 34 supernovae and nearly double the estimated IIb fraction.","keywords":["supernova classification","type IIb supernovae","type II supernovae","pseudo-equivalent width","H-alpha line","He I 5876","core-collapse supernova rates","random forest classifier"],"falsifier":"Take a sample of type II and IIb supernovae classified by an independent method, such as late-time spectra that clearly show either persistent hydrogen lines or helium-dominated ejecta, and re-measure pEW and FWHM in the first 10-20 days; if those two distributions overlap nearly completely, the paper's separation is an artifact of its training labels, whereas a clean offset would confirm the claim.","tokens_in":71279,"feed_emoji":"💥","tokens_out":9095,"duration_ms":101788,"temperature":0.7,"pith_summary":"This paper sets out to show that type II and type IIb supernovae, which look nearly identical in the first days after explosion, can be told apart quantitatively by how much their early spectra absorb. Using 866 public spectra from 393 supernovae, it measures the pseudo-equivalent width and full width at half maximum of H-alpha and He I 5876 absorption within 40 days of explosion and finds that type IIb events consistently show stronger and broader absorption, with the clearest separation between 10 and 20 days. The paper uses those measurements to train a random-forest classifier and applies it to low-resolution spectra, identifying 34 supernovae whose official classifications are probably wrong. Correcting for these raises the inferred fraction of type IIb events from 4.0% to 7.26%, so the practical stake is that published core-collapse supernova subtype rates may be biased by early-time misclassification.","feed_headline":"Spectra reclassify 34 'type II' supernovae as IIb","feed_subtitle":"Measuring H-alpha and helium absorption in the first 40 days raises the inferred IIb fraction from 4.0% to 7.26%.","key_machinery":"The load-bearing objects are the pseudo-equivalent width (pEW) and the full width at half maximum (FWHM) of the H-$\\alpha$ and He I $\\lambda5876$ absorption features. pEW, defined as $\\mathrm{pEW}=\\sum_i (1 - f(\\lambda_i)/f_0(\\lambda_i))\\,\\Delta\\lambda_i$, measures how much flux a line removes relative to the local continuum, and the FWHM, obtained from a Gaussian fit as $2\\sqrt{2\\ln 2}\\,\\sigma$, measures the velocity spread of the ejecta. The paper combines these four measurements with the epoch and feeds them to a Random Forest Classifier, whose classification probabilities provide the practical tool; Quadratic Discriminant Analysis supplies the decision regions used for visual checks, and t-SNE with LDA shows that the two classes form two clusters with a mixed continuum between them.","core_discovery":"On the paper's own terms, the central discovery is that the strength of the H-alpha and He I 5876 absorption features is a clear discriminator between SNe II and SNe IIb in the first 40 days, not just in individual well-known objects but across a large balanced sample. SNe IIb have higher pseudo-equivalent widths and broader profiles at all early phases, the difference is statistically significant up to about day 30, and it peaks in the 10-20 day window. A Random Forest Classifier built from the pEW and FWHM of both lines plus the epoch separates the two classes with precision and recall above 0.8, and when applied to 106 low-resolution spectra it flags 34 likely misclassifications: 26 objects listed as type II look like type IIb, one listed as IIb looks like type II, and six remain ambiguous. Reclassifying those events changes the low-resolution sample's IIb fraction from 4.0% to 7.26%, which the paper interprets as evidence that spectroscopic misclassification of early spectra has a measurable effect on estimated core-collapse supernova rates.","pith_inferences":["An extension the paper does not make: if the 7.26% fraction holds up, the true volumetric rate of SNe IIb may be roughly twice the commonly quoted value, and rate comparisons across surveys should re-derive completeness using the classifier's probabilities rather than hard labels.","The pEW continuum between the classes could be mapped directly to hydrogen-envelope mass using the models cited in the paper, turning a classification tool into a physical mass-loss estimator.","A direct testable extension would be to measure the same pEW/FWHM features on H-beta and the He I 6678 line; since the paper shows H-alpha and He I 5876 lose separation after day 30, additional lines may extend reliable classification later.","The classifier's success on low-resolution spectra suggests that real-time early classification from survey-grade spectra is feasible; a follow-up would be to run the same feature measurements on fully automatic continuum fitting and quantify how much the interactive continuum placement affects the probabilities."],"forward_implications":["SNe IIb will, on average, have stronger H-alpha and He I absorption than SNe II throughout the first 30 days, so an early spectrum alone can flag an object whose official type is doubtful.","The 10-20 day window is the most informative; spectra taken before day 10 or after day 30 lose discriminating power, so classification efforts should prioritize that window.","Line strength (pEW) separates the classes better than line width (FWHM), so coarse-resolution spectra that still resolve line depth may suffice for screening.","Applying the classifier to a magnitude-limited low-resolution sample shifts the inferred SNe IIb fraction from 4.0% to 7.26%, implying published subtype fractions for core-collapse supernovae are affected by misclassification.","The two populations form a continuum with overlapping regions rather than a clean gap, so the classifier's output is a probability, and objects near the boundary need additional information such as light-curve shape."],"supporting_citations":[{"why":"Supplies the explosion-date estimation method, the pEW measurement conventions, and a large set of the type II spectra used here.","marker":"Gutiérrez et al. (2017)"},{"why":"Establishes the median-spectrum comparison for stripped-envelope SNe and provides the type IIb median spectra that the authors compare their sample against.","marker":"Liu et al. (2016)"},{"why":"Documents how He I 6678 blends with H-alpha in SNe IIb, which is why the analysis focuses on the H-alpha absorption and He I 5876.","marker":"Holmbo et al. (2023)"},{"why":"Radiative-transfer models linking hydrogen-envelope stripping to H-alpha and helium line strengths; used to explain why IIb show stronger absorption.","marker":"Dessart et al. (2024)"},{"why":"Establishes the photometric gap and rise-time differences between SNe II and IIb used as an independent check in the reclassifications.","marker":"Pessi et al. (2019)"},{"why":"Provides the baseline CCSN subtype fractions (about 70% II and 11% IIb) against which the authors' revised IIb fraction is compared.","marker":"Shivvers et al. (2017)"},{"why":"Provides the magnitude-limited survey sample and its CCSN fraction estimates used for the relative-rate comparison.","marker":"Perley et al. (2020)"},{"why":"Provides the public spectral archive from which the main sample was drawn.","marker":"Yaron & Gal-Yam (2012)"},{"why":"SNID template matching is used to estimate explosion epochs for objects without photometric non-detections.","marker":"Blondin & Tonry (2007)"}],"fun_headline_variants":["Early spectra expose 34 misclassified supernovae as type IIb","H-alpha and helium absorption distinguish SNe II from IIb","Machine learning on early spectra corrects 34 supernova types","Supernova IIb fraction nearly doubles after spectral reclassification"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The classifications used to train the classifier were partly assigned by visually inspecting the same H-alpha and He I features that later serve as classifier inputs, so the measured separation may partly encode the authors' own judgments rather than an independent ground truth.","fun_headline_variants_meta":{"raw":{"variants":["Early spectra expose 34 misclassified supernovae as type IIb","H-alpha and helium absorption distinguish SNe II from IIb","Machine learning on early spectra corrects 34 supernova types","Supernova IIb fraction nearly doubles after spectral reclassification"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000401,"raw_usage":{"total_tokens":2194,"prompt_tokens":1145,"completion_tokens":1049,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":761,"completion_tokens_details":{"reasoning_tokens":977}},"tokens_in":761,"tokens_out":1049,"duration_ms":13131,"temperature":1.0,"reasoning_tokens":977,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:10:41.947352+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a sample of type II and IIb supernovae classified by an independent method, such as late-time spectra that clearly show either persistent hydrogen lines or helium-dominated ejecta, and re-measure pEW and FWHM in the first 10-20 days; if those two distributions overlap nearly completely, the paper's separation is an artifact of its training labels, whereas a clean offset would confirm the claim.","supporting_citations":[],"review_version":1}