{"id":"8c8347fa-6151-48bd-b691-011579ab88fa","arxiv_id":"2607.03532","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"ABC-SN classifies ten supernova subtypes with no performance loss down to R_λ=50 and SNR=5, and only minimal loss at R_λ=25.","lead":"Supernova subtype classification with a deep-learning model stays accurate down to spectral resolution R=50 and SNR=5, and is only mildly degraded at R=25. This maps the practical floor for spectroscopic follow-up of the millions of supernovae LSST will find, freeing high-resolution resources for detailed science.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The SNR=5 / R=50 thresholds rest on a non-standard, subtype- and phase-specific line SNR that required heavy manual curation and spectrum culling.","rationale":"The reader correctly isolated the custom SNR construction as the weakest assumption. That construction is load-bearing for the numerical claim, yet the paper is transparent about the manual steps, the removals, and the definition itself; the experimental design (homogeneous degradation, leakage-free 2-fold CV, macro-F1, public code) remains sound. The limitation does not overturn the empirical map or the practical utility of the result, so the ACCEPT verdict stands. The proposed test simply quantifies how sensitive the headline numbers are to the SNR definition; a large shift would convert the claim from “SNR=5” to “SNR≈5 under this particular metric,” which is still useful but less immediately actionable for instrument designers.","tokens_in":27386,"tokens_out":612,"duration_ms":21930,"concrete_test":"Re-generate the R_λ=50 and R_λ=100 grids at target SNR=5 and SNR=2.5 using a single, fully automatic continuum-based SNR (median flux / RMS in two line-free windows, e.g. 5100–5200 Å and 6800–6900 Å) instead of the line-specific S of Eq. 2; retrain ABC-SN with the identical protocol. If macro-F1 at the paper’s “SNR=5, R=50” cell falls by more than ~8–10 points relative to the published value, the quoted thresholds are definition-dependent and less transferable to real observing programs.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim quotes concrete numerical thresholds (no loss down to R_λ=50 and SNR=5; only minimal impact to R_λ=25). Those numbers are defined exclusively by the custom SNR of §3.2–3.3 and Table 3: S is the width-normalized area of one hand-chosen diagnostic feature (Si II 6355, He I 5876, Na I D/S blend, Hα, Fe lines, or a late-time blend) relative to a linear pseudo-continuum, after Gaussian smoothing whose σ was manually tuned per spectrum. ~10 % of the library (190 + 179 spectra) was discarded because the procedure failed or produced outliers. At low R the same fixed shoulder wavelengths become unreliable (explicitly noted by the authors), yet the grid still reports “SNR=5”. Because every cell of Figures 6–7 is generated from this definition, any systematic bias in how the definition ranks classification difficulty across subtypes or resolutions directly shifts the claimed thresholds. The degradation itself (white Gaussian noise added to the extracted signal) is also idealized relative to real Poisson + sky + host contamination.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper systematically measures how supernova subtype classification performance depends on spectral resolution R_λ and signal-to-noise ratio (SNR). Using a curated SNID-derived library of ten subtypes (Ia-norm, Ia-91T, Ia-91bg, Iax, Ib-norm, Ibn, IIb, Ic-norm, Ic-broad, IIP), the authors define a line-based, width-normalized SNR (Eq. 2 and Table 3), degrade the spectra to a grid of 16 SNR values and 14 resolutions, retrain the attention-based classifier ABC-SN on each of 476 datasets (2-fold CV that keeps all spectra of a given SN in one fold), and report macro-F1 heat-maps (Figs. 6–7). They conclude that refined subtype classification remains essentially undiminished down to R_λ = 50 and SNR = 5, and is only mildly degraded to R_λ = 25.","tokens_in":27695,"tokens_out":1226,"duration_ms":9531,"significance":"If the thresholds hold under the authors’ SNR definition, the result is immediately useful for LSST-era follow-up strategy and for the design of low-cost classification spectrographs. The work is the first systematic R_λ–SNR grid for a refined SN taxonomy, is fully reproducible (public GitHub repository with notebooks that regenerate the figures), and employs careful leakage control and macro-F1 reporting that correctly handle class imbalance. These strengths make the paper a concrete reference for observers and instrument builders even if the precise numerical thresholds later shift under alternative SNR conventions.","major_comments":[{"comment":"The central numerical claim (no loss to R_λ=50 / SNR=5; only minimal impact to R_λ=25) is defined exclusively by the custom, subtype- and phase-specific line SNR of §3.2–3.3 and Table 3. S is the width-normalized area of one hand-chosen diagnostic feature after Gaussian smoothing whose σ was manually tuned per spectrum; ~10 % of the library (190+179 spectra) was discarded because the procedure failed or produced outliers. At low R the fixed shoulder wavelengths become unreliable (authors note this explicitly), yet the grid still reports “SNR=5”. Because every cell of Figs. 6–7 is generated from this definition, any systematic bias in how it ranks classification difficulty across subtypes or resolutions directly shifts the claimed thresholds. The paper should either (i) re-run a subset of the grid with at least one alternative SNR estimator (e.g., continuum rms in a fixed line-free window","section":null},{"comment":"The noise model used to reach target SNR (§3.3) is additive white Gaussian noise scaled to the measured S and added to the extracted signal. Real SN spectra are dominated by Poisson statistics, wavelength-dependent sky, host-galaxy continuum, and residual tellurics. While the idealized model is a reasonable first step, the paper should quantify (or at least discuss with a small controlled experiment) whether the performance cliff moves when more realistic noise is injected. Without that check, the claim that “SNR=5 is sufficient” risks being optimistic for actual observing conditions.","section":null}],"minor_comments":[{"comment":"Abstract and §5 state “no loss … down to R_λ=50 and SNR=5” while Fig. 6 already shows a few-percent drop at R=50 for several SNR rows; “no statistically significant loss” or “within the scatter of the original-SNR row” would be more precise.","section":null},{"comment":"Table 2 and the appendix tables list removed spectra, but the text never states the final number of unique SNe remaining after both culling steps; a single sentence would help readers assess residual class imbalance.","section":null},{"comment":"Fig. 1 caption says spectra were normalized to [0,1] for display while ABC-SN trains on standardized data; a brief reminder in the main text would avoid confusion when comparing panels.","section":null},{"comment":"The wavelength cut 4500–7000 Å excludes the O I 7774 and Ca NIR triplet that are often decisive for late-time SESNe; a short paragraph on how this restriction may affect the low-R thresholds would strengthen the discussion.","section":null},{"comment":"Minor typos: “W e” in the title page, “diï¬cult” throughout, and “for arbitrary \\SNR{}” in the abstract.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a solid, reproducible methods contribution that belongs in the journal. The SNR-definition issue is real but fixable with either a short robustness test or clearer language; I do not see it as grounds for major revision or rejection. The idealized noise model is a secondary concern that can be addressed in discussion. I would accept after the authors respond to the two major points."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: with ABC-SN on a carefully degraded SNID-derived library, refined ten-class SN typing (including SESN subtypes) holds with essentially no macro-F1 loss down to R_\\lambda=50 and their SNR=5, and only mild degradation to R=25. That is the first dense, controlled grid of this kind, and it is immediately actionable for LSST/Argus follow-up strategy and for people designing cheap low-R spectrographs.\n\nWhat they did well is the experimental hygiene. They retrain from scratch on 476 homogeneous datasets (16 SNR × 14 R, plus original-SNR controls), keep whole SNe out of leakage via 2-fold CV, report macro-F1, document the ~10% culls, and ship the notebooks. The degradation operators themselves (Gaussian convolution for R, additive white noise scaled to measured S) do not bake in the labels. The heat-maps in Figs. 6–7 are clear, and the result is not just a restatement of SEDM folklore.\n\nThe soft spot is real but proportionate: the quoted numbers live entirely inside their custom SNR (Eq. 2 + Table 3). S is the width-normalized area of one hand-chosen, phase-dependent diagnostic line after per-spectrum manual Gaussian smoothing; shoulders are fixed from the high-R originals and become less trustworthy at low R (they say so). Roughly 369 spectra were removed because the procedure failed or produced outliers. The noise model is idealized relative to real Poisson+sky+host. So “SNR=5” is not a universal lab number; it is their definition. That does not invalidate the relative surface or the practical message that classification survives surprisingly low R and modest S/N, but anyone quoting the absolute thresholds should re-measure under their own SNR convention.\n\nTaxonomy is restricted (no IIn, no SLSNe, etc.) and only one classifier is tested; both are acknowledged. Citations and math look clean; code is public.\n\nThis is for instrument designers, survey strategists, and anyone allocating scarce spectroscopic time. It deserves a serious referee. I would cite the thresholds (with the SNR caveat) and I would accept it for peer review.","headline":"Solid empirical map of SN subtype classification vs R and SNR; the R=50/SNR=5 thresholds are usable but rest on a custom, hand-curated line SNR that is the main soft spot.","tokens_in":28296,"tokens_out":568,"would_cite":true,"duration_ms":6049,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Supernova subtype classification stays accurate down to spectral resolution Rλ=50 and SNR=5, and is only mildly hurt even at Rλ=25.","keywords":["supernova classification","spectral resolution","signal-to-noise ratio","LSST follow-up","stripped-envelope supernovae","deep learning","spectroscopic requirements","ABC-SN"],"falsifier":"Retrain the same classifier on an independent library that already has published uncertainty arrays (so SNR can be measured without manual Gaussian smoothing) and check whether the macro-F1 still stays flat down to Rλ=50 and SNR=5.","tokens_in":28285,"feed_emoji":"🔭","tokens_out":857,"duration_ms":7458,"temperature":0.7,"pith_summary":"Upcoming sky surveys will find millions of supernovae, far more than existing spectrographs can follow up at high quality. This paper measures how far spectral resolution and signal-to-noise can be lowered before an automated classifier can no longer tell the main supernova subtypes apart. Using a curated library of spectra and a deep-learning model, the authors create homogeneous datasets across many combinations of resolution and noise, then retrain and score the classifier on each one. They find that a refined taxonomy (including subtypes of stripped-envelope events) remains fully usable down to Rλ=50 and SNR=5, with only modest losses at Rλ=25. The practical payoff is clear: observers can free high-resolution instruments for rare or high-value targets, and smaller facilities can still contribute useful classifications by accepting lower resolution or shorter exposures.","feed_headline":"SN subtypes classifiable at R=50 and SNR=5","feed_subtitle":"Low-resolution, low-SNR spectra still separate refined supernova types, freeing big telescopes for rarer targets.","key_machinery":"A subtype- and phase-specific SNR definition that measures signal from a single emblematic spectral feature (e.g., Si II λ6355 for early Ia, He I λ5876 for Ib/Ibn) relative to a local pseudo-continuum, then injects controlled Gaussian noise and Gaussian-convolves the spectrum to any target Rλ before retraining the ABC-SN attention classifier.","core_discovery":"Classification of supernova spectra into a refined ten-subtype taxonomy is possible at low resolution and low SNR with no loss of model performance down to Rλ=50 and SNR=5; performance is only minimally reduced even at Rλ=25, and degrades rapidly only below Rλ≈20.","pith_inferences":["The same floor may apply to other modern spectral classifiers, not only the attention model used here, because the information content of the lines themselves is what is being degraded.","Including rare classes such as IIn (narrow lines) would likely push the useful resolution floor higher, so the present numbers are best read as optimistic for the included taxonomy.","A production pipeline that trains on mixed real-world SNR and resolution rather than uniform synthetic grids may still inherit the same practical thresholds if the median of the training set sits near SNR~20."],"forward_implications":["Spectrographs can deliberately trade resolution or exposure time for classification work without losing refined subtype purity.","High-resolution instruments can be reserved for detailed follow-up of rare or scientifically critical events rather than routine typing.","Smaller telescopes and lower-cost spectrographs become viable partners for LSST-scale classification campaigns.","Instrument designers can target Rλ~50 as a practical floor for classification-mode modes rather than pushing for higher resolving power.","Survey planners can set exposure-time calculators knowing that SNR~5 is already sufficient under this taxonomy."],"fun_headline_variants":["SN subtypes hold at R=50 SNR=5 with no model loss","Refined 10-type SN taxonomy works to Rλ=50 SNR=5","Low-res low-SNR spectra sort SN subtypes to R=50","Performance intact at R=50 SNR=5; slight drop at 25","Min specs for SN subtype class: Rλ=50 and SNR=5"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The custom line-based SNR measure, which needed heavy manual smoothing choices and the removal of roughly a tenth of the spectra, is assumed to give a fair, homogeneous ranking of classification difficulty across every subtype and every resolution.","fun_headline_variants_meta":{"raw":{"variants":["SN subtypes hold at R=50 SNR=5 with no model loss","Refined 10-type SN taxonomy works to Rλ=50 SNR=5","Low-res low-SNR spectra sort SN subtypes to R=50","Performance intact at R=50 SNR=5; slight drop at 25","Min specs for SN subtype class: Rλ=50 and SNR=5"]},"model":"grok-4.5","effort":"low","cost_usd":0.004238,"raw_usage":{"total_tokens":1337,"prompt_tokens":850,"num_sources_used":0,"completion_tokens":84,"cost_in_usd_ticks":42380000,"prompt_tokens_details":{"text_tokens":850,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":403,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":850,"tokens_out":84,"duration_ms":3491,"temperature":1.0,"reasoning_tokens":403,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T01:48:40.121488+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Retrain the same classifier on an independent library that already has published uncertainty arrays (so SNR can be measured without manual Gaussian smoothing) and check whether the macro-F1 still stays flat down to Rλ=50 and SNR=5.","supporting_citations":[],"review_version":1}