{"id":"5c13623b-87fb-4cce-8171-4019df0eac2d","arxiv_id":"1908.06190","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A few well-timed VLA observations can distinguish relativistic, CSM-interacting, and off-axis GRB supernovae, but only if real light curves resemble the small template bank used in the simulations.","lead":"This paper simulates radio follow-up of stripped-envelope supernovae to find the fewest, best-timed VLA observations for telling apart relativistic, CSM-interacting, and off-axis GRB explosions. It also reports a new late-time radio upper limit for iPTF17cw and estimates when X-ray follow-up would work.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"In-sample template matching likely inflates the 97% association rate; a leave-one-out test is needed to show the optimized cadence generalizes beyond the four relativistic SN templates.","rationale":"I agree with the reader's CONDITIONAL verdict. The weakest point is not the algebra or the simulation mechanics; it is that the evaluation protocol cannot distinguish 'the cadence works for the templates we happened to choose' from 'the cadence works for the population.' The 97% figure is computed by asking how often the exact generating template is uniquely recovered when the bank contains that template. Any real event will not be an exact template, so this is an upper bound unless the templates are representative. The paper acknowledges limitations for CSM-interacting SNe, but the relativistic-SNe headline inherits the same issue. A leave-one-out test is a natural, low-cost check that would quantify the overfitting. The new iPTF17cw upper limit is interesting but does not validate the strategy. Since the reader already requested conditional acceptance and our concern reinforces that rather than moving it, the verdict should remain UNCHANGED relative to the reader's CONDITIONAL.","tokens_in":14602,"tokens_out":7257,"duration_ms":74535,"concrete_test":"Leave-one-out cross-validation: for each of the four relativistic templates, re-run the Section 3.3 optimization using only the other three relativistic templates (keeping the same CSM and off-axis GRB contaminants), then measure the fraction of unique correct associations for the held-out template at z=0.01. If the leave-one-out efficiency drops by more than ~10 percentage points relative to the 97% in Table 2, the optimized cadence is tuned to the specific templates; if it remains near 97%, the strategy is robust to template choice. Also repeat with the early-time extrapolations of Section 4.1.1 to separate the effect of the extrapolation assumption from template-bank overfitting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline efficiency (97% at z=0.01 in Table 2, Section 4.1) is an in-sample template-retrieval score, not a measured out-of-sample classification rate. In the Monte Carlo procedure (Section 3.3), each simulated light curve is drawn from one of the four relativistic templates in the bank and is scored as 'correct' only if that exact generating template becomes the unique 3-sigma match. The optimized epochs are then selected on these same simulated data, so there is no penalty for overfitting to the particular shapes of SN 1998bw, SN 2006aj, SN 2009bb, and iPTF17cw. The bank is extremely small and contains mostly extreme, well-studied events; the true stripped-envelope SN population will include light curves with intermediate peak luminosities, different rise times, and early epochs that this work treats as non-detections before the first actual observation (Section 3.1). If a real event is not well represented by one of these four templates, the optimized cadence may fail to produce a unique correct association, making the reported 97% an optimistic upper bound for real follow-up campaigns. The paper is transparent about the small CSM sample (Section 4.2) but applies the same in-sample logic to the relativistic-SN headline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a Monte Carlo simulation framework to optimize radio follow-up observations of stripped-envelope core-collapse SNe with the VLA. The authors simulate light curves from template radio observations of known relativistic SNe (SN 1998bw, SN 2006aj, SN 2009bb, iPTF17cw), CSM-interacting SNe (PTF11qcj, SN2007bg), and BOXFIT models of off-axis GRBs, then determine the minimum number and timing of radio epochs that maximize the fraction of unique and correct associations between an observed target and its generating template. For relativistic SNe at z=0.01, they report that five observations at approximately 2, 8, 18, and 30 days after the first radio observation yield 97% efficiency. They also present a new late-time VLA upper limit on iPTF17cw and estimate X-ray detectability of the same sources. The paper includes a discussion of the importance of early-time observations and of the limited CSM-interacting sample.","tokens_in":14903,"tokens_out":6494,"duration_ms":58630,"significance":"If the reported efficiencies were robust, the proposed cadence would be a practical guideline for upcoming transient surveys (ZTF, LSST) and for VLA/ngVLA follow-up, using only 5 epochs. The simulation strategy is transparent, uses large Monte Carlo sets (10,000 realizations), and the authors explicitly test the effect of early-time extrapolations. The new upper limit on iPTF17cw is a modest but useful observational contribution. However, the central quantitative claims are currently based on in-sample template retrieval, which means the headline efficiencies are likely upper bounds. With a leave-one-out validation or appropriately softened claims, the paper would offer a useful planning tool, but in its present form the significance of the specific numbers is unclear.","major_comments":[{"comment":"The definition of a 'correct' association is that the unique matching template is 'the same template/model from which the observations were simulated' (Section 3.3). Since the simulated light curves are generated from the same templates that form the classification bank, every simulated event is, by construction, a member of the bank. The reported efficiencies (Table 2, e.g., 97% at z=0.01) are therefore retrieval rates from a closed set of known templates, not classification rates for real events. A real stripped-envelope SN with a light curve that is not well represented by one of the four relativistic templates (or two CSM templates) could be uniquely matched to a wrong template, and such failures are absent from the simulation. I recommend either (a) a leave-one-out test in which each template is excluded from the bank and then used to generate simulated targets, or (b) an explicit statement that the quoted efficiencies are conditional on the template bank being representative of the true population, and should be regarded as upper bounds. This point is load-bearing because the paper's abstract and conclusions present these efficiencies as the basis for the proposed follow-up cadence.","section":"Section 3.3, Table 2"},{"comment":"The relativistic SN template bank comprises only SN 1998bw, SN 2006aj, SN 2009bb, and iPTF17cw. These are extreme events: two are very bright nearby explosions, one is very faint and rapidly fading, and one is at higher redshift with sparse early-time sampling. The simulation assumes linear interpolation/extrapolation of fluxes (Section 3.1), including extrapolation of early-time behavior from other SNe in Section 4.1.1, but there is no test of how the optimized epochs perform for light curves with, e.g., intermediate peak luminosities, different rise times, or different spectral indices. The CSM-interacting class is even more limited (only PTF11qcj and SN2007bg; the authors acknowledge this in Section 4.2). Because the proposed cadence (Table 2) is selected to maximize discrimination among these specific templates, it may not generalize to the broader stripped-envelope SN population. I would like to see a sensitivity study with perturbed template shapes (e.g., varying peak luminosity, rise time, decay rate) to demonstrate that the optimized cadence is not tailored to the individual events.","section":"Section 3.1, Table 1"},{"comment":"The epoch-selection procedure optimizes the delays M2, M3, ... by maximizing the number of unique and correct associations computed on the same Monte Carlo realizations that are later used to evaluate the strategy. There is no separation between the data used to choose the epochs and the data used to measure the efficiency; thus the reported percentages are in-sample and expected to be optimistically biased. A cross-validation approach, in which the 10000 realizations are split into a training set for epoch selection and a test set for evaluation, would quantify the degree of optimism. Without this, the reader cannot tell how much of the 97% is a genuine property of the cadence and how much is overfitting to the particular realizations and templates.","section":"Section 3.3"}],"minor_comments":[{"comment":"There is a typo in the reference line for SN 2009bb: 'Strauss et al. (1992) anf Soderberg et al. (2010)' should read 'and'.","section":"Table 1"},{"comment":"The caption contains the typo 'campiagn'; it should be 'campaign'.","section":"Figure 2"},{"comment":"The source is referred to inconsistently as both 'iPTF17cw' and 'PTF 2017cw'; the standard naming 'iPTF 17cw' should be used throughout.","section":"Sections 2 and 4"},{"comment":"The notation for time delays is not fully consistent: the text defines 'Delta Tn = tn - tradio,0 = Mn x 2 d' but later uses 'Delta t2 = M2 x 2 d' and 'Delta t3'; harmonize the notation for readability.","section":"Section 3.3"},{"comment":"The spectral index convention F_nu proportional to nu^{-beta} is stated earlier for SN 2006aj but not near Eq. (1); it should be explicitly restated when the radio-to-X-ray spectral index is introduced.","section":"Section 5, Eq. (1)"},{"comment":"The efficiencies are computed after excluding non-detectable targets, which is a reasonable choice; however, it would also be informative to report the absolute fraction of all simulated targets that are correctly associated (including non-detections), since in a real follow-up one does not know detectability a priori.","section":"Sections 4.1-4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for an astro-ph.HE journal and the observational upper limit is a small but useful addition. The main concern is the in-sample nature of the efficiency claims; if the authors can add a leave-one-out analysis, a cross-validation of the epoch optimization, or substantially soften the claims, the paper could be acceptable. I would not reject on the basis of the template-representativeness issue alone, as the method itself is a reasonable planning tool, but the current wording overstates the reliability of the proposed cadence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a legitimate, useful extension of the authors' own earlier optimization work, and the optimized epoch tables are the kind of thing a PI planning VLA time would actually want. The headline 97% association rate should be read as an upper bound, not a measured field performance, because the simulated targets are drawn from the same template bank the classifier matches against. The paper is transparent about some of this but not all of it.\n\nWhat's new: application of the Carbone & Corsi method to stripped-envelope SNe; a genuinely new radio upper limit for iPTF17cw (late-time 3σ limits of 10 µJy at 5 GHz, 15 µJy at 3 GHz) that is consistent with the 1998bw-like decay; optimized cadence tables for three target classes; and X-ray detectability calculations that give some useful distance horizons. The Monte Carlo is straightforward and large (10,000 realizations per target), and the authors explicitly test high/medium/low urgency strategies. The early-time extrapolation section (4.1.1) is a nice touch—it addresses a real gap and shows the importance of early observations.\n\nSoft spots, in proportion: The main one is the in-sample nature of the efficiency numbers. The simulated observations are generated by adding noise to the same templates that form the association bank, so 'unique and correct association' is a retrieval test, not a generalization test. With only four relativistic templates, two CSM templates, and a grid of off-axis GRB models, the bank covers a tiny slice of the real stripped-envelope population. The authors acknowledge the CSM sample limitation but don't note this same concern for the relativistic SNe headline. A leave-one-out test—train on three templates, test on the fourth—would tell you how much the cadence depends on the exact shapes, and it's cheap to run. Also, the early-time treatment before the first detection is either non-detection or an extrapolation borrowed from a similar SN; both are reasonable, but they bracket rather than eliminate the uncertainty. The X-ray section is more of a rough estimate; the constant spectral index assumption is fine for a horizon estimate, but I wouldn't put much weight on the absolute numbers.\n\nThe citation pattern looks honest; this is a clear incremental step in their own program, not a repackaging. No invented entities, no load-bearing math errors that I could see.\n\nWho it's for: people planning radio follow-up of rare stripped-envelope SNe in the ZTF/LSST era. It deserves serious peer review, but I'd push for the leave-one-out test or at least a sentence in the abstract that the efficiencies are relative to the template bank. Recommend: send to review, conditional on that.","headline":"A practical, clearly-written extension of the Carbone & Corsi optimization framework to stripped-envelope SNe, with real new data (iPTF17cw upper limit) but classification efficiencies that are probably optimistic because they are in-sample retrieval rates.","tokens_in":15409,"tokens_out":2589,"would_cite":true,"duration_ms":22390,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Five radio observations, timed 2, 8, 18, and 30 days after the first detection, identify about 97% of nearby detectable relativistic supernovae.","keywords":["stripped-envelope supernovae","core-collapse supernovae","radio follow-up","relativistic supernovae","circumstellar medium","off-axis gamma-ray bursts","Very Large Array","Monte Carlo optimization"],"falsifier":"Apply the proposed cadence to a radio-detected stripped-envelope supernova of spectroscopically known subtype and compare the classification: a single event whose observed flux pattern is uniquely matched to the wrong template, or that is not detected at an epoch where its template predicts a clear detection, would show the 97% figure does not generalize.","tokens_in":14385,"feed_emoji":"📡","tokens_out":8855,"duration_ms":79422,"temperature":0.7,"pith_summary":"Stripped-envelope core-collapse supernovae include some of the rarest explosions known: those that launch relativistic jets, those seen slightly off the jet axis, and those that shock a dense shell of circumstellar gas. Radio emission tracks the fastest ejecta, but each type has a different radio rise-and-fall, so the paper asks how few radio observations can tell them apart. Simulating ten thousand light curves per template with VLA-like sensitivity, it finds that five observations placed about 2, 8, 18, and 30 days after the first radio detection uniquely and correctly identify 97% of detectable relativistic supernovae at $z=0.01$, and that CSM-interacting SNe need only three epochs. The value of the result is practical: it converts limited telescope time into reliable classification of events that optical surveys will soon find in large numbers.","feed_headline":"Five radio snapshots classify most relativistic supernovae","feed_subtitle":"A compact VLA cadence of 2, 8, 18, and 30 days separates engine-driven explosions from look-alikes.","key_machinery":"The load-bearing object is the template bank and the Monte Carlo classifier built around it: four observed relativistic SNe, two observed CSM-interacting SNe, and BOXFIT model light curves for off-axis long GRBs. For each simulated target the classifier matches noisy synthetic observations against this bank, counts a source as identified only when exactly one template fits all epochs within $3\\sigma$ and that template is the true one, then greedily picks the next observation delay (in two-day steps, up to ten epochs) that maximizes the number of such unique and correct associations. The early-time behavior of two templates is extrapolated from other SNe to test how much the recommended cadence depends on the rising part of the light curve.","core_discovery":"The paper's central claim is that a fixed, small number of radio epochs is enough to classify the explosion type of radio-emitting stripped-envelope SNe, provided the first observation is early and the cadence is chosen by simulation rather than by habit. With the sensitivity of a two-hour VLA observation, five epochs at delays of 2, 8, 18, and 30 days after the first detection identify $97\\%$ of the detectable relativistic (engine-driven) targets at $z=0.01$, while avoiding confusion with CSM-interacting supernovae and off-axis GRBs; at $z=0.1$ the same strategy identifies $78\\%$ of detectable targets. For CSM-interacting SNe, a low-urgency first observation plus delays of about 6 and 90 days classifies all detectable simulated sources. For off-axis GRBs, five epochs with a high-urgency start give $99\\%$ efficiency at $z=0.01$ for detectable sources. The paper also reports a new VLA upper limit on the late-time radio emission of iPTF 17cw and projects X-ray detectability under a synchrotron radio-to-X-ray extrapolation.","pith_inferences":["Inference: The same optimization machinery could be re-run for a next-generation array with roughly ten times the VLA sensitivity; the paper notes such an instrument would reach about three times farther and increase detections by a factor of order thirty, but it does not simulate its optimal cadence.","Inference: Because the recommended epochs are derived from a small set of bright historical events, the 2-8-18-30-day pattern should be read as a starting point rather than a law; the paper's own early-time extrapolation exercise shows how much the cadence moves when the rising part of the light curve is filled in.","Inference: Joint radio-X-ray scheduling is a natural extension of the paper's approach; a single campaign that picks radio epochs while also predicting X-ray visibility could break the degeneracy between environment density and the fraction of shock energy in magnetic fields."],"forward_implications":["A relativistic supernova needs its first radio observation within roughly 1 hour to 2 days of optical discovery; waiting a week or more lowers the identifiable fraction noticeably.","The five-epoch campaign for relativistic SNe at $z=0.01$ costs about ten hours of VLA time per target, so a season of follow-up can cover a sizable sample rather than one or two objects.","For CSM-interacting SNe the cadence is much slower: a first observation followed by delays near 6 and 90 days identifies every detectable simulated source, so these events do not force an urgent telescope response.","Off-axis GRB afterglows with isotropic energy above $10^{51}$ erg are detected and uniquely identified near 100% of the time at $z=0.01$, while low-energy models below $10^{49}$ erg are mostly undetectable.","If the radio-to-X-ray spectral index is about 0.7, Chandra-class 20 ks exposures should see the X-ray counterparts of nearby ($z=0.01$) relativistic and CSM-interacting SNe for hundreds of days."],"supporting_citations":[{"why":"Supplies the simulation and optimization method that this work extends to stripped-envelope supernovae.","marker":"Carbone & Corsi 2018"},{"why":"Provides the template radio light curve of SN 1998bw, the bright benchmark relativistic supernova.","marker":"Kulkarni et al. 1998"},{"why":"Provides the template radio light curve of SN 2009bb, a relativistic engine-driven supernova without a detected GRB.","marker":"Soderberg et al. 2010"},{"why":"Provides the template radio light curve of SN 2006aj, a faint and fast-fading relativistic supernova.","marker":"Soderberg et al. 2006c"},{"why":"Provides the template radio light curve of iPTF 17cw and the context for the new late-time upper limit.","marker":"Corsi et al. 2017"},{"why":"Provides the template radio light curve of PTF 11qcj for the CSM-interacting supernova class.","marker":"Corsi et al. 2014"},{"why":"Provides the template radio light curve of SN 2007bg for the CSM-interacting supernova class.","marker":"Salas et al. 2013"},{"why":"Supplies the BOXFIT models used to generate off-axis long GRB radio light curves.","marker":"van Eerten et al. 2012"}],"fun_headline_variants":["Five VLA looks tell jet supernovae from impostors","Radio cadence of 2-8-18-30 days decodes explosions","Optimal radio schedule identifies rarest supernova types","Compact VLA plan classifies relativistic explosions","Five radio epochs unmask engine-driven supernovae"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"All efficiencies rest on the template light curves being representative of the true stripped-envelope supernova population; where early-time data are missing, early rises are borrowed from other SNe, so a real event that rises or fades outside the template range could break the optimized cadence.","fun_headline_variants_meta":{"raw":{"variants":["Five VLA looks tell jet supernovae from impostors","Radio cadence of 2-8-18-30 days decodes explosions","Optimal radio schedule identifies rarest supernova types","Compact VLA plan classifies relativistic explosions","Five radio epochs unmask engine-driven supernovae"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00097,"raw_usage":{"total_tokens":4165,"prompt_tokens":1023,"completion_tokens":3142,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":639,"completion_tokens_details":{"reasoning_tokens":3060}},"tokens_in":639,"tokens_out":3142,"duration_ms":25736,"temperature":1.0,"reasoning_tokens":3060,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:53:38.527673+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the proposed cadence to a radio-detected stripped-envelope supernova of spectroscopically known subtype and compare the classification: a single event whose observed flux pattern is uniquely matched to the wrong template, or that is not detected at an epoch where its template predicts a clear detection, would show the 97% figure does not generalize.","supporting_citations":[{"cited_title":"2018, ApJ, 867, 135","cited_arxiv_id":null,"evidence_quote":"Supplies the simulation and optimization method that this work extends to stripped-envelope supernovae."},{"cited_title":"R., Frail, D","cited_arxiv_id":null,"evidence_quote":"Provides the template radio light curve of SN 1998bw, the bright benchmark relativistic supernova."},{"cited_title":"M., Chakraborti, S., Pignata, G., et al","cited_arxiv_id":null,"evidence_quote":"Provides the template radio light curve of SN 2009bb, a relativistic engine-driven supernova without a detected GRB."},{"cited_title":"B., Kasliwal, M","cited_arxiv_id":null,"evidence_quote":"Provides the template radio light curve of iPTF 17cw and the context for the new late-time upper limit."}],"review_version":1}