{"id":"094891b4-b523-48da-9fcc-225dcd8ed886","arxiv_id":"2501.03929","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A decade-long infrared survey of Cygnus-X reveals 68 candidate eruptive young stars, with large outbursts roughly ten times more frequent among embedded Class I protostars than among more evolved Class II objects.","lead":"Using ten years of infrared data from NASA's NEOWISE satellite, astronomers found 68 young stars in the Cygnus-X region that appear to undergo sudden, powerful brightening events, including 14 that resemble a rare class of outburst called FUor. The study suggests such large eruptions are far more common in the youngest, most embedded protostars than in older, more evolved young stars.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'order of magnitude' FUor-incidence claim is never derived; raw counts (13 vs 1) have a 95% Poisson lower bound below 10, and the two samples use asymmetric selection and discovery channels.","rationale":"Agree with the reader that selection functions are the principal confound, but the more immediate, internally checkable problem is that the abstract's central quantitative claim is not derived. The reader's weakest_assumption identifies the selection mismatch; this stress-test goes one step further and shows that even if selection were perfect, the claimed 'order of magnitude' is not secured by the numbers in the paper. The small-number issue is concrete: with 13 vs 1 events, the exact Poisson lower bound straddles 10. The paper deserves credit for releasing light curves, providing a useful candidate catalogue, and acknowledging limitations (e.g. the RMSE cut failing two genuine targets and the fitting method being 'of questionable utility' in Appendix A). These admissions do not, however, substitute for a rate calculation. The suggested test would settle whether the headline claim should be softened; if the exact CI remains entirely above 10 after removing UGPS-only discoveries, the claim could stand with qualification. The verdict remains CONDITIONAL pending such a revision, so no change from the reader's assessment is recommended.","tokens_in":28294,"tokens_out":9021,"duration_ms":87358,"concrete_test":"Compute exact Poisson 95% confidence intervals for the FUor-candidate incidence ratio from Table 4, with N=1332 and N=4935, both including and excluding the two UGPS-discovered sources (362, 1964). If the lower bound is <10, or if excluding UGPS sources drops the point estimate below ~10, the 'order of magnitude' wording must be revised. As a second check, recompute the high-amplitude incidence after splitting both samples by a uniform spectral-index α criterion and after restricting both to sources with (or without) a [24] detection, to test whether selection, not stage, drives the excess.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's headline claim that candidate FUors are 'an order of magnitude more common' among Class I systems is never quantified in the body. Section 4.3 compares only the incidence of ΔW1 > 1 mag variability (22.46% vs 4.94%); no FUor-specific rate, denominator, or uncertainty is presented. Using the paper's own sample sizes and Table 4, the raw counts are 13 FUor candidates among 1332 [24]-selected sources versus 1 among 4935 SPICY sources, a crude rate ratio of ~48. A Poisson 95% confidence interval for this ratio is roughly (6, 370): the lower bound is below 10, so 'order of magnitude' is not supported at standard significance even before selection effects. Selection makes this worse: the embedded sample was supplemented by a two-epoch UGPS search (Sources 362 and 1964) with no equivalent for the SPICY sample, and the [24]-detection selection criterion is physically tied to envelope luminosity. The claim therefore rests on an unperformed rate calculation, small counts, and asymmetric discovery channels.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a decade-long NEOWISE mid-infrared variability study of Spitzer-selected YSOs in Cygnus-X, comparing an embedded Class I sample (1332 sources with MIPS [24] detections from Kryukova et al. 2014) with a sample of more evolved flat-spectrum/Class II/III YSOs (4935 sources from SPICY without [24] detections). The authors identify 68 candidate eruptive variables (33 in the embedded sample and 35 in the SPICY sample), of which 14 are classified as likely FUor candidates and 20 as less-certain candidates. They report that candidate FUors are an order of magnitude more common among Class I systems, that embedded Class I sources have a higher incidence of high-amplitude (ΔW1 > 1 mag) variability than Class II sources (22.46% vs 4.94%), and that many short-duration eruptive YSOs become redder when brighter. The paper includes light curves of individual sources, spectra of two objects, structure-function analysis, and a public data release of the light curves.","tokens_in":28484,"tokens_out":3742,"duration_ms":38244,"significance":"If the central claims hold, the paper roughly doubles the available set of photometric FUor-like candidates and provides a statistical stage-dependence for large-amplitude MIR variability, with implications for the protostellar luminosity spread problem and for accretion-instability models. The study has notable strengths: the catalogue is built from public NEOWISE photometry with transparent cleaning cuts, the light curves and fitting code (aptare) are publicly released, the authors explicitly flag uncertain candidates, and they candidly acknowledge the limited utility of their template-fitting method. The comparison of amplitude distributions using structure functions and the redder-when-brighter colour behaviour of many SDEs are interesting observational results. However, the headline 'order of magnitude' claim for FUor incidence is not quantified in the body, and the rate comparison is affected by small counts and asymmetric discovery channels.","major_comments":[{"comment":"The claim that candidate FUors are 'an order of magnitude more common' among Class I systems is never derived in the body. Using Table 4 and the sample sizes in §2.1, the raw counts are 13 FUor candidates among 1332 [24]-selected sources versus 1 among 4935 SPICY sources, a crude rate ratio of about 48. A Poisson 95% confidence interval for this ratio is roughly (6, 370), whose lower bound is below 10, so the 'order of magnitude' statement is not supported at standard significance. The paper should either present an explicit rate comparison with confidence intervals and a significance test, or soften the claim to reflect the large uncertainty.","section":"Abstract and §4.3"},{"comment":"The discovery channels are asymmetric between the two samples. Sources 362 and 1964 were found via a two-epoch UGPS search that was applied only to the Kryukova [24]-selected sample, with no equivalent search of the SPICY sample. These two sources constitute 2 of the 13 embedded FUor candidates, so the rate comparison is inflated by an un-matched detection method. Additionally, the [24]-detection selection criterion is physically tied to envelope luminosity, while the SPICY sample is defined by the complement of that sample; the two samples may therefore differ in luminosity and extinction regimes as well as in evolutionary stage. The rate comparison should be repeated excluding the UGPS-discovered sources and with a discussion of how selection functions affect the result.","section":"§2.1 and §3.1"},{"comment":"The statistical support for the 'higher incidence' claim is marginal. The Mann-Whitney U test is performed on two equal-sized samples resampled from Gaussian KDEs of the amplitude distributions, yielding an average p-value of 3.46%; this is borderline, and the resampling procedure does not propagate the covariance between the samples. The headline percentages 22.46% versus 4.94% are quoted without confidence intervals. I recommend adding bootstrap confidence intervals for the two proportions, a direct two-sample test on the amplitude distributions, and a sensitivity test with respect to the ΔW1 > 1 mag threshold and the adopted completeness criteria.","section":"§4.3"},{"comment":"The candidate counts are internally inconsistent. The abstract and §3 report 68 candidate EVs with 14 FUor-like sources and 20 less-certain objects; §5 says 'we discovered 14 eruptive sources' that are FUor candidates, but also mentions 'up to 16 candidate members' and 'six other sources' that are plausible FUor candidates. Table 4 lists 14 'likely' candidates and 6 'potential' candidates. The numbers 13/1, 14/6, and 16 are used in different places without reconciliation. The authors should define a single, consistent hierarchy of candidate classes and make the sample counts in the abstract, body, tables, and summary agree.","section":"§3, §4.2, §5, and Table 4"}],"minor_comments":[{"comment":"The RMSE threshold is inconsistent: §2.1 states a cut-off of RMSE > 0.1 (10% of the normalized amplitude), while Appendix A states that candidates were marked if the reduced mean squared error was less than 0.15. Please unify the description.","section":"§2.1 and Appendix A"},{"comment":"Cross-references to tables in the text read 'Table 2.2' and 'Table 3.1.5', which do not correspond to the numbered tables; these should be corrected to the actual table numbers.","section":"Table 2 and Table 3"},{"comment":"The number of stars in each duration category is described differently in the text (e.g., '11 long-duration, 8 intermediate-duration, and 11 short-duration' in §3.1) and in Table 1; please ensure the counts and category labels in tables and text are fully consistent.","section":"§3.1 and §3.2"},{"comment":"The Kendall tau values for the ΔW2–colour correlation are reported without uncertainties or sample sizes; adding these would let the reader judge the claimed difference between the NIR-selected and MIR-selected samples.","section":"§4.1"},{"comment":"The reference list uses both 'Kryukova et al. 2014a' and 'Kryukova et al. 2014b' for the same work in different places; unify these citations.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest and data-rich, and the central observational catalogue is a useful contribution. My main concern is that the abstract's headline claim goes beyond what the body establishes: the FUor incidence ratio is not computed, the counts are small, and the discovery channels are asymmetric. These issues are fixable by adding a proper rate calculation with uncertainties and by softening or conditioning the claim. I would not reject the paper, but it needs a substantial revision of the statistical presentation before the claims can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful candidate-finding paper whose abstract oversells one number. The NEOWISE light-curve catalogue, the 14 FUor-like candidates, and the slow-rising Class I sources are new and worth having; the 'order of magnitude' FUor statement is not derived anywhere in the body.\n\nWhat it does well: the sample construction is transparent, and the authors are properly cautious throughout—they flag uncertain candidates, admit the FUor template fit is 'of questionable utility,' and even note that two likely genuine eruptive sources failed the RMSE cut. They also make the light curves available and the aptare code is public, so the fitting is reproducible. The redder-when-brighter trend for short-duration events and the five-to-seven-year rising sources are genuinely new observations, and the comparison of ΔW1 > 1 mag incidence (22.46% vs 4.94%) is backed by a MW-U test and structure functions. For the field, the candidate list itself is the contribution.\n\nSoft spots, in proportion: the abstract's strongest claim—FUors an order of magnitude more common among Class I—is never calculated. Section 4.3 only compares general high-amplitude variability rates; no FUor-specific denominator, rate, or uncertainty appears. Using the paper's own counts (13 FUor candidates among 1332 embedded sources vs 1 among 4935 SPICY sources) the crude ratio is ~48, but a Poisson 95% interval runs roughly 6–370, so 'order of magnitude' is not established. The selection asymmetry makes it worse: the embedded sample gets an extra two-epoch UGPS search with no SPICY equivalent, and the [24]-selection is physically tied to envelope luminosity. The RMSE threshold is also inconsistent (0.1 in Section 2, 0.15 in Appendix A), and there is no systematic contamination estimate from evolved stars or dippers. These are fixable problems, not fatal ones—the body already acknowledges most of them. The paper never claims more than candidate status for most sources, which is the right register.\n\nWho it's for: people working on eruptive YSOs, FUor statistics, and protostellar variability. It deserves a serious referee; an editor should send it out. The revision should add an explicit FUor rate with uncertainties, reconcile the threshold, and discuss selection biases. If they do that, this is a solid MNRAS-type contribution.","headline":"Useful candidate catalogue with a handful of genuinely new slow-risers, but the abstract's 'order of magnitude' FUor claim is not supported by the body and needs a proper rate calculation before this is ready.","tokens_in":29091,"tokens_out":3168,"would_cite":true,"duration_ms":31453,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A decade of NEOWISE mid-infrared light curves shows that embedded protostars in Cygnus-X erupt with high amplitude at about ten times the rate of more evolved young stars, and that FUor-like candidates are correspondingly an order of…","keywords":["young stellar objects","eruptive variables","FU Orionis stars","mid-infrared variability","NEOWISE","protostars","Cygnus-X","star formation"],"falsifier":"Recalculate the rate of $\\Delta W1 > 1$ mag variability for a sample of Class I sources selected without the 24 $\\mu$m requirement, matched to the SPICY sample in quiescent $W1$ brightness and extinction. If the 22.46% rate drops to the ~5% level of the Class II sample, the stage-dependence claim would be falsified.","tokens_in":28046,"feed_emoji":"⭐","tokens_out":6377,"duration_ms":51861,"temperature":0.7,"pith_summary":"The paper searches a decade of NEOWISE mid-infrared light curves for eruptive young stellar objects in the Cygnus-X star-forming complex, comparing embedded Class I protostars with more evolved flat-spectrum and Class II stars. It reports 48 candidate eruptive variables (plus 20 of less certain classification), 14 of which have photometric behaviour resembling FU Orionis outbursts. The key statistical claim is that FUor-like candidates are roughly an order of magnitude more common among the youngest Class I systems than among more evolved objects. A broader variability comparison finds 22.46% of Class I sources reach $\\Delta W1 > 1$ mag versus 4.94% of flat-spectrum/Class II sources, supporting the view that episodic accretion is most vigorous at the earliest stages.","feed_headline":"FUor outbursts 10x more common in youngest protostars","feed_subtitle":"A decade of NEOWISE mid-infrared data nearly doubles the known FUor-like photometric candidates and ties the rate to stellar age.","key_machinery":"The machine for the search is the NEOWISE time series in the 3.4 $\\mu$m ($W1$) and 4.6 $\\mu$m ($W2$) bands, cleaned and binned into ~6-month epochs over ten years, together with a template-fitting routine that models a FUor-type outburst as a two-component rising slope plus a linear decay, with an e-folding timescale $\\tau$ fit by Markov-chain Monte Carlo. The companion samples are the embedded YSO catalogue of Kryukova et al. (2014a), defined by a 24 $\\mu$m detection, and the SPICY catalogue of Kuhn et al. (2021) for the non-embedded comparison. The templates and the amplitude threshold of 1 mag in either WISE band are what separate candidate eruptive variables from ordinary MIR variability.","core_discovery":"On the basis of photometric light curves alone, the authors identify a population of 48 candidate eruptive variables in Cygnus-X, split between embedded Class I objects selected by a 24 $\\mu$m detection and non-embedded sources from the SPICY catalogue. Fourteen of these show FUor-like morphologies: a fast rise followed by a long decay or plateau, with several showing an unusually slow rise longer than five years. The central comparative result is that candidate FUors are about an order of magnitude more common among Class I systems than among flat-spectrum/Class II sources, and that high-amplitude ($\\Delta W1 > 1$ mag) mid-infrared variability is much more frequent in the embedded group (22.46%) than in the evolved group (4.94%). A large fraction of the short-duration eruptive sources become redder when brighter, in contrast with optically discovered EXors, and the paper presents three near-infrared spectra of two candidate variables indicating FUor-like red continua and CO bandhead absorption.","pith_inferences":["The redder-when-brighter colour behaviour in most short-duration eruptive sources may be a general MIR signature of accretion events, one that is invisible at optical wavelengths; this could be tested by extending the same colour analysis to other star-forming regions with NEOWISE data.","If even a fraction of the slow-rising Class I sources are pre-outburst FUors, the estimated recurrence time of FUor events (~10^5 yr) may need to be revised downward, with consequences for the total mass accreted in episodic bursts.","A direct spectroscopic follow-up of the 14 FUor candidates would turn a photometric classification into a physical taxonomy, and could reveal whether the MIR-selected embedded FUors are spectroscopically distinct from optically discovered FUors."],"forward_implications":["The candidate FUor population in Cygnus-X is roughly doubled, providing new targets for near-infrared spectroscopy to confirm outburst spectra.","The order-of-magnitude excess of FUor-like candidates among Class I systems means models of protostellar accretion and the luminosity spread problem should weight embedded-stage events much more heavily.","The slow-rising (>5 yr) Class I outbursts extend the known range of FUor rise times and imply that some FUors brighten in the mid-infrared long before any optical rise.","The 22.46% versus 4.94% rate for $\\Delta W1 > 1$ mag sets a quantitative benchmark for future time-domain surveys of star-forming regions."],"supporting_citations":[{"why":"defines the embedded Class I/FS sample via 24 $\\mu$m detection and provides the parent catalogue of the first sample.","marker":"Kryukova et al. (2014a)"},{"why":"provides the SPICY catalogue of flat-spectrum/Class II YSOs used as the comparison sample.","marker":"Kuhn et al. (2021)"},{"why":"gives the earlier statistical claim that accretion outbursts are more common at younger stages, which this work tests with MIR data.","marker":"Contreras Peña et al. (2024)"},{"why":"sets the classical FUor definition of high-amplitude, long-duration outbursts with ~1000-day rise times.","marker":"Hartmann & Kenyon (1996)"},{"why":"supplies the NIR-selected FUor comparison sample used in the amplitude-colour correlation analysis.","marker":"Guo et al. (2024a)"},{"why":"documents the NEOWISE survey and its time-series data products.","marker":"Mainzer et al. (2014)"}],"fun_headline_variants":["Youngest protostars host 10x more FUor-like outbursts","Class I YSOs show 10x more FUor candidates than Class II","NEOWISE: FUor candidates 10x more common in embedded YSOs","Cygnus-X eruptive YSOs: 14 FUor-like candidates found"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The rate comparison in Section 4.3 assumes that the embedded Class I sample, defined by a 24 $\\mu$m detection, and the SPICY comparison sample, selected by IRAC colours and a random forest classifier, probe the same underlying physics; if the selection functions differ in luminosity or extinction, the observed difference in high-amplitude variability could be a selection effect rather than an evolutionary stage effect.","fun_headline_variants_meta":{"raw":{"variants":["Youngest protostars host 10x more FUor-like outbursts","Class I YSOs show 10x more FUor candidates than Class II","NEOWISE: FUor candidates 10x more common in embedded YSOs","Cygnus-X eruptive YSOs: 14 FUor-like candidates found"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001179,"raw_usage":{"total_tokens":4931,"prompt_tokens":1063,"completion_tokens":3868,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":679,"completion_tokens_details":{"reasoning_tokens":3780}},"tokens_in":679,"tokens_out":3868,"duration_ms":27375,"temperature":1.0,"reasoning_tokens":3780,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:44:20.307783+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recalculate the rate of $\\Delta W1 > 1$ mag variability for a sample of Class I sources selected without the 24 $\\mu$m requirement, matched to the SPICY sample in quiescent $W1$ brightness and extinction. If the 22.46% rate drops to the ~5% level of the Class II sample, the stage-dependence claim would be falsified.","supporting_citations":[],"review_version":1}