{"id":"14e37893-3312-44eb-926b-872bd418d0e4","arxiv_id":"2501.04208","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"The eRO-ExTra catalog contains 304 extragalactic X-ray transients and variables not associated with known AGN, selected from the first two eROSITA all-sky surveys.","lead":"This paper presents a catalog of 304 X-ray flashes and brightness changes from outside our galaxy that look like they come from supermassive black holes, found in the first two eROSITA sky surveys. It is the first large clean sample of such events, with redshifts, light curves, and spectra, meant to help find and study rare events like stars being torn apart by black holes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Variability-selection statistics (Eqs. 1-2) rest on median error substitution and zero-error 3-sigma upper limits; no robustness test shows how many of the 304 sources survive alternative error/upper-limit choices.","rationale":"Good-faith reading: this is a careful catalog paper with transparent selection, and the authors do not overclaim; they explicitly acknowledge a residual AGN fraction. The reader's CONDITIONAL recommendation is appropriate. The most load-bearing condition for the central claim is not the AGN-cleaning (which is conservative and partly verified by the color-magnitude diagram and existing spectra) but the initial variability definition itself: every one of the 304 sources is in the catalog because it passed the S>4 and A>4 cuts, and those cuts depend on two ad hoc error treatments. The manuscript does not provide a perturbation analysis showing membership stability. This is a standard robustness check and can be done from the authors' own data. The unpublished eRASS2/eRASS:5 catalogs and lack of code make such a check more important, not less. I therefore agree with the reader's weakest-assumption identification and leave the verdict unchanged at CONDITIONAL, pending the robustness test and the already-requested data/code availability.","tokens_in":29046,"tokens_out":12291,"duration_ms":129772,"concrete_test":"Recompute S and A for all 2331 candidates using three alternatives: (i) asymmetric Poisson errors derived from raw eRASS1/eRASS2 counts instead of catalog/median symmetric errors; (ii) a non-zero lower-measurement error, e.g., F_ERR_min = FUL/3, for C1/C2 upper limits; and (iii) 3-sigma upper limits recomputed with ECFs for Gamma=1.5 and Gamma=3.0 (Appendix A). Tabulate how many of the 304 eRO-ExTra sources fail either S>4 or A>4 in each variant and how many new sources enter. If the membership changes by more than ~10% (about 30 sources), the catalog is not robust to the exact definitions in Eqs. (1)-(2), and the 'clean parent sample' claim would need to be weakened or re-derived. If membership changes by less than ~5%, the reader's concern is resolved and the catalog can be used as presented.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that eRO-ExTra is a relatively clean, systematically selected sample of 304 extragalactic X-ray transients and variables. The entire selection is defined by Eqs. (1)-(2) in Sect. 2.1: A=Fmax/Fmin and S=(Fmax-Fmin)/sqrt(F_ERR_max^2+F_ERR_min^2), with both A and S required to exceed 4. Two features of this definition are load-bearing and are not stress-tested in the paper. First, for the 404 C1/C2 sources with only one epoch detected, Fmin is a 3-sigma upper limit and F_ERR_min is set to zero. Setting the lower measurement's error to zero maximizes S by construction, so S is not a significance in the usual sense; it is a distance-in-units-of-the-upper-limit criterion. Second, for catalog detections without published errors, F_ERR is replaced by the median count error of sources within 10% in counts and exposure. If the true error for a faint, crowded, or slightly extended source is larger than this median, S is overestimated. The upper limits themselves also depend on a fixed absorbed power-law ECF (Gamma=2, NH=3e20; Appendix A), so a soft or hard transient gets a systematically wrong Fmin. The paper does not report how many of the 2331 initial candidates, or of the final 304, would change if any of these choices were varied. Since the catalog and all downstream statistics (light-curve class fractions, spectral slopes, Appendix E rate) inherit the S>4 and A>4 membership, this is the least secure link in the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents the eRO-ExTra catalog, a systematically selected sample of 304 extragalactic X-ray transients and variables in the eROSITA all-sky surveys eRASS1 and eRASS2, selected to have variability significance S>4 and fractional amplitude A>4 between the two surveys, located in the Legacy Survey DR10 footprint, and cleaned of Galactic stars and known AGN using multiwavelength criteria. The catalog provides, for each source, eROSITA light curves over eRASS1-4(5) with a classification into decline, flare, brightening, or other; peak-epoch X-ray spectral fits with an absorbed power law; archival X-ray detections and upper limits from Swift, ROSAT, and XMM-Newton; optical/near-infrared counterparts and redshifts (spectroscopic or photometric) for more than 80% of sources; and radio detections from RACS and VLASS for 31 sources. The authors propose the catalog as a 'relatively clean parent sample of non-AGN variability phenomena associated with massive black holes' and use it to compute an X-ray luminosity function and an integrated volumetric rate in Appendix E.","tokens_in":29326,"tokens_out":7531,"duration_ms":63558,"significance":"If the selection is robust, this is the largest systematically selected sample of extragalactic X-ray transients and variables without known AGN association, with a level of multiwavelength characterization that previous samples lack. The paper is transparent about methodology: selection steps are documented with flowcharts, thresholds are stated, purity/completeness of the p_any counterpart probability is quantified, upper limits and archival constraints are described in detail, and spectral fitting uses Bayesian methods (BXA/UltraNest). The authors also explicitly acknowledge that residual AGN contamination remains (Sect. 7), which is appropriate for a 'relatively clean' rather than pure sample. The catalog will be a useful resource for TDE, QPE, and other nuclear-transient population studies, and the paper helps bridge the flux gap between deep pencil-beam surveys and shallower all-sky variability studies.","major_comments":[{"comment":"For the 404 sources detected in only one of eRASS1/eRASS2 (C1 and C2), Fmin is a 3-sigma upper limit and F_ERR_min is set to zero. Thus S = (Fmax - 3sigma_UL)/F_ERR_max is not a significance in the usual sense; it is a ratio in which the lower measurement is treated as exact. The choice of a 3-sigma confidence, the fixed absorbed power law (Gamma=2, NH=3e20) used to convert counts to flux for the upper limit (Appendix A), and the zeroing of F_ERR_min all directly control which sources pass the S>4 cut. The paper does not report how many of the 2331 initial candidates, or of the final 304 sources, would survive if the upper-limit error were propagated, if a different confidence level were used, or if the spectral model were varied to, e.g., Gamma=1.5 or 2.5. Because the entire catalog and all downstream statistics (light-curve class fractions, spectral index distribution, and the Appendix E rate) inherit this membership, a robustness test of these choices is load-bearing and should be added.","section":"Sect. 2.1, Eqs. (1)-(2)"},{"comment":"The 1/Vmax XLF and the integrated rate are computed using sensitivity maps that account for the DET_LIKE>15 cut and the A>4 amplitude cut, but not for the S>4 significance cut. Since S depends on the flux errors, which scale with the number of detected counts, S is not equivalent to A>4 and will impose an additional distance-dependent constraint on the detectable volume. Omitting this criterion from the selection function can bias the maximum volume Vmax and hence the XLF and the reported rate of 1.8e-7 Mpc^-3 yr^-1. The authors should either include the S>4 condition in the sensitivity calculation or demonstrate that it is always less constraining than the amplitude cut for the sample under consideration.","section":"Appendix E"},{"comment":"For catalog detections without published errors, the paper assigns the median count error of sources within 10% in counts and exposure time. This is a reasonable first-order approximation, but the paper does not report how many of the 2331 sources are affected, nor does it test the sensitivity of the final catalog to the width of the matching bin (e.g., 5% or 20%) or to the use of a different percentile (e.g., 84th instead of 50th). If the true errors for faint, slightly extended, or crowded sources are larger than the adopted median, S will be overestimated and sources may enter the catalog spuriously. A brief robustness test showing the number of final sources under alternative error-assignment choices would strengthen the central claim.","section":"Sect. 2.1 (median error substitution)"}],"minor_comments":[{"comment":"The optimal p_any=0.17 threshold is determined as the intersection of purity and completeness functions calculated 'for our sample after the NW AY match.' This is a self-calibration on the same sample used to build the catalog; it would be helpful to state whether the purity/completeness curves were validated on an independent set or via cross-validation, so that the reported <5% chance-coincidence rate is not circular.","section":"Sect. 2.2"},{"comment":"Table 1 sums to 440 sources (144 detected, 296 not detected), which matches the number of sources entering the archival variability step, not the final 304. The caption should state that the table refers to the sample before the exclusion of the 136 archival-variable sources, to avoid confusion with the final catalog.","section":"Sect. 2.5, Table 1"},{"comment":"The projection name 'Aito ff' in the caption of Fig. 2 should be 'Aitoff'.","section":"Fig. 2 and throughout"},{"comment":"The text 'the function should have been corrected by a factor of 2.6' is unclear. Please specify that the factor accounts for both the LS10 areal coverage (76% of eROSITA_DE) and the fact that eROSITA_DE covers approximately half the sky, and clarify whether the resulting XLF is per unit volume of the full sky.","section":"Appendix E"},{"comment":"The light-curve classification uses the same S definition (Eq. 2) for comparisons involving later eRASS nondetections, so the F_ERR_min=0 issue also affects the 'decline' and 'flare' classes when upper limits are involved. A brief note acknowledging this, or a consistency check, would be appropriate.","section":"Sect. 3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the scope of A&A and the catalog is potentially a valuable community resource. The main issue is that the catalog membership is defined by variability statistics (S and A) that rely on admittedly approximate treatments of errors and upper limits, and the paper does not currently provide any robustness tests of these treatments. The requested tests are straightforward to perform and would substantially increase confidence in the 304-source sample and in the Appendix E rate. I also note that the paper cites several 'in prep' or 'submitted' works (e.g., eRASS:5, Salvato et al. 2024a) that are central to the counterpart identification; this is acceptable but should be flagged in the published version if not yet available."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a genuinely useful catalog paper: 304 eROSITA sources with S>4 and A>4 between eRASS1 and eRASS2, cleaned against stars and known AGN, with redshifts for more than 80%, light curve classes, peak spectral fits, and radio matches. It is the largest systematic sample of its kind and fills a real gap between the shallow ROSAT variability sample and the deep pencil-beam XMM-COSMOS study. The selection flow is documented in detail, and the authors are appropriately cautious: they call the sample \"relatively clean\" and explicitly note that some AGN remain. That is the right tone.\n\nSecond, the soft spot is exactly where the stress-test points: the S statistic in Eq. (2). For the 404 single-epoch sources, Fmin is a 3-sigma upper limit and F_ERR_min is set to zero. That makes S a ratio of the peak flux to the upper limit, not a significance in any standard sense, and it maximizes S by construction. The upper limits themselves assume a fixed absorbed power law with Gamma=2 and NH=3e20, so a soft transient gets a systematically wrong Fmin. The median-error substitution for catalog detections without errors is a reasonable patch, but it adds another layer of uncertainty. The paper does not report how many of the final 304 would survive alternative error treatments or upper-limit assumptions. That matters because the catalog membership is defined by those two numbers, and everything downstream—light-curve class fractions, spectral slopes, the Appendix E rate—inherits that selection.\n\nIs this fatal? Not for the catalog as a resource. The selection is explicit, and users can re-derive it once the eRASS2 catalog is public. But it is the least secure link. A robustness appendix with a few variants (e.g., using a 1-sigma upper limit, a softer spectral model, or propagating an error on the upper limit) would make the population conclusions much firmer. I would not block the paper on this, but I would ask for it.\n\nAlso worth flagging: the sample depends on the unpublished eRASS2 and eRASS:5 catalogs, which makes exact reproduction hard. That is typical for eROSITA pre-DR2, but it should be stated clearly and the product versions frozen. The XLF and rate in Appendix E are self-measurements under the selected criteria, not circular predictions, which is fine.\n\nBottom line: this deserves a serious referee and will be a widely used catalog. The central claim holds up—it is a relatively clean parent sample—provided the caveats are respected. I would accept it after a moderate revision that adds the robustness checks and makes the upper-limit treatment more defensible.","headline":"The largest clean-ish sample of extragalactic non-AGN X-ray transients from eROSITA, with rich data products; the variability selection has a real systematic that should be stress-tested before the catalog is used for precision population statistics.","tokens_in":30091,"tokens_out":3197,"would_cite":true,"duration_ms":31782,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"eROSITA finds 304 X-ray flares outside known active galaxies","keywords":["X-ray transients","eROSITA","tidal disruption events","supermassive black holes","active galactic nuclei","X-ray variability","all-sky surveys","extragalactic catalogs"],"falsifier":"Recompute fluxes, errors, and upper limits for all 2331 pre-cleaning variable candidates using independent forced photometry on the eRASS1 and eRASS2 images, then reapply the $S>4$ and $A>4$ cuts; if the 304-source catalog does not reproduce within a few percent, its completeness and derived rate are not robust. A complementary check is to search archival X-ray images for the 296 sources listed as not detected, looking for any at or above the claimed $3\\sigma$ upper limits; finding a significant number would break the claim that most eRO-ExTra sources are genuinely new X-ray transients.","tokens_in":28773,"feed_emoji":"🔭","tokens_out":12622,"duration_ms":107429,"temperature":0.7,"pith_summary":"Most extragalactic X-ray variability is the ordinary flickering of active galactic nuclei (AGN), galaxies whose central black hole is steadily accreting, but a small fraction of X-ray flares come from rarer events: stars torn apart by supermassive black holes, quasi-periodic eruptions, and other short-lived accretion phenomena. This paper argues that, after stripping away stars, galaxy clusters, and everything previously classified as an AGN, the first two eROSITA all-sky surveys still contain 304 genuinely extragalactic X-ray transients and variables. The resulting eRO-ExTra catalog is presented as a relatively clean parent sample of non-AGN variability associated with massive black holes, with optical counterparts for over 90% of sources, reliable redshifts for over 80%, peak-spectrum fits, light-curve classes, and radio identifications for 31 sources. More than 95% of the sources were discovered in X-rays for the first time. If the selection is right, this is the largest systematically selected sample of its kind and a direct resource for measuring how often otherwise quiet black holes flare.","feed_headline":"eROSITA finds 304 X-ray flares outside known active galaxies","feed_subtitle":"Most were never seen in X-rays before, giving astronomers a rare sample of nuclear flare candidates to study.","key_machinery":"The load-bearing machinery is a two-epoch variability selection expressed by two numbers per source. The fractional amplitude is $A = F_{\\mathrm{max}}/F_{\\mathrm{min}}$, and the variability significance is $S = (F_{\\mathrm{max}}-F_{\\mathrm{min}})/\\sqrt{F_{\\mathrm{err,max}}^2+F_{\\mathrm{err,min}}^2}$, with both required to exceed four; for sources detected in only one survey, the missing flux is a $3\\sigma$ upper limit computed from aperture photometry under an assumed absorbed power-law spectrum ($\\Gamma=2$, $N_{\\mathrm{H}}=3\\times10^{20}\\,\\mathrm{cm}^{-2}$). Everything downstream, including the final 304-source catalog and its derived luminosity function and rate, inherits these cuts. The second half of the machinery is the cleaning cascade: optical counterpart association, exclusion of stars by parallax, morphology, and catalog classifications, exclusion of extended cluster emission, a mid-infrared color cut against AGN, exclusion of known AGN and quasars from pre-eROSITA catalogs, visual inspection of archival optical spectra, and exclusion of sources with archival X-ray variability. The catalog is the intersection of this selection with a large optical survey footprint covering about three-quarters of the surveyed sky.","core_discovery":"The central discovery is the catalog itself: 304 X-ray sources that change by more than a factor of four in flux between the first two eROSITA all-sky surveys in the 0.2-2.3 keV band, with variability significance $S>4$ and fractional amplitude $A>4$, and that survive a deliberate campaign to exclude anything with a known AGN signature. The selection starts from the full eRASS1 and eRASS2 source lists, requires a detection-likelihood of at least 15 for the brighter detection of each pair, removes extended and spurious sources, assigns $3\\sigma$ upper limits to sources detected in only one survey, and then removes stars, galaxy clusters, mid-infrared-selected AGN, objects with pre-eROSITA active-galaxy classifications, broad-line AGN spectra, blazars, and sources with archival X-ray variability. The paper reports the resulting sample's properties: more than 90% have reliable optical counterparts, more than 80% have spectroscopic or photometric redshifts, each source has a peak-epoch power-law spectral fit and a light-curve class (flare, decline, brightening, or other), and 31 sources are radio detected. More than 95% of the sources have no prior X-ray detection, so they are new discoveries. The paper concludes that eRO-ExTra constitutes a relatively clean parent sample of non-AGN variability phenomena associated with massive black holes, suitable for population studies of tidal disruption events and related transients.","pith_inferences":["Editorial inference: the integrated rate of about $1.8\\times10^{-7}\\,\\mathrm{Mpc}^{-3}\\,\\mathrm{yr}^{-1}$ for this mixed population is higher than canonical tidal-disruption-event rates, suggesting the sample includes a broader class of low-luminosity nuclear accretion flares whose true occurrence rate in quiescent galaxies may have been underestimated.","Editorial inference: because the selection compares only eRASS1 with eRASS2, events that rise and fade within one six-month survey or peak outside that window are missed; applying the same cuts to the later all-sky surveys should multiply the sample and constrain the duty cycle of these events.","Editorial inference: the cleaning strategy predicts that the remaining sources should not preferentially sit in massive, actively growing host galaxies; measuring host stellar masses could test whether the parent population is genuinely inactive black holes.","Editorial inference: the radio-detected sources with luminosities above the star-formation expectation may include a systematically selected sample of radio-emitting nuclear transients; comparing radio brightness across the two radio epochs already available could separate newly launched jets from persistent low-level AGN activity."],"forward_implications":["The catalog gives a homogeneous parent population for studying rare nuclear transients such as tidal disruption events and quasi-periodic eruptions, with a known selection function rather than a set of serendipitous discoveries.","Population statistics follow directly: a sky density of about 0.03 sources per square degree per year, a double-power-law X-ray luminosity function, and an integrated volumetric rate of $1.8^{+0.5}_{-0.4}\\times10^{-7}\\,\\mathrm{Mpc}^{-3}\\,\\mathrm{yr}^{-1}$ for this selection.","Individual follow-up can now be targeted: the peak photon-index distribution, light-curve class, and radio data separate subpopulations worth multiwavelength study.","Because more than 95% of the sources are new X-ray discoveries, previous X-ray surveys lacked the sensitivity or cadence to catch this population, and future all-sky surveys can apply the same cuts directly.","The 31 radio-detected sources provide a sample for testing whether compact jets accompany these non-AGN nuclear flares."],"supporting_citations":[{"why":"Supplies the eRASS1 source catalog and the detection-likelihood and positional-error framework used to build the variability sample.","marker":"Merloni et al. 2024"},{"why":"Provides the eRASS1 short-timescale variability study whose selection and source classes are compared with eRO-ExTra.","marker":"Boller et al. 2024"},{"why":"Supplies the optical counterpart association method and the p_any reliability probabilities used to attach counterparts.","marker":"Salvato et al. 2018"},{"why":"Gives the W1-W2 mid-infrared color criterion used to exclude AGN from the sample.","marker":"Stern et al. 2012"},{"why":"Supplies the Million Quasars catalog used to identify and remove additional AGN and QSO contaminants.","marker":"Flesch 2023"},{"why":"Provides the Bayesian upper-limit method used both for archival flux limits and for the 3-sigma non-detection upper limits in the variability selection.","marker":"Kraft et al. 1991"},{"why":"Provides the COSMOS XMM-Newton variability baseline against which the eRO-ExTra flux, luminosity, and redshift ranges are compared.","marker":"Lanzuisi et al. 2014"},{"why":"Gives the canonical tidal-disruption-event luminosity and spectral expectations used to interpret the catalog population.","marker":"Gezari 2021"}],"fun_headline_variants":["eRO-ExTra: 304 new X-ray transients from eROSITA","304 extragalactic X-ray flares discovered by eROSITA","eROSITA catalog reveals 304 non-AGN X-ray variables","First two eROSITA surveys yield 304 exotic X-ray sources","eRO-ExTra: a clean sample of 304 X-ray transients"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The two variability numbers per source depend on estimated flux errors, and for sources not detected in one of the two surveys they depend on a $3\\sigma$ upper limit computed from an assumed spectral model; if those substituted errors or upper limits are systematically wrong, the set of sources crossing the $S>4$ and $A>4$ thresholds changes, taking the whole catalog and its derived statistics with it.","fun_headline_variants_meta":{"raw":{"variants":["eRO-ExTra: 304 new X-ray transients from eROSITA","304 extragalactic X-ray flares discovered by eROSITA","eROSITA catalog reveals 304 non-AGN X-ray variables","First two eROSITA surveys yield 304 exotic X-ray sources","eRO-ExTra: a clean sample of 304 X-ray transients"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000302,"raw_usage":{"total_tokens":1899,"prompt_tokens":1261,"completion_tokens":638,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":877,"completion_tokens_details":{"reasoning_tokens":539}},"tokens_in":877,"tokens_out":638,"duration_ms":5883,"temperature":1.0,"reasoning_tokens":539,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:39:01.376141+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute fluxes, errors, and upper limits for all 2331 pre-cleaning variable candidates using independent forced photometry on the eRASS1 and eRASS2 images, then reapply the $S>4$ and $A>4$ cuts; if the 304-source catalog does not reproduce within a few percent, its completeness and derived rate are not robust. A complementary check is to search archival X-ray images for the 296 sources listed as not detected, looking for any at or above the claimed $3\\sigma$ upper limits; finding a significant number would break the claim that most eRO-ExTra sources are genuinely new X-ray transients.","supporting_citations":[{"cited_title":"The eROSITA DR1 variability catalogue","cited_arxiv_id":"2401.17280","evidence_quote":"Provides the eRASS1 short-timescale variability study whose selection and source classes are compared with eRO-ExTra."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Million Quasars catalog used to identify and remove additional AGN and QSO contaminants."},{"cited_title":"2014, , 781, 105","cited_arxiv_id":null,"evidence_quote":"Provides the COSMOS XMM-Newton variability baseline against which the eRO-ExTra flux, luminosity, and redshift ranges are compared."}],"review_version":1}