{"id":"c0bb1df2-2f27-4fef-8152-0343225323a6","arxiv_id":"2412.14298","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Variability-based machine learning selection achieves 87% spectroscopic confirmation of active galactic nuclei in low-stellar-mass galaxies, reaching black hole masses near 10^6 solar masses.","lead":"Astronomers trained machine-learning models on ZTF light curves to pick out active black holes in small galaxies, then checked the picks against SDSS spectra. They report that 87% of 415 checked candidates are genuine AGN, a much higher success rate than earlier variability searches.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Training-set overlap may inflate the headline 87% confirmation rate; the paper never states the 506 candidates were excluded from the classifiers' training labels.","rationale":"The paper's central quantitative claim is the high spectroscopic confirmation rate of variability-selected AGN candidates in low-mass galaxies. That claim is only meaningful if the confirmation is an independent test of the classifier. Because the classifiers were trained on labeled AGN samples that likely include SDSS-confirmed objects of the same kind as the validation set, undetected training overlap would make the 87% figure partly a self-consistency check. This is precisely the reader's weakest assumption, and I agree it is the most load-bearing concern. The X-ray match rate provides partial independent support, but it does not directly test the spectroscopic broad-line purity fraction that headlines the paper. The proposed hold-out test would settle the matter; the verdict remains CONDITIONAL as the reader recommended, pending that verification.","tokens_in":37853,"tokens_out":5170,"duration_ms":53140,"concrete_test":"Cross-match the 506 candidate coordinates and unique object IDs against the positive training samples of Sánchez-Sáez et al. (2021a), Sánchez-Sáez et al. (2023), and the forced-photometry classifier (Arévalo et al., in prep.); obtain these training IDs from the authors or public catalogs. Then recompute the broad-line and 87% confirmation rates on the subset of candidates with no training-set match. If the held-out confirmation rate remains comparable (e.g., ≥85%) across all sets, the claim survives; if it drops substantially, the reported purity is inflated by leakage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2 adopts the random-forest classifiers of Sánchez-Sáez et al. (2021a, 2023) and Arévalo et al. (in prep.), whose positive training labels are, in standard practice, spectroscopically confirmed AGN drawn from SDSS and similar surveys. The validation in Sections 3–4 then uses SDSS DR17 spectra to measure how many of the 506 variability-selected candidates show broad Balmer lines or AGN-like BPT ratios. If a substantial subset of those 506 candidates was present in the training set, the quoted 86–98% broad-line fractions (Section 4.1) measure how well the classifiers reproduce their own training labels, not how well optical variability selects previously unknown AGNs in low-mass galaxies. The paper does not state that the candidates were held out from training. This concern is partly mitigated by the independent eROSITA X-ray match rate (Section 5: 67% of VCS objects, 75% of BEL objects), which is not part of the classifier inputs, but the central quantitative purity claim is spectroscopically defined and therefore directly affected by any training overlap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports on the use of random-forest classifiers applied to ZTF light curves to select AGN candidates in low-stellar-mass galaxies from the NASA-Sloan Atlas, using four different input data products: the ALeRCE alert stream, ZTF DR11 g-band and r-band light curves, and custom forced photometry. The 506 matched candidates are cross-checked against SDSS DR17 spectra; 415 pass the author's visual and redshift-quality cuts, and from these 362 (87%) are confirmed as AGNs, mainly through broad Balmer lines (357 objects) plus five additional BPT/WHAN-selected objects. The paper also derives black hole masses and Eddington ratios from spectral fits, compares BPT classifications with external catalogs, and reports X-ray counterpart fractions from eRASSv1.1 (67% of the VCS sample in the eROSITA-DE sky, rising to 75% for objects with broad lines). The central claim is that variability-based classification, especially when based on complete light curves, is a highly pure method for finding AGNs in low-mass galaxies.","tokens_in":38142,"tokens_out":7900,"duration_ms":75493,"significance":"If the reported purity holds, this is a valuable result for AGN censuses in low-mass galaxies and for planning IMBH searches with future facilities such as LSST. The paper's strengths include a careful spectral-fitting approach with Monte Carlo error estimation, visual inspection of borderline objects and light curves, comparison with MPA-JHU and Portsmouth emission-line catalogs, and an independent X-ray validation using eROSITA. The eROSITA match rates are particularly useful as an external check that does not depend on the optical classifiers. However, the main quantitative claim—the 87% confirmation rate—depends on the classifiers' training labels, and the paper does not establish that the candidates were held out from training. This must be addressed before the selection purity can be accepted as demonstrated.","major_comments":[{"comment":"The manuscript adopts the random-forest classifiers of Sánchez-Sáez et al. (2021a), Sánchez-Sáez et al. (2023), and Arévalo et al. (in prep.) but never states that the 506 NSA-matched candidates (or the 415 VCS objects) were excluded from those classifiers' training labels. If the training sets include spectroscopically confirmed SDSS AGN, as is standard practice for such classifiers, then the 86% broad-line fraction in Section 4.1 and the per-set 94-98% fractions partly measure how well the classifier recalls its own training labels rather than how well optical variability selects new AGNs in low-mass galaxies. Because this confirmation rate is the paper's central quantitative claim, please (a) report whether any of the 506 candidates appear in the training sets, (b) recompute the confirmation fractions after removing any overlapping objects, or (c) clearly justify that the classifiers' labels were obtained without using SDSS spectroscopy of these candidates. The eROSITA match rate in Section 5 is an important independent check, but it does not by itself validate the spectroscopically defined purity numbers.","section":"Section 2 and Sections 3-4"},{"comment":"The headline '87% confirmed' is computed only for the 415-object VCS sample after excluding 35 spectra that include 22 higher-redshift AGN and 5 suspected AGN, and the 52 candidates without spectra are not in the denominator. The paper is transparent about the 65 unconfirmed objects, but it should also state the unconditional confirmation rate among all 506 candidates (at least 362/506 = 71.5%, rising to about 77% if the 22 higher-z and 5 suspected AGN are counted) and discuss how the exclusion of the lowest-mass, lowest-redshift candidates without spectra (Section 3.2) affects the reported confirmation and mass statistics. This would prevent the abstract's conditional rate from being read as the overall selection success rate.","section":"Section 3.1, Table 2, Section 4.1"},{"comment":"The Forced Photometry selection, which produces the paper's highest confirmation rate (168/170 in the VCS sample), relies on a classifier and feature definitions described only as in Arévalo et al. (in prep.), with no public code, training-label description, or hyperparameters. The additional features (PAN-STARRS i-z color, Gaia proper motion, Mexican-Hat variances, flux asymmetry) are listed, but the labeled training set and classifier construction are not available for inspection. Because this set is central to the comparison in Section 4.1 and to the conclusion that forced photometry finds twice as many AGNs in the same sky region, please include the method details or a public reference that the reader can inspect; otherwise the main result for the most sensitive set is not reproducible.","section":"Section 2 (Forced Photometry)"}],"minor_comments":[{"comment":"There is a typo: 'Whit this reduction' should be 'With this reduction'; also, Sánchez-Sáez et al. 2021a and 2021b appear to be the same paper and should be merged in the reference list.","section":"Section 2"},{"comment":"The text 'below 1.8×16M⊙' appears to be a typo for '1.8×10^6 M⊙'; similarly, '6.4−6.4×10^6 M⊙' should likely be '6.4×10^6 M⊙' for both DR-g and DR-r.","section":"Conclusion, item 4"},{"comment":"The per-set confirmation percentages are computed only for objects with good spectra; adding the total candidate counts per set (e.g., 168/188 for Forced Photometry, 40/54 for DR-g, 40/54 for DR-r) would make the selection purity less dependent on SDSS spectral availability.","section":"Table 2 and Section 4.1"},{"comment":"The text gives both 5 arcsec and 10 arcsec match rates (123 and 150 objects) and later uses 10 arcsec in Table 8; please state explicitly in the table caption that the quoted percentages use a 10 arcsec matching radius.","section":"Section 5 and Table 8"},{"comment":"The green diamonds and the green square are described in the text, but the markers may be difficult to distinguish in printed grayscale; consider using distinct shapes or a colorblind-safe palette.","section":"Figure 9"},{"comment":"The sentence 'The dashed grey line correspond to zero Flux' should be 'The dashed grey line corresponds to zero flux.'","section":"Appendix D"}],"recommendation":"major_revision","confidential_remarks":"The training-set-overlap question is the key substantive issue: the paper's headline purity claim cannot be fully evaluated until the authors either demonstrate that the 506 candidates were not in the classifiers' training sets or recompute the confirmation rates after removing any overlapping objects. The dependence on an unpublished classifier description for the Forced Photometry set is a secondary but real reproducibility concern. The topic and approach are within the journal's scope, and the paper is otherwise well constructed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"One thing to know: this is a solid observational paper that makes a good case for high purity in variability-selected AGN in low-mass galaxies. The headline 87% spectroscopic confirmation rate is probably in the right ballpark, but the paper never says the 506 candidates were excluded from the classifiers' training labels, so the number should be treated with a grain of salt until that is clarified.\n\nWhat's new: the systematic comparison of four ZTF-based selections—alerts, DR g/r, and a custom forced-photometry light-curve set—on the same low-mass galaxy sample. The forced-photometry extension is genuinely useful; it finds lower-mass black holes and reaches 97% BEL detection among candidates with spectra. The spectral analysis is careful: Monte Carlo error estimation, visual inspection of edge cases, and a detailed decomposition with three models. The X-ray match rate of 67–75% with eROSITA is an independent validation that strengthens the case. The authors also do a service by explaining the 65 non-confirmations: most are bad subtractions or transients, which tells you where the method fails.\n\nSoft spots: the biggest is the training-set overlap. The random-forest classifiers come from prior work and were presumably trained on spectroscopically labeled AGN, many from SDSS. The confirmation here uses SDSS spectra. The paper does not state the 506 candidates were held out from training. If many were in the training set, the 87% is partly a memory test. The eROSITA results suggest the selection is genuinely finding AGN, but the authors need to address the hold-out question directly. The DR selections also rely on hand-tuned cuts whose optimization is deferred to Bauer et al. (in prep.), and the forced-photometry classifier is in Arévalo et al. (in prep.), so full reproducibility is not in this paper. That is a moderate issue, not a dealbreaker. Black hole masses rest on a virial relation not validated at the lowest masses, and the sample is biased toward brighter, more massive BHs; the authors acknowledge both.\n\nBottom line: this is a useful paper for anyone working on AGN demographics in low-mass galaxies or building variability selection for LSST. It deserves a serious referee. I would ask the authors to (1) state explicitly whether the candidates overlapped with training data and provide a hold-out estimate if possible, and (2) give enough detail about the filter optimization or cite a public description. If those are addressed, the 87% claim will be much easier to trust.","headline":"A high-purity variability-selected AGN sample in low-mass galaxies, but the 87% confirmation rate needs an explicit hold-out statement before it becomes fully convincing.","tokens_in":38608,"tokens_out":4126,"would_cite":true,"duration_ms":37852,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Applying random forest classifiers to optical light curves confirms 87% of low-mass galaxy AGN candidates, showing variability is a reliable route to black holes in dwarf galaxies.","keywords":["active galactic nuclei","low-stellar-mass galaxies","optical variability","random forest classification","intermediate-mass black holes","broad emission lines","Zwicky Transient Facility","X-ray counterparts"],"falsifier":"Compare the object IDs of the 506 candidates against the training sets used to build the random forest classifiers; if a substantial fraction of the 357 confirmed AGNs appear there, the 87% confirmation rate is not an independent validation, and a cleaner test would be to retrain on labels that explicitly exclude all 506 candidates and remeasure the broad-line confirmation rate on the held-out set.","tokens_in":37636,"feed_emoji":"🔭","tokens_out":8453,"duration_ms":67808,"temperature":0.7,"pith_summary":"This paper asks whether optical variability alone can reliably uncover actively accreting black holes in low-stellar-mass galaxies, the places where intermediate-mass black hole seeds are expected to hide. It reports yes: of 415 candidates selected by random forest classifiers on Zwicky Transient Facility light curves and cleaned by visual spectral inspection, 87% show broad Balmer emission lines or BPT/WHAN AGN diagnostics in archival SDSS spectra. Purity is higher when selection uses complete light curves rather than the alert stream, with 94–98% of complete-light-curve candidates showing broad lines and the forced-photometry set reaching 97%. Black hole masses run from about $2.2\\times10^6$ to $4.2\\times10^7\\,M_\\odot$ and cluster near 0.1% of host stellar mass, and two-thirds of candidates with X-ray coverage have eROSITA counterparts. The point that matters is that a variability-based census can populate the low-mass end of the black hole mass function, though it appears biased toward larger black holes for a given host.","feed_headline":"Variability uncovers black holes in 87% of dwarf galaxy candidates","feed_subtitle":"Follow-up spectra confirm broad emission lines in 94-98% when selection uses full light curves","key_machinery":"The mechanism that carries the argument is a set of hierarchical random forest classifiers that turn multi-epoch ZTF photometry into AGN classes: the alert-stream classifier, the ZTF data-release classifiers in g and r, and a forced-photometry classifier run on difference images. Their job is to separate AGN variability from ordinary stellar variability, transients, and artifacts using features such as damped random walk timescales, excess variance, colors, Gaia proper motion, and morphology. These classifiers select the 506 candidates; the confirmation step then uses pPXF spectral fitting of SDSS spectra with stellar-population templates, Gaussian narrow lines, Gauss-Hermite broad lines, Fe II pseudo-continuum, and power-law AGN continua to measure broad Balmer lines, black hole masses, and narrow-line ratios.","core_discovery":"The paper's central claim is that random-forest classifiers applied to multi-epoch optical light curves select genuine type I AGNs in low-stellar-mass galaxies with high purity. Starting from 506 variability-selected candidates matched to NASA-Sloan Atlas galaxies with $M_*<2\\times10^{10}\\,M_\\odot$ and $z<0.15$, the authors obtain 415 clean low-redshift spectra and confirm 362 (87%) as AGNs: 357 through significant broad Balmer lines (EW$_{\\mathrm{H}\\alpha}>5$ Å and SNR>3) and five more through BPT/WHAN narrow-line diagnostics. Candidates chosen from complete light curves, namely the ZTF data-release g and r sets and a custom forced-photometry set on difference images, are confirmed at 94–98%, while the alert-based set is 80% and contains essentially all the false positives, mostly bad image subtraction and nuclear transients. The inferred black hole masses span $2.2\\times10^6$ to $4.2\\times10^7\\,M_\\odot$, and the host-to-black-hole mass ratio clusters near $M_*/M_{\\mathrm{BH}}\\sim1000$, above the Reines & Volonteri relation but near the 0.1% ratio seen in more massive ellipticals. Among candidates in the eROSITA-DE footprint, 67% have X-ray counterparts, rising to 75% for those with broad lines, and this X-ray match rate does not depend on BPT class.","pith_inferences":["Editorial extension: verify whether the 506 candidates were excluded from the random forest training labels; without that exclusion, part of the 87% rate measures label reproduction rather than independent generalization.","Editorial extension: apply the forced-photometry classifier to the full NSA galaxy catalog, not just the mass-limited sample, to map how purity and mass bias vary with host mass.","Editorial extension: stack eROSITA images at the 33% of candidates without catalog counterparts to separate genuinely X-ray-faint AGNs from objects below the detection limit.","Editorial extension: high-spatial-resolution spectroscopy of the star-forming-classified broad-line objects would test whether their narrow lines are diluted by host galaxy emission, as the paper suggests."],"forward_implications":["Variability selection can be used to build a census of low-mass black holes in nearby galaxies, adding hundreds of objects to the small set of spectroscopically confirmed AGNs in dwarfs.","BPT-only searches miss a large fraction of type I AGNs: over half of the candidates classified as star-forming by narrow-line ratios still show broad Balmer lines.","Selection from complete, difference-image-based light curves is both purer and more sensitive to the lowest black hole masses than alert-based selection.","X-ray follow-up is most productive for candidates with broad lines: the fraction with eROSITA counterparts is 75% for broad-line objects and only 4% for those without.","Because variability selection preferentially finds the largest black holes for a given host mass, the low-mass end of the scaling relation remains incomplete until deeper, higher-cadence surveys such as LSST supply more sensitive light curves."],"supporting_citations":[{"why":"Supplies the alert-stream random forest classifier used to build the Alerts set.","marker":"Sánchez-Sáez et al. 2021a"},{"why":"Supplies the g- and r-band data-release classifiers and the filtering features that define the DR-g and DR-r sets.","marker":"Sánchez-Sáez et al. 2023"},{"why":"Describes the ZTF data-release PSF light curves from which the DR light curves are built.","marker":"Masci et al. 2019"},{"why":"Provides the single-epoch virial mass scaling relation used for all black hole mass estimates.","marker":"Mejía-Restrepo et al. 2016"},{"why":"Provides the eRASSv1.1 X-ray source catalog used for the X-ray counterpart analysis.","marker":"Merloni et al. 2024"},{"why":"Is the reference black hole mass versus stellar mass relation whose shallower normalization the paper contrasts with its own mass ratios.","marker":"Reines & Volonteri 2015"},{"why":"Is the closest previous variability-based low-mass AGN sample; the paper compares confirmation and X-ray match fractions against it.","marker":"Baldassare et al. 2020"},{"why":"Provides the prior ZTF variability-selected dwarf AGN sample cross-matched in detail to test agreement between selection methods.","marker":"Ward et al. 2022"},{"why":"Establishes the low X-ray match rates for earlier variability-selected samples that the paper explains by selection and depth.","marker":"Arcodia et al. 2024"}],"fun_headline_variants":["Machine learning on light curves finds 87% of dwarf AGNs","Variability selects dwarf galaxy black holes with 87% purity","Full light curves catch 98% of dwarf AGN candidates","Optical variability confirms black holes in 87% of dwarf galaxies","Random forests on ZTF data reveal dwarf AGNs at high rate"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The classifiers that pick the candidates were trained on previously labeled variable sources, and the paper does not show that the 506 candidates were excluded from that training; if they were not, the high confirmation rate partly reflects how well the classifier repeats its own training labels rather than an independent test of variability selection.","fun_headline_variants_meta":{"raw":{"variants":["Machine learning on light curves finds 87% of dwarf AGNs","Variability selects dwarf galaxy black holes with 87% purity","Full light curves catch 98% of dwarf AGN candidates","Optical variability confirms black holes in 87% of dwarf galaxies","Random forests on ZTF data reveal dwarf AGNs at high rate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000221,"raw_usage":{"total_tokens":1603,"prompt_tokens":1252,"completion_tokens":351,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":868,"completion_tokens_details":{"reasoning_tokens":262}},"tokens_in":868,"tokens_out":351,"duration_ms":3969,"temperature":1.0,"reasoning_tokens":262,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:21:08.536288+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the object IDs of the 506 candidates against the training sets used to build the random forest classifiers; if a substantial fraction of the 357 confirmed AGNs appear there, the 87% confirmation rate is not an independent validation, and a cleaner test would be to retrain on labels that explicitly exclude all 506 candidates and remeasure the broad-line confirmation rate on the held-out set.","supporting_citations":[{"cited_title":"J., Laher, R","cited_arxiv_id":null,"evidence_quote":"Describes the ZTF data-release PSF light curves from which the DR light curves are built."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the reference black hole mass versus stellar mass relation whose shallower normalization the paper contrasts with its own mass ratios."},{"cited_title":"F., Geha, M., & Greene, J","cited_arxiv_id":null,"evidence_quote":"Is the closest previous variability-based low-mass AGN sample; the paper compares confirmation and X-ray match fractions against it."},{"cited_title":"2024, A&A, 681, A97","cited_arxiv_id":null,"evidence_quote":"Establishes the low X-ray match rates for earlier variability-selected samples that the paper explains by selection and depth."}],"review_version":1}