{"id":"0b25e255-ec5f-41f0-998e-0b63b47975f5","arxiv_id":"2505.15819","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A random forest using r-band, g-band, and combined light-curve features selects AGN, with a tuned feature subset recovering 68% of obscured AGN in the VST-COSMOS field.","lead":"This paper trains random forest classifiers on optical light curves from the VLT Survey Telescope in the COSMOS field to pick out active galactic nuclei (AGN) from non-varying stars and galaxies. It reports that adding a second band can improve sample purity but not completeness, and that a tuned feature set raises the recovery of obscured AGN to 68%.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 68.1% obscured-AGN recall in Sect. 4.2 is likely optimistically biased because feature selection and ks8 model choice use the full labeled set before LOOCV; nested cross-validation or an external test is needed.","rationale":"The paper is clear and honest about many choices, and the engineering claim that multiband variability plus colors can select obscured AGN is plausible. The central quantitative result, however, rests on a validation procedure that leaks label information into feature and model selection. The K-S feature pre-filter, the iterative elimination, and the choice of eight features are all performed on the full labeled set before the reported LOOCV evaluation; this makes 68.1% an in-sample selection metric rather than an unbiased estimate of out-of-sample recall. The proposed nested-CV test directly settles whether the leak is material. The reader's weakest assumption already identified this same issue, so I agree with it; my read does not change the verdict, which should remain conditional until such a test is run.","tokens_in":23046,"tokens_out":4738,"duration_ms":40722,"concrete_test":"Recompute the ks8 obscured-AGN recall using nested leave-one-out cross-validation: for each held-out source, run the K-S feature pre-selection (D>0.25), the feature-importance ranking, and the iterative elimination to 8 features on the remaining 2,542 sources only, then predict the held-out source. If the resulting recall is materially below 68.1% (e.g., more than 5–10 percentage points), the headline figure is biased. A complementary check is to apply the ks8 feature set and random-forest threshold fixed on COSMOS to an independent sample, such as variable AGN in SDSS Stripe 82, and compare recall.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 builds the ks8 classifier in three steps that all use the full 2,543-source labeled set: (1) a K-S test between obscured AGN and inactive galaxies selects the 25 features with D>0.25; (2) a feature-importance ranking computed from those data drives iterative elimination from 25 down to 8 features; (3) the eight-feature model is chosen because it gives the highest obscured-AGN recall in Table 4. LOOCV is then applied to the already-fixed ks8 feature set. Each left-out source is therefore not used to fit the tree ensemble, but it has participated in deciding which features enter the model and how many features are kept. This is a classic selection-on-the-full-data leak: the reported 68.1% recall is an in-sample figure of merit, not an unbiased estimate of what ks8 would recover on new sources. The bias is likely largest for obscured AGN because the K-S pre-filter was explicitly optimized to separate exactly that class from inactive galaxies. A secondary inconsistency: the text says 68.1% 'almost doubles' the De Cicco et al. (2021) result, but the introduction quotes 21% recall for Type II AGN in that work; 68.1% is more than triple that value, so the 'almost doubles' comparison is not numerically accurate, although this does not by itself affect the validity of 68.1%.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a random forest classification pipeline for selecting AGN from VST-COSMOS optical light curves, using 162 variability, color, and morphology features measured in the r and g bands, with imputation to 33 visits per band. The labeled set contains 380 AGN (217 Type I, 104 Type II, 211 MIR-selected) and 2163 non-AGN. Using leave-one-out cross-validation (LOOCV), the authors compare several feature sets and then, in Sect. 4.2, use K-S feature selection and iterative elimination to build a compact eight-feature classifier (ks8) that reaches 68.1% recall for spectroscopically confirmed obscured AGN, described as almost doubling the previous result from De Cicco et al. (2021). The paper also investigates the effect of replacing real visits with synthetic imputed visits and finds no large impact on the main metrics.","tokens_in":23523,"tokens_out":8096,"duration_ms":70269,"significance":"If the 68.1% obscured-AGN recall were an unbiased estimate, the result would be genuinely useful for time-domain AGN surveys, since completeness for obscured AGN is a known weakness of optical variability selection. The study also contributes a broad feature set, careful class-imbalance handling, and a transparent comparison across classifiers. However, the headline performance is not yet supported because feature selection and model selection are performed on the full labeled set before LOOCV; the reported recall is an in-sample selection metric. With a nested cross-validation or external test, the method could be valuable for LSST-era applications.","major_comments":[{"comment":"The reported 68.1% obscured-AGN recall is likely an optimistic in-sample figure because feature selection and model selection use the full labeled set before LOOCV. Specifically, the K-S test between obscured AGN and inactive galaxies selects 25 of 162 features, the feature-importance ranking is computed from an RF trained on the same 2543 sources, and the iterative elimination from 25 to 8 features is driven by the recall values shown in Table 4. LOOCV is then applied only to the already-determined ks8 feature set, so every held-out source has contributed to the choice of features and to the number of features retained. Because the K-S pre-filter was explicitly optimized to separate obscured AGN from inactive galaxies, the bias is likely largest for the obscured-AGN recall. The authors should re-estimate performance with nested cross-validation, with feature selection inside each training loop, or on an external validation sample.","section":"Sect. 4.2 and Table 4"},{"comment":"The most important feature in ks8, and in all classifiers tested, is the MIR color ch21, which is the same color entering the Donley et al. (2012) criterion used to build part of the AGN labeled set (211 MIR-selected sources). This creates a risk that the classifier is partly reproducing its own label definition rather than learning a new selection rule. The obscured-AGN recall itself is less directly affected because the 104 Type II AGN are spectroscopically classified, but the general AGN-selection numbers and the feature-importance interpretation are not. I ask for a concrete test: recompute the classification and feature importances either after removing MIR-selected AGN from the positive class or after dropping ch21 from the feature set, and quantify how much of the reported performance depends on this potentially circular feature.","section":"Sect. 2.2 and Fig. 4"}],"minor_comments":[{"comment":"The claim that the 68.1% recall 'almost doubles' the De Cicco et al. (2021) value is not numerically correct: the Introduction quotes 21% recall for Type II AGN, and 68.1/21 ≈ 3.2; please correct the comparison (e.g., 'more than triples') or use an appropriate baseline.","section":"Sect. 4.2"},{"comment":"The synthetic-visit impact test is carried out on r-band data only; because the g-band light curves contain up to 16 synthetic points out of 33, a direct check of the g-band feature distributions would be useful, especially for the g-band features that appear in the K-S preselection (Autocor_length_g, Q31_g, MedianAbsDev_g).","section":"Sects. 3.1 and 4.1"},{"comment":"The reported uncertainties are standard deviations across ten simulations with different random seeds; they do not include the sampling uncertainty of the LOOCV estimate or the variability introduced by feature selection. The text should state this limitation explicitly.","section":"Table 4 and Appendix A"},{"comment":"The degree of overlap between the MIR-selected AGN and the spectroscopic Type I/II samples is not quantified; explicit counts or a Venn diagram would help the reader assess the ch21 circularity concern and interpret the per-subclass recalls.","section":"Sect. 2.2"},{"comment":"The text should be careful that ks8 is chosen for maximum obscured-AGN recall, not because it is uniformly the best classifier; ks9 has higher precision, accuracy, and F1, so the choice is a scientifically motivated trade-off rather than a global optimum.","section":"Sect. 4.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits A&A's scope and the core idea is interesting. The main technical blocker is the evaluation protocol: feature selection on the full data followed by LOOCV makes the headline 68.1% recall an optimistic in-sample estimate. If a nested-CV or external-test estimate confirms a large improvement over previous work, I would be happy to see the paper accepted after revision. The ch21 overlap with the MIR label should also be quantified. I did not find grounds for rejection beyond these fixable issues."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou should know: this is a clean, honest method paper with one load-bearing number that is likely too good. The ks8 classifier's 68.1% obscured-AGN recall is the headline, and I think it's optimistic because feature selection uses the full labeled set before leave-one-out cross-validation. The K-S pre-filter and the iterative elimination down to eight features both saw all 2,543 sources, so each LOOCV training fold already knows which features separate obscured AGN from inactive galaxies. That is classic selection-on-the-full-data leakage. The reported recall is an in-sample figure of merit, not an unbiased estimate for new sources.\n\nWhat the paper does well: it extends the authors' earlier RF pipeline to g-band data, tests bivariate features and finds they don't help, and runs a controlled synthetic-visit experiment showing that replacing real visits with imputed points has only a mild effect on obscured-AGN recall. Those are useful, clear results. The writing is transparent about most choices, and the comparison with the earlier De Cicco et al. (2021) baselines is helpful.\n\nThe soft spots beyond the CV leak: the text says 68.1% 'almost doubles' the 21% recall from De Cicco et al. (2021), but 68.1/21 is 3.24, so 'more than triples' is what they meant. Minor, but it matters for how the result is advertised. Also, the most important feature in every classifier is the MIR color ch21, which overlaps with the Donley et al. (2012) criterion used to select part of the AGN label set. The spectroscopic Type I/II labels are independent, so the obscured-AGN recall is not purely circular, but the feature-importance story is partly self-referential. The 15-day imputation window and the threshold D>0.25 for feature selection are reasonable but arbitrary; they should be varied to show robustness.\n\nThe math and data handling look sound in the sense that I don't see hidden errors or inflated error bars. The sample is small (104 obscured AGN), so the 1.2% uncertainty is just Poisson-like scatter, not a measure of the bias.\n\nBottom line: this is a useful forecasting paper for LSST preparation, and the community will want it in the literature. But the 68% number should be re-derived with nested cross-validation or an external test set before it becomes a planning number. I'd send it to review, with that as the requested revision.","headline":"Useful g-band extension of the VST-COSMOS RF pipeline, but the headline 68% obscured-AGN recall is likely inflated by feature selection on the full labeled set before LOOCV.","tokens_in":23922,"tokens_out":3629,"would_cite":true,"duration_ms":30306,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that reducing a random forest from 162 features to eight—chosen to separate obscured AGN from inactive galaxies—raises the recall of known obscured AGN to about 68%, nearly doubling the earlier result.","keywords":["active galactic nuclei","random forest","optical variability","obscured AGN","multiband photometry","feature selection","COSMOS field","VLT Survey Telescope"],"falsifier":"Take the ks8 feature set to an independent sample of spectroscopically classified AGN from a different survey and measure the obscured-AGN recall; if it falls to roughly 50 percent, the level of the single-band classifiers here, the 68 percent gain is mostly selection on the same data used for scoring rather than a portable advantage.","tokens_in":22887,"feed_emoji":"🔭","tokens_out":9891,"duration_ms":78915,"temperature":0.7,"pith_summary":"This paper asks whether optical variability alone—the stochastic brightness changes that set active galactic nuclei (AGN) apart from quiet stars and galaxies—can be turned into a more complete census of AGN, especially the obscured ones whose light is dimmed and reddened by dust. The authors train random forest classifiers on 2,543 labeled sources from the COSMOS field observed with the VLT Survey Telescope in $r$ and $g$ bands, starting from 162 variability, color, and morphology features. They find that most of those features are unnecessary: a classifier built from only eight features, chosen specifically to separate obscured AGN from inactive galaxies, recovers ($68.1\\pm1.2$)% of known obscured AGN, nearly doubling the recall of the earlier version of the same pipeline, while still finding about 99% of unobscured AGN. The paper presents this as a step toward variability-based AGN selection in upcoming wide-field surveys, where completeness for obscured AGN has been a known weakness.","feed_headline":"Eight features recover 68% of obscured AGN","feed_subtitle":"A random forest using variability, colors, and morphology nearly doubles earlier recall for obscured AGN.","key_machinery":"The engine of the result is a two-stage pruning procedure. Stage one is a Kolmogorov-Smirnov screen: for each of the 162 features, compare the distribution over obscured AGN with the distribution over inactive galaxies, and keep the 25 with distance $D>0.25$. Stage two is iterative elimination: train a random forest with class weights balanced, rank features by impurity-based importance, delete the least important, retrain, and repeat down to seven features. The random forest itself is standard, with leave-one-out cross-validation used to score each setting; the load-bearing element is the choice of the final eight features, which the paper argues separate obscured AGN from inactive galaxies better than the original 162-feature set. The K-S screen also shows which features change most when synthetic visits replace real ones, linking the feature selection to the cadence tests.","core_discovery":"The paper's central claim is that the difficulty of finding obscured AGN through optical variability is primarily a feature-selection problem, not a data problem. By measuring, for every feature, the Kolmogorov-Smirnov distance between the feature distributions of obscured AGN and inactive galaxies, the authors isolate 25 features worth keeping and then iteratively strip the least important ones until eight remain: the mid-infrared color ch21, three optical colors ($u-B$, $r-i$, $i-z$), the HST stellarity index, and three $r$-band variability features (ExcessVar, GP_DRW_sigma, GP_DRW_tau). On the same labeled set, the eight-feature random forest returns ($68.1\\pm1.2$)% recall for obscured AGN and 99.1% recall for unobscured AGN, at the cost of lower precision than feature-rich classifiers. The paper also reports that bivariate features combining $r$ and $g$ light curves add no benefit, that $g$-band features vanish from the final selection, and that replacing up to half of a light curve's visits with linearly interpolated synthetic points does not significantly change the results. These findings are presented as evidence that variability selection can be made competitive for obscured AGN and ready for application to LSST-era data.","pith_inferences":["Editor's inference: because the 25-then-8 feature choice was made on the same labeled set later used for leave-one-out evaluation, the 68.1% recall figure is likely an optimistic in-sample estimate; an independent transfer test could settle the true performance.","Editor's inference: the dominance of $r$-band variability features in the final set, despite the $g$-band being bluer and generally expected to show larger AGN variability, hints that the advantage comes more from the $r$-band's denser real sampling than from physics; a redder-band version of the experiment would test this.","Editor's inference: if the eight-feature set transfers, a practical extension would be to use it as a prior for anomaly detection in streaming alert streams from wide-field surveys, flagging obscured-AGN candidates before spectra are taken.","Editor's inference: the K-S distance table provides a feature ranking that could be repurposed as a cheap pre-filter to remove inactive galaxies, reducing the labeled-set size needed for training on new surveys."],"forward_implications":["Obscured AGN completeness is not locked to survey depth: better feature selection on the same light curves can raise recall from roughly 50% to about 68%.","The final eight-feature set—two variability timescales, an excess-variance measure, four colors, and morphology—is a concrete recipe to test on other surveys.","For LSST-like monitoring, the result implies that a modest number of well-chosen features should be used when completeness for obscured AGN is the goal, rather than the full feature zoo.","Because synthetic imputed visits did not degrade the classifiers, observing strategies that do not sample two bands simultaneously can still feed the same selection method.","The higher obscured-AGN recall is traded against higher contamination, so samples built with the eight-feature classifier will need follow-up to keep purity."],"supporting_citations":[{"why":"Defines the earlier random-forest AGN selection pipeline and the baseline recall figures that the ks8 result is claimed to nearly double.","marker":"De Cicco et al. (2021)"},{"why":"Supplies the random forest algorithm that is the classifier used throughout the paper.","marker":"Breiman 2001"},{"why":"Provides the Chandra-COSMOS Legacy Catalog with the spectroscopic Type I and Type II classes used as the unobscured and obscured ground truth in the labeled set.","marker":"Marchesi et al. (2016)"},{"why":"Defines the mid-infrared color criterion used to add MIR-selected AGN to the labeled set.","marker":"Donley et al. (2012)"},{"why":"Supplies the COSMOS2015 catalog used for inactive-galaxy labels, the optical/NIR colors, and the MIR photometry that feeds the ch21 feature.","marker":"Laigle et al. (2016)"},{"why":"Documents leave-one-out cross-validation, the evaluation protocol whose scores the paper reports.","marker":"Sammut & Webb 2010"},{"why":"Defines the structure-function variability features that appear among the 25 features screened for separating obscured AGN from inactive galaxies.","marker":"Schmidt et al. (2010)"},{"why":"Defines the DRW timescale and sigma features (GP_DRW_tau and GP_DRW_sigma) that survive into the final eight-feature classifier.","marker":"Graham et al. (2017)"}],"fun_headline_variants":["Eight features recover 68% of obscured AGN","Random forest finds 68% of obscured AGN with 8 features","Variability plus colors: 8 features boost obscured AGN recall to 68%","Feature choice, not survey size, unlocks hidden AGN in variability data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The performance numbers assume that choosing the best 25 and then eight features on the full labeled set and evaluating those same sources with leave-one-out cross-validation yields an unbiased estimate of classifier performance on new data.","fun_headline_variants_meta":{"raw":{"variants":["Eight features recover 68% of obscured AGN","Random forest finds 68% of obscured AGN with 8 features","Variability plus colors: 8 features boost obscured AGN recall to 68%","Feature choice, not survey size, unlocks hidden AGN in variability data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000553,"raw_usage":{"total_tokens":2710,"prompt_tokens":1091,"completion_tokens":1619,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":707,"completion_tokens_details":{"reasoning_tokens":1541}},"tokens_in":707,"tokens_out":1619,"duration_ms":10561,"temperature":1.0,"reasoning_tokens":1541,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:10:24.171078+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the ks8 feature set to an independent sample of spectroscopically classified AGN from a different survey and measure the obscured-AGN recall; if it falls to roughly 50 percent, the level of the single-band classifiers here, the 68 percent gain is mostly selection on the same data used for scoring rather than a portable advantage.","supporting_citations":[{"cited_title":"2001, Machine Learning, 45, 5, cited By 34434","cited_arxiv_id":null,"evidence_quote":"Supplies the random forest algorithm that is the classifier used throughout the paper."},{"cited_title":"& Webb, G","cited_arxiv_id":null,"evidence_quote":"Documents leave-one-out cross-validation, the evaluation protocol whose scores the paper reports."}],"review_version":1}