{"id":"743bac17-ff0b-4ea5-bf1a-418f313c1a65","arxiv_id":"2501.08196","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A probabilistic random forest classifies ~130 million VMC sources, yielding >77,000 extragalactic sources and >49,000 new AGN candidates behind the Magellanic Clouds.","lead":"Researchers trained a probabilistic random forest on millions of multi-wavelength observations to classify objects in the Magellanic Clouds into stars, galaxies, and active black holes. The result is a public catalogue with about 78,000 high-confidence extragalactic sources, including tens of thousands of newly identified galaxy and black hole candidates.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline accuracies are measured on a spectroscopically bright-biased, class-imbalanced test split and then applied to the 56.7M-source full-catalogue P>80% selection; without a faint-source validation, the claimed extragalactic counts are not quantitatively secure.","rationale":"The reader's weakest assumption is the same one I find most load-bearing: the training set is bright-biased, and the quoted accuracies are measured on a test split of that training set, not on the full VMC catalogue. I agree with the CONDITIONAL verdict because the paper is transparent, performs multiple independent checks, and does not overclaim in its discussion, but the headline quantitative accuracy is not directly measured on the faint, class-imbalanced catalogue population. The proposed synthetic dimming test is a decisive, low-cost check using existing data: if high-confidence AGN/galaxy accuracy degrades in faint bins, the >49,500 new AGN candidate count becomes unreliable; if it holds, the selection-bias concern is largely resolved. No verdict change is needed relative to the reader's CONDITIONAL assessment. If the check fails badly, the verdict should move toward REJECT; if it passes, ACCEPT would be defensible. UNCHANGED is therefore the appropriate response.","tokens_in":54348,"tokens_out":13205,"duration_ms":153633,"concrete_test":"Run a synthetic dimming experiment on the existing spectroscopic test set: take each source, rescale its fluxes to simulate Ks = 17-20 (or the faint half of the VMC catalogue), add realistic photometric scatter and non-detections matching VMC/SMASH/AllWISE errors, and apply the trained PRF with the same P>80% threshold. Recompute precision/recall for AGN and galaxy classes per magnitude bin and compare with the quoted 90/98%. If faint-bin accuracy falls substantially below the quoted values, the full-catalogue high-confidence counts are not supported; if it holds, the selection-bias objection is largely resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3.4 concedes that the spectroscopic training sets are biased towards sources bright enough for optical spectroscopy, and Section 5.8 shows that most of the faint VMC population lands in Unknown. The central accuracy claims (79%/87%, and 90%/98% at P>80%) are computed on a test split drawn from that same bright spectroscopic sample, not from the 130M-source catalogue. The Table 3 P>80% selection contains 56.7M sources, of which roughly 55.6M are Unknown; applying the Fig. 2 test-set accuracies to this population assumes the PRF probabilities remain calibrated under large covariate shift, for which there is no evidence. This matters directly for the >49,500 new AGN candidates: AGN training labels come mostly from Milliquas, GAMA/SDSS and new SALT/SAAO observations, which are biased toward brighter, often type-1 sources, while the faint tail of the catalogue is where AGB/PNe/YSO contaminants, the main AGN interlopers in Fig. 2, are most likely to be misclassified. The radio/X-ray/Quaia checks are useful but are AGN-selected and do not measure faint-source purity. The concern is not that the method is invalid; it is that the quantitative accuracy of the headline extragalactic counts is not established for the population to which it is applied.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains probabilistic random forest classifiers on spectroscopically labelled sources with 237 multi-wavelength features and applies them to the ~130 million sources of the VMC PSF catalogue, producing separate SMC and LMC classifications into ten astrophysical classes plus an Unknown class. The authors report test-set accuracies of 0.79±0.01 (SMC) and 0.87±0.01 (LMC), rising to 0.90 and 0.98 for sources with P_class > 80%. They report 707,939 (SMC) and 397,899 (LMC) high-confidence non-Unknown sources, including >77,600 extragalactic sources and >49,500 previously unknown AGN candidates, and validate the classifications using X-ray, radio, Quaia, and YSO catalogues.","tokens_in":54686,"tokens_out":6204,"duration_ms":58898,"significance":"If the quantitative claims hold for the full VMC population, the paper delivers a valuable public classification product for a large, deep survey and demonstrates a practical pipeline for separating extragalactic sources from Magellanic stellar populations. The use of held-out test splits, the publication of training data and feature importances, and the cross-checks with independent wavelength regimes are genuine strengths that go beyond a purely internal accuracy assessment. However, the headline counts of new AGN and galaxy candidates depend on unvalidated extrapolation to the faint, class-imbalanced catalogue population, and there is an internal inconsistency in the definition of the high-confidence sample. These issues affect the central quantitative claims and require attention before the results can be taken at face value.","major_comments":[{"comment":"The reported accuracies (79%/87% overall, 90%/98% at P_class>80%) are measured on a test split drawn from the spectroscopically labelled training set, which Section 3.3.4 admits is biased toward sources bright enough for optical spectroscopy. Section 5.8 shows that the majority of the faint catalogue population lands in the Unknown class, yet the headline counts (e.g., >49,500 new AGN candidates) are obtained by applying the classifier to the full catalogue, including the 56.7M sources with P_class>80%, of which roughly 55.6M are Unknown. The independent tests in Sections 5.4–5.6 (radio, X-ray, Quaia) are drawn from populations that are themselves brighter or extreme and do not measure the purity of the faint high-confidence extragalactic selection. The claim that the high-confidence extragalactic counts are accurate therefore rests on an assumption of robustness to covariate shift that is not tested. The authors should either restrict the quantitative claims to the magnitude range actually populated by the training set, or provide a faint-source validation (e.g., a spectroscopic or SED-based check on a random faint subset).","section":"§3.3.4, §5.8, Table 3"},{"comment":"The definition of the high-confidence catalogue is internally inconsistent. Section 4.2 states that sources with combined AGN+Galaxy probability ≥80% (56,134 for SMC and 140,212 for LMC) are 'moved to the high confidence catalogue', and that the rest of the paper refers to these high-confidence sources. However, the totals quoted in the Abstract and in Table 3—707,939 (SMC) and 397,899 (LMC) known sources, and >77,600 extragalactic sources—correspond exactly to the P_class>80% rows of Table 3 and do not include these combined-threshold sources. Moreover, the accuracy of the combined-threshold selection is not measured by the confusion matrices in Figure 2, which are restricted to P_class>80%. The paper thus mixes a validated subset (individual P_class>80%) with an unvalidated, differently defined subset in the description of the final catalogue, and the headline counts do not reflect the procedure described. Please clarify which catalogue the numbers refer to, and provide accuracy estimates for the combined-threshold selection if it is retained.","section":"§4.2, Table 3, Abstract"},{"comment":"The statement in the conclusions that 'all the AGN were classified correctly' for the dust-dominated AGN sample is an overstatement. Table 4 shows that most of the sources in this sample, including all of the SMC sources classified as AGN and many of the LMC AGN, were included in the training set (column 'T?' = Y). The held-out sources (T? = N) include several that are not confidently classified as AGN (e.g., SMCtSNE8 at P=0.45; LMCtSNE13 as Hii/YSO at P=0.21; LMCtSNE16 as Hii/YSO at P=0.41). Section 5.1 itself acknowledges the role of training-set inclusion, but the conclusion bullet in Section 6 does not carry this caveat. Please either restrict the claim to the held-out subset or report the performance separately for training and held-out sources.","section":"§5.1, Table 4, §6"}],"minor_comments":[{"comment":"The description of combining the validation and test sets into a single test set is acceptable given the repeated random splits, but the wording could be tightened to clarify that parameter tuning and final evaluation are not performed on the same fixed split.","section":"§3.3.1"},{"comment":"The construction of the Unknown class from randomly selected VMC sources is a key modelling assumption; it is described clearly in the text, but it would be helpful to state explicitly in the conclusions that the reliability of the Unknown class as a catch-all depends on this assumption, which is not directly testable with the current data.","section":"§2.2.6"},{"comment":"The caption should specify whether the columns labelled 'P_class>80%' include sources moved to the high-confidence catalogue via the combined AGN+Galaxy probability criterion of Section 4.2; the current text implies they do not, which contradicts the procedure described in the body.","section":"Table 3 caption"},{"comment":"Since Section 4.2 introduces a combined AGN+Galaxy probability threshold for the high-confidence catalogue, it would be helpful to show the confusion matrix for the resulting combined selection, or to state explicitly that the reported accuracies apply only to the individual P_class>80% selection.","section":"Figure 2"},{"comment":"There are a number of typographical errors, e.g., 'deccreasing' in Section 5.4, 'inclde' in Section 5.5, and 'anologue' in Section 3.2; a careful proofread is recommended.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a useful classification product and a generally careful ML pipeline, and the external validation tests are a strong point. The main issues are the extrapolation of test-set accuracies to the faint catalogue population and the inconsistency between the high-confidence definition in Section 4.2 and the numbers quoted in the Abstract and Table 3. Both are fixable: the authors could rephrase the claims to avoid overstating the faint-source accuracy, or provide a dedicated faint-source validation. The training-set circularity in Section 5.1 is acknowledged and can be handled by reporting held-out performance separately. The paper is of clear interest to the survey and extragalactic communities, but the headline quantitative claims should not be accepted without addressing these points."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou should know two things about this paper. First, it is a careful, transparent application of the probabilistic random forest to classify the full VMC PSF catalogue, producing a large new catalogue of >77,600 extragalactic sources and >49,500 AGN candidates. Second, the quoted accuracies (79%/87%, and 90%/98% at P>80%) are measured on a held-out split of the spectroscopically labelled training set, which is biased toward bright, optically observed sources. The accuracy for the ~130M catalogue sources, especially the faint ones that dominate the Unknown class, is directly unmeasured. That limitation is real and the paper does not hide it; sections 3.3.4 and 5.8 are explicit about it.\n\nWhat is genuinely good: the work combines UV-to-FIR photometry with careful construction of a training set including new SALT and SAAO spectroscopy; it runs the standard checks one would want—X-ray and radio matches behave as expected, Quaia AGN mostly come back as AGN, YSOs mostly do not leak into the extragalactic classes. The high-confidence subset does show improved precision in confusion matrices. The paper is honest about the Unknown class being dominant and about the training-set brightness bias. It releases training sets and will release the classification catalogue.\n\nThe soft spots are real but not fatal. The main one is exactly the stress-test concern: a 25% test split of the same biased training set cannot calibrate the classifier under covariate shift to the faint, mostly-missing-photometry population. The >49,500 new AGN count is likely a lower limit and the purity for faint sources is not established. A second, minor soft spot is the post-hoc threshold that moves AGN+Galaxy combined-probability sources into the high-confidence catalogue; this is useful but does not have an independent accuracy estimate. Section 5.1's validation on dust-dominated AGN is partly circular because some of those sources were in the training set—though the paper notes which were not.\n\nThe structure and writing are clear, and the citation pattern is appropriate; the paper does not oversell. It is a useful contribution for anyone working on background galaxies/AGN behind the Magellanic Clouds or on ML classification of large surveys. A serious referee can add value by pressing for a fainter validation sample or at least a clearer statement that the headline accuracies do not extrapolate to the full catalogue.\n\nI would send this to review, with the expectation of a revision that either adds a faint validation set or softens the accuracy claims. I would bring it to our reading group, and I would cite it with the accuracy caveat.\n\nBest.","headline":"A solid, honest ML classification of the VMC catalogue with useful new AGN/galaxy samples; headline accuracies are proven only for bright spectroscopic sources, so the faint-tail counts are the main caveat.","tokens_in":55275,"tokens_out":2765,"would_cite":true,"duration_ms":28446,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A probabilistic random forest sorts 130 million Magellanic Cloud sources into stars, galaxies, and AGN, recovering 77,600 extragalactic sources including 49,500 new AGN candidates.","keywords":["probabilistic random forest","Magellanic Clouds","active galactic nuclei","galaxy classification","multi-wavelength photometry","VMC survey","machine learning","young stellar objects"],"falsifier":"Spectroscopically observe a random sample of the PRF's high-confidence AGN and galaxy candidates in the faint regime where the training set is thinnest (for example $G > 18.8$ mag or $K_s > 18$ mag) and measure the confirmation rate. If that rate falls far below the ~90–98% accuracy quoted for the bright test set, the headline extragalactic counts would be inflated.","tokens_in":54179,"feed_emoji":"🌌","tokens_out":17648,"duration_ms":147198,"temperature":0.7,"pith_summary":"The VMC (VISTA Survey of the Magellanic Clouds) near-infrared survey has detected more than 130 million sources toward the Small and Large Magellanic Clouds (SMC and LMC), most of which have never been individually identified. This paper claims that a probabilistic random forest — a supervised machine-learning classifier that handles missing data and returns a probability per class — trained on spectroscopically confirmed examples and fed 237 photometric features from the ultraviolet to the far-infrared, can sort that entire catalogue into ten astrophysical classes (AGN, galaxies, OB stars, red-giant-branch stars, asymptotic-giant-branch stars, red supergiants, post-AGB/post-RGB stars, planetary nebulae, HII regions/young stellar objects, and foreground high-proper-motion stars) plus an 'Unknown' class. On held-out test sources the classifier reaches average accuracies of about 79% (SMC) and 87% (LMC), improving to about 90% and 98% when only sources with class probability $P_{\\rm class} > 80\\%$ are kept. The high-confidence sample contains more than 77,600 extragalactic sources behind the Clouds, including over 49,500 new AGN candidates and roughly 26,500 new galaxy candidates. If these labels are right, a survey built to study stars in two nearby galaxies doubles as a working source of background-galaxy and AGN science.","feed_headline":"Random forest finds 49,500 new AGN behind the Magellanic Clouds","feed_subtitle":"It flags 77,600 extragalactic sources behind the Clouds, including 49,500 new AGN candidates.","key_machinery":"The load-bearing mechanism is the probabilistic random forest (PRF), a random-forest variant that handles missing data by propagating each object down both branches of a decision node, weighted by the probability that each path is correct, and that outputs a probability per class rather than a single hard label. Around it the paper builds a 237-dimensional feature space: magnitudes, photometric errors, point-spread-function sharpness values, Gaia proper motions, colours formed from all pairwise band differences, and the local far-infrared background at 250 microns, which acts as a proxy for extinction and for how close a source sits to the centre of a Cloud. A deliberately unstructured 'Unknown' class, formed by randomly drawing sources from the VMC catalogues themselves, absorbs anything the spectroscopic training set does not represent; this is what lets the pipeline report honest confidence, because the faint majority of sources land in Unknown while the headline extragalactic counts come from the subset with class probability above 80%.","core_discovery":"The paper's central claim is that the probabilistic random forest assigns reliable astrophysical labels, with honest per-class probabilities, to essentially the whole VMC catalogue. The authors report average test-set accuracies of $0.79 \\pm 0.01$ (SMC) and $0.87 \\pm 0.01$ (LMC), rising to $0.90$ and $0.98$ when restricted to sources with $P_{\\rm class} > 80\\%$, and they report that 99.8–99.9% of true extragalactic test sources are placed into extragalactic classes at high confidence. After removing the Unknown class they classify 707,939 (SMC) and 397,899 (LMC) sources at high confidence, of which more than 77,600 are extragalactic: over 49,500 previously unknown AGN candidates, over 26,500 new galaxy candidates, and over 2,800 new young-stellar-object candidates. Independent checks support the physical content of the labels: the majority of X-ray sources (554/883) are classified as AGN, the majority of radio sources (1756/2694) as AGN with a further 659/2694 as galaxies, and about 86% of the spectroscopically confirmed quasars in the Quaia catalogue receive AGN labels.","pith_inferences":["Combining the PRF probabilities with the X-ray and radio flags that were deliberately kept out of the feature space would yield a cheap, testable obscured-AGN candidate list; the paper already shows that X-ray- or radio-detected Unknowns concentrate in AGN-like regions of magnitude–flux space.","The brightness bias of the spectroscopic training set means the high-confidence extragalactic counts are not directly usable as population statistics; converting them into number counts or luminosity functions would require a completeness correction calibrated on the Unknown class.","Retraining without the 250-micron far-infrared background, which ranks as the top or near-top feature and encodes position and extinction rather than the source's own emission, would reveal how much of the extragalactic separation rests on where a source sits in the Cloud, and how portable the classifier is to fields outside the Magellanic Clouds.","Because the published catalogue carries a probability for every class, the Unknown probability can serve as a continuous novelty score for prioritising unusual or underrepresented sources in future spectroscopic campaigns."],"forward_implications":["More than 49,500 previously unknown AGN candidates, over 26,500 new galaxy candidates, and over 2,800 new young-stellar-object candidates become prioritized targets for spectroscopic follow-up.","Radio and X-ray detected sources are predominantly classified as AGN or galaxies, providing reliable multi-wavelength counterparts and tentative AGN labels for X-ray- or radio-detected Unknowns.","The spatial distributions match physical expectation — extragalactic sources spread uniformly, Magellanic stellar sources concentrate toward the Cloud centres — which supports the validity of the labels for population studies.","Stellar classes absent from the training set (Wolf-Rayet stars, R Coronae Borealis stars, supernova remnants) are classified as other stellar classes rather than as extragalactic, implying low stellar contamination of the AGN and galaxy samples.","The Unknown class maps where the training set is incomplete — mostly faint main-sequence and RGB stars — indicating where new spectroscopy of faint sources would most improve the classifier."],"supporting_citations":[{"why":"Supplies the probabilistic random forest algorithm, the classifier at the heart of the work.","marker":"Reis et al. 2018"},{"why":"Defines the VMC survey whose ~130 million near-infrared PSF sources are the objects being classified.","marker":"Cioni et al. 2011a"},{"why":"Provides the SMASH optical photometry that contributes key features to the multi-wavelength dataset.","marker":"Nidever et al. 2017"},{"why":"Provides the SAGE Spitzer mid-infrared photometry of the LMC used as classifier features.","marker":"Meixner et al. 2006"},{"why":"Provides the SAGE Spitzer mid-infrared photometry of the SMC used as classifier features.","marker":"Gordon et al. 2011"},{"why":"SAGE-Spec spectroscopy that supplies SMC stellar training labels.","marker":"Ruffle et al. 2015"},{"why":"SAGE-Spec spectroscopy that supplies LMC stellar training labels.","marker":"Jones et al. 2017"},{"why":"The Milliquas catalogue from which the spectroscopically confirmed AGN training samples are drawn.","marker":"Flesch 2019b"},{"why":"The 6dFGS survey that supplies the galaxy training labels.","marker":"Jones et al. 2009"},{"why":"The Quaia quasar catalogue used as an independent check of the AGN classifications.","marker":"Storey-Fisher et al. 2023"}],"fun_headline_variants":["AI finds 49,500 new AGN behind Magellanic Clouds","Random forest uncovers 49,500 AGN behind Magellanic Clouds","Machine learning flags 77,600 galaxies behind Magellanic Clouds","Probabilistic random forest nets 49,500 new AGN","49,500 AGN candidates discovered behind Magellanic Clouds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The spectroscopically confirmed training sources are taken to represent the entire 130-million-source catalogue, even though spectroscopy is systematically biased toward bright objects; the reported ~79–87% and ~90–98% accuracies are measured only on that brighter test split, so the accuracy on the faint majority of the catalogue is not directly known.","fun_headline_variants_meta":{"raw":{"variants":["AI finds 49,500 new AGN behind Magellanic Clouds","Random forest uncovers 49,500 AGN behind Magellanic Clouds","Machine learning flags 77,600 galaxies behind Magellanic Clouds","Probabilistic random forest nets 49,500 new AGN","49,500 AGN candidates discovered behind Magellanic Clouds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000503,"raw_usage":{"total_tokens":2592,"prompt_tokens":1214,"completion_tokens":1378,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":830,"completion_tokens_details":{"reasoning_tokens":1285}},"tokens_in":830,"tokens_out":1378,"duration_ms":9973,"temperature":1.0,"reasoning_tokens":1285,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:29:04.209154+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Spectroscopically observe a random sample of the PRF's high-confidence AGN and galaxy candidates in the faint regime where the training set is thinnest (for example $G > 18.8$ mag or $K_s > 18$ mag) and measure the confirmation rate. If that rate falls far below the ~90–98% accuracy quoted for the bright test set, the headline extragalactic counts would be inflated.","supporting_citations":[],"review_version":1}