{"id":"2122e38f-710f-4d17-9242-02a6353bcce0","arxiv_id":"1909.00606","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"On Stripe 82X AGN with optical-to-MIR photometry, machine-learning photo-z from MLPQNA matches SED fitting in accuracy and outliers, and its redshift probability distributions are more conservative for outliers.","lead":"This paper tests a machine-learning method for estimating distances (photometric redshifts) of X-ray-selected active galaxies using optical, near-infrared and mid-infrared data, and compares it to template fitting. The authors find the two approaches perform comparably on data as deep as the upcoming eROSITA all-sky survey, and release a new photo-z catalogue.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The eROSITA projection is conditional on a representative training sample, and the paper's own blind test on 257 new fainter sources shows MLPQNA accuracy degrades sharply (sigma_NMAD 0.104-0.163, eta 32-41%) when that condition fails; the central extrapolation is therefore not yet supported.","rationale":"The reader identified the representative-training-sample assumption as the weakest link, and the paper's own blind test on new fainter sources confirms that this assumption is currently unmet. I agree that this is the load-bearing concern: the central comparison claim is about performance on a training-represented population, while the eROSITA projection requires generalization to a broader, fainter population. The in-field accuracy numbers in Table 8 and the comparison with LePhare on the original sample are credible and supported by the blind cross-validation design and the released catalogue. A secondary concern is that feature selection performed before the train/test split in Tables 3-4 could bias the reported gains from PhiLAB, but it does not directly affect the headline comparison in Table 8, which is based on full photometric sub-samples. The verdict CONDITIONAL is therefore appropriate: the paper's core comparison is likely robust, but the eROSITA-scale projection should be re-evaluated after a representative training sample is actually constructed and tested on fainter sources.","tokens_in":25168,"tokens_out":9299,"duration_ms":205672,"concrete_test":"Use the 257/258 LaMassa et al. (2019) redshifts as a held-out population. Train MLPQNA on the original A17 spectroscopic sample augmented with a randomly chosen half of the new sources, using the same photometric features and four-fold cross-validation protocol as in Section 6, then measure sigma_NMAD and eta on the held-out half, repeating over many random splits. If the metrics approach the original values (sigma_NMAD approximately 0.056, eta approximately 13%), the Table 10 degradation is caused by training-set representativeness and the eROSITA projection is plausible once representative spectroscopy is gathered. If the metrics remain at the Table 10 level (sigma_NMAD above 0.10, eta above 30%), the method itself fails on fainter sources and the eROSITA projection is not supported by the current analysis.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The core in-field comparison on Stripe 82X is reasonably supported by blind cross-validation, but the headline eROSITA claim, that reliable photo-z can be obtained for a large fraction of Southern-hemisphere sources before spectroscopic follow-up, depends on the ability to assemble a spectroscopic training sample representative of the full eROSITA population. That precondition is not demonstrated. The paper's own Table 10 is the decisive evidence: the 257 new, fainter spectroscopic redshifts from LaMassa et al. (2019) were not represented in the training set, and MLPQNA accuracy drops from sigma_NMAD approximately 0.056 and eta approximately 12.7% on the original sample to sigma_NMAD = 0.104-0.163 and eta = 32-41% across photometric subsets. Because eROSITA's all-sky sample will be dominated by faint sources, and existing spectroscopy is biased toward brighter, optically selected objects, the assumption that sufficient spectroscopy to build a representative training sample can be gathered is an external condition that the paper does not establish. The authors acknowledge this in Section 8, but acknowledging an unmet precondition does not evidence the projection. The proper status is CONDITIONAL: the comparison of MLPQNA and LePhare on the original Stripe 82X sample can stand, while the eROSITA-scale extrapolation remains unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper tests whether machine-learning photometric redshifts (photo-z), computed with the MLPQNA neural network, are competitive with SED-fitting (LePhare) for X-ray-selected AGN, using the multi-wavelength Stripe 82X catalogue. The authors perform a feature-selection analysis with PhiLAB, compute photo-z for various photometric sub-samples, derive redshift probability density functions with METAPHOR, and release a photo-z catalogue. Their main quantitative claim is that when optical, near-IR, and mid-IR photometry are all available, MLPQNA and LePhare perform comparably in accuracy, outlier fraction, and PDZ realism, with both methods degrading when photometric coverage is reduced. The paper also projects that reliable photo-z can be obtained for a large fraction of eROSITA sources in the Southern hemisphere before spectroscopic follow-up, conditional on assembling a representative spectroscopic training sample.","tokens_in":25442,"tokens_out":3703,"duration_ms":308545,"significance":"If the underlying comparison holds, the paper is a useful step toward producing photo-z for the roughly three million AGN expected from eROSITA, where SED fitting alone may be too slow or too template-dependent. The manuscript's strengths include a released photo-z catalogue, a blind four-fold cross-validation design on the original sample, and an external test on 257 newly obtained spectroscopic redshifts that the authors themselves use to expose the limitations of a non-representative training set. The explicit discussion of the training-sample requirement in Section 8 is honest and is a constructive contribution to the methodology for AGN photo-z. The feature-selection analysis with PhiLAB also provides practical guidance on which photometric bands and colours matter most. However, as discussed below, the feature-selection procedure is not fully blind and the eROSITA-scale extrapolation is not yet supported by the data; these issues are fixable but require revision.","major_comments":[{"comment":"The feature-selection step with PhiLAB is performed on the full BEST sample before the four-fold split into training and test folds. Because the same objects that later appear in the test folds contribute to the selection of the feature set, the reported sigma_NMAD and outlier fractions for MLPQNA are not obtained under a fully blind procedure: the model is blind to the test redshifts, but the feature engineering is not. This inflates the apparent performance of MLPQNA and weakens the direct comparison with LePhare. The authors should either perform feature selection inside each training fold (nested cross-validation) or provide evidence that the feature ranking is insensitive to the inclusion of the test-fold objects.","section":"Section 3.1 and Section 4, Tables 3-9"},{"comment":"The paper's central eROSITA projection is conditional on assembling a spectroscopic training sample representative of the eROSITA population, but that condition is not demonstrated. The paper's own Table 10 shows the consequence when the condition fails: for the 257 new fainter sources from LaMassa et al. (2019), sigma_NMAD rises to 0.104-0.163 and the outlier fraction to 32-41% for MLPQNA across photometric subsets. This is the most direct empirical evidence available about extrapolation to the faint eROSITA population, and it does not support the statement that reliable photo-z can be obtained for a large fraction of Southern-hemisphere sources before spectroscopic follow-up. The authors should either temper the conclusion to present this as a forward requirement rather than an achieved capability, or provide a concrete assessment of how the required representative training sample could be assembled and what sample size or completeness would mitigate the degradation seen in Table 10.","section":"Section 6, Table 10, and Section 8"}],"minor_comments":[{"comment":"The number of new spectroscopic redshifts is given as 257 in the text, 258 in Table 10, and 258 in the caption of Figure 2; please make these numbers consistent.","section":"Section 6 and Table 10"},{"comment":"The photometric-error cut experiments change both the sample size and the redshift distribution, so the small variations in sigma_NMAD and eta among the three cuts are not clearly significant; a brief statement about the statistical uncertainty of these differences would help.","section":"Section 5.1, Table 5"},{"comment":"The comparison of METAPHOR and LePhare PDZs depends on the chosen bin sizes (0.01 for both, but with different ranges and normalization conventions). The paper acknowledges this in part, but the claim that METAPHOR is 'superior' in all classes would be strengthened by a sensitivity test to binning.","section":"Section 7, Table 11"},{"comment":"Table 9 reports statistics for sources in common across all samples but does not state the number of such sources; adding this number would make the table clearer.","section":"Table 8 and Table 9"},{"comment":"The phrase 'they are not the same for the two methods' is vague; specifying whether the difference refers to the outlier populations, the PDZ widths, or the systematic offsets would make the abstract more informative.","section":"Section 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest about the training-sample limitation in Section 8, and the release of the photo-z catalogue is a positive feature. The main reasons for major revision are the non-blind feature selection and the gap between the blind-test degradation in Table 10 and the eROSITA-scale projection. Both are addressable within the manuscript's scope, so I do not recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this paper earns its place as a solid, honest comparison of machine-learning and SED-fitting photo-z for X-ray-selected AGN on Stripe 82X, and it ships a public catalogue. The core claim, that MLPQNA and LePhare are comparable when optical, NIR and MIR photometry are all available, is backed by 4-fold blind cross-validation and holds up. The eROSITA extrapolation is conditional and explicitly flagged as such, but the paper's own test on 257 fainter redshifts (Table 10) shows how quickly ML accuracy degrades when the training sample is not representative, so the extrapolation remains unverified.\n\nThe new content: a systematic comparison at eROSITA depth, the released photo-z catalogue, the use of PhiLAB feature selection for AGN, and a PDZ reliability analysis showing LePhare PDZs are overconfident for outliers when only broad-band photometry is available. That PDZ result is practically useful and worth following up, though it depends on binning choices the authors acknowledge.\n\nThe main soft spots, in order. First, the feature selection with PhiLAB is done on the full BEST sample before splitting into folds. That leaks test-fold information into the chosen feature set. The effect is probably not huge, given the features chosen are sensible and the performance difference across subsets is consistent, but it should be quantified by moving the selection inside the CV loop or showing the results are stable to re-selection. Second, the eROSITA projection is conditional on assembling a representative training sample, and Table 10 is the decisive caveat: sigma_NMAD jumps to 0.104-0.163 and outliers to 32-41% on the fainter LaMassa et al. sample. The authors flag this in Section 8, but a flag is not evidence. Third, the comparison to Ruiz et al. and Meshcheryakov et al. is a little thin; a few extra sentences on why results differ would help.\n\nWho it is for: anyone preparing photo-z for eROSITA, or working on AGN photo-z in wide surveys. The catalogue is a useful resource, and the explicit demonstration of where ML fails (missing bands, unrepresentative training, point-like sources) is worth having on record.\n\nI'd send it to peer review. The two things a referee should push on are the feature-selection leakage and the gap between the conditional eROSITA statement and the actual evidence. Both are fixable in revision.","headline":"Solid ML-vs-SED photo-z comparison with a released catalogue; the central claim holds, but the eROSITA projection is conditional on a training sample that the authors' own new test shows is not yet in hand.","tokens_in":26059,"tokens_out":3442,"would_cite":true,"duration_ms":30229,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Machine learning matches template fitting for photometric redshifts of X-ray-selected AGN, with more conservative error estimates.","keywords":["photometric redshifts","active galactic nuclei","eROSITA","machine learning","neural networks","SED fitting","Stripe 82X","redshift probability density functions"],"falsifier":"Go to the released photo-z catalogue and compare the merged MLPQNA photo-z against the 257 new spectroscopic redshifts from LaMassa et al. (2019): the paper reports $\\sigma_{\\rm NMAD}$ of 0.154 and an outlier fraction of 38.4%, versus 0.056 and 12.7% on the original training-representative sample. If that fainter sample is representative of what eROSITA will detect, the paper's claim that reliable photo-z can be obtained for a large fraction of eROSITA AGN is contradicted by its own numbers.","tokens_in":24938,"feed_emoji":"🔭","tokens_out":9679,"duration_ms":72608,"temperature":0.7,"pith_summary":"The paper sets out to show that machine learning can produce photometric redshifts for X-ray-selected active galactic nuclei (AGN) that are just as accurate and reliable as the traditional template-fitting approach, even though AGN are notoriously hard to fit because both the host galaxy and the nucleus contribute light. This matters because eROSITA will detect about three million AGN across the whole sky and needs quick, trustworthy distance estimates from patchy multi-wavelength data. The authors test this on the Stripe 82X field, which matches eROSITA's depth and data coverage, using the MLPQNA neural network. They find that when optical, near-infrared, and mid-infrared photometry are all available, machine learning and SED fitting perform comparably in accuracy, outlier fraction, and the realism of redshift probability distributions. They also find that the machine-learned probability distributions are more conservative, giving faint and unreliable sources appropriately low confidence, whereas template fitting can be overconfident for its outliers.","feed_headline":"Machine learning matches template fitting for AGN redshifts","feed_subtitle":"For eROSITA's 3 million AGN, machine-learned redshifts match templates — and their error bars are more honest.","key_machinery":"The central object is MLPQNA, a multi-layer perceptron neural network trained with a quasi-Newton algorithm, which maps photometric magnitudes and colours into a redshift estimate. Around it, the paper uses PhiLAB, a hybrid feature-selection algorithm based on shadow features and LASSO, to identify the most informative photometric features, and METAPHOR, a workflow that perturbs the photometry 999 times to build a redshift probability density function from 1000 machine-learning estimates.","core_discovery":"On the central claim: in Stripe 82X, when SDSS, VHS, WISE, and IRAC photometry are all available, the MLPQNA neural network reaches a normalized median absolute deviation of $\\sigma_{\\rm NMAD}=0.056$ and an outlier fraction of 12.7 percent, compared with 0.059 and 13.3 percent for the SED-fitting catalogue it is tested against. At the X-ray flux limit that eROSITA will reach, the machine-learning results are slightly better—fewer outliers and no systematic bias—when the training sample is restricted to the bright sources eROSITA will detect. The paper also claims that the redshift probability density functions produced by METAPHOR are reliable and generally more conservative than those from SED fitting, which often assigns high confidence to its own outliers, and it recommends that science use the full probability distributions rather than point estimates. It explicitly states that the remaining bottleneck is the representativeness of the training sample: on a blind test of 257 new fainter sources, both methods degrade, with MLPQNA's outlier fraction rising to 32–41 percent.","pith_inferences":["The disagreement between ML and SED-fitting outliers could be used as a flag for peculiar or variable sources, since the two methods rarely fail on the same objects.","The feature analysis's finding that colours dominate over single magnitudes suggests that other AGN surveys should prioritise multi-band colour coverage over deeper single-band photometry.","The overconfident PDZ from template fitting in broad-band-only data implies that luminosity functions and clustering measurements built from such fits may underestimate their redshift errors; this can be tested by comparing to samples with narrow-band photometry."],"forward_implications":["If eROSITA has a representative spectroscopic training set, reliable ML photometric redshifts can be computed for roughly two-thirds of its AGN using current all-sky photometry, before optical spectroscopy arrives.","The gap between the 12.7% outlier fraction on the original sample and the 32–41% on the fainter blind test means training representativeness is the deciding factor for the eROSITA forecast.","With the already-deeper unWISE data and future SpherEx coverage, the accuracy of ML photo-z should improve further.","When fewer photometric bands are available, SED fitting remains the more reliable method, so the two approaches are complementary rather than interchangeable."],"supporting_citations":[{"why":"Provides the Stripe 82X multi-wavelength catalogue and the SED-fitting photo-z and PDZ baseline used for comparison.","marker":"A17"},{"why":"Original catalogue of X-ray sources and counterparts in Stripe 82X that defines the sample.","marker":"LaMassa et al. 2016"},{"why":"New spectroscopic redshifts of 257 fainter sources used as the additional blind test that reveals training-set dependence.","marker":"LaMassa et al. 2019"},{"why":"Introduces the MLPQNA neural network algorithm that is the machine-learning method being tested.","marker":"Brescia et al. 2013"},{"why":"Introduces METAPHOR, the workflow used to compute machine-learned redshift probability density functions.","marker":"Cavuoti et al. 2017"},{"why":"Introduces PhiLAB, the feature-selection algorithm used to identify the most informative photometric features.","marker":"Delli Veneri et al. 2019"},{"why":"Referenced as the way to handle photometric errors within machine-learning photo-z, by adding errors as parameters.","marker":"Reis et al. 2019"},{"why":"Earlier machine-learning photo-z for X-ray-selected sources; the paper compares its results with this work.","marker":"Ruiz et al. 2018"}],"fun_headline_variants":["ML redshifts match templates for eROSITA AGN","Neural nets tie SED fitting for AGN photo-z","AGN photo-z: ML vs templates, a statistical dead heat","For eROSITA AGN, ML photo-z equals templates in accuracy","Training sample is the real limit for ML AGN redshifts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's central eROSITA projection assumes that a large, representative spectroscopic training sample covering the full range of AGN types, luminosities, and redshifts can be assembled; the blind test on the 257 new fainter sources shows that when this fails, ML accuracy degrades to outlier fractions of 32–41%.","fun_headline_variants_meta":{"raw":{"variants":["ML redshifts match templates for eROSITA AGN","Neural nets tie SED fitting for AGN photo-z","AGN photo-z: ML vs templates, a statistical dead heat","For eROSITA AGN, ML photo-z equals templates in accuracy","Training sample is the real limit for ML AGN redshifts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000255,"raw_usage":{"total_tokens":1644,"prompt_tokens":1092,"completion_tokens":552,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":708,"completion_tokens_details":{"reasoning_tokens":466}},"tokens_in":708,"tokens_out":552,"duration_ms":6090,"temperature":1.0,"reasoning_tokens":466,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:41:24.746747+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Go to the released photo-z catalogue and compare the merged MLPQNA photo-z against the 257 new spectroscopic redshifts from LaMassa et al. (2019): the paper reports $\\sigma_{\\rm NMAD}$ of 0.154 and an outlier fraction of 38.4%, versus 0.056 and 12.7% on the original training-representative sample. If that fainter sample is representative of what eROSITA will detect, the paper's claim that reliable photo-z can be obtained for a large fraction of eROSITA AGN is contradicted by its own numbers.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces METAPHOR, the workflow used to compute machine-learned redshift probability density functions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier machine-learning photo-z for X-ray-selected sources; the paper compares its results with this work."}],"review_version":1}