{"id":"0abf62d0-e538-417d-b96e-364e7dbf3c56","arxiv_id":"2603.19882","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":13,"one_line_summary":"For KiDS quasars, multi-component MDNs and BNNs reconstruct redshift distributions better than plain ANNs, but the best model depends on the data-quality regime and combined faint/missing-band cases degrade all models.","lead":"Using KiDS photometry and DESI spectra, this paper benchmarks three machine-learning approaches — plain networks, mixture-density networks, and Bayesian networks — for quasar photo-z uncertainties and redshift distribution reconstruction. It finds that no model wins in every regime and that uncertainty-aware models matter most for tomographic binning, which is directly relevant to upcoming surveys such as Euclid and LSST.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract claims single Gaussian is best-calibrated in faint+missing regime; Table 2 shows it has worst NLL (1.28 vs 0.69 for BNN 3).","rationale":"The reader's verdict is CONDITIONAL, and I agree that the paper cannot be accepted as-is. However, the reader's weakest_assumption focuses on DESI selection effects as an unbiased ground truth; while that is a legitimate concern, the more immediate and load-bearing issue is the direct contradiction between the abstract's claim that a single Gaussian is best-calibrated in the hardest regime and Table 2, where MDN 1 has the worst NLL there. The central claim that multi-component MDNs/BNNs are essential partly rests on the assertion that no single-component model is sufficient; if the abstract's claim were true, it would undermine that assertion. The body's Table 2 suggests the abstract is wrong, but the abstract is part of the submitted manuscript and must be internally consistent. This issue is more fundamental than the selection-effect concern because it affects the credibility of all reported numbers, not just the generalization to fainter samples. The concrete test—recomputing NLL and PIT on the faint+missing subset—would settle whether the abstract or the table is erroneous. The reader did note this inconsistency in the rationale but did not make it the weakest_assumption; hence partial agreement. The proper outcome is still CONDITIONAL (revise), so I keep the verdict unchanged rather than escalate to REJECT, as a corrected abstract or a reconciled metric could resolve the conflict without invalidating the underlying comparison.","tokens_in":10897,"tokens_out":5822,"duration_ms":55186,"concrete_test":"Recompute the NLL on the faint-extrapolation, missing-features subset (test set iv) for MDN 1, MDN 3, and BNN 3, using the same model checkpoints and imputation (Section 3.2). Also compute the PIT values for these models on that subset. If MDN 1's NLL is 1.28 as in Table 2, the abstract statement that 'the single Gaussian becomes the best-calibrated' is contradicted; if the NLL differs, Table 2 contains a reporting error. If the PIT indicates MDN 1 is best-calibrated despite the higher NLL, the paper must explicitly reconcile the two metrics; the current submission does neither.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is an internal consistency failure between the paper's abstract and its reported results. The abstract states: 'while in the hardest faint-plus-missing regime the single Gaussian becomes the best-calibrated.' Table 2 reports NLL for that exact regime: MDN 1 (single Gaussian) = 1.28, MDN 3 = 0.74, BNN 3 = 0.69. Thus the single-Gaussian model has the worst, not the best, NLL on that subset. Moreover, the abstract's opening claim that 'a single Gaussian is miscalibrated and produces more catastrophic outliers than the ANN' does not harmonize with calling it best-calibrated in any regime. The central conclusion—that MDNs with at least two components are essential and that the best model is regime-dependent—depends on the internal consistency of the evaluation. If the abstract's 'best-calibrated' claim is based on a different metric (e.g., PIT), that evaluation is absent from the body; if it is based on NLL, it is directly contradicted by Table 2. Either way, the current manuscript does not support the abstract's characterization of the hardest regime. This is not a subtle statistical quibble; it is a direct contradiction in the paper's headline finding.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares ANN, MDN (with 1–5 Gaussian mixture components), and BNN (with three components) for quasar photometric redshift estimation and uncertainty quantification using KiDS DR5 photometry and DESI DR1 spectroscopy. Four test regimes are constructed from the combination of \"random\" vs \"faint\" (r > 22.9) samples and complete vs missing photometric bands. The models are compared via NLL, point-estimate metrics, and the MSE* of reconstructed n(z) histograms; unsupervised clustering (t-SNE + HDBSCAN) of predicted PDFs is used to identify degenerate solutions. The body's results indicate that at least two MDN components are needed, BNNs improve NLL in the faint/missing regimes, and n(z) reconstruction degrades substantially when missing features and magnitude extrapolation are combined. The abstract included with the arXiv listing, however, makes several claims—SOM benchmarking, PIT calibration, and the single Gaussian being \"best-calibrated\" in the faint-plus-missing regime—that are absent from or contradicted by the body and its tables.","tokens_in":11279,"tokens_out":9938,"duration_ms":89244,"significance":"If properly supported, the paper would provide a useful benchmark for uncertainty-aware quasar photo-z estimation in upcoming surveys such as Euclid and LSST. The controlled construction of four test regimes, the use of external DESI spectroscopic redshifts as ground truth, and the explicit NLL comparisons in Table 2 are genuine strengths; the body's central conclusion that multi-component MDN/BNN outputs are needed for good n(z) reconstruction is broadly consistent with the reported numbers. However, the manuscript in its current form contains internal inconsistencies between the abstract and the body, missing analyses that the abstract promises, and a row-labeling error in the n(z) results section. These issues must be resolved before the headline claims can be evaluated.","major_comments":[{"comment":"The arXiv-listing abstract states that \"in the hardest faint-plus-missing regime the single Gaussian becomes the best-calibrated.\" Table 2, column 4, reports NLL = 1.28 for MDN1, 0.74 for MDN3, and 0.69 for BNN3 in that regime: the single Gaussian has the worst, not the best, NLL. If \"best-calibrated\" is meant to refer to a PIT analysis, no such analysis appears in the body. The same abstract also announces a SOM forecast and a PIT metric, neither of which is found in Sections 3–5. This is not a cosmetic mismatch; it directly affects the paper's headline claim that the best model is regime-dependent.","section":"Abstract vs §4.1, Table 2"},{"comment":"The text and the figure caption disagree on which row corresponds to which test regime. The caption defines row 1 = random complete, row 2 = random missing, row 3 = faint complete, row 4 = faint missing; the text describes the \"second row\" as the extrapolation test with all magnitudes available (i.e., row 3) and then refers to \"the random test (row 3)\" (i.e., row 2). Since the regime-by-regime MSE* rankings are central to the conclusion, this mislabeling must be corrected. Also, \"RSS values\" should read \"MSE* values.\"","section":"§4.2, Fig. 2"},{"comment":"Missing features are imputed with the training-set maximum value of that feature, and this imputation is applied uniformly to all missing cases. Because u-band gaps dominate (76% of incomplete cases), the choice of imputed value can strongly influence the \"missing feature\" test sets, especially the faint+missing regime that the abstract highlights. The Conclusions mention that missing-feature patterns are not separated, but the imputation scheme itself is not tested or discussed as a sensitivity choice. At minimum, the paper should test an alternative imputation (median, conditional mean, or missingness indicators) or explicitly restrict its missing-band conclusions to this imputation procedure.","section":"§3.2"},{"comment":"The key quantitative comparisons—e.g., BNN over MDN3 by 0.07/0.11 NLL and MSE* factors of 3 and 25—are reported without any measure of run-to-run or sampling variability. The authors train each model five times and select the best validation realization, but the spread across realizations is not reported. Several NLL differences are smaller than 0.1; without error bars or significance tests, the \"regime-dependent best model\" claim is not yet established. Please provide at least the standard deviation of each metric over the five training runs or a bootstrap over test objects.","section":"§4.1, Table 2; §4.2, Fig. 2"},{"comment":"The ground-truth redshift distribution is the spectroscopic distribution of the DESI DR1 subsample after the SPECTYPE='QSO', ZWARN==0 cuts and the 1-arcsec cross-match. DESI target selection and redshift success are not complete or unbiased at r>22.9 or for missing-band objects. The Conclusions appropriately note that the study is an upper bound because it does not probe training-inference domain mismatch, but the body repeatedly calls this the \"true\" redshift distribution. All n(z) claims should be qualified as applying to the DESI-matched spectroscopic sample; otherwise the abstract's strong statement about \"fundamentally unattainable\" reconstruction overreaches.","section":"§2.2 and Conclusions"}],"minor_comments":[{"comment":"\"The BNN achieves the lowest NLL in all three tests\" is imprecise: BNN is not the lowest on the random complete test (MDN3 has -0.69 vs BNN3 -0.62). It should read \"in the three tests involving additional sources of uncertainty\" or similar.","section":"§4.1"},{"comment":"The arXiv title and the internal manuscript title differ, and the two abstract versions are materially different (one mentions SOM/PIT, the other does not). These should be reconciled before submission.","section":"Title/Abstract"},{"comment":"Selecting the realization with the lowest validation loss from five training runs is a form of model selection; the manuscript should state explicitly whether this selection is applied independently for each test set or only on the validation set, and whether the reported NLL values are from that single realization.","section":"§3.4"},{"comment":"The text says clusters 8 and -1 show degeneracy, and cluster -1 \"usually represents outliers.\" Please define the cluster numbering convention and the criterion used to label a cluster as \"suspicious\" before reporting the 0.5%/0.2% fractions.","section":"§4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper appears to be at an early stage where the arXiv abstract and the body text have diverged. In addition to the listed major comments, I strongly encourage the editor to require the authors to unify the abstract with the reported experiments (remove SOM/PIT claims if not in the text), correct the Fig. 2 row labels, and provide uncertainty estimates on the headline metric differences. The body's core comparison table is informative and worth publishing after these revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful part of this paper is the body. The controlled split into four regimes — random vs. faint-extrapolated, complete vs. missing features — is a clean way to compare ANN, MDN, and BNN uncertainty models for quasar photo-zs. The main trends in Tables 1 and 2 are coherent: a single Gaussian MDN is worse than multi-component MDNs, and the BNN improves NLL on faint extrapolation while costing a bit on bright data. The t-SNE/HDBSCAN clustering of predicted PDFs is a nice diagnostic, and the two degenerate redshift pairs it finds are a genuine byproduct. The paper also honestly states its own limitation: because training and inference draw from the same KiDS–DESI overlap, the results are an upper bound on performance. That is the right kind of caveat.\n\nNow the problem. The first-page abstract says that in the faint-plus-missing regime the single Gaussian becomes the best-calibrated. Table 2 shows the opposite: MDN 1 has NLL 1.28 there, the worst of any model; BNN 3 has 0.69. That is not a subtle disagreement. The abstract also reports a self-organizing map and a probability integral transform evaluation. Neither appears anywhere in the methodology or results. The body's own abstract and conclusions do not repeat these claims, which suggests the front-page abstract was written for a different version of the paper. But as submitted, the paper is internally inconsistent at its headline level.\n\nOther soft spots are minor by comparison. There are no error bars on the NLL or MSE* values, and no code or data are released, so the differences between MDN 4 and MDN 5 (NLL 0.75 vs. 0.77, for instance) could be noise. The missing-magnitude imputation with training-set maxima is crude and could distort the missing-feature subsets, and the DESI ground truth may carry selection effects correlated with faintness or missing bands. These are worth a referee's attention, but none of them undercut the body's central conclusion that multi-component uncertainty models beat a single Gaussian on well-covered data and that no single model wins everywhere.\n\nWho is this for? Survey teams working with photometric quasars and anyone choosing between MDN and BNN-style uncertainty estimates for photo-z. It deserves a serious referee. My recommendation: send it to review, but tell the authors the abstract must be reconciled with the actual results before acceptance — either remove the unsupported SOM/PIT statements and the \"best-calibrated\" sentence, or add the missing analysis.","headline":"The body is a useful, internally consistent benchmark of quasar photo-z uncertainty models, but the arXiv abstract contradicts its own Table 2 and promises an SOM/PIT analysis that the text doesn't deliver.","tokens_in":11767,"tokens_out":3600,"would_cite":false,"duration_ms":32807,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Quasar photometric redshifts cannot be trusted without multi-component uncertainty models.","keywords":["photometric redshifts","quasars","uncertainty quantification","mixture density networks","Bayesian neural networks","redshift distribution","Kilo-Degree Survey","machine learning"],"falsifier":"Take a spectroscopically complete sample of faint quasars from a deeper survey that does not depend on DESI targeting, and repeat the faint-extrapolation NLL comparison between the three-component MDN and BNN. If the BNN's 0.11 NLL advantage disappears, the reported benefit of Bayesian networks is an artifact of the DESI selection function. A second test: recompute the 'missing features' rankings after dropping objects with missing u-band; if the rankings change, the results are driven by the imputation of the most important band.","tokens_in":10772,"feed_emoji":"🔭","tokens_out":6232,"duration_ms":55990,"temperature":0.7,"pith_summary":"The paper asks which machine-learning architecture best estimates quasar photometric redshifts and their uncertainties, and which can reconstruct the true redshift distribution n(z) from predicted probability densities. Using KiDS DR5 photometry matched to DESI DR1 spectroscopy, the authors train artificial neural networks (ANNs), mixture density networks (MDNs), Bayesian neural networks (BNNs), and a self-organizing map, testing them on data that are faint, missing photometric bands, or both. They find that the ANN point predictions reconstruct n(z) poorly, that MDNs need at least two Gaussian components to capture multi-modal photo-z solutions, and that BNNs improve performance in the faint extrapolation regime but degrade it for bright sources. No model wins in every regime; the combination of faintness and missing bands defeats all approaches. Unsupervised clustering of the predicted PDFs identifies two degenerate redshift pairs.","feed_headline":"Quasar photo-z accuracy demands at least two Gaussian components","feed_subtitle":"Point-estimate networks fail to recover quasar redshift distributions; the best uncertainty model depends on data quality.","key_machinery":"The load-bearing object is the Gaussian Mixture Model output layer attached to a neural network (MDN) or to a Bayesian network (BNN), which turns photo-z prediction into a full probability density function rather than a point estimate. The paper compares these PDFs through the negative log-likelihood at the true spectroscopic redshift, and reconstructs the redshift distribution by Monte Carlo sampling each predicted PDF. The BNN estimates epistemic uncertainty by averaging 100 stochastic forward passes. A separate diagnostic uses t-SNE and HDBSCAN clustering on the sampled PDF values to expose degenerate solutions.","core_discovery":"The central claim is that an accurate reconstruction of the true redshift distribution is unattainable without a proper uncertainty estimation framework, and that MDNs with at least two Gaussian components are essential to capture the intrinsic multi-modality of photo-z solutions. Concretely, on a clean random test set the ANN's distribution reconstruction error is about six times that of the best MDN; the single-Gaussian MDN is miscalibrated and produces more catastrophic outliers than the ANN. A three-component MDN is the best MDN overall, while a three-component BNN improves the negative log-likelihood by 0.11 on faint extrapolation data but is 0.07 worse for brighter objects. In the hard","pith_inferences":["Because 76% of missing-band cases are missing the u-band, the paper's grouped 'missing features' test is effectively a u-band dropout test; a per-band analysis would probably show that missing u-band dominates the degradation and that other bands are nearly harmless.","The two degeneracies (1.2,2.3) and (1.6,2.5) are consistent with quasar spectral-feature aliasing (e.g., Ly-alpha vs CIV line confusion); adding infrared photometry or a prior on quasar luminosity could break these degeneracies and would be a testable improvement.","The regime-dependent best model suggests operational pipelines could adopt a two-stage design: a cheap classifier to identify data-quality regime, then a dedicated uncertainty model for that regime, rather than relying on a single network.","Because the study deliberately restricts to a well-matched spectroscopic sample, the reported numbers are an upper bound; applying these models to the full KiDS quasar photometric sample, where training–inference domain mismatch is real, would likely degrade all metrics and could change the ranking."],"forward_implications":["Tomographic analyses with photometric quasars should be based on predicted probability distributions, not on single point estimates, since the ANN's reconstructed n(z) shows step-like artifacts and six times the error of the best uncertainty-aware model.","At least two Gaussian mixture components are required in the output layer; a single Gaussian is miscalibrated and yields more catastrophic outliers than a plain ANN.","Bayesian neural networks with three components are the best choice for samples extending fainter than the training set (NLL gain 0.11), but they come with a 0.07 NLL penalty for well-covered bright sources.","Reconstruction remains feasible for fainter data or missing magnitudes separately, but fails when both occur together, so survey samples must be quality-controlled for combined incompleteness.","Clustering predicted photo-z PDFs can identify degenerate redshift pairs ((1.2,2.3) and (1.6,2.5)) that could be removed from cosmological inference."],"fun_headline_variants":["Quasar photo-z needs two Gaussian components","Uncertainty models beat point estimates for quasar photo-z","Best quasar photo-z model depends on data quality"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The DESI DR1 spectroscopic quasars matched to KiDS are treated as an unbiased ground truth for the faint and missing-band test sets; if DESI target selection, redshift failures, or the 1-arcsec cross-match correlate with r>22.9 or missing u-band, the reported model rankings and n(z) reconstructions reflect selection effects rather than model capability.","fun_headline_variants_meta":{"raw":{"variants":["Quasar photo-z needs two Gaussian components","Uncertainty models beat point estimates for quasar photo-z","Best quasar photo-z model depends on data quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000634,"raw_usage":{"total_tokens":2836,"prompt_tokens":894,"completion_tokens":1942,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":638,"completion_tokens_details":{"reasoning_tokens":1892}},"tokens_in":638,"tokens_out":1942,"duration_ms":15924,"temperature":1.0,"reasoning_tokens":1892,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T05:43:07.652682+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a spectroscopically complete sample of faint quasars from a deeper survey that does not depend on DESI targeting, and repeat the faint-extrapolation NLL comparison between the three-component MDN and BNN. If the BNN's 0.11 NLL advantage disappears, the reported benefit of Bayesian networks is an artifact of the DESI selection function. A second test: recompute the 'missing features' rankings after dropping objects with missing u-band; if the rankings change, the results are driven by the imputation of the most important band.","supporting_citations":[],"review_version":1}