{"id":"e2df3498-c747-44fc-9b17-fb6c0349af93","arxiv_id":"2411.18769","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"DeepDISC photo-z, an instance-segmentation network with a mixture-density redshift head, produces better photometric redshifts than catalog-based BPZ and FlexZBoost on simulated Rubin LSST images.","lead":"A deep-learning photo-z estimator that reads raw multi-band images, detects objects, and outputs a redshift probability distribution outperforms catalog-based photo-z codes on simulated Rubin LSST data. It is most robust to blended sources, and its redshift scatter improves with image depth roughly as fast as the signal-to-noise grows.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline comparison is not fully controlled: the test sample is defined by DeepDISC detections, so the claimed outperformance may partly reflect selection rather than estimator quality.","rationale":"The reader's weakest assumption is that the DC2 benchmark provides a fair and representative test, citing test-set selection, the simulated dust-law flaw, and simulation realism. My concern is more specific: the test set is defined by DeepDISC detections, so the comparison in Table 2 is conditional on DeepDISC's detection behavior. This is acknowledged by the authors, which is good, but the paper does not provide a fully neutral control for the evaluation-selection effect. The BPZ template mismatch is real but less load-bearing because DeepDISC also outperforms FZB, a machine-learning method that is not handicapped by that flaw. The selection issue, however, could in principle affect the comparison against both baselines and the global metrics. I therefore agree with the conditional verdict but sharpen the condition: the outperformance claim should be demonstrated on a truth- or object-catalog-selected sample, not only on a DeepDISC-selected sample. The proposed concrete test would settle this. If it passes, the central claim is substantially strengthened; if it fails, the abstract should be scoped to the DeepDISC-selected sample. I do not see a basis for rejection, given the open-source code, the controlled DC2 setup, and the authors' explicit disclosure of limitations.","tokens_in":25351,"tokens_out":4339,"duration_ms":42346,"concrete_test":"Recompute all global metrics on a test set defined by the DC2 object catalog rather than by DeepDISC inference: take all gold-sample objects in the test footprints, run BPZ and FZB on their cModel photometry, and require DeepDISC to produce a photo-z prediction for each such object, counting non-detections as catastrophic failures. Compare the resulting bias, scatter, outlier fraction, and CDE loss to Table 2. If DeepDISC no longer beats FZB and BPZ on this neutrally selected sample, the headline claim fails; if it still does, the selection concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that DeepDISC outperforms catalog-based estimators on a common evaluation sample. In Section 4, BPZ and FZB are scored only on objects that DeepDISC detected and that cross-match to the DC2 object catalog. Section 5.1 then states that the test set selection is determined by DeepDISC inference matched with the object catalog, potentially privileging DeepDISC. This is not a minor caveat: if DeepDISC preferentially detects objects for which its photo-z head is accurate, or systematically misses a class of objects that catalog methods handle well, every global metric in Table 2 is biased in its favor. The Section 5.1 control retrains FZB on a DeepDISC-selected training sample, but it does not re-run BPZ on a neutrally selected sample, and no global metrics are reported on a truth-selected sample that includes DeepDISC non-detections. Because the claimed advantage is global and the paper does not yet establish it on a neutrally selected evaluation sample, the abstract's outperformance statement is not fully supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents DeepDISC-photoz, an extension of the DeepDISC instance-segmentation framework that adds a Mixture Density Network head producing photo-z PDFs directly from multi-band coadded images. Using the DESC DC2 simulations and the LSST gold sample, the authors train DeepDISC on 1048 cutouts and evaluate on 1925 test images, comparing against BPZ and FlexZBoost run on matched object-catalog photometry. They report that DeepDISC has lower bias, interquartile scatter, and outlier fraction, as well as better PIT calibration and lower CDE/KS/CvM/AD statistics, and that it is less sensitive to blending. The paper also examines depth dependence, model and training-set scaling, the information content of DeepDISC features versus photometry, stellar contamination, and the effect of image smoothing.","tokens_in":25549,"tokens_out":3693,"duration_ms":33272,"significance":"If the claims hold, this is a valuable step toward end-to-end image-based photo-z estimation for LSST, and the public code integrated in RAIL is a useful community resource. The paper's strengths include open release of code, candid disclosure of limitations, and a clean feature-versus-photometry control experiment. However, the headline comparison is weakened by the DeepDISC-selected test sample and by the acknowledged BPZ template mismatch, so the broad outperformance statement in the abstract is not yet fully supported by the presented evidence.","major_comments":[{"comment":"The global metrics in Table 2 are computed on a test set whose selection is determined by DeepDISC inference matched to the DC2 object catalog, as stated in Section 5.1. Because BPZ and FZB are evaluated only on objects that DeepDISC detected and cross-matched, a systematic detection preference could favor DeepDISC in every global metric. The Section 5.1 control retrains FZB on a DeepDISC-selected sample, but it does not re-run BPZ on a neutrally selected sample, and no global metrics are reported on a truth-selected sample that includes DeepDISC non-detections. Please add such an evaluation, or explicitly restrict the main outperformance claim to the DeepDISC-selected sample.","section":"Section 5.1 and Table 2"},{"comment":"The BPZ comparison is not fully controlled because the simulated dust extinction law is flawed and the BPZ template set cannot represent the resulting high-redshift colors; the paper states that this will degrade BPZ relative to Schmidt et al. (2020). Consequently, the global superiority of DeepDISC over BPZ in Table 2, and especially the advantage at 1.5 < z < 2.5 in Figure 5, may partly reflect template mismatch rather than estimator quality. Please either use a corrected template set, restrict the BPZ comparison to regimes where the templates are reliable, or clearly qualify the abstract's outperformance claim accordingly.","section":"Section 3.2 and Section 3.2.1"},{"comment":"DeepDISC never produces a PDF mode above z ~ 2.5, as the paper acknowledges in Section 5.3. This is a systematic failure in the high-redshift regime, yet Table 2 reports only global metrics. Because the abstract claims general outperformance in point-estimate metrics, the paper should quantify performance separately for z > 2.5 and explicitly state that the claimed advantage does not extend to that regime.","section":"Section 4 and Section 5.3"}],"minor_comments":[{"comment":"The sentence 'At 1.5 ≤ z ≥ 2.5, DeepDISC maintains...' contains a typo; it should read '1.5 ≤ z ≤ 2.5'.","section":"Section 4"},{"comment":"The caption 'DeepDISC outperforms BPZ and FZB in all cases' would be clearer as 'in all metrics', since the CDE loss is negative and the comparison is by magnitude rather than direction.","section":"Table 2 caption"},{"comment":"Global metrics are quoted without uncertainties; bootstrapped errors are shown for binned metrics in Figure 5, but the global values in Table 2 should also include at least bootstrap or jackknife uncertainties.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"This is a solid empirical paper whose authors are transparent about the two main threats to their headline comparison. The fixes are feasible within the manuscript's scope: adding a neutral selection evaluation and either correcting the BPZ templates or qualifying the BPZ comparison. No concerns about novelty disclosure or citation behavior."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the short version: this is a real contribution, and the paper is worth refereeing, but the abstract's \"outperforms traditional catalog-based estimators\" should be read as \"outperforms in this particular controlled setup.\" The two main caveats are known and disclosed by the authors, but they bite harder than the paper lets on.\n\nWhat's new and good: Adding an MDN photo-z head to the DeepDISC instance-segmentation framework is a clean, natural extension, and it gives you a single network that detects, deblends, and produces a redshift PDF. The empirical work is thoughtful: the depth comparison (1-year vs 5-year data) shows scatter scaling roughly with SNR; the scaling-law experiments are honestly reported with a null result; and the test of whether DeepDISC features carry more information than photometry (training FlexZBoost on PCA'd features vs colors) is a smart, direct probe. The code is public and integrated into RAIL, which makes it immediately usable. The paper is also unusually candid: it flags the dust-law flaw in the simulations, the potential selection effect, and the lack of high-z modes.\n\nThe soft spots. The main benchmark is on a test set defined by DeepDISC detections matched to the DC2 object catalog. That can systematically favor DeepDISC, because objects it fails to detect never enter the comparison. The authors acknowledge this in Section 5.1 and do a partial control (retraining FlexZBoost on a DeepDISC-selected sample on a small subset), but they don't provide global metrics on a neutrally selected sample that includes DeepDISC non-detections. The abstract's global claim is therefore not fully supported. Second, BPZ is knowingly handicapped by the flawed simulated dust extinction law, so the comparison to BPZ in particular is not a fair test. That doesn't invalidate the paper, but it does mean the \"outperforms\" claim is conditioned on that simulation flaw. Third, global metrics in Table 2 have no uncertainties, which matters when the margins are thin for some metrics.\n\nWho is this for? Photo-z practitioners in LSST DESC and anyone building image-based estimators. It deserves a serious referee. I'd send it out, and ask for either a neutral-selection validation or a scoped claim.","headline":"A solid, openly-released image-based photo-z benchmark, but the headline outperformance claim is not yet fully controlled because the test set is DeepDISC-selected and BPZ runs with a known dust-law handicap.","tokens_in":26149,"tokens_out":2696,"would_cite":true,"duration_ms":23386,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"On simulated Rubin images, a pixel-reading network beats catalog-based redshift estimators.","keywords":["photometric redshifts","deep learning","instance segmentation","mixture density networks","Rubin LSST","DC2 simulations","source blending","photo-z PDF calibration"],"falsifier":"Run DeepDISC, BPZ, and FlexZBoost on a common test sample selected independently of DeepDISC (e.g., purely from the LSST science pipeline object catalog) and on images with realistic irregular galaxy morphologies and a corrected dust extinction law; if the deep-learning model's advantage in $\\sigma_{\\rm IQR}$ and outlier fraction $\\eta$ shrinks or reverses under either change, the claimed outperformance is an artifact of the simulation or the selection scheme.","tokens_in":25138,"feed_emoji":"🔭","tokens_out":6634,"duration_ms":56877,"temperature":0.7,"pith_summary":"Photometric redshifts, distances estimated from a few broad-band brightness measurements rather than spectra, are needed for billions of Rubin LSST galaxies. This paper claims that a single deep-learning network, an extension of the DeepDISC instance-segmentation framework, can detect, segment, classify, and estimate a full redshift probability distribution directly from multi-band images, skipping the usual step of building a photometric catalog. On simulated LSST (DC2) images, the network reports lower bias, scatter, and outlier fraction than the template-based code BPZ and the machine-learning code FlexZBoost, together with better-calibrated uncertainty distributions. The paper argues that the performance gain comes from pixel-level morphological and color-gradient information, and shows that the network is particularly robust to blended sources.","feed_headline":"AI that reads pixels beats catalog-based galaxy redshift codes","feed_subtitle":"On simulated Rubin LSST images, the DeepDISC network reports lower scatter, bias, and outlier rates than BPZ and FlexZBoost.","key_machinery":"The load-bearing object is the DeepDISC architecture with a Mixture Density Network (MDN) as the redshift head. The backbone is a multi-scale vision transformer (MViTv2) with a feature pyramid and cascade Region of Interest heads, trained end-to-end on six-band $ugrizy$ coadded images. For each detected object, the MDN outputs the weights, means, and log standard deviations of five Gaussian components, which parameterize the photo-z probability density function $p(z)$; the network is trained with the negative log-likelihood of the true redshifts. Milky Way dust reddening $E(B-V)$ enters as an extra neuron input to the MDN. During training, the detection and segmentation branches learn from deblended ground truth produced by the scarlet algorithm, while the redshift branch trains only on the LSST 'gold' sample ($i<25.3$ mag).","core_discovery":"The central claim is that adding a redshift-estimation Region of Interest head, implemented as a Mixture Density Network, to the DeepDISC object-detection framework yields a photo-z estimator that outperforms catalog-based estimators on simulated LSST data in both point estimates and probabilistic metrics. On the DC2 Year-5 test set, DeepDISC photo-z achieves bias $e_z=0.0007$, scatter $\\sigma_{\\rm IQR}=0.0412$, outlier fraction $\\eta=0.1191$, and a CDE loss of $-4.249$, beating BPZ and FlexZBoost on every metric reported. The paper also establishes that photo-z scatter decreases roughly in proportion to the inverse of image signal-to-noise when comparing Year-1 and Year-5 coadds, and that pixel-level information is the carrier of the advantage: blurring the images degrades the model, and training FlexZBoost on DeepDISC's learned features recovers part of the gain over photometry.","pith_inferences":["If the claimed margin survives on real data, the practical implication is that the LSST photo-z pipeline could be simplified around one image-based estimator, reserving catalog-based codes as cross-checks; the paper's own selection caveat means this transfer must first be tested with an independent detection catalog.","The inverse-scatter-with-SNR scaling suggests a simple survey-design rule: photo-z quality in the i-band gold sample is set mostly by depth, so area-depth trade-offs can be evaluated before any new models are trained.","The feature-distillation experiment hints at a broader recipe: rather than discarding catalog methods, train them on features extracted by image models, which may offer a lower-cost hybrid with some of the image-based advantage; the paper only demonstrates this for FlexZBoost.","Testing on higher-fidelity simulations with irregular morphologies and a correct dust extinction law would sharpen the claim, since those are exactly the axes on which the paper admits the DC2 data depart from reality."],"forward_implications":["LSST photo-z production could become a single forward pass over coadded images that simultaneously detects, deblends, classifies, and assigns redshift PDFs, removing the separate forced-photometry catalog step.","Adding observing time improves photo-z scatter nearly proportionally to the gain in signal-to-noise, giving a quantitative forecast for how Year-1 to Year-10 coadds will improve redshift quality.","The strong robustness to blending implies that pixel-based estimators can partly bypass the hardest failure mode of catalog deblending, which should matter for weak-lensing shape and redshift analyses.","Secondary peaks in the DeepDISC PDFs carry genuine redshift information: evaluating the PDF at the secondary peak recovers a majority of the point-estimate outliers.","Since increasing model size and training-set size produced no clear gains, further progress is more likely to come from better pre-training or training-set augmentation than from scaling the current architecture."],"supporting_citations":[{"why":"Supplies the DeepDISC instance-segmentation framework and the transfer-learning strategy (MViTv2 backbone pre-trained on ImageNet) that this work extends with a photo-z head.","marker":"Merz et al. (2023)"},{"why":"Produces the DC2 simulated images and truth catalogs that are the sole dataset for training and benchmarking all codes.","marker":"LSST Dark Energy Science Collaboration (LSST DESC) et al. (2021)"},{"why":"Defines BPZ, the template-based photo-z code that is the primary catalog-based comparator.","marker":"Benítez (2000)"},{"why":"Defines FlexZBoost, the machine-learning catalog-based photo-z code used as the second comparator.","marker":"Izbicki & Lee (2017)"},{"why":"Establishes the mixture-density-network approach to photo-z PDFs from images and the choice of five Gaussian components.","marker":"D'Isanto & Polsterer (2018)"},{"why":"Provides the benchmark methodology and outlier definition used for the comparison metrics.","marker":"Schmidt & Malz et al. (2020)"},{"why":"Foundational reference for the mixture density network used as the redshift head.","marker":"Bishop (1994)"},{"why":"Supplies the scarlet deblending algorithm used to create the deblended ground truth for training detection and segmentation.","marker":"Melchior et al. (2018)"}],"fun_headline_variants":["Pixel-level AI photo-z beats catalog-based codes on LSST sims","DeepDISC photo-z: pixel redshifts top BPZ and FlexZBoost in Rubin sims","AI that reads pixels: photo-z wins over catalogs for LSST","Pixel-wise photo-z from DeepDISC outperforms BPZ and FlexZBoost","DeepDISC's pixel-based photo-z beats catalog methods on simulated LSST"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark assumes that the DC2 simulated images and the matched-catalog training and test scheme are a fair and representative stand-in for real LSST data; the paper itself notes that the simulated morphologies are simplified bulge+disk+knot profiles, that high-redshift SED diversity is limited, that a flaw in the simulated dust extinction law leaves BPZ with mismatched templates, and that the test set is selected by DeepDISC's own detections, which may privilege it.","fun_headline_variants_meta":{"raw":{"variants":["Pixel-level AI photo-z beats catalog-based codes on LSST sims","DeepDISC photo-z: pixel redshifts top BPZ and FlexZBoost in Rubin sims","AI that reads pixels: photo-z wins over catalogs for LSST","Pixel-wise photo-z from DeepDISC outperforms BPZ and FlexZBoost","DeepDISC's pixel-based photo-z beats catalog methods on simulated LSST"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001628,"raw_usage":{"total_tokens":6518,"prompt_tokens":1033,"completion_tokens":5485,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":649,"completion_tokens_details":{"reasoning_tokens":5381}},"tokens_in":649,"tokens_out":5485,"duration_ms":38469,"temperature":1.0,"reasoning_tokens":5381,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:53:38.784810+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run DeepDISC, BPZ, and FlexZBoost on a common test sample selected independently of DeepDISC (e.g., purely from the LSST science pipeline object catalog) and on images with realistic irregular galaxy morphologies and a corrected dust extinction law; if the deep-learning model's advantage in $\\sigma_{\\rm IQR}$ and outlier fraction $\\eta$ shrinks or reverses under either change, the claimed outperformance is an artifact of the simulation or the selection scheme.","supporting_citations":[{"cited_title":"L., 2018, @doi [ ] 10.1051/0004-6361/201731326 , https://ui.adsabs.harvard.edu/abs/2018A&A...609A.111D 609, A111","cited_arxiv_id":null,"evidence_quote":"Establishes the mixture-density-network approach to photo-z PDFs from images and the choice of five Gaussian components."}],"review_version":1}