{"id":"65791f90-1f3e-4017-a218-4584580192ac","arxiv_id":"2501.01942","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A hybrid deep-learning model combining four-band images and nine-band photometry reduces photometric redshift scatter for KiDS-Bright galaxies by about 20% compared to the previous ANNz2 catalog.","lead":"This paper uses deep learning on galaxy images and brightness measurements to estimate distances (redshifts) for 1.2 million bright galaxies from the Kilo-Degree Survey, improving precision by about 20% over the previous standard method. It matters because better distance estimates for these foreground galaxies reduce systematic errors in studies of how light from distant galaxies is bent by matter, and in measuring how galaxies cluster.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Final catalog is produced by a smoothed retrained model whose accuracy is never validated; the headline 0.014(1+z) scatter is measured on the unsmoothed model, so the released-catalog claim is not established.","rationale":"The reader's weakest assumption identifies exactly the gap I find most load-bearing: the released catalog is generated by a smoothed retrained model that is not validated against any spectroscopic sample, while the headline scatter is measured on the unsmoothed model. This is not a manufactured concern; the paper's own Sec. 5.3 describes the smoothing and retraining but provides no spectroscopic evaluation of the final model, and it acknowledges a residual peak at z~0.38 even after smoothing. The concern is about the product claim, not the methodology in general. The test-set result (SMAD ~0.014 on GAMA clean test) is credible and consistent with external validation on G23 and 2dFLenS for the unsmoothed model, and the paper gives appropriate caveats about bias versus scatter trade-offs. But because the final catalog is the paper's central deliverable and is produced by a different model, the reported statistics do not yet support the final product's accuracy. The reader already issued a CONDITIONAL verdict, and my analysis agrees rather than raising a new objection; no adjustment is needed.","tokens_in":24564,"tokens_out":2589,"duration_ms":28468,"concrete_test":"Retrain Hybrid-z on the smoothed 118k subsample described in Sec. 5.3 and evaluate it on the same held-out GAMA clean test set (20,965 galaxies) and on the blind G23 and 2dFLenS crossmatches. Compute SMAD(Δz), ⟨δz⟩, and dN/dz_phot and compare directly with Tables 1 and 2. If the smoothed model's SMAD(Δz) rises above ~0.016(1+z), or if |⟨δz⟩| exceeds ~0.005, or if a systematic bump at z~0.38 persists in the dN/dz comparison against spec-z, then the released catalog does not inherit the validated 0.014(1+z) precision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central product is the released 1.2M-galaxy KiDS-Bright photo-z catalog, but the paper's headline precision is not measured on that catalog. Section 5.3 describes retraining Hybrid-z on a GAMA subsample whose redshift distribution has been smoothed to remove cosmic-variance features, then applying this model to the full sample. This retrained model is never evaluated against any spectroscopic catalog; all reported statistics, including Table 1's 0.014(1+z) SMAD on clean data and Table 2's external-sample validations, use the unsmoothed model trained on full GAMA-equatorial. Smoothing by subsampling in z changes the training prior p(z); for a photo-z mapping with intrinsic scatter, the optimal conditional mean E[z|x] depends on that prior, so the smoothed model may systematically shift predictions for given colors/images even if p(color|z) is unchanged. The paper itself notes a persistent artifact at z_phot~0.38 in the smoothed model (Sec. 5.3, Fig. 7b), indicating the smoothing does not fully remove the problem. Separately, GAMA incompleteness at r>19.5 mag (cited from Jalan et al. 2024) means the final catalog includes galaxies beyond the well-sampled training regime; the smoothed model's faint-end behavior is unquantified. Thus the load-bearing weakness is that the only claim validated with spectroscopic data is for a different model than the one used to generate the released catalog.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Hybrid-z, a deep-learning photometric redshift model for the KiDS-Bright DR4 sample (r < 20 mag). The model combines a convolutional network operating on four-band KiDS image cutouts with a fully connected network operating on nine-band KiDS+VIKING magnitudes, and is trained on spectroscopic redshifts from the GAMA equatorial fields. Against a held-out GAMA test sample, the authors report SMAD(Δz) ≈ 0.014(1+z) for the clean sample, a roughly 20% improvement over the previous ANNz2 results (≈ 0.018), with comparable small mean residuals; they also validate the model on the external samples SDSS, 2dFGRS, GAMA G23, and 2dFLenS, finding consistently lower scatter than ANNz2. For the final catalog of ~1.2 million galaxies, the network is retrained on a GAMA subsample whose redshift distribution has been smoothed to suppress cosmic-variance features, and the resulting photo-z catalog is publicly released.","tokens_in":24898,"tokens_out":4463,"duration_ms":43747,"significance":"If the headline result holds, this is a useful and practical improvement for a widely used low-redshift foreground sample: it demonstrates that image-plus-magnitude deep learning beats feature-only ANNz2 in the bright, well-resolved regime, and it provides a public catalog. The paper's main methodological strengths are genuine: the internal comparison uses a held-out split, the external validation includes two spectroscopic samples (GAMA G23 and 2dFLenS) that are spatially disjoint from the training data, and the improvement over ANNz2 is reported across multiple independent samples with consistent statistics. The central weakness is that the released catalog is produced by a retrained, smoothed model whose accuracy is never spectroscopically validated, so the paper does not currently establish that the public catalog has the measured 0.014(1+z) scatter.","major_comments":[{"comment":"The headline precision, SMAD(Δz) ≈ 0.014 for the clean sample, and all external-sample statistics in Tables 1 and 2 are obtained with the model trained on the full GAMA-equatorial sample. The released 1.2M-galaxy catalog is instead generated by a model retrained on the smoothed 118k-galaxy subsample described in Section 5.3. No spectroscopic validation of this retrained model is presented; the only displayed output of the smoothed model is the dN/dzphot distribution in Fig. 7b. The paper therefore does not establish that the released catalog has the claimed scatter or the claimed near-zero mean residuals. Please validate the smoothed model on a held-out GAMA test set and, ideally, on the blind external samples G23 and 2dFLenS, reporting the same statistics as in Tables 1 and 2. If that is not possible, the claims and the catalog description should be restricted to the unsmoothed model.","section":"Section 5.3; Tables 1 and 2"},{"comment":"The abstract's statement of 'negligible mean residuals of O(10^-4)' is not supported by the external validation. Table 2 lists mean rescaled biases of -0.0031 for GAMA-Equatorial, -0.0023 for SDSS DR16, -0.0025 for 2dFLenS, and -0.0017 for GAMA G23. These offsets are small relative to the scatter, but they are of order 10^-3, not 10^-4. The O(10^-4) level applies to the internal GAMA test sample in Table 1. Please rephrase the abstract and Section 6 to attribute the O(10^-4) claim only to the internal test set and to state explicitly that external samples show systematic offsets of a few times 10^-3.","section":"Abstract; Table 2"},{"comment":"The smoothing procedure subsamples the training set in redshift to flatten dN/dzspec, changing the training prior p(z). For a regression model with intrinsic scatter, the optimal conditional mean prediction E[z|x] depends on that prior, so preserving the color-redshift relation p(color|z) does not guarantee that p(z|color) or the conditional photo-z mapping is preserved. The paper's assertion that smoothing does not affect the color-redshift relation is therefore not a substitute for end-to-end spectroscopic validation of the retrained model. This concern is reinforced by the paper's own observation that a peak at z_phot ≈ 0.38 persists in the smoothed model (Section 5.3, Fig. 7b), indicating that the mitigation is incomplete. Please provide bias and scatter as functions of true redshift, magnitude, and color for the smoothed model on spectroscopic data, and discuss how the smoothing shifts the conditional predictions relative to the validated unsmoothed model.","section":"Section 5.3, smoothing procedure"},{"comment":"The paper acknowledges that GAMA becomes incomplete at r ≳ 19.5 mag and that this causes covariate shift for the faintest KiDS-Bright galaxies, yet the magnitude dependence of the released (smoothed) model is never quantified. Fig. 5c shows growing scatter at r near 20 for the unsmoothed model, but the released catalog uses the smoothed model, whose faint-end behavior may differ. Given that the catalog is flux-limited to r < 20 and contains over a million galaxies, the faint-end extrapolation is a load-bearing uncertainty for the public product. Please report magnitude-binned statistics for the smoothed model on spectroscopic data, or explicitly quantify the resulting systematic uncertainty in the released catalog.","section":"Section 5.3, faint end"}],"minor_comments":[{"comment":"The text states that z_i is the predicted value and \\hat{z}_i is the true value, which is the opposite of the standard convention used elsewhere in the paper; please swap the notation or the definitions for consistency.","section":"Equations (3)-(4)"},{"comment":"The symbol m is used both for the magnitude and for the mean in the standardization formula; please use a different symbol for the mean, such as \\mu_m, to avoid ambiguity.","section":"Equation (2)"},{"comment":"The paper says about 125k galaxies are used for training, but the stated 70:15:15 split of the ~173k equatorial galaxies gives roughly 121k; please reconcile these numbers.","section":"Section 3.2"},{"comment":"The phrase 'we would such as to emphasize' should read 'we would like to emphasize'.","section":"Section 5.1"},{"comment":"The caption contains a typo: 'Input (36,36,,4)' should be 'Input (36,36,4)'.","section":"Fig. 2 caption"},{"comment":"The caption calls the GAMA test sample 'blind', but it is a random split from the same GAMA-equatorial distribution used for training; the genuinely blind samples are GAMA G23 and 2dFLenS. Please adjust the wording to avoid overstating independence.","section":"Fig. 5 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper's core comparison of the unsmoothed model against ANNz2 is sound and useful, and the catalog release is a valuable community resource. The main gap is that the released catalog is produced by a model that is never validated against spectroscopy, and the abstract overstates the bias claims relative to the external validation. These issues are fixable within the manuscript's scope by adding a validation of the smoothed model and softening the claims, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid applied photo-z paper that delivers a real, externally validated scatter improvement for KiDS-Bright and ships a public catalog. The main caveat, which the stress-test note gets right, is that the released catalog is produced by a retrained model whose accuracy is never directly measured.\n\nWhat is actually new: first DL photo-zs for KiDS-Bright DR4, a hybrid CNN+ONN combining four-band images with nine-band magnitudes, and a public 1.2M-galaxy catalog. The core test-set result is credible: SMAD ~0.014(1+z) versus ~0.018 for ANNz2 on the clean sample, and the external checks on 2dFLenS and GAMA G23, which are disjoint from training, show similar scatter gains. The blue-galaxy improvement is physically sensible and a nice result. The paper is also honest about the cosmic-variance artifacts inherited from GAMA and explains the smoothing mitigation clearly.\n\nThe soft spots are real but one is load-bearing. The final catalog comes from a model retrained on a smoothed GAMA subsample (Sect. 5.3), and that model is never evaluated against any spectroscopic catalog. All the headline statistics, including the external validations in Table 2, use the unsmoothed model. Smoothing the training redshift distribution changes the prior for a supervised regression, which can shift E[z|x] for fixed photometry; the paper even notes that a z_phot~0.38 artifact persists in the smoothed model. So the released catalog's accuracy is genuinely unquantified. That does not undermine the central comparison between Hybrid-z and ANNz2 on held-out data, but it does mean the product claim is weaker than the headline. Also, GAMA incompleteness at r>19.5 means faint-end extrapolation, mentioned but not quantified. Minor issues: the abstract's 'negligible mean residuals of O(10^-4)' holds for the GAMA test set but not for external samples, where mean rescaled biases are ~2-3e-3, still small but not 1e-4; and the reported metrics have no uncertainties.\n\nOverall, the core claim holds up. The released catalog is the one part the authors did not close the loop on. The paper deserves a serious referee: the fix is straightforward (validate the smoothed model on any spec-z sample, or at least flag the catalog as unvalidated), and the community will use this catalog. I'd send it out.","headline":"Genuine 20% scatter improvement over ANNz2 for KiDS-Bright, honestly documented, but the public catalog comes from a retrained model that never got a spectroscopic validation.","tokens_in":25480,"tokens_out":2496,"would_cite":false,"duration_ms":24847,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A deep-learning model that combines four-band galaxy images with nine-band magnitudes cuts photometric-redshift scatter for the KiDS-Bright sample by about 20% over the previous ANNz2 results, from SMAD($\\Delta z$) $\\approx 0.018$ to…","keywords":["photometric redshifts","deep learning","convolutional neural networks","KiDS-Bright","GAMA spectroscopy","galaxy surveys","cosmic variance"],"falsifier":"Take the released Hybrid-z catalog and compare its photo-$z$s with an independent spectroscopic sample covering the faint part of KiDS-Bright ($r$ between 19.5 and 20 mag) that was not used in training, for instance DESI DR1 or a dedicated follow-up; if the SMAD($\\Delta z$) there exceeds about $0.018(1+z)$ or the mean bias grows beyond a few times $10^{-3}$, the claim that the full 1.2-million-galaxy catalog achieves the reported precision would be refuted.","tokens_in":24376,"feed_emoji":"🔭","tokens_out":14585,"duration_ms":115400,"temperature":0.7,"pith_summary":"The paper builds a hybrid deep-learning model, Hybrid-z, that estimates photometric redshifts for the bright, flux-limited KiDS-Bright galaxy sample by feeding four-band KiDS images through a convolutional network and nine-band KiDS+VIKING magnitudes through an ordinary neural network, then concatenating the two. Tested on GAMA spectroscopy, the model reaches a scatter of SMAD($\\Delta z$) $\\approx 0.014(1+z)$ on the clean sample, about 20% smaller than the earlier ANNz2 value of $\\approx 0.018(1+z)$, while keeping the mean bias at a few times $10^{-4}$. The improvement is largest for blue galaxies, whose scatter drops by about 22%, consistent with image morphology carrying information that magnitudes alone do not. For the full 1.2-million-galaxy catalog the authors retrain on a smoothed subsample of GAMA redshifts to prevent cosmic-variance features in the training sky from being imprinted on the predicted redshift distribution, releasing this final catalog alongside the paper. If correct, the work shows that adding imaging to magnitude-based photo-$z$ networks gives a material precision gain for the bright foreground galaxies used in lensing and clustering analyses.","feed_headline":"Deep learning shaves 20% off bright-galaxy redshift errors","feed_subtitle":"A CNN on KiDS images plus nine-band magnitudes hits 0.014(1+z) scatter, sharpening lensing and clustering analyses.","key_machinery":"The load-bearing mechanism is the Hybrid-z architecture and its training procedure. A convolutional network built from inception modules, using parallel $1\\times1$, $3\\times3$, and $5\\times5$ convolutions with average pooling, extracts multi-scale features from $36\\times36$-pixel, four-band ($ugri$) KiDS cutouts; an ordinary fully-connected network processes nine standardized magnitudes; and the two outputs are depth-concatenated before the final dense layers, an architectural choice that lets image morphology directly influence the redshift prediction. The second mechanism is the smoothing algorithm applied for the final catalog: iteratively trimming spikes from the GAMA training redshift histogram to produce a more uniform $dz$ distribution, which the authors use to retrain the model and thereby suppress redshift-focusing artifacts in the full-sample prediction.","core_discovery":"On the paper's own terms, the central discovery is that a deep network processing galaxy images together with nine-band magnitudes yields photometric redshifts for KiDS-Bright that are substantially more precise than the previous feature-only ANNz2 estimates, while remaining equally unbiased. Specifically, on the fiducial clean GAMA test set, Hybrid-z achieves SMAD($\\Delta z$) $\\approx 0.014$ compared with $\\approx 0.018$ for ANNz2, a 20% reduction, with the mean of $\\delta z$ at most a few times $10^{-4}$; the same pattern holds across all external spectroscopic cross-checks, with scatter reductions of 17%–24% and, in the blind samples (GAMA G23 and 2dFLenS) that are disjoint from the training fields, reductions of 17% and 20%. The paper additionally reports that applying the model trained on the full GAMA data to the entire KiDS-Bright sample imprints GAMA's large-scale-structure features, such as a dip near $z \\sim 0.25$ and a peak near $z \\sim 0.38$, into the predicted redshift distribution, and that training on a smoothed version of the GAMA redshift histogram removes these artifacts while, per the authors, preserving the color–redshift relation.","pith_inferences":["The headline scatter of $0.014(1+z)$ is measured only for the unsmoothed GAMA-trained model; the released catalog's retrained model is never validated against spectroscopy, so a direct external check of the released redshifts could still reveal that the smoothing changes the photometry-to-redshift mapping in ways the reported statistics do not cover.","The larger gains for blue galaxies suggest that morphological information from the images is a major driver of the improvement; an ablation study comparing image-only, magnitude-only, and hybrid models on matched training would quantify this contribution and is an obvious next step.","The same hybrid architecture could plausibly transfer to other bright, low-redshift surveys such as SDSS or future LSST data, with the main risk being covariate shift at the faint end where the training sample is incomplete (roughly $r > 19.5$ for GAMA).","The redshift-focusing artifacts reveal that the network is sensitive to the training $dN/dz$; rather than only smoothing GAMA, training on a wider-area spectroscopic sample such as DESI DR1 would remove the cosmic-variance imprint at the source."],"forward_implications":["The KiDS-Bright photo-$z$ catalog recommended for science gains about 20% in scatter compared with the ANNz2 version, reducing redshift-error systematics in galaxy-galaxy lensing and clustering analyses that use these galaxies as a foreground.","Blue, typically spiral galaxies receive the largest improvement (about 22% scatter reduction), so samples with a high blue fraction will see the biggest gains in photo-$z$ precision.","Because the model performs consistently better across several independent spectroscopic samples, including two fully blind southern fields, the improvement is not an artifact of the particular GAMA equatorial test split.","The released full-sample catalog is trained on a smoothed redshift distribution, so its $dN/dz$ no longer mimics GAMA's cosmic-variance features, which is the relevant quantity for tomographic and clustering applications."],"supporting_citations":[{"why":"Defines the KiDS-Bright DR4 sample and supplies the ANNz2 photo-$z$ baseline that Hybrid-z improves on.","marker":"B21"},{"why":"The ANNz2 package that produced the previous nine-band photo-$z$s.","marker":"Sadeh et al. 2016"},{"why":"Final GAMA DR4 spectroscopic catalog providing the redshift labels for training and testing.","marker":"Driver et al. 2022"},{"why":"Introduces the GAMA survey whose flux-limited selection matches KiDS-Bright.","marker":"Driver et al. 2011"},{"why":"Earlier KiDS deep-learning photo-$z$ work combining images and magnitudes that Hybrid-z builds on.","marker":"Li et al. 2022"},{"why":"Source of the inception-module design for CNN photo-$z$s.","marker":"Henghes et al. 2022"},{"why":"Analogous CNN photo-$z$ derivation for SDSS at the same $r<20$ depth that this work's performance matches.","marker":"Treyer et al. 2024"},{"why":"KiDS DR4 release providing the imaging and photometry used as model inputs.","marker":"Kuijken et al. 2019"}],"fun_headline_variants":["Hybrid-z cuts KiDS-Bright photo-z scatter by 20%","Deep learning sharpens KiDS-Bright redshifts 20%","Images plus nine bands slash KiDS-Bright photo-z errors","Neural net photo-zs beat ANNz2 on KiDS-Bright by 20%","KiDS-Bright gets 20% better redshifts via deep learning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that smoothing the GAMA training redshifts removes cosmic-variance artifacts without changing the mapping from photometry to redshift, and it also assumes the model generalizes to the faintest KiDS-Bright galaxies where GAMA is incomplete, while the final model is never validated against spectroscopy.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid-z cuts KiDS-Bright photo-z scatter by 20%","Deep learning sharpens KiDS-Bright redshifts 20%","Images plus nine bands slash KiDS-Bright photo-z errors","Neural net photo-zs beat ANNz2 on KiDS-Bright by 20%","KiDS-Bright gets 20% better redshifts via deep learning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000252,"raw_usage":{"total_tokens":1689,"prompt_tokens":1205,"completion_tokens":484,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":821,"completion_tokens_details":{"reasoning_tokens":387}},"tokens_in":821,"tokens_out":484,"duration_ms":5094,"temperature":1.0,"reasoning_tokens":387,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:14:25.171099+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the released Hybrid-z catalog and compare its photo-$z$s with an independent spectroscopic sample covering the faint part of KiDS-Bright ($r$ between 19.5 and 20 mag) that was not used in training, for instance DESI DR1 or a dedicated follow-up; if the SMAD($\\Delta z$) there exceeds about $0.018(1+z)$ or the mean bias grows beyond a few times $10^{-3}$, the claim that the full 1.2-million-galaxy catalog achieves the reported precision would be refuted.","supporting_citations":[],"review_version":1}