{"id":"b7184b10-1482-4cf6-858f-152eeb0a6f12","arxiv_id":"2505.24175","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A CNN-MLP model fusing DESI Legacy images and photometric features estimates photometric redshifts of emission line galaxies with sigma_NMAD=0.014 and 2.57% outliers on a held-out test set.","lead":"Astronomers built a hybrid neural network that combines galaxy images with photometric measurements to estimate distances (redshifts) for emission line galaxies. On a test set of about 38,000 galaxies, it achieves a normalized median absolute deviation of 0.014 and an outlier rate of 2.57%, better than using either images or photometry alone.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline metrics are dominated by bright low-redshift SDSS galaxies; the DESI target-selection benefit is asserted without validation on the actual DESI ELG target population.","rationale":"The reader's weakest assumption correctly identifies the central risk: the paper's sample is not representative of the DESI ELG target population, and the model is never validated on DESI-selected ELG targets. The paper's own Table 5 provides the strongest evidence for this concern: the faint-source performance (sigma_NMAD = 0.0479, eta = 0.0691) is close to the DESI ELG regime and far from the headline values. This makes the abstract's conclusion that the model 'directly benefits' DESI target selection an unsupported extrapolation. I also note that the only direct model comparisons are to the authors' own single-modality baselines, and the external comparison to Zhou et al. (2025) in the introduction uses a different sample and selection, so the claim of improvement over other models is not substantiated. These issues are not fatal to the core methodological contribution — the architecture and internal ablations are competently executed — but they justify a conditional acceptance: the DESI target-selection claim and external comparison should be tested on a DESI ELG target sample before the abstract's conclusion is accepted. The reader's CONDITIONAL verdict therefore remains appropriate.","tokens_in":24672,"tokens_out":6094,"duration_ms":74437,"concrete_test":"Apply the trained CNN-MLP, with the exact same input pipeline (10-channel 64x64 cutouts and 85 photometric features), to an independent set of DESI EDR (sv1/sv3) or DESI Main Survey ELG targets that have reliable spectroscopic redshifts and satisfy the DESI ELG selection box (g-band magnitude cut and (g-r, r-z) colour criteria from Raichoor et al. 2023). Compute sigma_NMAD and eta separately for 0.6 < z < 1.6 and for the full DESI ELG target sample. If the metrics are close to the headline values (0.014/0.026), the domain-transfer claim holds; if they resemble the paper's faint-source row (0.048/0.069), the claimed direct benefit to DESI target selection is not established.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's headline numbers (sigma_NMAD = 0.0140, eta = 0.0257) are computed on a random split of the compiled sample, whose composition is 64% SDSS sources with median z ~ 0.11 and median r ~ 17.5, plus ~33% DESI EDR/SV sources at z ~ 1.0 and r ~ 22.7 (Table 1). Because the test set mirrors this bimodal distribution, the aggregate metrics are dominated by the bright, low-redshift SDSS population. For the faint subset (r > 21.5), which is the regime relevant to DESI ELG targets (0.6 < z < 1.6, r ~ 22-23), the paper itself reports sigma_NMAD = 0.0479 and eta = 0.0691 (Table 5) — several times worse than the headline. The conclusion that CNN-MLP 'directly benefits the target selection process for DESI' (Section 6) therefore extrapolates from a sample whose composition differs from the DESI ELG target population. No external validation is performed on DESI-selected ELG targets; the 15.78% outlier rate quoted from Zhou et al. (2025) refers to a different sample and evaluation context, so it cannot support the abstract's claim of improvement over other models without reproducing that evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a multimodal deep-learning model called CNN-MLP that combines convolutional neural networks on multi-band DESI Legacy Survey images with a multilayer perceptron on photometric features to estimate photometric redshifts of emission line galaxies. The model is trained and evaluated on a compiled sample of 192,375 ELGs with spectroscopic redshifts from 16 surveys, using a held-out test set. The authors report sigma_NMAD = 0.0140 and an outlier fraction of 2.57% on this test set, compare the full model against their own CNN-only and MLP-only baselines (Table 4), analyze performance as a function of redshift, magnitude, and ELG subtype, and examine outlier characteristics. The paper concludes that the method improves photo-z accuracy and 'directly benefits the target selection process for DESI.'","tokens_in":24948,"tokens_out":4876,"duration_ms":51384,"significance":"If the claims are properly supported, the paper would be a useful contribution to photometric redshift estimation for ELGs, a challenging population for DESI and future surveys. The evaluation uses standard metrics (bias, sigma_NMAD, outlier fraction) on a held-out test set with a redshift distribution matched to the training set, and the architecture and hyperparameter choices are described in sufficient detail to be reproducible. The ablation study against single-modality baselines is a reasonable first step. However, the central comparative claim in the abstract—'Compared to other models, CNN-MLP demonstrates a significant improvement'—is not substantiated by the experiments, which only compare against the authors' own baselines. Furthermore, the claim that the model directly benefits DESI target selection rests on an unvalidated domain-transfer assumption from a sample dominated by bright, low-redshift SDSS galaxies to the faint, high-redshift DESI ELG target population. These issues are load-bearing for the paper's main conclusions and require revision.","major_comments":[{"comment":"The abstract's claim that 'Compared to other models, CNN-MLP demonstrates a significant improvement' is not supported by the experiments reported in the paper. The only comparisons presented are against the authors' own MLP-only and CNN-only baselines (Table 4). No comparison is made to established photometric redshift codes such as EAZY, BPZ, or LePhare, nor to the DESI official photo-z used in target selection. The 15.78% outlier rate quoted from Zhou et al. (2025) in the Introduction refers to a different sample and a different evaluation context, and without reproducing that evaluation on the same test set it cannot substantiate a comparative improvement. The claim should either be supported by an external comparison on a common sample or revised to refer specifically to the single-modality baselines.","section":"Abstract; §5.1; Table 4"},{"comment":"The conclusion that the CNN-MLP model 'directly benefits the target selection process for DESI' is an extrapolation that is not validated by the data. The compiled sample is dominated by bright, low-redshift SDSS galaxies (median r = 18.41, median z = 0.16; Table 1), so the headline metrics on the test set do not represent performance on the DESI ELG target population (0.6 < z < 1.6, r ~ 22-23). Table 5 shows that for objects with MAG_R > 21.5, the one-part CNN-MLP model achieves sigma_NMAD = 0.0479 and an outlier fraction of 0.0691, several times worse than the headline values, and no validation is performed on a sample selected according to the DESI ELG target criteria of Raichoor et al. (2023). To support the DESI target-selection claim, the authors should either evaluate on a DESI-selected ELG sample or substantially qualify the statement.","section":"Section 6; Table 5; Section 2.3"}],"minor_comments":[{"comment":"The text states that the downloaded image data have shape (192375, 6, 64, 64) for the six bands, but the model input has 10 channels (g, r, i, z, g-r, r-i, i-z, W1, W2, W1-W2). Please clarify that the colour-difference channels are constructed from the downloaded images before being fed to the network.","section":"§2.3 and §2.4"},{"comment":"The summation notation 'NcX k=1' appears corrupted in the manuscript; it should read 'z_phot = sum_{k=1}^{N_c} z_k P(z_k)'.","section":"Equation (1)"},{"comment":"In the sentence 'along with the outliers', 'alone' should be 'along'.","section":"§5.4"},{"comment":"The column header 'Model Sources' is unclear; consider renaming it to 'Sample subset' or 'Magnitude range' for readability.","section":"Table 5"},{"comment":"The Data Availability statement provides access to the DESI LS10 catalogue but does not mention availability of the compiled training sample, the trained model weights, or the code. Sharing these would improve reproducibility.","section":"Data Availability"},{"comment":"There is a typo in the reference to York et al. (2000): 'Y ork' should be 'York'. Additionally, the header 'Publications of the Astronomical Society of Australia (2022)' appears to be a template artifact; the correct year should be used.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's core methodology and evaluation on the compiled sample are sound, but the abstract and conclusions overstate the comparative improvement and the direct applicability to DESI target selection. I would recommend requiring the authors to either add a comparison with at least one established photo-z method on a common sample, or revise the claims to match the scope of the experiments. The paper may be suitable for publication after these revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: it's a competent and well-documented application of an image-plus-photometry CNN-MLP to ELG photo-z, with a large and carefully assembled training sample. The headline numbers are probably real for the sample they test on, but they are dominated by bright SDSS galaxies at low redshift, and the paper overclaims the DESI target-selection benefit. The abstract's 'significant improvement' over other models is not supported by any external comparison.\n\nWhat is actually new: the dual-CNN design that keeps the low-resolution NEOWISE images separate from the optical images is a sensible fix for the resolution mismatch, and it appears to help. The sample itself—192k ELGs with spec-z from 16 surveys cross-matched to DESI LS10—is a useful resource. The staged training, permutation-importance feature selection, and classification-with-PDF approach are all standard but executed cleanly. The reported gains over their own single-modality baselines (12.5% and 14.6% in sigma_NMAD) are credible and consistent with earlier multimodal work.\n\nThe soft spots, in order. First, the only comparisons are their own ablations. No established photo-z code, no DESI official photo-z, no re-implementation of prior CNN-MLP work on the same survey. So the abstract's claim is empty as written. Second, the sample is 64% SDSS at median z~0.11 and r~17.5; the aggregate sigma_NMAD of 0.014 is mostly measuring that bright, low-z population. For the faint regime relevant to DESI (r>21.5), Table 5 shows sigma_NMAD~0.048 and outlier fraction~0.069—several times worse. The conclusion that CNN-MLP 'directly benefits' DESI target selection is an extrapolation; no validation on a DESI-selected ELG target sample is performed. Third, the 15.78% outlier rate quoted from Zhou et al. (2025) is from a different sample and evaluation context, so it can't support the improvement claim. Minor: no code released, and feature selection plus hyperparameter tuning both used the same validation set, which is a mild optimism risk.\n\nWho this is for: people working on photo-z for ELGs or faint galaxy samples generally. It's a solid engineering contribution, not a breakthrough. With external baselines and a properly DESI-selected test set, it could be a good paper. As is, the claims outrun the evidence.\n\nRecommendation: send to peer review with a clear request to add external comparisons, report stratified metrics for the faint/high-z DESI-relevant subset, and temper the abstract. It deserves referee time.","headline":"Competent multimodal ELG photo-z work undermined by an unsubstantiated abstract and lack of external comparisons; the headline metrics are dominated by bright low-z SDSS galaxies.","tokens_in":25529,"tokens_out":3372,"would_cite":false,"duration_ms":32103,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A fused CNN-MLP network estimates emission-line galaxy redshifts with 1.4% scatter and a 2.57% outlier fraction.","keywords":["photometric redshifts","emission line galaxies","CNN-MLP","DESI Legacy Surveys","deep learning","galaxies: distances and redshifts","methods: data analysis"],"falsifier":"Take the DESI SV1/SV3 spectroscopic ELG subsample, restrict it to the DESI ELG colour-selection box and $z > 0.6$, and compute $\\sigma_{\\mathrm{NMAD}}$ and the outlier fraction on that subset; if the faint high-redshift metrics come out near the paper's own Table 5 value of roughly $\\approx 0.048$ rather than the advertised $0.0140$, the central performance claim is shown not to transfer to the DESI target population.","tokens_in":24491,"feed_emoji":"🌌","tokens_out":15674,"duration_ms":129498,"temperature":0.7,"pith_summary":"Emission-line galaxies are the main dark-energy tracers for the DESI survey, yet their photometric redshifts are unusually hard to estimate, with catastrophic outlier fractions near 16% in recent DESI validation data. This paper argues that feeding a neural network both multi-band images and catalogue photometry at once, rather than either alone, substantially improves those estimates. On a held-out test set of 192,375 emission-line galaxies with spectroscopic redshifts, the proposed CNN-MLP model reaches $\\sigma_{\\mathrm{NMAD}}=0.0140$ and an outlier fraction of 2.57%, improving on its own image-only and photometry-only baselines by roughly 12.5% and 14.6%, respectively. If that accuracy transfers to the actual DESI ELG target population, it would make photometric pre-selection of targets more reliable and support the cosmological use of ELG clustering.","feed_headline":"Two-branch network cuts galaxy redshift scatter to 1.4%","feed_subtitle":"For emission-line galaxies, the target population of DESI dark-energy surveys, images plus photometry beat either alone.","key_machinery":"The load-bearing object is the CNN-MLP architecture with two deliberately separated image streams: one CNN processes seven-channel optical images (g, r, i, z plus optical colour differences) at 0.262 arcseconds per pixel, and a second, identically structured CNN processes three-channel infrared images (W1, W2, W1$-$W2) at 2.75 arcseconds per pixel. Each CNN uses an inception-module design adapted from earlier photometric-redshift networks, with the galactic extinction $E(B-V)$ concatenated into the image features; an MLP independently processes the 85 photometric features; and a final MLP fuses both streams into a 770-bin redshift classification. The mechanism that makes the fusion work is treating the resolution mismatch between optical and infrared images as a reason for separate branches: naive channel-wise concatenation of all ten image bands degrades performance below the optical-only case.","core_discovery":"The central claim is that a multimodal network, with convolutional branches for optical and infrared images plus a multilayer perceptron for 85 photometric features, estimates photometric redshifts of emission-line galaxies more accurately than either data type alone. With a classification head that divides the redshift range 0 to 3.85 into 770 bins and takes the probability-weighted bin midpoint as the prediction, the model reports bias $=0.0002$, $\\sigma_{\\mathrm{NMAD}}=0.0140$, and an outlier fraction $\\eta=0.0257$ on the test split, against $\\sigma_{\\mathrm{NMAD}}=0.0160$ and $\\eta=0.0284$ for photometry alone and $\\sigma_{\\mathrm{NMAD}}=0.0164$ and $\\eta=0.0316$ for images alone. The paper also finds that performance degrades for faint and high-redshift galaxies, that starforming galaxies are predicted most accurately, and that a single model trained on all magnitudes beats separately trained bright and faint models.","pith_inferences":["The paper's own Table 5 shows $\\sigma_{\\mathrm{NMAD}} \\approx 0.048$ for the faint $r > 21.5$ subset, so the advertised 0.0140 describes the heterogeneous sample as a whole, not the faint colour-selected DESI ELG population at $0.6 < z < 1.6$.","A decisive test the paper does not run is a holdout of DESI SV1/SV3 ELG spectra restricted to the DESI colour-selection box; it would settle whether the model transfers to the actual target population with the data already in hand.","The outlier pattern, where faint low-redshift galaxies are predicted as high-redshift, suggests a redshift-stratified training weighting or synthetic faint-galaxy augmentation could shrink that error tail.","Because the image-only branch still reaches $\\sigma_{\\mathrm{NMAD}}=0.0164$, image-based photometric redshifts could become viable for deep surveys that lack matched multi-band photometric catalogues, a direction the paper gestures at but does not demonstrate."],"forward_implications":["Adding images to photometry cuts $\\sigma_{\\mathrm{NMAD}}$ by about 12.5% relative to photometry alone and about 14.6% relative to images alone, with the outlier fraction dropping by roughly 0.3 to 0.6 percentage points.","The classification formulation outputs a full redshift probability distribution for every galaxy, not just a point estimate, so the same model can support uncertainty-aware target selection.","A single global model trained on all magnitudes outperforms a two-part bright/faint scheme, so the gains do not come from splitting the sample by brightness.","The same image-plus-photometry fusion should transfer to other multi-band surveys with comparable image scales, such as LSST, CSST, and Euclid, which the paper names as downstream applications.","The error analysis identifies faint galaxies and high-redshift sources as the residual problem, so future gains depend on better-represented faint training data rather than on further architectural changes."],"supporting_citations":[{"why":"Supplies the DESI Legacy Surveys DR10 catalogue and imaging from which the 10-channel image inputs and 85 photometric features are built.","marker":"Dey et al. (2019)"},{"why":"Provides the classification-based photometric-redshift framework and CNN design that the imaging branch adapts.","marker":"Pasquet et al. (2019)"},{"why":"Supplies the inception-module CNN architecture and the classification/regression strategy that the imaging branch integrates.","marker":"Treyer et al. (2024)"},{"why":"Demonstrates that combining imaging and photometric data improves photometric redshifts, the multimodal premise the CNN-MLP extends.","marker":"Henghes et al. (2022)"},{"why":"Defines the DESI ELG target-selection colour box and the $0.6 < z < 1.6$ science redshift range the model is meant to serve.","marker":"Raichoor et al. (2023)"},{"why":"Reports the 15.78% catastrophic outlier rate for DESI EDR ELGs that motivates the improved estimator.","marker":"Zhou et al. (2025)"},{"why":"Defines the $\\sigma_{\\mathrm{NMAD}}$ metric used to report the headline 0.0140 accuracy.","marker":"Brammer et al. (2008)"},{"why":"Defines the outlier fraction with the $|\\Delta z| > 0.15$ threshold used throughout.","marker":"Hildebrandt et al. (2012)"},{"why":"Baseline comparison of EAZY and CatBoost on DESI LS10 that frames the machine-learning photo-z problem on this catalogue.","marker":"Li et al. (2024)"}],"fun_headline_variants":["CNN-MLP merges images and photometry to sharpen galaxy redshifts","Multimodal net beats single-input for DESI emission-line galaxy redshifts","Better redshifts for DESI: CNN-MLP uses both images and photometry","One model, two data types: sharper galaxy redshift estimates","Fusing images and photometry cuts galaxy redshift scatter"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a random split of this 192,375-galaxy sample, dominated by bright SDSS galaxies near $z \\sim 0.1$ to $0.2$, represents the faint colour-selected DESI ELG population at $0.6 < z < 1.6$ for which the method is intended.","fun_headline_variants_meta":{"raw":{"variants":["CNN-MLP merges images and photometry to sharpen galaxy redshifts","Multimodal net beats single-input for DESI emission-line galaxy redshifts","Better redshifts for DESI: CNN-MLP uses both images and photometry","One model, two data types: sharper galaxy redshift estimates","Fusing images and photometry cuts galaxy redshift scatter"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000647,"raw_usage":{"total_tokens":3024,"prompt_tokens":1050,"completion_tokens":1974,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":666,"completion_tokens_details":{"reasoning_tokens":1885}},"tokens_in":666,"tokens_out":1974,"duration_ms":15958,"temperature":1.0,"reasoning_tokens":1885,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:34:47.737943+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the DESI SV1/SV3 spectroscopic ELG subsample, restrict it to the DESI ELG colour-selection box and $z > 0.6$, and compute $\\sigma_{\\mathrm{NMAD}}$ and the outlier fraction on that subset; if the faint high-redshift metrics come out near the paper's own Table 5 value of roughly $\\approx 0.048$ rather than the advertised $0.0140$, the central performance claim is shown not to transfer to the DESI target population.","supporting_citations":[{"cited_title":"A., et al","cited_arxiv_id":null,"evidence_quote":"Defines the DESI ELG target-selection colour box and the $0.6 < z < 1.6$ science redshift range the model is meant to serve."}],"review_version":1}