{"id":"b88bd997-ef75-40bb-8aed-4eacc7078ce9","arxiv_id":"2607.14525","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Conditional normalising flows trained on mock multi-band images recover the redshift distribution of unresolved galaxies with sub-percent accuracy in mean and width, under idealized simulation-matched conditions.","lead":"A new machine-learning framework estimates the redshift distribution of galaxies too faint to be detected individually, using only the diffuse background light they emit together, in simulated maps matching the Euclid and LSST surveys. If it holds up on real data, it would let cosmologists use the majority of galaxies that remain unresolved as an extra probe of cosmic structure.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unresolved sample is built from VIS-only SExtractor detection (§2.1), so simulated background maps contain galaxies that a real Euclid+LSST analysis would detect and remove; the sub-percent accuracy applies to a population different from the actual unresolved sample.","rationale":"After reviewing the paper, I find the internal simulation logic largely consistent: training and test are from independent regions, the flow is appropriately conditioned, and the robustness tests are honest. The reader's verdict of CONDITIONAL is appropriate. The single most load-bearing concern I see is not the generic fidelity of Flagship (which the paper acknowledges and conditions on), but a concrete mismatch between the unresolved-sample definition in §2.1 and the paper's multi-band survey premise. The unresolved population is constructed using VIS-only detection; because LSST r/g/i are ~1.4 mag deeper than VIS, a significant population of VIS-undetected galaxies will be LSST-detectable and would be removed in a real joint analysis. These galaxies are thus incorrectly included in the simulated unresolved background and in the target redshift distribution. The sub-percent accuracy is measured on this composite population; the actual unresolved population of the Euclid+LSST background would be different. This is not 'outside current consensus' but an internal mismatch that can be settled by a controlled simulation test. The lack of error bars on Table 2 is a lesser issue because of the very large number of galaxies/pixels, and the paper's tomographic results provide corroborating internal evidence. The proposed test—rebuilding the unresolved sample with multi-band detection and retraining—would directly determine whether the headline claim survives a more realistic definition of 'unresolved.' Until then, the claim should be read as conditional on the VIS-only detection recipe.","tokens_in":16680,"tokens_out":12337,"duration_ms":134346,"concrete_test":"Repeat the unresolved-sample construction of §2.1 using SExtractor detection on the coadded r-band (or on a detection image combining VIS and LSST g/r/i), removing every galaxy detected in any band; then retrain the conditional flow on the new unresolved sample and evaluate on an independent test region with the same matched-noise protocol as §4. Compare the resulting unresolved sample's VIS magnitude/redshift distributions and the fiducial Δμ/Δσ of the VIS flux-weighted redshift density to Table 2. If the sample loses a substantial fraction of galaxies (especially with 25≲VIS≲27, where LSST is deeper) and the error grows beyond ~1%, the headline result is specific to the VIS-only detection definition and must be qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.1 defines the unresolved population by running SExtractor on the Euclid VIS band only and then applying a VIS>24 cut. But the paper's stated target is the unresolved background light in the ten-band Euclid+LSST synergy. In real data, object detection/removal would use all available bands; a galaxy undetected in VIS can be detected in the deeper LSST bands (e.g., r-band 5σ depth 27.15 vs VIS 25.70, Table 1). Such galaxies are therefore present in the simulated 'unresolved' maps and in the training/evaluation target distributions, even though they would be resolved and removed in the actual survey. The model's sub-percent recovery is thus measured for a mixed sample of VIS-undetected galaxies, not for the true unresolved population of the combined survey. This is not a generic simulation-realism worry: it is a mismatch between the operational definition used in the validation and the definition implied by the multi-band survey context. It directly affects the external validity of the headline claim, regardless of how faithfully Flagship represents the galaxy population. The paper's caveat about 'future object-detection and removal techniques' (Section 2.1) assumes perfect subtraction, but the simulation does not actually subtract all sources detectable by the survey.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a machine-learning framework for estimating redshift distributions of unresolved galaxies from multi-band surface brightness maps, targeting the Euclid and LSST synergy. Using the Flagship mock, it constructs unresolved samples by simulating Euclid VIS images, detecting and removing sources with SExtractor, and projecting the remaining galaxies into HEALPix maps at Nside=4096. A conditional normalizing flow is trained on ten-band pixel fluxes to predict joint distributions of redshift, VIS magnitude, and VIS noise RMS. On a held-out region with matched imaging conditions, the model recovers the VIS flux-weighted redshift density with sub-percent accuracy in the mean and standard deviation (Table 2, Fiducial row). The paper also tests missing-band imputation, noise-level shifts, and pixel-level tomographic binning, and proposes using the predicted noise RMS as a diagnostic for train-target mismatch.","tokens_in":17044,"tokens_out":6243,"duration_ms":60878,"significance":"If the demonstrated accuracy transfers to real data, the method would enable new cosmological analyses of the unresolved cosmic optical background, which contains the majority of galaxies. The methodological novelty is the use of conditional normalizing flows for map-based redshift distribution inference, with a built-in observable diagnostic. Strengths include a null test, a held-out evaluation region, independent noise realizations, and reproducible open-source components (MultiBand_ImSim, FlowJAX). The central limitation is that validation is entirely within the Flagship simulation, so the sub-percent accuracy is a statement about the mock world; the paper is transparent about this. The most serious concern is the definition of the unresolved sample, which is based on VIS-only detection despite the ten-band Euclid+LSST context.","major_comments":[{"comment":"The unresolved sample is defined by running SExtractor on Euclid VIS images only, then applying a VIS>24 cut. The stated target is the unresolved background in the ten-band Euclid+LSST synergy. LSST bands are substantially deeper (e.g., r-band 5σ depth 27.15 vs VIS 25.70). Galaxies that are undetected in VIS but detected in LSST are therefore retained in the simulated unresolved maps and in the training/evaluation target. In a real analysis these galaxies would be resolved and removed. The sub-percent accuracy in Table 2 therefore applies to a mixed population of VIS-undetected galaxies, not to the actual unresolved population of the combined survey. This is a definitional mismatch, not only a simulation-realism concern. Please re-run the detection using all ten bands (or a combined detection image), or explicitly re-scope the claims to a VIS-only unresolved definition.","section":"§2.1, Table 1"},{"comment":"The validation is entirely internal to the Flagship simulation. The held-out region and independent noise realizations are out-of-sample only within the same galaxy-formation and SED model. The paper acknowledges that population mismatches 'cannot be fully captured within our current simulation framework' (Section 5), but this is the central source of external-validity risk for the headline sub-percent claim. Unless a test with an independent mock (e.g., hydrodynamical or semi-analytic) is added, the abstract and conclusions should state more prominently that the accuracy is demonstrated only for the Flagship galaxy population, and that transfer to real data is untested.","section":"§5 ('Several avenues')"},{"comment":"The ±3σ noise-shift tests show mean biases of -3.64e-2 and +3.67e-2 in the VIS flux-weighted redshift density, i.e., ~3.6% in the mean. The paper uses this to motivate the noise-RMS diagnostic, which is reasonable, but the term 'robust' in the abstract and conclusion should not be read as unbiased recovery under noise mismatch. The work demonstrates detection of the mismatch, not correction. Consider explicitly distinguishing 'diagnostic capability' from 'recovery accuracy' in the summary.","section":"Table 2 (Shift RMS rows)"}],"minor_comments":[{"comment":"The two-iteration blending safeguard is heuristic; please state how many galaxies remain after the VIS>24 cut and how sensitive the results are to the cut.","section":"§2.1"},{"comment":"Mean imputation for 20% missing pixels assumes missingness at random; real missing data (e.g., bad pixels, chip gaps) may be spatially correlated. A brief note would be useful.","section":"§4.1"},{"comment":"The chosen tomographic bin edges (0,1.2,1.4,1.6,3) are not motivated. The lowest bin is very broad and contains the sharp features the model smooths; report the source fraction per bin to aid interpretation.","section":"§4.3"},{"comment":"Dashed/solid line legend: clarify which colors correspond to which scenarios in the caption; currently the text refers to colors but the figure caption does not list them.","section":"Figure 6"},{"comment":"The reference to Bogachev et al. for universal approximation could be supplemented by a more standard normalizing-flow approximation theorem reference; not essential.","section":"Appendix A"},{"comment":"'Sub-per cent accuracy' appears before the caveat 'when the training and target samples are statistically well matched'; consider moving the caveat earlier to avoid overstatement.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper is well written and the experiments are thoughtfully designed. The main concern is the mismatch between the VIS-only unresolved sample definition and the ten-band survey context; this should be fixed or explicitly scoped. I do not have concerns about the integrity of the work; the sim-based validation is appropriate for a methods paper, but the claims should be calibrated to what is demonstrated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nShort version: genuinely new problem setup, careful mock validation, honest writing—but the headline accuracy is validated on a sample that is not the actual unresolved population of the combined Euclid+LSST survey. The stress-test concern is real and the paper does not fully answer it.\n\nWhat's new: first framework to estimate redshift distributions of unresolved galaxies directly from multi-band background-light maps, using conditional normalising flows. The problem reformulation is valuable—most galaxies will be unresolved even in Stage-IV surveys, and the diffuse COB is a potentially rich tracer. The simulations are careful: Flagship mock, Euclid Q1 and LSST ten-year depths, tile-to-tile noise and PSF variations, HEALPix projection. Validation is genuine out-of-sample: held-out sky region, independent noise realizations, a pure-noise null test, and sensible metrics (Δμ, Δσ, KL, EMD). The diagnostic idea—including the observable VIS noise RMS as a target variable to flag train-target mismatch—is clever and useful. The test of robustness to missing bands with imputation is also reasonable.\n\nThe soft spot: Section 2.1 defines the unresolved population by running SExtractor on Euclid VIS only, then applying a VIS>24 cut. But the paper's stated context is the ten-band Euclid+LSST synergy. In real data, object detection and removal would use all bands. LSST r-band reaches 27.15 (5σ) against VIS 25.70 (Table 1); a galaxy invisible in VIS can be detected in r and would be resolved and removed from the background map. The simulated 'unresolved' maps therefore include galaxies that a real multi-band pipeline would subtract, and the training and evaluation target distributions contain those galaxies. Sub-percent accuracy thus applies to a mixed sample of VIS-undetected sources, not to the true unresolved population of the combined survey. This is a mismatch between the operational definition used in validation and the stated multi-band application, not just a generic simulation-realism worry. The paper's caveat about 'future object-detection and removal techniques' (Section 2.1) assumes perfect subtraction, but the simulation does not actually subtract all survey-detectable sources. Section 5's admission that population mismatches can't be captured is a broader point; it doesn't cover this specific band-selection issue.\n\nMinor soft spots: no code or data shipped, no error bars on the headline metrics (single training run), and Flagship's mass resolution limits the faint end (no sources fainter than H~26). All addressable.\n\nWho it's for: people working on diffuse optical backgrounds, component separation, survey synergy, and map-level ML inference. It deserves a serious referee—the idea is novel and the internal evidence is consistent—but the claim needs to be either restricted to the VIS-defined unresolved population or re-run with a multi-band detection pipeline.\n\nRecommendation: send to peer review, and ask the authors to address the band-selection mismatch explicitly.\n\nBest,\n[You]","headline":"New and useful problem setup, honest simulation work, but the sub-percent claim is calibrated on a VIS-only unresolved sample that doesn't match the multi-band unresolved population the paper aims for.","tokens_in":17445,"tokens_out":3600,"would_cite":true,"duration_ms":35624,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A conditional normalising flow trained on simulated ten-band images can infer the redshift distribution of galaxies too faint to be individually detected, recovering the flux-weighted redshift density to sub-percent accuracy in mean and wid","keywords":["redshift estimation","unresolved galaxies","cosmic optical background","conditional normalising flows","background light maps","flux-weighted redshift distribution","tomographic binning","multi-band photometry"],"falsifier":"Take real multi-band images from an overlap region of a space survey and a ground-based survey, run the model to predict the VIS flux-weighted redshift distribution of the unresolved background, then measure the same quantity by stacking spectroscopic or high-confidence photometric redshifts of the same faint population (e.g., from deep spectroscopy or SED fitting aided by deeper space data). If the recovered mean redshift differs by more than about 1 percent while the noise-RMS diagnostic shows no mismatch, the claimed accuracy would be shown not to transfer. Alternatively, retrain the flow o","tokens_in":16627,"feed_emoji":"🌌","tokens_out":5760,"duration_ms":59962,"temperature":0.7,"pith_summary":"This paper argues that the redshift distribution of the vast, mostly unresolved galaxy population can be recovered statistically from maps of the unresolved optical background light alone, without resolving individual galaxies. The authors train a conditional normalising flow to output per-pixel probability distributions of redshift, VIS magnitude, and imaging noise, conditioned on ten-band pixel values from mock images built to match a space survey and a ground-based survey. On an independent matched test region, the recovered VIS flux-weighted redshift distribution matches the truth to better than one percent in its mean and standard deviation, with a missing-band imputation strategy keeping performance close to that even when photometric coverage is incomplete. The paper also shows that adding an observable like the noise RMS to the target variables gives a built-in check for mismatches between simulation and observation, and that the pixel-level predictions support tomographic binning. If it holds, this would turn the cosmic optical background into a quantitative cosmological tracer.","feed_headline":"Hidden galaxies' redshifts read from background glow","feed_subtitle":"Map-based normalising flows turn the cosmic optical background into a tomographic probe, at sub-percent accuracy.","key_machinery":"Conditional normalising flow with coupling layers and rational-quadratic splines: a generative model that turns a simple probability distribution into a target distribution through a stack of invertible transformations, with a conditioning network steering the transformations by observed data. In this work, the flow is conditioned on ten-band pixel values of unresolved background light and outputs a joint density over redshift, VIS magnitude, and VIS noise RMS, giving tractable per-pixel conditional distributions. The design choices doing the heavy lifting are using flux ratios between adjacent bands plus the VIS flux as conditioning inputs, standardising all variables, and dedicating one ta","core_discovery":"The central discovery is that the redshift and brightness distributions of unresolved galaxies—objects whose individual light is lost in the background—are encoded in the joint statistics of multi-band surface brightness fluctuations, and that a conditional normalising flow can decode this encoding. Trained by maximum likelihood on mock multi-band images with realistic point-spread functions, depths, and tile-to-tile noise variations, the model represents the target density p(z, m_VIS, sigma_RMS | ten-band fluxes). On the fiducial evaluation set (an independent sky region, independently generated noise), the aggregate VIS flux-weighted redshift distribution is recovered with a mean shift of","pith_inferences":["If the sub-percent accuracy persists on real data from a Euclid-like and LSST-like overlap, cross-correlating redshift-binned background-light maps with resolved galaxy surveys would test galaxy population models and potentially constrain the faint-end slope of the luminosity function—a step the paper leaves implicit.","The authors' diagnostic trick of targeting an observable systematic generalises beyond noise RMS: the same flow machinery could predict PSF size, astrometric residuals, or foreground cirrus amplitude, turning any known systematic into a built-in test for simulation-reality mismatch.","A concrete testable extension would be to train the same architecture on two independent mocks built from different galaxy-formation models; the spread in predicted n(z) would quantify the systematic error budget that the paper identifies as the main obstacle to real-data use.","The findings suggest that statistical smoothing inherent to normalising flows may suppress sharp low-redshift features; if such features matter, hybrid methods combining flow densities with explicit clustering priors could be needed."],"forward_implications":["Unresolved background light can be used as a tomographic large-scale-structure tracer: per-pixel redshift estimates allow pixel-level tomographic binning, recovering redshift distributions per bin with sub-percent or near sub-percent accuracy.","Missing photometric bands need not invalidate the method: with mean-based imputation from remaining bands or from the VIS band, redshift estimates stay close to the fiducial performance.","The noise-RMS diagnostic provides a practical safeguard: because the flow predicts an observable quantity as part of its target, users can detect when the training simulation does not match the observed imaging conditions before trusting the redshift output.","The framework extends naturally to flux-weighted redshift distributions in other bands by adding those magnitudes to the target vector, enabling band-by-band component separation for the cosmic optical background."],"fun_headline_variants":["Normalizing flows decode redshift from background glow","Background glow yields galaxy redshifts at sub-percent precision","Unresolved galaxies' redshifts mapped from cosmic background","AI decodes hidden galaxy redshifts from background light maps","Sub-percent redshift estimates for unresolved sources from glow maps"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The results stand or fall on whether the simulated galaxy population used for training faithfully represents the real unresolved population—including galaxies fainter than the survey limit—and on whether real source removal leaves a residual population matching the simulation's 'undetected' sample; the paper itself notes this mismatch cannot be fully captured or easily diagnosed.","fun_headline_variants_meta":{"raw":{"variants":["Normalizing flows decode redshift from background glow","Background glow yields galaxy redshifts at sub-percent precision","Unresolved galaxies' redshifts mapped from cosmic background","AI decodes hidden galaxy redshifts from background light maps","Sub-percent redshift estimates for unresolved sources from glow maps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000388,"raw_usage":{"total_tokens":1875,"prompt_tokens":729,"completion_tokens":1146,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":473,"completion_tokens_details":{"reasoning_tokens":1071}},"tokens_in":473,"tokens_out":1146,"duration_ms":8183,"temperature":1.0,"reasoning_tokens":1071,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T01:48:29.425923+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take real multi-band images from an overlap region of a space survey and a ground-based survey, run the model to predict the VIS flux-weighted redshift distribution of the unresolved background, then measure the same quantity by stacking spectroscopic or high-confidence photometric redshifts of the same faint population (e.g., from deep spectroscopy or SED fitting aided by deeper space data). If the recovered mean redshift differs by more than about 1 percent while the noise-RMS diagnostic shows no mismatch, the claimed accuracy would be shown not to transfer. Alternatively, retrain the flow o","supporting_citations":[],"review_version":1}