{"id":"d6b16af4-503b-4e1a-b029-747a61a40989","arxiv_id":"2607.06891","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":8,"one_line_summary":"An I2SB diffusion model translates DESI Legacy Survey images to Euclid VIS resolution, recovering structure down to 0.37'' and removing size-dependent biases in Petrosian radii and Sersic parameters.","lead":"A generative model translates blurry ground-based DESI galaxy images into sharper Euclid-quality images, recovering structural detail down to 0.37 arcseconds. This lets astronomers measure galaxy sizes and shapes without the bias caused by Earth's atmosphere, across sky areas where space telescope data doesn't exist yet.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Generalization from Q1 deep fields to DR1 wide survey is the load-bearing gap; a leave-one-field-out test could partially address it now.","rationale":"The reader's identification of the generalization gap as the load-bearing concern is correct. The paper's validation is entirely in-distribution: the test set is spatially disjoint from training but drawn from the same three deep fields with the same observational characteristics. The released E-BGS catalog applies to a different domain (DR1 wide survey), and no cross-domain test has been performed. The CONDITIONAL verdict appropriately reflects this gap. The paper's honesty in framing the release as a prediction to be validated once DR1 is public is commendable, and the methodology is sound within its tested domain. The FRC analysis provides genuine evidence that recovered structure is data-constrained rather than prior-fabricated, and the structural-parameter de-biasing is convincingly demonstrated on the test set. The concern is not about internal inconsistency or methodological error—it is about the scope of validation relative to the scope of the released product. I propose a leave-one-field-out test as a concrete check that could be done now with existing data, which would provide partial evidence for or against generalization before DR1 becomes available. If the leave-one-field-out test shows comparable performance to the current test set, it would strengthen confidence in the DR1 release; if it shows degradation, it would warn of domain-shift risk. Either way, the CONDITIONAL verdict stands: the in-distribution validation is solid, but out-of-distribution generalization remains untested.","tokens_in":11598,"tokens_out":4195,"duration_ms":186135,"concrete_test":"Perform leave-one-field-out cross-validation: train the I2SB model on pairs from two of the three Q1 deep fields and test on the held-out field. Recompute the FRC crossing scale and the three structural-parameter biases (ΔR_Pet, ΔR_e, Δn) on the held-out field. If the FRC crossing degrades by more than ~0.1'' or any parameter bias increases by more than ~50% relative to the current test-set values, the model has learned field-specific features and the DR1 release is at risk of domain-shift bias. This test is feasible now with existing Q1 data and would directly probe cross-sky-region generalization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader correctly identifies the most load-bearing concern. The paper's two validation criteria—FRC phase coherence down to 0.37'' and structural-parameter de-biasing—are both measured on a test set that is spatially disjoint from training but drawn from the same three Euclid Deep Fields at the same nominal Wide Survey depth (Section 2.3). The released E-BGS catalog, however, covers the full Euclid DR1 wide-survey footprint, which spans different sky regions with potentially different stellar densities, galactic extinction, and source distributions. The paper is transparent about this: it frames the DR1 release as a prediction 'to be blindly validated once DR1 is public.' This honesty is appropriate, but the gap between the validated domain and the released catalog is real. If the model has learned field-specific features (e.g., background characteristics, source distributions particular to the deep fields), the released predictions could carry unknown biases that would undermine the 'unbiased' claim in the title. The concern is not that the methodology is flawed—it is sound within the tested domain—but that the central claim of 'unbiased galaxy structures' is only validated in-distribution, while the released product applies out-of-distribution. No other concern is more load-bearing: the FRC analysis is methodologically sound, the Sérsic-fitting approach with a random PSF is reasonable given the predictions lack a true PSF, and the bandpass-offset handling is correct. The generalization gap is the single point where the claim is least secure.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This paper presents an Image-to-Image Schrödinger Bridge (I2SB) model that translates DESI Legacy Imaging Surveys r- and z-band images of Bright Galaxy Survey (BGS) targets into Euclid VIS-resolution images. The model is trained on 63.1 deg² of observed DESI–Euclid Q1 overlap and validated on a spatially disjoint test region. The authors demonstrate, via Fourier Ring Correlation (FRC), that recovered structure is phase-coherent with Euclid ground truth down to 0.37″ (from 1.41″ in DESI r), and that structural parameters (Petrosian radius, Sérsic radius, Sérsic index) measured from the predictions are substantially de-biased relative to DESI measurements. The authors release predictions over the full Euclid DR1 footprint as E-BGS, framing the release as a blind prediction to be validated once DR1 is public. The methodology is sound, the validation framework is well-designed, and the FRC-based test for data-constrained versus prior-invented structure is a genuine methodological strength.","tokens_in":12400,"tokens_out":1114,"duration_ms":118307,"significance":"The paper addresses a real and timely problem: the majority of BGS galaxies lack space-based imaging, and seeing-induced biases in structural parameters cannot be fully removed by post-hoc correction once galaxies are under-resolved. The generative approach, validated against two well-motivated criteria (data-constrained structure and unbiased parameters), is a credible contribution. Particular strengths include: (1) the FRC phase-coherence test, which directly addresses whether recovered structure is genuinely data-constrained rather than prior-fabricated; (2) the transparent framing of the DR1 release as a falsifiable prediction; (3) the release of both code and predictions, enabling community validation; and (4) the honest acknowledgment that scatter is not reduced, with a clear argument that bias removal is what matters for population-level scaling relations. The work is a meaningful step toward enabling unbiased structural measurements ahead of full Euclid coverage.","major_comments":[{"comment":"§2.3, §4: The most load-bearing concern is the gap between the validated domain and the released product. The model is trained and tested on data drawn entirely from the three Euclid Deep Fields (Q1), which, while at nominal Wide Survey depth, may differ from the DR1 wide-survey footprint in stellar density, galactic extinction, background characteristics, or source distributions. The released E-BGS catalog covers the full DR1 footprint, but no test of cross-field generalization is presented. The authors are transparent that this is a prediction to be validated later, which is appropriate. However, a leave-one-field-out cross-validation—training on two of the three deep fields and testing on the third—would directly assess whether field-specific features have been learned and would substantially strengthen the claim that the released predictions are expected to be unbiased. If this is in","section":null}],"minor_comments":[{"comment":"§4.3: The fraction of the full BGS sample retained after the Sérsic-fit quality cut (χ²_ν < 3 on the GT) is not reported. Since the de-biasing results for Re and n are presented only on this clean subset, the reader needs to know what fraction of the population it represents to assess generalizability of the structural-parameter claims to the full BGS.","section":null},{"comment":"Figure 4b: The three FWHM-binned FRC curves are described in the caption but are difficult to distinguish in the figure. Consider using more distinct line styles or a separate panel to make the ~0.04″ variation across PSF bins clearly visible.","section":null},{"comment":"§3.2: The statement that the per-band stretch ensures 'each step only adds structure toward the Euclid endpoint and never carries the DESI structure across' is a strong claim about the mechanism. A brief clarification of how the stretch parameters achieve this (beyond the Appendix A reference) would help the reader evaluate this design choice.","section":null},{"comment":"Abstract: 'approximately 3.8-fold improvement' could be stated more precisely as '~3.8×' for consistency with the body text, where the symbol is used.","section":null},{"comment":"§2.1: The statement that the g-band 'lies almost entirely below the Euclid VIS bandpass' could benefit from a quantitative overlap fraction for precision.","section":null},{"comment":"Table 1 caption: 'All values in native flux units' should specify the unit system (e.g., nanomaggies) for reproducibility.","section":null},{"comment":"Appendix A.5: The Optuna objective function is described as minimizing 'the bias and scatter of its recovered Sérsic and photometric parameters,' but the exact functional form of this score is not given. A brief specification would aid reproducibility.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The reader's stress-test concern about generalization from Q1 deep fields to the DR1 wide survey is well-founded and is the single most important issue. However, the authors' framing of the DR1 release as a blind prediction to be validated later is honest and appropriate for publication, provided that a leave-one-field-out test (or equivalent cross-field validation) is added to demonstrate that the model has not learned field-specific features. The methodology is otherwise sound, and the FRC-based validation is a genuine contribution to the methodology of generative models in astronomy."},"author_rebuttal":null,"desk_editor":{"model":"glm-5.2","letter":"The headline: this paper trains an Image-to-Image Schrödinger Bridge on real DESI–Euclid image pairs (63 deg²) to translate ground-based DESI imaging into Euclid VIS resolution, and validates it with a Fourier-domain framework that explicitly tests whether recovered structure is data-constrained rather than prior-invented. The FRC analysis showing phase coherence down to 0.37″ (from 1.41″ in DESI r-band) is the right diagnostic for the right question, and the structural parameter de-biasing—Petrosian radius, Sérsic radius, Sérsic index—is demonstrated against genuine Euclid ground truth on a spatially disjoint test set. Code and data are released. This is solid, reproducible work that earns its claims within the tested domain. The Sérsic-fit PSF handling (random PSF from the test set for predictions lacking a real one) is a reasonable workaround, and the bandpass-offset discussion is honest about what is and isn't being claimed. The per-band stretch parameters are tuned, but they're dataset-level and fixed before training—no smell of test-set leakage. The main soft spot is real but proportionate: the model is trained and tested on three Euclid Deep Fields, while the released E-BGS catalog covers the full DR1 wide-survey footprint. The paper is transparent about this—it frames the DR1 release as a prediction to be validated once DR1 is public. That's the right posture, but the title's 'unbiased galaxy structures' claim is only validated in-distribution. A leave-one-field-out test across the three deep fields could partially address this now and would strengthen the paper meaningfully. The scatter on Sérsic index (0.609) is large, but the paper correctly notes this reflects irreducible information loss in the under-resolved input and argues that bias, not scatter, is what shifts population-level relations like the mass–size relation. That argument holds. This is for extragalactic survey astronomers who need structural parameters for large ground-based samples ahead of space-based imaging. The methodology is sound, the validation is well-designed, and the released predictions are falsifiable. It deserves a serious referee.","headline":"Generative DESI-to-Euclid image translation works well on the test set; the released catalog's out-of-distribution generalization is untested but honestly framed.","tokens_in":12569,"tokens_out":526,"would_cite":true,"duration_ms":60835,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Generative model sharpens ground-based galaxy images, removing size bias","keywords":[],"falsifier":"If the released E-BGS predictions over the Euclid DR1 footprint, once compared against real Euclid DR1 images, show systematic biases in structural parameters that depend on galaxy size, type, or local noise conditions—or if the Fourier Ring Correlation with real Euclid data falls below the 0.5 threshold at scales coarser than 0.37 arcseconds—the claim of unbiased, data-constrained recovery would fail.","tokens_in":11834,"feed_emoji":"🔭","tokens_out":1189,"duration_ms":219460,"temperature":0.7,"pith_summary":"Ground-based telescope images of galaxies are blurred by Earth's atmosphere, which systematically distorts measurements of galaxy size and structure. This distortion is size-dependent: smaller galaxies are affected more, corrupting the scaling relations astronomers use to trace how galaxies grow. Space telescopes solve the problem but cover only a sliver of the sky. This paper trains a generative diffusion model called an Image-to-Image Schrödinger Bridge on the small overlap between ground-based DESI imaging and space-based Euclid imaging, learning to translate blurry DESI images into sharp Euclid-quality images. The key mechanism is a stochastic bridge: the model learns to reverse a gradual blurring process that connects a sharp Euclid image to its blurry DESI counterpart, recovering structure that the atmosphere destroyed. The authors show, using Fourier-domain analysis, that the recovered structure is genuinely constrained by the input data down to 0.37 arcseconds—a 3.8-fold improvement over the 1.41-arcsecond DESI baseline—rather than being fabricated by the model's learned prior. At this recovery level, the systematic biases in three standard structural parameters (Petrosian radius, Sérsic radius, and Sérsic index) are essentially eliminated, removing the size-dependent distortion that plagues ground-based measurements. The authors release their translations across the full Euclid DR1 footprint as a dataset called E-BGS, which can be validated once Euclid's real images become public.","feed_headline":"Generative model sharpens ground-based galaxy images, removing size bias","feed_subtitle":"A diffusion bridge trained on 63 sq deg of DESI–Euclid pairs recovers structure 3.8× finer and de-biases galaxy sizes ahead of Euclid's full","key_machinery":"The Image-to-Image Schrödinger Bridge (I2SB): a diffusion model that defines a stochastic path between two fixed image endpoints—a sharp Euclid image and its blurry DESI counterpart. The forward process gradually blurs the Euclid image into the DESI one; a neural network learns to reverse this path, recovering sharp structure from the blurry input. The bridge is stochastic (not a fixed interpolation), which lets it express the one-to-many nature of recovering fine structure from degraded data. A Fourier Ring Correlation test checks whether recovered structure matches the ground truth in phase, not just in power, distinguishing genuine data-constrained recovery from prior-driven invention.","core_discovery":"A bridge diffusion model trained on 63 square degrees of real DESI–Euclid image pairs can translate blurry ground-based galaxy images into near-space-based resolution well enough to remove the systematic, size-dependent biases in structural parameters. The recovery is data-constrained—confirmed by phase coherence in the Fourier domain down to 0.37 arcseconds—rather than prior-driven fabrication, and the de-biasing holds across the Petrosian radius, Sérsic radius, and Sérsic index simultaneously.","pith_inferences":["The claim that bias removal matters more than scatter for population studies is well-taken, but it implicitly assumes that the residual scatter is uncorrelated with galaxy properties. If the model's plausible-but-wrong predictions for unresolved central profiles cluster systematically at certain masses or redshifts, the scatter could still bias scaling-relation slopes even if the mean bias is zero","The model's trusted scale of 0.37 arcseconds is set by where the Fourier Ring Correlation drops to 0.5, but this is a population-level statistic. Individual galaxies—especially the smallest or most compact ones—may have substantially worse recovery, meaning the de-biasing may not be uniform across the full BGS sample.","The approach could be extended to recover structural parameters that are currently impossible to measure from the ground at all, such as bulge-to-disk decomposition or central velocity dispersion proxies, if the model can be shown to preserve the relevant small-scale information faithfully."],"forward_implications":["Population-level studies of galaxy structure—such as the mass–size relation—can proceed at near-Euclid resolution across the full DESI BGS footprint before Euclid imaging is available, using the released E-BGS dataset.","The phase-coherence validation method (Fourier Ring Correlation) provides a general framework for testing whether any generative image-translation model in astronomy recovers real structure versus fabricating plausible-looking but unconstrained output.","If the E-BGS predictions are confirmed against real Euclid DR1 data, the same bridge-diffusion approach could be applied to other ground-to-space translation problems in astronomy, such as HSC-to-Hubble or LSST-to-Roman."],"fun_headline_variants":["Diffusion model translates DESI galaxy images to near-Euclid resolution","Generative bridge removes size bias from ground-based galaxy structures","DESI-to-Euclid image translation de-biases galaxy structural parameters","Fourier-validated diffusion model recovers 0.37\" galaxy structure from DESI","Trained on real DESI-Euclid pairs, model de-biases galaxy sizes 3.8× sharper"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The model is trained on 63 square degrees of Euclid Deep Fields and applied to the full Euclid DR1 footprint, but the deep fields may have different noise properties, depth, or source distributions than the wide survey. The paper does not test whether the model generalizes across genuinely distinct sky regions beyond the spatially disjoint test set carved from the same deep-field data.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion model translates DESI galaxy images to near-Euclid resolution","Generative bridge removes size bias from ground-based galaxy structures","DESI-to-Euclid image translation de-biases galaxy structural parameters","Fourier-validated diffusion model recovers 0.37\" galaxy structure from DESI","Trained on real DESI-Euclid pairs, model de-biases galaxy sizes 3.8× sharper"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":670,"prompt_tokens":563,"completion_tokens":107,"prompt_tokens_details":null},"tokens_in":563,"tokens_out":107,"duration_ms":20417,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T23:27:33.503682+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If the released E-BGS predictions over the Euclid DR1 footprint, once compared against real Euclid DR1 images, show systematic biases in structural parameters that depend on galaxy size, type, or local noise conditions—or if the Fourier Ring Correlation with real Euclid data falls below the 0.5 threshold at scales coarser than 0.37 arcseconds—the claim of unbiased, data-constrained recovery would fail.","supporting_citations":[],"review_version":1}