{"id":"85d55664-87bc-4a94-9f4c-c832a0c2e229","arxiv_id":"2411.19158","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A diffusion-based deconvolution method turns HSC ground-based images into HST-like reconstructions and uses a posterior-to-prior variance ratio to flag hallucinated features.","lead":"This paper applies a diffusion model trained on simulated galaxies, together with diffusion posterior sampling, to deconvolve ground-based Subaru/HSC images and compares the results with Hubble Space Telescope images. It also proposes a variance-ratio metric meant to flag image features that come from the model prior rather than from the data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed HST-comparable resolution lacks a quantitative resolution test; if the deconvolved output is only compared visually, the central claim is unsupported.","rationale":"The reader identified the prior mismatch as the weakest point. I agree that is a genuine vulnerability, but the more immediate and more directly testable load-bearing gap is the lack of a quantitative resolution assessment. The prior mismatch affects reliability at high redshift and for faint galaxies, and the paper documents that failure. The central claim, however, is about achieving 'HST-comparable resolutions' for the cases shown, and that claim is supported only by qualitative side-by-side images and by a variance-ratio metric that has not been validated as a resolution measure. The variance-ratio metric measures consistency with the prior, not fidelity to the true object. The proposed concrete test would settle the resolution claim and also reveal where the prior mismatch has real consequences. If that test passes, the central claim is substantially strengthened. If it fails, the claim should be weakened to 'reconstructions that visually resemble HST images,' which is still a useful proof-of-concept. This is why I recommend keeping the CONDITIONAL verdict: the paper should be accepted only if a quantitative resolution test (or a clear quantitative comparison) is added. My concern is therefore not a full rejection, but a specific condition. The reader's proposed additional validation of the variance-ratio metric on known hallucinated structures is useful but secondary; the primary condition is that the resolution claim must be quantified.","tokens_in":9043,"tokens_out":1685,"duration_ms":13703,"concrete_test":"Use the TNG-HSC simulated image set (Section 3.1): take idealized TNG galaxies, convolve with the actual HSC PSF, add HSC noise, then run DPS. On the output, compute quantitative resolution metrics versus the known idealized ground truth: e.g., cross-correlation coefficient within the half-light radius, recovered half-light radius, and Fourier ring correlation with an information threshold. If the recovered half-light radius of faint or compact galaxies is systematically biased or the FRC indicates resolution only at scales larger than the HST PSF, then the HST-comparable resolution claim fails. The same metrics should be applied to the real HSC/HST pairs (Figures 4 and 6) for a small quantitative check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, stated in the abstract, is that DPS with a TNG-trained DDPM prior deconvolves real HSC images 'to resolutions comparable to those obtained by HST images.' The evidence for this is mainly visual comparison between deconvolved HSC and downsampled HST images (Figures 4, 6, 8), plus a posterior variance ratio. However, because the prior is trained on idealized TNG galaxies, the deconvolved images can contain galaxy-like structures that originate from the prior, not from the data. On small galaxies with limited resolution, the method does not necessarily recover HST-equivalent information; it can render features that resemble the prior's training examples. The paper acknowledges this failure at high redshift (Section 3.3, Figure 8), showing that prior-driven features dominate when the inverse problem demands upsampling. This weakness is particularly relevant because the two successful comparisons are low-redshift galaxies (z ~ 0.17 and ~ 0.22) where the HSC PSF may already resolve most of the structure. The claim therefore needs a quantitative test of spatial resolution on a sample with known ground truth, not merely a visual check. Actually, the HSC images and the HST images are compared only visually; no quantitative metric such as cross-correlation with the HST reference, resolution criteria (e.g., half-light radius recovery, peak signal-to-noise, Fourier ring correlation), or residual power spectrum is provided. Without such a metric, the central claim is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a Bayesian deconvolution pipeline for ground-based astronomical images, combining a DDPM prior trained on idealized TNG100 simulation images with the Diffusion Posterior Sampling (DPS) algorithm. The forward model includes a redshift- and pixel-scale-dependent operator that maps idealized TNG images to HSC-like observations. The method is applied to HSC PDR3 images, and the abstract claims that the deconvolved images reach resolutions comparable to HST images, while also proposing a posterior-to-prior variance-ratio metric to identify prior-driven features. The paper includes results on simulated TNG observations, visual comparisons of deconvolved HSC images with HST images for two low-redshift objects, a high-redshift failure case, and a small simulated test of the hallucination metric as a function of galaxy magnitude.","tokens_in":9287,"tokens_out":3669,"duration_ms":35921,"significance":"If the central claims are established, the work would be valuable: it offers a publicly available implementation, includes redshift and pixel scale in the forward model, and attempts to provide a pixel-level uncertainty-aware tool for deconvolution, which is important for scientific use of generative-model-based restoration. The paper is honest about some limitations, particularly the high-redshift behavior, and it makes a genuine effort to quantify prior-driven features. However, the current evidence does not yet establish the headline claim of HST-comparable resolution, and the hallucination-metric validation is weakened by its dependence on the same simulation family used to train the prior. These issues are load-bearing for the paper's main conclusions, so the manuscript requires substantial additional work rather than minor polishing.","major_comments":[{"comment":"The central claim that deconvolved HSC images reach 'resolutions comparable to those obtained by HST images' is supported only by visual side-by-side comparisons for two low-redshift objects. No quantitative metric (e.g., Fourier ring correlation, half-light radius recovery, cross-correlation with the HST reference, or residual power spectrum) is provided. Because the prior is trained on simulated galaxies, visual similarity to HST images can arise from prior hallucination rather than from genuine resolution recovery. A quantitative comparison against HST, or against simulated ground truth with known PSF and noise, is needed to substantiate this claim.","section":"3.2, Figures 4-6"},{"comment":"The paper acknowledges that at z > 0.7 the reconstruction is highly prior-dominated, with features hallucinated around the galaxy that are invisible in the HST counterpart. This is a load-bearing limitation because the abstract states the HST-comparable resolution claim without this qualification. The manuscript should either restrict the claim to the regime where the method is validated or provide a quantitative characterization of where the transition to prior domination occurs (e.g., in terms of object size, SNR, and upsampling factor). Without this, the general claim is not supported by the presented evidence.","section":"3.3, Figure 8"},{"comment":"The validation of the variance-ratio metric uses simulated TNG images that are drawn from the same simulation family used to train the DDPM prior. This can show that the metric responds to changes in SNR within the training distribution, but it does not establish that a low variance ratio indicates absence of hallucination for real HSC galaxies, whose morphologies may differ from idealized face-on TNG galaxies. The test also uses only a single galaxy with six magnitude values and reports no uncertainties or repeated draws. Validation on an independent simulation suite or on real HST images with artificially degraded input would be needed to support the metric as a general tool.","section":"3.4, Table 2"},{"comment":"The variance-ratio metric is not fully specified. The text says 'we take the pixel-wise posterior variance and prior variance over 256 samples, take their ratio and compute the average value on a centered aperture of radius r,' but it does not define what 'prior variance' means (variance over samples from the unconditional prior? over posterior samples with different seeds? over pixels?) nor whether the ratio is pixel-wise before aperture averaging. The subsequent claim that a central ratio close to zero means the model is 'not hallucinating' is an interpretive assumption that is not derived or tested. These details must be clarified for the metric to be reproducible and interpretable.","section":"3.4, variance-ratio definition"}],"minor_comments":[{"comment":"The text 'camera field-of-view face-on (also referred to asv0)' contains a typo: 'asv0' should be 'as v0' or the FITS keyword should be formatted consistently.","section":"2, TNG paragraph"},{"comment":"The schematic in Figure 1 is difficult to parse, particularly the meaning of the nested braces and the relation between the input size n and the observed size m(z,p). The text should state explicitly that when m(z,p) < n the inverse problem includes an upsampling component, since this is central to the high-redshift limitation.","section":"Figure 1"},{"comment":"The FID values are reported without uncertainties or a comparison baseline, so the statement that the model 'improves' from 300k to 525k steps is only weakly supported. The FID at 525k (132.6) is also considerably higher than the 450k value (119.4), which the text attributes to the random shift; a quantitative check of this explanation would be useful.","section":"Table 1"},{"comment":"The aperture radius r = 30 pixels is fixed, and no sensitivity analysis is provided. Since the variance-ratio values in Table 2 are small and monotonically increasing with magnitude, the result should be shown to be robust to the choice of r.","section":"3.4"},{"comment":"Appendix C presents visual results at all simulated magnitudes, but no quantitative residual between the posterior mean and the true idealized TNG image is given. A simple metric such as the residual RMS or peak SNR would strengthen the claim that the variance ratio correlates with reconstruction quality.","section":"3.4, Appendix C"},{"comment":"The abstract states that the method 'reaches resolutions comparable to those obtained by HST images,' while the conclusion is more cautious, saying only that the paper provides a foundation and a metric. The wording should be aligned so that the abstract does not overstate what the current evidence supports.","section":"Conclusions and abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper is promising and clearly written, with a useful public implementation and an honest discussion of limitations. The main gap is that the headline claim of HST-comparable resolution is not quantitatively validated, and the hallucination-metric validation is tied to the same simulation family as the prior. These are fixable with additional experiments and analysis, but they are load-bearing for the paper's conclusions. I would not recommend rejection, but the revision needs to be substantial."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is a solid, honest proof-of-concept for using diffusion posterior sampling with a simulation-trained prior to deconvolve real ground-based survey images. The central claim—that you reach HST-comparable resolution—is plausible but not actually demonstrated, because the only comparison to HST is visual side-by-side for two low-redshift galaxies. There is no quantitative resolution or fidelity metric in the paper. That is the main weak spot, and I think the stress-test note lands.\n\nWhat is genuinely new: adapting the forward operator to include redshift and pixel scale, which makes the tool portable to other surveys. The variance-ratio metric, posterior variance divided by prior variance, is a sensible idea for flagging regions where the prior is driving the reconstruction. The authors also deserve credit for being transparent about training difficulties (mode collapse, compute time) and for explicitly acknowledging that at z > 0.7 the model produces prior-dominated, hallucinated features. Code and data are public, which makes the work reproducible.\n\nThe soft spots are real but addressable. The quantitative validation of the variance-ratio metric is thin: six magnitude points on simulated TNG galaxies, no error bars, and those simulated images come from the same simulation suite that trained the prior, so the test is mildly circular. The paper also does not compare its metric to existing hallucination-detection approaches, such as Sampson and Melchior. And the prior is trained only on idealized, face-on TNG galaxies at z ~ 0.15–0.18, so the successful comparisons are exactly the regime where the HSC PSF already resolves most of the structure. That limits the scope of the claim.\n\nNone of this is fatal. The paper is a working prototype with honest limitations, not an overhyped claim. The math is standard DPS/DDPM, and the derivation is not circular. I would send it to peer review, but the referee should require a quantitative comparison against HST (or a realistic simulated reference) on a larger, systematic sample, and a proper validation of the variance-ratio metric against known hallucinated structures. With that, it could be a useful contribution.\n\nRecommendation: engage with it, cite it for the method and the cautionary example, but do not rely on the HST-comparable resolution claim until it is measured.","headline":"Honest proof-of-concept with a useful prior-dominated-feature flag, but the HST-resolution claim needs a quantitative test before it can be trusted.","tokens_in":9859,"tokens_out":3605,"would_cite":true,"duration_ms":31286,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a diffusion model trained on simulated galaxies can deconvolve ground-based survey images to Hubble-like resolution, and that a variance-ratio metric flags which reconstructed pixels are prior-driven rather than…","keywords":["astronomical image deconvolution","diffusion models","Diffusion Posterior Sampling","galaxy morphology","prior-driven features","hallucination quantification","HSC survey","TNG Illustris simulations"],"falsifier":"Take a sample of HSC galaxies with independent space-based imaging (for example JWST) at redshifts above 0.7, and compare each deconvolved reconstruction with the space-based truth. If regions flagged by the variance-ratio metric as prior-driven turn out to match the deeper space-based data, or if regions flagged as data-driven turn out to be absent in that data, the metric's trust map is falsified.","tokens_in":8802,"feed_emoji":"🔭","tokens_out":5226,"duration_ms":43036,"temperature":0.7,"pith_summary":"The paper tries to establish that a score-based diffusion model trained purely on idealized simulated galaxies can act as a Bayesian prior for deconvolving real ground-based images, and that the resulting posterior samples recover detail comparable to space-based Hubble imaging. It reports that applying diffusion posterior sampling to Hyper Suprime-Cam observations yields reconstructions whose re-convolved form matches the observed image, with residuals concentrated in background noise, and that sampling many posterior draws and averaging suppresses stochastic fluctuations. The accompanying variance-ratio metric is proposed as a way to see, pixel by pixel, which features are pinned down by the data and which are inherited from the training prior. If this holds, ground-based survey archives could be sharpened to near-space quality with an explicit map of where each reconstruction should and should not be trusted.","feed_headline":"Ground-based galaxy images sharpened to Hubble-level detail","feed_subtitle":"A diffusion model trained on simulations sharpens Subaru survey images and flags which pixels are prior-driven.","key_machinery":"The machinery is Diffusion Posterior Sampling (DPS): a reverse-time diffusion guided by a learned score plus an approximate likelihood term, with the unknown clean image estimated through Tweedie's formula from the noisy latent. The prior is a DDPM trained on idealized TNG galaxies. The forward operator A embeds a redshift-dependent apparent-size scaling and the HSC pixel scale before convolution with the HSC PSF, which makes the tool adaptable to any dataset. A variance-ratio metric compares pixel-wise posterior variance to prior variance over 256 samples; it is the instrument that identifies prior-driven features, and its calibration on simulated galaxies with known signal-to-noise is what connects it to hallucination.","core_discovery":"On the paper's own terms, the central discovery is that diffusion posterior sampling (DPS) with a DDPM prior trained on idealized TNG100 galaxy images can deconvolve HSC observations at moderate redshift to resolutions comparable to HST/ACS images, and that the faithfulness of the reconstruction can be audited. The reconstruction is judged by comparing the forward model A(x_diff0) with the observation y, by residual maps against known TNG truth in simulation, and visually against HST crops. Over 256 posterior samples, the pixel-wise mean is reported to be a robust output while the ratio of posterior variance to prior variance marks regions dominated by the prior; in the galaxy centers this ratio is near zero, while at higher redshifts, where the solver must simultaneously upsample, the ratio grows and hallucinated features appear. The paper also demonstrates, on simulated galaxies of varying magnitude, that the variance ratio rises monotonically as signal-to-noise falls, supporting its use as a hallucination indicator.","pith_inferences":["Extending beyond the paper: the variance-ratio map could be turned into a per-pixel uncertainty mask for downstream scientific catalogs, flagging measurements that should be down-weighted.","A natural test the paper does not run: compare variance-ratio flags against independent space-based imaging (for example JWST) at redshifts the TNG prior does not cover, to see whether low-ratio regions are indeed trustworthy.","The same DPS formulation with redshift- and scale-dependent forward operators could be applied to other inverse problems in astronomy, such as deblending crowded fields or super-resolution of low-surface-brightness features.","The dependence of the metric on the chosen aperture radius suggests it could be calibrated into a scalar quality score per galaxy, enabling selection functions for large statistical samples."],"forward_implications":["If the central claim holds, ground-based survey images of moderately distant galaxies can be deconvolved to resolutions comparable to Hubble without needing space-based observations for every target.","The variance-ratio metric gives a pixel-level trust map, so scientific users can exclude prior-dominated regions before measuring galaxy structure.","Because redshift and pixel scale enter the forward model explicitly, the same trained prior can be applied to other cameras and surveys, not only HSC.","At redshifts above about 0.7, the inverse problem becomes a joint deconvolution and upsampling task, and the paper's own results show that this regime is prone to hallucinated features.","Sampling multiple posterior draws and averaging yields a more reliable reconstruction than any single sample, at roughly 100 seconds per evaluation on a modern GPU."],"supporting_citations":[{"why":"Supplies the Diffusion Posterior Sampling algorithm that guides the reverse diffusion with the likelihood approximation.","marker":"[7]"},{"why":"Provides the score-based SDE formulation and the Tweedie-style estimate of x0 used to approximate the likelihood term.","marker":"[36]"},{"why":"Defines the DDPM framework used for the denoising diffusion prior.","marker":"[33]"},{"why":"Provides the improved DDPM architecture and cosine noise schedule used in training.","marker":"[23]"},{"why":"Provides the TNG-HSC simulated observations used to check that the reconstruction matches the observation and for residual tests.","marker":"[6]"},{"why":"Supplies the HSC PDR3 survey images, PSFs, and photometric redshifts used as the real data targets.","marker":"[3]"},{"why":"Establishes the prior application of diffusion models to astronomical deconvolution that this work extends with redshift and pixel-scale parameters.","marker":"[2]"}],"fun_headline_variants":["Diffusion deconvolution matches HST resolution from ground","Bayesian diffusion model flags prior-driven pixels in deconvolution","Uncertainty metric separates prior-driven artifacts in galaxy deconvolution","Deconvolving galaxy images: diffusion models with confidence","Diffusion-based deconvolution yields HST-like sharpness and uncertainty maps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole method relies on the assumption that idealized, face-on TNG100 simulated galaxies at a few low redshifts are a faithful prior for real HSC galaxies, including those at substantially higher redshift; if that prior is wrong, the reconstructions will be dominated by simulation features and the variance-ratio metric will not reveal whether those features are correct.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion deconvolution matches HST resolution from ground","Bayesian diffusion model flags prior-driven pixels in deconvolution","Uncertainty metric separates prior-driven artifacts in galaxy deconvolution","Deconvolving galaxy images: diffusion models with confidence","Diffusion-based deconvolution yields HST-like sharpness and uncertainty maps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001594,"raw_usage":{"total_tokens":6326,"prompt_tokens":893,"completion_tokens":5433,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":5346}},"tokens_in":509,"tokens_out":5433,"duration_ms":33772,"temperature":1.0,"reasoning_tokens":5346,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:28:03.767953+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a sample of HSC galaxies with independent space-based imaging (for example JWST) at redshifts above 0.7, and compare each deconvolved reconstruction with the space-based truth. If regions flagged by the variance-ratio metric as prior-driven turn out to match the deeper space-based data, or if regions flagged as data-driven turn out to be absent in that data, the metric's trust map is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the score-based SDE formulation and the Tweedie-style estimate of x0 used to approximate the likelihood term."},{"cited_title":"Bottrell, H","cited_arxiv_id":null,"evidence_quote":"Provides the TNG-HSC simulated observations used to check that the reconstruction matches the observation and for residual tests."},{"cited_title":"Aihara, Y","cited_arxiv_id":null,"evidence_quote":"Supplies the HSC PDR3 survey images, PSFs, and photometric redshifts used as the real data targets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the prior application of diffusion models to astronomical deconvolution that this work extends with redshift and pixel-scale parameters."}],"review_version":1}