{"id":"a1c27ee2-28e8-4e5f-9398-dab05bd15b2f","arxiv_id":"2411.08111","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A joint galaxy-dust model with two neural score-matching priors recovers spatially resolved stellar and dust morphologies from multi-band images of simulated and three SDSS galaxies.","lead":"Galaxy images are colored and dimmed by dust inside the galaxy, and most survey tools treat that dust as a single uniform screen. This paper introduces a method that uses two neural-network priors to recover separate maps of starlight and dust from ordinary multi-band images, tested on simulations and three SDSS galaxies.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulation tests reuse the same 65 NIHAO-SKIRT galaxies used to train the dust prior, so the reported f<90% recovery may reflect prior memorization rather than generalization; a held-out-galaxy retest is needed.","rationale":"The reader's verdict and weakest-assumption analysis identify exactly this issue, and I agree. The paper has real strengths: a novel conditional score-based dust prior, a non-parametric joint model, honest discussion of the thin-screen degeneracy and age-gradient confusion, and demonstration of stability to S/N on SDSS images. But the simulation experiment is the only quantitative validation of the central claim, and it is contaminated by construction unless a split was performed and not reported. This is not a disagreement with consensus; it is an internal validation gap. A held-out-galaxy retest is feasible because the catalog is public and the training pipeline is described. Until that retest is done, the correct verdict remains CONDITIONAL: the method is promising and internally consistent, but the claim that it 'recovers galaxy host and dust properties over a wide range of attenuation levels and geometries' is not yet supported for geometries outside the training distribution. I therefore keep the reader's verdict unchanged.","tokens_in":11184,"tokens_out":4492,"duration_ms":49147,"concrete_test":"Retrain the dust score model on a strictly disjoint subset of the NIHAO-SKIRT catalog—e.g., 52 galaxies for training and 13 fully held-out galaxies, excluding all viewing angles of those 13—and rerun the §3.1 red-start and blue-start fits on the held-out galaxies. Recompute the residual AV and (D_fit − D_truth)^2 metrics of Figure 4; if the median residual or map error worsens materially outside the thin-screen regime (f<90%), the central claim is not established for unseen geometries. The split must be by galaxy, not by viewing angle, because the 10 viewing angles of a given galaxy share the same physical dust geometry.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—accurate recovery of dust amplitude and morphology for covering fractions f<90%—rests on §3.1, but that test set is not independent of the dust prior. The dust score model in §2.2 is trained on the 65 NIHAO-SKIRT galaxies of Faucher et al. (2023); §3.1 then fits '65 simulated galaxies from Faucher et al. (2023)' and reports results for three orientation angles each. No train/test split is stated anywhere in the paper. Since the conditional prior p(D|S) is learned from exactly these galaxy–dust geometries, the Figure 4 residuals largely measure how well the prior has encoded the test distribution. This matters most in the blue-start and low-S/N regimes, where §2.3 states the model 'relies on our data-driven priors' and where the likelihood is too weak to correct an unrepresentative prior. On real SDSS galaxies, the same prior would then imprint simulation-like dust morphologies and bias the recovered dust maps and dereddened SEDs. The convergence criterion on the host spectrum does not mitigate this, because weakly constrained directions are governed by the prior.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a joint forward model for spatially resolved galaxy and dust decomposition from multi-band optical/NIR images. The host galaxy is described by a non-parametric morphology and a single positive spectrum, while dust attenuation is described by a spatial dust map and a fitted attenuation-curve slope. Two score-based neural priors regularize the inversion: a host morphology prior trained on HSC images and a conditional dust prior p(D|S) trained on NIHAO-SKIRT radiative-transfer simulations. MAP fits are obtained with gradient descent. The method is validated on simulated griz observations of 65 NIHAO-SKIRT galaxies under red-start and blue-start initialization, and applied to three SDSS galaxies. The headline claim is accurate recovery of dust amplitude and morphology for covering fractions f < 90%, with a known thin-screen degeneracy at higher covering fractions.","tokens_in":11382,"tokens_out":6441,"duration_ms":67603,"significance":"If the validation is sound, the method is a meaningful step toward spatially resolved dust attenuation maps and dereddened galaxy SEDs from ordinary multi-band images, which would be valuable for the large imaging surveys now coming online. The paper has several concrete strengths: the conditional dust prior is a principled way to regularize a severely ill-posed inverse problem; the host prior is independently trained on a large sample of real HSC galaxy images; the thin-screen degeneracy is acknowledged explicitly rather than hidden; and the authors point to public catalogs and code for the host prior and the simulation library. The main qualification is that the quantitative claim in Section 3.1 is currently based on a simulated test set whose independence from the dust-prior training set is not established. If that independence is demonstrated, the paper would be a solid methods contribution.","major_comments":[{"comment":"The central simulated validation is not independent of the dust prior. Section 2.2 trains the dust score model on the 65 NIHAO-SKIRT galaxies from Faucher et al. (2023), and Section 3.1 tests on \"65 simulated galaxies from Faucher et al. (2023)\" with three orientation angles each; no train/test split is stated anywhere in the paper. Because the public catalog contains ten viewing angles per galaxy, the reader cannot rule out that the tested orientations are drawn from the same set of images used in training. In that case, the Figure 4 residuals largely measure the dust prior's ability to recall the training distribution rather than its ability to generalize to unseen galaxy-dust geometries. This matters most in the blue-start and low-S/N regimes, where Section 2.3 states that the model \"relies on our data-driven priors\" and where the likelihood is too weak to correct an unrepresentative prior. Please provide a held-out validation: train the dust prior on a subset of the catalog and test on the remaining galaxies (or use leave-one-out cross-validation), and report the same recovery metrics for the held-out set. If a split was already used, state it explicitly and give the details.","section":"§2.2, §3.1"},{"comment":"The forward model as written does not include PSF convolution or a pixel response function, and the simulated validation in Section 3.1 appears to use PSF-free images. Real SDSS images are seeing-limited; even for the large galaxies considered here, the PSF is non-negligible after downsampling to 64×64 pixels, and it will mix attenuated and unattenuated light across pixel boundaries. If scarlet2 includes a PSF that is simply not shown in Eqs. (3)–(5), this should be stated explicitly and the PSF should be included in the simulated validation. If it is not included, the \"realistic simulations\" claim in the abstract is overstated, and the SDSS application needs to be tested with PSF-convolved images before the recovered dust maps can be interpreted at the claimed spatial resolution.","section":"§2.1, Eqs. (3)–(5), §3.2"}],"minor_comments":[{"comment":"The phrase \"bright (mg > 16 mag)\" appears to have the inequality sign reversed; all three galaxies are bright objects with apparent magnitudes below 16.","section":"§3.2"},{"comment":"The text alternates between \"isochromatic\" and \"monochromatic\" galaxies in the same paragraph; please use one term consistently.","section":"§3.1"},{"comment":"The blue-start initialization would benefit from a precise definition of the color used to rank pixels (e.g., g−i) and a description of how the bluest N% of pixels are selected within the galaxy footprint.","section":"§2.3"},{"comment":"No trained weights or code are provided for the dust prior network; a data/code availability statement would improve reproducibility, especially because the host prior and the NIHAO-SKIRT catalog are already public.","section":"§2.2, §5"},{"comment":"The symbol F is used both for the spectral flux density in Eq. (1) and for the vectorized filter-integrated spectrum in Eq. (3); this overloaded notation is confusing and should be disambiguated.","section":"Eq. (1) and Eq. (3)"},{"comment":"The residual panels in Figure 5 are only described qualitatively; reporting the RMS residual with and without the dust component would make the improvement concrete and easier to compare across S/N levels.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the overlap between the dust-prior training set and the simulated test set. This is a fixable issue but it is load-bearing for the paper's central quantitative claim, so I would not accept without a held-out validation. The PSF question is a second issue that needs either a clarification or an additional simulation test. The manuscript is honest about the thin-screen degeneracy, and I do not regard that degeneracy itself as a reason to reject, since the paper explicitly restricts its claims and proposes remedies. If the authors add a held-out test and clarify the PSF treatment, the paper could be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The new thing here is real: a score-based dust prior p(D|S) conditioned on the host morphology, trained on radiative-transfer simulations, and a joint MAP fit of host spectrum, host morphology, dust morphology, and attenuation slope with two neural priors. The blue-start initialization is also a practical contribution: it lets the method work from optical-only data, which is what LSST/Euclid/Roman will actually provide. The paper is honest about the thin-screen degeneracy and the age-gradient confusion, and the SDSS demonstrations, while qualitative, show the method does something sensible on real data.\n\nThe soft spot is the one the stress-test note identifies, and it is real. Section 2.2 trains the dust prior on the 65 NIHAO-SKIRT galaxies of Faucher et al.; Section 3.1 then validates on exactly those 65 galaxies, with three orientations each, and no train/test split is stated. The central claim—accurate dust recovery for f < 90%—therefore largely measures how well the prior has memorized the test distribution. In the blue-start and low-S/N regimes the paper explicitly says it relies on the priors, which is precisely where an overfit prior would do the most damage. This is not a fatal flaw in the method, but it is a load-bearing gap in the evidence. A held-out-galaxy retest (train on 50, test on 15, or leave-one-out) would settle it. Also worth asking for: baseline comparisons against a single-screen fit or a Sersic-dust model, and some form of uncertainty quantification on the MAP dust maps. The lack of those is not disqualifying, but it makes the quantitative claims harder to judge.\n\nOn balance, the paper deserves a serious referee. The method is timely and the conditional dust prior is a genuine step beyond scarlet and Sampson et al. I would send it to review with a clear request for re-validation on a disjoint sample. The core idea is solid; the validation protocol is not yet.","headline":"A genuinely new conditional dust prior and a promising joint modeling scheme, but the headline simulation result is contaminated by training/test overlap and needs a held-out retest.","tokens_in":600,"tokens_out":1684,"would_cite":true,"duration_ms":25492,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A pair of score-based neural priors lets multi-band galaxy images be decomposed into intrinsic starlight and a resolved dust map, recovering host and dust properties for dust covering fractions below about 90 percent.","keywords":["dust attenuation","spatially resolved galaxy morphology","data-driven priors","score matching","multi-band imaging","galaxy SED recovery","neural network priors","radiative transfer simulations"],"falsifier":"Run the method on images of galaxies that have independent spatially resolved attenuation measurements, for example Balmer-decrement maps from integral-field spectroscopy or far-infrared dust maps, and compare the inferred $A_V$ images pixel by pixel; if the residuals correlate with inclination, morphological type, or dust covering fraction in the claimed $f<90\\%$ regime, the simulation-based prior is not transferring.","tokens_in":10929,"feed_emoji":"🌌","tokens_out":8411,"duration_ms":99729,"temperature":0.7,"pith_summary":"The authors are trying to establish that the spatial distribution of dust inside a galaxy and the galaxy's unreddened starlight can both be recovered from a stack of ordinary multi-band optical and near-infrared images, with no integral-field spectroscopy or infrared maps. If true, this would scale resolved dust mapping to the billions of galaxies that the next generation of wide-field surveys will image, where current methods are limited to a few hundred well-observed nearby systems or assume an unrealistic uniform foreground dust screen. The model uses non-parametric stellar and dust morphologies, and two neural-network priors make the highly degenerate inverse problem tractable: one scores the plausibility of the stellar morphology, the other scores the dust morphology conditioned on the current stellar morphology. On simulated galaxies the method recovers dust amplitude and geometry for dust covering fractions below roughly 90 percent, and on three local galaxies it reproduces the prominent dust features and dereddened spectral energy distributions.","feed_headline":"Multi-band images alone can yield resolved galaxy dust maps","feed_subtitle":"Recovers unreddened starlight and dust geometry for galaxies too faint or distant for IFU and infrared follow-up.","key_machinery":"The load-bearing object is the score-matching neural-network prior, a network that outputs the gradient of the log-probability of a morphology, $\\nabla \\log p_{\\mathrm{data}}(x)$, for the host image and, in a conditional U-Net variant, for the dust image given the host image $p(D\\mid S)$. The conditional dust prior is what makes the non-parametric decomposition work: it is flat where the likelihood is flat, strongly penalizes uncorrelated pixel-scale color noise, and encodes the observed tendency for dust to track stellar light, so the optimizer can move dust where it is unconstrained by data without inventing arbitrary geometries. These score gradients are added to the likelihood gradient during gradient-descent optimization, and the fitted quantities are the host spectrum $F$, host morphology $S$, dust morphology $D$, and attenuation slope $\\delta$. A supporting mechanism is the initialization scheme, which starts near zero dust and, when no red filter is available, builds a crude dust map from the bluest pixels so the regularization can refine it into a coherent dust morphology.","core_discovery":"The central claim is that the joint model $Y = (F^{T}S) \\odot 10^{-0.4A}$, where $F$ is a one-dimensional host spectrum, $S$ is a non-parametric monochrome stellar morphology, and $A = K^{T}[-2.5\\log_{10}D]$ is the dust attenuation cube with a spatially uniform slope parameter $\\delta$, becomes identifiable in multi-band images once two score-based priors supply the missing regularization. The stellar prior $p(S)$ suppresses pixel noise and keeps the host morphology on the manifold of real galaxy shapes, while the conditional dust prior $p(D\\mid S)$ encodes the correlation between starlight and dust location learned from radiative-transfer simulations, preventing the likelihood from interpreting every red pixel or noise fluctuation as dust. Maximum-a-posteriori optimization with these priors recovers accurate attenuation maps and unreddened spectra across a wide range of attenuation amplitudes and geometries, with the known exception of the thin-screen limit in which every line of sight is covered and the attenuation level is degenerate with an intrinsically redder spectrum. The same behavior holds when the fit is initialized only from optical bands, where the starting dust map is noisy and the priors are primarily responsible for steering it to a plausible morphology.","pith_inferences":["The accuracy reported on simulations is probably optimistic because the same radiative-transfer catalog both trains the conditional dust prior and serves as the test set without a train/test split; performance on real galaxies outside that distribution is the open question.","A natural extension the authors do not develop is to distill the fitted host models into a fast feed-forward network that maps images directly to dust maps, removing the need for per-object optimization at survey scale.","The conditional prior could be recalibrated on real galaxies by cross-checking inferred $A_V$ maps against Balmer-decrement maps from integral-field surveys, turning the simulation-only prior into a genuinely data-driven one and probing whether dust geometries in nature match those in simulations."],"forward_implications":["Dust attenuation maps and dereddened host SEDs become obtainable for ordinary survey imaging, without requiring IFU observations, far-infrared data, or parametric galaxy models.","Because nonuniform attenuation currently biases stellar-population mass and redshift estimates, recovering the dust map first should reduce those biases for the bulk of survey galaxies.","The documented thin-screen failure defines a concrete development path: adding a stellar-population-synthesis prior on the host spectrum or external tracers such as the Balmer decrement and far-infrared emission should extend the method to fully covered, heavily obscured galaxies.","Survey-scale resolved dust maps would transform studies of dust production, transport, and destruction from a few hundred galaxies to statistically powerful samples."],"supporting_citations":[{"why":"Supplies the 65 simulated galaxies with ground-truth host and dust morphologies that train the conditional dust prior and constitute the simulation test set.","marker":"Faucher et al. 2023"},{"why":"Provides the host morphology score network and the differentiable modeling framework in which the joint fit is implemented.","marker":"Sampson et al. 2024"},{"why":"Defines the score-matching training objective used to build both neural priors.","marker":"Song & Ermon 2020"},{"why":"Supplies the empirical attenuation curve, with Noll et al. and Kriek & Conroy modifications, that parameterizes $A(\\lambda)$ in the model.","marker":"Calzetti et al. 2000"},{"why":"Provides the Sloan Digital Sky Survey observations to which the method is applied for three local galaxies.","marker":"York et al. 2000"},{"why":"Documents the bias in SED-based stellar-population inference caused by nonuniform dust attenuation, the problem the method addresses.","marker":"Hahn & Melchior 2024"},{"why":"The SKIRT radiative-transfer code generates the mock multi-band images with dust used for training and testing.","marker":"Camps & Baes 2020"}],"fun_headline_variants":[],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the dust morphologies appearing in the radiative-transfer simulations used to train the conditional prior are representative of real galaxy-dust geometry, and that the same simulation set can stand in for real galaxies as a test bed.","fun_headline_variants_meta":{"error":"Client error '402 Payment Required' for url 'https://api.deepseek.com/chat/completions'\nFor more information check: https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/402"},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:57:15.209870+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the method on images of galaxies that have independent spatially resolved attenuation measurements, for example Balmer-decrement maps from integral-field spectroscopy or far-infrared dust maps, and compare the inferred $A_V$ images pixel by pixel; if the residuals correlate with inclination, morphological type, or dust covering fraction in the claimed $f<90\\%$ regime, the simulation-based prior is not transferring.","supporting_citations":[{"cited_title":"Score-matching neural networks for improved multi-band source separation","cited_arxiv_id":"2401.07313","evidence_quote":"Provides the host morphology score network and the differentiable modeling framework in which the joint fit is implemented."},{"cited_title":"2020, in Advances in Neural Information Processing Systems, ed","cited_arxiv_id":null,"evidence_quote":"Defines the score-matching training objective used to build both neural priors."},{"cited_title":"Inhomogeneous Dust Biases Photometric Redshifts and Stellar Masses for LSST","cited_arxiv_id":"2409.19054","evidence_quote":"Documents the bias in SED-based stellar-population inference caused by nonuniform dust attenuation, the problem the method addresses."}],"review_version":1}