{"id":"8acc6bb0-8806-4b42-84da-97fd76b70a6e","arxiv_id":"2411.18399","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A proof-of-concept method de-baryonifies halos by sampling gravity-only maps at fixed optimal transport cost from the full-physics map, recovering the correct convergence power spectrum suppression in IllustrisTNG.","lead":"This paper proposes de-baryonifying simulated halos by using optimal transport to remove baryonic feedback, then shows that a trained generative model plus a fixed transport cost can recover the correct power spectrum suppression. It is a proof-of-concept for a field-level method to handle baryonic feedback in weak lensing analyses.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The de-baryonification result is currently demonstrated only with an oracle transport cost T; the paper does not show that T can be estimated from the noisy feedback proxies of Fig. 1 to the 0.01 dex precision used in Eq. (5), so the central claim is not yet a usable forward model.","rationale":"The reader's weakest assumption identifies the same load-bearing concern, and I agree. The paper's internal logic is sound: the prior is learned from independent gravity-only simulations, the squared-Euclidean versus sqrt-cost ablation shows the transport cost carries non-trivial information, and the code and training procedure are described in enough detail to be reproducible. The reason this is the most load-bearing concern is that the central claim is worded generically in the abstract, while the actual test uses an oracle value of T that would be unavailable in a lensing analysis. The missing piece is the calibration and error model connecting observable feedback proxies to T, which the paper explicitly defers to future work. A practical check is to replace true T with T_pred drawn from the Fig. 1 scatter and remeasure the suppression; this would settle whether the method survives outside the oracle setup. Other weaknesses, such as the maximum-posterior versus posterior-mean choice and the one-halo approximation, are acknowledged explicitly and are secondary because they are scoped within the proof-of-concept. The verdict should remain conditional, unchanged from the reader's assessment.","tokens_in":12524,"tokens_out":7493,"duration_ms":76746,"concrete_test":"Build a regression for log10 T from the Fig. 1 proxies (for example, fbar or Delta Y) with a Gaussian scatter estimated from the plotted points, then for the 3926 test halos draw T_pred from that conditional distribution and rerun the Hamiltonian Monte Carlo and maximum-posterior selection with Eq. (5) centered on T_pred instead of the true T. Compare the resulting one-halo Ch/Cd - 1 to ground truth and to the oracle-T curve. If the recovered suppression moves outside the scatter band of Fig. 3 at the observed proxy scatter, the method fails outside the oracle setting; if it survives, the missing calibration is less serious.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, in the abstract and Section III, is that the maximum-posterior de-baryonified halos reproduce the baryonic suppression of the convergence power spectrum. As Section II D states, this is achieved by conditioning the likelihood Eq. (5) on the true optimal transport cost T, evaluated on the matched gravity-only halo. In a real weak-lensing application T is not observed; it must come from the empirical correlation in Fig. 1 between T and feedback proxies such as Delta Y and fbar. That correlation is visibly noisy, and no scatter model or calibration is given, while the likelihood uses sigma_logT = 0.01 dex. The paper itself flags in Section IV that the T-to-feedback relation needs future work and in Section II C that the cost-matrix choice may not generalize. The result is therefore a conditional statement: if T is known to percent-level precision, the one-halo suppression is recovered. It is not yet a forward model from observable feedback indicators. This is a missing demonstration rather than an internal inconsistency; the normalizing-flow prior is trained on independent gravity-only simulations, so the core is not circular.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a field-level 'de-baryonification' method based on optimal transport: given a projected full-physics halo map, it samples gravity-only maps from a normalizing-flow prior subject to a likelihood that fixes the optimal transport cost between the full-physics and gravity-only maps. The transport cost is argued to correlate with baryonic feedback strength (Fig. 1). The method is applied to individual halos from IllustrisTNG-300, using a normalizing flow pre-trained on several gravity-only simulations and fine-tuned on miniUchuu. Hamiltonian Monte Carlo sampling of the likelihood (Eq. 5) yields posterior samples; the maximum-posterior samples are then used to compute the one-halo convergence power spectrum suppression. The paper reports that these MAP de-baryonified halos reproduce the true suppression Ch/Cd - 1 in Fig. 3, while the posterior average does not. The authors are careful to frame the work as a proof of concept and to list several open issues, including the need to connect transport cost to observable feedback proxies and the cost-matrix dependence.","tokens_in":12769,"tokens_out":2544,"duration_ms":26117,"significance":"If the result holds, it is a novel and interesting proof-of-concept: it demonstrates that a macroscopic, simulation-based prior plus an optimal-transport constraint can recover a nontrivial field-level summary of baryonic feedback, complementing existing baryonification approaches. The design has genuine strengths: the prior is trained on independent gravity-only simulations, so the central demonstration is not circular; the use of HMC with a normalizing-flow likelihood is technically sound; and the paper is admirably honest about its limitations, including the oracle nature of the transport cost and the heuristic likelihood. However, the current evidence for the central claim is conditional on knowing the true optimal transport cost T, and the absence of uncertainty estimates on Fig. 3 leaves the strength of the claim somewhat open. The significance is therefore substantial but presently partial: the method is demonstrated as an oracle-conditioned inversion, not yet as a usable forward model from observable feedback indicators.","major_comments":[{"comment":"The central result is obtained by conditioning on the true optimal transport cost T, evaluated on the matched gravity-only halo from the simulation. In a real weak-lensing application T is not observed; as the paper itself states in Sec. IV, the relation between optimal transport cost and observable feedback strength (e.g., tSZ Y or baryon fraction) is only an empirical correlation with visible scatter and no calibration. Thus, the claim that 'the set of de-baryonified halos reproduces the correct convergence power spectrum suppression' is currently demonstrated only for an oracle input. To make the forward-model claim load-bearing, the paper should either demonstrate end-to-end de-baryonification with T estimated from the proxies of Fig. 1 (including a scatter model), or explicitly re-frame the result as a conditional proof of concept with a quantitative sensitivity analysis to errors in T.","section":"Sec. II D and Sec. III (Eq. 5, Fig. 3)"},{"comment":"Fig. 3 shows the MAP suppression recovering the ground truth, but no error bars or uncertainty bands are provided. Given that the posterior average fails, it is important to establish that the MAP result is not an artifact of HMC noise, chain non-convergence, or the particular set of 3926 halos. At minimum, the authors should report uncertainties from multiple chains, bootstrap resampling of halos, or the posterior spread of the suppression; without this, the reader cannot judge whether the agreement is statistically significant or fortuitous.","section":"Sec. III (Fig. 3)"},{"comment":"The likelihood is a heuristic construction: a log-normal term in |T_hd| with sigma_logT = 0.01 dex, plus an ad hoc power-law volume correction with alpha = 10. The paper states that the results are not very sensitive to alpha, but no evidence is shown for this claim. In addition, the cost matrix choice (squared Euclidean versus square-rooted) changes the prediction qualitatively, and the paper notes it is unclear whether the chosen cost matrix generalizes to other feedback implementations. Since these choices are made after seeing the IllustrisTNG test set, the key result lacks a validation on an independent hydrodynamical simulation or a systematic sensitivity analysis. This is not a fatal flaw for a proof of concept, but it needs to be addressed or explicitly scoped to make the central claim robust.","section":"Sec. II C and II D (Eq. 4, Eq. 5)"}],"minor_comments":[{"comment":"The word 'dimesional' should be 'dimensional' in the discussion of the tSNE visualizations.","section":"Sec. II B"},{"comment":"There are several typos: 'oberved' should be 'observed', and 'astrohpysical' should be 'astrophysical'.","section":"Sec. IV"},{"comment":"The sentence 'This partial degeneracy may be responsible for the apparent mismatch in measurements of the clustering amplitude S8 between the cosmicmicrowavebackgroundandweakgravitational lensing' has missing spaces and should be reworded.","section":"Introduction"},{"comment":"The horizontal axis is labeled 'angular wavenumber', but the text refers to the convergence power spectrum; please clarify whether the axis is multipole ell or wavenumber k, and consistently use the corresponding notation.","section":"Fig. 3"},{"comment":"The sentence 'Thus, the presented methodology does indeed perform something non-trivial' would be clearer as 'Thus, the presented methodology does indeed perform something non-trivial, in that the result depends on the cost metric.'","section":"Sec. II C"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick private take. The genuinely new piece here is the pairing: optimal transport cost as a summary of baryonic feedback, plus a normalizing-flow prior over gravity-only halos, used backwards to de-baryonify full-physics halos. I don't see that combination in the baryonification literature or in the earlier optimal-transport galaxy-clustering papers, and the author is properly modest that this is a proof of concept. The training setup is the strongest part: the prior p(xd|Mvir) is learned on independent gravity-only simulations (miniUchuu, Quijote, MDPL, Uchuu), so the central loop is not circular. The citation pattern is appropriate—baryonification, OT clustering, feedback probes, and the simulation suites are all there.\n\nThe main soft spot is the oracle transport cost, and the stress-test note is right about it. Eq. (5) conditions on the true T from the simulated gravity-only halo, with sigma_logT = 0.01 dex. The paper shows that if you hand the method T at that precision, the MAP de-baryonified halos recover the one-halo suppression. It does not show T can be estimated from the noisy proxies in Fig. 1 (Delta Y, fbar) anywhere close to that precision. The author says this in Section IV, so it is an acknowledged missing demonstration, not an internal contradiction. That makes the headline claim conditional: given a faithful T, field-level de-baryonification works. A legitimate proof of concept, but not yet a forward model.\n\nTwo smaller issues. First, the cost-matrix choice is post hoc: squared Euclidean works, square-rooted doesn't, and no independent argument is given for why the chosen metric generalizes. The author admits this. Second, Fig. 3 has no error bars, and the claim that the noise matches between ground truth and MAP is hard to evaluate without a bootstrap or split. The failure of the posterior average is honest, but it means the method depends on a MAP selection that may not be stable in higher-dimensional settings.\n\nOverall the central argument holds up as a conditional proof of concept. It is clearly written, reproducible from public simulations, and worth a serious referee. The next version needs to close the T-to-proxy gap and add uncertainty quantification. I'd cite it if I work in this area and would bring it to reading group; with the oracle caveat explicit and one robustness test on T estimation, it would be a solid publication.","headline":"A sound, clearly written proof-of-concept for optimal-transport de-baryonification, but the headline power-spectrum recovery relies on an oracle transport cost and needs the T-to-proxy gap closed before it is a usable forward model.","tokens_in":13287,"tokens_out":2943,"would_cite":true,"duration_ms":27338,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["98.80.-k","98.62.Sb"],"model":"deepseek-v4-flash","headline":"The paper proposes a field-level de-baryonification method that selects the maximum-likelihood gravity-only halo on a fixed optimal-transport-cost hypersurface and shows it recovers the true convergence power-spectrum suppression in…","keywords":["baryonic feedback","optimal transport","weak gravitational lensing","de-baryonification","normalizing flow","convergence power spectrum","IllustrisTNG","field-level inference"],"falsifier":"Run the de-baryonification pipeline with the transport cost replaced by a value predicted from the thermal Sunyaev-Zel'dovich or baryon-fraction correlation plotted in Fig. 1, and check whether the resulting set of halos still recovers the true convergence power-spectrum suppression; if it does not, the method is not usable as a forward model. A second test: repeat the whole procedure on a hydrodynamic simulation with a different feedback implementation—if the squared-Euclidean cost matrix and the empirical cost–feedback correlation fail to reproduce that simulation's suppression, the method does not generalize beyond IllustrisTNG.","tokens_in":12294,"feed_emoji":"🔭","tokens_out":7525,"duration_ms":61226,"temperature":0.7,"pith_summary":"Baryonic feedback—gas being blown out of halos by supermassive black holes and supernovae—changes the projected matter density in a way that mimics weaker cosmological clustering in weak-lensing surveys. This paper proposes that the right way to remove that effect is optimal transport: instead of fitting nuisance parameters, find the gravity-only mass map that costs the least (in a transport sense) to reach from the observed full-physics map, subject to a fixed total cost. As a proof of concept on IllustrisTNG halos, the author samples gravity-only halos from a normalizing flow trained on dark-matter-only simulations, constrained to lie on the surface of fixed optimal transport cost around each full-physics halo. The maximum-posterior samples of this constrained distribution reproduce the true suppression of the convergence power spectrum; the posterior average does not. If this generalizes to full convergence maps, it would give weak-lensing analyses a field-level way to fold in baryonic feedback without assuming a specific feedback model.","feed_headline":"De-baryonified halos reproduce lensing suppression","feed_subtitle":"Maximum-posterior transport maps recover the one-halo power spectrum; posterior averages do not","key_machinery":"The central object is the entropy-regularized optimal transport plan $T_{hd}$ between the full-physics projected halo and a candidate gravity-only halo, computed with the Sinkhorn algorithm using the squared Euclidean cost matrix $M_{ab} = \\tfrac{1}{2}\\|r_a - r_b\\|^2$. The likelihood for a candidate is Eq. (5): the learned gravity-only density $p(x_d|M_{\\mathrm{vir}})$ times a narrow Gaussian in $\\log(|T_{hd}|/T)$ with $\\sigma_{\\log T} = 0.01\\,\\mathrm{dex}$ times a power law $|T_{hd}|^{-\\alpha}$ with $\\alpha = 10$ that counteracts the growing volume of transport plans. A masked autoregressive normalizing flow learns the gravity-only halo distribution from dark-matter-only simulations, pre-trained on a range of simulation boxes and fine-tuned on a run matched to the IllustrisTNG resolution, and Hamiltonian Monte Carlo samples the posterior; the maximum-likelihood sample is the de-baryonified halo.","core_discovery":"The central claim is that baryonic feedback can be undone at the field level without a detailed feedback model. For each full-physics halo $x_h$, the author defines the de-baryonified halo as the point of maximum posterior under the learned gravity-only distribution $p(x_d|M_{\\mathrm{vir}})$ on the hypersurface where the entropy-regularized optimal transport cost between $x_h$ and $x_d$ equals the true cost $T$ (Eq. 5). Across $3 \\times 3926$ halos from IllustrisTNG-300 at $z=0$, the maximum-posterior samples of this constrained distribution reproduce the true one-halo suppression $C_h/C_d - 1$ of the convergence power spectrum, while the posterior average does not. The paper interprets the match as evidence that the fixed-transport-cost slice through the gravity-only posterior is highly informative, and that individual-halo scatter is large because many gravity-only configurations share the same transport cost.","pith_inferences":["If the empirical cost–feedback correlation survives across different feedback models, de-baryonification could replace baryonification as the default field-level feedback model, since it starts from maximum ignorance and adds only the transport cost as an external input.","The sensitivity of the result to the cost matrix (squared versus square-rooted Euclidean) suggests the method is implicitly choosing a metric; finding a physically motivated metric, perhaps tied to gravitational potential energy, could make the method transferable to other feedback implementations.","The normalizing flow's poor out-of-distribution behavior, which the paper observes but says does not affect these results, is likely to matter more when applying the method to full convergence maps where the target lives far from the training distribution; energy-based models or diffusion alternatives may be needed.","A fully Bayesian de-baryonification without a fixed $T$ would require computing the volume element of the transport-cost hypersurface; the power-law approximation in Eq. (5) is a placeholder, and deriving that volume term would remove the need for an external feedback proxy."],"forward_implications":["If the method holds at full-map level, weak lensing field-level analyses can account for baryonic feedback by conditioning on an optimal transport cost instead of adding nuisance parameters to a baryonification model.","The posterior mean is not a valid de-baryonified map; only the maximum-posterior point recovers the power spectrum, so any field-level use of this approach must preserve the MAP solution.","Because the transport cost correlates with both thermal Sunyaev-Zel'dovich Y-deviation and baryon fraction, astrophysical feedback measurements could in principle provide the cost estimate needed to run the method on real data without knowing the gravity-only truth.","Individual halo de-baryonification is multimodal—there are many plausible gravity-only configurations at one transport cost—so aggregate statistics, not single-halo matching, are the right target for validation."],"supporting_citations":[{"why":"defines the baryonification concept that this work contrasts with and tries to replace","marker":"[24-28]"},{"why":"establishes the empirical link between retained gas fraction and power-spectrum suppression that motivates using feedback proxies","marker":"[20, 21]"},{"why":"supplies the halo finder that defines centers, masses, and radii used to extract the projected halo maps","marker":"[43]"},{"why":"provides the isometric log-ratio transform that maps the simplex-valued mass maps to real space for the density estimator","marker":"[44]"},{"why":"provides the IllustrisTNG simulation used as the full-physics test set and the matched gravity-only ground truth","marker":"[45-50]"},{"why":"supplies gravity-only simulations, especially miniUchuu, used to match TNG resolution and to pre-train and fine-tune the normalizing flow","marker":"[51-55]"},{"why":"gives the masked autoregressive flow architecture used to learn the gravity-only halo distribution","marker":"[58, 59]"},{"why":"provides the Sinkhorn algorithm that makes entropy-regularized optimal transport GPU-batchable","marker":"[63]"},{"why":"supplies the simulated-annealing Sinkhorn solver with efficient gradients used to evaluate the transport cost and its gradient in the likelihood","marker":"[64]"},{"why":"gives the one-halo term expression used to compute the convergence power spectrum suppression from the halo profiles","marker":"[65]"}],"fun_headline_variants":["Optimal transport reverses baryonic feedback in halos","Field-level de-baryonification recovers lensing signal","Maximum-posterior transport fixes halo baryon effect","Transport cost undoes baryonic feedback on halos","Posterior-constrained halos restore convergence power"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole pipeline succeeds only when the optimal transport cost between a full-physics halo and its gravity-only counterpart is taken from the true simulated gravity-only halo; in a real analysis that cost would have to be predicted from observable feedback proxies, and the paper does not show that prediction works.","fun_headline_variants_meta":{"raw":{"variants":["Optimal transport reverses baryonic feedback in halos","Field-level de-baryonification recovers lensing signal","Maximum-posterior transport fixes halo baryon effect","Transport cost undoes baryonic feedback on halos","Posterior-constrained halos restore convergence power"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000167,"raw_usage":{"total_tokens":1259,"prompt_tokens":952,"completion_tokens":307,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":568,"completion_tokens_details":{"reasoning_tokens":240}},"tokens_in":568,"tokens_out":307,"duration_ms":3643,"temperature":1.0,"reasoning_tokens":240,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:14:19.531130+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the de-baryonification pipeline with the transport cost replaced by a value predicted from the thermal Sunyaev-Zel'dovich or baryon-fraction correlation plotted in Fig. 1, and check whether the resulting set of halos still recovers the true convergence power-spectrum suppression; if it does not, the method is not usable as a forward model. A second test: repeat the whole procedure on a hydrodynamic simulation with a different feedback implementation—if the squared-Euclidean cost matrix and the empirical cost–feedback correlation fail to reproduce that simulation's suppression, the method does not generalize beyond IllustrisTNG.","supporting_citations":[{"cited_title":"Van der Maaten and G","cited_arxiv_id":null,"evidence_quote":"provides the Sinkhorn algorithm that makes entropy-regularized optimal transport GPU-batchable"},{"cited_title":"Nalisnick, A","cited_arxiv_id":null,"evidence_quote":"supplies the simulated-annealing Sinkhorn solver with efficient gradients used to evaluate the transport cost and its gradient in the likelihood"},{"cited_title":"Du and I","cited_arxiv_id":null,"evidence_quote":"gives the one-halo term expression used to compute the convergence power spectrum suppression from the halo profiles"}],"review_version":1}