{"id":"0f5dac5c-064b-4621-a195-79b405357404","arxiv_id":"2607.14206","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A conditional normalizing flow learned from two simulation suites maps HI halo-occupation parameters to EFT bias parameters, producing correlated non-Gaussian priors that are much tighter than flat priors for 21 cm analyses.","lead":"This paper uses simulations and a machine-learning density model to map small-scale hydrogen gas physics in dark matter halos to the large-scale bias parameters used in 21 cm cosmology. It shows these mappings are tight and curved, and that using them as priors could sharpen cosmological constraints from future neutral-hydrogen surveys.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Load-bearing concern: EFT bias parameters defined by unvalidated k→0 extrapolation of transfer functions (Eqs. 12–13) over an unstated k-range, with no error bars; fitting-systematic errors could create the apparent curved manifold and over-tight priors.","rationale":"The paper's central empirical claim is that the HI HOD-to-bias mapping is highly structured (curved, correlated, non-Gaussian), and that propagating 21 cm constraints through the learned flow yields much tighter priors than flat ones. Both statements rest on the 2000 paired samples (θ_HOD, θ_EFT). The weakest link in the chain is the measurement of θ_EFT: bias parameters are obtained from a parametric k→0 extrapolation of transfer functions, with no quoted uncertainties and no specified fit range. The paper itself states (Sec. II.D) that 'the transfer functions do not always exhibit a perfectly flat low-k plateau.' If the polynomial forms in Eqs. (12)–(13) do not exactly match the true scale dependence, the inferred limits are biased, and the bias depends on the HOD because the transfer-function shapes vary with HOD (Fig. 2). Such HOD-dependent systematics can create spurious correlations and curvature in bias–bias plots, which is exactly the evidence for the 'highly structured' manifold. The same pipeline is used for TNG300, so that cross-check cannot reveal common systematic errors; the analytic comparison is not precise enough to exclude them. No error bars are reported, so the reader cannot assess whether the scatter around the b2(b1) loci is physical or due to estimation noise. The CHORD-induced priors inherit any such error, and the claimed 'substantial tightening' over flat priors could be overconfident if the conditional flow learns an artificially narrow manifold. This is a load-bearing concern because it targets the foundation of the framework, rather than a secondary approximation like RSD or the Gaussian Fisher posterior. A decisive check would vary the fitting range and functional form and see if the structure and forecast results are stable. If they are unstable, the paper's central claim would need to be weakened. If stable, the concern is alleviated. Hence I agree with the reader's CONDITIONAL verdict; no verdict change is needed, but the condition—independent validation of the bias extraction—should be made explicit.","tokens_in":34193,"tokens_out":9677,"duration_ms":101262,"concrete_test":"Pick a subset of 50–100 HOD realizations from the HV ensemble at z=3. Re-estimate bias parameters from the same transfer functions using: (i) three different kmax values (e.g., 0.2, 0.4, 0.6 h/Mpc) and (ii) an alternative fitting form without the c1,1 k term in Eq. (12) (e.g., pure even polynomial). For each variant, retrain the conditional flow and recompute the CHORD-induced prior at z=3 (100-day case). If the resulting contours shift by more than ~30–50% of their width, or if the correlations among (b1,b2,bG2,b3) change sign, the k→0 extrapolation is not robust and the claimed structure is not trustworthy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Sec. II.C defines each (b1,b2,bG2,b3) as the k→0 limit of polynomial fits (Eqs. 12–13) to transfer functions measured on an unspecified 'quasi-linear' range. Fig. 2 shows β1(k), βG2(k), and βδ3(k) lacking a flat low-k plateau for representative HODs, yet no error bars are reported and the fitted kmax/kmin are not given. If the fitting forms (e.g., the unusual linear term c1,1 k in Eq. 12) absorb HOD-dependent scale dependence, the extrapolated limits carry HOD-dependent systematics. These systematics would artificially organize the 2000-point ensemble into a narrow curved manifold and create correlations among bias parameters, directly seeding the central claim of a 'highly structured' mapping and the tight non-Gaussian priors in Sec. VI. The TNG300 comparison does not resolve this: both simulations use the identical extraction pipeline, so common systematics would produce similar qualitative structure. The analytic comparison (Fig. 9) is qualitative and already deviates substantially for b3, which is attributed to the truncated basis, but the same logic applies to the other bias parameters. Without validation against an independent bias measurement (e.g., separate-universe responses or a larger-volume simulation with more low-k modes), the learned prior may be precisely the artifact of the extrapolation rather than a property of HI physics.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a framework for constructing simulation-based priors on EFT bias parameters for 21 cm intensity mapping. It paints a four-parameter HOD model onto Hidden Valley N-body halo catalogs, measures (b1, b2, bG2, b3) via field-level shifted-operator transfer functions extrapolated to k→0, trains a conditional normalizing flow p(θ_EFT | θ_HOD), and then uses a CHORD-like Fisher forecast to produce induced EFT priors. The same procedure is repeated with IllustrisTNG300 to assess simulation dependence. The central claim is that the HOD-to-bias mapping is highly structured—curved, correlated, non-Gaussian—and that flow-based priors derived from nonlinear-scale 21 cm measurements are substantially tighter than conventional flat priors, especially at z=3.","tokens_in":34560,"tokens_out":7523,"duration_ms":78540,"significance":"If the structured mapping is physical, the paper provides a timely and useful step toward informative EFT priors for current and future HI surveys. The authors make good use of existing field-level tools, include useful validation checks (mesh-resolution convergence, TNG300-Dark comparison, analytic baseline comparison), and are unusually careful in stating limitations. However, the quantitative claims rest on bias parameters extracted by an extrapolation procedure that is not validated against any independent bias estimator, and the b3 used in the forecasts is an effective coefficient in a truncated basis rather than a unique physical EFT parameter. Both points are load-bearing for the paper's headline conclusions and need to be addressed before the priors can be considered robust.","major_comments":[{"comment":"The bias parameters used as training labels are defined by polynomial fits extrapolated to k→0, but no error bars, fit range, or fit residuals are reported. The text itself notes that β1, βG2, and βδ3 do not always show flat low-k plateaus; the fitting forms, especially the linear term c1,1 k in Eq. (12), could absorb HOD-dependent scale dependence. Because the same extraction pipeline is used for both Hidden Valley and TNG300, common systematics would produce similar qualitative 'structure' regardless of the true physics. The mesh-resolution and TNG300-Dark tests do not break this degeneracy. I request (i) explicit kmin/kmax and residual diagnostics, (ii) a sensitivity test of the inferred bias parameters to the fitting form (e.g., dropping the linear term, adding odd terms, varying the fit range), and (iii) validation against an independent estimator (e.g., separate-universe responses","section":"Sec. II.C, Eqs. (12)–(14), Fig. 2"},{"comment":"The parameter b3 is explicitly defined as the large-scale coefficient of the local cubic operator within a truncated basis, and the paper states that it can absorb contributions from omitted third-order operators. Nevertheless, b3 is then used as an EFT bias parameter in the flow prior and in the CHORD-like Fisher forecast. A full-shape EFT likelihood that uses a complete operator basis would require the physical b3, not this effective coefficient. Please demonstrate that the learned prior on b3 is stable under extension of the operator basis (e.g., by adding the omitted third-order operators for a subset of realizations), or explicitly restrict the use case to the same truncated basis as the bias measurement. The analytic comparison in Fig. 9 already shows a qualitative failure for b3, so this is not a purely hypothetical concern.","section":"Sec. II.B, Sec. VII, Fig. 9"},{"comment":"The abstract claims that the induced priors are 'substantially tighter than conventional flat priors' across z=1–3, but Sec. VI does not contain any quantitative comparison to flat priors. The covariance-volume and error ratios in Fig. 24 are normalized to the 100-day case at the same redshift, not to a flat prior, and the flat-prior parameter ranges are never specified. This is the paper's headline quantitative claim. Please add a direct comparison (e.g., ratio of flow-prior volume to flat-prior volume over the same parameter ranges, or marginalized flat-to-flow width ratios) or revise the claim to something like 'visibly tighter in the examples shown.'","section":"Abstract and Sec. VI, Figs. 21–24"}],"minor_comments":[{"comment":"The sampling range is written as log10(Mmin/h−1M⊙) ∈ [5×10^10, 12], which mixes linear and logarithmic quantities. It should be [log10(5×10^10), 12] or equivalently [10.7, 12].","section":"Eq. (3)"},{"comment":"The text states that TNG300 generally gives 'more negative' values of bG2 than HV, while the Fig. 16 caption says TNG300 is 'less negative' than HV. From Figs. 15 and 25, the latter appears to be correct (TNG300 bG2 is higher/less negative). Please resolve this internal contradiction.","section":"Sec. V.C and Fig. 16 caption"},{"comment":"The fiducial HOD point is written as (M0, Mmin, α, β) = (3×10^10, 2×10^11, 0.8, 0.6) without units for M0 and Mmin. Please specify h−1M⊙.","section":"Eq. (29)"},{"comment":"Reference [43] (and a few other arXiv entries) is missing the publication year; please complete the bibliographic details for consistency.","section":"References"},{"comment":"The cubic polynomial coefficients in the bias–bias fits are quoted without uncertainties or a stated fit range. If these are intended as a practical reference, add a caveat about the limited HOD coverage and the lack of error bars.","section":"Figs. 8, 26, 27"}],"recommendation":"major_revision","confidential_remarks":"The paper is well-written, honest about limitations, and the framework is promising. The main risk is that the 'highly structured manifold' is partly a consequence of the shared bias-extraction pipeline rather than HI physics; this needs to be addressed with independent validation before the central claim is fully credible. The missing flat-prior comparison in Sec. VI is also important for the abstract's headline claim. The internal contradiction about the sign of the TNG300 bG2 offset should be corrected. These are fixable with additional analysis, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is the first HI-specific version of the simulation-based prior (SBP) idea that has already been used for galaxy surveys. The genuinely new content is the conditional normalizing flow mapping HOD parameters to EFT bias parameters for HI, the empirical finding that the high-mass slope α organizes a narrow, curved bias manifold, and a CHORD-like forecast showing that nonlinear-scale 21 cm data could induce substantially tighter EFT priors than flat ones. On the whole it is a well-executed proof of concept, and the authors are unusually candid about its limitations.\n\nThe paper does several things well. The field-level bias pipeline follows Refs. [36,37] and is internally consistent. The TNG300 comparison is honest, and the TNG300-Dark check in Appendix D is a good control showing baryons are not the main driver of simulation differences. The mesh-resolution test in Appendix C is also reassuring. The flow training seems standard, and the one-at-a-time conditional slices give a useful sense of what the mapping looks like. The qualitative claim that the mapping is structured rather than independent Gaussian scatter is well supported by the figures.\n\nThe soft spots are real but proportionate. The bias parameters are point estimates from polynomial fits extrapolated to k→0 with no error bars, and the fit range is not explicitly stated. The stress-test concern that this could artificially tighten the manifold is not resolved by the TNG comparison, since both simulations use the same extraction pipeline. However, the analytic HI-mass-weighted halo comparison in Fig. 9 shows broad consistency for b1 and b2, which gives some independent, albeit qualitative, support. The b3 discrepancy is honestly attributed to the truncated operator basis. I would not call this a load-bearing flaw, but it is the main thing I would ask the authors to address: report the fit range, estimate systematic uncertainties from the fit choice, and ideally validate against a separate-universe or larger-volume measurement.\n\nOther minor concerns: no code is released, which hurts reproducibility; the forecast relies on a Fisher approximation and ad hoc regularizing priors on M0 and Mmin, and the authors say these are optimistic. None of this undercuts the proof-of-concept.\n\nThis paper deserves a serious referee. The central claim is probably correct, and the method will be useful to the 21 cm and LSS communities even if the exact priors change after more careful treatment of the bias-extraction systematics. I would recommend acceptance after a revision that adds error bars or at least a sensitivity test on the k→0 extrapolation, and ideally releases the trained flows and code.","headline":"A credible proof-of-concept for HI simulation-based EFT priors; the central claim of a structured HOD-to-bias mapping is robust, but the bias-extraction step needs more validation before the tight priors are taken at face value.","tokens_in":35020,"tokens_out":1963,"would_cite":true,"duration_ms":22817,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HI bias parameters are not independent: a learned mapping from halo-occupation physics yields tight, non-Gaussian priors for 21 cm cosmology.","keywords":["21 cm intensity mapping","HI bias parameters","effective field theory of large-scale structure","halo occupation distribution","normalizing flows","simulation-based priors","full-shape analysis","nonlinear scales"],"falsifier":"Take a single set of HOD realizations and infer the bias parameters with a finer mesh and a more complete operator basis (e.g., including all independent third-order operators); if the extracted (b1,b2,bG2,b3) move by more than the width of the flow prior, the prior is not robust. Alternatively, increase the simulation volume and check whether the offsets between the two simulation suites shrink; if they persist, the simulation dependence is a genuine property, not finite-volume noise.","tokens_in":34075,"feed_emoji":"📡","tokens_out":4950,"duration_ms":44388,"temperature":0.7,"pith_summary":"The paper tries to establish that the small-scale physics of how neutral hydrogen occupies dark-matter halos controls, in a structured and learnable way, the large-scale bias parameters that enter full-shape analyses of 21 cm intensity maps. Concretely, it claims that the mapping from four halo-occupation parameters to the EFT bias parameters (b1, b2, b3, bG2) is curved, correlated, and non-Gaussian, not a set of independent distributions. By training a conditional normalizing flow on this mapping, the authors turn nonlinear-scale 21 cm power-spectrum forecasts into informative priors on the bias parameters, substantially tighter than flat priors across z=1–3, with the largest gain at z=3. They also show the mapping depends on which simulation supplies the halos, so future priors must account for that dependence. A careful reader would care because broad uninformative priors are currently a major source of conservatism in 21 cm cosmology; informative priors would translate directly into sharper cosmological constraints.","feed_headline":"Learned priors tighten 21 cm cosmology at z=1–3","feed_subtitle":"Simulation-trained flow maps halo-occupation physics to EFT bias parameters, sharpening constraints most at high redshift.","key_machinery":"The central object is the conditional density p(θ_EFT | θ_HOD), the distribution of the EFT bias parameters given the four parameters (M0, Mmin, α, β) of a simplified HI halo-occupation model. The paper estimates it with a conditional normalizing flow—an invertible neural map from a Gaussian base to the target density, conditioned on the HOD parameters—trained by maximum likelihood on paired samples generated from simulations. This flow does the work of encoding the curved, correlated, non-Gaussian support of the bias manifold and providing a sampleable prior; without it the structure visible in the HOD scan could not be turned into a usable prior.","core_discovery":"On the paper's own terms, the central discovery is that the HOD-to-bias relation for neutral hydrogen is not a loose cloud but a thin, curved manifold: across an ensemble of 2,000 HOD realizations per redshift, the four EFT bias parameters occupy a correlated, non-Gaussian region of parameter space, and the high-mass slope α acts as the main organizing direction. When a conditional normalizing flow is trained on paired HOD/bias samples, it reproduces this structure and can be sampled to yield simulation-based priors. Propagating a CHORD-like 21 cm power-spectrum measurement on nonlinear scales through the flow gives priors on (b1,b2,bG2,b3) that are dramatically tighter than flat priors, mos","pith_inferences":["Extending the pipeline to redshift space, with Finger-of-God damping and foreground wedges, would likely change the induced priors; the current real-space forecasts are explicitly optimistic.","The same flow architecture could learn p(θ_EFT | θ_HOD, θ_cosmo), letting cosmology and HI astrophysics vary jointly; this would produce priors directly usable in cosmological parameter estimation rather than fixed-cosmology proof-of-concept.","A multi-simulation training set with controlled variations in volume, resolution, gravity solver, and baryonic physics could turn the observed offsets between the two simulation suites into a quantitative uncertainty model, for example by adding simulation-level hyperparameters to the conditional density.","Because bispectrum and higher-order analyses involve a larger bias-parameter space, they stand to gain even more from informative simulation-based priors; the present b1–b3 focus is the first step."],"forward_implications":["Replacing broad, independent flat priors on HI bias parameters with the flow-based prior significantly tightens EFT-parameter posteriors in full-shape 21 cm analyses; at z=3 the covariance volume shrinks by roughly two orders of magnitude as observing time grows from 100 to 1000 days.","The learned conditional density can be used in two ways: directly as a simulation-based prior, or as a map that converts external or nonlinear-scale constraints on HOD parameters into an induced prior on EFT bias parameters.","The mapping is organized mainly by the high-mass slope α; M0 has almost no effect on the deterministic overdensity field, while Mmin and β mainly broaden the manifold rather than move it.","Because the mapping differs between the two simulation suites, quantitative HI bias priors must treat simulation choice as a systematic; the difference is largest in the tidal and high-bias sectors.","Simple mass-weighted analytic halo-bias formulae capture only the broad ordering, not the detailed bias relations, especially for the effective cubic response b3."],"fun_headline_variants":["Flow-mapped HOD priors sharpen 21 cm constraints","Simulation-based priors beat flat priors for HI bias","Tighter HI bias priors from HOD flows","Learning HOD-to-bias mapping tightens 21 cm forecasts","HI bias priors: from flat to flow-tightened at z=1-3"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the fitted k→0 limits of the transfer functions—obtained by extrapolating smooth polynomial fits over a quasi-linear k range—give unbiased values of (b1,b2,bG2,b3); if the true k-dependence is not captured by those forms, every bias estimate and the prior built from them is systematically wrong.","fun_headline_variants_meta":{"raw":{"variants":["Flow-mapped HOD priors sharpen 21 cm constraints","Simulation-based priors beat flat priors for HI bias","Tighter HI bias priors from HOD flows","Learning HOD-to-bias mapping tightens 21 cm forecasts","HI bias priors: from flat to flow-tightened at z=1-3"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000888,"raw_usage":{"total_tokens":3734,"prompt_tokens":877,"completion_tokens":2857,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":2766}},"tokens_in":621,"tokens_out":2857,"duration_ms":19408,"temperature":1.0,"reasoning_tokens":2766,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T02:46:33.282571+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a single set of HOD realizations and infer the bias parameters with a finer mesh and a more complete operator basis (e.g., including all independent third-order operators); if the extracted (b1,b2,bG2,b3) move by more than the width of the flow prior, the prior is not robust. Alternatively, increase the simulation volume and check whether the offsets between the two simulation suites shrink; if they persist, the simulation dependence is a genuine property, not finite-volume noise.","supporting_citations":[],"review_version":1}