{"id":"7654740d-e994-4b3a-886c-1ae7162160ae","arxiv_id":"2412.10310","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":10,"one_line_summary":"Using neural ratio estimation on mock HERA power spectra, the authors forecast 95% lower bounds on the thermal WDM mass that exceed the 5.3 keV Lyman-alpha limit when the galaxy threshold mass Mturn is below 1e8 M_sun.","lead":"This paper forecasts how well the HERA radio telescope could measure the warm dark matter particle mass using the cosmic dawn 21cm signal and simulation-based inference on mock observations. It finds that HERA could beat current Lyman-alpha forest limits only if the smallest star-forming galaxies are abundant, with a threshold mass below about 1e8 solar masses.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed HERA sensitivity is carried entirely by the WDM-suppressed halo mass function, but the sharp-k/SMT implementation used is not validated at cosmic dawn in the halo-mass range that drives the forecast.","rationale":"The paper is best read as a conditional sensitivity forecast: given a single galaxy population with a fixed Mturn below ~10^8 Msun, a Park et al. source prescription, the sharp-k WDM HMF, and 1000 h of HERA, the SBI pipeline finds that 95% lower bounds on mWDM can exceed the Lyman-alpha bound. Read this way, the forecast is internally coherent: the training and mock data share the same simulator, coverage tests show conservative posteriors, and the Mturn-mWDM degeneracy is explained and demonstrated. The reader's weakest-assumption statement is appropriate. My stress-test sharpens it: the only physical bridge between mWDM and the 21cm power spectrum is the WDM HMF, and the specific implementation (sharp-k, cSK = 2.5, SMT) is neither released nor directly validated at the redshifts and halo masses that set the quoted bounds. This is distinct from the galaxy-population limitation already flagged in Sec. 5. The concern is not that the paper is internally inconsistent; it is that the forecast's quantitative reach may be controlled by an uncalibrated modeling choice. A rerun with an alternative N-body-calibrated HMF would settle whether the 17/10 keV numbers and the Mturn < 10^8 keV boundary are robust, or whether they should be reported as conditional on the sharp-k HMF choice. Because the conditional verdict already requires such follow-up, my read does not change the reader's verdict.","tokens_in":28937,"tokens_out":16579,"duration_ms":155032,"concrete_test":"Re-run the Sec. 4.3 exclusion pipeline with the sharp-k/SMT WDM HMF replaced by an N-body-calibrated WDM HMF (e.g., Lovell et al. 2014 or Bose et al. 2016), using the same benchmark parameters and the same NRE training budget. Compare the resulting 95% lower bounds on mWDM at Mturn = 10^5 and 10^7 Msun against Figs. 10-11. If either bound shifts by more than ~20% or drops below the 5.3 keV Lyman-alpha line, the headline result is HMF-choice-dependent and must be reported as conditional on that choice.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline sensitivity to mWDM enters solely through the WDM-suppressed halo mass function that weights galactic sources: Eqs. (2.10)-(2.15) combine a sharp-k window with cSK = 2.5, the Sheth-Mo-Tormen first-crossing distribution, and the fitting-function cutoff M_cut. The paper cites low-redshift N-body calibrations (refs. [34, 35, 7]) for this choice, but it shows no comparison of the resulting WDM HMF against N-body simulations at z ~ 10-20 or at the halo masses 10^5-10^8 Msun that dominate the forecast. Because Mturn from Eq. (2.6) and the WDM cutoff scale enter multiplicatively through fduty x dn/dM, a shift in the HMF calibration moves both the 'surpass Lyman-alpha' boundary in Mturn and the quoted 17/10 keV bounds in Figs. 10-11. The Sec. 5 limitation about the single-population galaxy model is acknowledged; the corresponding WDM-HMF calibration uncertainty is not tested. If a WDM N-body-calibrated HMF has a shallower cutoff or a different effective cutoff mass, the forecast sensitivity could be substantially reduced even when Mturn is below 10^8 Msun.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a forecast of HERA's ability to constrain the mass of thermal warm dark matter (mWDM) using the 21cm power spectrum from Cosmic Dawn. The authors modify the public code 21cmFast to use a sharp-k window function for the halo mass function instead of the default top-hat, include a single-population galaxy model with a threshold mass Mturn, and use neural ratio estimation (NRE) as the simulation-based inference method. They generate 18k training simulations with 10 thermal-noise realizations each, validate the inferred posteriors with held-out test sets and coverage tests, and compare training budgets of 1k, 10k, and 18k simulations. After using an MCMC analysis of UV luminosity data to set informative priors, they compute 95% CL lower bounds on mWDM for two benchmark astrophysical models as a function of Mturn. The central claim is that HERA with 1000 hours of observation could surpass the Lyman-alpha forest bound of 5.3 keV if Mturn is below about 10^8 Msun, with lower bounds reaching approximately 17 keV (benchmark 1) or 10 keV (benchmark 2) at Mturn = 10^5 Msun.","tokens_in":29258,"tokens_out":8801,"duration_ms":79981,"significance":"The analysis is carefully executed: the existence of held-out test sets, coverage tests in Appendix B, ten noise realizations per simulated spectrum, and the explicit comparison of reconstruction quality at different training budgets make the statistical pipeline credible. The paper's central result is a conditional forecast and is honest about its dependence on the single-population galaxy model. If the forecast is correct, it would establish HERA as a competitive independent probe of non-cold dark matter, complementary to Lyman-alpha forest constraints. The main value of the paper is the transparent quantification of how the assumed threshold mass for star formation controls the WDM sensitivity, and the demonstration that SBI can be applied to this problem at scale. However, the forecast's dependence on the WDM-suppressed halo mass function is not tested against cosmic-dawn N-body simulations, which is the key unresolved calibration risk.","major_comments":[{"comment":"The WDM signal enters exclusively through the product fduty × dn/dM (Eqs. (2.6) and (2.10)), so the calibration of the sharp-k/SMT halo mass function with cSK = 2.5 is load-bearing for the forecast. The paper cites N-body calibrations from refs. [34, 35, 7] but does not show any comparison of the resulting WDM HMF to simulations at z ~ 10–20 in the mass range 10^5–10^8 Msun that dominates the forecast. Since M_cut in Eq. (2.15) and Mturn enter multiplicatively, a shift in the HMF calibration changes both the threshold Mturn below which HERA surpasses the Lyman-alpha bound and the 17/10 keV bounds in Figs. 10–11. I recommend adding a robustness test, for example varying cSK over its published plausible range or comparing against an N-body-calibrated WDM HMF at high redshift, and reporting how the forecast bounds shift.","section":"Section 2.3, Eqs. (2.10)–(2.15)"},{"comment":"The headline bounds are computed for only two fixed benchmark models of X-ray properties (Table 2), while the UV-luminosity MCMC in Sec. 3.1 leaves LX and E0 only weakly constrained. The statement that HERA could surpass Lyman-alpha constraints if Mturn < 10^8 Msun is therefore not demonstrated across the allowed astrophysical parameter space; the abstract itself notes that X-ray properties may influence the strength of the constraint, but the paper does not quantify how the threshold Mturn < 10^8 Msun shifts when LX, E0, alpha*, or f_star,10 vary within the MCMC-allowed prior. To make the central claim adequately robust, the authors should either show the bound as a function of Mturn for several representative points across the MCMC posterior or explicitly restrict the claim to the two benchmarks rather than phrasing it as a general condition.","section":"Section 4.3, Figs. 10–11"}],"minor_comments":[{"comment":"For z bins 16–19, the k columns are empty; the text states that 15 k-bins are used, but it is unclear whether those high-redshift bins contribute no k modes or whether the table simply omits them. Please clarify in the caption.","section":"Section 3.2, Table 3"},{"comment":"The sentence 'Considering low E0, we make the X-ray spectrum is softer' contains a grammatical error; it should read 'makes the X-ray spectrum softer.'","section":"Section 2.4, last paragraph"},{"comment":"The conclusion acknowledges the simplicity of the single-population galaxy model but does not mention the WDM-HMF calibration uncertainty described in Eq. (2.10)–(2.15); adding a sentence noting that the sharp-k prescription is calibrated at low redshift would help bracket the forecast's systematics.","section":"Section 5, Conclusion"},{"comment":"The coverage plots are informative, but it would help to state explicitly that the test simulations used for coverage are independent of the training set (the division is described in Sec. 3.3, but a reminder in the caption would avoid ambiguity).","section":"Appendix B, Figures 13–14"},{"comment":"The statement that the Planck optical depth measurement 'did not significantly affect the posteriors' is given without supporting evidence; a brief note or a supplementary figure would make this check reproducible.","section":"Section 3.1, last paragraph"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-executed forecast with a clearly stated conditional claim. The main technical concern is the lack of validation of the WDM halo mass function implementation at cosmic dawn, which is fixable with an additional robustness test rather than a full re-analysis. The benchmark-specific nature of the final bounds is disclosed, but showing the dependence across the allowed astrophysical prior would materially strengthen the central claim. Overall, the paper is within the scope of JCAP and suitable for publication after a major revision addressing the HMF calibration and clarifying the scope of the forecast."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, read this if you care about 21cm forecasts and WDM. The real content: they implement a sharp-k WDM halo mass function in 21cmFast, run a clean NRE/SBI pipeline with 18k training simulations, and map HERA's mWDM lower bound against Mturn. That map is new and useful. The headline result—HERA surpassing the 5.3 keV Lyman-alpha bound for Mturn < 1e8 Msun, up to ~17 keV at Mturn = 1e5—is honestly presented as conditional on a single galaxy population.\n\nWhat is done well: the pipeline is careful. Appendix B coverage tests, held-out test sets, ten noise realizations per spectrum, and a training-size convergence check are all there. They compare noisy and noiseless data, and they are explicit about degeneracies and about the model's limitations in Section 5. The comparison with prior SBI WDM work [18] is fair; this paper's sharp-k implementation and HERA focus do differ from that SKA/NPE/top-hat analysis.\n\nSoft spots, in proportion. The forecast's sensitivity to mWDM enters entirely through the WDM HMF at z ~ 10-20 and halo masses 1e5-1e8 Msun, where fduty x dn/dM multiplies the WDM cutoff. The sharp-k/SMT choice with cSK = 2.5 is cited to N-body calibrations, but the paper does not show a validation of the resulting HMF at those redshifts and masses. The stress-test concern lands here: a shallower cutoff or a different effective cutoff mass could shift both the 'surpass Lyman-alpha' boundary in Mturn and the quoted 17/10 keV bounds. This is the load-bearing uncertainty and it deserves a dedicated test or an explicit calibration-marginalized forecast. Second, the transfer-function prefactor alpha_WDM is set so that 5.3 keV saturates the current Lyman-alpha bound; that choice is inherited from [61] and is not circular in the inference, since the mock data contain no fitted WDM, but it does mean the benchmark comparison is partly tied to the very constraint being compared against. Third, the modified 21cmFast and training code are not released, which makes the HMF test harder to do independently.\n\nNone of this breaks the central claim as stated. The Mturn dependence is the real point, and the single-population caveat is acknowledged in the text. Who this is for: forecast readers in 21cm cosmology and WDM phenomenology. It deserves a serious referee, and with the HMF validation added or explicitly framed as a calibration uncertainty it would be a solid JCAP paper. I would send it to review and would cite it as the HERA SBI WDM forecast.","headline":"A careful, conditional forecast: HERA could beat Lyman-alpha on WDM only if cosmic dawn galaxies form in haloes below ~1e8 Msun, and the sharp-k HMF carries the whole sensitivity.","tokens_in":29821,"tokens_out":2701,"would_cite":true,"duration_ms":585921,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A fully built HERA, in 1000 hours of observation, could exclude warm dark matter masses up to about 17 keV, beating the current Lyman-alpha bound of 5.3 keV, provided cosmic-dawn galaxies form in halos below about 10^8 solar masses.","keywords":["warm dark matter","21cm cosmology","HERA","simulation-based inference","halo mass function","cosmic dawn","Lyman-alpha forest"],"falsifier":"The decisive check is to run the same trained neural ratio estimator on a real 1000-hour HERA data set in the same 19 redshift bins and 15 $k$-bins. If the recovered $M_{\\rm turn}$ posterior lies mostly above $10^8\\,M_\\odot$, or the 95% lower bound on $m_{\\rm WDM}$ falls below the 5.3 keV Lyman-$\\alpha$ value, the forecast's central prediction is falsified. A complementary check is to measure the power spectra at $z\\approx6$-$10$; if they show no delay relative to a cold-dark-matter model with $M_{\\rm turn}\\gtrsim10^8\\,M_\\odot$, the WDM imprint the paper relies on is absent.","tokens_in":28740,"feed_emoji":"📡","tokens_out":10011,"duration_ms":88916,"temperature":0.7,"pith_summary":"This paper asks whether the 21-centimeter power spectrum of neutral hydrogen at Cosmic Dawn, as it would be measured by a fully built HERA in 1000 hours, can determine the mass of warm dark matter (WDM) and beat existing probes. The authors use simulation-based inference, training a neural classifier on thousands of simulated power spectra from a semi-numerical cosmic dawn model, and update the model's WDM treatment to a sharp-$k$ halo mass function cutoff that better matches simulations. They find that HERA can surpass the Lyman-$\\alpha$ forest lower bound of 5.3 keV if the threshold mass for star-forming halos is at or below about $10^8\\,M_\\odot$, reaching bounds near 17 keV (or 10 keV for stronger X-ray heating) when $M_{\\rm turn}=10^5\\,M_\\odot$. The conclusion makes the galaxy-population assumption, rather than telescope sensitivity, the main condition for a competitive WDM constraint.","feed_headline":"HERA forecast beats Lyman-alpha WDM bounds if galaxies are small","feed_subtitle":"A 1000-hour observation could exclude warm dark matter up to 17 keV, rivaling Lyman-alpha forest limits","key_machinery":"The load-bearing mechanism is the halo mass function cutoff: WDM free-streaming suppresses the abundance of low-mass halos, and the paper computes this suppression with a sharp-$k$ window function in the variance integral, which reproduces N-body behavior better than the older top-hat filter. The threshold mass $M_{\\rm turn}$ enters through the duty cycle $f_{\\rm duty}=\\exp(-M_{\\rm turn}/M)$, cutting off star formation in the same low-mass halos. The analysis separates these two cutoffs using a neural ratio estimator, a binary classifier trained on matching versus mismatched parameter-spectrum pairs, which approximates the posterior without assuming a Gaussian likelihood. The competition between $M_{\\rm turn}$ and the WDM free-streaming cutoff mass is what determines whether HERA can see the WDM imprint.","core_discovery":"The paper argues that the 21cm power spectrum forecast for HERA contains enough information to place a 95% lower bound on the warm dark matter mass, $m_{\\rm WDM}$, that exceeds the current Lyman-$\\alpha$ forest bound of 5.3 keV whenever the threshold mass for star-forming halos satisfies $M_{\\rm turn}\\lesssim 10^8\\,M_\\odot$. For the most favorable single-population case, $M_{\\rm turn}=10^5\\,M_\\odot$, the forecast excludes $m_{\\rm WDM}\\lesssim 17$ keV with soft X-ray sources and $m_{\\rm WDM}\\lesssim 10$ keV with harder, more luminous X-ray sources. The signal that carries this sensitivity is the delay of the cosmic dawn features, namely Lyman-$\\alpha$ coupling, X-ray heating, and reionization, caused by WDM free-streaming suppressing low-mass halos; the analysis resolves this delay against HERA thermal noise in 19 redshift bins between $z\\approx6$ and $25$ and 15 $k$-bins between $0.15$ and $0.99\\,\\mathrm{Mpc}^{-1}$. The paper also finds positive degeneracies between $m_{\\rm WDM}$ and $M_{\\rm turn}$, $\\alpha_\\star$, and $t_\\star$, so colder dark matter can be mimicked by astrophysics that delays the signal, and these degeneracies, not the noise itself, are the main limiting factor.","pith_inferences":["Read as an effective population parameter, $M_{\\rm turn}$ in a single-population model may not equal the physical threshold for any real galaxy population; if cosmic dawn is a mix of molecular-cooling minihalos and atomic-cooling galaxies, the true duty cycle is a superposition and the WDM sensitivity could shift in either direction.","The same pipeline could be applied directly to the free-streaming scale, or to sterile-neutrino-like models, since it is the cutoff mass that the 21cm signal actually responds to; the WDM mass is only a derived label.","Because thermal noise dominates at $z\\gtrsim12$ in this forecast, most of the constraining power comes from the lower-redshift heating and reionization features; a more sensitive low-frequency array or longer integration would push the WDM reach upward more directly than better astrophysical priors.","If real HERA data later require a second galaxy population, the posteriors shown here should widen, and the 17 keV and 10 keV numbers should be treated as upper limits of what the single-population model can achieve."],"forward_implications":["A 1000-hour HERA observation would give an independent, likelihood-free WDM constraint that can exceed the Lyman-$\\alpha$ bound of 5.3 keV when $M_{\\rm turn}\\lesssim10^8\\,M_\\odot$.","The reach is controlled by $M_{\\rm turn}$: at $10^5\\,M_\\odot$ the forecast excludes $m_{\\rm WDM}$ up to roughly 17 keV (soft X-ray benchmark) or 10 keV (harder benchmark), while at $M_{\\rm turn}\\gtrsim10^8\\,M_\\odot$ the WDM signal is screened and HERA loses its edge.","Because $m_{\\rm WDM}$ correlates positively with $M_{\\rm turn}$, $\\alpha_\\star$, and $t_\\star$, any forecast or eventual measurement must quote the galaxy-population assumptions alongside the WDM bound, or the bound is not interpretable.","Training the neural ratio estimator with 10,000 simulated spectra gives reconstruction quality close to the 18,000-simulation budget, while 1,000 is too few, setting a practical simulation budget for similar forecasts.","Different WDM halo-mass-function treatments (top-hat versus sharp-$k$) change the expected signal, so comparing forecasts requires knowing which filter was used."],"supporting_citations":[{"why":"The semi-numerical 21cm simulation code whose WDM implementation the paper modifies.","marker":"[21]"},{"why":"The maintained Python/C version of that simulator used to generate lightcones and power spectra.","marker":"[22]"},{"why":"Provides the galaxy star-formation and X-ray parameterization that defines the astrophysics model.","marker":"[41]"},{"why":"Basis for replacing the top-hat filter with the sharp-k window in the WDM halo mass function.","marker":"[7]"},{"why":"Sets the Lyman-alpha forest 95% lower bound of 5.3 keV that the HERA forecast must surpass.","marker":"[37]"},{"why":"Supplies the HERA array sensitivity calculation used for thermal noise.","marker":"[42]"},{"why":"Extends the sensitivity treatment to power spectrum measurements and noise.","marker":"[43]"},{"why":"UV luminosity function data used to build the MCMC priors for the astrophysics parameters.","marker":"[68]"},{"why":"Provides the neural ratio estimation implementation used to approximate the posterior.","marker":"[81]"}],"fun_headline_variants":["HERA forecast excludes warm dark matter up to 17 keV","Small galaxies let HERA beat Lyman-alpha WDM bounds","HERA 21cm power spectrum could rule out 17 keV WDM","Cosmic dawn 21cm signal may rival Lyman-alpha for WDM","HERA's WDM sensitivity hinges on galaxy threshold mass"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forecast rests on the simulated 21cm power spectra being a faithful stand-in for cosmic dawn under the single-population galaxy model; if the true threshold mass lies above 1e8 solar masses, or feedback or a second galaxy population screens the WDM suppression, the claimed HERA sensitivity disappears.","fun_headline_variants_meta":{"raw":{"variants":["HERA forecast excludes warm dark matter up to 17 keV","Small galaxies let HERA beat Lyman-alpha WDM bounds","HERA 21cm power spectrum could rule out 17 keV WDM","Cosmic dawn 21cm signal may rival Lyman-alpha for WDM","HERA's WDM sensitivity hinges on galaxy threshold mass"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000249,"raw_usage":{"total_tokens":1620,"prompt_tokens":1085,"completion_tokens":535,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":701,"completion_tokens_details":{"reasoning_tokens":444}},"tokens_in":701,"tokens_out":535,"duration_ms":5231,"temperature":1.0,"reasoning_tokens":444,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:58:26.371962+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The decisive check is to run the same trained neural ratio estimator on a real 1000-hour HERA data set in the same 19 redshift bins and 15 $k$-bins. If the recovered $M_{\\rm turn}$ posterior lies mostly above $10^8\\,M_\\odot$, or the 95% lower bound on $m_{\\rm WDM}$ falls below the 5.3 keV Lyman-$\\alpha$ value, the forecast's central prediction is falsified. A complementary check is to measure the power spectra at $z\\approx6$-$10$; if they show no delay relative to a cold-dark-matter model with $M_{\\rm turn}\\gtrsim10^8\\,M_\\odot$, the WDM imprint the paper relies on is absent.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the neural ratio estimation implementation used to approximate the posterior."}],"review_version":1}