{"id":"0cf05a2f-0380-4f7b-99b0-5532a81232a8","arxiv_id":"2412.06879","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Forward-modeled 21-cm forest spectra show that a 1D power spectrum detection is feasible with 500 hr of uGMRT or 50 hr of SKA1-low time for a 25% neutral, cold IGM at z=6, and that null detections can constrain IGM heating and ionization.","lead":"This paper simulates how the 21-cm forest signal from neutral hydrogen during reionization would appear in spectra of distant radio-loud quasars observed with current and upcoming radio telescopes. It finds that a statistical detection is possible with long uGMRT or short SKA1-low observations if the intergalactic medium is cold and partially neutral, and that a non-detection could still set useful constraints.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The detection claim for P21 compares median signal to the noise floor, but the paper's own covariance analysis shows sample variance dominates; no detection S/N for P21 is ever computed.","rationale":"The reader identifies the finite 50 cMpc simulation box as the weakest assumption. I agree that this is a real limitation and is explicitly acknowledged in Sec. 2.1: the simulations do not capture HI islands larger than the box. However, the binned analysis uses k > 0.316 MHz^-1, corresponding to scales below about 45 cMpc, so missing islands larger than 50 cMpc may have limited direct impact on the quoted k range. The more immediate internal gap is that the detection claim rests on a noise-floor comparison while the paper's own covariance analysis says sample variance dominates. This is not a modeling disagreement; it is a missing statistic in the paper's argument. The concrete test requires no new physics, only a calculation from existing mock ensembles. If the S/N including sample variance is adequate, the verdict should remain conditional; if not, the central claim should be weakened. I would keep the reader's CONDITIONAL verdict, but for a different reason: the paper needs to demonstrate detection significance for P21, not just a favorable ratio of median signal to noise floor.","tokens_in":26951,"tokens_out":13746,"duration_ms":162735,"concrete_test":"Using the released code, compute the detection S/N for P21 in the fiducial model (log10 fX = -2, xHI = 0.25) for uGMRT 500 hr and SKA1-low 50 hr, for both 1 and 10 sightlines. Define the detection statistic as S/N = sqrt(d^T C^{-1} d) over k < 8.5 MHz^-1 (and k < 32.4 MHz^-1 for SKA1-low), where d is the binned P21 of a mock observation minus the mean noise-only band powers and C is the total (sample variance + noise) covariance used in Eq. 6 / Fig. 8. Generate 10^4 signal+noise and 10^4 noise-only realizations, and report the fraction with S/N > 3 and S/N > 5 along with the false-alarm fraction. If the uGMRT 500 hr detection fraction is below about 50%, the abstract's detectability claim is not supported by the present analysis.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central detection claim in Sec. 3.2 and the abstract is based on the median P21 lying above the noise-only power (Fig. 6 and Sec. 5). For a band-power measurement, however, the relevant uncertainty is the total variance, and Sec. 4.1 states that sample variance is 'orders of magnitude stronger' than telescope noise and dominates the inference (Fig. 8, Eq. 6). Comparing the median signal with the noise floor is necessary but not sufficient: it does not establish that a single 200 cMpc spectrum, or even the average of 10 such spectra, can be distinguished from noise. The paper reports no detection significance or detection probability for P21 itself; the KS-test detection fractions in Sec. 3.1 and the per-channel SNR in Appendix C probe different statistics and do not support the P21 claim. The Bayesian recovery in Sec. 4.2 starts from mock observations that already contain the signal and demonstrates parameter inference, not detection. Thus the headline statement that it is 'possible to detect the 1D power spectrum ... k < 8.5 MHz^-1' is under-supported as written.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper forward-models 21-cm forest spectra at z=6 using 21cmfast semi-numerical simulations, adds telescope noise and spectral resolution for uGMRT and SKA1-low, and studies two statistical observables: the differential number density of the transmitted flux and the 1D power spectrum P21. The authors claim that a statistical detection of P21 is possible with 500 hr of uGMRT or 50 hr of SKA1-low time under a late-end reionization model with <xHI> = 0.25 and spin temperature < 30 K, and that a null detection of 10 such spectra would constrain the neutral fraction and X-ray heating of the IGM more tightly than current Ly-alpha forest and 21-cm tomography limits. The paper also presents Bayesian MCMC forecasts for joint constraints on log10 fX and <xHI>.","tokens_in":27119,"tokens_out":6054,"duration_ms":66319,"significance":"If the detection claim holds, the paper would provide a timely and observationally actionable forecast: the 21-cm forest 1D power spectrum could be probed with existing or near-future telescopes at z=6, and a null detection would place competitive limits on a cold, neutral IGM. The forward-modelling is careful in several respects: the code is public, a resolution convergence test is included in Appendix A, the covariance matrix is estimated from 10^4 mock realizations, and the authors explicitly discuss omitted effects such as minihaloes, RFI, mode-mixing, and redshift evolution. The main shortfall is that the headline detectability statement is not backed by a detection significance computed from the total variance, and the null-detection criterion is not defined at the data level.","major_comments":[{"comment":"The detectability claim for P21 is based on comparing the median signal power with the noise-only power in Fig. 6, but §4.1 and Fig. 8 show that sample variance is orders of magnitude stronger than telescope noise and dominates the covariance. For a band-power measurement, the relevant uncertainty is the total variance, not the noise floor. The paper reports no detection S/N, no detection probability, and no false-detection rate for P21 itself: the KS-test fractions in §3.1 and the per-channel SNR in Appendix C concern different statistics, and the MCMC recovery in §4.2 starts from mock observations that already contain the signal. The abstract's statement that it is 'possible to detect the 1D power spectrum ... k < 8.5 MHz^{-1}' is therefore under-supported as written. Please compute a detection significance using the full covariance (e.g. S/N per k-bin or a detection fraction over mock realizations) or soften the claim accordingly.","section":"§3.2 and abstract"},{"comment":"The simulation volume is (50 cMpc)^3, and §2.1 explicitly states that coherent HI islands on scales larger than the box are not captured. This is a direct limitation for the low-k part of P21, where the claimed detection is strongest: a 50 cMpc spectrum corresponds to a fundamental mode around k ≈ 1 MHz^{-1}, so bins below this scale are influenced by modes outside the box. Late-end reionization models, which motivate the paper, predict large neutral islands. Please quantify how much of the low-k signal in Fig. 6 could come from scales larger than the box, for example by comparing with a larger-volume simulation at lower resolution or by estimating the contribution of modes k < 2π/L_box. Without this, both the detection forecast and the null-detection constraints in Figs. 10, 12, and 13 are uncertain.","section":"§2.1 and Figs. 6, 10"},{"comment":"The null-detection criterion is defined as the model power <P21_sim>10 falling below the noise-only power in at least one bin below 8.5 MHz^{-1}, and the 32% threshold is used to draw the solid black curves in Figs. 10, 12, and 13. This is not a data-level detection criterion: it does not include the sample-variance fluctuations of an actual mock observation, and it does not correspond to a likelihood-based or hypothesis-test decision. Please define a null detection operationally (for example, the fraction of mock observations for which the posterior excludes a nonzero signal, or a detection S/N threshold), and calibrate that criterion on mock realizations. The current definition makes the claimed null-detection constraints difficult to interpret.","section":"§4.2 (null-detection definition, Fig. 10)"},{"comment":"The multivariate Gaussian likelihood is used for all parameter inference, and the conclusions themselves note that the 21-cm forest is expected to be non-Gaussian (citing Wolfson et al. 2023). Because sample variance dominates the covariance matrix, the shape of the likelihood is what sets the credible intervals in Figs. 9–13. Please add a validation of the likelihood, for example a coverage test on mock observations across the parameter grid, to verify that the reported 1σ and 2σ contours are calibrated. Without such a test, the joint constraints and the comparisons with Ly-alpha forest and HERA limits are not fully supported.","section":"§4.1, Eq. (6)"}],"minor_comments":[{"comment":"The Fig. 6 caption says the power spectrum is from a single 50 cMpc spectrum, while §2.3 constructs spliced 200 cMpc spectra and states that they are used in Section 3.2 and Section 4; please clarify which spectrum length underlies the detectability statement, since the longer spectra lower the noise floor and add low-k bins.","section":"Fig. 6 caption vs. §2.3"},{"comment":"The text says the solid curves are 'mean values over 1000 LOS', while the Fig. 5 caption says 'solid lines correspond to the median values'; please make the terminology consistent.","section":"§3.1 and Fig. 5 caption"},{"comment":"The Gaussian likelihood as written omits the (2π)^{n/2} normalization factor; this is irrelevant for the MCMC implementation, but it would be clearer to state that the normalization is dropped.","section":"Eq. (6)"},{"comment":"The statement that 'the SNR achieved by the uGMRT (SKA1-low) ... is 0.37 (1.2)' refers to the per-channel absorption SNR from Appendix C, not to a detection significance of P21; please rename this quantity to avoid confusion with the power-spectrum detection claim.","section":"Conclusions, first bullet"}],"recommendation":"major_revision","confidential_remarks":"The forward-modeling and inference machinery are solid and the paper is a useful contribution to the 21-cm forest literature. The main issue is that the headline detection claim needs to be reworked: either compute a proper detection significance including sample variance, or substantially soften the abstract and conclusions. The null-detection definition also needs to be made operational. These are fixable within the scope of the manuscript, so I do not recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth reading and worth refereeing. It gives the most concrete forecast I have seen for a statistical 21-cm forest detection at z=6 with uGMRT and SKA1-low, and the null-detection constraints on log10 fX and <xHI> are a real new result. The forward modeling is careful: they include spectral resolution, telescope noise, a quasar SED, a convergence test, and they are openly honest about omitted effects like minihaloes and lightcone evolution. The code is public. The Bayesian recovery tests are appropriate for a forecasting study, and the finding that sample variance dominates over telescope noise is an important, well-documented point.\n\nThe soft spot is exactly what the stress-test note identifies. The abstract and conclusions claim that the 1D power spectrum P21 is detectable at k < 8.5 MHz^-1 with 500 hr of uGMRT, but the only evidence shown is the median signal lying above the noise-only power spectrum (Fig. 6). Their own covariance analysis (Fig. 8, Sec. 4.1) says sample variance is orders of magnitude stronger than telescope noise. For a band-power measurement, the relevant uncertainty is the total variance, not the noise floor. No detection significance or detection probability for P21 is computed anywhere. The KS-test detection fractions in Sec. 3.1 apply to the flux distribution, not to P21, and the Bayesian recovery starts from mock observations that already contain the signal, so it demonstrates parameter inference rather than detection. The phrase \"hence should be detectable\" is doing work that the analysis does not support.\n\nThe other caveats are real but secondary: the 50 cMpc box cannot capture HI islands larger than 50 cMpc, which matter at the low-k scales where detection is claimed; the Gaussian likelihood is likely suboptimal for a non-Gaussian field; and the post-hoc removal of k > 8.5 MHz^-1 bins is a modeling choice that affects the quoted thresholds. These do not break the null-detection constraints as a model-dependent forecast, but they do mean the headline detection thresholds are more fragile than the prose suggests.\n\nWho should read this: anyone planning 21-cm forest observations with uGMRT or SKA1-low, and anyone working on late-end reionization thermal state constraints. The paper deserves a serious referee. The main fix is to either compute a proper detection significance including sample variance, or rephrase the P21 claim as a signal-to-noise-ratio comparison with the noise floor, clearly separated from an actual detection claim. The null-detection constraint forecasts are strong enough to justify publication even if the detection claim is softened.","headline":"A genuinely useful forecast paper whose headline P21 detection claim is not backed by a detection statistic; the null-detection constraints are the stronger and more interesting result.","tokens_in":721,"tokens_out":1281,"would_cite":true,"duration_ms":37125,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A statistical detection of the 21-cm forest is within reach of uGMRT and SKA1-low, and a null detection would tighten constraints on the thermal state of the neutral IGM at the end of reionization.","keywords":["21-cm forest","Epoch of Reionization","intergalactic medium","21-cm power spectrum","neutral hydrogen","radio-loud quasars","uGMRT","SKA1-low"],"falsifier":"Take a 500-h uGMRT band-2 spectrum (or 50-h SKA1-low spectrum) of a z≈6 quasar with S147 ≈ 64 mJy in the cold, 25%-neutral IGM model; if the measured 1D power spectrum shows no excess above the noise-only power at k < 8.5 $MHz^{-1}$ (or k < 32.4 $MHz^{-1}$ for SKA1-low), the central detectability claim is falsified. A cleaner test is to run the same forward model in a simulation box larger than 50 cMpc, or with lightcone effects, and check whether the low-k power spectrum changes by more than the claimed noise margin.","tokens_in":26715,"feed_emoji":"📡","tokens_out":6148,"duration_ms":55770,"temperature":0.7,"pith_summary":"The paper argues that the 21-cm forest, the absorption lines imprinted on radio spectra of distant quasars by neutral hydrogen, can be detected statistically rather than feature-by-feature at the end of reionization ($z\\approx 6$). Forward-modelling semi-numerical reionization simulations with realistic telescope noise, it finds that the one-dimensional power spectrum of the forest rises above the noise at large scales: $k \\lesssim 8.5\\,\\mathrm{MHz}^{-1}$ with 500 h of uGMRT time and $k \\lesssim 32.4\\,\\mathrm{MHz}^{-1}$ with 50 h of SKA1-low, provided the IGM is 25% neutral and the neutral gas has spin temperature $\\lesssim 30$ K. It then shows that a measurement of ten such spectra can jointly constrain the X-ray heating efficiency and neutral fraction, and that a null detection would disfavour cold, significantly neutral IGM models more strongly than current Ly$\\alpha$ forest and 21-cm tomography limits. The practical pay-off is that an outcome of such a modest observing campaign, whether a detection or a null, would teach us about the thermal state of the IGM that other probes cannot see.","feed_headline":"uGMRT and SKA1-low can detect the 21-cm forest power spectrum","feed_subtitle":"Statistical detection needs a 25% neutral IGM colder than 30 K; a null result would tighten reionization constraints.","key_machinery":"The load-bearing object is the one-dimensional power spectrum $P_{21}(k)$ of the normalized 21-cm forest flux $F_{21}=e^{-\\tau_{21}}$, computed from mock spectra drawn from 21cmfast semi-numerical simulations of a $(50\\,\\mathrm{cMpc})^3$ volume at $z=6$. The power spectrum converts the forest from a set of rare individual absorption features into a statistically measurable quantity; the paper uses it to define detectability against the noise power of the telescope and to build a Gaussian likelihood whose covariance matrix includes both telescope noise and, dominantly, sample variance. The spin temperature is assumed equal to the kinetic temperature, which is valid when the Ly$\\alpha$ background is strong at this redshift, and the telescope noise is modelled from the sensitivity curves of uGMRT and SKA1-low.","core_discovery":"The central claim is that the 21-cm forest should be pursued as a statistical signal measured through the one-dimensional power spectrum of the transmitted flux, $P_{21}(k)$, rather than through direct detection of individual absorption lines. In the authors' forward model, for an IGM that is $\\langle x_{\\mathrm{HI}}\\rangle = 0.25$ neutral and preheated so that the spin temperature is $\\lesssim 30$ K, the signal exceeds the noise level at $k \\lesssim 8.5$ MHz$^{-1}$ for 500 h of uGMRT band-2 observations and $k \\lesssim 32.4$ MHz$^{-1}$ for 50 h of SKA1-low observations of a bright $z \\approx 6$ radio-loud quasar. The paper further claims that sample variance, not telescope noise, dominates the uncertainty in parameter inference, so SKA1-low does not dramatically outperform uGMRT for the same integration, and that a null detection of the power spectrum from ten such spectra would place upper limits on the neutral fraction and lower limits on X-ray heating that are competitive with or tighter than current Ly$\\alpha$ forest and 21-cm tomography constraints.","pith_inferences":["If late-reionization neutral islands extend beyond the 50 cMpc simulation box, the large-scale (low-$k$) power could be larger than forecast, making detection easier while strengthening the null-detection limits; the paper's convergence tests address resolution but not box size.","The Gaussian likelihood used may be suboptimal for the non-Gaussian 21-cm forest, so likelihood-free inference could extract more information from the same ten spectra and sharpen the thermal constraints.","The same $P_{21}$ statistic could also probe small-scale structure such as minihaloes or dark-matter candidates, since those imprint on the high-$k$ end of the power spectrum beyond the scales used here for the IGM constraints.","The requirement of a single ultra-bright quasar could be relaxed by targeting the predicted population of roughly 50 radio-loud quasars brighter than 100 mJy at $z>5.5$, if such a population is confirmed by forthcoming surveys."],"forward_implications":["With 500 h of uGMRT on one bright $z\\approx 6$ quasar, the 1D power spectrum is detectable at $k \\lesssim 8.5$ MHz$^{-1}$ if the IGM is 25% neutral and colder than roughly 30 K.","SKA1-low achieves a similar or better detection window, $k \\lesssim 32.4$ MHz$^{-1}$, in only 50 h per source.","Ten measured spectra of 22.1 MHz bandwidth can jointly constrain $\\log_{10} f_X$ and $\\langle x_{\\mathrm{HI}}\\rangle$, and for cold, neutral models these constraints are tighter than current Ly$\\alpha$ forest and 21-cm tomography measurements.","A null detection from ten 500-h uGMRT spectra would rule out $T_{\\mathrm{HI}} \\lesssim 54$ K at 25% neutral fraction, with correspondingly higher temperature limits at higher neutral fractions, a region of parameter space currently unconstrained.","Even 50 h of uGMRT on ten quasars, if null, would disfavour very cold IGMs with neutral fraction $\\gtrsim 0.18$, already competitive with existing probes."],"supporting_citations":[{"why":"Supplies the 21cmfast semi-numerical reionization simulation tool used to generate the IGM skewers.","marker":"Mesinger et al. (2011)"},{"why":"Provides the optical-depth formula and forward-modelling approach for the 21-cm forest spectra.","marker":"Šoltinský et al. (2021)"},{"why":"Supplies the A_eff/T_sys sensitivity curves for uGMRT and SKA1-low used to set the noise levels.","marker":"Braun et al. (2019)"},{"why":"Sets the mean free path of ionizing photons at z=6 used to calibrate the simulated neutral fraction.","marker":"Becker et al. (2021)"},{"why":"Defines the X-ray luminosity parametrization with efficiency f_X that controls the thermal state of the IGM.","marker":"Furlanetto (2006)"},{"why":"Supplies the discrete power-spectrum estimator used to compute P21 from mock spectra.","marker":"Khan et al. (2024)"},{"why":"Characterises PSO J0309+27, the bright z≈6 quasar with S147=64.2 mJy and alpha_R=-0.44 used as the background source template.","marker":"Belladitta et al. (2020)"},{"why":"Provides a Ly-alpha forest neutral-fraction measurement at z≈6 that the null-detection constraints are compared against.","marker":"Gaikwad et al. (2023)"},{"why":"Provides the 21-cm power-spectrum spin-temperature limits used to display the currently allowed thermal range.","marker":"HERA Collaboration (2023)"},{"why":"Forecasts a population of bright radio-loud quasars at z>5.5 that would make such observations feasible.","marker":"Niu et al. (2024)"}],"fun_headline_variants":["Statistical 21-cm forest detection within reach for uGMRT and SKA1-low","uGMRT and SKA1-low can detect 21-cm forest power spectrum","Null 21-cm forest detection rivals Lyα and 21-cm tomography","500h uGMRT or 50h SKA1-low: 21-cm forest power spectrum detection","21-cm forest P(k) detection: possible with uGMRT and SKA1-low"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forecasts assume that the 50 cMpc simulation box contains all the neutral-hydrogen structure that produces the 21-cm forest power at the claimed detection scales; the paper itself notes that larger neutral islands, which late-reionization models predict, lie outside the box and could change both the predicted signal and the null-detection limits.","fun_headline_variants_meta":{"raw":{"variants":["Statistical 21-cm forest detection within reach for uGMRT and SKA1-low","uGMRT and SKA1-low can detect 21-cm forest power spectrum","Null 21-cm forest detection rivals Lyα and 21-cm tomography","500h uGMRT or 50h SKA1-low: 21-cm forest power spectrum detection","21-cm forest P(k) detection: possible with uGMRT and SKA1-low"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00275,"raw_usage":{"total_tokens":10612,"prompt_tokens":1203,"completion_tokens":9409,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":819,"completion_tokens_details":{"reasoning_tokens":9287}},"tokens_in":819,"tokens_out":9409,"duration_ms":58806,"temperature":1.0,"reasoning_tokens":9287,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:19:26.617875+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a 500-h uGMRT band-2 spectrum (or 50-h SKA1-low spectrum) of a z≈6 quasar with S147 ≈ 64 mJy in the cold, 25%-neutral IGM model; if the measured 1D power spectrum shows no excess above the noise-only power at k < 8.5 $MHz^{-1}$ (or k < 32.4 $MHz^{-1}$ for SKA1-low), the central detectability claim is falsified. A cleaner test is to run the same forward model in a simulation box larger than 50 cMpc, or with lightcone effects, and check whether the low-k power spectrum changes by more than the claimed noise margin.","supporting_citations":[{"cited_title":"K., Kulkarni G., Bolton J","cited_arxiv_id":null,"evidence_quote":"Supplies the discrete power-spectrum estimator used to compute P21 from mock spectra."}],"review_version":1}