{"id":"352e609f-94ba-46ab-adee-fa2fe4485fbe","arxiv_id":"2507.10533","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"NenuFAR observations of the NT04 field set the deepest 21-cm power spectrum upper limits to date at z=20.3 and z=17.0, with the z=20.3 limit more than an order of magnitude deeper than any previous Cosmic Dawn limit.","lead":"Astronomers report the deepest upper limits yet on the power spectrum of the 21-cm hydrogen signal from the Cosmic Dawn, using four nights of NenuFAR observations of a carefully chosen field. The limits at redshift 20.3 are over an order of magnitude stronger than all previous Cosmic Dawn limits and begin to rule out the most extreme EDGES-motivated models.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ML-GPR may absorb 21-cm signal shapes outside the VAE latent space; the §5.2 injection tests only use VAE-space signals, so the quoted 2σ limits could be biased low.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the GPR signal kernel, together with the excess kernels, could absorb a portion of the 21-cm signal, making the quoted 2σ upper limits invalid. I agree this is the most critical threat to the central claim. The paper does perform injection tests, but these reuse the VAE latent space, making them circular with respect to the signal model. The 1.33% of z-scores below −2 in the in-space tests show that even for signals inside the model, absorption happens in some k bins; out-of-space signals could suffer much worse. Since the headline limit at z=20.3 is a factor of only ~2 above the residual power, even a modest negative bias at that k bin would break the 'deepest limit' claim. The concrete test I propose directly probes this failure mode by injecting signals with shapes that the VAE cannot represent. If the test passes, the concern is mitigated; if it fails, the paper must soften its claims or add a conservative bias correction. This does not change the verdict: CONDITIONAL remains appropriate until the test is performed. I also considered the field/night selection effect, but that is secondary: the field selection is based on known foreground properties rather than the final noise realization, and the paper already treats the excess variance conservatively by not subtracting it.","tokens_in":32163,"tokens_out":6448,"duration_ms":78809,"concrete_test":"Run injection tests with synthetic 21-cm signals deliberately outside the VAE latent space, e.g., a flat power spectrum P(k) ∝ k^0, a steep spectrum P(k) ∝ k^{-2}, and a signal with a Lorentzian frequency coherence of width 0.5 MHz, each injected at amplitudes comparable to the quoted A²≈2×10^5 mK² at k=0.038 for Z20. Measure the z-score at the headline k bin. If any out-of-space shape is recovered with z < −2 (i.e., biased low by >2σ), the GPR is not robust to unknown signal shapes and the upper limits must be qualified as model-dependent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—the deepest 21-cm power-spectrum upper limits at z=20.3 and z=17.0—rests on the residual power spectrum after ML-GPR foreground removal. In the GPR model (Eqs. 1–3, §5.1), the 21-cm component is described by a VAE kernel trained on 21cmFAST simulations, while two Matern 3/2 excess kernels (Kex1, Kex2) with short coherence scales (Table 3: l≈0.045 and 0.251 MHz for Z20) absorb residual small-coherence power. A real 21-cm signal whose spectral shape is not representable by the VAE latent space—for example, with a flatter or steeper power-spectrum slope than any 21cmFAST model, or with a narrow frequency feature—would be attributed to the excess or foreground components and subtracted, biasing the residual power spectrum low. The signal injection tests in §5.2 draw injected signals from the same VAE latent space (x1, x2 ∈ [−2, 2]) used by the GPR signal kernel; they cannot detect this failure mode. Even within that space, 1.33% of z-scores fall below −2, indicating some absorption. The headline limit at z=20.3 (Δ² < 4.6×10^5 mK² at k=0.038) is only modestly above the measured residual (2.1×10^5 mK²), so a downward bias of even ~2σ at that k bin would invalidate the limit as an upper bound on the true 21-cm signal. Without an out-of-sample injection test, the claimed 'deepest limits' are conditional on the VAE latent-space assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports 2σ upper limits on the 21-cm power spectrum during the Cosmic Dawn using four nights of NenuFAR observations of a newly selected NT04 field, analyzed in two spectral windows centered at z=20.3 and z=17.0. The analysis pipeline includes improved A-team and 3C source subtraction, uv-plane RFI flagging, an elevation cut to confine the foreground wedge, and machine-learning Gaussian process regression (ML-GPR, the same framework as M24) for residual foreground removal. After GPR, the authors obtain noise-bias-subtracted residual power spectra and set upper limits of Δ²21 < 4.6×10^5 mK² at k=0.038 h/cMpc for z=20.3 and Δ²21 < 5.0×10^6 mK² at k=0.041 h/cMpc for z=17.0, claiming these are the deepest Cosmic Dawn limits at these redshifts. They also present signal injection tests, an analysis of the coherence of the excess variance, and a comparison of the z=20.3 limits against simulated exotic 21-cm models.","tokens_in":32515,"tokens_out":5868,"duration_ms":72824,"significance":"If the quoted limits are robust, this is an important advance for 21-cm cosmology at Cosmic Dawn: the z=20.3 limit improves on previous NenuFAR/NCP results by more than an order of magnitude and begins to approach predictions of extreme EDGES-inspired models. The paper is thorough in its calibration description, uses a conservative approach by not subtracting the excess variance, and provides a useful analysis of the temporal coherence of the excess. The public pipelines (Nenuflow, pspipe), explicit calibration parameters, and signal injection tests are strengths. However, the central claim of 'deepest upper limits' is conditional on the ML-GPR model not absorbing the 21-cm signal, and the current injection tests only probe signals drawn from the same VAE latent space used in the GPR model, leaving the most important failure mode untested.","major_comments":[{"comment":"The central claim of deepest upper limits rests on the residual power spectrum after ML-GPR subtraction, in which the 21-cm component is modelled by a VAE kernel trained on 21cmFAST while two Matern 3/2 excess kernels (Kex1, Kex2; Table 3) absorb short-coherence power. A real 21-cm signal whose spectral shape lies outside the VAE latent space—e.g. a flatter or steeper power-spectrum slope than any 21cmFAST realization, or a narrow spectral feature—would be attributed to Kex or the foreground component and subtracted, biasing the residual power spectrum low. The injection tests in §5.2 draw injected signals from the same VAE latent space (x1, x2 ∈ [−2, 2]) used by the GPR signal kernel, so they cannot detect this failure mode; the tests also cover only the central part of the prior range [−4, 4] and already show 1.33% of z-scores below −2. Because the quoted z=20.3 limit at k=0.038 (4.6×10^5 mK²) is only about 2.2 times the measured residual (2.1×10^5 mK²), even a modest downward bias would invalidate the limit as an upper bound on the true 21-cm signal. I ask for out-of-sample injection tests with signals not representable by the VAE (e.g. power-law spectra of varying slope, narrow band features, or the 21cmSPACE/LICORICE models used in §7.1) and a quantification of the resulting bias at the k bins of Table 4.","section":"§5.1–5.2 (Eqs. 1–3, Table 3, Fig. 13)"},{"comment":"The field (NT04) and the four nights were selected after the fact using criteria directly tied to the measured power spectra: the field was identified as optimal from a survey using power spectra after subtraction of brightest sources as a metric, and the four nights were chosen 'based on superior RFI statistics' (Table 1, Fig. 2). This selection on the dependent variable can bias the quoted upper limits downward relative to a pre-defined observing strategy. To support the claim that these are the deepest limits to date, the authors should either present the corresponding limits for all four nights without selection, report limits for the full survey of fields, or quantify the selection bias (e.g. by bootstrap over the six survey fields/nights).","section":"§2.2"},{"comment":"The 21-cm hyperparameters are unconstrained (x1, x2) and σ²21 has only an upper limit (< −3.208 for Z20, < −2.749 for Z17), which the text correctly interprets as no detection. However, the upper limits are not constructed from the fitted 21-cm component but from the total residual power after subtracting the posterior-mean foreground. The uncertainty on the residual therefore depends on the assumed GPR model; if the model is misspecified, the ensemble of 1000 posterior samples provides underestimates of the true systematic uncertainty. The paper should state this explicitly and, if possible, quantify the model-marginalized uncertainty, e.g. by varying the choice of excess kernels or priors and recomputing the limits.","section":"§5.1, Table 3"}],"minor_comments":[{"comment":"The reference 'Bennett A., Simth F., 1962' contains a typo in the second author's name; it should be 'Smith'.","section":"References"},{"comment":"The field-selection criteria are deferred to Mertens et al. (in prep.), which makes it difficult for the reader to assess the selection bias discussed above; please include the relevant selection metrics or a summary in an appendix.","section":"§2.2"},{"comment":"The caption of the left panel says the z-scores are shown as a function of 'injected signal variances', but the x-axis appears to be k; please clarify the meaning of the axis and the color/line coding.","section":"Fig. 13"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about out-of-sample signal shapes is real and is the main reason for the major-revision recommendation. I do not think the paper should be rejected, because the pipeline is detailed, the limits are computed conservatively, and the missing out-of-sample tests are a well-defined addition rather than a fundamental flaw. I would also encourage the editor to ask the authors to address the field/night selection bias explicitly, as the current 'deepest limits to date' claim is tied to a post-hoc selection that, if not quantified, could be seen as over-claiming."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a solid observational paper. The new field (NT04), the adaptive calibration, the uv-plane RFI flagging, and the elevation selection are concrete pipeline improvements, and the 50-fold reduction in excess variance relative to the NCP analysis is credible. The z=20.3 limit is the deepest Cosmic Dawn power-spectrum limit to date, and the z=17.0 bin is genuinely new. The paper is also unusually honest: it describes the limitations, does not subtract the excess variance when setting limits, and includes injection tests and a coherence analysis. I believe the headline numbers.\n\nThe soft spots are real but manageable. The ML-GPR foreground removal uses a VAE kernel trained on 21cmFAST shapes; if the true 21-cm signal has spectral structure outside that latent space, the excess kernels could absorb part of it and bias the residual power low. The injection tests only draw from the same VAE space, so they cannot catch this failure mode, and even within that space 1.33% of z-scores fall below -2. The z=20.3 upper limit sits only modestly above the measured residual, so a downward bias of about 2 sigma at that k bin would indeed erode the claim. But this is a modeling assumption common in the field, not a sloppy one, and the paper flags the relevant posterior behavior (21-cm variance hitting the prior boundary). The field and night selection also used power-spectrum and RFI statistics from the same type of data, so there is some selection risk, though I do not see it as fatal. Data availability on request is a minor reproducibility annoyance, not a scientific defect.\n\nI largely agree with the reader's conditional verdict, but I would push back on the circularity score: the limits are measured residuals, not fitted VAE predictions, and the injection tests reuse the VAE kernel mainly as a practical necessity. The circularity is minor, not central.\n\nThis paper deserves serious peer review. It will be useful to anyone working on 21-cm upper limits or on GPR-based foreground removal, and the careful discussion of where the excess variance comes from is worth engaging with. I would cite it and bring it to a reading group.","headline":"A careful, genuinely improved upper-limit paper from NenuFAR whose headline numbers are solid, though the quoted limits inherit a real but not fatal caveat about the ML-GPR signal model.","tokens_in":718,"tokens_out":708,"would_cite":true,"duration_ms":28784,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Four nights of NenuFAR observations of a carefully chosen field produce the deepest 21-cm power spectrum upper limits yet from the Cosmic Dawn, more than an order of magnitude below all previous limits at these redshifts.","keywords":["21-cm cosmology","Cosmic Dawn","power spectrum upper limits","NenuFAR","radio interferometry","foreground subtraction","Gaussian process regression","EDGES exotic models"],"falsifier":"A decisive test is to inject a synthetic 21-cm signal with a deliberately short frequency coherence scale (outside the VAE latent space) into the real calibrated visibility cubes and run the GPR; if the recovered power is biased low by more than 2σ in any k-bin used for the limits, the quoted upper limits do not hold for that signal class.","tokens_in":31880,"feed_emoji":"📡","tokens_out":9847,"duration_ms":91466,"temperature":0.7,"pith_summary":"This paper reports the deepest upper limits so far on the 21-cm signal power spectrum from the Cosmic Dawn, using four nights of NenuFAR observations of the NT04 field: a 2σ limit of $\\Delta^2_{21} < 4.6\\times10^5\\,\\mathrm{mK}^2$ at $k = 0.038\\,h\\,\\mathrm{cMpc}^{-1}$ at $z=20.3$, and $\\Delta^2_{21} < 5.0\\times10^6\\,\\mathrm{mK}^2$ at $k = 0.041\\,h\\,\\mathrm{cMpc}^{-1}$ at $z=17.0$. The $z=20.3$ result improves on all previous Cosmic Dawn power spectrum limits by more than an order of magnitude, and reduces the excess variance over the earlier north celestial pole analysis by roughly a factor of 50. The authors attribute the improvement to selecting an optimal field that minimizes sidelobe leakage from bright off-axis sources, plus pipeline upgrades in A-team and 3C source subtraction, RFI mitigation in the uv plane, and an elevation cut that confines foregrounds to low line-of-sight modes. Comparing the limits with simulated exotic 21-cm signals, the $z=20.3$ limit begins to exclude only the most extreme models (those predicting global signals stronger than the EDGES detection); an order-of-magnitude deeper limit would start to constrain EDGES-compatible signals. The paper argues that the excess variance in the $z=20.3$ bin is largely incoherent across nights, so continued integration could push the limits substantially deeper.","feed_headline":"NenuFAR sets deepest 21-cm Cosmic Dawn limits yet","feed_subtitle":"At z≈20.3, the 2σ 21-cm power spectrum limit is Δ²<4.6×10⁵ mK², an order of magnitude deeper than all earlier Cosmic Dawn limits.","key_machinery":"The argument rests on a multi-stage calibration and foreground-removal pipeline. First, direction-dependent calibration with adaptive solution intervals subtracts the bright A-team sources (Cas A, Cyg A, and others) as they move through NenuFAR's primary-beam grating lobes, using apparent sky models built from a simulated primary beam. A more conservative 3C-source subtraction avoids overfitting on the short baselines used for the power spectrum, and clustered target-field sources are subtracted down to the confusion limit. Two data-selection steps then confine foregrounds: five uv cells around the $u=0$ line (where stationary RFI accumulates for a non-NCP phase centre) are flagged, and only time ranges with phase-centre elevation above $50^\\circ$ are kept, which keeps the full-sky horizon line shallow. Finally, machine-learning-enhanced Gaussian process regression (ML-GPR) models the residual data as the sum of foreground, 21-cm signal, excess, and noise Gaussian processes: the foreground kernels are RBF, the excess uses two Matern 3/2 kernels, and the 21-cm kernel is a variational autoencoder (VAE) trained on 21cmFAST simulations, with its variance left free so that boosted exotic-model signals can in principle be absorbed. The GPR is run on each spectral window separately, and the residual power spectrum is estimated from an ensemble of 1000 posterior hyperparameter samples.","core_discovery":"The paper's central claim is that a better-chosen target field and a refocused calibration pipeline allow NenuFAR to set the deepest 21-cm power spectrum upper limits yet obtained from the Cosmic Dawn. After calibration-based sky-model subtraction and Gaussian-process-regression foreground removal, the noise-bias-subtracted residual power spectrum gives a best 2σ upper limit of $\\Delta^2_{21} < 4.6\\times10^5\\,\\mathrm{mK}^2$ at $k = 0.038\\,h\\,\\mathrm{cMpc}^{-1}$ in the $z=20.3$ bin, and $\\Delta^2_{21} < 5.0\\times10^6\\,\\mathrm{mK}^2$ at $k = 0.041\\,h\\,\\mathrm{cMpc}^{-1}$ in the $z=17.0$ bin. The $z=20.3$ limit improves on all previous Cosmic Dawn power spectrum limits by more than an order of magnitude and on the earlier NenuFAR NCP analysis by roughly a factor of 50. The paper further claims that these limits begin to exclude the most extreme exotic-model 21-cm signals invoked to explain the EDGES absorption feature, and that the remaining excess variance in the $z=20.3$ bin is largely incoherent across nights, so continued integration should deepen the limits substantially.","pith_inferences":["Beyond the paper: if the $z=20.3$ excess variance stays incoherent over hundreds of hours, the 2σ limit should scale roughly as the inverse square root of integration time, bringing the EDGES-compatible model range within reach of the ~500 hours of NT04 data already collected.","Beyond the paper: the GPR verification tests only injected signals drawn from the VAE latent space; a test with signals of deliberately short frequency coherence would directly probe whether the Matern excess kernels can absorb part of the true 21-cm signal.","Beyond the paper: the elevation-cut strategy suggests that future Cosmic Dawn surveys could optimise their LST scheduling to keep the phase centre high, effectively trading integration time for foreground confinement."],"forward_implications":["The $z=20.3$ limit is more than an order of magnitude below all previous Cosmic Dawn power spectrum limits, and roughly 50 times below the earlier NenuFAR NCP analysis at the same redshift.","These limits begin to exclude the most extreme exotic-model 21-cm signals — those predicting a global absorption signal stronger than the EDGES detection.","An order-of-magnitude deeper limit, which the paper argues is attainable because the $z=20.3$ excess variance integrates down incoherently across nights, would probe models with signal strengths comparable to EDGES.","The $z=17.0$ result is the deepest limit at that redshift, but it is dominated by coherent RFI-related excess variance, so longer integrations alone will not improve it much without additional RFI mitigation."],"supporting_citations":[{"why":"Previous NenuFAR NCP analysis whose limits and excess variance this work improves by a factor of roughly 50.","marker":"Munshi et al. 2024 (M24)"},{"why":"Introduced the Gaussian process regression foreground-removal method adapted here.","marker":"Mertens et al. 2018"},{"why":"Supplies the ML-GPR framework with the VAE-based 21-cm signal kernel used in residual foreground removal.","marker":"Mertens et al. 2024"},{"why":"21cmFAST simulations used to train the VAE kernel and to generate the injected-signal shapes for GPR tests.","marker":"Mesinger et al. 2011"},{"why":"Full-sky horizon line equations that motivate the elevation-dependent foreground-wedge arguments and the elevation cut.","marker":"Munshi et al. 2025a"},{"why":"Forward simulations showing RFI signature along the u=0 line, supporting the uv-plane flagging strategy.","marker":"Munshi et al. 2025b"},{"why":"The EDGES detection whose exotic-model interpretation sets the astrophysical benchmark that the new limits begin to test.","marker":"Bowman et al. 2018"}],"fun_headline_variants":["NenuFAR's improved 21-cm limits reach deep Cosmic Dawn","NenuFAR sets deepest 21-cm upper limits yet at z≈20","NenuFAR's z=20.3 21-cm limit beats all by 10x","21-cm Cosmic Dawn limits improved by NenuFAR's new field","NenuFAR achieves best 21-cm constraints on Cosmic Dawn"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The GPR foreground removal assumes that none of the true 21-cm signal is absorbed by the foreground or excess kernels, meaning the signal's frequency coherence is always large enough and its shape close enough to the trained VAE to stay in the signal kernel; if part of the signal leaks into the excess or foreground components, the quoted 2σ limits would be too low.","fun_headline_variants_meta":{"raw":{"variants":["NenuFAR's improved 21-cm limits reach deep Cosmic Dawn","NenuFAR sets deepest 21-cm upper limits yet at z≈20","NenuFAR's z=20.3 21-cm limit beats all by 10x","21-cm Cosmic Dawn limits improved by NenuFAR's new field","NenuFAR achieves best 21-cm constraints on Cosmic Dawn"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000438,"raw_usage":{"total_tokens":2388,"prompt_tokens":1272,"completion_tokens":1116,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":888,"completion_tokens_details":{"reasoning_tokens":1008}},"tokens_in":888,"tokens_out":1116,"duration_ms":11618,"temperature":1.0,"reasoning_tokens":1008,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:29:37.331541+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test is to inject a synthetic 21-cm signal with a deliberately short frequency coherence scale (outside the VAE latent space) into the real calibrated visibility cubes and run the GPR; if the recovered power is biased low by more than 2σ in any k-bin used for the limits, the quoted upper limits do not hold for that signal class.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Previous NenuFAR NCP analysis whose limits and excess variance this work improves by a factor of roughly 50."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduced the Gaussian process regression foreground-removal method adapted here."},{"cited_title":"G., Bobin J., Carucci I","cited_arxiv_id":null,"evidence_quote":"Supplies the ML-GPR framework with the VAE-based 21-cm signal kernel used in residual foreground removal."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"21cmFAST simulations used to train the VAE kernel and to generate the injected-signal shapes for GPR tests."}],"review_version":1}