{"id":"9330f62a-ddc6-415f-885c-a3fa481c83ec","arxiv_id":"2608.13153","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A mock-observation forecast finds that WEAVE-QSO measurements of the Lyman-alpha forest power spectrum at z≥4 can detect the relic large-scale power from patchy hydrogen reionization at 4.5 sigma.","lead":"Using simulated spectra and a detailed noise model, the authors forecast that the WEAVE-QSO survey will detect the imprint of patchy hydrogen reionization in the Lyman-alpha forest power spectrum at z=4.0 to 4.6 at 4.5 sigma significance. This matters because it would turn a tentative hint from existing data into a robust, large-scale measurement of how the first galaxies reionized the universe.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Detection significance hinges on the lowest-k bin, whose forecast covariance is dominated by HCD damping-wing and continuum systematics; the paper's own no-lowest-k test drops the preference from 4.5σ to 1.2σ.","rationale":"The reader and I identify the same weakest assumption, so agreement is 'agree'. The paper is careful and transparent; the forecast pipeline is detailed, and the authors themselves flag the lowest-k sensitivity and the HCD/continuum dominance. That honesty, however, does not remove the load-bearing nature of the assumption. The strongest claim in the abstract is an unconditional-sounding 'should be detectable at ~4.5σ', but the evidence is conditional on an unvalidated model of HCD damping-wing residual and continuum errors at z ≥ 4. The Wilks-theorem boundary issue raised by the reader is not, in my view, a weakening concern: with four Ap parameters at their lower boundary under the null, the correct asymptotic null is a chi-bar-square mixture, which for Δχ² = 29.21 yields a smaller p-value than the quoted chi-square(4) tail, so the quoted significance is conservative rather than optimistic. The load-bearing issue is the external systematics. Because the paper's own no-lowest-k test shows the significance collapses to 1.2σ, and because the HCD completeness function is an external calibration not tested by the mocks, the verdict remains conditional: the forecast is publishable and useful, but the headline significance should be advertised as conditional on HCD/continuum systematics, and ideally the abstract should carry the caveat. No change to the reader's CONDITIONAL verdict is needed.","tokens_in":20890,"tokens_out":8965,"duration_ms":82339,"concrete_test":"Recompute the Section 4.2 model comparison after perturbing the HCD completeness function in Section 3.5 within the published uncertainty of Wang et al. (2022), including a deliberately conservative case with 50% completeness at log N_HI = 19–20 at z ≈ 4.5, and recompute the lowest-k covariance term. If the resulting Δχ² drops below about 10 (or the Gaussian-equivalent significance below about 3σ), the headline 4.5σ detection is not robust to HCD modeling; if Δχ² remains above about 10, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The 4.5σ claim is an injection-recovery forecast in which the mock data vector and the covariance are both built from Sherwood-Relics and the adopted instrument model; it cannot validate the astrophysical systematics that dominate the modes carrying the signal. Section 4.2 shows that removing the lowest-k bin (k ≲ 0.0033 s/km) reduces Δχ² from 29.21 to 5.40 (1.2σ), and Section 4.2 states that the uncertainties in those bins are dominated by incomplete HCD identification and, secondarily, continuum placement. The HCD residual term in Section 3.5 is computed as the fractional power difference using the Wang et al. (2022) completeness function for N_HI > 10^19 cm^-2, applied at z ≈ 4–4.6 with R ≈ 5000; if the real completeness is lower, or the HCD column-density distribution differs from Prochaska et al. (2014), the residual damping-wing power at k ~ 10^-3 s/km is underestimated. Because the patchy reionization signal is itself a smooth large-scale enhancement, residual HCD power can partially mimic or absorb it, shrinking the apparent Δχ². The paper's quadrature-summed systematic term is a forecast-level estimate, not a measured calibration. Hence the central claim is only as secure as the HCD completeness and continuum models at z ≥ 4, which are external inputs rather than outputs of the pipeline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Ma et al. present a forecast for detecting the large-scale imprint of patchy hydrogen reionization in the 1D Lyman-alpha forest power spectrum with the WEAVE-QSO survey at z=4.0-4.6. They generate mock spectra from the Sherwood-Relics simulations, process them with the WEAVEify instrument forward model, and construct a forecast covariance that includes statistical errors and systematics from spectral resolution, noise subtraction, metals, high-column-density (HCD) damping wings, and continuum placement. They then perform an MCMC injection-recovery analysis with an A_p amplitude parameter for the patchy reionization correction, finding Delta chi^2 = 29.21 between models with and without A_p (about 4.5 sigma). Removing the lowest-k bin reduces this to Delta chi^2 = 5.40 (about 1.2 sigma), which the paper reports transparently.","tokens_in":21147,"tokens_out":12909,"duration_ms":117626,"significance":"This is a useful and unusually transparent forecast for a key observable. Its strengths include an end-to-end forward model for WEAVE-QSO spectra, a systematic budget that goes beyond most Lyman-alpha forest forecasts, explicit tests of robustness to thermal priors, and honest reporting of the critical role of the lowest-k bin. If the adopted systematics and the Sherwood-Relics reionization model are correct, the paper provides a concrete, falsifiable prediction: WEAVE-QSO will recover the redshift-dependent A_p amplitude at about 4.5 sigma. However, the significance is conditional on external HCD completeness and continuum models and on the Sherwood-Relics patchy reionization morphology; it is an injection-recovery forecast rather than a model-independent detection claim. With revisions that address the treatment of large-scale systematics and the framing of the headline claim, this would be a valuable benchmark for the community.","major_comments":[{"comment":"The headline significance of about 4.5 sigma is almost entirely carried by the lowest-wavenumber bin in each redshift bin. Removing the bin at k less than about 0.0033 s/km reduces Delta chi^2 from 29.21 to 5.40, corresponding to about 1.2 sigma. The paper itself states that the uncertainties in those bins are dominated by incomplete HCD identification and, to a lesser extent, continuum placement. The HCD residual term in Section 3.5 is computed using the Wang et al. (2022) completeness function applied at z about 4-4.6 with R about 5000 and the Prochaska et al. (2014) column-density distribution; if the real completeness is lower, or the HCD population differs, the residual damping-wing power at k about 10^-3 s/km is underestimated. Because the patchy reionization signal is itself a smooth large-scale enhancement, residual HCD power can partially mimic or absorb it. The 4.5 sigma claim is therefore only as secure as these external inputs, and the paper should demonstrate robustness to plausible variations, for example by treating an HCD normalization amplitude as a nuisance parameter in the fits.","section":"Section 4.2, Figs 10-11"},{"comment":"The covariance matrix used for the Delta chi^2 statistic appears to be diagonal, with the HCD and continuum systematic terms added in quadrature to the statistical variance in each k bin. The HCD and continuum systematics are smooth functions of wavenumber (Figs 7 and 8), so an error in their amplitude or shape produces correlated shifts across the large-scale bins that dominate the A_p constraint. A diagonal covariance underweights such a coherent systematic and will tend to overstate the significance of a smooth signal. The authors should either marginalize over the amplitudes of these systematics in the MCMC or construct a non-diagonal systematic covariance before quoting a single detection significance.","section":"Section 3.7, Fig. 9"},{"comment":"The mock data vector is generated from the same Sherwood-Relics simulations and the same Molaro et al. (2022, 2023) patchy reionization correction used to define A_p in the inference. The reported significance is therefore an injection-recovery consistency check under the assumed model, not an external validation of the model or of the systematic inputs. This is not itself an error, since injection-recovery is a standard forecasting tool, but the abstract's statement that WEAVE-QSO should be able to directly detect the relic imprint should be explicitly conditioned on the fidelity of the Sherwood-Relics reionization morphology and on the adopted systematic models, so that the headline significance is not misread as a model-independent detection claim.","section":"Sections 2.1 and 4.1"}],"minor_comments":[{"comment":"The text refers to a survey configuration of approximately 10,000 deg2 with m_r < 22, but the adopted cut throughout the paper is m_r < 21.9 (Section 2.4); please reconcile this inconsistency.","section":"Section 4.2, last paragraph"},{"comment":"The caption reads 'The posterior for A_p is does not change significantly'; this should read 'does not change significantly'.","section":"Appendix A, Fig. A2 caption"},{"comment":"The synthetic HCDs are inserted at locations where the Lyman-alpha optical depth is largest; this clustering assumption should be justified or tested, because it affects the large-scale residual power that enters the systematic budget.","section":"Section 3.5"},{"comment":"The flat-continuum test is a useful bounding exercise, but it does not necessarily bracket the wavelength-dependent PCA continuum errors of WEAVEify; the continuum systematic term should be described as a forecast-level estimate rather than a calibrated error.","section":"Section 3.6"},{"comment":"The use of Wilks' theorem for the Delta chi^2 statistic is not strictly regular because the A_p parameters are constrained to A_p >= 0 and the null model lies on the boundary. In the present case the best-fit values are well away from the boundary, so the correction would not weaken the conclusion, but the statement should be qualified.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"This is a solid forecast paper with an unusually complete systematic budget for the Lyman-alpha forest. The main risk is the fragility of the headline significance to the lowest-k bin and to the HCD and continuum systematics that set the error budget in that bin; the authors acknowledge this but need to add robustness tests and reframe the claim before the 4.5 sigma statement can be taken at face value. The work fits the scope of MNRAS well."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing to know: this is a well-executed forecast paper, but the headline 4.5σ detectability of patchy reionization is an injection-recovery that leans almost entirely on the lowest-k bin and on the Sherwood-Relics model being the truth. The paper itself is admirably honest about this — removing that bin drops the preference to 1.2σ — but the abstract still leads with the 4.5σ.\n\nWhat's genuinely new: WEAVEify, the forward-modeling pipeline that turns hydro skewers into WEAVE-like spectra with detector-level noise, and the survey-specific forecast built on it. The covariance budget is unusually thorough for a forecast: resolution, noise subtraction, metals (correlated and uncorrelated), HCD damping wings, and continuum placement all get separate treatment, and the relative contributions are shown explicitly. I also give credit for the loose-thermal-prior test, which shows the A_p recovery isn't just an artifact of tight priors on T0 and gamma.\n\nThe soft spots are real but mostly acknowledged in the text. First, the circularity burden is high: the mock data vector is built using the exact patchy correction and A_p parameterization that the inference is designed to recover, both from the authors' own prior work. That's fine for a statistical precision forecast, but it means the 4.5σ is conditional on the Sherwood-Relics patchy model, not a validation of it. The abstract should say that more prominently. Second, the significance is driven by the lowest-k bin, and the forecast uncertainty in that bin is dominated by HCD damping-wing incompleteness and, secondarily, continuum placement. Those are external inputs — Wang et al. (2022) completeness, Prochaska et al. (2014) CDDF — not outputs of the pipeline. If the real completeness at z≈4 is worse, the residual HCD power can absorb part of the large-scale signal and the Δχ² shrinks. Third, applying Wilks' theorem with A_p bounded at zero is questionable; the likelihood ratio statistic may not be chi-square with 4 dof, so 4.5σ is probably optimistic. The paper's own \"two-sided Gaussian-equivalent\" phrasing hints at this without addressing it.\n\nThese issues don't sink the paper. The central forecast — that WEAVE-QSO has the statistical precision to turn the 2.7σ hint into a real measurement if the systematics are controlled at the level modeled — holds up as a forecast. I'd send it to peer review, with requests that the authors (1) state the model-conditional nature of the significance in the abstract, (2) address the boundary/Wilks point, and (3) release the code and mocks. Code on \"reasonable request\" is not great for a forecast paper whose whole value is the pipeline.\n\nWho is this for? Anyone planning Ly-alpha forest science with WEAVE-QSO or similar surveys (DESI, 4MOST, PFS). It's a serious forecast paper, not a breakthrough detection claim. Worth a reading group if you care about reionization probes or survey forecasting methodology.\n\nRecommendation: accept after minor-to-moderate revision, with the caveats above. The forecast machinery is solid, but the headline significance needs to be framed as conditional and the statistical test needs scrutiny.","headline":"A careful, honest forecast whose headline 4.5σ is an injection-recovery that depends on the lowest-k bin and on Sherwood-Relics being right; deserves peer review, but the abstract needs to state the model-conditional nature.","tokens_in":21851,"tokens_out":4118,"would_cite":false,"duration_ms":31473,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"With the forecast uncertainties adopted here, WEAVE-QSO measurements of the 1D Lyman-$\\alpha$ forest power spectrum at $z \\geq 4$ should detect the large-scale relic imprint of patchy hydrogen reionization at $4.5\\sigma$ significance.","keywords":["intergalactic medium","Lyman-alpha forest","patchy reionization","one-dimensional power spectrum","WEAVE-QSO","cosmological simulations","quasar absorption lines","forecasts"],"falsifier":"Compare the real WEAVE-QSO measurement in the lowest-wavenumber bin ($k \\sim 10^{-3}\\,\\mathrm{s\\,km^{-1}}$) against the forecast covariance: if the scatter of that bin across independent line-of-sight subsamples exceeds the forecast systematic error, or if the recovered $A_p$ posterior at $z=4.6$ does not include 0.45 within the 95% credible interval, the central claim is falsified. A direct null test is to fit the observed $P_{\\mathrm{Ly}\\alpha}$ at $z=4$ with and without $A_p$ and check whether $\\Delta\\chi^2$ remains above about 25.","tokens_in":20644,"feed_emoji":"🔭","tokens_out":9949,"duration_ms":71003,"temperature":0.7,"pith_summary":"The paper argues that the forthcoming WEAVE-QSO survey should be able to detect the large-scale relic imprint of patchy hydrogen reionization in the one-dimensional Lyman-$\\alpha$ forest power spectrum at $z=4.0$--$4.6$. Using the Sherwood-Relics reionization simulations and a survey-level mock pipeline, the authors forecast the full uncertainty budget---sample size, spectral resolution, noise subtraction, continuum placement, metal lines, and damping wings from high-column-density absorbers---and then fit mock power spectra with a Bayesian inference framework. Allowing one patchy reionization amplitude parameter $A_p$ per redshift bin improves the fit by $\\Delta\\chi^2 = 29.21$ for four extra parameters, which Wilks' theorem converts to a two-sided Gaussian-equivalent $4.5\\sigma$ preference. The signal is driven almost entirely by the lowest-wavenumber bin in each redshift bin; removing that bin reduces the preference to about $1.2\\sigma$. If the forecast covariance is realized in the data, WEAVE-QSO would turn the tentative $2.7\\sigma$ evidence already reported in current data into a direct detection.","feed_headline":"WEAVE-QSO can spot patchy reionization at 4.5σ","feed_subtitle":"Forecast mocks show the large-scale Lyα power bump survives noise and systematics at z = 4.0–4.6.","key_machinery":"The mechanical core is the patchy reionization correction of Molaro et al. (2022, 2023), encoded as a dimensionless amplitude $A_p = (z-3.45)/2.55$ that multiplies the ratio of $P_{\\mathrm{Ly}\\alpha}$ from hybrid radiative-transfer reionization simulations to the uniform-background model. The correction is applied to mock sightlines drawn from the Sherwood-Relics simulations and processed by the WEAVEify instrument model, which adds resolution, noise, and continuum effects. The forecast covariance is built from bootstrap resampling for the statistical errors plus quadrature-added systematics for resolution, noise subtraction, uncorrelated metals, incomplete high-column-density absorber removal, and continuum shape; the detection statistic is the $\\Delta\\chi^2$ between fits with and without the four $A_p$ parameters.","core_discovery":"The central claim is that a WEAVE-QSO measurement of the 1D Lyman-$\\alpha$ forest power spectrum at $z \\geq 4$, with the covariance forecast in this paper, can recover the redshift-dependent amplitude of patchy reionization at $4.5\\sigma$ significance. The mock analysis finds $\\chi^2/\\mathrm{dof}$ improving from $65.28/32$ without the patchy correction to $36.07/28$ with one $A_p$ parameter per redshift bin. The expected amplitudes $A_p = 0.216, 0.294, 0.373, 0.451$ at $z=4.0, 4.2, 4.4, 4.6$ fall inside the recovered 68% credible intervals, and the recovery persists when the thermal-history priors are broadened. The largest-scale modes ($k \\lesssim 0.0033\\,\\mathrm{s\\,km^{-1}}$) carry most of the constraining power, and the lowest-wavenumber bin is where systematic uncertainties from high-column-density absorber damping wings and continuum placement dominate.","pith_inferences":["Inference: a non-detection at $z=4.0$ combined with a detection at $z=4.6$ would effectively measure the redshift at which relic thermal fluctuations fade, adding a reionization-timing observable complementary to 21-cm and CMB probes.","Inference: re-running the same pipeline with reionization models that differ in bubble size or duration would translate the $A_p$ posterior width into a constraint on reionization geometry rather than amplitude alone.","Inference: because high-column-density absorber completeness dominates the lowest-wavenumber systematics, independent spectroscopic calibration of the completeness function toward the same QSOs could directly shrink the error bar that sets the detection significance.","Inference: the forecast implies that survey strategy---exposure time, magnitude cut, and sky coverage---could be optimised for lowest-wavenumber sensitivity rather than total sightline count."],"forward_implications":["WEAVE-QSO should detect the large-scale relic imprint of patchy reionization at $z \\geq 4$ at roughly $4.5\\sigma$ if the forecast covariance is achieved.","Per-redshift-bin recovery of $A_p$ means the survey can characterise how the imprint fades from $z=4.6$ to $z=4.0$, not just detect it.","Because the signal lives at $k \\lesssim 0.0033\\,\\mathrm{s\\,km^{-1}}$, survey pipelines should concentrate systematic control on that single largest-scale bin.","If the lowest-wavenumber bin cannot be protected from damping-wing and continuum systematics, the patchy-reionization signature would not rise above about $1.2\\sigma$.","Other forthcoming massive spectroscopic surveys targeting $k \\sim 10^{-3}\\,\\mathrm{s\\,km^{-1}}$ at $z>4$ should be able to make the same measurement."],"supporting_citations":[{"why":"Supplies the $A_p$ parameterisation, the wavenumber binning, and the previous $2.7\\sigma$ tentative detection that this forecast is designed to supersede.","marker":"Molaro et al. (2023)"},{"why":"Provides the tabulated Sherwood-Relics patchy reionization correction to $P_{\\mathrm{Ly}\\alpha}$ that the mock signal is injected from.","marker":"Molaro et al. (2022)"},{"why":"Describes the Sherwood-Relics simulation suite from which the optical-depth skewers used to build the mocks are drawn.","marker":"Puchwein et al. (2023)"},{"why":"Gives the HCD completeness function used to model incomplete removal of high-column-density absorbers in the covariance.","marker":"Wang et al. (2022)"},{"why":"Provides the high-column-density absorber column density distribution used to inject synthetic absorbers into the mock spectra.","marker":"Prochaska et al. (2014)"},{"why":"Supplies the QSO luminosity function used to Monte Carlo sample the WEAVE-QSO source population and number counts.","marker":"Palanque-Delabrouille et al. (2016)"},{"why":"Sets the effective optical depth normalisation used for each mock sightline and for mean-transmission subtraction.","marker":"Becker et al. (2013)"},{"why":"Provides the SiIII correlated-metal model and the skewer-combination procedure used in the mocks and inference tests.","marker":"Ma et al. (2026)"},{"why":"Defines the wavenumber binning and the high-resolution data set that the forecast covariance and comparison are built around.","marker":"Karaçaylı et al. (2022)"},{"why":"Supplies the Monte Carlo method for injecting uncorrelated metal lines from column-density distribution functions into the mock spectra.","marker":"Viel et al. (2013)"}],"fun_headline_variants":["Patchy reionization detectable in WEAVE-QSO Lyα power at 4.5σ","Forecast: WEAVE-QSO Lyα forest will reveal patchy reionization at 4.5σ","WEAVE-QSO Lyα power spectrum forecasts 4.5σ detection of patchy reionization","Patchy reionization's large-scale imprint visible in WEAVE-QSO mocks at 4.5σ","Reionization patchiness detectable in WEAVE-QSO Lyα forest at 4.5σ"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forecast significance assumes that the systematic uncertainty in the lowest-wavenumber bin---dominated by damping wings from high-column-density absorbers and by continuum placement---is correctly modeled; if the true completeness of absorber removal or the continuum errors differ from the adopted models, the $4.5\\sigma$ preference would drop toward the $1.2\\sigma$ the paper finds when that bin is omitted.","fun_headline_variants_meta":{"raw":{"variants":["Patchy reionization detectable in WEAVE-QSO Lyα power at 4.5σ","Forecast: WEAVE-QSO Lyα forest will reveal patchy reionization at 4.5σ","WEAVE-QSO Lyα power spectrum forecasts 4.5σ detection of patchy reionization","Patchy reionization's large-scale imprint visible in WEAVE-QSO mocks at 4.5σ","Reionization patchiness detectable in WEAVE-QSO Lyα forest at 4.5σ"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000801,"raw_usage":{"total_tokens":3573,"prompt_tokens":1049,"completion_tokens":2524,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":665,"completion_tokens_details":{"reasoning_tokens":2391}},"tokens_in":665,"tokens_out":2524,"duration_ms":16157,"temperature":1.0,"reasoning_tokens":2391,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:35:51.072532+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the real WEAVE-QSO measurement in the lowest-wavenumber bin ($k \\sim 10^{-3}\\,\\mathrm{s\\,km^{-1}}$) against the forecast covariance: if the scatter of that bin across independent line-of-sight subsamples exceeds the forecast systematic error, or if the recovered $A_p$ posterior at $z=4.6$ does not include 0.45 within the 95% credible interval, the central claim is falsified. A direct null test is to fit the observed $P_{\\mathrm{Ly}\\alpha}$ at $z=4$ with and without $A_p$ and check whether $\\Delta\\chi^2$ remains above about 25.","supporting_citations":[],"review_version":1}