{"id":"22e07087-4e7f-4a34-baea-137e63383c1f","arxiv_id":"2512.11353","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A Fisher-matrix forecast says galaxy-21cm cross-correlation can measure Ω_HI at z~1 to sub-percent and improve growth constraints by ~2x, but the forecast's precision appears overstated.","lead":"This paper forecasts how combining a galaxy survey like DESI with a 21-cm hydrogen survey could measure cosmic growth and the amount of neutral hydrogen. It claims the combination could pin down the hydrogen density to sub-percent accuracy, far better than today's measurements.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fisher forecast uses full DESI volume (V=8.4 h^-3 Gpc^3) while the 21-cm survey has fsky≈0.014; with the actual overlap volume the headline xHI error grows by ~4x, from 0.32% to ~1.3%, so the sub-percent claim is unsupported.","rationale":"The single most load-bearing point is exactly the survey-volume mismatch. The paper's headline quantitative claim is the sub-percent xHI forecast; that forecast is computed from Np = V_survey/(...) in Eq. (34). The paper never states which V_survey enters, but the only z=1.0–1.2 volume given is the DESI full volume, while §2.3 fixes the 21-cm survey to fsky≈0.014. The overlap volume is ~0.5 h^-3 Gpc^3, a factor ~16 smaller, so all constraints widen by a factor ~4. This directly moves Table 2's 0.32% to ~1.3%, above the 'sub-percent' threshold. Under the same correction, the Nfore=10^4 case becomes ~10%, and the σ8/fσ8 gains would also shrink. The reader's REJECT is therefore appropriate. I do not see a need to rely on more speculative objections (optimistic Tsys, exact N-body template): the volume issue is an internal inconsistency and sufficient by itself. A revised forecast with the correct overlap volume could be conditionally acceptable; as written, the central claim is unsupported.","tokens_in":21547,"tokens_out":8300,"duration_ms":79590,"concrete_test":"Rerun the Fisher analysis of §2.5/§3 exactly as in the paper, but replace Np in Eq. (34) with the overlap volume V_overlap = (fsky_21cm / fsky_DESI) × 8.4 h^-3 Gpc^3 ≈ 0.5 h^-3 Gpc^3 for the z=1.0–1.2 bin, using the same k, μ bins, fiducial spectra, and Fisher parameter list. If the no-foreground σ(xHI)/xHI rises above 1% (as expected, ≈1.3%), then the paper's central sub-percent claim is invalidated; report the corrected Table 2 rows and the fσ8/σ8 errors. To be fully decisive, also recompute Np directly from the overlap footprint's fundamental-mode grid to confirm the factor ~16 volume suppression.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Table 2; §3.3) that the joint gg+gH+HH power spectra measure xHI to 0.32% depends on the mode count Np in Eq. (34). The paper uses one V_survey per k-µ bin, and the only volume quoted for the z=1.0–1.2 ELG bin is V=8.4 h^-3 Gpc^3 (Table 1). But the 21-cm survey specified in §2.3 is a 24°×24° patch, fsky≈0.014; its comoving volume at z=1.0–1.2 is ≈0.5 h^-3 Gpc^3, i.e. ~16 times smaller than the full DESI volume. Because Eq. (31) is explicitly the covariance of gg, gH, and HH measured from the same Fourier modes, all three spectra must be evaluated in the overlap volume; if the patch lies inside DESI, this is the 21-cm volume. Replacing V_survey with V_overlap reduces Np by ~16 and inflates every Fisher error by ~4. Table 2's no-foreground xHI error becomes ~1.3% (not 0.32%), and the Nfore=10^4 entry becomes ~10% (not 2.6%). The claimed 'sub-percent' xHI measurement and the stated σ8/fσ8 precision therefore do not follow from the specified experimental setup. This is an internal inconsistency between §2.3 and Eq. (34), not merely a conservative-versus-optimistic modeling choice. Appendix C's analytic sanity check verifies the algebra of the xHI Fisher element but keeps the same inflated Np.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a Fisher-matrix forecast for combining DESI ELG galaxy clustering auto-power, 21-cm line-intensity mapping auto-power, and their cross-power at z = 1.0–1.2. The model includes a nonlinear RSD treatment with Finger-of-God damping, higher-order bias terms, and the Alcock–Paczynski effect. The central claims are that the joint analysis constrains the neutral hydrogen fraction xHI (and hence ΩHI) to sub-percent accuracy—σ(xHI)/xHI = 0.32% in the no-foreground case—while improving growth constraints by roughly a factor of two over galaxy clustering alone. Appendix C presents an analytic reduction of the xHI Fisher element as a sanity check.","tokens_in":21994,"tokens_out":10184,"duration_ms":103585,"significance":"If the forecast were correct as stated, the claimed sub-percent ΩHI measurement at z ≈ 1 would be a substantial advance over current stacked and clustering constraints, which have order-unity uncertainties, and would demonstrate a practical route to breaking the ΩHI–bHI degeneracy. The two-tracer extension of the nonlinear RSD formalism is nontrivial, and the analytic Fisher reduction in Appendix C is a useful internal consistency check. However, the headline quantitative claim depends on a survey-volume assignment that appears internally inconsistent with the stated 21-cm survey geometry. The proposed methodology is plausible, but the numerical results presented in Table 2 and the abstract do not follow from the experimental setup described in §2.3.","major_comments":[{"comment":"The central forecast uses inconsistent survey volumes. Section 2.3 specifies the 21-cm survey as a roughly 24°×24° patch with fsky≈0.014, corresponding to a comoving volume of about 0.3–0.5 h^-3 Gpc^3 at z = 1.0–1.2. The Fisher mode count in Eq. (34) is instead evaluated with V_survey = 8.4 h^-3 Gpc^3, the full DESI ELG volume quoted in Table 1. Since Eq. (31) is the covariance of Pgg, PgH, and PHH for the same Fourier modes, all three spectra should use the overlap volume, or the non-overlapping DESI-only modes must be separated block-diagonally. The present implementation overcounts modes by a factor of roughly 15–25, so all Fisher errors are underestimated by a factor of about 4–5. In particular, the headline σ(xHI)/xHI = 0.32% in Table 2 becomes approximately 1.3–1.6% in the no-foreground case, and the Nfore = 10^4 entry becomes approximately 10–13%, not 2.6%. The abstract’s sub-perc","section":"§2.3, Eq. (34), Table 1, Table 2, §3.3"},{"comment":"If the 21-cm patch is embedded inside the DESI footprint, the galaxy auto-spectrum can in principle use the full DESI volume, while the gH and HH spectra are limited to the overlap volume and the gg auto-spectrum in the overlap region has a different effective mode count. The current single-Np implementation does not specify which of these cases is intended and therefore cannot be interpreted as either an optimistic or conservative forecast. A corrected Fisher matrix must either set V_survey to the overlap volume for the full 3×3 block or construct a block-diagonal covariance that treats overlap and non-overlap DESI modes separately. This is a necessary correction regardless of the final numerical values.","section":"§2.5, Eq. (31)"},{"comment":"The claimed breaking of the bHI–xHI degeneracy relies on the nonlinear RSD terms A, B, T, and F, which are calibrated from the authors’ previous N-body work (Zheng & Song 2016; Zheng et al. 2019). The paper does not provide an independent validation of these templates for a HI tracer at z ≈ 1, nor does it propagate their calibration uncertainty into the Fisher forecast. Because the factor-of-several improvement over the HI auto-spectrum comes precisely from these template terms, the robustness of Table 2 should be tested, e.g. by varying the template amplitudes by their reported calibration uncertainty or by comparing with an independent HI simulation. This is a concrete correctness-risk concern and should be addressed even after the volume issue is corrected.","section":"Appendix B, §4.2"}],"minor_comments":[{"comment":"The column label “galaxy-21cm” is ambiguous. It should state explicitly whether the column is the full joint analysis (gg + gH + HH) or only the cross-spectrum contribution; the text sometimes refers to the joint analysis as “cross+auto.”","section":"Table 2, §3.3"},{"comment":"Typographical issue: “xHI parameter is most degenerate with galaxy bias bHI” should read “HI bias bHI,” not galaxy bias.","section":"§4.2"},{"comment":"The integral upper limit is written as MHI, but the intended upper limit is presumably the maximum hydrogen clump mass (or infinity). Please clarify the notation.","section":"Eq. (36)"},{"comment":"The abstract states that the cosmic-expansion constraint is “slightly improved,” while §3.2 concludes that DA is “not much improved” and H^-1 is actually more difficult to constrain. The wording should be harmonized to avoid overstating the AP benefit.","section":"Abstract, §3.2"}],"recommendation":"reject","confidential_remarks":"The volume mismatch is the decisive issue: it invalidates the headline sub-percent ΩHI claim as presented. A corrected calculation may still support a percent-level forecast, but the manuscript’s central quantitative conclusion would change. If the authors intended the 21-cm survey to cover the same sky as DESI, then §2.3 must be revised and the instrument-noise calculation rederived; if they intended the small patch, all forecasts must be recomputed with the overlap volume. I would encourage the authors to resubmit after making this explicit and rerunning the Fisher analysis, ideally with a robustness test of the N-body-calibrated RSD templates."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper's headline number—sub-percent ΩHI from DESI+SKA-MID—doesn't survive a check of the survey volumes. The 21-cm survey is specified as a 24x24 degree patch (fsky≈0.014), but the Fisher matrix in Eq. (34) uses V_survey = 8.4 h⁻³ Gpc³, the full DESI ELG volume. Those can't both be right. With the actual overlap volume, the xHI error goes from 0.32% to about 1.3%, and the sub-percent claim is gone. This is the thing to know.\n\nThe paper is not a nothing-burger. It extends the authors' hybrid RSD formalism to the double-tracer case (gg, gH, HH), lays out the covariance explicitly, and includes an analytic check in Appendix C that verifies the Fisher algebra for xHI. The idea that nonlinear RSD anisotropy can break the ΩHI–bHI degeneracy is physically reasonable and worth pursuing. The treatment of foreground residuals is also honest—they vary Nfore over a plausible range.\n\nThe volume issue is load-bearing, not a minor quibble. It comes from using the same N_p for all three spectra when the cross-spectrum can only be measured where both surveys overlap. The 21-cm patch is about 1/16 of the DESI volume, so all three errors scale by roughly a factor of four. Appendix C's sanity check keeps the same inflated N_p, so it doesn't catch the problem. Also, the nonlinear template comes from prior N-body calibrations and is treated as exact; the forecast doesn't propagate any uncertainty from that. That's not circular—xHI isn't an input to the calibration—but it does mean the quoted precision is conditional on the template being perfect.\n\nWho should read it: anyone doing galaxy–21cm cross-correlation forecasts, especially with SKA-MID. The formalism and the degeneracy-breaking mechanism are genuinely useful. But the quantitative claims need to be redone with the correct overlap volume. I'd send it to review—a referee can force the fix—and I'd want to see the revised numbers before believing the sub-percent story.","headline":"The double-tracer Fisher formalism is a real step forward, but the sub-percent ΩHI claim is built on using the full DESI volume for a 1.4% sky 21-cm patch; with the true overlap volume the headline error quadruples and the paper needs revision.","tokens_in":22517,"tokens_out":2977,"would_cite":false,"duration_ms":29863,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["98.80.-k"],"model":"deepseek-v4-flash","headline":"A joint analysis of galaxy clustering and 21-cm intensity mapping can pin the neutral hydrogen fraction to sub-percent accuracy, breaking the long-standing degeneracy with hydrogen bias.","keywords":["21-cm intensity mapping","galaxy clustering","cross-correlation power spectrum","neutral hydrogen fraction","redshift-space distortions","structure growth","cosmological distances","forecast covariance"],"falsifier":"Recompute the information-matrix forecast using the true overlap volume of the 21-cm survey (fsky≈0.014 at z≈1) instead of the full galaxy-survey volume; if σ(xHI)/xHI exceeds 1% in the no-foreground case, the headline sub-percent claim fails. A second check: run the identical forecast on mock catalogs with a known input xHI and see whether the 1σ scatter matches the predicted 0.32% error.","tokens_in":21405,"feed_emoji":"📡","tokens_out":8386,"duration_ms":80684,"temperature":0.7,"pith_summary":"The paper argues that adding a 21-cm intensity-mapping survey to a spectroscopic galaxy survey, and analyzing the galaxy auto-spectrum, the 21-cm auto-spectrum, and their cross-spectrum together, can measure the neutral hydrogen fraction xHI (equivalently ΩHI) to sub-percent precision at z≈1, far better than current determinations. The decisive move is that the cross-correlation, modeled through nonlinear redshift-space distortions, breaks the degeneracy between xHI and the hydrogen bias bHI that has limited previous power-spectrum analyses. If the forecast holds, the same data would also constrain the growth-rate combination fσ8 to about 1% and the amplitude σ8 to about 5%, twice as tightly on fσ8 as galaxy clustering alone. This matters because it offers a practical route to probe post-reionization astrophysics and to separate cosmic expansion from structure growth without relying on small-scale perturbation theory. The results come from an information-matrix forecast, so their precision depends on how much sky the 21-cm survey actually overlaps with the galaxy survey.","feed_headline":"Galaxy–21-cm cross-correlation maps cosmic hydrogen to 0.32%","feed_subtitle":"Combining galaxy clustering with 21-cm maps breaks the hydrogen-bias degeneracy and sharpens growth-of-structure forecasts.","key_machinery":"The machinery is a two-tracer extension of a nonlinear redshift-space distortion model. The anisotropic power spectra Pab(k, μ) are written as the sum of a perturbative part—including density–density, density–velocity-divergence, and velocity-divergence auto-spectra—plus a non-perturbative Finger-of-God damping factor, with higher-order correction terms calibrated in a fiducial cosmology and rescaled by growth functions. An information matrix then combines the three spectra with their full joint covariance, including galaxy shot noise, 21-cm shot noise, instrument noise, and foreground noise. This structure allows xHI, a multiplicative amplitude, to be separated from bHI, which changes the μ","core_discovery":"The central claim is that the three power spectra measured jointly—galaxy auto-correlation, 21-cm auto-correlation, and the galaxy–21-cm cross-correlation—contain enough anisotropic information in the quasi-linear regime to determine xHI to 0.32% in the no-foreground case and to 2.6% even with strong foreground contamination. Because xHI enters the spectra as a pure overall multiplier while bHI modulates the angular dependence through the line-of-sight anisotropy of redshift-space distortions, the joint fit separates the two and breaks the ΩHI–bHI degeneracy that limits linear-regime analyses. The same joint fit improves the coherent-velocity constraint fσ8 by a factor of about two relative","pith_inferences":["The quoted sub-percent xHI error appears to use the full galaxy-survey volume for all three spectra; if only the overlapping sky area of the 21-cm survey (about 1.4% of the sky at z≈1) enters the calculation, the errors grow by roughly a factor of four, and the no-foreground constraint may rise above 1%. Verification of the overlap volume is the first thing to check.","Because xHI is a pure amplitude, the same cross-correlation technique should extend to other line-intensity tracers at higher redshift, such as CO or [CII], to break analogous amplitude–bias degeneracies wherever a galaxy sample provides the cross-correlation.","The forecast's realism hinges on how many large-scale line-of-sight modes survive foreground removal; the paper parametrizes this with a foreground noise term, but real spectrally smooth foregrounds may remove a broader set of modes than a simple noise term captures."],"forward_implications":["If the forecast is correct, the neutral hydrogen fraction xHI at z≈1 can be measured to roughly 0.3% with clean foreground subtraction, and to a few percent even with residual foregrounds—far beyond current stacked-emission constraints.","The long-standing degeneracy between ΩHI and the hydrogen bias bHI, which has limited clustering-based estimates, would be broken by the anisotropic quasi-linear spectra.","The growth-rate combination fσ8 would be constrained to about 1% and σ8 to about 5% from a single galaxy-plus-21-cm program, roughly doubling the fσ8 precision available from galaxy clustering alone.","Distance constraints tied to geometric distortion would see little improvement from adding the 21-cm data, so this method complements rather than replaces galaxy BAO analyses.","With finer redshift bins, the same analysis could map xHI(z) across z≈0–2 at roughly percent-level accuracy per bin."],"fun_headline_variants":["Galaxy–21-cm cross-correlation pins HI to 0.32%","Joint galaxy-21cm analysis doubles growth constraints, measures HI to 0.32%","Cross-correlating galaxies and 21-cm breaks HI-bias degeneracy","Galaxy-21cm cross-correlation: 0.32% HI precision","0.32% HI from galaxy–21-cm cross-correlation, twofold growth boost"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing assumption is that the full galaxy-survey volume is available to the combined three-spectrum measurement, even though the 21-cm survey covers only a small sky patch; if the actual overlapping volume is used, the quoted sub-percent xHI precision weakens considerably.","fun_headline_variants_meta":{"raw":{"variants":["Galaxy–21-cm cross-correlation pins HI to 0.32%","Joint galaxy-21cm analysis doubles growth constraints, measures HI to 0.32%","Cross-correlating galaxies and 21-cm breaks HI-bias degeneracy","Galaxy-21cm cross-correlation: 0.32% HI precision","0.32% HI from galaxy–21-cm cross-correlation, twofold growth boost"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001523,"raw_usage":{"total_tokens":5971,"prompt_tokens":811,"completion_tokens":5160,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":5049}},"tokens_in":555,"tokens_out":5160,"duration_ms":36115,"temperature":1.0,"reasoning_tokens":5049,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T16:53:49.698617+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the information-matrix forecast using the true overlap volume of the 21-cm survey (fsky≈0.014 at z≈1) instead of the full galaxy-survey volume; if σ(xHI)/xHI exceeds 1% in the no-foreground case, the headline sub-percent claim fails. A second check: run the identical forecast on mock catalogs with a known input xHI and see whether the 1σ scatter matches the predicted 0.32% error.","supporting_citations":[],"review_version":1}