{"id":"6b74acd3-2045-448a-be9e-f37393ba3efa","arxiv_id":"2511.15072","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A two-channel UNet combining frequency differencing and PCA preprocessing recovers the 21-cm HI power spectrum at large scales under realistic beam effects, improving cross-correlation by 5-8% over single-channel baselines.","lead":"This paper trains a U-shaped neural network on two preprocessed views of simulated 21-cm sky maps — a frequency-differenced map and a principal-component-cleaned map — to recover the cosmological hydrogen signal buried under bright radio foregrounds and telescope beam distortions. The two-channel network keeps the large-scale reconstructed power spectrum closer to unity under a realistic MeerKAT-like beam, a 5-8% gain over either preprocessing alone.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's mismatched-beam robustness claim is unsupported: no train/test beam-mismatch experiment exists, and the FD channel is explicitly beam-dependent (Eq. 8).","rationale":"The reader's weakest_assumption correctly identifies that the evaluation is strictly in-distribution: training and test samples share the same CRIME simulation pipeline, the same foreground model, and the same beam model. The abstract's explicit claim of robustness to imperfect or mismatched beams is therefore unsupported by the supplied body. I agree with this as the primary load-bearing concern. I would add a concrete technical reason why the concern is more than a generic out-of-distribution worry: the FD preprocessing step in Eq. 8 requires the assumed beam FWHM to match adjacent-frequency maps to a common angular resolution. A mismatched beam would change the statistics of the FD input channel in a way the network has not seen, so the claimed robustness is not a harmless overstatement but a potentially broken assumption. The body's central result — hybrid outperforms single-channel variants under a matched Cosine beam — is plausible and internally consistent, but it does not establish the abstract's broader conclusion. The appropriate verdict remains CONDITIONAL: the paper's main contribution is promising, but the advertised generalization claim must be either demonstrated or removed. No change to the reader's verdict is needed, hence UNCHANGED.","tokens_in":20754,"tokens_out":3607,"duration_ms":41444,"concrete_test":"Add a mismatched-beam experiment using the existing CRIME pipeline: train FD+UNet, PCA+UNet, and Hybrid+UNet on Gaussian-beam-convolved cubes, then test on Cosine-beam-convolved cubes, and the reverse. Also include a perturbed-beam variant (e.g., Cosine beam parameters with 10–20% FWHM or sidelobe-amplitude errors) to directly probe the sensitivity of Eq. 8. If the hybrid large-scale R_cross(k<0.1 h/Mpc) deviates from unity by more than ~1σ, or degrades by more than the 5–8% margin claimed, the abstract's robustness claim fails; if it remains consistent, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central advertised conclusion — that the method 'can robustly recover the HI signal even when the beam model is imperfect and differs between training and testing' — is not backed by any experiment in Sections IV or V. All reported results train and test on the same beam model (Gaussian or Cosine). This is not merely a missing generality: the FD preprocessing itself relies on knowing the beam. Section III.B.2 states that adjacent-frequency maps are smoothed to a common angular resolution using Δθ_FWHM computed from the assumed beam FWHM (Eq. 8). If the test-time beam differs from the training beam, the frequency-differencing channel will contain residual chromatic beam structure with statistics the network never saw, and PCA cleaning will also remove a different number/pattern of modes. Therefore the robustness claim is not just an untested extrapolation; the method's input construction is coupled to the assumed beam model. The in-body matched-beam comparison (Figure 16) may support the relative improvement of the hybrid method under a Cosine beam, but it cannot support the abstract's generalization statement. The paper should either add an explicit mismatched-beam experiment or soften the abstract and conclusions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a deep-learning foreground and beam mitigation pipeline for 21-cm intensity mapping, comparing three preprocessing strategies feeding a UNet: frequency differencing (FD), principal component analysis (PCA), and a hybrid two-channel combination. Using CRIME simulations of HI plus foregrounds, with optional Gaussian or Cosine beam convolution, the authors report that under the Cosine beam the single-channel FD+UNet and PCA+UNet underestimate the cross-correlation power spectrum by roughly 5–8% at k < 0.1 h Mpc^{-1} and by more than 20% at k ~ 0.2 h Mpc^{-1}, while the hybrid method remains consistent with unity within 1σ at large scales. The paper also claims in the abstract that the method is robust to imperfect or mismatched beams between training and testing, but this claim is not supported by any experiment in the body.","tokens_in":20982,"tokens_out":2950,"duration_ms":32695,"significance":"If the in-body results are correct, the hybrid two-channel preprocessing is a useful, practical contribution: it combines the large-scale fidelity of FD with the small-scale fidelity of PCA, and the improvement under a realistic Cosine beam is directly relevant to ongoing MeerKAT-era intensity mapping analyses. The paper is honest in its in-body comparison and reports concrete power-spectrum ratios with error bars. However, the advertised headline claim of robustness to train/test beam mismatch is absent from the experiments, and the evaluation is entirely in-distribution with respect to the simulation pipeline, so the significance of the paper as written is lower than the abstract suggests.","major_comments":[{"comment":"The abstract states that the method 'can robustly recover the HI signal even when the beam model is imperfect and differs between training and testing,' but no such experiment appears anywhere in Sections IV or V. All reported results train and test on the same beam model (no beam, Gaussian, or Cosine) in matched conditions. This is a load-bearing overclaim: the central advertised generalization is unsupported. The authors should either add an explicit mismatched-beam experiment (e.g., train on Gaussian, test on Cosine, or perturb the Cosine beam parameters between train and test) or soften the abstract and conclusions to describe matched-beam robustness only.","section":"Abstract and Sections IV–V"},{"comment":"The mismatched-beam claim is not merely an untested extrapolation; the FD preprocessing itself depends on the assumed beam. Eq. (8) requires smoothing adjacent-frequency maps to a common resolution using Δθ_FWHM computed from the beam FWHM. If the test-time beam differs from the training beam, the FD channel will contain residual chromatic structure with statistics the network has not seen, and the PCA channel will also change. Thus the advertised robustness is coupled to a beam-dependent input construction. At minimum, the paper should quantify the sensitivity of the hybrid method to beam-model mismatch, or explicitly restrict the claim to the matched-beam case.","section":"Section III.B.2, Eq. (8)"},{"comment":"The evaluation is entirely in-distribution with respect to the simulation pipeline: training and test samples come from the same CRIME lognormal HI fields, the same Haslam-based foreground model, and the same beam models. The held-out test set measures generalization across random realizations of the same process, not robustness to different foreground morphology, spectral-index variations, or beam systematics. The numbers of PCA modes (fixed at 3) and the FD spacing (fixed at 1 MHz) are free parameters, but no sensitivity analysis is provided. These choices may affect the magnitude of the hybrid improvement. The paper should at least discuss this limitation and, ideally, vary the number of PCA modes or the FD spacing in a robustness check.","section":"Section IV and Section III.B"}],"minor_comments":[{"comment":"The caption of Figure 12 says 'under Cosine beam convolution,' but the surrounding text (Section IV.B) and the third panel describe the Gaussian beam case; Figure 16 is the Cosine-beam counterpart. The caption should be corrected.","section":"Figure 12 caption"},{"comment":"The text 'which is already visible at large scales (k < hMpc^{-1})' appears to be missing a factor of 0.1; it should read k < 0.1 h Mpc^{-1} for consistency with the rest of the paper.","section":"Section IV.C"},{"comment":"The caption says 'correlation coefficient (left panel)' but the correlation coefficient is the right panel; the left panel is the temperature distribution.","section":"Figure 11 caption"},{"comment":"The phrase 'Followingdeep21' should cite the deep21 paper (reference [51]) explicitly, as it currently appears as a plain text mention without a citation.","section":"Section III.B.1"},{"comment":"The polynomial coefficients in Eq. (6) are given with a very wide dynamic range; a table or scientific-notation formatting would improve readability.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The in-body comparison is careful and the hybrid improvement under the Cosine beam is a genuine contribution, but the abstract overstates the robustness claim in a way that would mislead readers. Adding a mismatched-beam experiment or softening the claim is essential. The paper is otherwise a reasonable incremental advance over Paper I and related deep21 work, though the absence of any sensitivity analysis for the PCA-mode number and FD spacing should also be addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The two-channel hybrid does deliver the claimed in-body result: under the cosine beam, the cross-correlation power ratio stays consistent with unity at k<0.1 h/Mpc, while FD+UNet and PCA+UNet each drop 5-8% there and >20% at k~0.2 h/Mpc. That is a real, directly usable improvement for 21-cm intensity mapping, and the comparison across no-beam, Gaussian, and cosine-beam conditions is honest and reasonably thorough.\n\nWhat's new is limited but real: FD+UNet and PCA+UNet separately are prior art (the authors' own Paper I and deep21/Ni et al.), and the novel element is feeding both preprocessed cubes as two channels to one UNet plus the systematic large-scale comparison under the cosine beam. That's a legitimate extension.\n\nThe main problem is the abstract. It claims robust recovery when the beam model is imperfect and differs between training and testing, but there is no mismatched-beam experiment anywhere in Sections IV or V. All tests train and test on the same beam model. This is not a minor omission, because the FD input is constructed with a beam-dependent smoothing step (Eq. 8) — if the test-time beam differs, the FD channel has chromatic structure the network never saw, and PCA mode count shifts too. The authors either need to run the mismatched-beam test or soften the abstract and conclusions substantially.\n\nOther soft spots are minor: no ablation of the two channels, no sensitivity to the number of PCA modes, and no code or data release. Those are addressable and not fatal. The reader's in-distribution worry is fair but applies to most simulation-based ML work in this area; here it is less an independent flaw than the same missing robustness test.\n\nWho gets value: 21-cm IM people working on foreground cleaning and beam mitigation, especially for BAO-scale measurements with MeerKAT/SKA. It builds on the authors' own line of work and cites deep21 correctly.\n\nRecommendation: send it to peer review. The in-body result is solid and the overclaim is fixable by experiment or by revising the abstract. I'd expect a moderate revision, not a reject.","headline":"The hybrid FD+PCA UNet is a real improvement under the cosine beam, but the abstract's mismatched-beam robustness claim has no experiment behind it.","tokens_in":21515,"tokens_out":3223,"would_cite":true,"duration_ms":29573,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-channel deep network — fed both frequency-differenced and PCA-cleaned maps — keeps the 21-cm cross-power spectrum unbiased at large scales under a realistic cosine beam, where either preprocessing alone loses 5–8% or more.","keywords":["21-cm cosmology","intensity mapping","foreground removal","beam effects","deep learning","U-Net","frequency differencing","PCA cleaning"],"falsifier":"Train the two-channel UNet on cosine-beam maps and test it on maps made with a different beam (e.g., a cosine beam with altered ripple amplitude/period, or a beam from real holographic measurements). If the large-scale cross-correlation ratio departs from unity beyond the 1σ band in that mismatched setting, the paper's robustness claim fails. A cheap version is to train on Gaussian-beam maps and test on cosine-beam maps; the abstract would predict near-unity recovery on large scales, but the body currently contains no such test.","tokens_in":20619,"feed_emoji":"📡","tokens_out":6182,"duration_ms":54668,"temperature":0.7,"pith_summary":"The paper sets out to show that the two standard ways of suppressing smooth foregrounds in 21-cm intensity mapping — spectral differencing between adjacent frequency channels, and subtracting the leading principal components of the data cube — preserve different parts of the cosmological signal, and that a U-shaped convolutional network fed both as separate input channels outperforms a network fed either one. In simulations with realistic foregrounds and a frequency-dependent cosine beam, each single-channel network underestimates the cross-correlation power spectrum by 5–8% on large scales (k < 0.1 h Mpc^-1) and by more than 20% near k = 0.2 h Mpc^-1, while the two-channel network stays consistent with unity within 1σ. The payoff, if correct, is that large-scale modes needed for baryon acoustic oscillation measurements survive a realistic chromatic beam without the signal loss that aggressive linear cleaning causes.","feed_headline":"Hybrid UNet recovers unbiased 21-cm large-scale power","feed_subtitle":"Feeding a network both frequency-differenced and PCA-cleaned maps beats either alone under realistic beams.","key_machinery":"The mechanism is a two-channel input cube for a 13-layer UNet: channel one is the frequency-differenced cube (adjacent 1 MHz channels subtracted after smoothing the higher-frequency map to the lower-frequency resolution), channel two is the cube after subtracting three principal components along frequency. The contrast between the channels — FD retaining diffuse large-scale structure, PCA retaining compact small-scale structure — is what gives the network the information to avoid the 5–8% large-scale bias each channel alone would bias it toward.","core_discovery":"The central claim is that frequency differencing and PCA are not competing preprocessing options but complementary views: differencing preserves diffuse large-scale emission but adds striping artifacts at bright pixels, while PCA preserves compact bright structures but subtracts large-scale modes. Under the cosine beam, the UNet trained on either view alone systematically suppresses the recovered cross-power spectrum on large scales; the hybrid two-channel network does not, and the paper attributes this to the network learning to draw on FD's large-scale fidelity and PCA's small-scale fidelity simultaneously. The result is stated as a bias-free large-scale HI reconstruction, with the caveat","pith_inferences":["A natural reading is that any pair of spectral-smoothness filters with opposite scale biases could be combined this way; ICA and SVD variants are obvious candidates, though the paper does not test them.","The abstract claims robustness to imperfect or mismatched beams, but the body only reports matched training/test conditions; this is the paper's gap to close, and it is directly testable.","The fact that the hybrid gain appears only under the cosine beam hints that the network is using the FD channel as a 'large-scale anchor' that resists the sidelobe-induced mode mixing; that interpretation could be probed by ablating the FD channel at selected scales.","The fixed 3-component PCA subtraction and fixed 1 MHz differencing are hyperparameters of the preprocessing; varying them in tandem with the loss function might push the residual large-scale bias even closer to zero, but no such exploration is reported."],"forward_implications":["Large-scale HI power, the part most affected by beam-induced spectral structure, can be recovered without the mode-subtraction cost of PCA-only cleaning.","The hybrid advantage is specific to the realistic cosine beam: with a Gaussian beam all three U-Net variants are comparable, so chromatic sidelobe structure is the regime where the two-channel design matters.","Because the network is trained on matched beam models in this work, the gain at k<0.1 h Mpc^-1 should be re-checked when the beam model is varied between training and inference.","If this transfers to real data, it would reduce one systematic in 21-cm auto-power measurements, which currently depend heavily on cross-correlations with galaxy surveys.","The 5–8% improvement at fixed network size and training cost suggests that other blind cleaning outputs, not just PCA, could be combined in the same two-channel way."],"fun_headline_variants":["Hybrid FD+PCA UNet recovers unbiased 21-cm large-scale power","Two-channel deep learning improves 21-cm HI recovery over single views","Beam-robust HI signal via hybrid frequency differencing and PCA","Complementary preprocessing yields 5-8% better 21-cm large-scale recovery","UNet leverages FD and PCA to preserve large-scale 21-cm signal"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire evaluation is in-distribution: training and test cubes come from the same simulation pipeline with the same foreground model and the same beam models, so the claimed accuracy and the abstract's assertion of robustness to imperfect beams have not been tested against genuinely different beam or foreground conditions.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid FD+PCA UNet recovers unbiased 21-cm large-scale power","Two-channel deep learning improves 21-cm HI recovery over single views","Beam-robust HI signal via hybrid frequency differencing and PCA","Complementary preprocessing yields 5-8% better 21-cm large-scale recovery","UNet leverages FD and PCA to preserve large-scale 21-cm signal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00118,"raw_usage":{"total_tokens":4693,"prompt_tokens":708,"completion_tokens":3985,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":452,"completion_tokens_details":{"reasoning_tokens":3882}},"tokens_in":452,"tokens_out":3985,"duration_ms":30286,"temperature":1.0,"reasoning_tokens":3882,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T21:26:48.516048+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the two-channel UNet on cosine-beam maps and test it on maps made with a different beam (e.g., a cosine beam with altered ripple amplitude/period, or a beam from real holographic measurements). If the large-scale cross-correlation ratio departs from unity beyond the 1σ band in that mismatched setting, the paper's robustness claim fails. A cheap version is to train on Gaussian-beam maps and test on cosine-beam maps; the abstract would predict near-unity recovery on large scales, but the body currently contains no such test.","supporting_citations":[],"review_version":1}