{"id":"440d75a6-f406-4d37-b1ff-a3d641cce858","arxiv_id":"2506.06386","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Restoring RFI-contaminated pixels with the LaMa inpainting network reduces post-foreground-removal RMS and improves recovery of the large-scale 21-cm power spectrum in simulations.","lead":"A deep neural network for image inpainting is used to fill radio-frequency interference contaminated pixels in 21-cm intensity mapping data. The authors report that restoration lowers residual noise after foreground removal and moves mock angular power spectra closer to the true 21-cm signal.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Training loss in Eqs. (1)-(5) appears to target the raw contaminated data x, so the network may learn to reconstruct RFI rather than remove it; the RMS improvements would then be spurious.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the loss functions in Section 2.2 are written against the full data x, which includes RFI in masked regions, with no explicit masking of the loss or substitution of a clean foreground target. This directly threatens the central claim that the network restores contaminated data, because the training objective could teach the network to reproduce the very contamination the pipeline is supposed to remove. The concern is not merely technical: it determines whether the reported RMS improvements and power-spectrum shifts are meaningful or artifacts of fitting the training target. The paper's own evaluation metric (Cm/Cu) and SSIM/PSNR use foregrounds from Li et al. as ground truth, but that ground truth is not stated to be used in training, so the training pipeline is either inconsistent or underspecified. A retraining test with a clean masked-region target would settle the issue. The reader's conditional verdict is appropriate given this unresolved ambiguity; I do not see a reason to move the verdict, so it remains UNCHANGED.","tokens_in":14493,"tokens_out":6138,"duration_ms":68920,"concrete_test":"Obtain the training code and check the target image used in Eqs. (1)-(5): specifically, inspect whether the target patch x is the raw data with RFI values in masked regions or a mask-weighted clean version. Then run a decisive ablation: retrain LaMa-dilated from scratch with the target in masked regions replaced by the foreground model from Li et al. (2021), or with the loss restricted to unmasked pixels, keeping all other hyperparameters identical. Recompute the RMS comparisons in Figures 5-7 and the angular power spectra in Figure 11. If the improvements vanish or reverse, the original training objective was fitting RFI and the central claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.2 defines the loss against x, which the text identifies as the input data, while the network input is the masked data x⊙m. In standard LaMa inpainting, x would be a clean ground-truth image; here, however, the data in masked regions is exactly the RFI-contaminated signal that the paper aims to remove. The paper nowhere states that the loss is restricted to unmasked pixels or that the masked-region target has been replaced by a foreground model. Equation (1) compares φ_HRF(x) and φ_HRF(ˆx) over the full image; Equations (3)-(4) use the raw data as ground truth. Only in Section 5.3 do the authors mention taking foregrounds from Li et al. as ground truth, and that is for SSIM/PSNR evaluation, not for training. If the target retains RFI, the network is trained to reproduce the contamination, and the lower RMS in Sections 4.1-4.3 would reflect the network's learned ability to regenerate RFI-like structures rather than genuine recovery of the underlying foreground+HI. The angular power spectrum test in Section 6 could still show improvement if the network happens to smooth the data, but the reported training objective undercuts the interpretation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using a LaMa-dilated inpainting neural network to restore RFI-masked pixels in 21-cm intensity mapping data from Parkes P913, and then evaluates the effect of this restoration on foreground removal by polynomial fitting, SVD, and ICA. The authors report lower RMS residuals after restoration (Sections 4.1-4.3), select LaMa-dilated via a Cm/Cu metric (Section 3.3), and report in a CRIME-based mock observation that angular power spectra are closer to the true HI signal (Section 6). The central claim is that DNN restoration before foreground removal improves data quality and preserves the large-scale HI signal.","tokens_in":14718,"tokens_out":6655,"duration_ms":73419,"significance":"If the pipeline works as claimed, restoring flagged pixels before foreground subtraction would be a practically useful pre-processing step for 21-cm IM surveys, and the use of real Parkes data with realistic RFI masks is a valuable feature. The paper includes an ablation of loss terms, patch sizes, and metrics, and reports a mock-observation test; these are appropriate steps and go beyond a purely qualitative demonstration. However, the current manuscript does not establish that the network learns to remove rather than reproduce RFI, and the evaluation metrics are tied to the same foreground model used to demonstrate improvement. The scientific conclusions therefore rest on assumptions that are not yet verified.","major_comments":[{"comment":"The loss functions are written against the full, contaminated data: Eq. (1) compares phi_HRF(x) and phi_HRF(x_hat) over the full image, and Eq. (3) uses a quantity called ground truth without stating that the masked regions are excluded from the loss or that the target in masked regions is a cleaned foreground model. Because the masked regions are precisely where RFI is present, a network trained with this objective can minimize the loss by reproducing the contamination rather than removing it. The RMS reductions in Sections 4.1-4.3 and the angular power spectrum shifts in Section 6 would then not establish genuine restoration of the foreground+HI signal. Please define the ground truth concretely, restrict the loss to unmasked pixels if that is the intention, and retrain and re-evaluate with the corrected objective before interpreting the downstream results.","section":"Sec. 2.2, Eqs. (1)-(5)"},{"comment":"The Cm/Cu metric used for model selection compares predicted data with \"foregrounds\", but the operational definition of Cm and Cu is never given, and Section 5.3 reveals that foregrounds extracted by the polynomial-fitting pipeline of Li et al. (2021) are used as ground truth in the SSIM/PSNR evaluation. If the same foreground model underlies Cm/Cu, then selecting LaMa-dilated because it best matches those polynomial-fitted foregrounds and then demonstrating improvement over polynomial fitting in Section 4.1 is circular. Please specify how Cm and Cu are computed and evaluate the restoration against a target that is independent of the foreground-removal method used as the baseline.","section":"Sec. 3.3 and Sec. 5.3"},{"comment":"The claim in Section 4 that \"a lower root-mean-square (RMS) level of data invariably means cleaner observational data with less contamination\" is not justified. Filling masked regions with smooth model predictions can lower the residual RMS after polynomial, SVD, or ICA subtraction even when the HI signal is distorted, so RMS alone cannot certify restoration quality. The mock-observation test in Section 6 is the right idea but currently lacks a quantitative measure of the distance to the fiducial C_l and, in the polynomial-fitting panel, shows a restored curve that lies above the HI curve rather than unambiguously matching it. Please add a signal-preservation diagnostic and quantify the agreement with the fiducial power spectrum, for example with error bars or a chi-square-like statistic.","section":"Sec. 4"},{"comment":"Figure 11 reports angular power spectra for a single CRIME realization without error bars, and the conclusion that the restored curves are \"closer\" to the HI fiducial is drawn from visual inspection in log space over several orders of magnitude. The simulation also omits observational noise and beam effects and downsamples the frequency axis by averaging 20 channels; these simplifications should be stated as limitations and their effect on the conclusion discussed. Without uncertainties, the central claim that restoration moves the measured C_l toward the true signal is not quantitatively supported.","section":"Sec. 6.4, Fig. 11"}],"minor_comments":[{"comment":"The sentence \"the channels with mean data values outside of the range ... e considered as the contaminated by RFI and are masked\" contains a typo: \"e considered\" should be \"are considered\", and \"the contaminated\" should be \"contaminated\".","section":"Sec. 3.2"},{"comment":"The optimization setup is underspecified: the optimizer, learning rate, learning-rate schedule, and convergence criterion are not reported. Please include these details or a reference to the exact LaMa configuration used.","section":"Sec. 3.2"},{"comment":"The relation between the notation (x, x_hat) in Eq. (1) and (I_hat, I_tilde) in Eqs. (3)-(4) is unclear; please define all symbols consistently.","section":"Sec. 2.2"},{"comment":"The caption states that the panels show M128, M256, and M256; the third panel should presumably be M512.","section":"Fig. 9 caption"},{"comment":"The label \"polinomial fitting\" in Figure 11 should be \"polynomial fitting\".","section":"Sec. 6.4, Fig. 11"}],"recommendation":"major_revision","confidential_remarks":"The central training-loss ambiguity is the main risk: if the authors can demonstrate that the loss is evaluated only on unmasked pixels with a clean target, and that the Cm/Cu metric is not circular, the paper merits serious consideration. The work uses a real observation and a standard inpainting architecture, and the ablation study is useful. I recommend major revision rather than rejection because the issues are, in principle, addressable by rerunning the pipeline with a corrected objective and an independent target."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper does something genuinely useful in spirit: it applies LaMa inpainting to restore RFI-contaminated data in 21-cm intensity mapping, using real Parkes data with realistic masks, and it tests the downstream effect on foreground removal with three different methods. The CRIME simulation is a nice touch. If the restoration actually works, this could reclaim flagged data for SKA-era surveys.\n\nBut there is a load-bearing problem with the training objective. In Section 2.2, the losses in Eqs. (1) through (5) are written against x, which the text identifies as the input data — the raw observed data that includes the RFI. The network input is the masked version x⊙m. So the perceptual loss, adversarial loss, and feature matching all push the output toward the contaminated values in the masked regions. No sentence anywhere says the loss is restricted to unmasked pixels or that the masked-region target is a foreground model. That means the network is taught to reconstruct the RFI, not remove it. The lower RMS after foreground removal then isn't evidence of genuine restoration; it could just reflect the network's skill at reproducing the kind of structure that the foreground removers later subtract as low-rank modes. This is the central claim, and it's currently not supported.\n\nThe model selection metric Cm/Cu is also a bit circular: it compares predictions to foregrounds extracted by the same polynomial pipeline (Li et al. 2021) that later demonstrates the improvement. That doesn't sink the paper, but it weakens the choice of LaMa-dilated over LaMa-Fourier, especially since SSIM/PSNR prefer LaMa-Fourier.\n\nMinor: the simulation power spectra are single realizations with no error bars, so the claimed improvement toward the HI curve is not statistically quantified. The abstract's 'significantly increases' is too strong for a single mock run.\n\nThe paper deserves a serious referee because the problem is real and the approach is plausible, but it needs major revision. The authors must clarify or fix the training loss — ideally by using a clean foreground model as the target in masked regions or by masking the loss on unmasked pixels. Then re-evaluate the model selection and add error bars to the simulation. As written, I wouldn't cite it.","headline":"Uses LaMa to inpaint RFI-flagged 21-cm IM data; the idea is promising but the loss as written trains the network to reproduce the RFI, so the headline results are not yet trustworthy.","tokens_in":15271,"tokens_out":5280,"would_cite":false,"duration_ms":61086,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A deep neural network that inpaints RFI-contaminated pixels before foreground removal lowers the residual noise in 21-cm intensity mapping and moves the measured angular power spectrum closer to the true neutral-hydrogen signal.","keywords":["21-cm intensity mapping","radio frequency interference","deep neural network restoration","image inpainting","foreground removal","angular power spectrum","LaMa-dilated network"],"falsifier":"In the mock observation the clean no-RFI foreground and HI are known in every masked pixel, so one can directly measure the error between the network's restored values and that known truth; if the restored pixels are no closer to the truth than a simple interpolation of surrounding unmasked pixels, the reported RMS and power-spectrum improvements would reflect smoothing rather than genuine signal recovery.","tokens_in":14236,"feed_emoji":"📡","tokens_out":10290,"duration_ms":92748,"temperature":0.7,"pith_summary":"Radio-frequency interference makes chunks of 21-cm intensity-mapping data unusable, and simply flagging those chunks throws away any chance of recovering the faint cosmological signal hidden underneath. This paper asks whether a deep neural network can fill in the flagged regions before foreground removal, so that the following cleaning steps see a complete data cube. Using real observations to train an inpainting network, it compares the residual root-mean-square after polynomial fitting, SVD, and FastICA with and without restoration, and finds that restoration lowers the residuals, most strongly when the contaminated fraction is large. In a mock observation with a known true HI signal, the restored data yield an angular power spectrum closer to the fiducial spectrum than the unrestored data do. The paper concludes that learned restoration can increase the signal-to-noise ratio available from contaminated intensity-mapping surveys.","feed_headline":"Lower residual noise in 21-cm intensity maps with neural inpainting","feed_subtitle":"Restored pixels lower residual noise after foreground cleaning and pull the power spectrum closer to the true HI signal.","key_machinery":"The central object is the LaMa-dilated inpainting network, a convolutional image-inpainting architecture whose residual blocks use dilated convolutions to combine context from different scales when filling masked regions. It is trained with the LaMa loss—a high-receptive-field perceptual loss plus adversarial, feature-matching, and gradient-penalty terms—on 256×256 patches of real observations with realistic RFI masks. The network's job is to replace masked pixels with an estimate of the underlying smooth foreground and HI signal before polynomial, SVD, or ICA foreground removal. A second load-bearing piece is the mock observation built with a public intensity-mapping simulation, which supplies a known true HI angular power spectrum for comparing restored and unrestored analyses.","core_discovery":"The central claim is that a DNN-based inpainting pipeline, the LaMa-dilated network, can restore RFI-contaminated pixels in 21-cm intensity mapping observations well enough that subsequent foreground-removal methods—polynomial fitting, SVD, and FastICA—recover a cleaner residual than they do from the flagged data. The paper further claims that in a mock observation generated by a public intensity-mapping simulation, the angular power spectrum of restored data lies closer to the true HI spectrum than that of unrestored data, especially at low multipoles. The restoration therefore is not just cosmetic; it is claimed to preserve large-scale astrophysical information that flagging would otherwise remove.","pith_inferences":["A natural next test is to train the same network on simulated data with a clean foreground target in masked pixels; measuring per-pixel error there would separate genuine signal recovery from smoothing.","Because the selection metric favors correlation with the foreground, the ranking of network architectures could change if the metric also penalized spurious HI-like structure in restored regions.","For future large surveys, the practical payoff will depend on whether the restored pixels are unbiased; cross-correlation of restored 21-cm data with optical galaxy surveys is the direct observational check."],"forward_implications":["Surveys can retain more flagged data: restoring masked pixels before foreground removal lowers residual noise instead of discarding those pixels.","The improvement grows with the contaminated fraction, so the method is most valuable in the data that current pipelines handle worst.","The low-multipole part of the angular power spectrum shifts toward the true HI spectrum, meaning the large-scale cosmological signal is better preserved.","The restoration step is method-agnostic: it improves polynomial fitting, SVD, and ICA foreground removal alike without changing those algorithms."],"supporting_citations":[{"why":"Supplies the LaMa inpainting architecture and the loss function that the paper adapts for radio data restoration.","marker":"Suvorov et al. (2021)"},{"why":"Supplies the observational data, the RFI masks, and the foreground model used for training and evaluation.","marker":"Li et al. (2021)"},{"why":"Supplies the public intensity-mapping simulation code used to build the mock observation for the power-spectrum comparison.","marker":"Alonso et al. (2014)"},{"why":"Supplies the foreground power-spectrum model and parameter values used by the simulation.","marker":"Santos et al. (2005)"},{"why":"Supplies the FastICA algorithm used as one of the foreground-removal methods in the comparison.","marker":"Hyvarinen (1999)"},{"why":"Provides the ICA foreground-removal approach that the paper's ICA comparison is built on.","marker":"Chapman et al. (2012)"}],"fun_headline_variants":["Deep learning restores contaminated 21-cm intensity data","Neural inpainting lowers noise in 21-cm maps","DNN fixes RFI pixels in 21-cm intensity mapping","Restoring 21-cm data with neural networks","Neural nets clean 21-cm maps and recover power spectrum"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the network learns to fill masked regions with the true underlying foreground-plus-HI signal rather than reproducing the RFI that was in those pixels, even though the training loss is written against data that still contain the contamination in the masked areas.","fun_headline_variants_meta":{"raw":{"variants":["Deep learning restores contaminated 21-cm intensity data","Neural inpainting lowers noise in 21-cm maps","DNN fixes RFI pixels in 21-cm intensity mapping","Restoring 21-cm data with neural networks","Neural nets clean 21-cm maps and recover power spectrum"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000202,"raw_usage":{"total_tokens":1343,"prompt_tokens":865,"completion_tokens":478,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":392}},"tokens_in":481,"tokens_out":478,"duration_ms":5082,"temperature":1.0,"reasoning_tokens":392,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:35:10.658480+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In the mock observation the clean no-RFI foreground and HI are known in every masked pixel, so one can directly measure the error between the network's restored values and that known truth; if the restored pixels are no closer to the truth than a simple interpolation of surrounding unmasked pixels, the reported RMS and power-spectrum improvements would reflect smoothing rather than genuine signal recovery.","supporting_citations":[{"cited_title":"G., Cooray A., Knox L., 2005, , 625, 575","cited_arxiv_id":null,"evidence_quote":"Supplies the foreground power-spectrum model and parameter values used by the simulation."}],"review_version":1}