{"id":"9733a647-0e99-4d4a-8738-0c5211c45b3e","arxiv_id":"2608.05546","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A compact morphology-aware U-Net with an astronomy-preservation loss achieves F1 0.978 on synthetic RFI masks, keeps 97.6% of injected dispersed-signal fluence, and runs 6.2-7.0x faster than filtool in compute-only benchmarks.","lead":"MARS is a lightweight GPU neural network that flags radio-frequency interference in pulsar and fast radio burst data while preserving dispersed astrophysical pulses. It claims near-real-time mitigation performance comparable to established CPU tools, which matters as next-generation telescopes produce data faster than offline processing can handle.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Patch-level F1/fluence results are in-distribution with the SPECTRALib-style training family; real-data validation is limited to two bright GMRT pulsars, so the central real-world generalization claim rests on unmeasured transfer.","rationale":"The paper is internally coherent: the architecture is plausible, the controlled benchmarks are clearly specified, the ablation isolates the contribution of the preservation loss, and the runtime comparison is honestly labelled as compute-only. The central weakness is external validity. The reader's weakest assumption identifies the same issue: the evaluation is in-distribution with respect to the SPECTRALib-style RFI family used in training, and the real-data support is thin. Two bright-GMRT-pulsar recoveries do not constrain mask-level accuracy or weak-signal retention, and the paper's own Limitations section and Appendix E concede the simulation-to-real gap. A positive result on the HERA real-RFI dataset used for the RFDL comparison would directly address this gap; a negative result would mean the headline numbers should be interpreted as synthetic-benchmark performance rather than deployment performance. I therefore see no reason to move the reader's CONDITIONAL verdict, which already requires broader validation and released artifacts before full acceptance.","tokens_in":22555,"tokens_out":7389,"duration_ms":70822,"concrete_test":"Release the MARS PyTorch/TensorRT checkpoint and run it with the fixed tau=0.5 inference on the HERA real-RFI dataset used for the RFDL comparison (van Zyl & Grobler 2024), reporting pixel-level precision/recall/F1 against the labeled masks. If real-data F1 stays within about 0.03 of the 0.978 synthetic value and precision within about 0.01 of 0.995, the transfer concern is largely resolved; a materially larger drop would re-scope the headline numbers as synthetic-only and require broader multi-telescope validation before the near-real-time deployment claim is accepted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing link is the transfer from SPECTRALib synthetic RFI to real observations. Training augmentation (Section 3) samples the same axis-aligned morphology families that SPECTRALib injects, and Appendix E describes SPECTRALib RFI as additive constant-amplitude regions clipped to 8-bit, explicitly conceding it 'does not represent every drifting, curved, stochastic, or instrument-specific morphology present in real observations.' The headline F1=0.978, precision=0.995, and fluence retentions 0.976/0.964 are measured on patches drawn from this same generator family, so they are near-in-distribution scores for the learned morphology prior, not evidence of coverage of real RFI diversity. The only real-data validation is two GMRT observations (Table 7): both known pulsars are bright (PRESTO significance 13.89 and 17.09), so their recovery is a weak test of mask accuracy or weak-signal preservation; it cannot bound real-RFI mask F1, and it says little about fluence retention for faint dispersed pulses in contaminated frequency-time planes. No code or model artifacts are released, so the transfer cannot currently be checked independently. If real RFI departs from the SPECTRALib family, the quantitative claims describe a simulator benchmark rather than a deployment-ready pipeline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents MARS, a GPU-oriented RFI mitigation pipeline for pulsar and FRB search workflows. The neural component is a compact 270k-parameter U-Net with a morphology-aware bottleneck (local, horizontal, and vertical convolutional branches) designed to segment RFI in the frequency-time plane, and the full pipeline implements normalisation, patch inference, mask reconstruction, replacement, baseline removal, and rescaling on GPU with TensorRT-accelerated inference. Training uses synthetic patches with a weighted BCE + Dice + astronomical-preservation loss (L_astro), and evaluation is staged at three levels: a 15,000-patch SPECTRALib RFI-mask benchmark versus a retrained RFDL baseline; a 18,720-patch dispersed-pulse fluence-retention benchmark with and without mixed injected RFI; and a filterbank-level comparison against filtool via matched PRESTO searches on synthetic binary-pulsar files and two real GMRT observations. Reported headline results are RFI-mask F1 of 0.978, precision 0.995, retained fluence fractions of 0.976/0.964, PRESTO significance ratios of 0.90-0.99, and compute-only speedups of 6.2-7.0x over the fastest tested filtool configuration.","tokens_in":22848,"tokens_out":7891,"duration_ms":73067,"significance":"If the transfer to real RFI is established, the contribution is significant: the architecture is genuinely lightweight and TensorRT-deployable, the L_astro preservation loss is a useful and well-motivated design, and the ablation models are retrained from scratch rather than disabled at inference time. The evaluation protocol is carefully staged, the threshold tau=0.5 is fixed before all reported evaluations, the reported arithmetic is internally consistent (F1=0.978 with precision 0.995 and recall 0.963), and the paper is unusually candid in listing its own limitations. The main unresolved question is whether the SPECTRALib-derived quantitative results, which dominate the abstract and conclusions, transfer to real RFI; the current real-data evidence is too thin to support deployment-ready claims.","major_comments":[{"comment":"The headline patch-level numbers are near-in-distribution measurements rather than evidence of real-RFI coverage. The training augmentation in Section 3 injects horizontal, vertical, compact, block-like, periodic, and composite RFI morphologies, while Appendix E says that SPECTRALib's RFI model is based on additive axis-aligned constant-amplitude regions clipped to 8-bit and that the training augmentation is 'SPECTRALib-style' implemented in the normalised input domain. Section 4.1 then evaluates on SPECTRALib-generated patches from the same morphology families. Consequently, F1=0.978, precision=0.995, and the fluence-retention values 0.976/0.964 measure how well the model fits its own training distribution, not how well it covers real drifting, curved, stochastic, or instrument-specific RFI. Because the paper's practical claim of near-real-time deployment rests on transfer beyond this family, this is load-bearing. Please add an out-of-distribution evaluation, for example synthetic RFI with curved, drifting, or non-axis-aligned structures, or real observations with independently obtained ground-truth masks, and report the resulting F1 and fluence gap; alternatively, explicitly re-scope the quantitative claims as simulator-benchmark results.","section":"Section 3, Section 4.1, Appendix E"},{"comment":"The PRESTO significance comparison is conditional on the target candidate being detected in both cleaned outputs. The text states that the significance comparison includes only realizations in which the injected target candidate is detected in both the filtool- and GPU-NN-cleaned files. The reported median significance ratios (0.946, 0.913, 0.898, 0.982, and 0.988 at the five sampling times) therefore cannot reveal cases where the GPU pipeline suppresses the candidate below the search threshold while filtool detects it. Please report the per-method detection counts out of 200 for each sampling time (Figure 12 gives n but not the detection fraction) and provide an unconditional comparison that treats non-detections explicitly, for example through detection-rate ratios or sensitivity upper limits.","section":"Section 4.3 and Section 5.3"},{"comment":"The real-data validation is limited to two bright, known-DM, non-accelerated GMRT pulsar searches, with recovered PRESTO significances of 13.89 and 17.09. These tests show that a strong periodic signal survives mitigation, but they cannot bound real-RFI mask accuracy, cannot measure false-positive masking of faint dispersed pulses, and cannot support generalisation across telescopes, observing bands, or RFI environments. A concrete strengthening would be to inject faint dispersed pulses with controlled DM and S/N into these real observations before mitigation and measure recovered S/N and fluence as a function of proximity to real RFI, and to report the flagging fraction and mask occupancy of the MARS output on real data relative to filtool. Without such evidence, the real-data section remains an existence check rather than validation of the central real-world generalization claim.","section":"Section 5.4 and Table 7"}],"minor_comments":[{"comment":"The phrase 'minimum valid scaled min' should read 'minimum valid scaled median absolute deviation' for consistency with the definition of d_c,s.","section":"Section 2.2"},{"comment":"The name 'MARS' is used both for the neural mask predictor and for the complete GPU mitigation pipeline, which is elsewhere called 'GPU RFI mitigation pipeline'; please define the naming once and use it consistently, since the performance claims apply to different entities.","section":"Sections 2-5"},{"comment":"The retained-fluence metric excludes injected-signal pixels that overlap ground truth RFI and limits overlap to at most 15% of the signal support; this is transparent in the text, but reporting the actually realised mean overlap fraction would help readers interpret the 0.964 figure.","section":"Section 4.2"},{"comment":"The runtime comparison is compute-only and compares different hardware (GH200 GPU versus EPYC CPU); the text is explicit about this, but the conclusion would benefit from a sentence restating that the 6.2-7.0x figure is not a hardware-independent algorithmic speedup.","section":"Section 4.6"},{"comment":"No code, trained checkpoint, or model artifact is released, which prevents independent reproduction of the central transfer claims; please add a data/code availability statement or a reproducibility appendix with exact artifact URLs and versioned hashes.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"In my assessment the paper is close to being publishable: the architecture is sensible, the loss design is interesting, and the internal consistency of the controlled benchmarks is good. However, the headline quantitative claims are essentially in-distribution with the training-time morphology generator, and the real-data evidence is thin. The authors' own limitations section is honest, but the abstract and conclusions currently overstate deployment readiness. The conditional PRESTO comparison is a further issue that should be addressed by reporting detection rates. If the authors add an out-of-distribution evaluation and unconditional detection-rate reporting, I would be willing to support acceptance. The lack of code/checkpoint release is a reproducibility concern for a methods paper of this type."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline result is real but narrower than the abstract suggests. The new work is the MARS architecture—the small-channel U-Net with the shape-aware bottleneck (1x9 and 9x1 branches), decoder refinement blocks, additive skips—and the astronomical-signal preservation loss L_astro. That loss is a genuinely nice idea, and the ablation shows it buys real fluence retention, especially for low-DM bright pulses. The experimental protocol is also better than most in this subfield: threshold fixed at 0.5 before evaluation, RFDL retrained on the same data, PRESTO settings matched, ablation models independently retrained, and the limitations section is honest.\n\nThe soft spot is not hidden, but it is load-bearing. The training augmentation is 'SPECTRALib-style,' and the patch-level benchmark uses SPECTRALib directly. So the F1 of 0.978 and the 0.976/0.964 fluence numbers are measured on the same morphology family the model was trained on. They are controlled in-distribution scores. That does not make them worthless—the comparison with RFDL on the same family is fair, and the filtool filterbank comparison is more operational—but it means the numbers do not yet support broad real-world generalization. The two real GMRT observations are a start, but both are bright pulsars (PRESTO sigma 13.9 and 17.1), so recovering them is a weak test of mask accuracy or weak-signal preservation. There is no code or model release, so nobody can check the transfer independently right now. The compute speedup is compute-only, on a GH200 against a CPU, so treat the 6–7x as a pipeline-stage number, not an end-to-end wall-clock speedup.\n\nNone of this is fatal. The paper is internally coherent, the baselines are fair, and the limitations are stated rather than buried. It is a practical engineering contribution that deserves referee time. What I would want from a review is pressure to release the artifacts and broaden real-data validation. If that happens, the claims will be much stronger.\n\nRecommendation: send it to review. It benefits from engagement, not desk rejection.","headline":"A careful, well-scoped RFI-mitigation paper whose headline numbers are still in-distribution with its synthetic training family—worth refereeing, but real-world transfer claims need artifacts and more data.","tokens_in":23385,"tokens_out":2477,"would_cite":true,"duration_ms":21604,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A compact 270,769-parameter morphology-aware U-Net segments radio frequency interference with 0.978 F1 and 0.995 precision, preserves 97.6% of injected dispersed-signal fluence, and runs 6.2–7.0x faster than filtool in compute time.","keywords":["radio frequency interference","RFI mitigation","pulsar search","fast radio bursts","segmentation network","U-Net","GPU computing","filterbank processing"],"falsifier":"Take real filterbank observations containing RFI with steeply drifting or curved frequency-time tracks, label the RFI by careful manual inspection or an independent method, and measure MARS's patch-level F1 and retained fluence on those patches. If F1 falls well below the 0.978 synthetic number, or if a bright dispersed pulse overlapped by such RFI loses most of its fluence, the central generalization claim fails; with only two real GMRT observations, this is currently untested.","tokens_in":22347,"feed_emoji":"📡","tokens_out":6471,"duration_ms":52392,"temperature":0.7,"pith_summary":"This paper argues that RFI mitigation for pulsar and fast radio burst searches can be done accurately and fast on a GPU by a compact, morphology-aware segmentation network rather than a large generic backbone or a CPU tool. The proposed MARS pipeline predicts full-resolution RFI masks, replaces masked samples in the original filterbank data, and rescales the output for downstream searches. The authors report near-perfect mask quality on synthetic patches (F1 0.978, precision 0.995), high preservation of injected dispersed signals (97.6% fluence in clean patches), PRESTO candidate significance close to the filtool baseline, and a 6.2–7.0x compute-only speedup. If these controlled results transfer to real observations, MARS would make near-real-time GPU-resident RFI mitigation practical for next-generation arrays while protecting the dispersed transients that searches are designed to find.","feed_headline":"Tiny GPU network cleans RFI at 6.2x speed while saving FRBs","feed_subtitle":"A 270k-parameter morphology-aware U-Net scores 0.978 F1 on RFI masks and preserves 97.6% of injected dispersed-signal fluence.","key_machinery":"The load-bearing object is the shape-aware bottleneck and decoder refinement: three parallel convolutional branches (local 3x3, horizontal 1x9, vertical 9x1) on the deepest feature map, plus horizontal and vertical residual refinement blocks (1x31 or 31x1 followed by 3x3) at each decoder scale, giving an explicit inductive bias for compact, time-extended, and frequency-extended interference without a large backbone. The second key mechanism is the astronomical-signal preservation loss, a quadratic penalty on high RFI probability over clean injected-pulse pixels, which is what raises retained fluence from 0.888 to 0.976 in clean patches. Around these sits a GPU-resident processing chain: per-channel median and median-absolute-deviation normalisation with tanh compression, 512x512 patch inference, mask reconstruction with channel-occupancy promotion, zero replacement in a mean/std-normalised stream, baseline correction, and 8-bit rescaling, all executed on GPU with TensorRT FP16 inference.","core_discovery":"The central claim is that the recurring local and anisotropic time-frequency morphology of RFI in filterbank data can be captured by a small deployment-oriented network, and that the resulting masks, applied through an explicit replacement pipeline, preserve the science content of pulsar and FRB searches. Concretely, a 270,769-parameter full-resolution U-Net with a shape-aware bottleneck achieves 0.978 F1 and 0.995 precision on SPECTRALib synthetic RFI, retains 97.6% of injected dispersed-signal fluence in clean patches and 96.4% in mixed-RFI patches, yields median PRESTO target-significance ratios of 0.90–0.99 relative to filtool on synthetic binary-pulsar filterbanks, and recovers the known pulsars in two GMRT observations. The drop to 0.888 and 0.837 retained fluence when the astronomical-signal preservation loss is removed identifies that loss term as the mechanism protecting compact, bright, low-dispersion-measure pulses.","pith_inferences":["Editorial: the same architecture-and-loss recipe should generalize to other telescopes if retrained on their RFI statistics; the design principle of small anisotropic morphology biases plus an explicit penalty on flagging dispersed power is not tied to GMRT or to SPECTRALib.","Editorial: a concrete testable extension is to train on synthetic RFI with drifting or curved frequency-time tracks and evaluate F1 and fluence retention; the paper's own limitation note predicts that performance would depend on how far real RFI departs from the axis-aligned family.","Editorial: the lower-tail PRESTO losses at 128–512 microseconds point to the replacement policy, not the mask, as the next bottleneck; replacing zero-fill with local interpolation or inpainting from adjacent channels could recover some lost significance.","Editorial: at the measured compute rates, streaming I/O rather than compute is likely to become the limiting factor for online mitigation, so the practical next test is an end-to-end online benchmark that includes disk and network transfer."],"forward_implications":["RFI mitigation can move inside GPU pulsar and FRB search pipelines: the full mitigation computation takes 0.34–3.31 seconds for a roughly 100-second, 4096-channel filterbank, a 6.2–7.0x compute-only reduction versus the fastest filtool configuration tested.","Dispersed transients are preserved much better than with a generic segmentation baseline: 97.6% and 96.4% injected fluence retained for clean and mixed-RFI patches, versus 57.7% and 53.6% for the RFDL baseline.","Cleaned filterbanks remain search-compatible: period-matched PRESTO candidates recover median significance ratios of 0.90–0.99 relative to filtool, approaching equality at 1024 and 1310 microseconds sampling times.","The astronomical-signal preservation loss is a necessary component, not a nicety: removing it drops retained fluence to 0.888 in clean patches and 0.837 in mixed-RFI patches, with the largest loss on low-DM, high-S/N pulses.","Real-data checks pass in two GMRT observations, recovering the known pulsars with comparable or slightly higher significance than filtool."],"supporting_citations":[{"why":"Establishes the time-frequency RFI morphology classes and the long history of mitigation that MARS targets.","marker":"P. Fridman & W. Baan 2001"},{"why":"Supplies the broader RFI mitigation context and the class of offline CPU tools that MARS is compared against operationally.","marker":"A. Offringa 2010"},{"why":"Provides SIGPROC, the fake filterbank generator and format used for the synthetic full-filterbank PRESTO benchmark.","marker":"D. Lorimer 2011"},{"why":"Provides PRESTO, the periodicity and acceleration search engine used to evaluate both MARS-cleaned and filtool-cleaned filterbanks.","marker":"S. M. Ransom et al. 2002"},{"why":"Defines the filtool baseline from PulsarX and the command template used for the filterbank-level comparison.","marker":"Y. Men et al. 2023"},{"why":"Supplies the U-Net architecture that MARS's compact encoder-decoder is derived from.","marker":"O. Ronneberger et al. 2015"},{"why":"Hosts the RFDL baseline comparison context and the HERA dataset evaluation that identifies strong neural RFI detectors.","marker":"D. J. van Zyl & T. L. Grobler 2024"},{"why":"Reports RFDL as the strongest ANN baseline in a recent comparison, justifying its use as the neural mask-prediction baseline.","marker":"N. J. Pritchard et al. 2025"},{"why":"Supplies SPECTRALib, the injection library used to create all synthetic RFI patches and full-filterbank contamination.","marker":"J. White 2026"}],"fun_headline_variants":["GPU U-Net cuts RFI 6.2x faster, keeps 97.6% of FRB signal","Lightweight U-Net scores 0.978 F1 on RFI masks, saves FRBs","MARS: 270k-param network cleans RFI at 6.2x speed","Morphology-aware network cleans RFI, preserves dispersed pulses","Real-time RFI cleanup: 6.2x speedup, 0.978 F1 mask accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that SPECTRALib-style synthetic RFI, which consists of additive, axis-aligned, constant-amplitude regions clipped to 8-bit, represents the RFI that real telescopes produce; the paper itself concedes this does not cover every drifting, curved, stochastic, or instrument-specific morphology present in real observations.","fun_headline_variants_meta":{"raw":{"variants":["GPU U-Net cuts RFI 6.2x faster, keeps 97.6% of FRB signal","Lightweight U-Net scores 0.978 F1 on RFI masks, saves FRBs","MARS: 270k-param network cleans RFI at 6.2x speed","Morphology-aware network cleans RFI, preserves dispersed pulses","Real-time RFI cleanup: 6.2x speedup, 0.978 F1 mask accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00072,"raw_usage":{"total_tokens":3322,"prompt_tokens":1125,"completion_tokens":2197,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":741,"completion_tokens_details":{"reasoning_tokens":2074}},"tokens_in":741,"tokens_out":2197,"duration_ms":14315,"temperature":1.0,"reasoning_tokens":2074,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T11:07:29.531701+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take real filterbank observations containing RFI with steeply drifting or curved frequency-time tracks, label the RFI by careful manual inspection or an independent method, and measure MARS's patch-level F1 and retained fluence on those patches. If F1 falls well below the 0.978 synthetic number, or if a bright dispersed pulse overlapped by such RFI loses most of its fluence, the central generalization claim fails; with only two real GMRT observations, this is currently untested.","supporting_citations":[{"cited_title":"2001, Astronomy & Astrophysics, 378, 327","cited_arxiv_id":null,"evidence_quote":"Establishes the time-frequency RFI morphology classes and the long history of mitigation that MARS targets."},{"cited_title":"2011, Astrophysics Source Code Library, ascl","cited_arxiv_id":null,"evidence_quote":"Provides SIGPROC, the fake filterbank generator and format used for the synthetic full-filterbank PRESTO benchmark."},{"cited_title":"M., Eikenberry, S","cited_arxiv_id":null,"evidence_quote":"Provides PRESTO, the periodicity and acceleration search engine used to evaluate both MARS-cleaned and filtool-cleaned filterbanks."},{"cited_title":"J., Carli, E., & Desvignes, G","cited_arxiv_id":null,"evidence_quote":"Defines the filtool baseline from PulsarX and the command template used for the filterbank-level comparison."},{"cited_title":"2015, in International Conference on Medical image computing and computer-assisted intervention, Springer, 234–241 sigpyproc developers","cited_arxiv_id":null,"evidence_quote":"Supplies the U-Net architecture that MARS's compact encoder-decoder is derived from."},{"cited_title":"J., Wicenec, A., Bennamoun, M., & Dodson, R","cited_arxiv_id":null,"evidence_quote":"Reports RFDL as the strongest ANN baseline in a recent comparison, justifying its use as the neural mask-prediction baseline."},{"cited_title":"2026, SPECTRALib: Synthetic Pulsar Emission, Contamination, and Transients Radio Astronomy Library,, GitHub software repository https://github.com/jack-white1/SPECTRALib","cited_arxiv_id":null,"evidence_quote":"Supplies SPECTRALib, the injection library used to create all synthetic RFI patches and full-filterbank contamination."}],"review_version":1}