{"id":"c1a59a5a-a8ab-437f-b77b-1a0f9c1488fb","arxiv_id":"2506.05489","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"F2T2-HiT combines FFT Transformer and hierarchical Transformer blocks in a UNet for single image reflection removal, reporting the highest average PSNR but not the highest SSIM.","lead":"A team at OPPO proposes a U-shaped network, F2T2-HiT, that mixes FFT-based transformer blocks and hierarchical window transformers for removing reflections from single photographs. The paper reports the highest average PSNR on three reflection benchmarks, but its evaluation is undermined because the training set appears to include the same real-world images used for testing.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The state-of-the-art claim is not supported by the paper's own Table 1, and the unstated Real/Nature train-test split leaves the generalization claim unresolved.","rationale":"The reader's weakest assumption (possible Real/Nature train-test overlap) is legitimate and is explicitly flagged in Section 4.2, but it is not by itself the most decisive problem: the 494-image average is dominated by SIR2 (454 images), which is a test-only set, so a disjoint Real/Nature split would not by itself address the SIR2 advantage that drives the average. The more direct, internally checkable problem is that the paper's own Table 1 contradicts the unqualified 'state-of-the-art' claim: the method loses to Zhu et al. on Real PSNR by 2.18 dB and on SSIM on every row (Nature, SIR2, average). Thus the evidence supports at most 'best average PSNR on these three sets,' not a general superiority claim. The data-leakage issue remains an aggravating, unverified assumption, and the lack of code, full configuration, and error bars further limits verification. I therefore keep the reader's REJECT verdict unchanged. This is a critique of the support for the empirical claim, not an accusation; the arithmetic in Table 1 appears internally consistent, but the interpretation overreaches.","tokens_in":8336,"tokens_out":11975,"duration_ms":134818,"concrete_test":"Run an evaluation audit: (a) compute the file-level intersection between the 89+200 real training pairs and the 40 images used in the Real and Nature test splits; if any overlap exists, retrain on a strictly disjoint split and rerun Table 1. (b) On the same split, recompute the comparison against Zhu et al. with paired bootstrap 95% confidence intervals and with a rank-based aggregate over the three benchmarks' PSNR and SSIM. If any test identity appears in training, or if the 0.17 dB average-PSNR advantage is not significant or does not rank first under the aggregate, the state-of-the-art claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of state-of-the-art performance does not follow from Table 1, even before considering data leakage. On Real(20), the method is substantially worse than Zhu et al.: PSNR 21.64 vs 23.82. On SSIM it is lower on Nature (0.837 vs 0.843), SIR2 (0.903 vs 0.910), and the 494-image average (0.894 vs 0.904). The only headline metric favoring the method is average PSNR (25.57 vs 25.40), a 0.17 dB gap dominated by the 454-image SIR2 subset and reported without error bars or significance tests. Independently, Section 4.2 states that 'Real and Nature include both training and test data' while the training set includes 89 real pairs from [30] and 200 real pairs from [28]; the paper never states that the test images were excluded from training. This unresolved overlap means the generalization claim is not fully supported, although SIR2 itself is test-only. Together these issues mean the empirical evidence does not license the unqualified 'state-of-the-art' conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes F2T2-HiT, a U-shaped network for single-image reflection removal that combines FFT-based Transformer blocks and hierarchical window Transformer blocks inside a NAFNet-style encoder-decoder. The training set includes 89 real pairs from [30], 200 real pairs from [28], and synthesized data. Evaluation is reported on the Real, Nature, and SIR2 benchmarks, and Table 1 reports the highest average PSNR (25.57 dB) among the compared methods. The paper's central claim is that this constitutes state-of-the-art performance, supported by the quantitative comparison in Table 1 and the ablations in Table 2.","tokens_in":8522,"tokens_out":4122,"duration_ms":51571,"significance":"If the reported results were obtained under a clean and properly split evaluation protocol, the architecture would be a plausible contribution: it combines two existing Transformer designs (F2T2 and HiT blocks) within a well-known UNet/NAFNet framework, and the ablation shows monotone gains from adding each block. However, the empirical evaluation as written does not establish the headline claim. The unresolved overlap between the training set and two of the three test benchmarks, together with a comparison table that contradicts the 'state-of-the-art' conclusion, mean the paper's main contribution is not supported by its evidence. The paper does not provide code, model size, FLOPs, or significance tests, so the efficiency and superiority claims are also not fully documented.","major_comments":[{"comment":"The training and test sets are not shown to be disjoint. The text states that 'Real and Nature include both training and test data' while the training set is described as including 89 real pairs from [30] and 200 real pairs from [28]. Since Real and Nature are benchmarks drawn from the same sources as these training pairs, the results on those benchmarks cannot be interpreted as generalization unless the test images were explicitly excluded from training and the split was documented. No such exclusion is stated anywhere in the manuscript. Tables 1 and 2 therefore contain numbers on Nature and Real that may reflect memorization of the training distribution rather than reflection-removal ability.","section":"Section 4.2"},{"comment":"The state-of-the-art claim is not supported by the paper's own comparison. On Real(20), the proposed method is substantially worse than Zhu et al. [15] in PSNR (21.64 vs. 23.82). On SSIM, the proposed method is lower than Zhu et al. on Nature (0.837 vs. 0.843), on SIR2 (0.903 vs. 0.910), and on the 494-image average (0.894 vs. 0.904). The only metric that favors the method is average PSNR (25.57 vs. 25.40), a 0.17 dB gap that is dominated by the 454-image SIR2 subset and is reported without error bars or statistical significance tests. The unqualified 'state-of-the-art performance' in the abstract and conclusion therefore does not follow from the evidence.","section":"Table 1"},{"comment":"The ablation study inherits the same evaluation-protocol problem. Because the Real and Nature portions of the test set may overlap with the training data, the incremental gains from the HiT and F2T2 blocks shown in Table 2 cannot be attributed to the architectural designs rather than to overfitting of those training examples. Without a clean disjoint split, the ablation does not establish the effectiveness of either proposed block.","section":"Section 4.3 / Table 2"}],"minor_comments":[{"comment":"The sentence 'Real and Nature include both training and test data' needs to be replaced with a precise statement of how many test images there are, which sources they come from, and how they were held out from the 89+200 real training pairs.","section":"Section 4.2"},{"comment":"The caption contains the typo 'architecutre'; also, '1x1, conv' and similar shorthand should be expanded or defined in the figure caption for clarity.","section":"Figure 1"},{"comment":"The paper claims computational efficiency in the introduction and conclusion, but no model parameter count, FLOPs, inference time, or memory usage is reported anywhere in Section 4.","section":"Section 4.1"},{"comment":"The closing 'state-of-the-art results' phrasing should be qualified or removed until the evaluation protocol is corrected and the comparison with Zhu et al. [15] is reconciled with the numbers in Table 1.","section":"Conclusion"},{"comment":"References [6]-[10] are largely self-citations of preprints and workshop papers; the authors should clarify the relationship of the current work to these papers and provide peer-reviewed versions where available.","section":"References"}],"recommendation":"reject","confidential_remarks":"The central problem is the unresolved overlap between training and test data, which is explicitly acknowledged in Section 4.2 but never resolved with a stated exclusion or split. Even setting that aside, Table 1 does not support the paper's state-of-the-art claim on the metrics that matter. A revised version would need to retrain with a clearly disjoint test protocol, rerun all comparisons and ablations, and substantially soften the claims. Given that the headline contribution depends entirely on the empirical evaluation, I cannot recommend acceptance of the manuscript in its current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I'll give you the short version: this is a workmanlike architecture paper whose headline claim doesn't survive contact with its own numbers. The authors take NAFNet's U-Net, drop in HiT and F2T2 blocks from super-resolution and restoration work, and test on the standard reflection-removal benchmarks. That combination is new for SIRR, and the ablations in Table 2 show each addition helps on most rows. Good: the writing is clear, the block descriptions are adequate, and the work is properly situated in the literature. No code is released, but that's not fatal.\n\nThe soft spots are real and central. First, the 'state-of-the-art' claim in the abstract and conclusion doesn't match Table 1. On the Real benchmark, your method is more than 2 dB behind Zhu et al. (21.64 vs 23.82), and the average SSIM is lower (0.894 vs 0.904). The only metric where you lead is average PSNR, by 0.17 dB, and that's driven by the 454-image SIR2 subset. With no error bars and no significance test, that's not a convincing lead.\n\nSecond, and more worrying: Section 4.2 says the training set includes 89 real pairs from Zhang et al. and 200 from IBCLN, and then says, without apparent irony, that 'Real and Nature include both training and test data.' The paper never states that the specific test images were held out. If they weren't, the reported numbers on Real and Nature are memorization, not generalization. The authors flag the overlap but don't resolve it. That's exactly the kind of thing a critical reader is entitled to demand before accepting the performance claims.\n\nThe ablation study is well-formed and does show a consistent gain from the HiT and F2T2 blocks, so the architectural contribution is probably real, just incremental. The citation pattern is fine; the blocks are properly attributed to [19] and [20].\n\nNet: a competent but incremental engineering paper whose evaluation is not trustworthy as presented. A serious referee would need to see a clear statement of how Real and Nature test images were excluded from training, plus a re-run of Table 1 with error bars and a fair comparison on the terms the authors themselves define. As-is, I wouldn't send this to peer review; the central empirical claim is contradicted by the paper's own data. If you want an example of how not to report train/test overlap, this is a good case study for the reading group.","headline":"The architecture is a sensible repackaging of known blocks, but the paper's own Table 1 undercuts its state-of-the-art claim, and the Real/Nature training overlap is left unresolved.","tokens_in":9063,"tokens_out":2545,"would_cite":false,"duration_ms":25876,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"F2T2-HiT, a U-shaped network pairing FFT Transformer blocks with hierarchical window Transformer blocks, reports state-of-the-art reflection removal, averaging 25.57 dB PSNR across the Nature, Real, and SIR2 benchmarks.","keywords":["reflection removal","single image reflection removal","transformer","fast Fourier transform","hierarchical attention","image restoration","U-Net","frequency domain"],"falsifier":"A decisive check is to verify the split: list the file names of the Real and Nature test images and check them, plus near-duplicates, against the 89 and 200 real training pairs. If any overlap exists and removing it drops the average PSNR below the 25.40 dB of the strongest prior method, the state-of-the-art claim rests on memorization rather than the architecture.","tokens_in":8131,"feed_emoji":"🪞","tokens_out":6012,"duration_ms":58525,"temperature":0.7,"pith_summary":"F2T2-HiT is a U-shaped network for removing reflections from a single photograph taken through glass. The paper's claim is that replacing the plain blocks of a NAFNet-style encoder–decoder with two custom Transformer blocks—an FFT Transformer block that reads global frequency-domain structure and a hierarchical window Transformer block that gathers multi-scale local detail—gives the best average performance yet reported on three real-world reflection benchmarks. The headline number is 25.57 dB average PSNR over 494 test images from the Nature, Real, and SIR2 datasets. If true, this would give photographers and downstream vision systems a practical way to recover the transmitted scene behind glass from one image alone.","feed_headline":"U-shaped FFT transformer tops reflection benchmarks at 25.57 dB","feed_subtitle":"Frequency-domain global context plus multi-scale window attention outperforms prior methods on three real-world datasets.","key_machinery":"The central objects are the two named blocks. The F2T2 (Fast Fourier Transform Transformer) block follows LayerNorm – FFT layer – LayerNorm – multi-kernel convFFN; its FFT layer runs a spatial-domain branch (dilated convolutions) and a frequency-domain branch (2D FFT, frequency-conditional positional embedding, and frequency dynamic convolution) and fuses them with channel attention, giving the network an image-wide receptive field early on. The HiT (Hierarchical Transformer) block follows LayerNorm – window self-attention – LayerNorm – channel-wise FFN, where the window self-attention uses expanding windows of sizes 4, 8, and 16 and computes spatial and channel self-correlations with linear complexity in window size. These blocks replace the convolutional blocks of a NAFNet U-shaped encoder–decoder, which is what lets the paper test the blocks' individual contributions in ablation.","core_discovery":"On the paper's own terms, the discovery is that a single-stage U-shaped architecture does not need bespoke reflection priors or multi-stage refinement to be state of the art: it needs a global receptive field from the first layers and attention that operates at several window scales. The F2T2 block supplies the first property by mixing a spatial-domain branch with a frequency-domain branch built on 2D FFT and frequency-conditional positional encoding, so every feature can see the whole image; the HiT block supplies the second by partitioning features into windows of sizes 4, 8, and 16 and aggregating them with spatial and channel self-correlation. Combined inside a NAFNet U-Net with skip connections, the model reports an average PSNR of 25.57 dB across the three benchmarks, with per-dataset values of 26.08 dB on Nature, 21.64 dB on Real, and 25.72 dB on SIR2. The authors interpret the ablation results—NAFNet alone at 23.97 dB, plus HiT at 25.00 dB, plus F2T2 at 25.57 dB—as evidence that each block carries a distinct part of the improvement.","pith_inferences":["Because the SIR2 subset contributes 454 of the 494 averaged images, the average PSNR is dominated by that benchmark; a method can top the average while being weaker on the small Real set, which is what the table shows, so per-benchmark numbers rather than the average should drive method comparisons.","The same dual-domain recipe—global frequency context plus hierarchical windows—could transfer to other single-image restoration tasks where artifacts span large areas, such as deraining, dehazing, or glare removal, though the paper does not test those.","A strict test of the generalization claim would retrain with every Real and Nature test image explicitly excluded and report the resulting PSNR; the current paper does not state that this exclusion was done."],"forward_implications":["A single trained F2T2-HiT model can be applied directly to three different real-world benchmarks without per-dataset tuning, and the paper reports it beats all listed prior methods on the average PSNR over 494 images.","Large-area reflections, which local receptive fields miss, become addressable because the FFT layers give even the early encoder layers a view of the whole image.","The hierarchical windows let the model keep local edge detail while still forming long-range dependencies, so reflections of mixed sizes can be handled in one pass.","The ablations show the gains are additive: replacing NAFNet blocks with HiT adds about 1 dB on average, and adding F2T2 adds about 0.6 dB more.","Training takes roughly 24 hours on eight A100 GPUs, indicating the architecture is practical to reproduce at academic scale."],"supporting_citations":[{"why":"Supplies the NAFNet U-shaped baseline architecture whose blocks this paper replaces.","marker":"[1]"},{"why":"Fast Fourier Convolutions are the source of the image-wide receptive field used in the F2T2 block.","marker":"[2]"},{"why":"The hierarchical window Transformer block is adapted from this method.","marker":"[19]"},{"why":"The FFT-meets-Transformer block design is adapted from this method.","marker":"[20]"},{"why":"Supplies the training-data recipe and is the strongest prior baseline (25.40 dB average) that this paper claims to beat.","marker":"[15]"},{"why":"Provides the component-synergy architecture whose synthesis protocol generates the simulated training pairs.","marker":"[14]"},{"why":"Supplies 89 of the real training pairs used for training.","marker":"[30]"},{"why":"Supplies 200 of the real training pairs and provides the IBCLN baseline.","marker":"[28]"},{"why":"Restormer block is the baseline replaced in ablation to show HiT's gain.","marker":"[32]"}],"fun_headline_variants":["FFT + hierarchical transformers hit 25.57 dB on reflection removal","One-stage transformer beats multi-stage rivals in reflection removal","Global FFT meets multi-scale windows: new SOTA in SIRR","Single U-Net transformer achieves 25.57 dB without reflection priors","FFT transformer removes reflections at 25.57 dB, no priors needed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument for state-of-the-art generalization assumes that no image from the Real or Nature test benchmarks was included in the 89+200 real training pairs, but the paper only says these datasets contain both training and test data and never states that the test images were left out.","fun_headline_variants_meta":{"raw":{"variants":["FFT + hierarchical transformers hit 25.57 dB on reflection removal","One-stage transformer beats multi-stage rivals in reflection removal","Global FFT meets multi-scale windows: new SOTA in SIRR","Single U-Net transformer achieves 25.57 dB without reflection priors","FFT transformer removes reflections at 25.57 dB, no priors needed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000962,"raw_usage":{"total_tokens":4120,"prompt_tokens":993,"completion_tokens":3127,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":3031}},"tokens_in":609,"tokens_out":3127,"duration_ms":26432,"temperature":1.0,"reasoning_tokens":3031,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:20:43.543173+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check is to verify the split: list the file names of the Real and Nature test images and check them, plus near-duplicates, against the 89 and 200 real training pairs. If any overlap exists and removing it drops the average PSNR below the 25.40 dB of the strongest prior method, the state-of-the-art claim rests on memorization rather than the architecture.","supporting_citations":[{"cited_title":"Swin transformer: Hierarchical vision transformer using shifted windows,","cited_arxiv_id":null,"evidence_quote":"Supplies 200 of the real training pairs and provides the IBCLN baseline."},{"cited_title":"Single image reflection re- moval with physically-based training images,","cited_arxiv_id":null,"evidence_quote":"Restormer block is the baseline replaced in ablation to show HiT's gain."},{"cited_title":"F2T2-HiT: A U-Shaped FFT Transformer and Hierarchical Transformer for Reflection Removal","cited_arxiv_id":"2506.05489","evidence_quote":"Supplies the NAFNet U-shaped baseline architecture whose blocks this paper replaces."},{"cited_title":"ghosting","cited_arxiv_id":null,"evidence_quote":"Fast Fourier Convolutions are the source of the image-wide receptive field used in the F2T2 block."},{"cited_title":"Single image reflection separation via component synergy,","cited_arxiv_id":null,"evidence_quote":"The hierarchical window Transformer block is adapted from this method."},{"cited_title":"Revisiting single image reflection removal in the wild,","cited_arxiv_id":null,"evidence_quote":"The FFT-meets-Transformer block design is adapted from this method."},{"cited_title":"Degradation-aware image enhancement via vision- language classification,","cited_arxiv_id":null,"evidence_quote":"Supplies the training-data recipe and is the strongest prior baseline (25.40 dB average) that this paper claims to beat."},{"cited_title":"Single image reflection removal beyond linearity,","cited_arxiv_id":null,"evidence_quote":"Supplies 89 of the real training pairs used for training."}],"review_version":1}