{"id":"7b76cf61-6e56-4fbd-8868-e04dca0cdde3","arxiv_id":"2506.05482","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"OpenRR-5k offers 5,300 pixel-aligned real-world image pairs for reflection removal, but its clean labels are generated by an AI tool and manual editing rather than captured true scenes.","lead":"OpenRR-5k is a new 5,300-pair benchmark for single-image reflection removal, built by running real photos through a commercial reflection-removal AI and then hand-editing the clean results. It matters because reflection removal research needs larger real-world training sets, and existing public datasets are smaller or synthetic.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"OpenRR-5k's clean labels and validation targets are both produced by the same OPPO AI plus manual-editing pipeline, so the reported 26.59 dB validation PSNR may measure consistency with the commercial tool rather than true reflection-free scenes.","rationale":"The reader's weakest-assumption analysis identifies exactly the same point: ground truth from commercial AI plus manual editing is not validated. My review confirms it is the load-bearing assumption. The reported validation metrics cannot distinguish true reflection removal from reproduction of the AI's artifacts, because both inputs and labels pass through the same commercial pipeline. Cross-dataset numbers in Table II are a step in the right direction but are not sufficient: without a same-architecture control trained on existing datasets, 'improves generalization' is not established. I find no other central-claim threat that is more severe. The dataset's non-release compounds the problem but is secondary to the circular supervision issue. The paper does not commit fraud; it simply provides no evidence that the AI-generated labels are accurate, and the evaluation design is not independent. Therefore the reader's REJECT verdict remains appropriate, with the path to acceptance being physical-capture spot checks and a controlled cross-training comparison.","tokens_in":6249,"tokens_out":4185,"duration_ms":49440,"concrete_test":"Take a random subset of 100 OpenRR-5k pairs (50 train, 50 val). For each scene, physically recapture the same viewpoint with the reflective glass removed or with a polarizer to obtain an independent clean image, allowing small manual alignment if needed. Compute PSNR/SSIM between the OpenRR-5k clean label and the physical clean capture. If median PSNR/SSIM is significantly below the level seen in physically captured paired datasets such as RRW, or if the largest errors correlate with regions where the OPPO tool hallucinates content, the Section III.A ground-truth assumption fails. Additionally, train the same NAFNet variant separately on OpenRR-5k and on RRW training splits and evaluate both on a common external test set; if OpenRR-5k does not beat the RRW-trained model, the generalization claim in Table II is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that OpenRR-5k 'enhances the performance and generalization' of SIRR models depends on the assumption in Section III.A that the OPPO AI reflection remover plus manual editing produces clean transmission layers that faithfully approximate the true reflection-free scenes. The paper gives no quantitative evidence for this assumption: no comparisons against physically captured clean images, no residual-artifact statistics, no inter-editor consistency, and no before/after analysis. Worse, the same pipeline is used to create both the training targets and the validation pairs. Thus Table III's PSNR/SSIM/LPIPS/DISTS/NIQE values on OpenRR-5kval measure consistency with the commercial software's output distribution, not fidelity to ground-truth scenes. A model trained on these labels can score highly while reproducing the AI's hallucinations or smoothing artifacts, and it will generalize poorly to true glass scenes. Table II does evaluate on Nature/Real/SIR2 external sets, but it lacks the necessary control: we do not see the same NAFNet trained on RRW or existing datasets and evaluated on the same external test sets, nor any comparison to prior state-of-the-art. Since the dataset is not released, the claimed benchmark value cannot be checked independently. The single load-bearing weakness is therefore an unvalidated, likely circular supervision signal; if it fails, the dataset is not a trustworthy benchmark for reflection removal.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces OpenRR-5k, a dataset of 5,300 real-world reflection-contaminated images paired with clean counterparts for single image reflection removal (SIRR). The clean images are generated by applying the OPPO smartphone AI reflection removal tool and then manually refining the results with image editing software (Section III.A). The authors train a NAFNet-based baseline on the 5,000 training pairs and report quantitative metrics (PSNR, SSIM, LPIPS, DISTS, NIQE) on the 300 validation pairs plus PSNR on three external datasets (Nature, Real, SIR2). The central claim is that OpenRR-5k provides a large-scale, pixel-aligned, diverse real-world dataset that enhances the performance and generalization of SIRR models. The paper also provides a 100-image test set without ground truth for qualitative evaluation.","tokens_in":6525,"tokens_out":3599,"duration_ms":35979,"significance":"If the dataset's label-generation pipeline were validated, the scale and diversity of OpenRR-5k could make it a valuable training resource for the SIRR community. The proposed protocol is convenient and scalable compared to physical capture setups. However, the paper provides no evidence that the AI-generated clean images faithfully represent true reflection-free scenes, and the experimental evaluation is too sparse to support the claimed benefits. The utility of the dataset therefore rests entirely on an unvalidated, likely circular supervision signal. The authors promise to release the dataset and code, which would allow independent checks, but as submitted the manuscript does not demonstrate that the dataset is a trustworthy benchmark.","major_comments":[{"comment":"The ground-truth generation pipeline, which uses the OPPO AI reflection remover followed by manual editing, is not validated. There are no quantitative comparisons against physically captured clean images, no residual-artifact statistics, no inter-editor consistency analysis, and no before/after evaluation. This is the load-bearing assumption of the entire paper: if the AI-generated clean images contain hallucinations or smoothing artifacts, models trained on this dataset will learn to reproduce those artifacts rather than perform true reflection removal. The absence of any such validation makes the dataset's supervision unreliable.","section":"III.A"},{"comment":"The claim that the dataset 'effectively enhances the performance and generalization' of SIRR models is unsupported by the reported experiments. Table II lists only NAFNet trained on OpenRR-5k evaluated on Nature, Real, and SIR2, with no control condition: we do not see the same architecture trained on existing datasets (e.g., RRW, Nature, or Real) and evaluated on the same external test sets, nor do we see comparisons with prior state-of-the-art methods. Without these baselines, the PSNR values of 25.62, 21.16, and 24.52 cannot be interpreted as evidence of improvement or generalization.","section":"IV.A, Table II"},{"comment":"The validation set (OpenRR-5kval) is generated by the same OPPO-AI-plus-manual-editing pipeline as the training set. Therefore, the reported PSNR of 26.59 dB on OpenRR-5kval measures the model's consistency with the commercial tool's output distribution, not its fidelity to true reflection-free scenes. A meaningful evaluation of the dataset's utility requires held-out real-world pairs with physically captured ground truth, which are absent from this paper.","section":"IV.A, Tables II and III"},{"comment":"The claim that the dataset provides 'True Real-World Data' is overstated. While the input reflection images are real photographs, the clean ground-truth images are the output of a proprietary AI model plus manual editing, not real captures. The statement that the method 'eliminates the need for collecting ground-truth data' conflates avoiding physical capture with obtaining authentic ground truth. This wording could mislead readers about the nature of the supervision.","section":"III.A, item 3"}],"minor_comments":[{"comment":"The OpenRR-1k dataset is cited inconsistently: reference [8] is used in the text on multiple occasions, but Table I cites it as [10]. Please unify the citations.","section":"References"},{"comment":"The abstract says the test set consists of '100 real-world testing images without ground truth,' while Section III.B says '100 image pairs without Ground Truth for the test set.' Please clarify whether the test set contains images or pairs.","section":"Abstract and III.B"},{"comment":"The implementation details do not specify the loss function used to train NAFNet, the total number of parameters, or the exact architecture modifications beyond increasing block counts. These details are needed for reproducibility.","section":"IV.B"},{"comment":"The pie chart percentages for scene content (60% landscapes, 7% animals, 14% humans, 19% objects) sum to 100%, but the figure is not visible in the submitted text; please ensure it is legible and the caption clearly defines the categories.","section":"Fig. 3"}],"recommendation":"reject","confidential_remarks":"This paper is a dataset contribution whose central value depends on the validity of AI-generated ground-truth labels. The lack of any validation of those labels, combined with the circular use of the same pipeline for training and validation, makes the benchmark's core claim unsupportable as written. Even a major revision would need to collect physical ground truth for at least a subset, quantify label fidelity, and add extensive cross-dataset training comparisons. The authors' prior OpenRR-1k paper and the NTIRE challenge paper are heavily self-referenced, which is not itself a problem, but the present manuscript does not advance the evidence base beyond those works."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: OpenRR-5k is a large-scale real-world reflection removal dataset with a genuinely new collection protocol, but the paper never validates the key asset—its ground truth. As submitted, the central claim does not hold up.\n\nWhat is good: 5,300 pairs is a real scale jump over Nature, Real, and SIR2, and the protocol of using a commercial AI reflection remover plus manual editing to produce clean images is pragmatic and potentially scalable. The paper correctly identifies the weaknesses of synthetic and semi-synthetic datasets. The diversity of scenes (lighting, glass types) is a plus.\n\nThe soft spots are not minor. The ground truth is produced by the OPPO AI software plus manual editing, and the same pipeline generates both the training targets and the validation pairs. So the 26.59 dB PSNR on OpenRR-5kval measures how well NAFNet reproduces the commercial tool's output distribution, not how close it gets to the true reflection-free scene. There are no quantitative checks: no comparisons against physically captured clean images, no residual-artifact statistics, no inter-editor consistency. Table II reports PSNR on Nature, Real, and SIR2, but lacks the essential control—we never see the same NAFNet trained on RRW or another existing dataset and evaluated on those same external sets, nor any comparison to prior state-of-the-art. The claim of 'notable improvements' is therefore unsupported. And the dataset is not released, so the resource cannot be independently checked.\n\nThe collection protocol could be a real contribution if the authors validated it—for example, by comparing AI-generated clean images to true captures on a small set, or at least showing that manual editing removes residual artifacts. Without that, the load-bearing assumption is unexamined, and no amount of training results can fix it.\n\nI would not send this to peer review in its current form. The flaw is fundamental and identifiable at a desk screen. If the authors return with a validation study and a released dataset, it becomes a different, potentially useful paper. For now, treat it as an interesting but unverified resource idea.","headline":"OpenRR-5k is a large-scale dataset with a novel collection protocol, but the ground truth is unvalidated and circular, so the benchmark claims don't hold as submitted.","tokens_in":6997,"tokens_out":3592,"would_cite":false,"duration_ms":36229,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that OpenRR-5k, a new dataset of 5,300 real-world, pixel-aligned reflection/clean image pairs produced with AI-assisted editing, supplies the large, diverse training data that single-image reflection removal models lack…","keywords":["single image reflection removal","real-world dataset","pixel-aligned image pairs","benchmark","dataset collection protocol","image restoration","reflection removal","ground truth generation"],"falsifier":"Collect scenes where the true clean image is independently known (for example, photograph the same subject through glass and again after removing the glass, or use a black-cloth capture), run the paper's AI-plus-editing protocol on the through-glass images, and measure per-pixel differences between the produced 'clean' images and the true clean images. If systematic discrepancies appear in regions such as text, fine textures, or faces that the commercial tool tends to blur or hallucinate, the central assumption that OpenRR-5k pairs are faithful ground truth would be refuted.","tokens_in":6093,"feed_emoji":"🪟","tokens_out":8401,"duration_ms":88686,"temperature":0.7,"pith_summary":"OpenRR-5k is a new large-scale benchmark for single image reflection removal (SIRR), containing 5,300 pixel-aligned pairs of reflection-contaminated and clean real-world images, split into 5,000 training, 300 validation, and 100 ground-truth-free test images. The paper's central claim is that this dataset, collected by passing real photos through a commercial AI reflection-removal tool and then manually refining the results, provides training data that improves both performance and generalization compared with existing reflection-removal datasets. What makes the claim worth attention is that it offers an inexpensive, scalable route to diverse real-world training pairs without synthetic blending or physical glass-removal setups. The authors support it by training a U-Net-style model on OpenRR-5k alone and reporting 26.59 dB PSNR on the validation set plus cross-dataset scores on Nature, Real, and SIR2. If the dataset works as claimed, it would let SIRR models train on thousands of authentic scenes rather than a few hundred controlled captures.","feed_headline":"5,300 real pairs target reflection removal's data gap","feed_subtitle":"AI-edited clean images, aligned to originals, train on real scenes instead of synthetic blends.","key_machinery":"The load-bearing mechanism is the paired-data generation pipeline: each training pair consists of an observed reflection image and a clean image obtained by applying a commercial AI reflection remover followed by manual fine-detail editing, so that the two images are aligned at the pixel level and can be used for direct supervised training. The second component is the dataset structure itself, with 5,000/300/100 train/validation/test splits and explicit coverage of scene types (humans, animals, objects, landscapes) and lighting conditions (daytime, nighttime, indoor), designed so that models trained on the training split face realistic in-the-wild variation at evaluation time. The empirical argument is carried by a NAFNet-based baseline: trained only on OpenRR-5k, it reaches 26.59 dB PSNR on its own validation set and the cross-dataset PSNR figures above, which the paper reads as evidence of improved generalization without any additional training data.","core_discovery":"The central discovery the paper argues for is that a high-quality SIRR training set can be mass-produced from real-world photographs: a proven commercial AI reflection remover produces an initial clean estimate, human editors refine away residual reflections and artifacts, and the refined image is paired with the original photograph as pixel-aligned ground truth. According to the paper, this protocol removes the pixel-level misalignment that plagues physically captured pairs and the domain gap that plagues fully synthetic pairs, while enabling data collection at a scale and diversity that previous real-world datasets do not match. The supporting experiment trains a NAFNet-based restoration network on OpenRR-5k and reports PSNR 26.59 on OpenRR-5kval, together with 25.62 on Nature, 21.16 on Real, and 24.52 on SIR2, which the authors present as evidence that the dataset 'enhances the performance and generalization' of reflection removal models.","pith_inferences":["The authors do not run a matched comparison against RRW's 14,952 pairs, so a direct test of whether OpenRR-5k's advantage comes from diversity, alignment, or sheer scale would be to train the same baseline on both datasets under identical settings and evaluate on a shared test set.","Because the ground truth is produced by a commercial AI plus human editing, the dataset implicitly teaches models to reproduce that tool's notion of a clean image; systematic errors in the tool, such as over-smoothing text or fine texture, could be inherited by trained models, a risk that perceptual or human-preference studies on the 100 no-GT test images could expose.","The same edit-the-AI-output protocol could be transferred to other ill-posed restoration tasks, such as dehazing or low-light enhancement, where true clean targets are similarly hard to capture, making the approach a general template for dataset construction."],"forward_implications":["Training a SIRR model on OpenRR-5k rather than on earlier small real-world datasets should improve transfer to unseen reflection scenes, because the training pairs come from genuine photographs across many lighting conditions and glass types.","The 300 validation pairs give the community a pixel-aligned real-world benchmark for comparable evaluation, while the 100 no-ground-truth test images allow practical performance checks on fully unlabeled scenes.","The collection protocol can be scaled further through crowdsourcing, since it requires no specialized capture equipment, meaning larger benchmarks of the same type are straightforward to build.","Because OpenRR-5k's training split includes all OpenRR-1k pairs, results on the new benchmark directly extend and are comparable with the earlier 1k dataset."],"supporting_citations":[{"why":"Supplies the earlier 1k-pair OpenRR dataset, whose training pairs are folded into OpenRR-5k's training split.","marker":"[10]"},{"why":"The 454-image SIR2 benchmark used as a cross-dataset evaluation target.","marker":"[19]"},{"why":"The Real dataset with 89/20 pairs used as a cross-dataset evaluation target.","marker":"[20]"},{"why":"The Nature dataset with 200/20 pairs used as a cross-dataset evaluation target and as the main example of physically captured semi-synthetic data.","marker":"[5]"},{"why":"The RRW dataset with 14,952 pairs, the larger in-the-wild comparison whose physical capture pipeline the paper argues is impractical.","marker":"[23]"},{"why":"The NAFNet image-restoration baseline that the paper adapts and trains to validate the dataset.","marker":"[26]"},{"why":"Documents pixel-alignment and residual-reflection problems in semi-synthetic capture, motivating the new protocol.","marker":"[22]"}],"fun_headline_variants":["Real-world reflection dataset: 5,300 aligned pairs","OpenRR-5k brings real reflection removal training data","Mass-produced real ground truth for reflection removal","5,300 real photo pairs solve reflection data scarcity","AI plus human edits yield 5,300 pixel-aligned reflection pairs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the clean images produced by the commercial AI reflection remover plus manual editing are faithful approximations of the true reflection-free scenes, so that models trained on them learn real reflection removal rather than the tool's own artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Real-world reflection dataset: 5,300 aligned pairs","OpenRR-5k brings real reflection removal training data","Mass-produced real ground truth for reflection removal","5,300 real photo pairs solve reflection data scarcity","AI plus human edits yield 5,300 pixel-aligned reflection pairs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000855,"raw_usage":{"total_tokens":3736,"prompt_tokens":988,"completion_tokens":2748,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":604,"completion_tokens_details":{"reasoning_tokens":2668}},"tokens_in":604,"tokens_out":2748,"duration_ms":27517,"temperature":1.0,"reasoning_tokens":2668,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:19:57.882666+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect scenes where the true clean image is independently known (for example, photograph the same subject through glass and again after removing the glass, or use a black-cloth capture), run the paper's AI-plus-editing protocol on the through-glass images, and measure per-pixel differences between the produced 'clean' images and the true clean images. If systematic discrepancies appear in regions such as text, fine textures, or faces that the commercial tool tends to blur or hallucinate, the central assumption that OpenRR-5k pairs are faithful ground truth would be refuted.","supporting_citations":[{"cited_title":"Benchmarking single-image reflection removal algorithms,","cited_arxiv_id":null,"evidence_quote":"The 454-image SIR2 benchmark used as a cross-dataset evaluation target."},{"cited_title":"Single image reflection separation with perceptual losses,","cited_arxiv_id":null,"evidence_quote":"The Real dataset with 89/20 pairs used as a cross-dataset evaluation target."},{"cited_title":"Single image reflection removal through cascaded refinement,","cited_arxiv_id":null,"evidence_quote":"The Nature dataset with 200/20 pairs used as a cross-dataset evaluation target and as the main example of physically captured semi-synthetic data."},{"cited_title":"Revisiting single image reflection removal in the wild,","cited_arxiv_id":null,"evidence_quote":"The RRW dataset with 14,952 pairs, the larger in-the-wild comparison whose physical capture pipeline the paper argues is impractical."},{"cited_title":"Simple baselines for image restoration,","cited_arxiv_id":null,"evidence_quote":"The NAFNet image-restoration baseline that the paper adapts and trains to validate the dataset."},{"cited_title":"A categorized reflection removal dataset with diverse real-world scenes,","cited_arxiv_id":null,"evidence_quote":"Documents pixel-alignment and residual-reflection problems in semi-synthetic capture, motivating the new protocol."}],"review_version":1}