{"id":"3aaa3271-17da-43bc-8aed-e912752fad17","arxiv_id":"2506.08299","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"OpenRR-1k is a 1,000-pair real-world reflection removal dataset generated by off-the-shelf AI removal plus manual refinement, with benchmarks showing fine-tuning gains.","lead":"Researchers collected 1,000 real-world photo pairs to train AI models that remove reflections from glass, using a phone's built-in reflection removal tool plus manual cleanup to create the 'reflection-free' versions. The paper argues this pipeline is cheaper and more scalable than existing ways of building such datasets, and that fine-tuning models on it improves their reflection removal.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground-truth labels come from an unvalidated OPPO AI tool plus subjective cleanup; Table 2 shows input beats all S1 methods on OpenRR-1k, so fine-tuning gains may just fit the AI-generated label mapping rather than true reflection removal.","rationale":"The reader's weakest assumption—that the OPPO AI output plus manual Photoshop cleanup yields accurate ground-truth transmission images—is indeed the load-bearing point. My stress-test adds a specific quantitative red flag from Table 2: on the OpenRR-1k validation and test sets, the unprocessed input image has higher PSNR/SSIM and lower LPIPS than all S1 methods, while on external benchmarks the input scores much lower relative to methods. This is exactly what one would expect if the OpenRR-1k labels are close to the inputs, either because strong reflections remain in the 'clean' label or because the AI tool's mapping is what the metric rewards. The fine-tuning gains in S2 are therefore not independently validated; they could reflect learning the OPPO tool's particular output distribution rather than true reflection removal. The paper also lacks any external validation of label quality, inter-annotator agreement, or comparison with a physically captured clean image. My proposed test—collecting a small new set with an independent physical ground truth—would directly settle whether the protocol produces valid labels. Since the reader already reached a CONDITIONAL verdict on essentially these grounds, I do not recommend changing the verdict; the condition should be explicit release of such validation or independent benchmark results.","tokens_in":7910,"tokens_out":4966,"duration_ms":63540,"concrete_test":"Collect 50 new scenes with the same no-restriction protocol, but for each scene also capture an independent ground truth by physically removing the reflective glass (or using a polarizing filter/second camera) to record true transmission. Compute PSNR/SSIM/LPIPS between (a) the OPPO+manual label and this true transmission and (b) the original input and this true transmission. If (a) is not substantially better than (b) (e.g., <3 dB PSNR gain) or if the label contains visible residual reflections in >5% of scenes by independent expert review, the label-generation protocol fails to produce valid ground truth, and the fine-tuning improvements on OpenRR-1k cannot be interpreted as real reflection-removal gains.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 describes the entire label-generation process as: run OPPO's commercial AI reflection-removal software, then manually edit in Photoshop/MeituPic to 'eliminate any remaining reflection residues.' The only quality check is three subjective criteria (Cleanliness, Artifacts, Overall Image Quality) with no inter-annotator agreement, no quantitative residual-reflection measurement, and no comparison against a physically captured clean transmission image. This matters because the central empirical claim—fine-tuning on OpenRR-1k improves reflection removal by >7 dB PSNR—is measured on test labels produced by the same OPPO+manual pipeline. The internal metric is therefore partially circular: a model can score higher simply by reproducing the OPPO tool's output distribution. Table 2 already contains a red flag: on OpenRR-1k val/test the unprocessed input image achieves PSNR 26.37/26.20 and SSIM 0.943/0.939, higher than every S1 method, whereas on Nature/Real/SIR2 input scores are much lower (20.46/19.07/22.76 PSNR). This asymmetry is consistent with OpenRR-1k labels being much closer to the input than true reflection-free transmissions, e.g., because the AI tool leaves reflections partially intact or because manual cleanup over-smooths and the evaluator rewards conservative outputs. If labels are biased, the dataset's claimed value for training and benchmarking is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a scalable protocol for constructing real-world reflection-removal training data. Instead of physically capturing clean transmission images, the protocol runs OPPO's commercial AI reflection-removal tool on in-the-wild blended images and then manually cleans residual artifacts with Photoshop/MeituPic, yielding 1,000 pixel-aligned transmission-reflection pairs (OpenRR-1k). The authors benchmark four existing SIRR methods plus a NAFNet-based baseline under two strategies: S1 (off-the-shelf or baseline training without OpenRR-1k) and S2 (fine-tuning on OpenRR-1k train). They report that S2 improves performance, most dramatically on the OpenRR-1k validation/test sets, with the baseline's PSNR increasing by over 7 dB, and conclude that the dataset improves generalization in challenging real-world environments. The dataset is publicly released and the paper includes comparisons on Nature, Real, and SIR2.","tokens_in":8277,"tokens_out":4895,"duration_ms":64881,"significance":"If the label-quality concerns are resolved, OpenRR-1k would be a useful contribution: it is larger and more diverse than Nature/Real, its pairs are aligned by construction, and its release plus the S1/S2 benchmark provide a common evaluation protocol. The paper honestly reports that all S1 methods underperform the input image on OpenRR-1k, which is a thought-provoking negative result. The external benchmark gains after fine-tuning (e.g., DSRNet's Nature PSNR rising from 25.07 to 26.33) offer non-circular evidence that the data may transfer. However, the central claim of 'high-quality ground truth' is not validated: the labels are generated by a commercial AI model plus subjective manual cleanup, with no independent quantitative error analysis, no inter-annotator checks, and no comparison against physically captured clean transmissions. Because the training labels, test labels, and the AI tool itself all come from the same source, the dataset's value for benchmarking and training is currently not established.","major_comments":[{"comment":"The ground-truth transmission images are produced solely by OPPO's AI reflection-removal software followed by manual cleanup in Photoshop/MeituPic. The paper asserts that 'the final processed images are of high quality and suitable for training and evaluation purposes,' but it provides no quantitative validation: no comparison against independently captured clean transmission images, no measurement of residual reflections, no inter-annotator agreement on the three subjective criteria (Cleanliness, Artifacts, Overall Image Quality). Since the dataset's claimed high quality is load-bearing for both the training and benchmarking claims, this missing validation is a central issue. Please add a physically captured ground-truth subset (e.g., black-cloth or two-shot captures), report residual-reflection statistics, and quantify annotator agreement.","section":"Section 3.1"},{"comment":"On OpenRR-1k val/test, the Input Image row (PSNR 26.37/26.20, SSIM 0.943/0.939, LPIPS 0.077/0.078) is better than every S1 method, whereas on Nature/Real/SIR2 the input PSNRs are much lower (20.46/19.07/22.76). This asymmetry is consistent with the OpenRR-1k labels being much closer to the input than true reflection-free transmissions, for example because the AI tool under-removes reflections or the manual cleanup favors conservative edits. The paper's interpretation that existing methods 'perform poorly' on true real-world data is only one possible explanation; the alternative is that the labels are biased toward the input. Please disambiguate these hypotheses, for instance by measuring residual reflection energy in the labels or by validating on a physically captured subset.","section":"Table 2 and Section 4.3"},{"comment":"The S2 versus S1 comparison is not a controlled experiment. For the proposed baseline, S1 uses 60 training epochs, while S2 adds 100 additional epochs on OpenRR-1k; for the pretrained methods, S2 fine-tuning also adds extra optimization. Thus the reported improvements, including the headline 'PSNR increased by over 7 dB' on OpenRR-1k, could be due to additional training steps rather than to the dataset's content. Moreover, the OpenRR-1k test labels are generated by the same OPPO+manual pipeline as the training labels, making the gain partly circular: a model can improve by reproducing the label-generation distribution. Please include a control that trains S1 for the same total number of epochs without OpenRR-1k data, and report gains on external benchmarks separately with confidence intervals or paired significance tests.","section":"Sections 4.2 and 4.3"},{"comment":"The protocol claims 'perfectly aligned' / 'pixel-level alignment' because no physical capture setup is needed. However, the OPPO AI tool and the subsequent manual editing can modify image content (inpainting, over-smoothing, color shifts), not just remove reflections. Alignment of the glass-rendering geometry is therefore not the only possible source of misalignment; content-level changes can make the label inconsistent with the transmission layer. Please provide evidence that the manual and AI editing do not introduce systematic content changes, for example by reporting per-pixel alignment statistics on a subset with physically captured ground truth or by showing that the edited labels preserve the original texture frequencies.","section":"Section 3.1"}],"minor_comments":[{"comment":"There are several typographical issues, including 'DA TASET' in the title and 'Speficially' in Section 3.1; please proofread the manuscript.","section":"Throughout"},{"comment":"The table reports 'Pair Number' and average resolution for previous datasets, but the split information for RRW is not shown; adding the train/test split and license information would improve comparability.","section":"Table 1"},{"comment":"The 80/10/10 split is mentioned, but the paper does not state whether the split was random, stratified by scene category, or balanced by lighting condition; please clarify to help future users reproduce the benchmark.","section":"Section 3.2"},{"comment":"For the S2 experiments, the fine-tuning protocol for the pretrained methods is not described in enough detail (learning rates, number of epochs, data augmentation); please provide the exact fine-tuning settings so that the comparison is reproducible.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper has a useful dataset idea and a reasonable benchmark, but the central label-quality and circularity concerns are substantive. The authors appear to be from the same organization that produces the AI tool used to generate the ground truth, so independent validation is especially important. I recommend major revision rather than rejection because the core issues could be addressed with additional experiments and a careful re-analysis, but the current manuscript does not yet support the claimed 'high-quality' label."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper's core idea is genuinely new: instead of synthetic composites or controlled glass/cloth setups, they generate real-world transmission-reflection pairs by running OPPO's commercial AI reflection removal on in-the-wild reflection photos and then manually cleaning the output in Photoshop. That is a scalable, low-cost protocol, and it does produce perfectly aligned pairs by construction. The dataset of 1,000 pairs is a real resource, and the authors show that fine-tuning several existing methods on it improves PSNR/SSIM/LPIPS on their own test set and modestly on some external benchmarks (e.g., DSRNet on Real, RDNet on Nature). The paper is clearly written and the related work is handled honestly.\n\nBut the central claim that the labels are high-quality ground truth is not supported. The ground truth comes from the same AI tool plus subjective manual cleanup, with no quantitative error analysis, no inter-annotator agreement, and no comparison to a physically captured reflection-free image. That matters because the headline gain — over 7 dB on OpenRR-1k test — is measured against labels produced by that same pipeline. The model can score higher simply by learning the OPPO tool's output distribution. And Table 2 contains a red flag: the unprocessed input image achieves PSNR 26.20 and SSIM 0.939 on OpenRR-1k test, higher than every S1 method, while on Nature/Real/SIR2 the input scores much lower (20.46/19.07/22.76). That asymmetry strongly suggests the OpenRR-1k labels are closer to the input than true reflection-free transmissions, either because the AI tool leaves reflections partially intact or because the cleanup over-smooths and the metrics reward conservative outputs. Without external validation, the dataset's value for benchmarking is not established.\n\nThere is a sliver of evidence in the authors' favor: fine-tuning on OpenRR-1k also improves results on Real and Nature for most methods, which is harder to explain by label-fitting alone. But those gains are small and there is no control (e.g., fine-tuning on an equivalent-size synthetic dataset). So the empirical story is promising but not conclusive.\n\nWho should read it: anyone working on reflection removal dataset collection. It deserves a serious referee because the protocol is novel and the dataset could be valuable — but the referee should require quantitative label validation, a non-circular evaluation protocol, and a comparison with a controlled training data increase before taking the fine-tuning claims at face value.\n\nI'd send it to review, but with the expectation of major revision.","headline":"Novel scalable protocol for real-world reflection data, but the ground-truth labels are unvalidated and the evaluation is partly circular; deserves review with major revision.","tokens_in":8718,"tokens_out":2705,"would_cite":false,"duration_ms":29724,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces OpenRR-1k, a dataset of 1,000 pixel-aligned real-world transmission-reflection image pairs built by cleaning smartphone AI outputs, and shows that fine-tuning existing reflection removal models on it raises PSNR by…","keywords":["reflection removal","single image reflection removal","real-world dataset","pixel-aligned image pairs","dataset collection protocol","fine-tuning","benchmark evaluation"],"falsifier":"Select a random subset of OpenRR-1k test pairs, revisit the depicted locations, photograph the same scenes without any glass present, align those physically clean captures with the published transmission images, and measure their difference; if a substantial fraction of pairs show residuals or details that the no-glass captures lack (for example, PSNR well below 30 dB), the claim of high-quality, perfectly aligned ground truth fails.","tokens_in":7677,"feed_emoji":"🪞","tokens_out":6212,"duration_ms":73372,"temperature":0.7,"pith_summary":"This paper claims that the main bottleneck in single-image reflection removal is data, not model architecture, and proposes a cheap way to get that data: photograph real reflections, run each photo through a commercial AI reflection-removal tool, then manually clean up residual reflections in an image editor. The result is OpenRR-1k, 1,000 perfectly pixel-aligned pairs of blended and transmission images drawn from everyday scenes with varied lighting and glass surfaces. The authors show that after fine-tuning on this dataset, all five tested methods improve, with their own baseline gaining over 7 dB PSNR on the OpenRR-1k test set and also improving on older benchmarks. If the labels are trustworthy, the claim that matters is that the field's data bottleneck can be removed at low cost and at scale.","feed_headline":"1,000 real-world pairs lift reflection removal by 7+ dB","feed_subtitle":"Fine-tuning on OpenRR-1k also improves older-benchmark scores, evidence that the real-world pairs add useful signal.","key_machinery":"The central mechanism is the clean-the-blend collection protocol: instead of capturing a ground-truth transmission image by removing glass or covering the scene with black cloth, the protocol starts with a real blended photo, applies a commercial AI reflection-removal tool to get a rough clean image, and then uses manual image-editing cleanup to remove residual reflections and artifacts. This guarantees pixel-level alignment because the ground truth is edited from the exact same image as the input, and it permits arbitrary real-world capture, yielding diversity in reflections, scenes, and lighting. The paper also introduces a NAFNet-based baseline with an expanded bottleneck of 12 middle blocks, used both to benchmark the dataset and to demonstrate the fine-tuning gains.","core_discovery":"The paper's central claim is that high-quality, perfectly aligned, in-the-wild reflection removal training data can be produced without physically removing glass or blocking light: take real photos containing reflections, use an off-the-shelf AI reflection removal tool to produce an initial clean image, then manually edit away residual reflections with professional image editing software. Because the transmission image is derived from the same photograph as the blended image, the pairs are aligned by construction, and because no capture restrictions are imposed, the data reflect natural diversity in lighting, glass type, and scene content. Using this protocol the authors assembled OpenRR-1k with 1,000 pairs split 800/100/100 for training, validation, and testing. Benchmarking four existing reflection removal methods plus their own NAFNet-based baseline, they find that all methods perform noticeably worse on OpenRR-1k than on older datasets when using pretrained weights, but after fine-tuning on OpenRR-1k's training set performance improves across almost all methods and benchmarks, with their baseline's PSNR rising from 24.15 dB to 31.93 dB on the OpenRR-1k test set.","pith_inferences":["A consequence the authors leave implicit: since the ground truth is produced by a commercial AI tool plus manual editing, the dataset's quality ceiling is set by that tool's ability to reconstruct the true scene, and any systematic hallucination or smoothing it introduces could be baked into the labels.","A testable extension would be to independently verify label fidelity by physically capturing the same scenes without glass for a subset of pairs and measuring agreement with the published transmission images.","The protocol could transfer to other layered-image tasks such as dehazing or shadow removal wherever a commercial or heuristic cleaner produces a plausible first estimate that human annotators can refine."],"forward_implications":["Existing single-image reflection removal methods generalize poorly to OpenRR-1k's real-world images without fine-tuning, so their strong scores on older synthetic and semi-synthetic benchmarks have overstated real-world robustness.","Fine-tuning on OpenRR-1k improves all five tested methods on the OpenRR-1k test set and often on older benchmarks, indicating the dataset provides useful training signal, not just a harder evaluation set.","Because the protocol needs no specialized equipment, controlled lighting, or physical manipulation of glass, it can be crowdsourced and scaled beyond 1,000 pairs.","The fixed 80/10/10 split gives the community a standard benchmark for comparing future real-world reflection removal methods."],"supporting_citations":[{"why":"Supplies the largest prior real-world reflection removal dataset (RRW) and the RDNet-RRNet baseline method, providing the main comparison for scale, protocol, and performance.","marker":"[10]"},{"why":"Provides the ERRNet baseline method and the training recipe (60 epochs, Adam, learning rate schedule) that the paper adapts for its own baseline.","marker":"[15]"},{"why":"Supplies the DSRNet state-of-the-art baseline used in the benchmark evaluation.","marker":"[16]"},{"why":"Supplies the RAGNet baseline method evaluated on OpenRR-1k and older benchmarks.","marker":"[17]"},{"why":"Provides the NAFNet architecture that the paper modifies for its baseline model by increasing bottleneck blocks from 1 to 12.","marker":"[18]"},{"why":"Supplies the synthetic data generation approach used to create part of the training data for the baseline model.","marker":"[3]"},{"why":"Supplies the Nature dataset used as a real-world benchmark and as part of the real training data, and represents the prior manual glass-removal collection approach.","marker":"[4]"},{"why":"Supplies the Real dataset used as a benchmark and real training source, representing an earlier semi-synthetic collection method.","marker":"[7]"},{"why":"Supplies the SIR2 benchmark dataset used to evaluate generalization to existing reflection removal test sets.","marker":"[6]"}],"fun_headline_variants":["OpenRR-1k: real-world data lifts reflection removal","Fine-tune on OpenRR-1k for better reflection removal","1,000 real pairs boost reflection removal by 7+ dB","New dataset sharpens reflection removal with real pairs","Real-world pairs improve reflection removal benchmarks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's labels are trustworthy as ground truth only if the commercial AI tool plus manual cleanup truly recovers the reflection-free scene, and the paper provides no independent verification that residual reflections or hallucinated details are absent from the transmission images.","fun_headline_variants_meta":{"raw":{"variants":["OpenRR-1k: real-world data lifts reflection removal","Fine-tune on OpenRR-1k for better reflection removal","1,000 real pairs boost reflection removal by 7+ dB","New dataset sharpens reflection removal with real pairs","Real-world pairs improve reflection removal benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000218,"raw_usage":{"total_tokens":1432,"prompt_tokens":933,"completion_tokens":499,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":419}},"tokens_in":549,"tokens_out":499,"duration_ms":5535,"temperature":1.0,"reasoning_tokens":419,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:13:18.817184+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Select a random subset of OpenRR-1k test pairs, revisit the depicted locations, photograph the same scenes without any glass present, align those physically clean captures with the published transmission images, and measure their difference; if a substantial fraction of pairs show residuals or details that the no-glass captures lack (for example, PSNR well below 30 dB), the claim of high-quality, perfectly aligned ground truth fails.","supporting_citations":[{"cited_title":"Ro- bust single image reflection removal against adversarial attacks,","cited_arxiv_id":null,"evidence_quote":"Supplies the largest prior real-world reflection removal dataset (RRW) and the RDNet-RRNet baseline method, providing the main comparison for scale, protocol, and performance."},{"cited_title":"Revisiting single image reflection removal in the wild,","cited_arxiv_id":null,"evidence_quote":"Provides the ERRNet baseline method and the training recipe (60 epochs, Adam, learning rate schedule) that the paper adapts for its own baseline."},{"cited_title":"Benchmarking Ultra-High-Definition Image Reflection Removal","cited_arxiv_id":"2308.00265","evidence_quote":"Supplies the DSRNet state-of-the-art baseline used in the benchmark evaluation."},{"cited_title":"Robust sepa- ration of reflection from multiple images,","cited_arxiv_id":null,"evidence_quote":"Supplies the RAGNet baseline method evaluated on OpenRR-1k and older benchmarks."},{"cited_title":"NTIRE 2025 challenge on single image reflection removal in the wild: Datasets, methods and results,","cited_arxiv_id":null,"evidence_quote":"Provides the NAFNet architecture that the paper modifies for its baseline model by increasing bottleneck blocks from 1 to 12."},{"cited_title":"Dataset Collection Protocol As shown in Fig","cited_arxiv_id":null,"evidence_quote":"Supplies the synthetic data generation approach used to create part of the training data for the baseline model."},{"cited_title":"conv2 2”, “conv3 2","cited_arxiv_id":null,"evidence_quote":"Supplies the Nature dataset used as a real-world benchmark and as part of the real training data, and represents the prior manual glass-removal collection approach."},{"cited_title":"Deep learning based end-to-end specular reflection re- moval for medical endoscopic images,","cited_arxiv_id":null,"evidence_quote":"Supplies the Real dataset used as a benchmark and real training source, representing an earlier semi-synthetic collection method."},{"cited_title":"Automatic acetowhite lesion segmenta- tion via specular reflection removal and deep attention network,","cited_arxiv_id":null,"evidence_quote":"Supplies the SIR2 benchmark dataset used to evaluate generalization to existing reflection removal test sets."}],"review_version":1}