Pith. sign in

REVIEW 4 major objections 4 minor 26 references

OpenRR-5k: A Large-Scale Benchmark for Reflection Removal in the Wild

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper claims that OpenRR-5k, a new dataset of 5,300 real-world, pixel-aligned reflection/clean image pairs produced with AI-assisted editing, supplies the large, diverse training data that single-image reflection removal models lack…

desk verdict OpenRR-5k is a large-scale dataset with a novel collection protocol, but the ground truth is unvalidated and circular, so the benchmark claims don't hold as submitted. read the letter →

arxiv 2506.05482 v1 pith:ZPWSFATC submitted 2025-06-05 cs.CV

classification cs.CV
keywords singleimagereflectionremovalreal-worlddatasetpixel-alignedpairsbenchmarkcollectionprotocolrestorationgroundtruthgeneration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

OpenRR-5k is a new large-scale benchmark for single image reflection removal (SIRR), containing 5,300 pixel-aligned pairs of reflection-contaminated and clean real-world images, split into 5,000 training, 300 validation, and 100 ground-truth-free test images. The paper's central claim is that this dataset, collected by passing real photos through a commercial AI reflection-removal tool and then manually refining the results, provides training data that improves both performance and generalization compared with existing reflection-removal datasets. What makes the claim worth attention is that it offers an inexpensive, scalable route to diverse real-world training pairs without synthetic blending or physical glass-removal setups. The authors support it by training a U-Net-style model on OpenRR-5k alone and reporting 26.59 dB PSNR on the validation set plus cross-dataset scores on Nature, Real, and SIR2. If the dataset works as claimed, it would let SIRR models train on thousands of authentic scenes rather than a few hundred controlled captures.

What carries the argument

The load-bearing mechanism is the paired-data generation pipeline: each training pair consists of an observed reflection image and a clean image obtained by applying a commercial AI reflection remover followed by manual fine-detail editing, so that the two images are aligned at the pixel level and can be used for direct supervised training. The second component is the dataset structure itself, with 5,000/300/100 train/validation/test splits and explicit coverage of scene types (humans, animals, objects, landscapes) and lighting conditions (daytime, nighttime, indoor), designed so that models trained on the training split face realistic in-the-wild variation at evaluation time. The empirical argument is carried by a NAFNet-based baseline: trained only on OpenRR-5k, it reaches 26.59 dB PSNR on its own validation set and the cross-dataset PSNR figures above, which the paper reads as evidence of improved generalization without any additional training data.

What would settle it

Collect scenes where the true clean image is independently known (for example, photograph the same subject through glass and again after removing the glass, or use a black-cloth capture), run the paper's AI-plus-editing protocol on the through-glass images, and measure per-pixel differences between the produced 'clean' images and the true clean images. If systematic discrepancies appear in regions such as text, fine textures, or faces that the commercial tool tends to blur or hallucinate, the central assumption that OpenRR-5k pairs are faithful ground truth would be refuted.

Watch

Extended reading notes

Core claim

The central discovery the paper argues for is that a high-quality SIRR training set can be mass-produced from real-world photographs: a proven commercial AI reflection remover produces an initial clean estimate, human editors refine away residual reflections and artifacts, and the refined image is paired with the original photograph as pixel-aligned ground truth. According to the paper, this protocol removes the pixel-level misalignment that plagues physically captured pairs and the domain gap that plagues fully synthetic pairs, while enabling data collection at a scale and diversity that previous real-world datasets do not match. The supporting experiment trains a NAFNet-based restoration network on OpenRR-5k and reports PSNR 26.59 on OpenRR-5kval, together with 25.62 on Nature, 21.16 on Real, and 24.52 on SIR2, which the authors present as evidence that the dataset 'enhances the performance and generalization' of reflection removal models.

Load-bearing premise

The load-bearing premise is that the clean images produced by the commercial AI reflection remover plus manual editing are faithful approximations of the true reflection-free scenes, so that models trained on them learn real reflection removal rather than the tool's own artifacts.

Editorial extensions

If this is right

  • Training a SIRR model on OpenRR-5k rather than on earlier small real-world datasets should improve transfer to unseen reflection scenes, because the training pairs come from genuine photographs across many lighting conditions and glass types.
  • The 300 validation pairs give the community a pixel-aligned real-world benchmark for comparable evaluation, while the 100 no-ground-truth test images allow practical performance checks on fully unlabeled scenes.
  • The collection protocol can be scaled further through crowdsourcing, since it requires no specialized capture equipment, meaning larger benchmarks of the same type are straightforward to build.
  • Because OpenRR-5k's training split includes all OpenRR-1k pairs, results on the new benchmark directly extend and are comparable with the earlier 1k dataset.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors do not run a matched comparison against RRW's 14,952 pairs, so a direct test of whether OpenRR-5k's advantage comes from diversity, alignment, or sheer scale would be to train the same baseline on both datasets under identical settings and evaluate on a shared test set.
  • Because the ground truth is produced by a commercial AI plus human editing, the dataset implicitly teaches models to reproduce that tool's notion of a clean image; systematic errors in the tool, such as over-smoothing text or fine texture, could be inherited by trained models, a risk that perceptual or human-preference studies on the 100 no-GT test images could expose.
  • The same edit-the-AI-output protocol could be transferred to other ill-posed restoration tasks, such as dehazing or low-light enhancement, where true clean targets are similarly hard to capture, making the approach a general template for dataset construction.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces OpenRR-5k, a dataset of 5,300 real-world reflection-contaminated images paired with clean counterparts for single image reflection removal (SIRR). The clean images are generated by applying the OPPO smartphone AI reflection removal tool and then manually refining the results with image editing software (Section III.A). The authors train a NAFNet-based baseline on the 5,000 training pairs and report quantitative metrics (PSNR, SSIM, LPIPS, DISTS, NIQE) on the 300 validation pairs plus PSNR on three external datasets (Nature, Real, SIR2). The central claim is that OpenRR-5k provides a large-scale, pixel-aligned, diverse real-world dataset that enhances the performance and generalization of SIRR models. The paper also provides a 100-image test set without ground truth for qualitative evaluation.

Significance. If the dataset's label-generation pipeline were validated, the scale and diversity of OpenRR-5k could make it a valuable training resource for the SIRR community. The proposed protocol is convenient and scalable compared to physical capture setups. However, the paper provides no evidence that the AI-generated clean images faithfully represent true reflection-free scenes, and the experimental evaluation is too sparse to support the claimed benefits. The utility of the dataset therefore rests entirely on an unvalidated, likely circular supervision signal. The authors promise to release the dataset and code, which would allow independent checks, but as submitted the manuscript does not demonstrate that the dataset is a trustworthy benchmark.

major comments (4)
  1. [III.A] The ground-truth generation pipeline, which uses the OPPO AI reflection remover followed by manual editing, is not validated. There are no quantitative comparisons against physically captured clean images, no residual-artifact statistics, no inter-editor consistency analysis, and no before/after evaluation. This is the load-bearing assumption of the entire paper: if the AI-generated clean images contain hallucinations or smoothing artifacts, models trained on this dataset will learn to reproduce those artifacts rather than perform true reflection removal. The absence of any such validation makes the dataset's supervision unreliable.
  2. [IV.A, Table II] The claim that the dataset 'effectively enhances the performance and generalization' of SIRR models is unsupported by the reported experiments. Table II lists only NAFNet trained on OpenRR-5k evaluated on Nature, Real, and SIR2, with no control condition: we do not see the same architecture trained on existing datasets (e.g., RRW, Nature, or Real) and evaluated on the same external test sets, nor do we see comparisons with prior state-of-the-art methods. Without these baselines, the PSNR values of 25.62, 21.16, and 24.52 cannot be interpreted as evidence of improvement or generalization.
  3. [IV.A, Tables II and III] The validation set (OpenRR-5kval) is generated by the same OPPO-AI-plus-manual-editing pipeline as the training set. Therefore, the reported PSNR of 26.59 dB on OpenRR-5kval measures the model's consistency with the commercial tool's output distribution, not its fidelity to true reflection-free scenes. A meaningful evaluation of the dataset's utility requires held-out real-world pairs with physically captured ground truth, which are absent from this paper.
  4. [III.A, item 3] The claim that the dataset provides 'True Real-World Data' is overstated. While the input reflection images are real photographs, the clean ground-truth images are the output of a proprietary AI model plus manual editing, not real captures. The statement that the method 'eliminates the need for collecting ground-truth data' conflates avoiding physical capture with obtaining authentic ground truth. This wording could mislead readers about the nature of the supervision.
minor comments (4)
  1. [References] The OpenRR-1k dataset is cited inconsistently: reference [8] is used in the text on multiple occasions, but Table I cites it as [10]. Please unify the citations.
  2. [Abstract and III.B] The abstract says the test set consists of '100 real-world testing images without ground truth,' while Section III.B says '100 image pairs without Ground Truth for the test set.' Please clarify whether the test set contains images or pairs.
  3. [IV.B] The implementation details do not specify the loss function used to train NAFNet, the total number of parameters, or the exact architecture modifications beyond increasing block counts. These details are needed for reproducibility.
  4. [Fig. 3] The pie chart percentages for scene content (60% landscapes, 7% animals, 14% humans, 19% objects) sum to 100%, but the figure is not visible in the submitted text; please ensure it is legible and the caption clearly defines the categories.

Circularity Check

1 steps flagged · score 7.0 of 10

OpenRR-5k's clean labels and validation targets are both generated by the same OPPO-AI-plus-manual-editing pipeline, so the reported validation metrics measure consistency with that tool rather than true reflection-free fidelity.

  1. self definitional [Section III.A (Dataset Collection Protocol); Section IV.A / Table III]
    "We adopted the OPPO smartphone’s AI-based reflection removal software to obtain the initial reflection removal results. ... After precise manual adjustments, the final processed images are of high quality and suitable for training and evaluation purposes. ... Our method eliminates the need for collecting ground-truth data, allowing us to capture authentic reflection scenarios directly from real-world environments."

    The 'clean version' in every OpenRR-5k pair is not an independent measurement of the reflection-free transmission layer; it is the output of a reflection-removal system (OPPO AI plus manual editing). The same protocol is used to produce the 300 validation pairs. Therefore the NAFNet baseline is trained to approximate this pipeline and then evaluated on the same pipeline's outputs: PSNR 26.59, SSIM, LPIPS, etc. on OpenRR-5kval measure agreement with the commercial tool's output distribution, not fidelity to a true scene. The paper supplies no captured clean images, no residual-artifact statistics, and no cross-check showing the edited AI output equals ground truth; it defines 'ground truth' as the pipeline result.

full rationale

The paper's central claim is that OpenRR-5k is a large-scale, high-quality benchmark that improves reflection-removal training and generalization. That claim rests entirely on Section III.A's assumption that OPPO-AI-plus-manual-editing outputs are faithful clean transmission images. Because both the training targets and the validation targets come from this same unverified pipeline, the headline validation metrics in Table III are self-referential: a model trained to predict the pipeline's output is scored against the pipeline's own output. The external Nature/Real/SIR2 evaluations in Table II are independent of the OPPO pipeline, but the paper provides no control—such as the same NAFNet trained on RRW or another existing dataset and evaluated on the same external sets—so those numbers cannot establish the stated claim that the dataset 'enhances the performance and generalization' of SIRR models. The self-citations to OpenRR-1k are not load-bearing here: the circularity is in the label-generation protocol itself, not in a citation chain. Overall, the benchmark's central validation evidence is defined by the same process it purports to evaluate, a substantial but partial circularity because external evaluation sets do provide some independent signal.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim rests on the assumption that AI-generated plus manually edited clean images are valid ground truth, and that alignment is preserved. No independent evidence is provided. There are no fitted free parameters in the paper's framework; the proprietary reflection removal model acts as an external black box defining the target.

assumptions (3)
  • domain assumption OPPO AI reflection removal plus manual editing yields faithful clean ground truth.
    Section III.A: the clean versions are generated by a commercial tool and human editors, never validated against true reflection-free captures.
  • domain assumption Pixel-level alignment holds between the reflection image and the processed clean image.
    Section III.A point 2 asserts alignment but provides no alignment metric or qualitative verification.
  • domain assumption The quality of the manual refinement is sufficient for training and evaluation.
    No inter-annotator agreement or quality control for the editing step is reported, yet the final images are used as supervision.

how reviews work

0 comments
Cite this review

Pith. "Pith review of OpenRR-5k: A Large-Scale Benchmark for Reflection Removal in the Wild." pith.science (2026). https://pith.science/paper/ZPWSFATC

@misc{pith2026250605482,
  author       = {Pith},
  title        = {Pith review of: OpenRR-5k: A Large-Scale Benchmark for Reflection Removal in the Wild},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZPWSFATC}},
  note         = {Machine review of arXiv:2506.05482}
}
read the original abstract

Removing reflections is a crucial task in computer vision, with significant applications in photography and image enhancement. Nevertheless, existing methods are constrained by the absence of large-scale, high-quality, and diverse datasets. In this paper, we present a novel benchmark for Single Image Reflection Removal (SIRR). We have developed a large-scale dataset containing 5,300 high-quality, pixel-aligned image pairs, each consisting of a reflection image and its corresponding clean version. Specifically, the dataset is divided into two parts: 5,000 images are used for training, and 300 images are used for validation. Additionally, we have included 100 real-world testing images without ground truth (GT) to further evaluate the practical performance of reflection removal methods. All image pairs are precisely aligned at the pixel level to guarantee accurate supervision. The dataset encompasses a broad spectrum of real-world scenarios, featuring various lighting conditions, object types, and reflection patterns, and is segmented into training, validation, and test sets to facilitate thorough evaluation. To validate the usefulness of our dataset, we train a U-Net-based model and evaluate it using five widely-used metrics, including PSNR, SSIM, LPIPS, DISTS, and NIQE. We will release both the dataset and the code on https://github.com/caijie0620/OpenRR-5k to facilitate future research in this field.

Figures

Figures reproduced from arXiv: 2506.05482 by the authors.

Figure 1
Figure 1. Visualization of paired data generation pipeline for reflection removal. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Overview of our OpenRR-5k dataset. TABLE I: Comparison of Existing Datasets with Our OpenRR-5k Dataset Dataset Year Usage Pair Number Average Resolution SIR2 [19] 2017 Test 454 540 x 400 Real [20] 2018 Train/Test 89/20 1152 x 930 Nature [5] 2020 Train/Test 200/20 598 x 398 RRW [23] 2023 Train 14952 2580 × 1460 OpenRR-1k [10] 2025 Train/Val/Test 800/100/100 922 x 917 OpenRR-5k 2025 Train/Val/Test 5,000/300/100 874 x … view at source ↗
Figure 3
Figure 3. The category distribution of our OpenRR-5ktest dataset TABLE II: Quantitative Comparisons of Real-World Reflection Removal Datasets Method Nature (20) Real (20) SIR2 (454) OpenRR-5kval (300) NAFNet 25.62 21.16 24.52 26.59 TABLE III: Comprehensive Quantitative Comparisons on OpenRR-5kval Metrics PSNR SSIM LPIPS DISTS NIQE NAFNet 26.59 0.9418 0.0911 0.0538 3.4066 scores indicating better performance. This simple train… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

26 extracted references · 20 canonical work pages

  1. [1]

    Online multi-object tracking using multi-function integration and tracking simulation training,

    J. Yang, H. Ge, J. Yang, Y . Tong, and S. Su, “Online multi-object tracking using multi-function integration and tracking simulation training,” Applied Intelli- gence, 2022

  2. [2]

    Modern approaches to augmented reality,

    O. Bimber and R. Raskar, “Modern approaches to augmented reality,” in ACM SIGGRAPH 2006 Courses, 2006

  3. [3]

    Deep learning based end-to-end specular reflection removal for medical endoscopic images,

    Chi-Sheng Shih, Yu-Cheng Liao, and Ching-Ting Tan, “Deep learning based end-to-end specular reflection removal for medical endoscopic images,” in Proceed- ings of the 2023 International Conference on Research in Adaptive and Convergent Systems , 2023

  4. [4]

    A generic deep architecture for single image reflection removal and image smoothing,

    Q. Fan, J. Yang, G. Hua, B. Chen, and D. Wipf, “A generic deep architecture for single image reflection removal and image smoothing,” in ICCV, 2017

  5. [5]

    Single image reflection removal through cascaded refinement,

    C. Li, Y . Yang, K. He, S. Lin, and J. E Hopcroft, “Single image reflection removal through cascaded refinement,” in CVPR, 2020

  6. [6]

    Robust single image reflection removal against adversarial attacks,

    Z. Song, Z. Zhang, K. Zhang, W. Luo, Z. Fan, W. Ren, and J. Lu, “Robust single image reflection removal against adversarial attacks,” in CVPR, 2023

  7. [7]

    Language-guided image reflection separation,

    H. Zhong, Y . Hong, S. Weng, J. Liang, and B. Shi, “Language-guided image reflection separation,” in CVPR, 2024, pp. 24913–24922

  8. [8]

    Ntire 2025 challenge on single image reflection removal in the wild: Datasets, methods and results,

    K. Yang, J. Cai, L. Ouyang, F. Vasluianu, R. Timofte, et al., “Ntire 2025 challenge on single image reflection removal in the wild: Datasets, methods and results,” in CVPR Workshops, 2025

Show all 26 references
  1. [9]

    Survey on single-image reflection removal using deep learning techniques,

    K. Yang, H. Sun, J. Cai, L. Fu, J. Ding, J. Li, and Z. Meng, “Survey on single-image reflection removal using deep learning techniques,” in MIPR, 2025

  2. [10]

    Openrr-1k: A scalable dataset for real-world reflection removal,

    K. Yang, L. Ouyang, H. Sun, J. Cai, L. Fu, J. Ding, C. M. Ho, and Z. Meng, “Openrr-1k: A scalable dataset for real-world reflection removal,” in ICIP, 2025

  3. [11]

    F2t2-hit: A u-shaped fft transformer and hierarchical transformer for reflection removal,

    J. Cai, K. Yang, L. Ouyang, L. Fu, J. Ding, H. Sun, C. M. Ho, and Z. Meng, “F2t2-hit: A u-shaped fft transformer and hierarchical transformer for reflection removal,” in ICIP, 2025

  4. [12]

    Degradation-aware image enhancement via vision-language classification,

    J. Cai, K. Yang, J. Ding, L. Fu, L. Ouyang, J. Li, J. Shen, and Z. Meng, “Degradation-aware image enhancement via vision-language classification,” in MIPR, 2025

  5. [13]

    Reversible de- coupling network for single image reflection removal,

    H. Zhao, M. Li, Q. Hu, and X. Guo, “Reversible de- coupling network for single image reflection removal,” in CVPR, 2025

  6. [14]

    Sin- gle image reflection removal exploiting misaligned training data and network enhancements,

    K. Wei, J. Yang, Y . Fu, D. Wipf, and H. Huang, “Sin- gle image reflection removal exploiting misaligned training data and network enhancements,” in CVPR, 2019

  7. [15]

    User assisted separation of reflections from a single image using a sparsity prior,

    Anat Levin and Yair Weiss, “User assisted separation of reflections from a single image using a sparsity prior,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2007

  8. [16]

    Crrn: Multi-scale guided concurrent reflection removal network,

    Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, and Alex C Kot, “Crrn: Multi-scale guided concurrent reflection removal network,” in CVPR, 2018

  9. [17]

    Seeing deeply and bidirectionally: A deep learning approach for single image reflection removal,

    Jie Yang, Dong Gong, Lingqiao Liu, and Qinfeng Shi, “Seeing deeply and bidirectionally: A deep learning approach for single image reflection removal,” in ECCV, 2018

  10. [18]

    Location-aware single image reflection removal,

    Zheng Dong, Ke Xu, Yin Yang, Hujun Bao, Weiwei Xu, and Rynson WH Lau, “Location-aware single image reflection removal,” in ICCV, 2021

  11. [19]

    Benchmarking single-image reflection removal algorithms,

    Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, and Alex C Kot, “Benchmarking single-image reflection removal algorithms,” in ICCV, 2017, pp. 3922–3930

  12. [20]

    Single image reflection separation with perceptual losses,

    Xuaner Zhang, Ren Ng, and Qifeng Chen, “Single image reflection separation with perceptual losses,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018

  13. [21]

    Polarized reflection removal with perfect alignment in the wild,

    Chenyang Lei, Xuhua Huang, Mengdi Zhang, Qiong Yan, Wenxiu Sun, and Qifeng Chen, “Polarized reflection removal with perfect alignment in the wild,” in CVPR, 2020, pp. 1750–1758

  14. [22]

    A categorized reflection removal dataset with diverse real-world scenes,

    Chenyang Lei, Xuhua Huang, Chenyang Qi, Yankun Zhao, Wenxiu Sun, Qiong Yan, and Qifeng Chen, “A categorized reflection removal dataset with diverse real-world scenes,” in CVPR, 2022, pp. 3040–3048

  15. [23]

    Revisiting single image reflection removal in the wild,

    Y . Zhu, X. Fu, Peng-Tao Jiang, H. Zhang, Q. Sun, J. Chen, Zheng-Jun Zha, and B. Li, “Revisiting single image reflection removal in the wild,” in CVPR, 2024

  16. [24]

    Benchmarking ultra-high-definition image reflection removal,

    Zhenyuan Zhang, Zhenbo Song, Kaihao Zhang, Zhaoxin Fan, and Jianfeng Lu, “Benchmarking ultra-high-definition image reflection removal,” arXiv preprint arXiv:2308.00265, 2023

  17. [25]

    Robust separation of reflection from multiple images,

    Xiaojie Guo, Xiaochun Cao, and Yi Ma, “Robust separation of reflection from multiple images,” in CVPR, 2014, pp. 2187–2194

  18. [26]

    Simple baselines for image restoration,

    L. Chen, X. Chu, X. Zhang, and J. Sun, “Simple baselines for image restoration,” in ECCV, 2022

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.