Pith. sign in

REVIEW 2 major objections 2 minor 37 references

SR-Prominence: A Crowdsourced Protocol and Dataset Suite for Perceptually-Weighted Super-Resolution Artifact Evaluation

T0 review · 2 major / 2 minor · reviewed 2026-06-30 · grok-4.3

Pith's one-line read Classical full-reference metrics like SSIM and DISTS predict perceptual prominence of super-resolution artifacts more reliably than no-reference methods or specialized detectors.

desk verdict The paper defines artifact prominence via crowdsourcing and finds SSIM/DISTS track it better than other methods, but the labels lack any reported stability checks. read the letter →

arxiv 2605.14847 v1 pith:4C76CKFX submitted 2026-05-14 cs.CV

classification cs.CV
keywords super-resolutionartifactprominencecrowdsourcedevaluationperceptualimagequalityfull-referencemetricsSSIMDISTSdetection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper defines artifact prominence as the fraction of viewers who notice a defect in a highlighted region of a super-resolved image. It releases a crowdsourced dataset suite of 3,935 such annotated masks drawn from multiple sources, including a realistic setting without ground-truth references. Tests across the suite show that standard full-reference metrics supply localized signals that match human judgments of noticeability, while many other approaches do not hold up when the dataset or reference condition changes. The release includes an objective scoring protocol that lets new metrics be checked against the prominence labels without repeating human annotation. This changes evaluation from counting defects to weighing their visible effect on viewers.

What carries the argument

artifact prominence, defined as the fraction of viewers who judge a highlighted region to contain a noticeable artifact, used as the target measure for perceptual impact instead of binary defect presence

What would settle it

A fresh crowdsourced annotation round on the same image regions with a different viewer pool produces prominence values that diverge substantially from the original labels and from the predictions of SSIM or DISTS.

Watch

Extended reading notes

Core claim

Artifact prominence is defined as the fraction of viewers who judge a highlighted region to contain a noticeable artifact. A crowdsourced protocol produces the SR-Prominence dataset suite with 3,935 masks from DeSRA, Open Images, Urban100, and a no-ground-truth Urban100-HR setting. Re-annotation of DeSRA shows 48.2 percent of its prior binary artifacts are not noticed by a majority of viewers. Classical full-reference metrics, especially SSIM and DISTS, provide strong localized prominence signals while no-reference IQA methods and specialized artifact detectors fail to generalize across datasets and reference settings.

Load-bearing premise

The crowdsourced majority-vote prominence values remain stable when applied to new annotator pools and image sources beyond those used in the study.

Editorial extensions

If this is right

  • 48.2 percent of binary artifacts previously labeled in DeSRA are not noticed by a majority of viewers under the new protocol.
  • SSIM and DISTS can supply localized predictions of where artifacts will stand out to people.
  • The released objective scoring protocol lets any new metric be benchmarked on the suite without further crowdsourcing.
  • Super-resolution methods can be compared on the basis of perceptual impact rather than binary defect counts.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Training objectives for future super-resolution networks could use prominence estimates from SSIM or DISTS to suppress only the defects that most viewers notice.
  • The crowdsourced protocol could be applied to measure perceptual impact in related tasks such as image denoising or compression artifact evaluation.
  • The consistent failure of no-reference methods points to a need for metrics that better capture localized viewer attention without a clean reference image.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper introduces artifact prominence as the fraction of viewers who notice an artifact in a highlighted region of a super-resolved image. It presents a crowdsourced annotation protocol and the SR-Prominence dataset suite (3,935 masks from DeSRA, Open Images, Urban100, and no-ground-truth Urban100-HR), re-annotates DeSRA to find 48.2% of its binary artifacts unnoticed by a majority, audits detectors and metrics, and reports that classical full-reference metrics (especially SSIM and DISTS) yield strong localized prominence signals while no-reference IQA methods and specialized artifact detectors often fail to generalize. The suite is released with an objective scoring protocol for benchmarking without new crowdsourcing.

Significance. If the prominence labels are shown to be stable, the dataset and protocol would usefully shift SR artifact evaluation from binary presence to perceptual impact. The reported strength of SSIM and DISTS as localized signals is a concrete, falsifiable observation that could influence metric selection. The public release of the dataset together with a reproducible scoring protocol is a clear strength supporting community benchmarking.

major comments (2)
  1. [Section 3] Protocol description (Section 3 / Dataset Construction): no annotation instructions, inter-annotator agreement (Fleiss’ κ, Krippendorff’s α), bootstrap stability of majority votes, or cross-pool replication are reported. These checks are load-bearing for treating the prominence fractions as reliable ground truth, especially for the Urban100-HR no-GT subset.
  2. [Section 5] Metric audit results (Section 5): the claim that SSIM/DISTS provide strong localized signals while other methods fail to generalize rests directly on the crowdsourced prominence values as ground truth. Without the missing agreement and stability statistics, the reported superiority and generalization failures cannot be evaluated for robustness.
minor comments (2)
  1. [Abstract] Abstract and Section 4: the phrase 'objective scoring protocol' is used without a concrete description or pseudocode; adding a short formal definition would improve clarity.
  2. [Dataset release] Dataset release statement: confirming that raw per-annotator votes (not only aggregated prominence) are included would strengthen transparency and allow independent re-analysis.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their thoughtful review and for highlighting the importance of validating the crowdsourced protocol. We address each major comment below and commit to revisions that will strengthen the manuscript's claims regarding label reliability and metric evaluation.

read point-by-point responses
  1. Referee: [Section 3] Protocol description (Section 3 / Dataset Construction): no annotation instructions, inter-annotator agreement (Fleiss’ κ, Krippendorff’s α), bootstrap stability of majority votes, or cross-pool replication are reported. These checks are load-bearing for treating the prominence fractions as reliable ground truth, especially for the Urban100-HR no-GT subset.

    Authors: We agree these statistics are essential to substantiate the prominence labels as reliable ground truth. In the revised manuscript, we will include the complete annotation instructions in Section 3. We will compute and report inter-annotator agreement metrics including Fleiss’ κ and Krippendorff’s α. We will also add analyses of bootstrap stability for the majority votes and cross-pool replication results. For the Urban100-HR no-ground-truth subset, we will provide dedicated replication details to confirm consistency across annotation pools. revision: yes

  2. Referee: [Section 5] Metric audit results (Section 5): the claim that SSIM/DISTS provide strong localized signals while other methods fail to generalize rests directly on the crowdsourced prominence values as ground truth. Without the missing agreement and stability statistics, the reported superiority and generalization failures cannot be evaluated for robustness.

    Authors: We acknowledge that the metric audit results in Section 5 depend on the quality of the prominence ground truth. Incorporating the agreement, stability, and replication statistics as described in our response to the Section 3 comment will enable a more rigorous evaluation of the robustness of the findings on SSIM, DISTS, and the generalization failures of other methods. The revised manuscript will include these validations to support the claims. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity

full rationale

The paper introduces a new crowdsourced protocol and dataset (SR-Prominence) for artifact prominence, defined directly from viewer annotations as the fraction noticing artifacts in highlighted regions. No equations, fitted parameters, or predictions are described that reduce to the same inputs by construction. Central claims about metric performance (SSIM/DISTS vs. others) rest on these independent new labels across multiple datasets, including no-GT settings, without self-definitional loops, load-bearing self-citations, or renaming of known results. The derivation chain is self-contained as empirical data collection followed by external benchmarking.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

The work rests on the empirical assumption that crowdsourced majority votes reliably capture perceptual noticeability; no mathematical axioms, free parameters, or invented physical entities are invoked.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SR-Prominence: A Crowdsourced Protocol and Dataset Suite for Perceptually-Weighted Super-Resolution Artifact Evaluation." pith.science (2026). https://pith.science/paper/4C76CKFX

@misc{pith2026260514847,
  author       = {Pith},
  title        = {Pith review of: SR-Prominence: A Crowdsourced Protocol and Dataset Suite for Perceptually-Weighted Super-Resolution Artifact Evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4C76CKFX}},
  note         = {Machine review of arXiv:2605.14847}
}
read the original abstract

Modern image super-resolution methods generate detailed, visually appealing results, but they often introduce visual artifacts: unnatural patterns and texture distortions that degrade perceived quality. These defects vary widely in perceptual impact--some are barely noticeable, while others are highly disturbing--yet existing detection methods treat them equally. We propose artifact prominence as an evaluative target, defined as the fraction of viewers who judge a highlighted region to contain a noticeable artifact. We design a crowdsourced annotation protocol and construct SR-Prominence, a dataset suite containing 3,935 artifact masks from DeSRA, Open Images, Urban100, and a realistic no-ground-truth Urban100-HR setting, annotated with prominence. Re-annotating DeSRA reveals that 48.2% of its in-lab binary artifacts are not noticed by a majority of viewers. Across the suite, we audit SR artifact detectors, image-quality metrics, and SR methods. We find that classical full-reference metrics, especially SSIM and DISTS, provide surprisingly strong localized prominence signals, whereas no-reference IQA methods and specialized artifact detectors often fail to generalize across datasets and reference settings. SR-Prominence is released with an objective scoring protocol that allows new metrics to be benchmarked on our suite without further crowdsourcing. Together, the data and protocols enable SR artifact evaluation to move from binary defect presence toward perceptual impact. SR-Prominence is available at https://huggingface.co/datasets/imolodetskikh/sr-artifact-prominence.

Figures

Figures reproduced from arXiv: 2605.14847 by the authors.

Figure 1
Figure 1. SR-Prominence artifact examples. Rows show Open Images (top) and DeSRA (bottom) [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Viewer interface for subjective data collection. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Example of mask preprocessing for human visual assessment. An artifact-detection method should output a tight mask around an artifact, since such masks are more useful for subsequent analysis and downstream tasks such as automatic correction. However, tight masks make it harder to visually judge whether the masked area contains an artifact. Additionally, the raw output from some methods is sparse, making it extra ch… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Bootstrap-analysis results for an image with a highly prominent artifact (left) and barely [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: Architecture of the reference artifact-prominence baseline. The input image is upscaled [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: Example artifacts detected by the baseline. (a): low-resolution input image; (b): target SR [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]
Figure 7
Figure 7. Figure 7: Example false detections by the baseline. (a): low-resolution input image; (b): target SR [PITH_FULL_IMAGE:figures/full_fig_p019_7.png]
Figure 8
Figure 8. Figure 8: Example false detections by the baseline due to inaccurate restoration from pseudo-GT [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

37 extracted references · 37 canonical work pages

  1. [1]

    Scaling up to excellence: Practicing model scaling for photo-realistic image restoration in the wild,

    F. Yu, J. Gu, Z. Li, J. Hu, X. Kong, X. Wang, J. He, Y . Qiao, and C. Dong, “Scaling up to excellence: Practicing model scaling for photo-realistic image restoration in the wild,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 25669–25680, June 2024

  2. [2]

    Exploiting diffusion prior for real-world image super-resolution,

    J. Wang, Z. Yue, S. Zhou, K. C. Chan, and C. C. Loy, “Exploiting diffusion prior for real-world image super-resolution,”International Journal of Computer Vision, 2024

  3. [3]

    Details or artifacts: A locally discriminative learning approach to realistic image super-resolution,

    J. Liang, H. Zeng, and L. Zhang, “Details or artifacts: A locally discriminative learning approach to realistic image super-resolution,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5657–5666, 2022

  4. [4]

    arXiv preprint arXiv:2307.02457 , year=

    L. Xie, X. Wang, X. Chen, G. Li, Y . Shan, J. Zhou, and C. Dong, “DeSRA: detect and delete the artifacts of gan-based real-world super-resolution models,”arXiv preprint arXiv:2307.02457, 2023

  5. [5]

    Perceptual artifacts localization for image synthesis tasks,

    L. Zhang, Z. Xu, C. Barnes, Y . Zhou, Q. Liu, H. Zhang, S. Amirghodsi, Z. Lin, E. Shechtman, and J. Shi, “Perceptual artifacts localization for image synthesis tasks,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 7579–7590, October 2023. 10

  6. [6]

    Photo-realistic single image super-resolution using a generative adversarial network,

    C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang,et al., “Photo-realistic single image super-resolution using a generative adversarial network,” inProceedings of the IEEE conference on computer vision and pattern recognition, pp. 4681–4690, 2017

  7. [7]

    Real-ESRGAN: Training real-world blind super- resolution with pure synthetic data,

    X. Wang, L. Xie, C. Dong, and Y . Shan, “Real-ESRGAN: Training real-world blind super- resolution with pure synthetic data,” inProceedings of the IEEE/CVF international conference on computer vision, pp. 1905–1914, 2021

  8. [8]

    Perceptual artifacts localization for inpainting,

    L. Zhang, Y . Zhou, C. Barnes, S. Amirghodsi, Z. Lin, E. Shechtman, and J. Shi, “Perceptual artifacts localization for inpainting,”arXiv preprint arXiv:2208.03357, 2022

Show all 37 references
  1. [9]

    Hallucination score: Towards mitigating hallucinations in generative image super-resolution,

    W. Ren, R. Goyal, Z. Hu, T. T. Aumentado-Armstrong, I. Mohomed, and A. Levinshtein, “Hallucination score: Towards mitigating hallucinations in generative image super-resolution,” 2025

  2. [10]

    Image quality assessment: from error visibility to structural similarity,

    Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,”IEEE transactions on image processing, vol. 13, no. 4, pp. 600–612, 2004

  3. [11]

    The unreasonable effectiveness of deep features as a perceptual metric,

    R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” inProceedings of the IEEE conference on computer vision and pattern recognition, pp. 586–595, 2018

  4. [12]

    Image quality assessment: Unifying structure and texture similarity,

    K. Ding, K. Ma, S. Wang, and E. P. Simoncelli, “Image quality assessment: Unifying structure and texture similarity,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 5, pp. 2567–2581, 2022

  5. [13]

    ERQA: Edge-restoration quality as- sessment for video super-resolution,

    A. Kirillova, E. Lyapustin, A. Antsiferova, and D. Vatolin, “ERQA: Edge-restoration quality as- sessment for video super-resolution,” inProceedings of the 17th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications - V olume 4:...

  6. [14]

    MSU video super-resolution quality metrics bench- mark

    A. Borisov, E. Bogatyrev, and D. Vatolin, “MSU video super-resolution quality metrics bench- mark.” https://videoprocessing.ai/benchmarks/super-resolution-metrics. html, 2025. Accessed: 2026-01-28

  7. [15]

    PIPAL: A large-scale image quality assessment dataset for perceptual image restoration,

    G. Jinjin, C. Haoming, C. Haoyu, Y . Xiaoxing, J. S. Ren, and D. Chao, “PIPAL: A large-scale image quality assessment dataset for perceptual image restoration,” inEuropean Conference on Computer Vision, pp. 633–651, 2020

  8. [16]

    KonIQ-10k: An ecologically valid database for deep learning of blind image quality assessment,

    V . Hosu, H. Lin, T. Sziranyi, and D. Saupe, “KonIQ-10k: An ecologically valid database for deep learning of blind image quality assessment,”IEEE Transactions on Image Processing, vol. 29, pp. 4041–4056, 2020

  9. [17]

    Rich human feedback for text-to-image generation,

    Y . Liang, J. He, G. Li, P. Li, A. Klimovskiy, N. Carolan, J. Sun, J. Pont-Tuset, S. Young, F. Yang, et al., “Rich human feedback for text-to-image generation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 19401–19411, 2024

  10. [18]

    The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale,

    A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov,et al., “The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale,”International journal of co...

  11. [19]

    Single image super-resolution from transformed self- exemplars,

    J.-B. Huang, A. Singh, and N. Ahuja, “Single image super-resolution from transformed self- exemplars,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015

  12. [20]

    SwinIR: Image restoration using swin transformer,

    J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte, “SwinIR: Image restoration using swin transformer,” inProceedings of the IEEE/CVF international conference on computer vision, pp. 1833–1844, 2021. 11

  13. [21]

    Swift parameter-free attention network for efficient super-resolution,

    C. Wan, H. Yu, Z. Li, Y . Chen, Y . Zou, Y . Liu, X. Yin, and K. Zuo, “Swift parameter-free attention network for efficient super-resolution,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6246–6256, 2024

  14. [22]

    Residual local feature network for efficient super-resolution,

    F. Kong, M. Li, S. Liu, D. Liu, J. He, Y . Bai, F. Chen, and L. Fu, “Residual local feature network for efficient super-resolution,” in2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 765–775, 2022

  15. [23]

    JPEG AI image compression visual artifacts: Detection methods and dataset,

    D. Tsereh, M. Mirgaleev, I. Molodetskikh, R. Kazantsev, and D. S. Vatolin, “JPEG AI image compression visual artifacts: Detection methods and dataset,”ArXiv, vol. abs/2411.06810, 2024

  16. [24]

    Towards real-world blind face restoration with generative facial prior,

    X. Wang, Y . Li, H. Zhang, and Y . Shan, “Towards real-world blind face restoration with generative facial prior,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9168–9178, 2021

  17. [25]

    One-step effective diffusion network for real-world image super-resolution,

    R. Wu, L. Sun, Z. Ma, and L. Zhang, “One-step effective diffusion network for real-world image super-resolution,”Advances in Neural Information Processing Systems, vol. 37, pp. 92529– 92553, 2024

  18. [26]

    Real-world super-resolution via kernel estimation and noise injection,

    X. Ji, Y . Cao, Y . Tai, C. Wang, J. Li, and F. Huang, “Real-world super-resolution via kernel estimation and noise injection,” inThe IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2020

  19. [27]

    Pixel-aware stable diffusion for realistic image super-resolution and personalized stylization,

    T. Yang, R. Wu, P. Ren, X. Xie, and L. Zhang, “Pixel-aware stable diffusion for realistic image super-resolution and personalized stylization,” inEuropean conference on computer vision, pp. 74–91, Springer, 2024

  20. [28]

    SeeSR: Towards semantics-aware real-world image super-resolution,

    R. Wu, T. Yang, L. Sun, Z. Zhang, S. Li, and L. Zhang, “SeeSR: Towards semantics-aware real-world image super-resolution,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 25456–25467, 2024

  21. [29]

    SinSR: diffusion-based image super-resolution in a single step,

    Y . Wang, W. Yang, X. Chen, Y . Wang, L. Guo, L.-P. Chau, Z. Liu, Y . Qiao, A. C. Kot, and B. Wen, “SinSR: diffusion-based image super-resolution in a single step,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 25796–25805, 2024

  22. [30]

    ResShift: Efficient diffusion model for image super-resolution by residual shifting,

    Z. Yue, J. Wang, and C. C. Loy, “ResShift: Efficient diffusion model for image super-resolution by residual shifting,”Advances in Neural Information Processing Systems, vol. 36, pp. 13294– 13307, 2023

  23. [31]

    HAT: Hybrid attention transformer for image restoration,

    X. Chen, X. Wang, W. Zhang, X. Kong, Y . Qiao, J. Zhou, and C. Dong, “HAT: Hybrid attention transformer for image restoration,”arXiv preprint arXiv:2309.05239, 2023

  24. [32]

    DRCT: Saving image super-resolution away from information bottleneck,

    C.-C. Hsu, C.-M. Lee, and Y .-S. Chou, “DRCT: Saving image super-resolution away from information bottleneck,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pp. 6133–6142, June 2024

  25. [33]

    BSDS200 dataset

    D. Martin, C. Fowlkes, and J. Malik, “BSDS200 dataset.” https://huggingface.co/ datasets/goodfellowliu/BSDS200, 2001. License: Apache 2.0

  26. [34]

    Historical dataset

    X. Wang, K. Yu,et al., “Historical dataset.” https://github.com/xinntao/BasicSR, 2018. Part of BasicSR library (license not explicitly specified)

  27. [35]

    General-100 dataset

    C. Dong and C. C. Loy, “General-100 dataset.”http://mmlab.ie.cuhk.edu.hk/projects/ FSRCNN.html, 2016. License: OpenRail

  28. [36]

    Set5 and set14 datasets

    M. Bevilacqua, A. Roumy, and C. Guillemot, “Set5 and set14 datasets.”https://figshare. com/articles/dataset/BSD100_Set5_Set14_Urban100/21586188, 2012. License: CC0 1.0 Universal

  29. [37]

    T91 super-resolution dataset

    J. Yang, J. Wright, and T. Huang, “T91 super-resolution dataset.” https://github.com/ open-mmlab/mmsr, 2010. License: DbCL 1.0. 12 1 5 101520253035404550556065707580859095100 Assessors, # 0% 0% 20% 20% 40% 40% 60% 60% 80% 80% 100% 100% Artifact Prominence95% confidence interva...

Pith tools

Reviewed June 30, 2026 · model on record in the stance chip above.