REVIEW 6 major objections 8 minor 62 references
RAP-SR: RestorAtion Prior Enhancement in Diffusion Models for Realistic Image Super-Resolution
T0 review · 6 major / 8 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read By fine-tuning a pretrained diffusion model on a curated 5,000-image aesthetic dataset with learned positive and negative restoration identifiers, RAP-SR produces a restoration prior that plugs into existing diffusion super-resolution…
desk verdict RAP-SR is a sensible plug-and-play recipe that likely improves diffusion SR realism, but the evidence is weakened by metric curation overlap and the absence of a human study. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a pair of learned restoration identifiers, [X] (positive) and [V] (negative), combined with classifier-free guidance. These tokens are optimized during the Restoration-Oriented Prompt Optimization (ROPO) phase, which appends the positive token to high-quality images and the negative token to RealESRGAN-degraded images together with the semantic caption, then fine-tunes the diffusion model (Restoration Priors Refinement, RPR) on the curated HFAID dataset. At inference the negative token's prediction is subtracted from the positive token's prediction (Equation 3), so the model learns a direction in output space that leads away from degradation; the dataset's role is to make this direction correspond to genuinely high-fidelity, aesthetic detail.
What would settle it
Train RAP-SR exactly as described, then test the plugged-in models on real low-resolution photos whose degradations are known to be unlike RealESRGAN's synthetic pipeline, such as long-kernel motion blur or heavy sensor noise; if the no-reference metric gains over the unmodified base methods disappear or reverse, the learned negative token does not transfer to unseen degradations.
Extended reading notes
Core claim
The central discovery is that the quality-tuning stage of pretrained text-to-image diffusion models, normally used to improve aesthetics, can be repurposed to build a restoration prior: fine-tuning on a tiny dataset of 5,000 ultra-high-quality images, with captions augmented by learned restoration identifiers, teaches the model to associate those identifiers with image quality. During training the model sees both high-quality images paired with a positive identifier and, with probability 0.2, degraded versions of the same images paired with a negative identifier, where the degradations come from the RealESRGAN pipeline. At inference, classifier-free guidance with the positive and negative identifiers produces a restoration direction that pushes the generated image away from degraded features. Swapping the fine-tuned backbone into existing SR methods is enough to improve their no-reference perceptual metrics across synthetic and real-world datasets.
Load-bearing premise
The negative restoration identifier is trained only on degradations synthesized by the RealESRGAN pipeline, so if real-world image degradation lies outside that synthetic distribution, the classifier-free guidance may not steer generation away from the actual defects.
Editorial extensions
If this is right
- StableSR, DiffBIR, and SeeSR all improve on no-reference quality metrics when their Stable Diffusion backbone is replaced with the RAP-SR fine-tuned model, with no additional fine-tuning of the SR methods.
- Training converges in about 3,000 steps with batch size 40, so the restoration-prior enhancement is cheap relative to training a full SR diffusion model.
- The gains come with a tradeoff: full-reference metrics such as PSNR, SSIM, and LPIPS may worsen on some datasets, so RAP-SR is tuned for perceptual realism rather than pixel fidelity.
- Ablations show that both the positive and negative identifiers matter, with the negative prompt contributing more to the no-reference score gains.
Reading between the lines
- A testable extension is to train the negative identifier on multiple degradation pipelines rather than only RealESRGAN; if the plug-in gains then transfer to benchmarks with unknown degradations, the restoration direction is not tied to one synthetic degradation family.
- Because the paper's evidence relies on no-reference metrics, a direct human-perception study on paired real-world images would clarify whether the improved scores correspond to subjectively better restoration, since the appendix shows these metrics can disagree with full-reference scores.
- Since RAP-SR only replaces the T2I backbone, it may combine with future SR adapters unchanged, making it a reusable component rather than a one-off method.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes RAP-SR, a plug-and-play enhancement for diffusion-based real-world image super-resolution. The authors curate a 5,000-image High-Fidelity Aesthetic Image Dataset (HFAID) using a four-stage Quality-Driven Aesthetic Image Selection Pipeline (QDAISP) that filters on no-reference quality metrics and a multi-modal aesthetic model. They then fine-tune Stable Diffusion 2.1 on HFAID with a restoration-oriented prompt optimization scheme: positive and negative restoration identifier tokens are appended to semantic captions, with negative samples generated by the RealESRGAN degradation pipeline (Algorithm 1). At inference, classifier-free guidance combines positive and negative identifier prompts (Eq. 3). The fine-tuned model is used to replace the base diffusion model in StableSR, DiffBIR, and SeeSR, and is evaluated on DIV2K, RealSR, and DrealSR using both full-reference and no-reference metrics. The central claim is that this prior enhancement improves the realism of diffusion-based super-resolution, with consistent no-reference metric gains and state-of-the-art results.
Significance. If validated, RAP-SR would be a useful orthogonal contribution: it is a lightweight, method-agnostic way to inject restoration-aware priors into existing diffusion-based SR pipelines, and HFAID could be reused by the community. The paper's strengths include the simple plug-and-play design, the consistent CLIPIQA and BRISQUE improvements across three base methods in most rows of Table 1, the ablations on dataset size and prompt design, and the commitment to releasing code and data. However, the evidence supporting the central realism claim is currently incomplete: the no-reference metrics used for evaluation overlap with the metrics used to curate the training set, no human perceptual study is provided, LPIPS degrades in most rows, and there is no comparison with current state-of-the-art Real-SR methods. These issues are addressable in revision, but they are load-bearing for the paper's conclusions.
major comments (6)
- [Section 3.1, Appendix B, Table 1] The evaluation is partly circular. HFAID is curated by selecting images with high CLIPIQA, MANIQA, and NIQE scores, and the main results in Table 1 are reported on the same no-reference metric family, including MANIQA, CLIPIQA, and BRISQUE. Fine-tuning on metric-selected images can inflate those metrics without improving perceived realism, because the model is trained to reproduce images that already score well on those metrics. The authors should add a human perceptual study (e.g., pairwise preference) and/or evaluate on held-out no-reference metrics that were not part of QDAISP, and report whether the gains persist.
- [Table 1] The full-reference fidelity results undermine the realism claim as currently presented. LPIPS worsens in 7 of 9 rows (e.g., SeeSR on DIV2K: 0.3168 to 0.3501; SeeSR on RealSR: 0.3096 to 0.3388; SeeSR on DrealSR: 0.3243 to 0.3542; StableSR on DIV2K: 0.3118 to 0.3542), while PSNR/SSIM improve only in some rows. Since the paper claims more realistic output and provides no human evaluation, the no-reference gains alone do not establish realism; they may indicate a shift toward a preferred style rather than improved restoration. A user study or a detailed fidelity-realism trade-off analysis is needed.
- [Section 4.1, Table 1] Table 1 reports single values with no error bars, no number of runs, and no significance tests. Several improvements are very small relative to expected run-to-run variation (e.g., SeeSR MUSIQ on DIV2K: 68.66 to 68.86), and some no-reference metrics actually decrease (e.g., StableSR MANIQA on DIV2K: 0.6193 to 0.6017). The statement that the method achieves 'great improvements in almost all no-reference metrics' is not statistically supported without repeated runs and variability measures.
- [Section 4.2, Table 2] The dataset-size ablation does not support the claims made in the text. The text says that at 8,000 images 'both reference and no-reference metrics show a decline,' but Table 2 reports only MANIQA and CLIPIQA; no full-reference metrics are shown. In addition, the choice of the default 5,000-image size, the positive ratio r=0.8, and the guidance scale lambda_s is not accompanied by a described selection protocol, so it is unclear whether these values were tuned on the DrealSR test set. The authors should report the full metric table and clarify the hyperparameter selection procedure.
- [Section 3.2, Algorithm 1, Eq. (3)] The negative restoration identifier [V] is trained only on degradations synthesized by the RealESRGAN pipeline, while the method is evaluated on real-world LR images with unknown and mixed degradations. If the synthetic degradation model does not cover the real test distribution, the learned association between [V] and 'low quality' may not transfer, and the classifier-free guidance in Eq. (3) would not steer correctly on real inputs. The authors should analyze the degradation gap, for example by testing on additional synthetic degradation types or including real LR/HR pairs, to demonstrate that the negative identifier generalizes.
- [Abstract, Section 4.1, Conclusion] The paper claims 'state-of-the-art results,' but Section 4.1 compares only each base method with and without RAP-SR. No comparison is made with other published Real-SR methods such as SUPIR, PASD, CoSeR, or recent GAN-based baselines. The authors should either add such comparisons or restrict the claim to 'improves the base methods,' which is all the current evidence supports.
minor comments (8)
- [Section 4] The text contains a typo: 'RealESRAGN' should be 'RealESRGAN.'
- [Section 4.1] The paragraph under Quantitative Comparison refers to 'PAR-SR' instead of 'RAP-SR.'
- [Appendix A.2] The dataset name is spelled 'HAFID' in the qualitative comparison section; it should be 'HFAID' for consistency.
- [Table 3] The row header 'Filck2K' is a typo for 'Flickr2K.'
- [Table 1] The DrealSR MUSIQ value for SeeSR + RAP-SR is reported as '70.0644,' which appears to have an extra digit; please verify the formatting.
- [Section 3.1, Figure 4] The text says 'As illustrated in figure 2, our dataset excels in both image quality and detail richness,' but the dataset-quality comparison is shown in Figure 4, not Figure 2; the cross-reference should be corrected.
- [Eq. (1)-(3)] The notation z_t_lr is used without definition; please define the latent variable and its relationship to the low-resolution image x_lr.
- [Appendix B and D] There is an inconsistency: Appendix B reports that MUSIQ and BRISQUE correlate less with human aesthetic preferences, while Appendix D states that no-reference metrics 'such as MANIQA, MUSIQ, and CLIPIQA' align more closely with human perception. This should be reconciled.
Circularity Check
Metric selection and metric evaluation overlap: the no-reference gains that carry the realism claim are partly built into HFAID curation, making the central evidence partially circular.
-
self definitional
[Section 3.1, QDAISP and 'Comparison With Other Datasets' (Figure 4); Appendix A.1]
"We first select 200 images with the best and worst performance under each metric from the LSDIR dataset and conduct a 10-person user evaluation, ultimately selecting MANIQA (Yang et al. 2022), CLIP-IQA (Wang, Chan, and Loy 2023), and NIQE (Zhang, Zhang, and Bovik 2015) as the core evaluation metrics. ... We evaluate the quality of our dataset using five no-reference image quality assessment metrics: MANIQA, MUSIQ, CLIP-IQA, BRISQUE, and NIQE. Figure 4 presents the results, showing that our dataset significantly outperforms others across all metrics."
HFAID is defined by the QDAISP selection rule, whose core metrics are MANIQA, CLIP-IQA, and NIQE. Claiming that HFAID exceeds other datasets in fidelity and aesthetic quality using those same metrics (Figure 4 and Table 3) is therefore true by construction of the selection rule and is not independent evidence of higher quality. The extra metrics MUSIQ and BRISQUE were also examined in the same user study and explicitly rejected as less aligned with human preferences, so they do not provide a fully external check.
-
fitted input called prediction
[Section 4, Evaluation Metrics and Section 4.1, Quantitative Comparison (Table 1); Appendix B]
"Evaluation Metrics. To better align with human perception, we use seven evaluation metrics: ... MANIQA (Yang et al. 2022), MUSIQ (Ke et al. 2021), BRISQUE (Shao and Mou 2021) and CLIPIQA (Hessel et al. 2021). ... Firstly, the method we proposed has achieved great improvements in almost all no-reference metrics such as MANIQA, MUSIQ, CLIPIQA, and BRISQUE on all three data sets. This shows that our method significantly improves the image generation capabilities of the original method and can generate richer details."
The model whose restoration prior is enhanced is fine-tuned on HFAID, which was curated using MANIQA, CLIP-IQA, and NIQE as core selection metrics. The central quantitative evidence for improved realism is gains on the same family of no-reference metrics (Table 1), so those gains are partly a consequence of shifting the training distribution toward metric-selected images rather than an independent validation of restoration fidelity. The paper itself acknowledges the overlap in Appendix B by adding MUSIQ and BRISQUE 'to avoid data bias', yet the headline conclusion still relies on the overlapping metrics.
full rationale
The paper's derivation chain is not formally circular in the sense of an equation reducing to its own inputs, and there is no load-bearing self-citation: the cited base methods (StableSR, DiffBIR, SeeSR) and tools (DreamBooth, Florence-2, mPLUG-Owl2) are external prior work. The circularity is in the evaluation instrument. HFAID is curated by ranking images with MANIQA, CLIP-IQA, and NIQE, and the paper's main evidence that RAP-SR improves realism is gains on the same no-reference metric family. Because the fine-tuned model is trained to reproduce HFAID's distribution, improvements on those metrics are at least partially built into the training-data choice. The added MUSIQ and BRISQUE metrics mitigate but do not eliminate this, since they were also inspected and rejected in the same user study used to pick the core metrics. The small 10-person user study anchors the metrics to human preferences on 200 LSDIR images, and the selected HFAID images are human-verified, but neither of these checks evaluates the restored outputs in Table 1; no human perceptual study of the final SR results is reported. The consistent LPIPS degradation in Table 1 further indicates that the metric gains come with a loss of perceptual fidelity, so interpreting them as enhanced realism requires an external anchor that the paper does not provide. Separately, the 'state-of-the-art results' claim is under-supported because Table 1 only compares each base method against itself plus RAP-SR and omits current SOTA baselines; that is a completeness concern rather than a circularity one. The synthetic-degradation gap between the RealESRGAN-trained negative identifier and real test distributions is a generalizability concern, not a circularity one. Overall, the central claim has independent components, including the plug-and-play integration and the identifier-based prompt optimization, but its headline evidence is partially circular due to the metric overlap between curation and evaluation.
Assumptions & free parameters
free parameters (5)
- HFAID dataset size =
5,000 images
- Positive sample ratio r =
0.8
- CFG guidance scale lambda_s =
not reported
- Training hyperparameters (learning rate, batch size, iterations) =
5e-5, 40, 3000
- QDAISP quality metric thresholds =
not reported
assumptions (5)
- domain assumption Fine-tuning Stable Diffusion on 5,000 curated images enhances its restoration prior without harming its generative capacity
- domain assumption RealESRGAN degradation covers the real-world degradation distribution seen in evaluation
- domain assumption CLIPIQA, MANIQA, and NIQE are valid proxies for human aesthetic preference
- domain assumption Florence-2 captions are accurate and descriptive enough for fine-tuning
- standard math Standard diffusion training objective and classifier-free guidance are valid
invented entities (1)
-
Restoration identifier tokens [X] (positive) and [V] (negative)
Cite this review
Pith. "Pith review of RAP-SR: RestorAtion Prior Enhancement in Diffusion Models for Realistic Image Super-Resolution." pith.science (2026). https://pith.science/paper/C3NEOI5O
@misc{pith2026241207149,
author = {Pith},
title = {Pith review of: RAP-SR: RestorAtion Prior Enhancement in Diffusion Models for Realistic Image Super-Resolution},
year = {2026},
howpublished = {\url{https://pith.science/paper/C3NEOI5O}},
note = {Machine review of arXiv:2412.07149}
}
read the original abstract
Benefiting from their powerful generative capabilities, pretrained diffusion models have garnered significant attention for real-world image super-resolution (Real-SR). Existing diffusion-based SR approaches typically utilize semantic information from degraded images and restoration prompts to activate prior for producing realistic high-resolution images. However, general-purpose pretrained diffusion models, not designed for restoration tasks, often have suboptimal prior, and manually defined prompts may fail to fully exploit the generated potential. To address these limitations, we introduce RAP-SR, a novel restoration prior enhancement approach in pretrained diffusion models for Real-SR. First, we develop the High-Fidelity Aesthetic Image Dataset (HFAID), curated through a Quality-Driven Aesthetic Image Selection Pipeline (QDAISP). Our dataset not only surpasses existing ones in fidelity but also excels in aesthetic quality. Second, we propose the Restoration Priors Enhancement Framework, which includes Restoration Priors Refinement (RPR) and Restoration-Oriented Prompt Optimization (ROPO) modules. RPR refines the restoration prior using the HFAID, while ROPO optimizes the unique restoration identifier, improving the quality of the resulting images. RAP-SR effectively bridges the gap between general-purpose models and the demands of Real-SR by enhancing restoration prior. Leveraging the plug-and-play nature of RAP-SR, our approach can be seamlessly integrated into existing diffusion-based SR methods, boosting their performance. Extensive experiments demonstrate its broad applicability and state-of-the-art results. Codes and datasets will be available upon acceptance.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Agustsson, E.; and Timofte, R. 2017. NTIRE 2017 Challenge on Single Image Super-Resolution: Dataset and Study. In IEEE Conference on Computer Vision and Pattern Recognition Workshops, CVPRW 2017 , 1122--1131
work page 2017
-
[2]
Avrahami, O.; Lischinski, D.; and Fried, O. 2022. Blended Diffusion for Text-driven Editing of Natural Images. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 18187--18197
work page 2022
-
[3]
Brooks, T.; Holynski, A.; and Efros, A. A. 2023. InstructPix2Pix: Learning to Follow Image Editing Instructions. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023 , 18392--18402
work page 2023
-
[4]
Cai, J.; Zeng, H.; Yong, H.; Cao, Z.; and Zhang, L. 2019. Toward Real-World Single Image Super-Resolution: A New Benchmark and a New Model. In IEEE/CVF International Conference on Computer Vision, ICCV 2019 , 3086--3095
work page 2019
-
[5]
Chen, C.; Xiong, Z.; Tian, X.; Zha, Z.; and Wu, F. 2019. Camera Lens Super-Resolution. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019 , 1652--1660
work page 2019
-
[6]
T.; Luo, P.; Lu, H.; and Li, Z
Chen, J.; Yu, J.; Ge, C.; Yao, L.; Xie, E.; Wang, Z.; Kwok, J. T.; Luo, P.; Lu, H.; and Li, Z. 2024. PixArt- \( \) : Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis. In The Twelfth International Conference on Learning Representations, ICLR 2024
work page 2024
-
[7]
Chung, H.; Sim, B.; and Ye, J. C. 2022. Come-Closer-Diffuse-Faster: Accelerating Conditional Diffusion Models for Inverse Problems through Stochastic Contraction. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 12403--12412
work page 2022
-
[8]
Dai, X.; Hou, J.; Ma, C.-Y.; Tsai, S.; Wang, J.; Wang, R.; Zhang, P.; Vandenhende, S.; Wang, X.; Dubey, A.; et al. 2023. Emu: Enhancing image generation models using photogenic needles in a haystack. arXiv preprint arXiv:2309.15807
arXiv 2023
Show all 62 references
-
[9]
Deng, J.; Dong, W.; Socher, R.; Li, L.; Li, K.; and Li, F. 2009. ImageNet: A large-scale hierarchical image database. In IEEE Computer Society Conference on Computer Vision and Pattern Recognition, CVPR 2009 , 248--255
2009
-
[10]
C.; He, K.; and Tang, X
Dong, C.; Loy, C. C.; He, K.; and Tang, X. 2014. Learning a Deep Convolutional Network for Image Super-Resolution. In Fleet, D. J.; Pajdla, T.; Schiele, B.; and Tuytelaars, T., eds., Proceedings of European Conference on Computer Vision, Part IV , volume 8692, 184--199
2014
-
[11]
Fei, B.; Lyu, Z.; Pan, L.; Zhang, J.; Yang, W.; Luo, T.; Zhang, B.; and Dai, B. 2023. Generative Diffusion Prior for Unified Image Restoration and Enhancement. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023 , 9935--9946
2023
-
[12]
J.; Pouget - Abadie, J.; Mirza, M.; Xu, B.; Warde - Farley, D.; Ozair, S.; Courville, A
Goodfellow, I. J.; Pouget - Abadie, J.; Mirza, M.; Xu, B.; Warde - Farley, D.; Ozair, S.; Courville, A. C.; and Bengio, Y. 2014. Generative Adversarial Nets. In Ghahramani, Z.; Welling, M.; Cortes, C.; Lawrence, N. D.; and Weinberger, K. Q., eds., Advances in Neural Informatio...
2014
-
[13]
Haris, M.; Shakhnarovich, G.; and Ukita, N. 2018. Deep Back-Projection Networks for Super-Resolution. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018 , 1664--1673
2018
-
[14]
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016 , 770--778
2016
-
[15]
Hessel, J.; Holtzman, A.; Forbes, M.; Le Bras, R.; and Choi, Y. 2021. CLIPS core: A Reference-free Evaluation Metric for Image Captioning. In Moens, M.-F.; Huang, X.; Specia, L.; and Yih, S. W.-t., eds., Proceedings of the 2021 Conference on Empirical Methods in Natural Langua...
2021
-
[16]
Ho, J.; Jain, A.; and Abbeel, P. 2020. Denoising diffusion probabilistic models. In Advances in neural information processing systems, NeurIPS 2020, volume 33, 6840--6851
2020
-
[17]
Ho, J.; and Salimans, T. 2022. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598
2022 arXiv
-
[18]
Kawar, B.; Elad, M.; Ermon, S.; and Song, J. 2022. Denoising Diffusion Restoration Models. In Koyejo, S.; Mohamed, S.; Agarwal, A.; Belgrave, D.; Cho, K.; and Oh, A., eds., Advances in Neural Information Processing Systems, NeurIPS 2022
2022
-
[19]
Ke, J.; Wang, Q.; Wang, Y.; Milanfar, P.; and Yang, F. 2021. MUSIQ: Multi-scale Image Quality Transformer. In IEEE/CVF International Conference on Computer Vision, ICCV 2021 , 5128--5137
2021
-
[20]
K.; and Lee, K
Kim, J.; Lee, J. K.; and Lee, K. M. 2016. Deeply-Recursive Convolutional Network for Image Super-Resolution. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016 , 1637--1645
2016
-
[21]
C.; Lo, W.; Doll \' a r, P.; and Girshick, R
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.; Doll \' a r, P.; and Girshick, R. B. 2023. Segment Anything. In IEEE/CVF International Conference on Computer Vision, ICCV 2023 , 3992--4003
2023
-
[22]
Li, C.; Wong, C.; Zhang, S.; Usuyama, N.; Liu, H.; Yang, J.; Naumann, T.; Poon, H.; and Gao, J. 2023 a . LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day. In Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; and Levine, S., eds., Ad...
2023
-
[23]
Li, W.; Zhou, K.; Qi, L.; Lu, L.; and Lu, J. 2022. Best-Buddy GANs for Highly Detailed Image Super-resolution. In AAAI Conference on Artificial Intelligence, AAAI 2022, , 1412--1420
2022
-
[24]
Li, Y.; Hunt, S.; Park, J.; O'Toole, M.; and Kitani, K. 2023 b . Azimuth Super-Resolution for FMCW Radar in Autonomous Driving. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023 , 17504--17513
2023
-
[25]
Li, Y.; Zhang, K.; Liang, J.; Cao, J.; Liu, C.; Gong, R.; Zhang, Y.; Tang, H.; Liu, Y.; Demandolx, D.; Ranjan, R.; Timofte, R.; and Gool, L. V. 2023 c . LSDIR: A Large Scale Dataset for Image Restoration. In IEEE/CVF Conference on Computer Vision and Pattern Recognition Worksh...
2023
-
[26]
V.; and Timofte, R
Liang, J.; Cao, J.; Sun, G.; Zhang, K.; Gool, L. V.; and Timofte, R. 2021. SwinIR: Image Restoration Using Swin Transformer. In IEEE/CVF International Conference on Computer Vision Workshops, ICCVW 2021 , 1833--1844
2021
-
[27]
Liang, J.; Zeng, H.; and Zhang, L. 2022. Details or Artifacts: A Locally Discriminative Learning Approach to Realistic Image Super-Resolution. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 5647--5656
2022
-
[28]
J.; Hays, J.; Perona, P.; Ramanan, D.; Doll \' a r, P.; and Zitnick, C
Lin, T.; Maire, M.; Belongie, S. J.; Hays, J.; Perona, P.; Ramanan, D.; Doll \' a r, P.; and Zitnick, C. L. 2014. Microsoft COCO: Common Objects in Context. In Fleet, D. J.; Pajdla, T.; Schiele, B.; and Tuytelaars, T., eds., Proceedings of European Conference on Computer Visio...
2014
-
[29]
Lin, X.; He, J.; Chen, Z.; Lyu, Z.; Dai, B.; Yu, F.; Ouyang, W.; Qiao, Y.; and Dong, C. 2024. DiffBIR: Towards Blind Image Restoration with Generative Diffusion Prior. arXiv:2308.15070
2024 arXiv
-
[30]
Loshchilov, I.; and Hutter, F. 2019. Decoupled Weight Decay Regularization. In 7th International Conference on Learning Representations, ICLR 2019
2019
-
[31]
Lugmayr, A.; Danelljan, M.; Romero, A.; Yu, F.; Timofte, R.; and Gool, L. V. 2022. RePaint: Inpainting using Denoising Diffusion Probabilistic Models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 11451--11461
2022
-
[32]
Meng, C.; He, Y.; Song, Y.; Song, J.; Wu, J.; Zhu, J.; and Ermon, S. 2022. SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations. In The Tenth International Conference on Learning Representations, ICLR 2022
2022
-
[33]
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022. High-Resolution Image Synthesis with Latent Diffusion Models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022 , 10674--10685
2022
-
[34]
Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; and Aberman, K. 2023. DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023 , 22500--22510
2023
-
[35]
A.; Ho, J.; Salimans, T.; Fleet, D
Saharia, C.; Chan, W.; Chang, H.; Lee, C. A.; Ho, J.; Salimans, T.; Fleet, D. J.; and Norouzi, M. 2022. Palette: Image-to-Image Diffusion Models. In Nandigjav, M.; Mitra, N. J.; and Hertzmann, A., eds., SIGGRAPH '22: Special Interest Group on Computer Graphics and Interactive ...
2022
-
[36]
Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; Schramowski, P.; Kundurthy, S.; Crowson, K.; Schmidt, L.; Kaczmarczyk, R.; and Jitsev, J. 2022. LAION-5B: An open large-scale dataset for training ne...
2022
-
[37]
Shang, S.; Shan, Z.; Liu, G.; Wang, L.; Wang, X.; Zhang, Z.; and Zhang, J. 2024. ResDiff: Combining CNN and Diffusion Model for Image Super-resolution. In Wooldridge, M. J.; Dy, J. G.; and Natarajan, S., eds., Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024...
2024
-
[38]
Shao, W.; and Mou, X. 2021. No-Reference Image Quality Assessment Based on Edge Pattern Feature in the Spatial Domain. IEEE Access , 9: 133170--133184
2021
-
[39]
Song, J.; Meng, C.; and Ermon, S. 2021. Denoising Diffusion Implicit Models. In 9th International Conference on Learning Representations, ICLR 2021
2021
-
[40]
P.; Kumar, A.; Ermon, S.; and Poole, B
Song, Y.; Sohl - Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2021. Score-Based Generative Modeling through Stochastic Differential Equations. In 9th International Conference on Learning Representations, ICLR 2021
2021
-
[41]
Sun, H.; Li, W.; Liu, J.; Chen, H.; Pei, R.; Zou, X.; Yan, Y.; and Yang, Y. 2024. CoSeR: Bridging Image and Language for Cognitive Super-Resolution. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 , 25868--25878
2024
-
[42]
V.; Yang, M.; Zhang, L.; Lim, B.; Son, S.; Kim, H.; Nah, S.; Lee, K
Timofte, R.; Agustsson, E.; Gool, L. V.; Yang, M.; Zhang, L.; Lim, B.; Son, S.; Kim, H.; Nah, S.; Lee, K. M.; Wang, X.; Tian, Y.; Yu, K.; Zhang, Y.; Wu, S.; Dong, C.; Lin, L.; Qiao, Y.; Loy, C. C.; Bae, W.; Yoo, J.; Han, Y.; Ye, J. C.; Choi, J.; Kim, M.; Fan, Y.; Yu, J.; Han, ...
2017
-
[43]
Wang, J.; Chan, K. C. K.; and Loy, C. C. 2023. Exploring CLIP for Assessing the Look and Feel of Images. In Williams, B.; Chen, Y.; and Neville, J., eds., Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023 , 2555--2563
2023
-
[44]
Wang, J.; Yue, Z.; Zhou, S.; Chan, K. C. K.; and Loy, C. C. 2024. Exploiting Diffusion Prior for Real-World Image Super-Resolution. Int. J. Comput. Vis., 132(12): 5929--5949
2024
-
[45]
Wang, R.; Zhang, D.; Li, Q.; Zhou, X.; and Lo, B. 2021 a . Real-time Surgical Environment Enhancement for Robot-Assisted Minimally Invasive Surgery Based on Super-Resolution. In IEEE International Conference on Robotics and Automation, ICRA 2021 , 3434--3440
2021
-
[46]
Wang, X.; Xie, L.; Dong, C.; and Shan, Y. 2021 b . Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data. In IEEE/CVF International Conference on Computer Vision Workshops, ICCVW 2021 , 1905--1914
2021
-
[47]
Wang, Y.; Yu, J.; and Zhang, J. 2023. Zero-Shot Image Restoration Using Denoising Diffusion Null-Space Model. In The Eleventh International Conference on Learning Representations, ICLR 2023
2023
-
[48]
C.; Sheikh, H
Wang, Z.; Bovik, A. C.; Sheikh, H. R.; and Simoncelli, E. P. 2004. Image quality assessment: from error visibility to structural similarity. IEEE Trans. Image Process. , 13(4): 600--612
2004
-
[49]
Wei, P.; Xie, Z.; Lu, H.; Zhan, Z.; Ye, Q.; Zuo, W.; and Lin, L. 2020. Component Divide-and-Conquer for Real-World Image Super-Resolution. In Vedaldi, A.; Bischof, H.; Brox, T.; and Frahm, J., eds., Proceedings of European Conference on Computer Vision, Part VIII , volume 1235...
2020
-
[50]
Wu, R.; Yang, T.; Sun, L.; Zhang, Z.; Li, S.; and Zhang, L. 2024. SeeSR: Towards Semantics-Aware Real-World Image Super-Resolution. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 , 25456--25467
2024
-
[51]
Xiao, B.; Wu, H.; Xu, W.; Dai, X.; Hu, H.; Lu, Y.; Zeng, M.; Liu, C.; and Yuan, L. 2024. Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 , 4818--4829
2024
-
[52]
Yang, S.; Wu, T.; Shi, S.; Lao, S.; Gong, Y.; Cao, M.; Wang, J.; and Yang, Y. 2022. MANIQA: Multi-dimension Attention Network for No-Reference Image Quality Assessment. In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, CVPRW 2022 , 1190--1199
2022
-
[53]
Yang, T.; Wu, R.; Ren, P.; Xie, X.; and Zhang, L. 2024. Pixel-Aware Stable Diffusion for Realistic Image Super-Resolution and Personalized Stylization. In Leonardis, A.; Ricci, E.; Roth, S.; Russakovsky, O.; Sattler, T.; and Varol, G., eds., Proceedings of European Conference ...
2024
-
[54]
Ye, Q.; Xu, H.; Ye, J.; Yan, M.; Hu, A.; Liu, H.; Qian, Q.; Zhang, J.; and Huang, F. 2024. mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 , 13040--13051
2024
-
[55]
Yu, F.; Gu, J.; Li, Z.; Hu, J.; Kong, X.; Wang, X.; He, J.; Qiao, Y.; and Dong, C. 2024. Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 , 25669--25680
2024
-
[56]
Yue, Z.; Wang, J.; and Loy, C. C. 2023. ResShift: Efficient Diffusion Model for Image Super-resolution by Residual Shifting. In Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; and Levine, S., eds., Advances in Neural Information Processing Systems, NeurIPS 2023
2023
-
[57]
V.; and Timofte, R
Zhang, K.; Liang, J.; Gool, L. V.; and Timofte, R. 2021. Designing a Practical Degradation Model for Deep Blind Image Super-Resolution. In IEEE/CVF International Conference on Computer Vision, ICCV 2021 , 4771--4780
2021
-
[58]
Zhang, L.; Rao, A.; and Agrawala, M. 2023. Adding Conditional Control to Text-to-Image Diffusion Models. In IEEE/CVF International Conference on Computer Vision, ICCV 2023 , 3813--3824
2023
-
[59]
Zhang, L.; Zhang, L.; and Bovik, A. C. 2015. A Feature-Enriched Completely Blind Image Quality Evaluator. IEEE Trans. Image Process. , 24(8): 2579--2591
2015
-
[60]
A.; Shechtman, E.; and Wang, O
Zhang, R.; Isola, P.; Efros, A. A.; Shechtman, E.; and Wang, O. 2018. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018 , 586--595
2018
-
[61]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...
-
[62]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.