REVIEW 2 major objections 6 minor 1 cited by
TASR: Timestep-Aware Diffusion Model for Image Super-Resolution
T0 review · 2 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A timestep-aware adapter that fuses low-resolution structure early and diffusion detail late improves super-resolution quality, with gains over state-of-the-art methods on no-reference perceptual metrics.
desk verdict Timestep-aware adapter for ControlNet SR is a clean, well-ablated engineering contribution; the CLIPIQA-as-reward-and-metric circularity is real but the paper handles it better than most. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the timestep-aware adapter: two convolutional layers with ReLU and normalization, with the timestep injected through an adaptive layer-norm (AdaLN) layer, ending in a sigmoid that predicts a per-pixel control weight $\alpha$ in $[0,1]$. It is inserted at each decoder block between ControlNet's skip features and the pre-trained Stable Diffusion features, computing $f_d + f_{\mathrm{cond}} \cdot \alpha$. This design does two jobs: it implements the temporal policy (structure early, detail late) and it constrains the optimization space so that the late-stage CLIPIQA reward cannot push the model into reward-hacking configurations. The complementary machinery is the timestep-aware training strategy, which applies the L1 loss to ControlNet when $t \le 800$ and the CLIPIQA reward only to the adapter when $t \le 200$, alternating optimization of the two modules.
What would settle it
Train the adapter alone with only the CLIPIQA reward on RealSR and compare its outputs to the paper's full model: if the adapter-only outputs develop the same style artifacts as ControlNet-only training, then the sigmoid-bounded space does not exclude reward hacking and the reported perceptual gains are not a general result.
Extended reading notes
Core claim
The central claim is that ControlNet's low-resolution conditioning is not equally useful across the denoising trajectory: it mainly fixes low-frequency structure early, and in late steps it can suppress high-frequency detail. TASR acts on this by inserting a lightweight timestep-aware adapter into each decoder block of the denoising U-Net; the adapter takes the current timestep, the U-Net feature, and the ControlNet skip feature, and outputs a sigmoid-bounded control weight map $\alpha$, fusing features as $f_d + f_{\mathrm{cond}} \cdot \alpha$. Because $\alpha$ lies in $[0,1]$, the adapter's optimization space is constrained, so optimizing it with the CLIPIQA reward avoids the reward-hacking artifacts that appear when the same reward is applied to ControlNet's larger parameter space. The paper's experiments claim consistent gains over state-of-the-art methods on no-reference metrics across DIV2K-val, RealSR, DRealSR, and RealLR200.
Load-bearing premise
The method depends on the assumption that the reward-hacking configurations that fool CLIPIQA lie outside the adapter's sigmoid-bounded optimization space, a claim supported by a single ablation on DIV2K-val rather than by a formal argument or cross-dataset tests.
Editorial extensions
If this is right
- ControlNet-based super-resolution pipelines can devote late denoising steps to the diffusion prior rather than conditional control, potentially saving computation without losing fidelity.
- Timestep-aware loss scheduling (L1 early, perceptual reward late) is an effective training recipe when applied to disjoint modules.
- Constraining the reward-optimized module, the sigmoid-bounded adapter, avoids the reward-hacking artifacts observed when the same perceptual reward is applied to the larger ControlNet.
- The method improves no-reference quality metrics across DIV2K-val, RealSR, DRealSR, and RealLR200, with the largest gains on MANIQA and CLIPIQA.
Reading between the lines
- A testable extension is applying the same temporal split to other conditional diffusion tasks, such as depth-to-image, inpainting, or image editing, where the condition may also matter mainly in the semantic-planning stage.
- The sigmoid-bounded reward optimization argument suggests a general design rule: when using a learned perceptual reward, confine training to a small-capacity module whose output space provably excludes known reward-hacking directions.
- The temporal-control analysis implies that late-step ControlNet computation could be skipped entirely at inference, which would directly reduce latency for diffusion super-resolution; the paper does not test this speed-latency tradeoff.
- Replacing CLIPIQA with a stronger perceptual reward at late steps could push quality further, as the paper itself notes, but the safe-space argument would need to be re-verified for each new reward.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes TASR, a latent-diffusion image super-resolution method built on Stable Diffusion and ControlNet. The authors first observe with DiffBIR that ControlNet conditioning mainly affects the early denoising steps (high noise levels) and has little effect at the end. From this they design a timestep-aware adapter that predicts a sigmoid-bounded weight map to fuse ControlNet features into each SD decoder block, and a two-stage training scheme: ControlNet is trained first with a denoising loss; then ControlNet is optimized with an L1 fidelity loss for intermediate timesteps and the adapter with a CLIPIQA perceptual reward for late timesteps, alternating between the two modules. On DIV2K-val, RealSR, DRealSR and RealLR200, TASR reports the best no-reference MANIQA, MUSIQ and CLIPIQA scores among the compared GAN and diffusion methods, at lower PSNR/SSIM. The paper includes ablations of adapter architecture, loss schedule, and reward choice.
Significance. If the results hold, the contribution is useful: a lightweight adapter that reweights ControlNet features by timestep is simple, the training recipe is clearly specified, the code is released, and the ablations include important controls such as a reward-hacking condition and alternative reward functions. The temporal analysis of ControlNet is a plausible and testable insight. However, the evaluation is weakened by the use of CLIPIQA as both the training reward and a headline evaluation metric, and by an unsupported claim that the sigmoid-bounded adapter cannot reach reward-hacking configurations. The independent support from unrewarded metrics MANIQA and MUSIQ, from GPT-4o preference, and from Gram-matrix statistics is partial and not fully quantified in the tables.
major comments (2)
- [Sec. 3.3 / Eq. (5) / Table 3] The safeguard against CLIPIQA reward hacking is not established, and the evaluation is partly circular. CLIPIQA appears both in the training reward L_clipiqa and as a headline metric in Table 1. The argument in Sec. 3.3 and Fig. 3 that the sigmoid-bounded adapter cannot reach reward-hacking configurations is not supported by a reachability analysis; the mapping from adapter weights through 20 sequential DDPM steps to final images is highly nonlinear, and no formal or empirical characterization of the reachable set is given. More directly, Table 3, row 2 shows that training with L_clipiqa alone attains CLIPIQA 0.7687 and MUSIQ 69.73 while LPIPS degrades to 0.4110 from 0.3762 in the full model. This is the same fidelity-for-CLIPIQA trade that the paper attributes to reward hacking, and it is achieved within the adapter's optimization space. Consequently, the CLIPIQA gains in Table 1 may partially reflect optimization of the evaluation metric rather than genuine perceptual quality. The MANIQA and MUSIQ gains are less directly circular because those metrics are not used as rewards in the final model, but Table 4 shows that they are also optimizable when used as rewards. The authors should add a human study or a set of fully unrewarded quality metrics to support the claim that the perceptual gains reflect quality rather than metric gaming.
- [Sec. 3.3 / Eq. (6)] The timestep-aware training schedule is stated inconsistently with the paper's temporal premise. In standard DDPM notation, t=1000 is the noisiest latent and t=0 is the clean image, so the early stages of the reverse denoising process correspond to large t. The text says L1 is applied in early denoising stages (0<=t<=800) and CLIPIQA in later stages (0<=t<=200), but Eq. (6) applies L1 only for t in [200,800], applies no L1 for t in (800,1000], and applies CLIPIQA for t in [0,200]. Under the standard convention this is the reverse of the stated design: the very noisy early steps receive no fidelity supervision, and the perceptual reward is applied only at the end. If the authors are using a nonstandard convention for 'early' and 'late', it must be stated explicitly. As written, the schedule does not operationalize the motivating observation from Fig. 1 and makes the central contribution hard to evaluate.
minor comments (6)
- [Sec. 4.4 / Table 3] Table 3 does not include PSNR or SSIM, but the text states that adding L_clipiqa alone decreases PSNR by 4.3% and SSIM by 10.6%, and that adding L1 alone improves reference metrics. Please add the missing PSNR/SSIM columns so these claims can be verified.
- [Fig. 3] The labels and metric scales in Fig. 3 appear inconsistent: the right panel reports CLIPIQA 0.9308 and LPIPS 0.6745 with very low PSNR, which matches the reward-hacking 'ControlNet' row of Table 2 (CLIPIQA 0.9345, LPIPS 0.4632), yet the figure labels it 'Results in Adapter'. In addition, the metric values are on inconsistent scales (e.g., MANIQA 75.15 vs 0.6158, MUSIQ 0.5810 vs 70.55). Please correct the labels and units.
- [Sec. 4.2] The description of the adapter reports only that it takes f_d, f_cond, and timestep t as inputs, but it does not state how these are combined before the convolutions (e.g., concatenation versus summation). Please clarify the exact tensor operations in the adapter.
- [Sec. 4.3] The claim that TASR 'significantly outperforms' other methods is not accompanied by any statistical testing or error bars. Given the variability of no-reference IQA metrics, please report standard deviations over multiple evaluation runs or a paired significance test.
- [Sec. 2.1 / References] Reference [3] has a garbled author name ('ao Yang') and is missing venue details; please correct the entry. Also, the associated 'PASD' method is cited in a way that makes it difficult to identify the published version.
- [Fig. 1] The experimental setup for Fig. 1 should state explicitly how 'conditional inference steps' is implemented (e.g., zeroing the ControlNet adapter output after a given step) and whether the comparison is on a fixed seed or multiple samples; this is important for reproducing the temporal observation.
Circularity Check
CLIPIQA is both training reward and headline evaluation metric, so the reported CLIPIQA gains are partly forced by construction; MANIQA/MUSIQ and the adapter design provide independent support.
-
fitted input called prediction
[Sec. 3.3, Eq. (5); Sec. 4.3, Tab. 1]
"In the later stages of denoising (i.e., 0 ≤ t ≤ 200), we use the non-reference metric CLIP-IQA [38] to evaluate the visual quality of the generated HR images and use its results as a perception reward... L_clipiqa = 1 − R(Ĩ). ... our method significantly outperforms other state-of-the-art (SOTA) methods on non-reference metrics such as MANIQA, MUSIQ, and CLIPIQA across all datasets."
The CLIPIQA score reported as evidence of quality in Tab. 1 is the same function R(Ĩ) that Eq. (5) optimizes as a training reward: the model is explicitly trained to increase CLIPIQA, so higher CLIPIQA on test sets is, at least in part, a direct consequence of the training objective rather than an independent confirmation of perceptual quality. The paper's own loss ablation (Tab. 3, row 2) confirms this: training with CLIPIQA alone yields CLIPIQA 0.7687 (above the final 0.7681) while LPIPS degrades from 0.3762 to 0.4110, showing CLIPIQA can be inflated at the cost of fidelity. This makes the headline CLIPIQA advantage partially self-referential. The claim is only partially circular because MANIQA and MUSIQ are not optimized in the final model and are also improved.
full rationale
Step 1 is the only circular step found. The paper's central technical contribution — the temporal analysis of ControlNet infusion, the timestep-aware adapter, and the staged L1/CLIPIQA training — is not a tautology: the adapter is a concrete architectural modification, and the L1 supervision, MANIQA/MUSIQ improvements, and qualitative comparisons are independent of the CLIPIQA reward. The temporal claim that LR information mainly influences early denoising is supported by the paper's own DiffBIR experiment (Fig. 1), not assumed by construction. Self-citations are limited to contextual mentions in Related Work (e.g., [36]) and are not load-bearing. A separate, non-circular correctness risk (not counted as a circular step) is the Sec. 3.3/Fig. 3 assertion that the reward-hacking point lies outside the adapter's sigmoid-bounded optimization space; this reachability claim is asserted rather than proved, and the paper's Tab. 3, row 2 suggests the adapter can in fact trade LPIPS for CLIPIQA. Overall, the CLIPIQA-as-reward-and-metric entanglement is a genuine but partial circularity, giving score 4.
Assumptions & free parameters
free parameters (6)
- t1 (L1 loss threshold) =
800
- t_clipiqa (CLIPIQA reward threshold) =
200
- lambda1 (L1 loss weight) =
0.01
- lambda_clipiqa (CLIPIQA reward weight) =
0.01
- Classifier-free guidance scale =
4.5
- Number of inference steps =
20
assumptions (6)
- standard math The DDPM forward process z_t = sqrt(alpha_bar_t) z_0 + sqrt(1 - alpha_bar_t) epsilon (Eq. 1) is the correct noise model for the SD latent space.
- domain assumption Pre-trained Stable Diffusion v2.1 and ControlNet provide a strong generative prior for general image SR.
- domain assumption CLIPIQA is a reliable proxy for human perceptual quality when used as a reward at late denoising timesteps (t <= 200).
- domain assumption The temporal pattern measured on DiffBIR (Fig. 1) transfers to the TASR-trained ControlNet and adapter.
- domain assumption The sigmoid-bounded adapter output space (alpha in [0,1]) excludes the CLIPIQA reward-hacking configuration.
- domain assumption Alternating optimization of ControlNet and adapter converges.
Cite this review
Pith. "Pith review of TASR: Timestep-Aware Diffusion Model for Image Super-Resolution." pith.science (2026). https://pith.science/paper/INL6YEQC
@misc{pith2026241203355,
author = {Pith},
title = {Pith review of: TASR: Timestep-Aware Diffusion Model for Image Super-Resolution},
year = {2026},
howpublished = {\url{https://pith.science/paper/INL6YEQC}},
note = {Machine review of arXiv:2412.03355}
}
read the original abstract
Diffusion models have recently achieved outstanding results in the field of image super-resolution. These methods typically inject low-resolution (LR) images via ControlNet.In this paper, we first explore the temporal dynamics of information infusion through ControlNet, revealing that the input from LR images predominantly influences the initial stages of the denoising process. Leveraging this insight, we introduce a novel timestep-aware diffusion model that adaptively integrates features from both ControlNet and the pre-trained Stable Diffusion (SD). Our method enhances the transmission of LR information in the early stages of diffusion to guarantee image fidelity and stimulates the generation ability of the SD model itself more in the later stages to enhance the detail of generated images. To train this method, we propose a timestep-aware training strategy that adopts distinct losses at varying timesteps and acts on disparate modules. Experiments on benchmark datasets demonstrate the effectiveness of our method. Code: https://github.com/SleepyLin/TASR
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
Semantic-Guided Cross-Sensor Super Resolution of Remote Sensing Images: A Gated Dual Conditioning Flow Matching Model
A gated dual-conditioning flow-matching model achieves 10 m→2 m cross-sensor super-resolution with a 38% FID reduction over the best baseline on a rare-landform (retrogressive thaw slump) benchmark.
Reference graph
Works this paper leans on
-
[1]
Aishwarya Agarwal, Srikrishna Karanam, Tripti Shukla, and Balaji Vasan Srini- vasan. 2023. An image is worth multiple words: Multi-attribute inversion for constrained text-to-image synthesis.arXiv preprint arXiv:2311.11919(2023)
arXiv 2023
-
[2]
Eirikur Agustsson and Radu Timofte. 2017. Ntire 2017 challenge on single image super-resolution: Dataset and study. InProceedings of the IEEE conference on computer vision and pattern recognition workshops. 126–135
2017
-
[3]
ao Yang, Rongyuan Wu, Peiran Ren, Xuansong Xie, and Lei Zhang. [n. d.]. Pixel- Aware Stable Diffusion for Realistic Image Super-Resolution and Personalized Stylization
-
[4]
Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Qinsheng Zhang, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, et al. 2022. ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers.arXiv preprint arXiv:2211.01324(2022)
arXiv 2022
-
[5]
Haolan Chen, Jinhua Hao, Kai Zhao, Kun Yuan, Ming Sun, Chao Zhou, and Wei Hu. 2024. CasSR: Activating Image Power for Real-World Image Super-Resolution. arXiv preprint arXiv:2403.11451(2024)
arXiv 2024
-
[6]
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhong- dao Wang, James Kwok, Ping Luo, Huchuan Lu, et al. 2023. Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis.arXiv preprint arXiv:2310.00426(2023)
arXiv 2023
-
[7]
Jooyoung Choi, Jungbeom Lee, Chaehun Shin, Sungwon Kim, Hyunwoo Kim, and Sungroh Yoon. 2022. Perception prioritized training of diffusion models. In IEEE Conf. Comput. Vis. Pattern Recog.11472–11481
work page 2022
-
[8]
Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion models beat gans on image synthesis.Adv. Neural Inform. Process. Syst.34 (2021), 8780–8794
work page 2021
Show all 55 references
-
[9]
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. 2024. Scaling rectified flow transformers for high-resolution image synthesis
2024
-
[10]
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2020. Generative adversarial networks.Commun. ACM63, 11 (2020), 139–144
2020
-
[11]
Shuhang Gu, Andreas Lugmayr, Martin Danelljan, Manuel Fritsche, Julien Lam- our, and Radu Timofte. 2019. Div8k: Diverse 8k resolution image dataset. In2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW). IEEE, 3512–3516
2019
-
[12]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models.Adv. Neural Inform. Process. Syst.33 (2020), 6840–6851
2020
-
[13]
Xiwei Hu, Rui Wang, Yixiao Fang, Bin Fu, Pei Cheng, and Gang Yu. 2024. Ella: Equip diffusion models with llm for enhanced semantic alignment.arXiv preprint arXiv:2403.05135(2024)
2024 arXiv
-
[14]
Tero Karras, Samuli Laine, and Timo Aila. 2019. A style-based generator archi- tecture for generative adversarial networks. InIEEE Conf. Comput. Vis. Pattern Recog.4401–4410
2019
-
[15]
D Kinga, Jimmy Ba Adam, et al. 2015. A method for stochastic optimization. In Int. Conf. Learn. Represent., Vol. 5. San Diego, California;, 6
2015
-
[16]
Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. 2021. Swinir: Image restoration using swin transformer. InInt. Conf. Comput. Vis.1833–1844
2021
-
[17]
Xinqi Lin, Jingwen He, Ziyan Chen, Zhaoyang Lyu, Bo Dai, Fanghua Yu, Wanli Ouyang, Yu Qiao, and Chao Dong. 2024. DiffBIR: Towards Blind Image Restora- tion with Generative Diffusion Prior. arXiv:2308.15070 [cs.CV]
2024 arXiv
-
[18]
Anran Liu, Yihao Liu, Jinjin Gu, Yu Qiao, and Chao Dong. 2022. Blind image super-resolution: A survey and beyond.IEEE Trans. Pattern Anal. Mach. Intell. 45, 5 (2022), 5461–5480
2022
-
[19]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruc- tion tuning.Advances in neural information processing systems36 (2024)
2024
-
[20]
Haozhe Liu, Wentian Zhang, Jinheng Xie, Francesco Faccio, Mengmeng Xu, Tao Xiang, Mike Zheng Shou, Juan-Manuel Perez-Rua, and Jürgen Schmidhuber. 2024. Faster Diffusion via Temporal Attention Decomposition.arXiv e-prints(2024), arXiv–2404
2024
-
[21]
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2021. Glide: Towards photorealistic image generation and editing with text-guided diffusion models.arXiv preprint arXiv:2112.10741(2021)
2021 arXiv
-
[22]
Alexander Quinn Nichol and Prafulla Dhariwal. 2021. Improved denoising diffu- sion probabilistic models. PMLR, 8162–8171
2021
-
[23]
William Peebles and Saining Xie. 2023. Scalable diffusion models with transform- ers. InInt. Conf. Comput. Vis.4195–4205
2023
-
[24]
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. 2018. Film: Visual reasoning with a general conditioning layer. In AAAI, Vol. 32
2018
-
[25]
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023. Sdxl: Improving latent diffusion models for high-resolution image synthesis.arXiv preprint arXiv:2307.01952 (2023)
2023 arXiv
-
[26]
Chenyang Qi, Zhengzhong Tu, Keren Ye, Mauricio Delbracio, Peyman Milanfar, Qifeng Chen, and Hossein Talebi. 2023. TIP: Text-Driven Image Processing with Semantic and Restoration Instructions.arXiv:2312.11595(2023)
2023 arXiv
-
[27]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. PMLR, 8748–8763
2021
-
[29]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In IEEE Conf. Comput. Vis. Pattern Recog.10684–10695
2022
-
[30]
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. 2022. Photorealistic text-to-image diffusion models with deep language understanding.Adv. Neural Inform. Proce...
2022
-
[31]
Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger. 2022. Defining and characterizing reward gaming.Adv. Neural Inform. Process. Syst.35 (2022), 9460–9471
2022
-
[32]
Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020. Denoising Diffusion Implicit Models.arXiv:2010.02502(October 2020). https://arxiv.org/abs/2010. 02502
2020 arXiv
-
[33]
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2020. Score-based generative modeling through stochastic differential equations.arXiv preprint arXiv:2011.13456(2020)
2020 arXiv
-
[34]
Haoze Sun, Wenbo Li, Jianzhuang Liu, Haoyu Chen, Renjing Pei, Xueyi Zou, Youliang Yan, and Yujiu Yang. 2024. Coser: Bridging image and language for cognitive super-resolution. InIEEE Conf. Comput. Vis. Pattern Recog.25868–25878
2024
-
[35]
Lingchen Sun, Rongyuan Wu, Zhengqiang Zhang, Hongwei Yong, and Lei Zhang
-
[36]
Xiaopeng Sun, Qinwei Lin, Yu Gao, Yujie Zhong, Chengjian Feng, Dengjie Li, Zheng Zhao, Jie Hu, and Lin Ma. 2024. Rfsr: Improving isr diffusion models via reward feedback learning.arXiv preprint arXiv:2412.03268(2024)
2024 arXiv
-
[37]
Radu Timofte, Eirikur Agustsson, Luc Van Gool, Ming-Hsuan Yang, and Lei Zhang. 2017. Ntire 2017 challenge on single image super-resolution: Methods and results. InProceedings of the IEEE conference on computer vision and pattern recognition workshops. 114–125
2017
-
[38]
Jianyi Wang, Kelvin CK Chan, and Chen Change Loy. 2023. Exploring CLIP for Assessing the Look and Feel of Images. InAAAI
2023
-
[39]
Chan, and Chen Change Loy
Jianyi Wang, Zongsheng Yue, Shangchen Zhou, Kelvin C.K. Chan, and Chen Change Loy. 2024. Exploiting Diffusion Prior for Real-World Image Super- Resolution. (2024)
2024
-
[40]
Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan. 2021. Real-esrgan: Training real-world blind super-resolution with pure synthetic data. InInt. Conf. Comput. Vis.1905–1914
2021
-
[41]
Xintao Wang, Ke Yu, Chao Dong, and Chen Change Loy. 2018. Recovering realistic texture in image super-resolution by deep spatial feature transform. In IEEE Conf. Comput. Vis. Pattern Recog.606–615
2018
-
[42]
Xintao Wang, Ke Yu, Shixiang Wu, Jinjin Gu, Yihao Liu, Chao Dong, Yu Qiao, and Chen Change Loy. 2018. Esrgan: Enhanced super-resolution generative adversarial networks. InProceedings of the European conference on computer vision (ECCV) workshops. 0–0
2018
-
[43]
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity.IEEE Trans. Image Process.13, 4 (2004), 600–612
2004
-
[44]
Pengxu Wei, Ziwei Xie, Hannan Lu, Zongyuan Zhan, Qixiang Ye, Wangmeng Zuo, and Liang Lin. 2020. Component divide-and-conquer for real-world image super-resolution. InEur. Conf. Comput. Vis.Springer, 101–117
2020
-
[45]
Rongyuan Wu, Tao Yang, Lingchen Sun, Zhengqiang Zhang, Shuai Li, and Lei Zhang. 2024. Seesr: Towards semantics-aware real-world image super-resolution. InIEEE Conf. Comput. Vis. Pattern Recog.25456–25467
2024
-
[46]
Rui Xie, Ying Tai, Chen Zhao, Kai Zhang, Zhenyu Zhang, Jun Zhou, Xiaoqian Ye, Qian Wang, and Jian Yang. 2024. Addsr: Accelerating diffusion-based blind super- resolution with adversarial diffusion distillation.arXiv preprint arXiv:2404.01717 (2024)
2024 arXiv
-
[47]
Sidi Yang, Tianhe Wu, Shuwei Shi, Shanshan Lao, Yuan Gong, Mingdeng Cao, Jiahao Wang, and Yujiu Yang. 2022. Maniqa: Multi-dimension attention network for no-reference image quality assessment. InIEEE Conf. Comput. Vis. Pattern MM ’25, October 27–31, 2025, Dublin, Ireland Qinwe...
2022
-
[48]
Fanghua Yu, Jinjin Gu, Zheyuan Li, Jinfan Hu, Xiangtao Kong, Xintao Wang, Jingwen He, Yu Qiao, and Chao Dong. 2024. Scaling up to excellence: Practicing model scaling for photo-realistic image restoration in the wild. InIEEE Conf. Comput. Vis. Pattern Recog.25669–25680
2024
-
[49]
Zongsheng Yue, Jianyi Wang, and Chen Change Loy. 2024. Resshift: Efficient diffusion model for image super-resolution by residual shifting.Adv. Neural Inform. Process. Syst.36 (2024)
2024
-
[50]
Kai Zhang, Jingyun Liang, Luc Van Gool, and Radu Timofte. 2021. Designing a Practical Degradation Model for Deep Blind Image Super-Resolution. InInt. Conf. Comput. Vis.4791–4800
2021
-
[51]
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding conditional control to text-to-image diffusion models. InInt. Conf. Comput. Vis.3836–3847
2023
-
[52]
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang
-
[53]
Wenlong Zhang, Yihao Liu, Chao Dong, and Yu Qiao. 2019. Ranksrgan: Generative adversarial networks with ranker for image super-resolution. InInt. Conf. Comput. Vis.3096–3105
2019
-
[54]
Youcai Zhang, Xinyu Huang, Jinyu Ma, Zhaoyang Li, Zhaochuan Luo, Yanchun Xie, Yuzhuo Qin, Tong Luo, Yaqian Li, Shilong Liu, et al. 2024. Recognize anything: A strong image tagging model. InIEEE Conf. Comput. Vis. Pattern Recog.1724– 1732
2024
-
[2018]
In IEEE Conf
The unreasonable effectiveness of deep features as a perceptual metric. In IEEE Conf. Comput. Vis. Pattern Recog.586–595
-
[2023]
Improving the stability of diffusion models for content consistent super- resolution.arXiv e-prints(2023), arXiv–2401
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.