REVIEW 3 major objections 7 minor 73 references
Latent Posterior-Mean Rectified Flow for Higher-Fidelity Perceptual Face Restoration
T0 review · 3 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Latent-PMRF moves posterior-mean rectified flow into VAE latent space, bounding minimum distortion by the VAE's reconstruction error and converging 5.79x faster in FID.
desk verdict Latent-PMRF is a useful empirical extension of PMRF to VAE space, but the paper's central distortion-bound claim only applies to the source, not the flow endpoint, and the speedup number needs a precise definition. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the choice of source distribution $z_0 = E(E[X\mid Y=y])$—the latent encoding of the posterior mean—combined with a rectified flow that transports it to the target latent $z_1 = E(x)$. Rectified flow trains a velocity network $v^\theta_t$ on linear interpolations $z_t = (1-t)z_0 + t z_1$ via the conditional flow-matching loss, so that sampling solves the ODE $dz_t/dt = v^\theta_t(z_t)$. A second component is Sim-VAE, a simplified VQGAN-style autoencoder with pixel-wise layer normalization, convolutional layers in place of middle self-attention, and channel adjustments folded into resizing layers; lowering the VAE's reconstruction error lowers the distortion bound directly.
What would settle it
On a fixed low-quality test set with known ground truth, compute the RMSE between Latent-PMRF's decoded output $D(z_1)$ and the posterior-mean estimate $\hat{x}=E[X\mid Y=y]$; if that error systematically exceeds the VAE's reconstruction error on clean natural images by more than the flow's training residual explains, the claimed distortion bound is violated.
Extended reading notes
Core claim
Latent-PMRF claims that running Posterior-Mean Rectified Flow in the latent space of a VAE, rather than in pixel space, is both distortion-safe and perceptually more efficient. The key design decision is the source distribution: instead of the posterior mean of latent encodings, $E[E(X)\mid Y]$, the paper uses the latent encoding of the posterior mean, $z_0 = E(E[X\mid Y=y])$. Because decoding this source reproduces the posterior mean whenever the VAE reconstructs perfectly, the minimum distortion of the whole pipeline is bounded by the VAE's reconstruction error, a guarantee the alternative source distribution does not carry. On blind face restoration benchmarks, the method matches or exceeds PMRF's fidelity while reaching better perceptual scores with a 5.79x speedup in FID convergence, and the paper attributes part of the gain to its proposed Sim-VAE, which outperforms Stable Diffusion-style VAEs in both reconstruction and restoration.
Load-bearing premise
The distortion bound assumes the latent flow transports the encoded posterior mean to the encoded high-quality distribution exactly, and that the VAE reconstructs posterior-mean estimates as faithfully as natural images, because the paper's proof relies on perfect VAE reconstruction.
Editorial extensions
If this is right
- Better VAE reconstruction directly lowers the minimum distortion Latent-PMRF can achieve, so future restoration gains can come from improving the autoencoder rather than scaling up the flow model.
- The 5.79x FID-convergence speedup means high-quality restoration can be trained with a fraction of the compute needed by pixel-space PMRF, making the approach practical at larger scales.
- Using the posterior mean of latent encodings as the source distribution drops the distortion guarantee, which explains why the paper predicts fidelity loss in that variant and keeps fidelity in its own design.
- On CelebA-Test, Latent-PMRF stays above 26.3 dB PSNR with top-tier statistical-distance scores, placing it closer to the perception-distortion frontier than GAN-based and diffusion-based baselines.
Reading between the lines
- Beyond the paper: the same source-distribution design should transfer to other restoration tasks with a posterior-mean estimator, such as super-resolution and deblurring, since the distortion-bound argument does not use face-specific structure.
- Beyond the paper: the theory predicts a monotone link between VAE reconstruction error and downstream restoration fidelity; a clean test would train Latent-PMRF on VAEs with graded reconstruction quality and check whether PSNR and identity metrics follow the same ordering.
- Beyond the paper: the 'human perception' claim is operationally supported by feature-space metrics such as FID, LPIPS, MUSIQ, and CLIP-IQA; a direct human-rating study would test whether the latent-space advantage survives subjective evaluation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Latent-PMRF, a latent-space variant of Posterior-Mean Rectified Flow for blind face restoration. A pretrained posterior-mean estimator produces x̂ = E[X|Y=y]; the VAE encoder maps x̂ to a source latent z0; a rectified-flow velocity network transports z0 to z1 in the latent distribution of high-quality images; and the decoder outputs D(z1). The authors argue that using E(x̂) as the source bounds the final distortion by the VAE's reconstruction error, propose a modified VAE (Sim-VAE), and report faster perceptual convergence and an improved perception-distortion tradeoff over PMRF on face benchmarks.
Significance. If the central theoretical claim is correct, the paper offers a practical way to reduce training compute for flow-based restoration while improving perceptual quality and retaining fidelity. The empirical study is reasonably broad: it includes synthetic and real-world benchmarks, convergence curves, and ablations of VAE architecture. The paper also evaluates with external feature-space metrics (FID, LPIPS, DINOv2) rather than optimizing those objectives directly, which mitigates circularity concerns. However, the headline distortion bound is not proven for the actual output D(z1), and the reported dependence on VAE reconstruction quality does not clearly support the bound.
major comments (3)
- [Section 3.2, Eq. (4)] The derivation only shows that, for a perfect VAE, the decoded source image D(z0) equals the posterior mean x̂, where z0=E(x̂). The actual restoration output is D(z1), with z1 obtained by integrating the learned velocity vθ from z0 (Section 3.3). Equation (5) matches the marginal distribution of z1 to p_{E(X)}; marginal matching does not constrain the coupling between z1 and the particular y that produced z0, so z1 may encode a different face than x. Therefore the abstract's and Section 3.2's claim that the final distortion is bounded by the VAE's reconstruction error is not established. A valid proof needs an explicit bound on E||D(z1)-X||^2, for instance via conditional flow matching or an optimal-transport coupling term; without it, the result should be stated only as a property of the latent source.
- [Table 3] The empirical channel ablations are in tension with the proposed bound. Raising the latent channels from 16 to 48 improves VAE reconstruction PSNR from 37.90 to 45.06 and LPIPS from 0.0261 to 0.0033, yet the restoration PSNR remains almost unchanged (26.44, 26.39, 26.38, 26.46) and restoration LPIPS is best at 16 channels and worsens slightly at 48. If the final distortion were dominated by the VAE reconstruction error, one would expect a clear monotone gain in restoration fidelity; the flat PSNR suggests that the flow's stochastic coupling, not the VAE, controls the final distortion. The authors should either provide an explanation for this decoupling or modify the claimed bound.
- [Section 6.2 and Table 4] The headline '5.79× speedup' is not precisely defined. The text says the speedup is in terms of FID, but does not state the FID threshold, the number of training steps or wall-clock time to reach it, or how the curves in Figure 1 and Figure 6 were evaluated (e.g., validation set, smoothing, number of seeds). In Table 4, PMRF* is described as trained under the same compute budget, but the training configurations are not given in the table or text; please report iterations, batch size, and ideally GPU-hours or FLOPs. Without this, the convergence-efficiency claim cannot be independently verified.
minor comments (7)
- [References] Reference [42] is miscited: the CodeFormer entry points to a GNN-based binary code similarity paper, not the face restoration method by Zhou et al. Please correct.
- [Section 4.2] The statement that the adversarial loss is unnecessary 'with sufficient model capacity' is not supported by an ablation; please add a comparison with and without L_adv, or flag this as a heuristic.
- [Section 6.2] The sentence 'we train both PMRF and Latent-PMRF using Sim-VAE' is ambiguous because PMRF operates in pixel space; clarify how Sim-VAE is involved in each baseline.
- [Figure 8] The term 'IndRMSE' is introduced without definition and the caption says it represents the RMSE of each method [46]; please define the metric and state how FID is computed on the real-world datasets without ground truth.
- [Table 1] The notation 'f8c4' is not explained; please define downsampling factor and channel count. Also 'Sim-VAEf8c32' is missing a space.
- [Section 5, concurrent works] The claim that ELIR 'leads to significant fidelity degradation' is made without a quantitative comparison; either add a comparison or soften the statement.
- [Section 1 and Section 3.1] The definition of perceptual quality as 'how humans distinguish between two image distributions' is informal; please state the formal definition (e.g., statistical distance as in [2]) and clarify the relationship to feature-space metrics.
Circularity Check
No circular step in the derivation chain; the only rubric-level concern is a peripheral non-load-bearing self-citation, and the Section 3.2 distortion-bound gap is a correctness issue rather than an input-equivalence.
full rationale
Latent-PMRF's empirical claims are checked against external benchmarks and standard objectives rather than against quantities fitted by the model. The flow training loss in Eq. (5) is a conventional unconditional conditional-flow-matching objective, and the restoration metrics (FID, LPIPS, DINOv2, PSNR) are not used as training losses for the restoration network. Sim-VAE is trained with reconstruction, perceptual, and KL losses, not with the evaluation metrics, so the reconstruction-versus-restoration comparisons in Tables 1-3 are not fitted-input predictions. The central theoretical claim in Section 3.2 is a gap, not a circularity: Eq. (4) bounds the decoded source D(E(E[X|Y])) by the VAE's reconstruction error, while the actual restoration output is the ODE endpoint D(z1) (Section 3.3); no equation equates these two quantities, and Eq. (5) does not couple z1 to the same scene as z0. This is an unsupported inference about the final distortion, but it is not equivalent to an input by construction. The only self-citation is [45], cited in the list of typical degradations ('downsampling [12, 39, 45]'), and it is not load-bearing; no uniqueness theorem or ansatz is imported from the authors' prior work. Accordingly, no prediction in the paper reduces to a fitted parameter or a self-citation chain. The score of 2 reflects only the presence of that minor peripheral self-citation under the rubric; there are no circular steps.
Assumptions & free parameters
free parameters (3)
- VAE latent channels =
32
- Euler sampling steps =
25
- KL regularization weight =
1e-6
assumptions (4)
- domain assumption Rectified flow with conditional flow matching can transport the source latent distribution to the target latent distribution exactly with sufficient capacity and training
- domain assumption Distances in VAE latent space align with human perception at least as well as pixel distances
- ad hoc to paper The VAE reconstructs posterior-mean images, not just natural images, near-perfectly
- domain assumption The pretrained DifFace posterior-mean estimator provides an accurate E[X|Y=y]
Cite this review
Pith. "Pith review of Latent Posterior-Mean Rectified Flow for Higher-Fidelity Perceptual Face Restoration." pith.science (2026). https://pith.science/paper/LLGZSQNE
@misc{pith2026250700447,
author = {Pith},
title = {Pith review of: Latent Posterior-Mean Rectified Flow for Higher-Fidelity Perceptual Face Restoration},
year = {2026},
howpublished = {\url{https://pith.science/paper/LLGZSQNE}},
note = {Machine review of arXiv:2507.00447}
}
read the original abstract
The Perception-Distortion tradeoff (PD-tradeoff) theory suggests that face restoration algorithms must balance perceptual quality and fidelity. To achieve minimal distortion while maintaining perfect perceptual quality, Posterior-Mean Rectified Flow (PMRF) proposes a flow based approach where source distribution is minimum distortion estimations. Although PMRF is shown to be effective, its pixel-space modeling approach limits its ability to align with human perception, where human perception is defined as how humans distinguish between two image distributions. In this work, we propose Latent-PMRF, which reformulates PMRF in the latent space of a variational autoencoder (VAE), facilitating better alignment with human perception during optimization. By defining the source distribution on latent representations of minimum distortion estimation, we bound the minimum distortion by the VAE's reconstruction error. Moreover, we reveal the design of VAE is crucial, and our proposed VAE significantly outperforms existing VAEs in both reconstruction and restoration. Extensive experiments on blind face restoration demonstrate the superiority of Latent-PMRF, offering an improved PD-tradeoff compared to existing methods, along with remarkable convergence efficiency, achieving a 5.79X speedup over PMRF in terms of FID. Our code will be available as open-source.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Build- ing normalizing flows with stochastic interpolants
Michael Samuel Albergo and Eric Vanden-Eijnden. Build- ing normalizing flows with stochastic interpolants. In ICLR,
-
[2]
The perception-distortion tradeoff
Yochai Blau and Tomer Michaeli. The perception-distortion tradeoff. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6228–6237, 2018. 1, 5
work page 2018
-
[3]
Deep compression autoencoder for efficient high-resolution diffusion models
Junyu Chen, Han Cai, Junsong Chen, Enze Xie, Shang Yang, Haotian Tang, Muyang Li, and Song Han. Deep compression autoencoder for efficient high-resolution diffusion models. In ICLR, 2025. 4
work page 2025
-
[4]
Simple baselines for image restoration
Liangyu Chen, Xiaojie Chu, Xiangyu Zhang, and Jian Sun. Simple baselines for image restoration. In European confer- ence on computer vision, pages 17–33. Springer, 2022. 4
work page 2022
-
[5]
Towards real-world blind face restoration with generative diffusion prior
Xiaoxu Chen, Jingfan Tan, Tao Wang, Kaihao Zhang, Wen- han Luo, and Xiaochun Cao. Towards real-world blind face restoration with generative diffusion prior. IEEE Transac- tions on Circuits and Systems for Video Technology , 2024. 5
work page 2024
-
[6]
Xception: Deep learning with depthwise separable convolutions
Franc ¸ois Chollet. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE con- ference on computer vision and pattern recognition , pages 1251–1258, 2017. 5
2017
-
[7]
Efficient image restoration via latent con- sistency flow matching
Elad Cohen, Idan Achituve, Idit Diamant, Arnon Netzer, and Hai Victor Habi. Efficient image restoration via latent con- sistency flow matching. arXiv preprint arXiv:2502.03500 ,
-
[8]
Scalable high-resolution pixel-space image syn- thesis with hourglass diffusion transformers
Katherine Crowson, Stefan Andreas Baumann, Alex Birch, Tanishq Mathew Abraham, Daniel Z Kaplan, and Enrico Shippole. Scalable high-resolution pixel-space image syn- thesis with hourglass diffusion transformers. In Forty-first International Conference on Machine Learning, 2024. 5
work page 2024
Show all 73 references
-
[9]
Arcface: Additive angular margin loss for deep face recognition
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 4690–4699, 2019. 6
2019
-
[10]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in neural informa- tion processing systems, 34:8780–8794, 2021. 5
2021
-
[11]
Image quality assessment: Unifying structure and texture similarity
Keyan Ding, Kede Ma, Shiqi Wang, and Eero P Simoncelli. Image quality assessment: Unifying structure and texture similarity. IEEE transactions on pattern analysis and ma- chine intelligence, 44(5):2567–2581, 2020. 6
2020
-
[12]
Learning a deep convolutional network for image super-resolution
Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. Learning a deep convolutional network for image super-resolution. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part IV 13 , pages 184–199. Springer,
2014
-
[13]
The vae used for stable diffusion 1.x/2.x and other models (kl-f8) has a critical flaw
drhead. The vae used for stable diffusion 1.x/2.x and other models (kl-f8) has a critical flaw. https : / / www . reddit . com / r / StableDiffusion / comments / 1ag5h5s / the _ vae _ used _ for _ stable _ diffusion _ 1x2x _ and _ other / ?utm _ source=share&utm_medium=web3x&u...
2024
-
[14]
Image denoising: The deep learning revolution and beyond—a sur- vey paper
Michael Elad, Bahjat Kawar, and Gregory Vaksman. Image denoising: The deep learning revolution and beyond—a sur- vey paper. SIAM Journal on Imaging Sciences, 16(3):1594– 1654, 2023. 1
2023
-
[15]
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12873–12883, 2021. 4, 5
2021
-
[16]
Scaling recti- fied flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling recti- fied flow transformers for high-resolution image synthesis. In Forty-first International Conference on Mach...
2024
-
[17]
A theory of the distortion-perception tradeoff in wasserstein space
Dror Freirich, Tomer Michaeli, and Ron Meir. A theory of the distortion-perception tradeoff in wasserstein space. Advances in Neural Information Processing Systems , 34: 25661–25672, 2021. 1
2021
-
[18]
Lumina-t2x: Scalable flow- based large diffusion transformer for flexible resolution gen- eration
Peng Gao, Le Zhuo, Dongyang Liu, Ruoyi Du, Xu Luo, Longtian Qiu, Yuhang Zhang, Rongjie Huang, Shijie Geng, Renrui Zhang, Junlin Xie, Wenqi Shao, Zhengkai Jiang, Tianshuo Yang, Weicai Ye, Tong He, Jingwen He, Junjun He, Yu Qiao, and Hongsheng Li. Lumina-t2x: Scalable flow- base...
2025
-
[19]
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. Commu- nications of the ACM, 63(11):139–144, 2020. 2, 5
2020
-
[20]
Vqfr: Blind face restoration with vector-quantized dictionary and parallel de- coder
Yuchao Gu, Xintao Wang, Liangbin Xie, Chao Dong, Gen Li, Ying Shan, and Ming-Ming Cheng. Vqfr: Blind face restoration with vector-quantized dictionary and parallel de- coder. In European Conference on Computer Vision, pages 126–143. Springer, 2022. 2, 5, 7
2022
-
[21]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 4
2016
-
[22]
Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium. Advances in neural information processing systems , 30, 2017. 2, 3, 6
2017
-
[23]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 1, 5
2020
-
[24]
Labeled faces in the wild: A database forstudying face recognition in unconstrained environments
Gary B Huang, Marwan Mattar, Tamara Berg, and Eric Learned-Miller. Labeled faces in the wild: A database forstudying face recognition in unconstrained environments. In Workshop on faces in’Real-Life’Images: detection, align- ment, and recognition, 2008. 8
2008
-
[25]
Re- thinking fid: Towards a better evaluation metric for image generation
Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner, Ayan Chakrabarti, and Sanjiv Kumar. Re- thinking fid: Towards a better evaluation metric for image generation. In Proceedings of the IEEE/CVF Conference 9 on Computer Vision and Pattern Recognition , pages 9...
2024
-
[26]
Towards flex- ible blind jpeg artifacts removal
Jiaxi Jiang, Kai Zhang, and Radu Timofte. Towards flex- ible blind jpeg artifacts removal. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 4997–5006, 2021. 1
2021
-
[27]
Percep- tual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Percep- tual losses for real-time style transfer and super-resolution. In Computer Vision–ECCV 2016: 14th European Confer- ence, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14, pages 694–711. Springer, 2016. 5
2016
-
[28]
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 4401–4410, 2019. 5
2019
-
[29]
Analyzing and improv- ing the image quality of stylegan
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improv- ing the image quality of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8110–8119, 2020. 4
2020
-
[30]
Musiq: Multi-scale image quality transformer
Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. Musiq: Multi-scale image quality transformer. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5148–5157, 2021. 6
2021
-
[31]
Auto-encoding variational bayes
Diederik P Kingma. Auto-encoding variational bayes. In ICLR, 2014. 2
2014
-
[32]
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 5
2015
-
[33]
Ensembling off-the-shelf models for gan training
Nupur Kumari, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. Ensembling off-the-shelf models for gan training. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 10651–10662, 2022. 2, 3
2022
-
[34]
Black Forest Labs. Flux. https://github.com/ black- forest- labs/flux, 2024. Accessed: 2025- 02-07. 3, 4, 5
2024
-
[35]
Photo- realistic single image super-resolution using a generative ad- versarial network
Christian Ledig, Lucas Theis, Ferenc Husz´ar, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, et al. Photo- realistic single image super-resolution using a generative ad- versarial network. In Proceedings of the IE...
-
[36]
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. In NeurIPS, 2016. 4
2016
-
[37]
Srdiff: Single image super-resolution with diffusion probabilistic models
Haoying Li, Yifan Yang, Meng Chang, Shiqi Chen, Huajun Feng, Zhihai Xu, Qi Li, and Yueting Chen. Srdiff: Single image super-resolution with diffusion probabilistic models. Neurocomputing, 479:47–59, 2022. 1
2022
-
[38]
Lsdir: A large scale dataset for image restoration
Yawei Li, Kai Zhang, Jingyun Liang, Jiezhang Cao, Ce Liu, Rui Gong, Yulun Zhang, Hao Tang, Yun Liu, Denis Deman- dolx, et al. Lsdir: A large scale dataset for image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1775–17...
2023
-
[39]
Swinir: Image restoration us- ing swin transformer
Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. Swinir: Image restoration us- ing swin transformer. InProceedings of the IEEE/CVF inter- national conference on computer vision , pages 1833–1844,
-
[40]
Diff- bir: Toward blind image restoration with generative diffusion prior
Xinqi Lin, Jingwen He, Ziyan Chen, Zhaoyang Lyu, Bo Dai, Fanghua Yu, Yu Qiao, Wanli Ouyang, and Chao Dong. Diff- bir: Toward blind image restoration with generative diffusion prior. In European Conference on Computer Vision , pages 430–448. Springer, 2024. 2, 5, 7
2024
-
[41]
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maxim- ilian Nickel, and Matthew Le. Flow matching for genera- tive modeling. In The Eleventh International Conference on Learning Representations, 2023. 2
2023
-
[42]
Codeformer: A gnn-nested trans- former model for binary code similarity detection
Guangming Liu, Xin Zhou, Jianmin Pang, Feng Yue, Wenfu Liu, and Junchao Wang. Codeformer: A gnn-nested trans- former model for binary code similarity detection. Electron- ics, 12(7):1722, 2023. 2, 5, 7
2023
-
[43]
Flow straight and fast: Learning to generate and transfer data with rectified flow
Xingchao Liu, Chengyue Gong, and qiang liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. In The Eleventh International Conference on Learning Representations, 2023. 2
2023
-
[44]
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feicht- enhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition , pages 11976–11986,
-
[45]
On the effectiveness of spectral discriminators for perceptual qual- ity improvement
Xin Luo, Yunan Zhu, Shunxin Xu, and Dong Liu. On the effectiveness of spectral discriminators for perceptual qual- ity improvement. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, pages 13243–13253,
-
[46]
Posterior- mean rectified flow: Towards minimum MSE photo-realistic image restoration
Guy Ohayon, Tomer Michaeli, and Michael Elad. Posterior- mean rectified flow: Towards minimum MSE photo-realistic image restoration. In The Thirteenth International Confer- ence on Learning Representations, 2025. 1, 5, 6, 7, 8
2025
-
[47]
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 6
2023 arXiv
-
[48]
SDXL: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M ¨uller, Joe Penna, and Robin Rombach. SDXL: Improving latent diffusion models for high-resolution image synthesis. In The Twelfth Interna- tional Conference on Learning Representations, 2024. 2, 3, 4, 5
2024
-
[49]
Train short, test long: Attention with linear biases enables input length ex- trapolation
Ofir Press, Noah Smith, and Mike Lewis. Train short, test long: Attention with linear biases enables input length ex- trapolation. In ICLR, 2022. 4
2022
-
[50]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2, 3, 4, 5, 7
2022
-
[51]
Seyedmorteza Sadat, Jakob Buhmann, Derek Bradley, Otmar Hilliges, and Romann M. Weber. LiteV AE: Lightweight and efficient variational autoencoders for latent diffusion models. In The Thirty-eighth Annual Conference on Neural Informa- tion Processing Systems, 2024. 4 10
2024
-
[52]
Projected gans converge faster
Axel Sauer, Kashyap Chitta, Jens M ¨uller, and Andreas Geiger. Projected gans converge faster. Advances in Neural Information Processing Systems, 34:17480–17492, 2021. 2, 3
2021
-
[53]
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International confer- ence on machine learning, pages 2256–2265. PMLR, 2015. 1
2015
-
[54]
Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models
George Stein, Jesse Cresswell, Rasa Hosseinzadeh, Yi Sui, Brendan Ross, Valentin Villecroze, Zhaoyan Liu, Anthony L Caterini, Eric Taylor, and Gabriel Loaiza-Ganem. Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models. Advances i...
2024
-
[55]
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 1–9, 2015. 2
2015
-
[56]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 4
2017
-
[57]
Ex- ploring clip for assessing the look and feel of images
Jianyi Wang, Kelvin CK Chan, and Chen Change Loy. Ex- ploring clip for assessing the look and feel of images. InPro- ceedings of the AAAI conference on artificial intelligence , pages 2555–2563, 2023. 6
2023
-
[58]
Exploiting diffusion prior for real-world image super-resolution
Jianyi Wang, Zongsheng Yue, Shangchen Zhou, Kelvin CK Chan, and Chen Change Loy. Exploiting diffusion prior for real-world image super-resolution. International Journal of Computer Vision, 132(12):5929–5949, 2024. 1, 2
2024
-
[59]
A survey of deep face restoration: Denoise, super-resolution, deblur, artifact re- moval
Tao Wang, Kaihao Zhang, Xuanxi Chen, Wenhan Luo, Jiankang Deng, Tong Lu, Xiaochun Cao, Wei Liu, Hong- dong Li, and Stefanos Zafeiriou. A survey of deep face restoration: Denoise, super-resolution, deblur, artifact re- moval. arXiv preprint arXiv:2211.02831, 2022. 1
2022
-
[60]
To- wards real-world blind face restoration with generative fa- cial prior
Xintao Wang, Yu Li, Honglun Zhang, and Ying Shan. To- wards real-world blind face restoration with generative fa- cial prior. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9168–9178,
-
[61]
Real-esrgan: Training real-world blind super-resolution with pure synthetic data
Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan. Real-esrgan: Training real-world blind super-resolution with pure synthetic data. In Proceedings of the IEEE/CVF inter- national conference on computer vision , pages 1905–1914,
1905
-
[62]
Mul- tiscale structural similarity for image quality assessment
Zhou Wang, Eero P Simoncelli, and Alan C Bovik. Mul- tiscale structural similarity for image quality assessment. In The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003, pages 1398–1402. Ieee, 2003. 5
2003
-
[63]
Restoreformer: High-quality blind face restoration from undegraded key-value pairs
Zhouxia Wang, Jiawei Zhang, Runjian Chen, Wenping Wang, and Ping Luo. Restoreformer: High-quality blind face restoration from undegraded key-value pairs. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 17512–17521, 2022. 2, 5, 7
2022
-
[64]
Q-align: Teaching LMMs for visual scoring via discrete text-defined levels
Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, Qiong Yan, Xiongkuo Min, Guang- tao Zhai, and Weisi Lin. Q-align: Teaching LMMs for visual scoring via discrete text-defined levels. InProceedings of ...
2024
-
[65]
Group normalization
Yuxin Wu and Kaiming He. Group normalization. In Pro- ceedings of the European conference on computer vision (ECCV), pages 3–19, 2018. 4
2018
-
[66]
Consistency flow matching: Defin- ing straight flows with velocity consistency
Ling Yang, Zixiang Zhang, Zhilong Zhang, Xingchao Liu, Minkai Xu, Wentao Zhang, Chenlin Meng, Stefano Er- mon, and Bin Cui. Consistency flow matching: Defin- ing straight flows with velocity consistency. arXiv preprint arXiv:2407.02398, 2024. 5
2024 arXiv
-
[67]
Difface: Blind face restoration with diffused error contraction
Zongsheng Yue and Chen Change Loy. Difface: Blind face restoration with diffused error contraction. IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 2024. 5, 6, 7
2024
-
[68]
Resshift: Efficient diffusion model for image super- resolution by residual shifting
Zongsheng Yue, Jianyi Wang, and Chen Change Loy. Resshift: Efficient diffusion model for image super- resolution by residual shifting. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. 2, 5, 7
2023
-
[69]
Effi- cient diffusion model for image restoration by residual shift- ing
Zongsheng Yue, Jianyi Wang, and Chen Change Loy. Effi- cient diffusion model for image restoration by residual shift- ing. IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(1):116–130, 2025. 7
2025
-
[70]
Deep image deblurring: A survey
Kaihao Zhang, Wenqi Ren, Wenhan Luo, Wei-Sheng Lai, Bj¨orn Stenger, Ming-Hsuan Yang, and Hongdong Li. Deep image deblurring: A survey. International Journal of Com- puter Vision, 130(9):2103–2130, 2022. 1
2022
-
[71]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 586–595, 2018. 2, 3, 6
2018
-
[72]
Towards robust blind face restora- tion with codebook lookup transformer
Shangchen Zhou, Kelvin Chan, Chongyi Li, and Chen Change Loy. Towards robust blind face restora- tion with codebook lookup transformer. Advances in Neural Information Processing Systems, 35:30599–30611, 2022. 8
2022
-
[73]
Flowie: Efficient image enhancement via rectified flow
Yixuan Zhu, Wenliang Zhao, Ao Li, Yansong Tang, Jie Zhou, and Jiwen Lu. Flowie: Efficient image enhancement via rectified flow. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 13–22, 2024. 1, 2, 5, 7 11
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.