REVIEW 3 major objections 6 minor 2 cited by
SILO: Solving Inverse Problems with Latent Operators
T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read SILO claims that a learned latent degradation operator lets latent diffusion models solve inverse problems entirely in latent space, improving perceptual quality while cutting runtime roughly 3-10x.
desk verdict SILO's latent-operator idea is fresh and the speedups look real, but the appendix's own CPSNR numbers show it is not actually enforcing measurement consistency, so the inverse-problem claim needs major qualification. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the learned latent degradation operator $H_{\theta}$, a small neural network (a Readout Guidance network in the main experiments, a plain CNN in an ablation) trained with loss (20) to map denoised latents $\hat{z}_t^0$ to the encoding of the degraded measurement. It carries the argument because it converts the image-space data-consistency term into a latent-space norm: minimizing $\|w - H_{\theta}(\hat{z}_t^0, t)\|$ is used as a proxy for the pixel-space likelihood, justified by near-perfect autoencoder reconstruction on degraded images and a Lipschitz bound on the decoder.
What would settle it
Measure the correlation between the latent guidance gradient $\nabla_z \|w - H_{\theta}(\hat{z}_t^0)\|$ and the pixel-space likelihood gradient $\nabla_z \|y - A(D(\hat{z}_t^0))\|$ across diffusion timesteps for a degradation such as phase retrieval or very large noise. If the two gradients point in largely unrelated directions, the proxy fails and SILO's reconstructions should degrade accordingly.
Extended reading notes
Core claim
The central claim is that the likelihood score of a latent diffusion posterior can be approximated by a learned latent operator. Given a measurement $y$, a clean latent $z_0$, and a trained $H_{\theta}$ with $H_{\theta}(E(x)) \approx E(A(x))$, the paper derives the guidance gradient $\nabla_{z_t} \ln p(y|z_t) \approx -\tfrac{\text{Const}}{\sigma_y^2} \nabla_{z_t} \|w - H_{\theta}(\hat{z}_t^0)\|^2$, where $w = E(y)$. The argument rests on two steps: the autoencoder is near-lossless on degraded images, so pixel consistency can be rewritten in latent space, and the decoder's Lipschitz continuity lets the latent residual bound the pixel residual. Training $H_{\theta}$ by Eq. (20) minimizes the $\ell^1$ distance between $H_{\theta}(\hat{z}_t^0, t)$ and $E(y)$ over timesteps, so the operator learns not just the degradation but also the denoiser's time-dependent effect on latents. In Algorithm 1 the encoder and decoder are each invoked exactly once, with all gradient steps passing through the denoiser and $H_{\theta}$.
Load-bearing premise
The method assumes that the latent-space residual $\|E(y) - H_{\theta}(z)\|$ is a faithful proxy for the true pixel-space mismatch $\|y - A(x)\|$, an assumption supported only empirically by the autoencoder's near-reconstruction on degraded images and a Lipschitz bound with an unknown constant.
Editorial extensions
If this is right
- The autoencoder is used only twice per restoration: once to encode the measurement and once to decode the final latent; every intermediate step happens in latent space.
- On FFHQ and COCO, SILO reports lower LPIPS, FID, and KID than LDPS, GML-DPS, PSLD, and ReSample for blur, super-resolution, inpainting, and JPEG tasks.
- Restoration runtime drops by roughly 3 times versus PSLD and about 10 times versus ReSample in the reported settings.
- A separate $H_{\theta}$ must be trained for each degradation operator, but that one training run supports unlimited restorations for that operator.
- The method benefits from better text conditioning and classifier-free guidance, so perceptual quality improves when the latent diffusion prior is stronger.
Reading between the lines
- The paper leaves implicit that the same latent-operator recipe should transfer to other nonlinear degradations, such as learned camera pipelines or MRI undersampling, as long as a latent surrogate can be trained.
- A natural extension not explored in the paper is to train $H_{\theta}$ to match gradients rather than just outputs, which could tighten the latent-space proxy and improve consistency where the Lipschitz bound is loose.
- Because $E(y)$ must be a meaningful representation, SILO's success for a given degradation is tied to the autoencoder's behavior on degraded images; alternative encoders or learned measurement-to-latent maps could extend it to phase retrieval and other cases where $y$ is far from natural images.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SILO, a latent-space inverse problem solver for latent diffusion models. Instead of repeatedly decoding latents and differentiating through the decoder to enforce measurement consistency, SILO learns a small network Hθ that emulates the degradation operator A directly in the latent space. The measurement is encoded once, and the diffusion sampling is guided by minimizing ||E(y) - Hθ(ẑ_t0, t)||. The paper reports experiments on FFHQ and COCO with Gaussian blur, super-resolution ×4/×8, box inpainting, and JPEG, claiming improvements in LPIPS, FID, and KID and 3–10x speedups over PSLD, LDPS, GML-DPS, and ReSample.
Significance. If the consistency gap were resolved, the idea of learning a latent operator would be a valuable contribution to LDM-based inverse problems, since it removes the costly and artifact-prone differentiation through the autoencoder. The paper provides a systematic comparison, ablations (t-dependence, CNN vs. RG), and additional diversity results in the appendix. However, the central claim that SILO solves the inverse problem by posterior sampling is undermined by the measurement-consistency results reported only in the appendix.
major comments (3)
- [Sec. 5.3 and Appendix F (Tables 13–15)] The appendix reports CPSNR values that directly contradict the claim that the guidance in Eq. (19) enforces consistency with the measurement. For FFHQ super-resolution ×8, SILO(RV) reaches CPSNR 32.60 dB while PSLD and LDPS reach 40.91 and 38.73 dB; for Gaussian blur the gap is 32.17 dB vs. 44.18 dB and 42.69 dB. Since CPSNR = PSNR(A(x), A(ẋ̂)), a gap of 8–12 dB means the reconstructions, after re-applying the known degradation, do not match the measurements that define the inverse problem. The main tables (Tables 2 and 3) omit CPSNR, so the reported perceptual gains are not accompanied by evidence that the method samples p(x|y). The authors should include CPSNR in the main tables and either modify the method to enforce measurement consistency or substantially temper the claim that SILO is a posterior sampler.
- [Sec. 4.2, Eqs. (16)–(19)] The derivation of Eq. (19) is a proxy that is not sufficient for the guidance step. Eq. (18) states ||D(E(y*))−D(Hθ(z))||² ≤ C ||E(y*)−Hθ(z)||², but Algorithm 1 uses the gradient of the latent residual, not of the pixel residual. The Lipschitz constant C is unknown, and the gradient of the left-hand side is not controlled by the gradient of the right-hand side without additional assumptions on the Jacobian of D. Moreover, the approximation D(E(y))≈y in Eq. (16) is poor for inpainting and JPEG (PSNR ≈ 31 dB in Table 1). The paper should either prove a gradient bound or empirically verify that minimizing the latent residual reduces the pixel-space residual during the sampling trajectory; the current CPSNR results suggest it does not.
- [Sec. 4.4, Algorithm 1, step 9, and Eq. (19)] There is an inconsistency between the formula in Eq. (19), which uses the squared norm ||w−Hθ(ẑ_t0)||², and the algorithm, which computes the gradient of the unsquared norm ||w−ŵ_t||₂ (as described in the text: 'taking a gradient of the square root of the RHS in Eq. (18)'). The gradient of the norm differs from the gradient of the squared norm by a factor of 1/(2||·||), which changes the effective step size and the behavior of the guidance. The authors should clarify which objective is actually minimized and align Eq. (19), the text, and Algorithm 1.
minor comments (6)
- [Sec. 5.1] There are typos in the manuscript, for example 'groundn-truth' in Sec. 5.1 and 'degredations' in the caption of Fig. 4; these should be corrected.
- [Sec. 5.2] The network name 'Readout-Guidence' is likely a typo for 'Readout Guidance' (reference [35]); please fix.
- [Sec. 4.1 and Algorithm 1] The clamping operation w = clamp(E(y),−4,4) is introduced without explanation; a brief justification of the range would improve reproducibility.
- [Table 1] The column headers in Table 1 (e.g., 'x, f(x)', 'ynl, f(ynl)') are not self-explanatory; the caption should define these pairs and describe how PSNR is computed for each.
- [Appendix D and Sec. 5.3] The ReSample comparison uses an adapted version of the code with a different prior and image size, and the paper notes discrepancies with the originally reported ReSample numbers; this caveat should also appear in the main text near Tables 2 and 3.
- [Sec. 5.1] No seed variation or error bars are reported; since the algorithm is stochastic and the seed is fixed at 1000, it would strengthen the paper to run multiple seeds and report means and standard deviations for the key metrics.
Circularity Check
No significant circularity: the latent operator is trained on known degradations and evaluated on held-out data, and the derivation rests on an explicit Lipschitz surrogate rather than on its own conclusion.
full rationale
The paper's derivation chain is not circular under the quoted-equation test. Hθ is trained by Eq. (20), L = E[||Hθ(ẑ_t0, t) − E(y)||_1], on training images with known degradations, and its only use at inference time is as the guidance term in Algorithm 1, step 9. All reported quality metrics are computed on held-out FFHQ and COCO test images that were not used to train Hθ or the diffusion prior, so the central empirical claim is not statistically forced by construction. The key theoretical step, Eq. (18), is a Lipschitz bound on the decoder showing that the latent residual controls the pixel residual up to an unknown constant C; it does not assume the posterior-sampling conclusion, and the paper explicitly flags the D(E(y)) ≈ y approximation and leaves C unspecified. The self-citations in the background sections are standard references to prior work by the same group and are not load-bearing; the Hθ architecture is imported from external work [35] and is independently ablated with a simple CNN in Table 7. The reviewer-identified weakness, namely large CPSNR gaps in the appendix tables, is an empirical correctness concern about whether the learned operator faithfully enforces measurement consistency; it is not a circularity in the derivation. Therefore no specific circular step can be exhibited, and the honest finding is no significant circularity.
Assumptions & free parameters
free parameters (2)
- consistency scale eta =
0.5 (most tasks); 1.0 (inpainting)
- classifier-free guidance scale CFG =
1 or 4
assumptions (5)
- standard math The decoder D is Lipschitz continuous
- domain assumption D(E(y)) approximately equals y for degraded images
- domain assumption Htheta trained by Eq. (20) approximates E(A(x)) for test latents
- ad hoc to paper Gradient of ||w - Htheta(zhat_t0)|| is a valid proxy for the score likelihood grad_z ln p(y|z)
- domain assumption Stable Diffusion v1.5 and Realistic Vision v5.1 provide a good latent prior
Cite this review
Pith. "Pith review of SILO: Solving Inverse Problems with Latent Operators." pith.science (2026). https://pith.science/paper/ZO3AZ7GP
@misc{pith2026250111746,
author = {Pith},
title = {Pith review of: SILO: Solving Inverse Problems with Latent Operators},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZO3AZ7GP}},
note = {Machine review of arXiv:2501.11746}
}
read the original abstract
Consistent improvement of image priors over the years has led to the development of better inverse problem solvers. Diffusion models are the newcomers to this arena, posing the strongest known prior to date. Recently, such models operating in a latent space have become increasingly predominant due to their efficiency. In recent works, these models have been applied to solve inverse problems. Working in the latent space typically requires multiple applications of an Autoencoder during the restoration process, which leads to both computational and restoration quality challenges. In this work, we propose a new approach for handling inverse problems with latent diffusion models, where a learned degradation function operates within the latent space, emulating a known image space degradation. Usage of the learned operator reduces the dependency on the Autoencoder to only the initial and final steps of the restoration process, facilitating faster sampling and superior restoration quality. We demonstrate the effectiveness of our method on a variety of image restoration tasks and datasets, achieving significant improvements over prior art.
Figures
Figures from the paper (13 more)
Forward citations
Cited by 2 Pith papers
-
Compressed Image Generation with Denoising Diffusion Codebook Models
Using fixed codebooks of noise vectors in diffusion sampling yields images that carry their own compressed bit-streams and enables a strong perceptual image codec.
-
Inverse Problem Sampling in Latent Space Using Sequential Monte Carlo
LD-SMC uses sequential Monte Carlo in latent diffusion space with auxiliary per-timestep observations to improve posterior sampling for inverse problems, showing strong gains on inpainting.
Reference graph
Works this paper leans on
- [1]
-
[2]
Improving Image Genera- tion with Better Captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, Wesam Manassra, Prafulla Dhariwal, Casey Chu, Yunxin Jiao, and Aditya Ramesh. Improving Image Genera- tion with Better Captions. 1
-
[3]
Sutherland, Michael Arbel, and Arthur Gretton
Mikołaj Bi ´nkowski, Danica J. Sutherland, Michael Arbel, and Arthur Gretton. Demystifying MMD GANs. In Inter- national Conference on Learning Representations, 2018. 6, 3
work page 2018
-
[4]
The Perception-Distortion Tradeoff
Yochai Blau and Tomer Michaeli. The Perception-Distortion Tradeoff. In Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition , pages 6228–6237,
-
[5]
Stanley H. Chan, Xiran Wang, and Omar A. Elgendy. Plug- and-Play ADMM for Image Restoration: Fixed-Point Con- vergence and Applications. IEEE Transactions on Computa- tional Imaging, 3(1):84–98, 2017. 3
work page 2017
-
[6]
Yunjin Chen and Thomas Pock. Trainable Nonlinear Reac- tion Diffusion: A Flexible Framework for Fast and Effective Image Restoration. IEEE Trans. Pattern Anal. Mach. Intell., 39(6):1256–1272, 2017. 1, 3
work page 2017
-
[7]
Diffusion Pos- terior Sampling for General Noisy Inverse Problems
Hyungjin Chung, Jeongsol Kim, Michael Thompson Mc- cann, Marc Louis Klasky, and Jong Chul Ye. Diffusion Pos- terior Sampling for General Noisy Inverse Problems. In The Eleventh International Conference on Learning Representa- tions, 2022. 1, 3, 5
work page 2022
-
[8]
Decom- posed Diffusion Sampler for Accelerating Large-Scale In- verse Problems
Hyungjin Chung, Suhyeon Lee, and Jong Chul Ye. Decom- posed Diffusion Sampler for Accelerating Large-Scale In- verse Problems. In The Twelfth International Conference on Learning Representations, 2023. 3
work page 2023
Show all 67 references
-
[9]
Prompt-tuning latent diffusion models for inverse problems, 2023
Hyungjin Chung, Jong Chul Ye, Peyman Milanfar, and Mauricio Delbracio. Prompt-tuning latent diffusion models for inverse problems, 2023. 2, 4, 6
2023
-
[10]
Image Denoising by Sparse 3-D Transform-Domain Collaborative Filtering
Kostadin Dabov, Alessandro Foi, Vladimir Katkovnik, and Karen Egiazarian. Image Denoising by Sparse 3-D Transform-Domain Collaborative Filtering. IEEE Transac- tions on Image Processing, 16(8):2080–2095, 2007. 1
2007
-
[11]
Diffusion Models Beat GANs on Image Synthesis
Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion Models Beat GANs on Image Synthesis. In Advances in Neural Information Processing Systems, 2021. 1, 2
2021
-
[12]
Tweedie’s Formula and Selection Bias
Bradley Efron. Tweedie’s Formula and Selection Bias. Jour- nal of the American Statistical Association, 106(496):1602– 1614, 2011. 3
2011
-
[13]
Image Denoising Via Sparse and Redundant Representations Over Learned Dic- tionaries
Michael Elad and Michal Aharon. Image Denoising Via Sparse and Redundant Representations Over Learned Dic- tionaries. IEEE Transactions on Image Processing, 15(12): 3736–3745, 2006. 1
2006
-
[14]
Image Denoising: The Deep Learning Revolution and Beyond—A Survey Paper
Michael Elad, Bahjat Kawar, and Gregory Vaksman. Image Denoising: The Deep Learning Revolution and Beyond—A Survey Paper. SIAM Journal on Imaging Sciences , 16(3): 1594–1654, 2023. 1, 3
2023
-
[15]
Adaptive Compressed Sensing with Diffusion-Based Posterior Sam- pling
Noam Elata, Tomer Michaeli, and Michael Elad. Adaptive Compressed Sensing with Diffusion-Based Posterior Sam- pling. In Computer Vision – ECCV 2024 , pages 290–308, Cham, 2025. Springer Nature Switzerland. 1
2024
-
[16]
Sparsity based Poisson de- noising
Raja Giryes and Michael Elad. Sparsity based Poisson de- noising. In 2012 IEEE 27th Convention of Electrical and Electronics Engineers in Israel, pages 1–5, 2012. 3
2012
-
[17]
Weighted Nuclear Norm Minimization with Applica- tion to Image Denoising
Shuhang Gu, Lei Zhang, Wangmeng Zuo, and Xiangchu Feng. Weighted Nuclear Norm Minimization with Applica- tion to Image Denoising. In 2014 IEEE Conference on Com- puter Vision and Pattern Recognition , pages 2862–2869,
2014
-
[18]
Alarc´on
Javier Gurrola-Ramos, Oscar Dalmau, and Teresa E. Alarc´on. A Residual Dense U-Net Neural Network for Im- age Denoising. IEEE Access, 9:31742–31754, 2021. 1
2021
-
[19]
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In Advances in Neural Information Processing Systems. Curran Associates, Inc., 2017. 6, 3
2017
-
[20]
Classifier-Free Diffusion Guidance
Jonathan Ho and Tim Salimans. Classifier-Free Diffusion Guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021. 8, 3
2021
-
[21]
Denoising Dif- fusion Probabilistic Models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising Dif- fusion Probabilistic Models. In Advances in Neural Infor- mation Processing Systems, pages 6840–6851. Curran Asso- ciates, Inc., 2020. 1, 3
2020
-
[22]
Fu Jie Huang and Y . LeCun. Large-scale Learning with SVM and Convolutional for Generic Object Categorization. In 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06), pages 284–291,
2006
-
[23]
What is the best multi-stage architecture for object recognition? In 2009 IEEE 12th International Conference on Computer Vision, pages 2146–2153, 2009
Kevin Jarrett, Koray Kavukcuoglu, Marc’Aurelio Ranzato, and Yann LeCun. What is the best multi-stage architecture for object recognition? In 2009 IEEE 12th International Conference on Computer Vision, pages 2146–2153, 2009. 8
2009
-
[24]
Towards Flex- ible Blind JPEG Artifacts Removal
Jiaxi Jiang, Kai Zhang, and Radu Timofte. Towards Flex- ible Blind JPEG Artifacts Removal. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 4997–5006, 2021. 3
2021
-
[25]
A Style- Based Generator Architecture for Generative Adversarial Networks
Tero Karras, Samuli Laine, and Timo Aila. A Style- Based Generator Architecture for Generative Adversarial Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 4401– 4410, 2019. 6
2019
-
[26]
Denoising Diffusion Restoration Models
Bahjat Kawar, Michael Elad, Stefano Ermon, and Jiaming Song. Denoising Diffusion Restoration Models. Advances in Neural Information Processing Systems, 35:23593–23606,
-
[27]
Regularization by Texts for Latent Diffusion Inverse Solvers, 2024
Jeongsol Kim, Geon Yeong Park, Hyungjin Chung, and Jong Chul Ye. Regularization by Texts for Latent Diffusion Inverse Solvers, 2024. 2, 4
2024
-
[28]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A Method for Stochastic Optimization, 2017. 5
2017
-
[29]
Kingma and Max Welling
Diederik P. Kingma and Max Welling. Auto-Encoding Vari- ational Bayes, 2022. 2, 6 9
2022
-
[30]
Im- ageNet Classification with Deep Convolutional Neural Net- works
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Im- ageNet Classification with Deep Convolutional Neural Net- works. In Advances in Neural Information Processing Sys- tems. Curran Associates, Inc., 2012. 6, 8, 2
2012
-
[31]
LSDIR: A Large Scale Dataset for Image Restoration
Yawei Li, Kai Zhang, Jingyun Liang, Jiezhang Cao, Ce Liu, Rui Gong, Yulun Zhang, Hao Tang, Yun Liu, Denis De- mandolx, Rakesh Ranjan, Radu Timofte, and Luc Van Gool. LSDIR: A Large Scale Dataset for Image Restoration. In 2023 IEEE/CVF Conference on Computer Vision and Pat- ter...
2023
-
[32]
SwinIR: Image Restoration Using Swin Transformer
Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. SwinIR: Image Restoration Using Swin Transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1833– 1844, 2021. 3
2021
-
[33]
Lawrence Zitnick
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C. Lawrence Zitnick. Microsoft COCO: Common Objects in Context. In Computer Vision – ECCV 2014 , pages 740–755, Cham,
2014
-
[34]
Decoupled Weight De- cay Regularization
Ilya Loshchilov and Frank Hutter. Decoupled Weight De- cay Regularization. In International Conference on Learning Representations, 2018. 4
2018
-
[35]
Gold- man, and Aleksander Holynski
Grace Luo, Trevor Darrell, Oliver Wang, Dan B. Gold- man, and Aleksander Holynski. Readout Guidance: Learn- ing Control from Diffusion Features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8217–8227, 2024. 6, 4
2024
-
[36]
High-Perceptual Quality JPEG Decoding via Posterior Sam- pling
Sean Man, Guy Ohayon, Theo Adrai, and Michael Elad. High-Perceptual Quality JPEG Decoding via Posterior Sam- pling. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 1272–1282,
-
[37]
A Variational Perspective on Solving Inverse Problems with Diffusion Models
Morteza Mardani, Jiaming Song, Jan Kautz, and Arash Vah- dat. A Variational Perspective on Solving Inverse Problems with Diffusion Models. In The Twelfth International Confer- ence on Learning Representations, 2023. 3
2023
-
[38]
High Perceptual Quality Image De- noising With a Posterior Sampling CGAN
Guy Ohayon, Theo Adrai, Gregory Vaksman, Michael Elad, and Peyman Milanfar. High Perceptual Quality Image De- noising With a Posterior Sampling CGAN. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 1805–1813, 2021. 1
2021
-
[39]
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning Transferable Visual Models From Natural Language Supervision. InProceedings of the...
2021
-
[40]
Inverse Methods for Atmospheric Sound- ing: Theory and Practice
Clive D Rodgers. Inverse Methods for Atmospheric Sound- ing: Theory and Practice. In Inverse Methods for Atmo- spheric Sounding: Theory and Practice . WORLD SCIEN- TIFIC, 2000. 1
2000
-
[41]
The Little Engine That Could: Regularization by Denoising (RED)
Yaniv Romano, Michael Elad, and Peyman Milanfar. The Little Engine That Could: Regularization by Denoising (RED). SIAM Journal on Imaging Sciences , 10(4):1804– 1844, 2017. 3
2017
-
[42]
High-Resolution Image Synthesis With Latent Diffusion Models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-Resolution Image Synthesis With Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022. 2, 3, 6
2022
-
[43]
Dimakis, and Sanjay Shakkottai
Litu Rout, Negin Raoof, Giannis Daras, Constantine Cara- manis, Alexandros G. Dimakis, and Sanjay Shakkottai. Solv- ing Linear Inverse Problems Provably via Posterior Sampling with Latent Diffusion Models, 2023. 2, 4, 6, 1
2023
-
[44]
Beyond First-Order Tweedie: Solving Inverse Problems using La- tent Diffusion
Litu Rout, Yujia Chen, Abhishek Kumar, Constantine Cara- manis, Sanjay Shakkottai, and Wen-Sheng Chu. Beyond First-Order Tweedie: Solving Inverse Problems using La- tent Diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 947...
2024
-
[45]
Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Pho- torealistic Text-to-Image Diffusion Models with Deep Lan- gu...
2022
-
[46]
Stablediffusionapi/realistic- vision-51 · Hugging Face
Evgeny SG161222. Stablediffusionapi/realistic- vision-51 · Hugging Face. https://huggingface.co/stablediffusionapi/realistic-vision-
-
[47]
Very Deep Convo- lutional Networks for Large-Scale Image Recognition, 2015
Karen Simonyan and Andrew Zisserman. Very Deep Convo- lutional Networks for Large-Scale Image Recognition, 2015. 6, 2
2015
-
[48]
Deep Unsupervised Learning using Nonequilibrium Thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep Unsupervised Learning using Nonequilibrium Thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning, pages 2256–2265. PMLR, 2015. 1
2015
-
[49]
Solving Inverse Problems with Latent Diffusion Models via Hard Data Consistency
Bowen Song, Soo Min Kwon, Zecheng Zhang, Xinyu Hu, Qing Qu, and Liyue Shen. Solving Inverse Problems with Latent Diffusion Models via Hard Data Consistency. In The Twelfth International Conference on Learning Representa- tions, 2023. 2, 4, 6, 1
2023
-
[50]
Generative Modeling by Es- timating Gradients of the Data Distribution
Yang Song and Stefano Ermon. Generative Modeling by Es- timating Gradients of the Data Distribution. In Advances in Neural Information Processing Systems. Curran Associates, Inc., 2019. 3
2019
-
[51]
Generative Modeling by Estimating Gradients of the Data Distribution
Yang Song and Stefano Ermon. Generative Modeling by Estimating Gradients of the Data Distribution. Advances in Neural Information Processing Systems, 32, 2019. 1
2019
-
[52]
Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-Based Generative Modeling through Stochastic Differential Equa- tions. In International Conference on Learning Representa- tions, 2020. 1, 3
2020
-
[53]
Solv- ing Inverse Problems in Medical Imaging with Score-Based Generative Models
Yang Song, Liyue Shen, Lei Xing, and Stefano Ermon. Solv- ing Inverse Problems in Medical Imaging with Score-Based Generative Models. In International Conference on Learn- ing Representations, 2021. 1
2021
-
[54]
Venkatakrishnan, Charles A
Singanallur V . Venkatakrishnan, Charles A. Bouman, and Brendt Wohlberg. Plug-and-Play priors for model based re- 10 construction. In 2013 IEEE Global Conference on Signal and Information Processing, pages 945–948, 2013. 3
2013
-
[55]
A Connection Between Score Matching and Denoising Autoencoders
Pascal Vincent. A Connection Between Score Matching and Denoising Autoencoders. Neural Computation, 23(7):1661– 1674, 2011. 3
2011
-
[56]
Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution, 2024
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Jun- yang Lin. Qwen2-VL: Enhancing Vision-Language Model’s ...
2024
-
[57]
ESRGAN: En- hanced Super-Resolution Generative Adversarial Networks
Xintao Wang, Ke Yu, Shixiang Wu, Jinjin Gu, Yihao Liu, Chao Dong, Yu Qiao, and Chen Change Loy. ESRGAN: En- hanced Super-Resolution Generative Adversarial Networks. In Proceedings of the European Conference on Computer Vi- sion (ECCV) Workshops, pages 0–0, 2018. 3
2018
-
[58]
Zero-Shot Image Restoration Using Denoising Diffusion Null-Space Model
Yinhuai Wang, Jiwen Yu, and Jian Zhang. Zero-Shot Image Restoration Using Denoising Diffusion Null-Space Model. In The Eleventh International Conference on Learning Rep- resentations, 2022. 1
2022
-
[59]
RestoreFormer: High-Quality Blind Face Restoration From Undegraded Key-Value Pairs
Zhouxia Wang, Jiawei Zhang, Runjian Chen, Wenping Wang, and Ping Luo. RestoreFormer: High-Quality Blind Face Restoration From Undegraded Key-Value Pairs. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17512–17521, 2022. 3
2022
-
[60]
On Single Image Scale-Up Using Sparse-Representations
Roman Zeyde, Michael Elad, and Matan Protter. On Single Image Scale-Up Using Sparse-Representations. In Curves and Surfaces , pages 711–730, Berlin, Heidelberg, 2012. Springer. 3
2012
-
[61]
Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising
Kai Zhang, Wangmeng Zuo, Yunjin Chen, Deyu Meng, and Lei Zhang. Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising. IEEE Transactions on Image Processing, 26(7):3142–3155, 2017. 3
2017
-
[62]
Efros, Eli Shecht- man, and Oliver Wang
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 586–595, 2018. 6
2018
-
[63]
PerceptualSimilarity, 2018 https://github.com/richzhang/PerceptualSimilarity
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. PerceptualSimilarity, 2018 https://github.com/richzhang/PerceptualSimilarity. 6, 2
2018
-
[64]
Denoising Dif- fusion Models for Plug-and-Play Image Restoration
Yuanzhi Zhu, Kai Zhang, Jingyun Liang, Jiezhang Cao, Bi- han Wen, Radu Timofte, and Luc Van Gool. Denoising Dif- fusion Models for Plug-and-Play Image Restoration. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1219–1229, 2023. 3 ...
2023
-
[66]
bi-domain
(22) We categorize such solutions as part of the “bi-domain” family, as they compute gradients in both the pixel and la- tent domains. To the best of our knowledge, all existing methods leveraging LDMs for inverse problems (apart from SILO) fall within this category. In this s...
-
[256]
Hence, to give a fair comparison to ReSample, we had to adapt their publicly available code. PSLD. We made no modifications to the PSLD code, ex- cept for adapting the data-loading process to enable sam- pling from the COCO dataset. The hyperparameters used were identical to t...
-
[2014]
Springer International Publishing. 6
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.