Pith. sign in

REVIEW 3 major objections 5 minor 64 references

Unpaired Joint Distribution Modeling via Multi-Scale Image Representations

T0 review · 3 major / 5 minor · reviewed 2026-07-10 · grok-4.5

Pith's one-line read A multi-scale image map lets unpaired clean and noisy images share a joint distribution well enough to train strong denoisers, including for cryo-EM.

desk verdict Solid methods paper with a real error decomposition and strong cryo-EM numbers; the main theoretical claim rests on an unchecked tail-equivalence assumption, but the work still deserves a referee. read the letter →

arxiv 2607.08198 v1 pith:Y4SSEDMJ submitted 2026-07-09 cs.CV

classification cs.CV
keywords unpairedjointdistributionmodelingmulti-scaleimagerepresentationevidencelowerboundnoisereal-worlddenoisingcryo-EMprobabilisticgraphicalmodeldomainconsistency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Learning a joint distribution from unpaired marginals is ill-posed: many couplings can match the same clean and noisy distributions. The authors introduce LUD-MSR, a latent-variable graphical model whose training losses are evidence lower bounds that use only unpaired samples. They prove that the KL gap to the true joint is controlled by reconstruction quality of the two domains plus how closely the auxiliary representations of a true pair agree under the model’s inference map. That bound exposes a concrete trade-off: the auxiliaries must be nearly identical for true pairs (domain consistency) while still retaining enough of the original image (information preservation). Multi-Scale image Representation (MSR) maps—built from invertible wavelet-flow layers that keep only the coarsest coefficients—achieve a better balance of this trade-off than noise injection or linear projections. Clean images pushed through the learned clean-to-noisy pipeline produce synthetic pairs that train denoisers competitive with fully supervised noise models and deliver large SNR gains on three real cryo-EM particle datasets.

What carries the argument

LUD-MSR: a hierarchical latent graphical model with frozen multi-scale auxiliaries hx, hy together with the MSR hypothesis class of invertible wavelet-flow maps that retain only the coarsest coefficients; Theorem 1 converts the resulting ELBO losses into an explicit upper bound on joint approximation error.

What would settle it

On a real paired denoising set, measure whether the empirical densities p(y|x) and p(y|h(y)) (and the symmetric pair) differ by more than a constant factor e^η whenever either density falls below a fixed ε0; if the ratio routinely exceeds that factor, Theorem 1’s bound is inapplicable.

Watch

Extended reading notes

Core claim

Under a mild tail-equivalence assumption, the sum of KL divergences between the true joint and the two generative joints of LUD-MSR is upper-bounded by the negative expected ELBOs of the auxiliary-conditioned likelihoods, the square-root KL distances between the inference distributions of each domain and its auxiliary, and the expected L1 distance between the two auxiliary inference maps; Multi-Scale image Representation mappings minimize that last term while losing far less information than previous auxiliary constructions, yielding higher-fidelity unpaired joint models.

Load-bearing premise

The generative conditionals given a true pair and given its multi-scale auxiliaries must stay within a fixed multiplicative factor of each other in the low-probability tails; if that regularity fails, the KL error bound no longer holds.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes LUD-MSR, a two-stage latent-variable framework for learning a joint distribution from unpaired marginals. Auxiliary representations h_x, h_y are first learned (via Multi-Scale image Representation maps for images), then a hierarchical probabilistic graphical model with shared content z and domain-specific z_n is trained by maximizing ELBOs of log p(x|h_x) and log p(y|h_y) that use only unpaired samples. Theorem 1 bounds the joint KL gap by reconstruction ELBOs plus an inference-invariance term under a tail-equivalence assumption; Theorems 2–3 compare hypothesis classes and claim that the nonlinear MSR class achieves a superior consistency–preservation trade-off (via an intrinsic intersection rank r*). Clean-to-noisy samples are used to train denoisers, with strong unpaired/semi-supervised results on SIDD-family benchmarks and large estimated-SNR gains on three cryo-EM datasets.

Significance. If the analysis holds, the work supplies a rare explicit error decomposition for unpaired joint modeling and a concrete inductive bias (MSR) that improves the consistency–preservation trade-off relative to noise injection and orthogonal projections. The hierarchical architecture, ELBO derivations, and full proofs of Theorems 1–3 / Propositions 1–3 are written out carefully; the empirical package is broad (unpaired and semi-supervised noise generation, multiple real denoising benchmarks, cryo-EM SNR, and ablations on T and wavelet choice). The cryo-EM application is a genuine high-impact use case where paired data are scarce. These strengths make the contribution of clear interest to the unpaired translation and scientific imaging communities, provided the load-bearing regularity is better supported.

major comments (3)
  1. Theorem 1 (and its proof steps (i)–(iii), eqs. 5–11) converts density differences |p(y|x)−p(y|h_y)| into log-likelihood gaps only by invoking Assumption 1(i) tail-equivalence (bounded ratio e^η whenever either density is below ε0). The constants ε0, η are never estimated, and no diagnostic (e.g., histogram or scatter of log p(y|x)/p(y|h_y) on held-out true pairs from SIDD or cryo-EM) is provided for the learned hierarchical conditionals. Without this check the claimed upper bound on joint approximation error remains conditional; a short verification or a discussion of when the assumption can fail would make the central guarantee load-bearing rather than formal.
  2. Theorems 2–3 analyze the consistency–preservation trade-off under linear-Gaussian inference q(z|h)=N(Ah,Γ) with full-row-rank A. The actual model (Appendix A) is a deep hierarchical VAE with layer-wise Gaussians and residual dense blocks. The paper does not bridge this gap: it is unclear whether the ranking e*_HNI ≳ K²ε⁻² vs. e*_HOP ≲ K−ε² vs. e*_HMSR ≲ r*−ε² continues to order the practical hypothesis classes once the inference model is nonlinear and hierarchical. A short remark or a controlled numerical check of the L1 inference-invariance term under the true architecture would strengthen the claim that MSR “significantly reduces the distribution approximation error.”
  3. Section 5.3 constructs the “clean” cryo-EM domain from homologous PDB structures (6HVR, 7ZJW, 8UTJ) projected and CTF-modulated to match real micrographs. The quantitative claim of large SNR gains (Table 3) therefore rests on how well these homologues share low-frequency content with the true particles. The manuscript does not quantify residual structural mismatch or report sensitivity to the choice of reference; without that, it is hard to separate genuine noise-modeling gains from domain-shift artifacts in the pseudo-clean set.
minor comments (5)
  1. Notation inconsistency: the text sometimes writes LUD-MSD (end of §3.1) while the title, abstract and algorithms use LUD-MSR.
  2. Figure 1 caption and Algorithm 1 refer to Stage I / Stage II; a one-sentence reminder that h is frozen after Stage I would help readers who skip the algorithms.
  3. In Proposition 2 the low-pass filter is fixed to (1/4,1/2,1/4); a brief note that the eigenvalue bounds continue to hold (up to constants) for other tight-frame low-pass filters would clarify generality.
  4. Table 1 reports “–” for DeFlow AKLD; either compute the missing entries or state explicitly that the metric is unavailable for that baseline.
  5. Appendix E lists M=8 Glow steps and the linear B-spline filters; adding the precise channel dimensions of the affine coupling networks would improve reproducibility.

Circularity Check

0 steps flagged · score 1.0 of 10

No load-bearing circularity: Theorem 1 bounds and Theorems 2–3 trade-off comparisons are derived from explicit assumptions and first-principles calculations on hypothesis classes; self-citations supply architecture lineage only.

full rationale

The derivation chain is self-contained. Theorem 1 obtains an upper bound on the joint KL gaps by applying the tail-equivalence of Assumption 1(i), triangle inequality on the L1 posterior differences, Pinsker, and the variational ELBOs (12)–(13); the subsequent unpaired loss simply minimizes the controllable terms of that bound and does not redefine them. Propositions 1–3 and Theorems 2–3 then compare three concrete hypothesis classes (noise-injection, orthogonal projections, MSR diffeomorphisms) under linear-Gaussian inference models, using the closed-form L1 distance of Gaussians, eigenvalue bounds on multi-scale filters, and the intrinsic intersection rank r*; none of these equalities reduce to a fitted identity or to a prior result of the same authors. The only self-citations (LUD-VAE, SeNM-VAE, the multi-scale flow preprint [3]) motivate the graphical-model skeleton and the concrete wavelet+flow realization of H_MSR; they are not invoked as uniqueness theorems or as the sole justification of any claimed bound. Empirical choices (α=0.5, T=3) are free hyperparameters validated by ablation, not predictions forced by construction. The paper therefore contains only ordinary architectural lineage, not circular reasoning.

Assumptions & free parameters 5 free parameters · 5 assumptions · 3 invented entities

The load-bearing content is a variational graphical model plus a designed multi-scale map. Theory rests on a tail-equivalence regularity assumption, Gaussian hierarchical latents, and the domain-consistency / information-preservation design principles. Several scalar and architectural knobs (α, T, σ schedule, flow depth) are chosen by hand. Invented objects are the LUD-MSR pipeline, the MSR hypothesis class, and the intrinsic intersection rank used to state the nonlinear trade-off.

free parameters (5)
  • mixture weight α in inference models (2),(25)
    Empirically fixed to 0.5; controls how much posterior mass comes from raw observations versus auxiliaries and appears in ELBOs and paired losses.
  • MSR scale depth T (and d=K/2^T)
    Default T=3 chosen by ablation; directly sets the consistency–preservation operating point of h.
  • Gaussian noise level σ in unpaired auxiliary loss (18)
    Sampled from hand-chosen ranges ([0,40/255] for SIDD; [0,0.99] for cryo-EM) to approximate domain consistency without pairs.
  • Hierarchical VAE depth L and flow blocks M
    L=7 latent layers and M=8 Glow steps per scale are architectural free choices that affect capacity and reported metrics.
  • ε0, η in Assumption 1
    Existential constants in the tail-equivalence hypothesis that make Theorem 1 go through; not estimated from data.
assumptions (5)
  • ad hoc to paper Assumption 1: tail-equivalence of p(y|x) vs p(y|h_y) and p(x|y) vs p(x|h_x) on low-probability regions, plus essential boundedness of conditional densities given z.
    Stated before Theorem 1; used to convert density gaps into log-likelihood and KL bounds (eqs. 8–11).
  • domain assumption Inference and generative conditionals are parameterized as (hierarchical) Gaussians with shared networks for p(z|h)=q(z|h).
    Standard VAE modeling choice enabling closed-form KL terms; appears throughout §3.1 and Appendix A.
  • domain assumption For denoising, y=x+n with well-defined covariances Σ_X, Σ_n; natural-image energy concentrates in low frequencies so multi-scale low-pass structure is shared.
    Used to justify H and H_MSR in §3.2 and Proposition 2; cites classical natural-image statistics.
  • ad hoc to paper Linear Gaussian inference q(z|h)=N(Ah,Γ) with full-row-rank A for the trade-off analysis in Theorems 2–3.
    Simplifies L1 Gaussian distance to erf expressions; analysis may not transfer quantitatively to the nonlinear hierarchical nets used in experiments.
  • standard math Standard variational ELBO inequalities and Pinsker's inequality.
    Used in Theorem 1 proof and Appendix B derivations.
invented entities (3)
  • LUD-MSR probabilistic graphical model with frozen auxiliaries h_x,h_y and latents (z,z_n)
    purpose: Factor unpaired joint modeling into representation learning plus ELBO training on marginals only.
    Core framework of the paper; extends prior LUD-style graphs with multi-scale auxiliaries.
  • Multi-Scale image Representation class H_MSR (invertible multi-scale maps with truncated coordinates)
    purpose: Realize a favorable domain-consistency vs information-preservation trade-off for images.
    Defined in (20)–(21); compared theoretically to H_NI and H_OP.
  • Intrinsic intersection rank r* of transformed signal and noise images
    purpose: Characterize when MSR can achieve zero or low information loss under inference-invariance constraints (Theorem 3).
    Defined in (32); not measured on real datasets, only used as an analytic quantity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Unpaired Joint Distribution Modeling via Multi-Scale Image Representations." pith.science (2026). https://pith.science/paper/Y4SSEDMJ

@misc{pith2026260708198,
  author       = {Pith},
  title        = {Pith review of: Unpaired Joint Distribution Modeling via Multi-Scale Image Representations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Y4SSEDMJ}},
  note         = {Machine review of arXiv:2607.08198}
}
read the original abstract

This paper studies the problem of learning a joint distribution from marginal observations, which is inherently ill-posed due to the ambiguity of feasible couplings. We propose LUD-MSR, a latent-variable probabilistic framework that models the joint distribution via auxiliary representations and optimizes evidence lower bounds using only marginal data. Under mild assumptions, we establish an upper bound on the distribution approximation error. This analysis reveals a trade-off in representation learning between domain consistency and information preservation. To address this trade-off, we introduce a Multi-Scale image Representation (MSR) mapping that exploits structural similarity at coarse scales while suppressing domain-specific variations. We show that MSR achieves a more favorable balance of this trade-off compared to existing approaches. Experiments on real-world denoising benchmarks, including cryo-electron microscopy (cryo-EM), demonstrate the effectiveness of the proposed framework.

Figures

Figures reproduced from arXiv: 2607.08198 by the authors.

Figure 1
Figure 1. Overview of the LUD-MSR training pipeline. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Probabilistic graphical model of LUD-MSR. The generative processes for the [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Visual comparison of noisy images generated by various noise modeling methods on the SIDD [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Visual comparison of denoising results for DnCNN models trained on synthetic data from various [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: Visual comparison of denoising results across three real-world cryo-EM datasets: EMPIAR-10025, [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: Hierarchical architecture of the probabilistic graphical model of LUD-MSR. [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

64 extracted references · 64 canonical work pages

  1. [1]

    Ntire 2020 challenge on real image denoising: Dataset, methods and results

    Abdelrahman Abdelhamed, Mahmoud Afifi, Radu Timofte, and Michael S Brown. Ntire 2020 challenge on real image denoising: Dataset, methods and results. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2020

  2. [2]

    A high-quality denoising dataset for smartphone cameras

    Abdelrahman Abdelhamed, Stephen Lin, and Michael S Brown. A high-quality denoising dataset for smartphone cameras. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 1692–1700, 2018

  3. [3]

    Enhancing low-resolution image representation through normalizing flows.arXiv preprint arXiv:2601.06834, 2026

    Chenglong Bao, Tongyao Pang, Zuowei Shen, Dihan Zheng, and Yihang Zou. Enhancing low-resolution image representation through normalizing flows.arXiv preprint arXiv:2601.06834, 2026

  4. [4]

    Topaz-denoise: general deep denoising models for cryoem and cryoet.Nature communications, 11(1):5208, 2020

    Tristan Bepler, Kotaro Kelley, Alex J Noble, and Bonnie Berger. Topaz-denoise: general deep denoising models for cryoem and cryoet.Nature communications, 11(1):5208, 2020

  5. [5]

    Un- processing images for learned raw denoising

    Tim Brooks, Ben Mildenhall, Tianfan Xue, Jiawen Chen, Dillon Sharlet, and Jonathan T Barron. Un- processing images for learned raw denoising. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11036–11045, 2019

  6. [6]

    2.8 ˚ a reso- lution reconstruction of the thermoplasma acidophilum 20s proteasome using cryo-electron microscopy

    Melody G Campbell, David Veesler, Anchi Cheng, Clinton S Potter, and Bridget Carragher. 2.8 ˚ a reso- lution reconstruction of the thermoplasma acidophilum 20s proteasome using cryo-electron microscopy. Elife, 4:e06380, 2015

  7. [7]

    Exploring efficient asymmetric blind-spots for self-supervised denoising in real-world scenarios

    Shiyan Chen, Jiyuan Zhang, Zhaofei Yu, and Tiejun Huang. Exploring efficient asymmetric blind-spots for self-supervised denoising in real-world scenarios. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2814–2823. IEEE, 2024

  8. [8]

    Unpaired deep image dehazing using contrastive disentanglement learning

    Yang Chen et al. Unpaired deep image dehazing using contrastive disentanglement learning. In European Conference on Computer Vision. Springer, 2022

Show all 64 references
  1. [9]

    Very deep vaes generalize autoregressive models and can outperform them on images

    Rewon Child. Very deep vaes generalize autoregressive models and can outperform them on images. arXiv preprint arXiv:2011.10650, 2020

  2. [10]

    Image denoising by sparse 3-d transform-domain collaborative filtering.IEEE Transactions on image processing, 16(8):2080–2095, 2007

    Kostadin Dabov, Alessandro Foi, Vladimir Katkovnik, and Karen Egiazarian. Image denoising by sparse 3-d transform-domain collaborative filtering.IEEE Transactions on image processing, 16(8):2080–2095, 2007

  3. [11]

    Framelets: Mra-based wavelet frames and applications.Applied and Computational Harmonic Analysis, 14(1):1–46, 2003

    Ingrid Daubechies, Bin Han, Amos Ron, and Zuowei Shen. Framelets: Mra-based wavelet frames and applications.Applied and Computational Harmonic Analysis, 14(1):1–46, 2003. 21

  4. [12]

    Unsupervised visual representation learning by context prediction

    Carl Doersch, Abhinav Gupta, and Alexei A Efros. Unsupervised visual representation learning by context prediction. InProceedings of the IEEE International Conference on Computer Vision (ICCV), pages 1422–1430, 2015

  5. [13]

    Relations between the statistics of natural images and the response properties of cortical cells.Journal of the Optical Society of America A, 4(12):2379–2394, 1987

    David J Field. Relations between the statistics of natural images and the response properties of cortical cells.Journal of the Optical Society of America A, 4(12):2379–2394, 1987

  6. [14]

    Bock, Cristina Maracci, Zhe Wang, Alena Paleskava, Andrey L

    Niels Fischer, Piotr Neumann, Lars V. Bock, Cristina Maracci, Zhe Wang, Alena Paleskava, Andrey L. Konevega, Gunnar F. Schr¨ oder, Helmut Grubm¨ uller, Ralf Ficner, Marina V. Rodnina, and Holger Stark. The pathway to gtpase activation of elongation factor selb on the ribosome....

  7. [15]

    Srgb real noise synthesizing with neighboring correlation- aware noise model

    Zixuan Fu, Lanqing Guo, and Bihan Wen. Srgb real noise synthesizing with neighboring correlation- aware noise model. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1683–1691, 2023

  8. [16]

    Academic Press, 2nd edition, 1990

    Keinosuke Fukunaga.Introduction to statistical pattern recognition. Academic Press, 2nd edition, 1990

  9. [17]

    Toward convolutional blind denoising of real photographs

    Shi Guo, Zifei Yan, Kai Zhang, Wangmeng Zuo, and Lei Zhang. Toward convolutional blind denoising of real photographs. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1712–1722, 2019

  10. [18]

    A class of fast gaussian binomial filters for speech and image processing.IEEE Transactions on Signal Processing, 39(3):723–727, 1991

    Richard A Haddad, Ali N Akansu, et al. A class of fast gaussian binomial filters for speech and image processing.IEEE Transactions on Signal Processing, 39(3):723–727, 1991

  11. [19]

    Masked autoen- coders are scalable vision learners

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ ar, and Ross Girshick. Masked autoen- coders are scalable vision learners. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16000–16009, 2022

  12. [20]

    Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

  13. [21]

    Cycada: Cycle-consistent adversarial domain adaptation

    Judy Hoffman, Eric Tzeng, Taesung Park, Jun-Yan Zhu, Phillip Isola, Kate Saenko, Alexei Efros, and Trevor Darrell. Cycada: Cycle-consistent adversarial domain adaptation. InInternational conference on machine learning, pages 1989–1998. PMLR, 2018

  14. [22]

    Qs-attn: Query-selected attention for contrastive learning in i2i translation

    Xiangpeng Hu, Xin Zhou, Qisheng Huang, Zhiwen Shi, Lin Sun, and Qingqi Li. Qs-attn: Query-selected attention for contrastive learning in i2i translation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7188–7197, 2022

  15. [23]

    Neighbor2neighbor: Self- supervised denoising from single noisy images

    Tao Huang, Songjiang Li, Xu Jia, Huchuan Lu, and Jianzhuang Liu. Neighbor2neighbor: Self- supervised denoising from single noisy images. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), 2021

  16. [24]

    Multimodal unsupervised image-to-image translation

    Xun Huang, Ming-Yu Liu, Serge Belongie, and Jan Kautz. Multimodal unsupervised image-to-image translation. InProceedings of the European conference on computer vision (ECCV), pages 172–189, 2018

  17. [25]

    C2n: Practical generative noise modeling for real-world denoising

    Geonwoon Jang, Wooseok Lee, Sanghyun Son, and Kyoung Mu Lee. C2n: Practical generative noise modeling for real-world denoising. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 2350–2359, 2021

  18. [26]

    EnlightenGAN: Deep light enhancement without paired supervision.IEEE Transactions on Image Processing, 30:2340–2349, 2021

    Yifan Jiang, Xinyu Gong, Ding Liu, Yu Cheng, Chen Fang, Xiaohui Shen, Jianchao Yang, Pan Zhou, and Zhangyang Wang. EnlightenGAN: Deep light enhancement without paired supervision.IEEE Transactions on Image Processing, 30:2340–2349, 2021. 22

  19. [27]

    srgb real noise modeling via noise-aware sampling with normalizing flows

    Dongjin Kim, Donggoo Jung, Sungyong Baik, and Tae Hyun Kim. srgb real noise modeling via noise-aware sampling with normalizing flows. InThe Twelfth International Conference on Learning Representations, 2024

  20. [28]

    Adam: A method for stochastic optimization.arXiv preprint arXiv:1412.6980, 2014

    Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization.arXiv preprint arXiv:1412.6980, 2014

  21. [29]

    Glow: Generative flow with invertible 1x1 convolutions

    Durk P Kingma and Prafulla Dhariwal. Glow: Generative flow with invertible 1x1 convolutions. Advances in neural information processing systems, 31, 2018

  22. [30]

    Modeling srgb camera noise with normalizing flows

    Shayan Kousha, Ali Maleky, Michael S Brown, and Marcus A Brubaker. Modeling srgb camera noise with normalizing flows. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17463–17471, 2022

  23. [31]

    Imagenet classification with deep convolu- tional neural networks

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolu- tional neural networks. InAdvances in neural information processing systems, volume 25, 2012

  24. [32]

    Deep learning.nature, 521(7553):436–444, 2015

    Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning.nature, 521(7553):436–444, 2015

  25. [33]

    Drit++: Diverse image-to-image translation via disentangled representations.International Journal of Computer Vision, 128(10):2402–2417, 2020

    Hsin-Ying Lee, Hung-Yu Tseng, Qi Mao, Jia-Bin Huang, Yu-Ding Lu, Maneesh Singh, and Ming-Hsuan Yang. Drit++: Diverse image-to-image translation via disentangled representations.International Journal of Computer Vision, 128(10):2402–2417, 2020

  26. [34]

    W. Lee, S. Son, and K. M. Lee. AP-BSN: Self-supervised denoising for real-world images via asymmetric PD and blind-spot network. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 17725–17734, 2022

  27. [35]

    Noise-transfer2clean: denoising cryo-em images based on noise modeling and transfer

    Hongjia Li, Hui Zhang, Xiaohua Wan, Zhidong Yang, Chengmin Li, Jintao Li, Renmin Han, Ping Zhu, and Fa Zhang. Noise-transfer2clean: denoising cryo-em images based on noise modeling and transfer. Bioinformatics, 38(7):2022–2029, 2022

  28. [36]

    Dual contrastive learning for unsupervised image-to-image translation

    Nianjian Liu et al. Dual contrastive learning for unsupervised image-to-image translation. InPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3156–3164, 2021

  29. [37]

    Multi-level wavelet-cnn for image restoration

    Pengju Liu, Hongzhi Zhang, Kai Zhang, Liang Lin, and Wangmeng Zuo. Multi-level wavelet-cnn for image restoration. InProceedings of the IEEE conference on computer vision and pattern recognition workshops, pages 773–782, 2018

  30. [38]

    Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017

  31. [39]

    A theory for multiresolution signal decomposition: the wavelet representation

    St´ ephane G Mallat. A theory for multiresolution signal decomposition: the wavelet representation. IEEE transactions on pattern analysis and machine intelligence, 11(7):674–693, 1989

  32. [40]

    Cyclenet: Rethinking cycle consistency in text-guided diffusion for image manipulation

    Chenlin Meng, Yutong He, Yang Gu, Shuang Yuan, Jiaming Song, Prafulla Ding, et al. Cyclenet: Rethinking cycle consistency in text-guided diffusion for image manipulation. InAdvances in Neural Information Processing Systems, volume 36, 2023

  33. [41]

    A holistic approach to cross-channel image noise modeling and its application to image denoising

    Seonghwan Nam, Youngbae Hwang, Yasuyuki Matsushita, and Seon Joo Kim. A holistic approach to cross-channel image noise modeling and its application to image denoising. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 1683–1691, 2016

  34. [42]

    DINOv2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193, 2023

    Maxime Oquab, Timoth´ ee Darcet, Th´ eo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Vasiljevic, Pulin Sun, Leonid Pishchulin, Peter Tulloch, Rene Ranftl, et al. DINOv2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193, 2023. 23

  35. [43]

    Recorrupted-to-recorrupted: Unsupervised deep learning for image denoising

    Tongyao Pang, Huan Zheng, Yuhui Quan, and Hui Ji. Recorrupted-to-recorrupted: Unsupervised deep learning for image denoising. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2043–2052, 2021

  36. [44]

    Contrastive learning for unpaired image-to-image translation

    Taesung Park, Alexei A Efros, Richard Zhang, and Jun-Yan Zhu. Contrastive learning for unpaired image-to-image translation. InEuropean conference on computer vision, pages 319–345. Springer, 2020

  37. [45]

    One-step image translation with text-to-image models

    Gaurav Parmar, Krishna Kumar Singh, Richard Zhang, Yijun Li, Jingwan Lu, and Jun-Yan Zhu. One-step image translation with text-to-image models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22709–22719, 2024

  38. [46]

    Benchmarking denoising algorithms with real photographs

    Tobias Plotz and Stefan Roth. Benchmarking denoising algorithms with real photographs. InProceed- ings of the IEEE conference on computer vision and pattern recognition, pages 1586–1595, 2017

  39. [47]

    Unsupervised representation learning with deep con- volutional generative adversarial networks

    Alec Radford, Luke Metz, and Soumith Chintala. Unsupervised representation learning with deep con- volutional generative adversarial networks. InInternational Conference on Learning Representations (ICLR), 2016

  40. [48]

    Statistics of natural images: Scaling in the woods.Advances in neural information processing systems, 6, 1993

    Daniel Ruderman and William Bialek. Statistics of natural images: Scaling in the woods.Advances in neural information processing systems, 6, 1993

  41. [49]

    Comprehensive integration of single-cell data.Cell, 177(7):1888–1902, 2019

    Tim Stuart, Andrew Butler, Paul Hoffman, Christoph Hafemeister, Efthymia Papalexi, William M Mauck III, Yuhan Hao, Marlon Stoeckius, Peter Smibert, and Rahul Satija. Comprehensive integration of single-cell data.Cell, 177(7):1888–1902, 2019

  42. [50]

    Dual diffusion implicit bridges for image- to-image translation

    Xuan Su, Jiaming Song, Chenlin Meng, and Stefano Ermon. Dual diffusion implicit bridges for image- to-image translation. InInternational Conference on Learning Representations, 2023

  43. [51]

    Deep image prior

    Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. Deep image prior. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 9446–9454, 2018

  44. [52]

    Nvae: A deep hierarchical variational autoencoder.Advances in neural information processing systems, 33:19667–19679, 2020

    Arash Vahdat and Jan Kautz. Nvae: A deep hierarchical variational autoencoder.Advances in neural information processing systems, 33:19667–19679, 2020

  45. [53]

    Deflow: Learning complex image degradations from unpaired data with conditional flows

    Valentin Wolf, Andreas Lugmayr, Martin Danelljan, Luc Van Gool, and Radu Timofte. Deflow: Learning complex image degradations from unpaired data with conditional flows. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 94–103, 2021

  46. [54]

    Cryo-em structure of thePlasmodium falciparum 80s ribosome bound to the anti-protozoan drug emetine.eLife, 3:e03080, jun 2014

    Wilson Wong, Xiao-chen Bai, Alan Brown, Israel S Fernandez, Eric Hanssen, Melanie Condron, Yan Hong Tan, Jake Baum, and Sjors HW Scheres. Cryo-em structure of thePlasmodium falciparum 80s ribosome bound to the anti-protozoan drug emetine.eLife, 3:e03080, jun 2014

  47. [55]

    Xie et al

    E. Xie et al. Unpaired image-to-image translation with shortest path regularization. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023

  48. [56]

    Real-world noisy image denoising: A new benchmark

    Jun Xu, Hui Li, Zhetong Liang, David Zhang, and Lei Zhang. Real-world noisy image denoising: A new benchmark. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 168–178, 2018

  49. [57]

    Self-augmented unpaired image dehazing via density and depth decomposition

    Yang Yang, Chaoyue Wang, Risheng Liu, Lin Zhang, Xiaojie Guo, and Dacheng Tao. Self-augmented unpaired image dehazing via density and depth decomposition. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2037–2046, 2022

  50. [58]

    Dual adversarial network: Toward real-world noise removal and noise generation

    Zongsheng Yue, Qian Zhao, Lei Zhang, and Deyu Meng. Dual adversarial network: Toward real-world noise removal and noise generation. InComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part X 16, pages 41–58. Springer, 2020. 24

  51. [59]

    Plug-and-play image restoration with deep denoiser prior.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):6360–6376, 2021

    Kai Zhang, Yawei Li, Wangmeng Zuo, Lei Zhang, Luc Van Gool, and Radu Timofte. Plug-and-play image restoration with deep denoiser prior.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):6360–6376, 2021

  52. [60]

    Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising.IEEE transactions on image processing, 26(7):3142– 3155, 2017

    Kai Zhang, Wangmeng Zuo, Yunjin Chen, Deyu Meng, and Lei Zhang. Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising.IEEE transactions on image processing, 26(7):3142– 3155, 2017

  53. [61]

    Residual dense network for image super-resolution

    Yulun Zhang, Yapeng Tian, Yu Kong, Bineng Zhong, and Yun Fu. Residual dense network for image super-resolution. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 2472–2481, 2018

  54. [62]

    Learn from unpaired data for image restoration: A variational bayes approach.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(5):5889–5903, 2022

    Dihan Zheng, Xiaowen Zhang, Kaisheng Ma, and Chenglong Bao. Learn from unpaired data for image restoration: A variational bayes approach.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(5):5889–5903, 2022

  55. [63]

    Senm-vae: Semi-supervised noise modeling with hierarchical variational autoencoder

    Dihan Zheng, Yihang Zou, Xiaowen Zhang, and Chenglong Bao. Senm-vae: Semi-supervised noise modeling with hierarchical variational autoencoder. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 25889–25899, 2024

  56. [64]

    Unpaired image-to-image translation using cycle-consistent adversarial networks

    Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. InProceedings of the IEEE international conference on computer vision, pages 2223–2232, 2017. 25

Pith tools

Reviewed July 10, 2026 · model on record in the stance chip above.