REVIEW 3 major objections 4 minor 68 references
Leveraging the Powerful Attention of a Pre-trained Diffusion Model for Exemplar-based Image Colorization
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A pre-trained diffusion model's self-attention, used as cross-image attention with zero fine-tuning, delivers state-of-the-art exemplar-based image colorization on a 335-pair benchmark.
desk verdict A genuinely training-free diffusion-attention colorizer with consistent metric wins, but the headline margin is partly a product of tuning on the test set. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the self-attention module of the pre-trained Stable Diffusion U-Net, repurposed as cross-image attention. In each applied layer, query features are taken from the input image's denoising pipeline while key and value features are taken from the reference image's inverted latents, so the scaled dot-product attention map assigns reference colors to regions of the input that are semantically similar. The machinery has four load-bearing parts: (1) DDIM inversion collects query and key latents for both images; (2) dual attention combines gray-to-gray attention (input vs. grayscale reference) and colorized-to-color attention (output vs. color reference) post-softmax as A_dual = softmax(S_g2g · γ + S_c2c · (1−γ)); (3) classifier-free colorization guidance extrapolates the color-transferred denoising output beyond the non-transferred output with weight w>1; and (4) auxiliary components, including self-attention injection, early stopping of inversion at T=5, Lab post-processing, repetition N=3, and initial latent AdaIN, stabilize structure and reference fidelity.
What would settle it
Build a test pair in which the input and reference have identical luminance structure but opposite semantics, for example a reference sky that is red and an input whose similarly shaped region is actually a wall known to be red; if the method still transfers the sky's red into the wall region, the correspondence is driven by low-level appearance rather than semantics. A cleaner verdict would come from a benchmark with per-pixel semantic-correspondence annotations, measuring whether the transferred color lands inside the same semantic label.
Extended reading notes
Core claim
The central claim is that the self-attention layers of a diffusion model trained on billions of images already encode the semantic correspondences needed for exemplar colorization, and that this capacity can be unlocked by a training-free modification of the denoising process. Concretely, the paper converts query, key, and value computation in Stable Diffusion's self-attention into cross-image attention: queries come from the input image, keys and values from the reference, so the attention map serves as a semantic guide for where to place reference colors. Because a direct gray-to-color similarity is unreliable, the paper uses dual attention: gray-to-gray similarity between the input and the grayscale reference, combined with colorized-to-color similarity between the current colorized output and the color reference, fused post-softmax with a scalar weight. A classifier-free colorization guidance step, extrapolating with weight w=10 between the color-transferred and non-color-transferred noise predictions, sharpens the transferred colors. The output is the best of the compared methods on all six metrics on the 335-pair benchmark (FID 95.27, SI-FID 5.51, HIS 0.792, LPIPS 0.634) and on the cross-domain painting dataset.
Load-bearing premise
The load-bearing premise is that the attention maps of a diffusion model trained on RGB images give correct semantic correspondences between a grayscale input and a color reference even after inversion is stopped early at five steps; the paper offers no direct check that attention aligns with human-perceived object boundaries.
Editorial extensions
If this is right
- A training-free colorizer built on Stable Diffusion attention matches or beats trained VGG- and attention-based colorizers on the 335-pair benchmark across all six reported metrics.
- The same dual-attention and guidance recipe can be lifted to other exemplar-based image translation tasks, such as style transfer, relighting, or normal transfer, without retraining.
- Because the method requires no fine-tuning, it inherits improvements in base diffusion models automatically: a stronger pretrained model should directly yield stronger colorization.
- Early stopping of DDIM inversion at five steps keeps quality nearly unchanged while cutting runtime from 36.7 seconds to 5.2 seconds and GPU memory from 31 GB to 14 GB.
- The method degrades on input-reference pairs with no semantic correspondence, for example a color palette, drawing a clear boundary between when this approach works and when it does not.
Reading between the lines
- If the attention maps are doing the semantic matching, the same self-attention-as-cross-attention trick could replace learned correspondence modules in other dense prediction tasks, such as semantic segmentation transfer or keypoint transfer.
- A natural testable extension is to weight the two attention branches by their own confidence, for example inverse entropy, instead of a fixed scalar γ; the paper's fixed-weight ablation suggests this could further improve alignment.
- The finding that early stopping at T=5 retains most of the benefit hints that semantic alignment is already available in the earliest denoising steps; if verified, a single-step distilled variant could give most of the gain at a fraction of the runtime.
- One could combine this colorizer with an automatic reference retriever that selects references with high semantic overlap, since the method explicitly fails when semantics are absent.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a fine-tuning-free exemplar-based image colorization method that uses the self-attention of a pre-trained Stable Diffusion model as a cross-image attention mechanism. Two contributions are presented: dual attention-guided color transfer, which combines gray-to-gray and colorized-to-color attention to move reference colors to semantically matching regions, and classifier-free colorization guidance, which extrapolates between color-transferred and non-color-transferred denoising outputs. The method is evaluated on 335 input-reference pairs from prior work and on a newly collected 100-pair dataset of historical paintings with contemporary photo references, reporting the best FID, SI-FID, ARNIQA, MUSIQ, HIS, and LPIPS among the compared methods.
Significance. If the reported quantitative results are reliable, the paper makes a useful empirical contribution: it shows that self-attention features of a large pre-trained diffusion model can support training-free exemplar colorization, with public code and a clear pipeline. The ablation coverage is broad, and the new cross-domain dataset is a practical addition. The core novelty, adapting self-attention into dual cross-image attention and amplifying it with a classifier-free-style guidance, is plausible and well motivated. The main risk is not the method itself but the evaluation protocol: hyperparameters and decoder-layer choices are selected on the same 335-pair benchmark on which the headline comparison is made, and the reported margins over strong baselines are small. As a result, the paper's central quantitative superiority claim is not yet established to the standard needed for a journal publication.
major comments (3)
- [§V-A, Appendix II, Appendix III, Tables I, VII, VIII] The headline claim of 'outperforms existing techniques' rests on metric margins that may be artifacts of model selection on the test benchmark. The adopted decoder-layer group (Table VII) and the hyperparameters γ, β, T, w, and N (Table VIII) are chosen by evaluating many configurations on the same 335-pair dataset used for the final comparison in Table I. The reported FID of 95.27 is therefore the maximum of a noisy objective over dozens of configurations, not an unbiased estimate of performance on new pairs. The margins are small: FID 95.27 vs. 95.85 for He et al., and SI-FID 5.51 vs. 6.11 for Lu et al. No confidence intervals, bootstrap estimates, or held-out validation are reported, and FID computed on only 335 images is known to have high sampling variability. The authors should either provide a validation split or cross-validation for hyperparameter selection, report uncertainty intervals, or explicitly reframe the numbers as configuration-tuned results rather than evidence of universal superiority.
- [§V-D, Appendix I, Table V] The ablation narrative is contradicted by the paper's own table. Section V-D states that the original condition 'outperforms the ablation settings,' but Table V shows that 'Ours w/o self-attention injection' achieves better FID (93.64 vs. 95.27), better SI-FID (5.41 vs. 5.51), and better HIS (0.816 vs. 0.792) than the proposed full method. Since self-attention injection is presented as a component of the method, this trade-off needs a concrete explanation, for example a stated preference for perceptual-quality metrics, or a revised claim that the component improves some quality metrics at the expense of others. As written, the categorical claim in §V-D is not supported by the evidence in the same paper.
- [§IV-A, Eq. (10), Fig. 4] The paper's central mechanism is that the dual attention map captures semantic correspondences, but no direct evaluation of correspondence accuracy is provided. The attention map is used both to transfer color and as the evidence that semantic matching works, which is circular if the only validation is the resulting colorized images. A direct evaluation on a small set of labeled region or keypoint correspondences, or an analysis of attention maps against human-annotated semantic regions, would substantially strengthen the claim that the pre-trained self-attention is finding semantically meaningful matches rather than relying on low-level luminance or texture cues.
minor comments (4)
- [Abstract] The sentence 'The color features from the reference image is then transferred' contains a subject-verb agreement error; it should be 'are then transferred.'
- [§IV-A, Eq. (13)] The notation in Eq. (13) is slightly confusing: A_in is defined in Eq. (12) as an attention map computed from the input query and key features, but the equation then multiplies A_in by value features v_out of the output. Clarifying that v_out is the output feature projection would help readability.
- [§V-B] The description of FID as 'the embedding distance measured between the output image set and the reference image set' is imprecise, since FID compares distributions and is not a single distance computed directly between two sets of embeddings. Rephrasing would avoid confusion.
- [Appendix VI, Table XI] The claim that the method achieves the best or second-best performance on 'almost all metrics' should be checked against the table: for MANIQA in the standard scenario, Ours (0.489) is below Xiao et al. (0.517) and several others. The average-rank statement is fine, but the sentence as written is too strong.
Circularity Check
No significant circularity: the colorization pipeline is built from an external pre-trained model and evaluated on external benchmarks; no central prediction reduces to its inputs by construction.
full rationale
The derivation chain is self-contained in the relevant sense. The method constructs color-transferred latents via dual self-attention maps (Eqs. 8-13) using a fixed pre-trained Stable Diffusion model; no parameters are learned from, or defined by, the target colorization benchmark. The classifier-free colorization guidance (Eqs. 14-15) is a designed extrapolation between the method's own color-transferred and non-color-transferred denoising outputs, not a fitted parameter renamed as a prediction. Post-processing replaces the output L channel with the input L channel and normalizes the ab channels to the reference ab statistics; these are intentional task constraints that influence fidelity metrics, but they are not hidden re-derivations of the headline result because the semantic placement of color is not forced by them. The main quantitative claim (FID 95.27, SI-FID 5.51) is vulnerable to the fact that hyperparameters (γ, β, T, w, N) and decoder-layer groups were selected on the same 335-pair benchmark (Appendix II-III, Tables VII-VIII), so the reported margins may be optimistic; however, this is a statistical validity and selection-bias concern about the evaluation, not a circularity of the derivation. Self-citations in the introduction and related work are peripheral and not load-bearing. The attention map is used both to transfer color and to illustrate semantic matching, but the paper's central claims are evaluated on external benchmarks and against external baselines, so the reasoning does not reduce to its own output.
Assumptions & free parameters
free parameters (6)
- gamma (gray-to-gray vs colorized-to-color weight) =
1 for layers 6-8, 0.5 for layers 9-11
- beta (self-attention injection weight) =
0 for layers 6-8, 0.5 for layers 9-11
- w (classifier-free colorization guidance scale) =
10
- T (DDIM inversion early-stop step) =
5
- N (colorization repetition count) =
3
- decoder layer selection =
6th to 11th decoder layers
assumptions (4)
- domain assumption Stable Diffusion self-attention features encode semantic correspondences across grayscale and color domains.
- domain assumption DDIM inversion of a grayscale latent through an RGB-trained Stable Diffusion VAE produces a valid noise trajectory that preserves input structure.
- domain assumption FID, SI-FID, HIS, LPIPS, and the no-reference IQA metrics are appropriate proxies for colorization quality and reference fidelity.
- ad hoc to paper Hyperparameters selected on the main 335-pair benchmark generalize to other pairs.
Cite this review
Pith. "Pith review of Leveraging the Powerful Attention of a Pre-trained Diffusion Model for Exemplar-based Image Colorization." pith.science (2026). https://pith.science/paper/FNZTEUGH
@misc{pith2026250515812,
author = {Pith},
title = {Pith review of: Leveraging the Powerful Attention of a Pre-trained Diffusion Model for Exemplar-based Image Colorization},
year = {2026},
howpublished = {\url{https://pith.science/paper/FNZTEUGH}},
note = {Machine review of arXiv:2505.15812}
}
read the original abstract
Exemplar-based image colorization aims to colorize a grayscale image using a reference color image, ensuring that reference colors are applied to corresponding input regions based on their semantic similarity. To achieve accurate semantic matching between regions, we leverage the self-attention module of a pre-trained diffusion model, which is trained on a large dataset and exhibits powerful attention capabilities. To harness this power, we propose a novel, fine-tuning-free approach based on a pre-trained diffusion model, making two key contributions. First, we introduce dual attention-guided color transfer. We utilize the self-attention module to compute an attention map between the input and reference images, effectively capturing semantic correspondences. The color features from the reference image is then transferred to the semantically matching regions of the input image, guided by this attention map, and finally, the grayscale features are replaced with the corresponding color features. Notably, we utilize dual attention to calculate attention maps separately for the grayscale and color images, achieving more precise semantic alignment. Second, we propose classifier-free colorization guidance, which enhances the transferred colors by combining color-transferred and non-color-transferred outputs. This process improves the quality of colorization. Our experimental results demonstrate that our method outperforms existing techniques in terms of image quality and fidelity to the reference. Specifically, we use 335 input-reference pairs from previous research, achieving an FID of 95.27 (image quality) and an SI-FID of 5.51 (fidelity to the reference). Our source code is available at https://github.com/satoshi-kosugi/powerful-attention.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
H. Yue, J. Liu, J. Yang, X. Sun, T. Q. Nguyen, and F. Wu, ``Ienet: Internal and external patch matching convnet for web image guided denoising,'' IEEE Trans. Circuits Syst. Video Technol., vol. 30, no. 11, pp. 3928--3942, Nov. 2020
work page 2020
-
[2]
L. Guo, S. Huang, H. Liu, and B. Wen, ``Towards robust image denoising via flow-based joint image and noise model,'' IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 7, pp. 6105--6115, Jul. 2024
work page 2024
- [3]
- [4]
- [5]
-
[6]
------, ``An improved gamma correction model for image dehazing in a multi-exposure fusion framework,'' J. Vis. Commun. Image Representation, vol. 78, p. 103122, Jul. 2021
work page 2021
-
[7]
S. Kosugi and T. Yamasaki, ``Unpaired image enhancement featuring reinforcement-learning-controlled image editing software,'' in Proc. AAAI Conf. Artif. Intell., vol. 34, no. 07, Feb. 2020, pp. 11\,296--11\,303
work page 2020
- [8]
Show all 68 references
-
[9]
Kosugi and T
S. Kosugi and T. Yamasaki, ``Crowd-powered photo enhancement featuring an active learning based local filter,'' IEEE Trans. Circuits Syst. Video Technol., vol. 33, no. 7, pp. 3145--3158, Jul. 2023
2023
-
[10]
Kosugi, ``Prompt-guided image-adaptive neural implicit lookup tables for interpretable image enhancement,'' in Proc
S. Kosugi, ``Prompt-guided image-adaptive neural implicit lookup tables for interpretable image enhancement,'' in Proc. ACM Int. Conf. Multimedia, Oct. 2024, pp. 6463--6471
2024
-
[11]
Kosugi and T
S. Kosugi and T. Yamasaki, ``Personalized image enhancement featuring masked style modeling,'' IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 1, pp. 140--152, Jan. 2024
2024
-
[12]
Rombach, A
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, ``High-resolution image synthesis with latent diffusion models,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2022, pp. 10\,684--10\,695
2022
-
[13]
Schuhmann, R
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al., ``Laion-5b: An open large-scale dataset for training next generation image-text models,'' in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 35, D...
2022
-
[14]
P. Lu, J. Yu, X. Peng, Z. Zhao, and X. Wang, ``Gray2colornet: Transfer more colors from reference image,'' in Proc. ACM Int. Conf. Multimedia, Oct. 2020, pp. 3210--3218
2020
-
[15]
attention is all you need
W. Yin, P. Lu, Z. Zhao, and X. Peng, ``Yes," attention is all you need", for exemplar based colorization,'' in Proc. ACM Int. Conf. Multimedia, Oct. 2021, pp. 2243--2251
2021
-
[16]
Zhang, C
J. Zhang, C. Xu, J. Li, Y. Han, Y. Wang, Y. Tai, and Y. Liu, ``Scsnet: An efficient paradigm for learning simultaneously image colorization and super-resolution,'' in Proc. AAAI Conf. Artif. Intell., vol. 36, no. 3, Feb. 2022, pp. 3271--3279
2022
-
[17]
Y. Bai, C. Dong, Z. Chai, A. Wang, Z. Xu, and C. Yuan, ``Semantic-sparse colorization network for deep exemplar-based colorization,'' in Proc. Eur. Conf. Comput. Vis. (ECCV), Oct. 2022, pp. 505--521
2022
-
[18]
Carrillo, M
H. Carrillo, M. Cl \'e ment, and A. Bugeau, ``Super-attention for exemplar-based image colorization,'' in Proc. Asian Conf. Comput. Vis. (ACCV), Dec. 2022, pp. 4548--4564
2022
-
[19]
H. Wang, D. Zhai, X. Liu, J. Jiang, and W. Gao, ``Unsupervised deep exemplar colorization via pyramid dual non-local attention,'' IEEE Trans. Image Process., vol. 32, pp. 4114--4127, Jul. 2023
2023
-
[20]
C. Zou, S. Wan, M. G. Blanch, L. Murn, M. Mrak, J. Sock, F. Yang, and L. Herranz, ``Lightweight deep exemplar colorization via semantic attention-guided laplacian pyramid,'' IEEE Trans. Vis. Comput. Graph., pp. 1--12, May 2024
2024
-
[21]
J. Song, C. Meng, and S. Ermon, ``Denoising diffusion implicit models,'' arXiv preprint arXiv:2010.02502, Oct. 2020
2010 arXiv
-
[22]
Ho and T
J. Ho and T. Salimans, ``Classifier-free diffusion guidance,'' in Proc. NeurIPS Workshop DGMs Appl, Dec. 2021
2021
-
[23]
M. He, D. Chen, J. Liao, P. V. Sander, and L. Yuan, ``Deep exemplar-based colorization,'' ACM Trans. Graph., vol. 37, no. 4, pp. 1--16, Jul. 2018
2018
-
[24]
Welsh, M
T. Welsh, M. Ashikhmin, and K. Mueller, ``Transferring color to greyscale images,'' ACM Trans. Graph., vol. 21, no. 3, pp. 277--280, Jul. 2002
2002
-
[25]
Morimoto, Y
Y. Morimoto, Y. Taguchi, and T. Naemura, ``Automatic colorization of grayscale images using multiple images on the web,'' in Proc. SIGGRAPH, Aug. 2009, pp. 59--60
2009
-
[26]
Oliva and A
A. Oliva and A. Torralba, ``Modeling the shape of the scene: A holistic representation of the spatial envelope,'' Int. J. Comput. Vis., vol. 42, pp. 145--175, May 2001
2001
-
[27]
Irony, D
R. Irony, D. Cohen-Or, and D. Lischinski, ``Colorization by example,'' in Proc. Eurographics Symp. Rendering, Jul. 2005, pp. 201--210
2005
-
[28]
Charpiat, M
G. Charpiat, M. Hofmann, and B. Sch \"o lkopf, ``Automatic image colorization via multimodal predictions,'' in Proc. Eur. Conf. Comput. Vis. (ECCV), Oct. 2008, pp. 126--139
2008
-
[29]
R. K. Gupta, A. Y.-S. Chia, D. Rajan, E. S. Ng, and H. Zhiyong, ``Image colorization using similar images,'' in Proc. ACM Int. Conf. Multimedia, Oct. 2012, pp. 369--378
2012
-
[30]
H. Bay, T. Tuytelaars, and L. Van Gool, ``Surf: Speeded up robust features,'' in Proc. Eur. Conf. Comput. Vis. (ECCV), May 2006, pp. 404--417
2006
-
[31]
A. Y.-S. Chia, S. Zhuo, R. K. Gupta, Y.-W. Tai, S.-Y. Cho, P. Tan, and S. Lin, ``Semantic colorization with internet images,'' ACM Trans. Graph., vol. 30, no. 6, pp. 1--8, Dec. 2011
2011
-
[32]
X. Liu, L. Wan, Y. Qu, T.-T. Wong, S. Lin, C.-S. Leung, and P.-A. Heng, ``Intrinsic colorization,'' ACM Trans. Graph., vol. 27, no. 5, pp. 1--9, Dec. 2008
2008
-
[33]
D. G. Lowe, ``Object recognition from local scale-invariant features,'' in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), vol. 2, Sep. 1999, pp. 1150--1157
1999
-
[34]
F. Fang, T. Wang, T. Zeng, and G. Zhang, ``A superpixel-based variational model for image colorization,'' IEEE Trans. Vis. Comput. Graph., vol. 26, no. 10, pp. 2931--2943, Oct. 2019
2019
-
[35]
Dalal and B
N. Dalal and B. Triggs, ``Histograms of oriented gradients for human detection,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), vol. 1, Jun. 2005, pp. 886--893
2005
-
[36]
E. Tola, V. Lepetit, and P. Fua, ``Daisy: An efficient dense descriptor applied to wide-baseline stereo,'' IEEE Trans. Pattern Anal. Mach. Intell., vol. 32, no. 5, pp. 815--830, May 2009
2009
-
[37]
Z. Xu, T. Wang, F. Fang, Y. Sheng, and G. Zhang, ``Stylization-based architecture for fast deep exemplar colorization,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2020, pp. 9363--9372
2020
-
[38]
H. Li, B. Sheng, P. Li, R. Ali, and C. P. Chen, ``Globally and locally semantic colorization via exemplar-based broad-gan,'' IEEE Trans. Image Process., vol. 30, pp. 8526--8539, Oct. 2021
2021
-
[39]
Huang, N
Z. Huang, N. Zhao, and J. Liao, ``Unicolor: A unified framework for multi-modal colorization with transformer,'' ACM Trans. Graph., vol. 41, no. 6, pp. 1--16, Nov. 2022
2022
-
[40]
Leduc, H
R. Leduc, H. Carrillo, and N. Papadakis, ``Non-local matching of superpixel-based deep features for color transfer and colorization,'' Image Process. On Line, vol. 14, pp. 232--249, Oct. 2024
2024
-
[41]
Simonyan and A
K. Simonyan and A. Zisserman, ``Very deep convolutional networks for large-scale image recognition,'' arXiv preprint arXiv:1409.1556, Sep. 2014
2014 arXiv
-
[42]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, ``Imagenet: A large-scale hierarchical image database,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2009, pp. 248--255
2009
-
[43]
Liang, Z
Z. Liang, Z. Li, S. Zhou, C. Li, and C. C. Loy, ``Control color: Multimodal diffusion-based interactive image colorization,'' arXiv preprint arXiv:2402.10855, Feb. 2024
2024 arXiv
-
[44]
Hertz, R
A. Hertz, R. Mokady, J. Tenenbaum, K. Aberman, Y. Pritch, and D. Cohen-Or, ``Prompt-to-prompt image editing with cross attention control,'' arXiv preprint arXiv:2208.01626, Aug. 2022
2022 arXiv
-
[45]
Tumanyan, M
N. Tumanyan, M. Geyer, S. Bagon, and T. Dekel, ``Plug-and-play diffusion features for text-driven image-to-image translation,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2023, pp. 1921--1930
2023
-
[46]
M. Cao, X. Wang, Z. Qi, Y. Shan, X. Qie, and Y. Zheng, ``Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing,'' in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Oct. 2023, pp. 22\,560--22\,570
2023
-
[47]
Chung, S
J. Chung, S. Hyun, and J.-P. Heo, ``Style injection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2024, pp. 8795--8805
2024
-
[48]
B. Liu, C. Wang, T. Cao, K. Jia, and J. Huang, ``Towards understanding cross and self-attention in stable diffusion for text-guided image editing,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2024, pp. 7817--7826
2024
-
[49]
Ronneberger, P
O. Ronneberger, P. Fischer, and T. Brox, ``U-net: Convolutional networks for biomedical image segmentation,'' in Proc. Int. Conf. Med. Image Comput. Comput.-Assist. Intervent (MICCAI), Oct. 2015, pp. 234--241
2015
-
[50]
Connolly and T
C. Connolly and T. Fleiss, ``A study of efficiency and accuracy in the transformation from rgb to cielab color space,'' IEEE Trans. Image Process., vol. 6, no. 7, pp. 1046--1048, Jul. 1997
1997
-
[51]
C. Xiao, C. Han, Z. Zhang, J. Qin, T.-T. Wong, G. Han, and S. He, ``Example-based colourization via dense encoding pyramids,'' in Comput. Graph. Forum, vol. 39, no. 1, Apr. 2020, pp. 20--33
2020
-
[52]
Heusel, H
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, ``Gans trained by a two time-scale update rule converge to a local nash equilibrium,'' in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 30, Dec. 2017, pp. 6629--6640
2017
-
[53]
Agnolucci, L
L. Agnolucci, L. Galteri, M. Bertini, and A. Del Bimbo, ``Arniqa: Learning distortion manifold for image quality assessment,'' in Proc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), Jan. 2024, pp. 189--198
2024
-
[54]
Z. Ying, H. Niu, P. Gupta, D. Mahajan, D. Ghadiyaram, and A. Bovik, ``From patches to pictures (paq-2-piq): Mapping the perceptual space of picture quality,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2020, pp. 3575--3585
2020
-
[55]
J. Ke, Q. Wang, Y. Wang, P. Milanfar, and F. Yang, ``Musiq: Multi-scale image quality transformer,'' in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Oct. 2021, pp. 5148--5157
2021
-
[56]
Y. Fang, H. Zhu, Y. Zeng, K. Ma, and Z. Wang, ``Perceptual quality assessment of smartphone photography,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2020, pp. 3677--3686
2020
-
[57]
T. R. Shaham, T. Dekel, and T. Michaeli, ``Singan: Learning a generative model from a single natural image,'' in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Oct. 2019, pp. 4570--4580
2019
-
[58]
Isola, J.-Y
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, ``Image-to-image translation with conditional adversarial networks,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2017, pp. 1125--1134
2017
-
[59]
Zhang, P
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, ``The unreasonable effectiveness of deep features as a perceptual metric,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2018, pp. 586--595
2018
-
[60]
M. Xia, W. Hu, T.-T. Wong, and J. Wang, ``Disentangled image colorization via global anchors,'' ACM Trans. Graph., vol. 41, no. 6, pp. 1--13, Nov. 2022
2022
-
[61]
X. Kang, T. Yang, W. Ouyang, P. Ren, L. Li, and X. Xie, ``Ddcolor: Towards photo-realistic image colorization via dual decoders,'' in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Oct. 2023, pp. 328--338
2023
-
[62]
Sauer, D
A. Sauer, D. Lorenz, A. Blattmann, and R. Rombach, ``Adversarial diffusion distillation,'' in Proc. Eur. Conf. Comput. Vis. (ECCV), Oct. 2024, pp. 87--103
2024
-
[63]
Redmon and A
J. Redmon and A. Farhadi, ``Yolov3: An incremental improvement,'' arXiv preprint arXiv:1804.02767, Apr. 2018
2018 arXiv
-
[64]
C. Chen, J. Mo, J. Hou, H. Wu, L. Liao, W. Sun, Q. Yan, and W. Lin, ``Topiq: A top-down approach from semantics to distortions for image quality assessment,'' IEEE Trans. Image Process., vol. 33, pp. 2404--2418, Mar. 2024
2024
-
[65]
Zhang, G
W. Zhang, G. Zhai, Y. Wei, X. Yang, and K. Ma, ``Blind image quality assessment via vision-language correspondence: A multitask learning perspective,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2023, pp. 14\,071--14\,081
2023
-
[66]
S. A. Golestaneh, S. Dadsetan, and K. M. Kitani, ``No-reference image quality assessment via transformers, relative ranking, and self-consistency,'' in Proc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), Jan. 2022, pp. 1220--1230
2022
-
[67]
S. Yang, T. Wu, S. Shi, S. Lao, Y. Gong, M. Cao, J. Wang, and Y. Yang, ``Maniqa: Multi-dimension attention network for no-reference image quality assessment,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), Jun. 2022, pp. 1191--1200
2022
-
[68]
H. Lin, V. Hosu, and D. Saupe, ``Kadid-10k: A large-scale artificially distorted iqa database,'' in Proc. Eleventh Int. Conf. Quality Multimedia Exp. (QoMEX), Jun. 2019, pp. 1--3
2019
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.