Pith. sign in

REVIEW 3 major objections 4 minor 68 references

Leveraging the Powerful Attention of a Pre-trained Diffusion Model for Exemplar-based Image Colorization

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A pre-trained diffusion model's self-attention, used as cross-image attention with zero fine-tuning, delivers state-of-the-art exemplar-based image colorization on a 335-pair benchmark.

desk verdict A genuinely training-free diffusion-attention colorizer with consistent metric wins, but the headline margin is partly a product of tuning on the test set. read the letter →

arxiv 2505.15812 v1 pith:FNZTEUGH submitted 2025-05-21 cs.CV

classification cs.CV
keywords exemplar-basedcolorizationdiffusionmodelself-attentioncross-imageattentionclassifier-freeguidanceDDIMinversionsemanticcorrespondencetraining-free
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to show that the self-attention mechanism of a pre-trained diffusion model can be reused, without any fine-tuning, to match semantically corresponding regions between a grayscale image and a color reference image. It proposes a colorization pipeline that turns self-attention into cross-image attention during denoising, using two complementary attention maps: one between the grayscale input and a grayscale version of the reference, and one between the colorized output and the color reference. A classifier-free guidance step then amplifies the difference between color-transferred and non-color-transferred outputs to make the transferred colors look more natural. On the established 335-pair benchmark the method reports FID 95.27 and SI-FID 5.51, the best of the compared methods on every reported metric, and it also leads on a new 100-pair dataset pairing grayscale historical paintings with contemporary color photos. If correct, this means expensive task-specific training for exemplar colorization can be replaced by leveraging the semantic knowledge already stored in large pretrained diffusion models.

What carries the argument

The central object is the self-attention module of the pre-trained Stable Diffusion U-Net, repurposed as cross-image attention. In each applied layer, query features are taken from the input image's denoising pipeline while key and value features are taken from the reference image's inverted latents, so the scaled dot-product attention map assigns reference colors to regions of the input that are semantically similar. The machinery has four load-bearing parts: (1) DDIM inversion collects query and key latents for both images; (2) dual attention combines gray-to-gray attention (input vs. grayscale reference) and colorized-to-color attention (output vs. color reference) post-softmax as A_dual = softmax(S_g2g · γ + S_c2c · (1−γ)); (3) classifier-free colorization guidance extrapolates the color-transferred denoising output beyond the non-transferred output with weight w>1; and (4) auxiliary components, including self-attention injection, early stopping of inversion at T=5, Lab post-processing, repetition N=3, and initial latent AdaIN, stabilize structure and reference fidelity.

What would settle it

Build a test pair in which the input and reference have identical luminance structure but opposite semantics, for example a reference sky that is red and an input whose similarly shaped region is actually a wall known to be red; if the method still transfers the sky's red into the wall region, the correspondence is driven by low-level appearance rather than semantics. A cleaner verdict would come from a benchmark with per-pixel semantic-correspondence annotations, measuring whether the transferred color lands inside the same semantic label.

Watch

Extended reading notes

Core claim

The central claim is that the self-attention layers of a diffusion model trained on billions of images already encode the semantic correspondences needed for exemplar colorization, and that this capacity can be unlocked by a training-free modification of the denoising process. Concretely, the paper converts query, key, and value computation in Stable Diffusion's self-attention into cross-image attention: queries come from the input image, keys and values from the reference, so the attention map serves as a semantic guide for where to place reference colors. Because a direct gray-to-color similarity is unreliable, the paper uses dual attention: gray-to-gray similarity between the input and the grayscale reference, combined with colorized-to-color similarity between the current colorized output and the color reference, fused post-softmax with a scalar weight. A classifier-free colorization guidance step, extrapolating with weight w=10 between the color-transferred and non-color-transferred noise predictions, sharpens the transferred colors. The output is the best of the compared methods on all six metrics on the 335-pair benchmark (FID 95.27, SI-FID 5.51, HIS 0.792, LPIPS 0.634) and on the cross-domain painting dataset.

Load-bearing premise

The load-bearing premise is that the attention maps of a diffusion model trained on RGB images give correct semantic correspondences between a grayscale input and a color reference even after inversion is stopped early at five steps; the paper offers no direct check that attention aligns with human-perceived object boundaries.

Editorial extensions

If this is right

  • A training-free colorizer built on Stable Diffusion attention matches or beats trained VGG- and attention-based colorizers on the 335-pair benchmark across all six reported metrics.
  • The same dual-attention and guidance recipe can be lifted to other exemplar-based image translation tasks, such as style transfer, relighting, or normal transfer, without retraining.
  • Because the method requires no fine-tuning, it inherits improvements in base diffusion models automatically: a stronger pretrained model should directly yield stronger colorization.
  • Early stopping of DDIM inversion at five steps keeps quality nearly unchanged while cutting runtime from 36.7 seconds to 5.2 seconds and GPU memory from 31 GB to 14 GB.
  • The method degrades on input-reference pairs with no semantic correspondence, for example a color palette, drawing a clear boundary between when this approach works and when it does not.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the attention maps are doing the semantic matching, the same self-attention-as-cross-attention trick could replace learned correspondence modules in other dense prediction tasks, such as semantic segmentation transfer or keypoint transfer.
  • A natural testable extension is to weight the two attention branches by their own confidence, for example inverse entropy, instead of a fixed scalar γ; the paper's fixed-weight ablation suggests this could further improve alignment.
  • The finding that early stopping at T=5 retains most of the benefit hints that semantic alignment is already available in the earliest denoising steps; if verified, a single-step distilled variant could give most of the gain at a fraction of the runtime.
  • One could combine this colorizer with an automatic reference retriever that selects references with high semantic overlap, since the method explicitly fails when semantics are absent.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a fine-tuning-free exemplar-based image colorization method that uses the self-attention of a pre-trained Stable Diffusion model as a cross-image attention mechanism. Two contributions are presented: dual attention-guided color transfer, which combines gray-to-gray and colorized-to-color attention to move reference colors to semantically matching regions, and classifier-free colorization guidance, which extrapolates between color-transferred and non-color-transferred denoising outputs. The method is evaluated on 335 input-reference pairs from prior work and on a newly collected 100-pair dataset of historical paintings with contemporary photo references, reporting the best FID, SI-FID, ARNIQA, MUSIQ, HIS, and LPIPS among the compared methods.

Significance. If the reported quantitative results are reliable, the paper makes a useful empirical contribution: it shows that self-attention features of a large pre-trained diffusion model can support training-free exemplar colorization, with public code and a clear pipeline. The ablation coverage is broad, and the new cross-domain dataset is a practical addition. The core novelty, adapting self-attention into dual cross-image attention and amplifying it with a classifier-free-style guidance, is plausible and well motivated. The main risk is not the method itself but the evaluation protocol: hyperparameters and decoder-layer choices are selected on the same 335-pair benchmark on which the headline comparison is made, and the reported margins over strong baselines are small. As a result, the paper's central quantitative superiority claim is not yet established to the standard needed for a journal publication.

major comments (3)
  1. [§V-A, Appendix II, Appendix III, Tables I, VII, VIII] The headline claim of 'outperforms existing techniques' rests on metric margins that may be artifacts of model selection on the test benchmark. The adopted decoder-layer group (Table VII) and the hyperparameters γ, β, T, w, and N (Table VIII) are chosen by evaluating many configurations on the same 335-pair dataset used for the final comparison in Table I. The reported FID of 95.27 is therefore the maximum of a noisy objective over dozens of configurations, not an unbiased estimate of performance on new pairs. The margins are small: FID 95.27 vs. 95.85 for He et al., and SI-FID 5.51 vs. 6.11 for Lu et al. No confidence intervals, bootstrap estimates, or held-out validation are reported, and FID computed on only 335 images is known to have high sampling variability. The authors should either provide a validation split or cross-validation for hyperparameter selection, report uncertainty intervals, or explicitly reframe the numbers as configuration-tuned results rather than evidence of universal superiority.
  2. [§V-D, Appendix I, Table V] The ablation narrative is contradicted by the paper's own table. Section V-D states that the original condition 'outperforms the ablation settings,' but Table V shows that 'Ours w/o self-attention injection' achieves better FID (93.64 vs. 95.27), better SI-FID (5.41 vs. 5.51), and better HIS (0.816 vs. 0.792) than the proposed full method. Since self-attention injection is presented as a component of the method, this trade-off needs a concrete explanation, for example a stated preference for perceptual-quality metrics, or a revised claim that the component improves some quality metrics at the expense of others. As written, the categorical claim in §V-D is not supported by the evidence in the same paper.
  3. [§IV-A, Eq. (10), Fig. 4] The paper's central mechanism is that the dual attention map captures semantic correspondences, but no direct evaluation of correspondence accuracy is provided. The attention map is used both to transfer color and as the evidence that semantic matching works, which is circular if the only validation is the resulting colorized images. A direct evaluation on a small set of labeled region or keypoint correspondences, or an analysis of attention maps against human-annotated semantic regions, would substantially strengthen the claim that the pre-trained self-attention is finding semantically meaningful matches rather than relying on low-level luminance or texture cues.
minor comments (4)
  1. [Abstract] The sentence 'The color features from the reference image is then transferred' contains a subject-verb agreement error; it should be 'are then transferred.'
  2. [§IV-A, Eq. (13)] The notation in Eq. (13) is slightly confusing: A_in is defined in Eq. (12) as an attention map computed from the input query and key features, but the equation then multiplies A_in by value features v_out of the output. Clarifying that v_out is the output feature projection would help readability.
  3. [§V-B] The description of FID as 'the embedding distance measured between the output image set and the reference image set' is imprecise, since FID compares distributions and is not a single distance computed directly between two sets of embeddings. Rephrasing would avoid confusion.
  4. [Appendix VI, Table XI] The claim that the method achieves the best or second-best performance on 'almost all metrics' should be checked against the table: for MANIQA in the standard scenario, Ours (0.489) is below Xiao et al. (0.517) and several others. The average-rank statement is fine, but the sentence as written is too strong.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the colorization pipeline is built from an external pre-trained model and evaluated on external benchmarks; no central prediction reduces to its inputs by construction.

full rationale

The derivation chain is self-contained in the relevant sense. The method constructs color-transferred latents via dual self-attention maps (Eqs. 8-13) using a fixed pre-trained Stable Diffusion model; no parameters are learned from, or defined by, the target colorization benchmark. The classifier-free colorization guidance (Eqs. 14-15) is a designed extrapolation between the method's own color-transferred and non-color-transferred denoising outputs, not a fitted parameter renamed as a prediction. Post-processing replaces the output L channel with the input L channel and normalizes the ab channels to the reference ab statistics; these are intentional task constraints that influence fidelity metrics, but they are not hidden re-derivations of the headline result because the semantic placement of color is not forced by them. The main quantitative claim (FID 95.27, SI-FID 5.51) is vulnerable to the fact that hyperparameters (γ, β, T, w, N) and decoder-layer groups were selected on the same 335-pair benchmark (Appendix II-III, Tables VII-VIII), so the reported margins may be optimistic; however, this is a statistical validity and selection-bias concern about the evaluation, not a circularity of the derivation. Self-citations in the introduction and related work are peripheral and not load-bearing. The attention map is used both to transfer color and to illustrate semantic matching, but the paper's central claims are evaluated on external benchmarks and against external baselines, so the reasoning does not reduce to its own output.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The method has no learned weights, but it depends on seven tuned hyperparameters or layer choices and a strong post-processing step. The most consequential choice is the ab-channel normalization to the reference mean and standard deviation, which by itself moves SI-FID from 5.51 to 6.63 and HIS from 0.792 to 0.675 when removed. Under the ledger view, the attention mechanism contributes less than the headline numbers suggest, because a cheap histogram-style correction carries much of the reference-fidelity score.

free parameters (6)
  • gamma (gray-to-gray vs colorized-to-color weight) = 1 for layers 6-8, 0.5 for layers 9-11
    Blend weight in Eq. (10); chosen by ablation in Appendix III, Table VIII.
  • beta (self-attention injection weight) = 0 for layers 6-8, 0.5 for layers 9-11
    Weight in Eq. (13); chosen by ablation in Appendix III, Table VIII.
  • w (classifier-free colorization guidance scale) = 10
    Extrapolation weight in Eq. (14); chosen by ablation in Appendix III.
  • T (DDIM inversion early-stop step) = 5
    Early stopping step; chosen by ablation in Appendix III.
  • N (colorization repetition count) = 3
    Number of repeated colorization passes; chosen by ablation in Appendix III.
  • decoder layer selection = 6th to 11th decoder layers
    Layer group chosen by experiments in Appendix II, Table VII; an architectural free choice rather than a scalar.
assumptions (4)
  • domain assumption Stable Diffusion self-attention features encode semantic correspondences across grayscale and color domains.
    Central premise of Section IV.A; no independent validation that the attention maps correspond to human semantic regions.
  • domain assumption DDIM inversion of a grayscale latent through an RGB-trained Stable Diffusion VAE produces a valid noise trajectory that preserves input structure.
    Required for Eqs. (6) and (7); the paper does not analyze the domain shift of encoding a replicated grayscale image.
  • domain assumption FID, SI-FID, HIS, LPIPS, and the no-reference IQA metrics are appropriate proxies for colorization quality and reference fidelity.
    Section V-B adopts these metrics without validation; FID here compares outputs to the reference image set, which is unusual.
  • ad hoc to paper Hyperparameters selected on the main 335-pair benchmark generalize to other pairs.
    The headline comparison and hyperparameter search use the same benchmark, so generalization is assumed rather than shown.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Leveraging the Powerful Attention of a Pre-trained Diffusion Model for Exemplar-based Image Colorization." pith.science (2026). https://pith.science/paper/FNZTEUGH

@misc{pith2026250515812,
  author       = {Pith},
  title        = {Pith review of: Leveraging the Powerful Attention of a Pre-trained Diffusion Model for Exemplar-based Image Colorization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FNZTEUGH}},
  note         = {Machine review of arXiv:2505.15812}
}
read the original abstract

Exemplar-based image colorization aims to colorize a grayscale image using a reference color image, ensuring that reference colors are applied to corresponding input regions based on their semantic similarity. To achieve accurate semantic matching between regions, we leverage the self-attention module of a pre-trained diffusion model, which is trained on a large dataset and exhibits powerful attention capabilities. To harness this power, we propose a novel, fine-tuning-free approach based on a pre-trained diffusion model, making two key contributions. First, we introduce dual attention-guided color transfer. We utilize the self-attention module to compute an attention map between the input and reference images, effectively capturing semantic correspondences. The color features from the reference image is then transferred to the semantically matching regions of the input image, guided by this attention map, and finally, the grayscale features are replaced with the corresponding color features. Notably, we utilize dual attention to calculate attention maps separately for the grayscale and color images, achieving more precise semantic alignment. Second, we propose classifier-free colorization guidance, which enhances the transferred colors by combining color-transferred and non-color-transferred outputs. This process improves the quality of colorization. Our experimental results demonstrate that our method outperforms existing techniques in terms of image quality and fidelity to the reference. Specifically, we use 335 input-reference pairs from previous research, achieving an FID of 95.27 (image quality) and an SI-FID of 5.51 (fidelity to the reference). Our source code is available at https://github.com/satoshi-kosugi/powerful-attention.

Figures

Figures reproduced from arXiv: 2505.15812 by the authors.

Figure 1
Figure 1. Overview of our method. First, z in 0 , z ref 0 , and z refL 0 are given; these are the latent variables corresponding to the input, reference, and grayscale reference, respectively. We apply DDIM inversion to these latent variables to obtain z in T , z ref T , and z refL T . Next, we set z in T as the initial value for z out T and apply the colorization process to z out T . This process consists of two key componen… view at source ↗
Figure 2
Figure 2. Illustration of (a) general usage of self-attention [12] and (b) dual [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Illustration of classifier-free colorization guidance. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Qualitative comparison. we evaluate the colorization results from two perspectives: 1) image quality and 2) fidelity to the reference. To assess 1) image quality, we use three metrics. • FID [52]: the embedding distance is measured between the output image set and the …
Figure 5
Figure 5. Figure 5: Qualitative comparison in ablation studies. [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Qualitative comparison in the cross-domain reference scenario. The references are contemporary photos, and the inputs are historical paintings. [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 8
Figure 8. Figure 8: Results using a color palette reference. The reference and input are [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 11
Figure 11. Figure 11: Our method demonstrates strong and consistent per [PITH_FULL_IMAGE:figures/full_fig_p012_11.png]
Figure 9
Figure 9. Figure 9: Qualitative comparison in ablation studies. [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Qualitative comparison using different color spaces. [PITH_FULL_IMAGE:figures/full_fig_p015_10.png]
Figure 11
Figure 11. Figure 11: Qualitative comparison using a complex input. [PITH_FULL_IMAGE:figures/full_fig_p016_11.png]
Figure 12
Figure 12. Figure 12: Qualitative comparison [PITH_FULL_IMAGE:figures/full_fig_p017_12.png]
Figure 13
Figure 13. Figure 13: Qualitative comparison in the cross-domain reference scenario. [PITH_FULL_IMAGE:figures/full_fig_p018_13.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

68 extracted references · 63 canonical work pages

  1. [1]

    H. Yue, J. Liu, J. Yang, X. Sun, T. Q. Nguyen, and F. Wu, ``Ienet: Internal and external patch matching convnet for web image guided denoising,'' IEEE Trans. Circuits Syst. Video Technol., vol. 30, no. 11, pp. 3928--3942, Nov. 2020

  2. [2]

    L. Guo, S. Huang, H. Liu, and B. Wen, ``Towards robust image denoising via flow-based joint image and noise model,'' IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 7, pp. 6105--6115, Jul. 2024

  3. [3]

    Zhang, J

    J. Zhang, J. Pan, D. Wang, S. Zhou, X. Wei, F. Zhao, J. Liu, and J. Ren, ``Deep dynamic scene deblurring from optical flow,'' IEEE Trans. Circuits Syst. Video Technol., vol. 32, no. 12, pp. 8250--8260, Dec. 2022

  4. [4]

    Zhang, T

    K. Zhang, T. Wang, W. Luo, W. Ren, B. Stenger, W. Liu, H. Li, and M.-H. Yang, ``Mc-blur: A comprehensive benchmark for image deblurring,'' IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 5, pp. 3755--3767, May 2024

  5. [5]

    Kumar, R

    A. Kumar, R. K. Jha, and N. K. Nishchal, ``Dynamic stochastic resonance and image fusion based model for quality enhancement of dark and hazy images,'' J. Electron. Imag., vol. 30, no. 6, pp. 063\,008--1--063\,008--20, Nov. 2021

  6. [6]

    ------, ``An improved gamma correction model for image dehazing in a multi-exposure fusion framework,'' J. Vis. Commun. Image Representation, vol. 78, p. 103122, Jul. 2021

  7. [7]

    Kosugi and T

    S. Kosugi and T. Yamasaki, ``Unpaired image enhancement featuring reinforcement-learning-controlled image editing software,'' in Proc. AAAI Conf. Artif. Intell., vol. 34, no. 07, Feb. 2020, pp. 11\,296--11\,303

  8. [8]

    Liang, Y

    J. Liang, Y. Xu, Y. Quan, B. Shi, and H. Ji, ``Self-supervised low-light image enhancement using discrepant untrained network priors,'' IEEE Trans. Circuits Syst. Video Technol., vol. 32, no. 11, pp. 7332--7345, Nov. 2022

Show all 68 references
  1. [9]

    Kosugi and T

    S. Kosugi and T. Yamasaki, ``Crowd-powered photo enhancement featuring an active learning based local filter,'' IEEE Trans. Circuits Syst. Video Technol., vol. 33, no. 7, pp. 3145--3158, Jul. 2023

  2. [10]

    Kosugi, ``Prompt-guided image-adaptive neural implicit lookup tables for interpretable image enhancement,'' in Proc

    S. Kosugi, ``Prompt-guided image-adaptive neural implicit lookup tables for interpretable image enhancement,'' in Proc. ACM Int. Conf. Multimedia, Oct. 2024, pp. 6463--6471

  3. [11]

    Kosugi and T

    S. Kosugi and T. Yamasaki, ``Personalized image enhancement featuring masked style modeling,'' IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 1, pp. 140--152, Jan. 2024

  4. [12]

    Rombach, A

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, ``High-resolution image synthesis with latent diffusion models,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2022, pp. 10\,684--10\,695

  5. [13]

    Schuhmann, R

    C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al., ``Laion-5b: An open large-scale dataset for training next generation image-text models,'' in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 35, D...

  6. [14]

    P. Lu, J. Yu, X. Peng, Z. Zhao, and X. Wang, ``Gray2colornet: Transfer more colors from reference image,'' in Proc. ACM Int. Conf. Multimedia, Oct. 2020, pp. 3210--3218

  7. [15]

    attention is all you need

    W. Yin, P. Lu, Z. Zhao, and X. Peng, ``Yes," attention is all you need", for exemplar based colorization,'' in Proc. ACM Int. Conf. Multimedia, Oct. 2021, pp. 2243--2251

  8. [16]

    Zhang, C

    J. Zhang, C. Xu, J. Li, Y. Han, Y. Wang, Y. Tai, and Y. Liu, ``Scsnet: An efficient paradigm for learning simultaneously image colorization and super-resolution,'' in Proc. AAAI Conf. Artif. Intell., vol. 36, no. 3, Feb. 2022, pp. 3271--3279

  9. [17]

    Y. Bai, C. Dong, Z. Chai, A. Wang, Z. Xu, and C. Yuan, ``Semantic-sparse colorization network for deep exemplar-based colorization,'' in Proc. Eur. Conf. Comput. Vis. (ECCV), Oct. 2022, pp. 505--521

  10. [18]

    Carrillo, M

    H. Carrillo, M. Cl \'e ment, and A. Bugeau, ``Super-attention for exemplar-based image colorization,'' in Proc. Asian Conf. Comput. Vis. (ACCV), Dec. 2022, pp. 4548--4564

  11. [19]

    H. Wang, D. Zhai, X. Liu, J. Jiang, and W. Gao, ``Unsupervised deep exemplar colorization via pyramid dual non-local attention,'' IEEE Trans. Image Process., vol. 32, pp. 4114--4127, Jul. 2023

  12. [20]

    C. Zou, S. Wan, M. G. Blanch, L. Murn, M. Mrak, J. Sock, F. Yang, and L. Herranz, ``Lightweight deep exemplar colorization via semantic attention-guided laplacian pyramid,'' IEEE Trans. Vis. Comput. Graph., pp. 1--12, May 2024

  13. [21]

    J. Song, C. Meng, and S. Ermon, ``Denoising diffusion implicit models,'' arXiv preprint arXiv:2010.02502, Oct. 2020

  14. [22]

    Ho and T

    J. Ho and T. Salimans, ``Classifier-free diffusion guidance,'' in Proc. NeurIPS Workshop DGMs Appl, Dec. 2021

  15. [23]

    M. He, D. Chen, J. Liao, P. V. Sander, and L. Yuan, ``Deep exemplar-based colorization,'' ACM Trans. Graph., vol. 37, no. 4, pp. 1--16, Jul. 2018

  16. [24]

    Welsh, M

    T. Welsh, M. Ashikhmin, and K. Mueller, ``Transferring color to greyscale images,'' ACM Trans. Graph., vol. 21, no. 3, pp. 277--280, Jul. 2002

  17. [25]

    Morimoto, Y

    Y. Morimoto, Y. Taguchi, and T. Naemura, ``Automatic colorization of grayscale images using multiple images on the web,'' in Proc. SIGGRAPH, Aug. 2009, pp. 59--60

  18. [26]

    Oliva and A

    A. Oliva and A. Torralba, ``Modeling the shape of the scene: A holistic representation of the spatial envelope,'' Int. J. Comput. Vis., vol. 42, pp. 145--175, May 2001

  19. [27]

    Irony, D

    R. Irony, D. Cohen-Or, and D. Lischinski, ``Colorization by example,'' in Proc. Eurographics Symp. Rendering, Jul. 2005, pp. 201--210

  20. [28]

    Charpiat, M

    G. Charpiat, M. Hofmann, and B. Sch \"o lkopf, ``Automatic image colorization via multimodal predictions,'' in Proc. Eur. Conf. Comput. Vis. (ECCV), Oct. 2008, pp. 126--139

  21. [29]

    R. K. Gupta, A. Y.-S. Chia, D. Rajan, E. S. Ng, and H. Zhiyong, ``Image colorization using similar images,'' in Proc. ACM Int. Conf. Multimedia, Oct. 2012, pp. 369--378

  22. [30]

    H. Bay, T. Tuytelaars, and L. Van Gool, ``Surf: Speeded up robust features,'' in Proc. Eur. Conf. Comput. Vis. (ECCV), May 2006, pp. 404--417

  23. [31]

    A. Y.-S. Chia, S. Zhuo, R. K. Gupta, Y.-W. Tai, S.-Y. Cho, P. Tan, and S. Lin, ``Semantic colorization with internet images,'' ACM Trans. Graph., vol. 30, no. 6, pp. 1--8, Dec. 2011

  24. [32]

    X. Liu, L. Wan, Y. Qu, T.-T. Wong, S. Lin, C.-S. Leung, and P.-A. Heng, ``Intrinsic colorization,'' ACM Trans. Graph., vol. 27, no. 5, pp. 1--9, Dec. 2008

  25. [33]

    D. G. Lowe, ``Object recognition from local scale-invariant features,'' in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), vol. 2, Sep. 1999, pp. 1150--1157

  26. [34]

    F. Fang, T. Wang, T. Zeng, and G. Zhang, ``A superpixel-based variational model for image colorization,'' IEEE Trans. Vis. Comput. Graph., vol. 26, no. 10, pp. 2931--2943, Oct. 2019

  27. [35]

    Dalal and B

    N. Dalal and B. Triggs, ``Histograms of oriented gradients for human detection,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), vol. 1, Jun. 2005, pp. 886--893

  28. [36]

    E. Tola, V. Lepetit, and P. Fua, ``Daisy: An efficient dense descriptor applied to wide-baseline stereo,'' IEEE Trans. Pattern Anal. Mach. Intell., vol. 32, no. 5, pp. 815--830, May 2009

  29. [37]

    Z. Xu, T. Wang, F. Fang, Y. Sheng, and G. Zhang, ``Stylization-based architecture for fast deep exemplar colorization,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2020, pp. 9363--9372

  30. [38]

    H. Li, B. Sheng, P. Li, R. Ali, and C. P. Chen, ``Globally and locally semantic colorization via exemplar-based broad-gan,'' IEEE Trans. Image Process., vol. 30, pp. 8526--8539, Oct. 2021

  31. [39]

    Huang, N

    Z. Huang, N. Zhao, and J. Liao, ``Unicolor: A unified framework for multi-modal colorization with transformer,'' ACM Trans. Graph., vol. 41, no. 6, pp. 1--16, Nov. 2022

  32. [40]

    Leduc, H

    R. Leduc, H. Carrillo, and N. Papadakis, ``Non-local matching of superpixel-based deep features for color transfer and colorization,'' Image Process. On Line, vol. 14, pp. 232--249, Oct. 2024

  33. [41]

    Simonyan and A

    K. Simonyan and A. Zisserman, ``Very deep convolutional networks for large-scale image recognition,'' arXiv preprint arXiv:1409.1556, Sep. 2014

  34. [42]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, ``Imagenet: A large-scale hierarchical image database,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2009, pp. 248--255

  35. [43]

    Liang, Z

    Z. Liang, Z. Li, S. Zhou, C. Li, and C. C. Loy, ``Control color: Multimodal diffusion-based interactive image colorization,'' arXiv preprint arXiv:2402.10855, Feb. 2024

  36. [44]

    Hertz, R

    A. Hertz, R. Mokady, J. Tenenbaum, K. Aberman, Y. Pritch, and D. Cohen-Or, ``Prompt-to-prompt image editing with cross attention control,'' arXiv preprint arXiv:2208.01626, Aug. 2022

  37. [45]

    Tumanyan, M

    N. Tumanyan, M. Geyer, S. Bagon, and T. Dekel, ``Plug-and-play diffusion features for text-driven image-to-image translation,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2023, pp. 1921--1930

  38. [46]

    M. Cao, X. Wang, Z. Qi, Y. Shan, X. Qie, and Y. Zheng, ``Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing,'' in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Oct. 2023, pp. 22\,560--22\,570

  39. [47]

    Chung, S

    J. Chung, S. Hyun, and J.-P. Heo, ``Style injection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2024, pp. 8795--8805

  40. [48]

    B. Liu, C. Wang, T. Cao, K. Jia, and J. Huang, ``Towards understanding cross and self-attention in stable diffusion for text-guided image editing,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2024, pp. 7817--7826

  41. [49]

    Ronneberger, P

    O. Ronneberger, P. Fischer, and T. Brox, ``U-net: Convolutional networks for biomedical image segmentation,'' in Proc. Int. Conf. Med. Image Comput. Comput.-Assist. Intervent (MICCAI), Oct. 2015, pp. 234--241

  42. [50]

    Connolly and T

    C. Connolly and T. Fleiss, ``A study of efficiency and accuracy in the transformation from rgb to cielab color space,'' IEEE Trans. Image Process., vol. 6, no. 7, pp. 1046--1048, Jul. 1997

  43. [51]

    C. Xiao, C. Han, Z. Zhang, J. Qin, T.-T. Wong, G. Han, and S. He, ``Example-based colourization via dense encoding pyramids,'' in Comput. Graph. Forum, vol. 39, no. 1, Apr. 2020, pp. 20--33

  44. [52]

    Heusel, H

    M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, ``Gans trained by a two time-scale update rule converge to a local nash equilibrium,'' in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 30, Dec. 2017, pp. 6629--6640

  45. [53]

    Agnolucci, L

    L. Agnolucci, L. Galteri, M. Bertini, and A. Del Bimbo, ``Arniqa: Learning distortion manifold for image quality assessment,'' in Proc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), Jan. 2024, pp. 189--198

  46. [54]

    Z. Ying, H. Niu, P. Gupta, D. Mahajan, D. Ghadiyaram, and A. Bovik, ``From patches to pictures (paq-2-piq): Mapping the perceptual space of picture quality,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2020, pp. 3575--3585

  47. [55]

    J. Ke, Q. Wang, Y. Wang, P. Milanfar, and F. Yang, ``Musiq: Multi-scale image quality transformer,'' in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Oct. 2021, pp. 5148--5157

  48. [56]

    Y. Fang, H. Zhu, Y. Zeng, K. Ma, and Z. Wang, ``Perceptual quality assessment of smartphone photography,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2020, pp. 3677--3686

  49. [57]

    T. R. Shaham, T. Dekel, and T. Michaeli, ``Singan: Learning a generative model from a single natural image,'' in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Oct. 2019, pp. 4570--4580

  50. [58]

    Isola, J.-Y

    P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, ``Image-to-image translation with conditional adversarial networks,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2017, pp. 1125--1134

  51. [59]

    Zhang, P

    R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, ``The unreasonable effectiveness of deep features as a perceptual metric,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2018, pp. 586--595

  52. [60]

    M. Xia, W. Hu, T.-T. Wong, and J. Wang, ``Disentangled image colorization via global anchors,'' ACM Trans. Graph., vol. 41, no. 6, pp. 1--13, Nov. 2022

  53. [61]

    X. Kang, T. Yang, W. Ouyang, P. Ren, L. Li, and X. Xie, ``Ddcolor: Towards photo-realistic image colorization via dual decoders,'' in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Oct. 2023, pp. 328--338

  54. [62]

    Sauer, D

    A. Sauer, D. Lorenz, A. Blattmann, and R. Rombach, ``Adversarial diffusion distillation,'' in Proc. Eur. Conf. Comput. Vis. (ECCV), Oct. 2024, pp. 87--103

  55. [63]

    Redmon and A

    J. Redmon and A. Farhadi, ``Yolov3: An incremental improvement,'' arXiv preprint arXiv:1804.02767, Apr. 2018

  56. [64]

    C. Chen, J. Mo, J. Hou, H. Wu, L. Liao, W. Sun, Q. Yan, and W. Lin, ``Topiq: A top-down approach from semantics to distortions for image quality assessment,'' IEEE Trans. Image Process., vol. 33, pp. 2404--2418, Mar. 2024

  57. [65]

    Zhang, G

    W. Zhang, G. Zhai, Y. Wei, X. Yang, and K. Ma, ``Blind image quality assessment via vision-language correspondence: A multitask learning perspective,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2023, pp. 14\,071--14\,081

  58. [66]

    S. A. Golestaneh, S. Dadsetan, and K. M. Kitani, ``No-reference image quality assessment via transformers, relative ranking, and self-consistency,'' in Proc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), Jan. 2022, pp. 1220--1230

  59. [67]

    S. Yang, T. Wu, S. Shi, S. Lao, Y. Gong, M. Cao, J. Wang, and Y. Yang, ``Maniqa: Multi-dimension attention network for no-reference image quality assessment,'' in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), Jun. 2022, pp. 1191--1200

  60. [68]

    H. Lin, V. Hosu, and D. Saupe, ``Kadid-10k: A large-scale artificially distorted iqa database,'' in Proc. Eleventh Int. Conf. Quality Multimedia Exp. (QoMEX), Jun. 2019, pp. 1--3

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.