REVIEW 2 cited by
Control Color: Multimodal Diffusion-based Interactive Image Colorization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Despite the existence of numerous colorization methods, several limitations still exist, such as lack of user interaction, inflexibility in local colorization, unnatural color rendering, insufficient color variation, and color overflow. To solve these issues, we introduce Control Color (CtrlColor), a multi-modal colorization method that leverages the pre-trained Stable Diffusion (SD) model, offering promising capabilities in highly controllable interactive image colorization. While several diffusion-based methods have been proposed, supporting colorization in multiple modalities remains non-trivial. In this study, we aim to tackle both unconditional and conditional image colorization (text prompts, strokes, exemplars) and address color overflow and incorrect color within a unified framework. Specifically, we present an effective way to encode user strokes to enable precise local color manipulation and employ a practical way to constrain the color distribution similar to exemplars. Apart from accepting text prompts as conditions, these designs add versatility to our approach. We also introduce a novel module based on self-attention and a content-guided deformable autoencoder to address the long-standing issues of color overflow and inaccurate coloring. Extensive comparisons show that our model outperforms state-of-the-art image colorization methods both qualitatively and quantitatively.
Forward citations
Cited by 2 Pith papers
-
Leveraging the Powerful Attention of a Pre-trained Diffusion Model for Exemplar-based Image Colorization
A training-free colorization method repurposing Stable Diffusion self-attention as cross-image attention, with dual attention and classifier-free guidance, reports the best FID and SI-FID on standard benchmarks.
-
Consistent Video Colorization via Palette Guidance
A palette-guided fine-tuning of Stable Video Diffusion produces more saturated and temporally consistent video colorization than three GAN-based baselines.
Discussion (0). Continue with ORCID to comment.