Pith. sign in

DiffArtist: Towards Structure and Appearance Controllable Image Stylization

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Artistic styles are defined by both their structural and appearance elements. Existing neural stylization techniques primarily focus on transferring appearance-level features such as color and texture, often neglecting the equally crucial aspect of structural stylization. To address this gap, we introduce \textbf{DiffArtist}, the first 2D stylization method to offer fine-grained, simultaneous control over both structure and appearance style strength. This dual controllability is achieved by representing structure and appearance generation as separate diffusion processes, necessitating no further tuning or additional adapters. To properly evaluate this new capability of dual stylization, we further propose a Multimodal LLM-based stylization evaluator that aligns significantly better with human preferences than existing metrics. Extensive analysis shows that DiffArtist achieves superior style fidelity and dual-controllability compared to state-of-the-art methods. Its text-driven, training-free design and unprecedented dual controllability make it a powerful and interactive tool for various creative applications. Project homepage: https://diffusionartist.github.io.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

REJECT 1

roles

background 1

polarities

support 1

representative citing papers

Controllable Coupled Image Generation via Diffusion Models

cs.CV · 2025-06-07 · reject · novelty 6.0

A cross-attention control method that couples backgrounds across multiple generated images by blending LLM-extracted background and entity prompts with a time-varying weight optimized for background similarity and text-image alignment.

citing papers explorer

Showing 1 of 1 citing paper.

  • Controllable Coupled Image Generation via Diffusion Models cs.CV · 2025-06-07 · reject · none · ref 23 · internal anchor

    A cross-attention control method that couples backgrounds across multiple generated images by blending LLM-extracted background and entity prompts with a time-varying weight optimized for background similarity and text-image alignment.