REVIEW 4 cited by
DiffArtist: Towards Structure and Appearance Controllable Image Stylization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Artistic styles are defined by both their structural and appearance elements. Existing neural stylization techniques primarily focus on transferring appearance-level features such as color and texture, often neglecting the equally crucial aspect of structural stylization. To address this gap, we introduce \textbf{DiffArtist}, the first 2D stylization method to offer fine-grained, simultaneous control over both structure and appearance style strength. This dual controllability is achieved by representing structure and appearance generation as separate diffusion processes, necessitating no further tuning or additional adapters. To properly evaluate this new capability of dual stylization, we further propose a Multimodal LLM-based stylization evaluator that aligns significantly better with human preferences than existing metrics. Extensive analysis shows that DiffArtist achieves superior style fidelity and dual-controllability compared to state-of-the-art methods. Its text-driven, training-free design and unprecedented dual controllability make it a powerful and interactive tool for various creative applications. Project homepage: https://diffusionartist.github.io.
Forward citations
Cited by 4 Pith papers
-
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
A zero-shot two-stage prompting baseline (ArtCoT) makes multimodal LLMs' aesthetic judgments align substantially better with human expert rankings in pairwise artwork comparisons.
-
Controllable Coupled Image Generation via Diffusion Models
A cross-attention control method that couples backgrounds across multiple generated images by blending LLM-extracted background and entity prompts with a time-varying weight optimized for background similarity and tex...
-
Training-Free Style and Content Transfer by Leveraging U-Net Skip Connections in Stable Diffusion
Injecting the fourth and fifth U-Net skip connections from one Stable Diffusion image into another transfers content or style without any training.
-
WikiStyle+: A Multimodal Approach to Content-Style Representation Disentanglement for Artistic Image Stylization
A multimodal dataset and diffusion method that explicitly separates content from style in artistic images, reducing content leakage during stylization.
Discussion (0). Continue with ORCID to comment.