Pith. sign in

Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Recent controllable generation approaches such as FreeControl and Diffusion Self-Guidance bring fine-grained spatial and appearance control to text-to-image (T2I) diffusion models without training auxiliary modules. However, these methods optimize the latent embedding for each type of score function with longer diffusion steps, making the generation process time-consuming and limiting their flexibility and use. This work presents Ctrl-X, a simple framework for T2I diffusion controlling structure and appearance without additional training or guidance. Ctrl-X designs feed-forward structure control to enable the structure alignment with a structure image and semantic-aware appearance transfer to facilitate the appearance transfer from a user-input image. Extensive qualitative and quantitative experiments illustrate the superior performance of Ctrl-X on various condition inputs and model checkpoints. In particular, Ctrl-X supports novel structure and appearance control with arbitrary condition images of any modality, exhibits superior image quality and appearance transfer compared to existing works, and provides instant plug-and-play functionality to any T2I and text-to-video (T2V) diffusion model. See our project page for an overview of the results: https://genforce.github.io/ctrl-x

citation-role summary

other 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

other 1

polarities

unclear 1

representative citing papers

Domain Generalizable Portrait Style Transfer

cs.CV · 2025-07-06 · conditional · novelty 6.0

A diffusion-based portrait style transfer method that uses semantic face alignment and an AdaIN-Wavelet latent blend to transfer style across photo, cartoon, sketch, and animation domains while preserving identity.

citing papers explorer

Showing 1 of 1 citing paper.

  • Domain Generalizable Portrait Style Transfer cs.CV · 2025-07-06 · conditional · none · ref 16 · internal anchor

    A diffusion-based portrait style transfer method that uses semantic face alignment and an AdaIN-Wavelet latent blend to transfer style across photo, cartoon, sketch, and animation domains while preserving identity.