REVIEW 7 cited by
Implicit Style-Content Separation using B-LoRA
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Image stylization involves manipulating the visual appearance and texture (style) of an image while preserving its underlying objects, structures, and concepts (content). The separation of style and content is essential for manipulating the image's style independently from its content, ensuring a harmonious and visually pleasing result. Achieving this separation requires a deep understanding of both the visual and semantic characteristics of images, often necessitating the training of specialized models or employing heavy optimization. In this paper, we introduce B-LoRA, a method that leverages LoRA (Low-Rank Adaptation) to implicitly separate the style and content components of a single image, facilitating various image stylization tasks. By analyzing the architecture of SDXL combined with LoRA, we find that jointly learning the LoRA weights of two specific blocks (referred to as B-LoRAs) achieves style-content separation that cannot be achieved by training each B-LoRA independently. Consolidating the training into only two blocks and separating style and content allows for significantly improving style manipulation and overcoming overfitting issues often associated with model fine-tuning. Once trained, the two B-LoRAs can be used as independent components to allow various image stylization tasks, including image style transfer, text-based image stylization, consistent style generation, and style-content mixing.
Forward citations
Cited by 7 Pith papers
-
Palette Aligned Image Diffusion
Palette-Adapter conditions text-to-image diffusion on a sparse color palette treated as a histogram, with entropy and distance controls and a negative-color guidance mechanism.
-
OmniStyle: Filtering High Quality Style Transfer Data at Scale
A new million-triplet dataset and a diffusion transformer model that performs text-guided and image-guided style transfer, with a filtering pipeline used to curate high-quality training examples.
-
A LoRA is Worth a Thousand Pictures
LoRA weight vectors, projected with PCA and a per-PC calibration, cluster and retrieve artistic styles more accurately than CLIP, DINO, and style-specialized image features.
-
StyleMaster: Stylize Your Video with Artistic Generation and Translation
StyleMaster improves reference-image video stylization by extracting global and local style separately, training on model-illusion pairs, and adding a motion adapter and gray tile ControlNet.
-
LayerFusion: Harmonized Multi-Layer Text-to-Image Generation with Generative Priors
LayerFusion creates harmonized foreground, background, and blended images at once by blending the attention outputs of two diffusion models, with no extra training.
-
Style-Friendly SNR Sampler for Style-Driven Generation
Fine-tuning diffusion models for style-driven generation is improved by shifting the training noise-level distribution toward high-noise steps, where stylistic features are shown to emerge.
-
Stable Flow: Vital Layers for Training-Free Image Editing
An automatic vital-layer selection for FLUX enables training-free, stable text-driven image editing via selective attention injection.
Discussion (0). Continue with ORCID to comment.