Pith. sign in

REVIEW 7 cited by

Implicit Style-Content Separation using B-LoRA

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.14572 v2 pith:H7GC2TD5 submitted 2024-03-21 cs.CV

classification cs.CV
keywords imagestylecontentseparationstylizationb-loralorastyle-content
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Image stylization involves manipulating the visual appearance and texture (style) of an image while preserving its underlying objects, structures, and concepts (content). The separation of style and content is essential for manipulating the image's style independently from its content, ensuring a harmonious and visually pleasing result. Achieving this separation requires a deep understanding of both the visual and semantic characteristics of images, often necessitating the training of specialized models or employing heavy optimization. In this paper, we introduce B-LoRA, a method that leverages LoRA (Low-Rank Adaptation) to implicitly separate the style and content components of a single image, facilitating various image stylization tasks. By analyzing the architecture of SDXL combined with LoRA, we find that jointly learning the LoRA weights of two specific blocks (referred to as B-LoRAs) achieves style-content separation that cannot be achieved by training each B-LoRA independently. Consolidating the training into only two blocks and separating style and content allows for significantly improving style manipulation and overcoming overfitting issues often associated with model fine-tuning. Once trained, the two B-LoRAs can be used as independent components to allow various image stylization tasks, including image style transfer, text-based image stylization, consistent style generation, and style-content mixing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Palette Aligned Image Diffusion

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Palette-Adapter conditions text-to-image diffusion on a sparse color palette treated as a histogram, with entropy and distance controls and a negative-color guidance mechanism.

  2. OmniStyle: Filtering High Quality Style Transfer Data at Scale

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new million-triplet dataset and a diffusion transformer model that performs text-guided and image-guided style transfer, with a filtering pipeline used to curate high-quality training examples.

  3. A LoRA is Worth a Thousand Pictures

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LoRA weight vectors, projected with PCA and a per-PC calibration, cluster and retrieve artistic styles more accurately than CLIP, DINO, and style-specialized image features.

  4. StyleMaster: Stylize Your Video with Artistic Generation and Translation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    StyleMaster improves reference-image video stylization by extracting global and local style separately, training on model-illusion pairs, and adding a motion adapter and gray tile ControlNet.

  5. LayerFusion: Harmonized Multi-Layer Text-to-Image Generation with Generative Priors

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LayerFusion creates harmonized foreground, background, and blended images at once by blending the attention outputs of two diffusion models, with no extra training.

  6. Style-Friendly SNR Sampler for Style-Driven Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Fine-tuning diffusion models for style-driven generation is improved by shifting the training noise-level distribution toward high-noise steps, where stylistic features are shown to emerge.

  7. Stable Flow: Vital Layers for Training-Free Image Editing

    cs.CV 2024-11 conditional novelty 6.0 of 10

    An automatic vital-layer selection for FLUX enables training-free, stable text-driven image editing via selective attention injection.

Pith tools