Pith. sign in

REVIEW 16 cited by

StyleAdapter: A Unified Stylized Image Generation Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.01770 v2 pith:34HB6Z2A submitted 2023-09-04 cs.CV

classification cs.CV
keywords styleimagescontentpromptgenerationimagemodelsemantic
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This work focuses on generating high-quality images with specific style of reference images and content of provided textual descriptions. Current leading algorithms, i.e., DreamBooth and LoRA, require fine-tuning for each style, leading to time-consuming and computationally expensive processes. In this work, we propose StyleAdapter, a unified stylized image generation model capable of producing a variety of stylized images that match both the content of a given prompt and the style of reference images, without the need for per-style fine-tuning. It introduces a two-path cross-attention (TPCA) module to separately process style information and textual prompt, which cooperate with a semantic suppressing vision model (SSVM) to suppress the semantic content of style images. In this way, it can ensure that the prompt maintains control over the content of the generated images, while also mitigating the negative impact of semantic information in style references. This results in the content of the generated image adhering to the prompt, and its style aligning with the style references. Besides, our StyleAdapter can be integrated with existing controllable synthesis methods, such as T2I-adapter and ControlNet, to attain a more controllable and stable generation process. Extensive experiments demonstrate the superiority of our method over previous works.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DreamStyle3D: Efficient 3D Stylized Asset Generation via Dual-Attention Disentanglement

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Decoupled dual cross-attention plus style/content augmentations let a TRELLIS-based model inject image style into 3D assets in ~10s while better preserving geometry than prior 2D-to-3D pipelines.

  2. PoseAlign: Sculpting Pose-Consistent Meshes via Text-Guided Deformation

    cs.GR 2026-07 conditional novelty 6.0 of 10

    Two-stage text-guided mesh deformation (Laplacian CLIP scaling + attention-shared SDS Jacobian sculpting) better preserves source pose while aligning to text than TextDeformer or MeshUp.

  3. DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A diffusion-based disentanglement method that separates object content from domain style achieves state-of-the-art unsupervised cross-domain image retrieval on three benchmarks.

  4. Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Layout-Togglable storytelling is introduced: diffusion transformers conditioned on layout enable precise control over character position and appearance, supported by a new large-scale dataset and benchmark.

  5. OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data

    cs.CV 2025-05 conditional novelty 6.0 of 10

    OmniConsistency is a style-agnostic consistency module for Flux that preserves structure and details during stylization with arbitrary LoRAs, reaching GPT-4o-level content consistency.

  6. CDST: Color Disentangled Style Transfer for Universal Style Reference Customization

    cs.CV 2025-05 conditional novelty 6.0 of 10

    CDST disentangles color from style via greyscale style input and a color histogram stream, enabling zero-shot style transfer with separate color control and a new characteristics-preserved mode.

  7. StyleBlend: Enhancing Style-Specific Content Creation in Text-to-Image Diffusion Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    StyleBlend learns few-shot artistic style as separate layout and texture components and blends them during diffusion sampling to improve text-aligned, style-specific image generation.

  8. StyleMaster: Stylize Your Video with Artistic Generation and Translation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    StyleMaster improves reference-image video stylization by extracting global and local style separately, training on model-illusion pairs, and adding a motion adapter and gray tile ControlNet.

  9. FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    The paper introduces FiVA, a ~1M-image synthetic dataset with fine-grained visual attribute labels, and FiVA-Adapter, a diffusion adapter that transfers and combines attributes like lighting, color, and motion from re...

  10. Numerical Study of Oblique Detonation Initiation Assisted by Local Energy Deposition

    physics.flu-dyn 2025-08 unverdicted novelty 5.0 of 10

    Pulsatile local energy deposition can initiate sustainable oblique detonation on a finite wedge with less than 10% of the average power needed by continuous deposition.

  11. Less is More: Masking Elements in Image Condition Features Avoids Content Leakages in Style Transfer Diffusion Models

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Masking the image-feature dimensions most correlated with the style reference's content text reduces content leakage and improves text fidelity in text-to-image style transfer diffusion models.

  12. WikiStyle+: A Multimodal Approach to Content-Style Representation Disentanglement for Artistic Image Stylization

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A multimodal dataset and diffusion method that explicitly separates content from style in artistic images, reducing content leakage during stylization.

  13. StyleStudio: Text-Driven Style Transfer with Selective Control of Style Elements

    cs.CV 2024-12 conditional novelty 5.0 of 10

    StyleStudio improves text-driven style transfer with cross-modal AdaIN, a negative-style-image classifier-free guidance, and teacher-model layout stabilization.

  14. StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation

    cs.CV 2025-05 reject novelty 4.0 of 10

    StyleAR enables autoregressive image generation models to do style-aligned text-to-image generation using only binary text-image data, via self-reconstruction training and style-enhanced tokens.

  15. ICAS: IP Adapter and ControlNet-based Attention Structure for Multi-Subject Style Transfer Optimization

    cs.CV 2025-04 reject novelty 3.0 of 10

    ICAS is a style-transfer pipeline that freezes IP-Adapter's style path, lightly tunes its content path, and adds ControlNet structure conditioning to preserve multi-subject layouts.

  16. Parameter-Efficient Fine-Tuning for Foundation Models

    cs.CL 2025-01 conditional novelty 2.0 of 10

    A survey that categorizes and summarizes parameter-efficient fine-tuning methods across large language, vision, and multimodal models.

Pith tools