Pith. sign in

REVIEW 7 cited by

BLIP-Diffusion: Pre-trained Subject Representation for Controllable Text-to-Image Generation and Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.14720 v2 pith:RQKUTSSD submitted 2023-05-24 cs.CV cs.AI

classification cs.CVcs.AI
keywords subjectgenerationblip-diffusionrepresentationsubject-drivenmodelsmodelmultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Subject-driven text-to-image generation models create novel renditions of an input subject based on text prompts. Existing models suffer from lengthy fine-tuning and difficulties preserving the subject fidelity. To overcome these limitations, we introduce BLIP-Diffusion, a new subject-driven image generation model that supports multimodal control which consumes inputs of subject images and text prompts. Unlike other subject-driven generation models, BLIP-Diffusion introduces a new multimodal encoder which is pre-trained to provide subject representation. We first pre-train the multimodal encoder following BLIP-2 to produce visual representation aligned with the text. Then we design a subject representation learning task which enables a diffusion model to leverage such visual representation and generates new subject renditions. Compared with previous methods such as DreamBooth, our model enables zero-shot subject-driven generation, and efficient fine-tuning for customized subject with up to 20x speedup. We also demonstrate that BLIP-Diffusion can be flexibly combined with existing techniques such as ControlNet and prompt-to-prompt to enable novel subject-driven generation and editing applications. Code and models will be released at https://github.com/salesforce/LAVIS/tree/main/projects/blip-diffusion. Project page at https://dxli94.github.io/BLIP-Diffusion-website/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Delta-Adapter: Scalable Exemplar-Based Image Editing with Single-Pair Supervision

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    Delta-Adapter extracts a semantic delta from a single image pair via a pre-trained vision encoder and injects it through a Perceiver adapter to enable scalable single-pair supervised editing.

  2. MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    MaSC is a masked similarity metric that decomposes concept-driven image generation evaluation into subject-specific preservation and background-based prompt following using SigLIP2 embeddings, outperforming global bas...

  3. PostureObjectstitch: Anomaly Image Generation Considering Assembly Relationships in Industrial Scenarios

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    PostureObjectStitch generates assembly-aware anomaly images by decoupling multi-view features into high-frequency, texture and RGB components, modulating them temporally in a diffusion model, and applying conditional ...

  4. NanoControl: A Lightweight Framework for Precise and Efficient Control in Diffusion Transformer

    cs.CV 2025-08 conditional novelty 5.0 of 10

    NanoControl injects condition-specific key-value pairs into every attention block of Flux via a LoRA-style branch, claiming state-of-the-art controllability at 0.024% extra parameters and 0.029% extra FLOPs.

  5. Compressible boundary layers over isotropic porous surfaces

    physics.flu-dyn 2025-08 unverdicted novelty 5.0 of 10

    A self-similar solution for compressible laminar boundary layers over isotropic porous substrates predicts reduced adiabatic recovery temperature and interface shear at high porosity, large grains, and high Mach numbers.

  6. StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization

    cs.CV 2025-07 unverdicted novelty 5.0 of 10

    A training-free inference-time pipeline uses masked cross-image attention sharing and region harmonization to keep subjects consistent across generated story images.

  7. SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation

    cs.CV 2025-06 unverdicted novelty 5.0 of 10

    SynMotion combines disentangled semantic embeddings, parameter-efficient motion adapters, and alternate subject-motion training on a new SPV dataset to improve motion customization in text-to-video and image-to-video ...

Pith tools