Pith. sign in

REVIEW 7 cited by

DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.03374 v4 pith:S4O2PITO submitted 2023-05-05 cs.CV

DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation

classification cs.CV
keywords embeddinggenerationdisenboothinformationentangledsubject-driventext-to-imageidentity-irrelevant
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Subject-driven text-to-image generation aims to generate customized images of the given subject based on the text descriptions, which has drawn increasing attention. Existing methods mainly resort to finetuning a pretrained generative model, where the identity-relevant information (e.g., the boy) and the identity-irrelevant information (e.g., the background or the pose of the boy) are entangled in the latent embedding space. However, the highly entangled latent embedding may lead to the failure of subject-driven text-to-image generation as follows: (i) the identity-irrelevant information hidden in the entangled embedding may dominate the generation process, resulting in the generated images heavily dependent on the irrelevant information while ignoring the given text descriptions; (ii) the identity-relevant information carried in the entangled embedding can not be appropriately preserved, resulting in identity change of the subject in the generated images. To tackle the problems, we propose DisenBooth, an identity-preserving disentangled tuning framework for subject-driven text-to-image generation. Specifically, DisenBooth finetunes the pretrained diffusion model in the denoising process. Different from previous works that utilize an entangled embedding to denoise each image, DisenBooth instead utilizes disentangled embeddings to respectively preserve the subject identity and capture the identity-irrelevant information. We further design the novel weak denoising and contrastive embedding auxiliary tuning objectives to achieve the disentanglement. Extensive experiments show that our proposed DisenBooth framework outperforms baseline models for subject-driven text-to-image generation with the identity-preserved embedding. Additionally, by combining the identity-preserved embedding and identity-irrelevant embedding, DisenBooth demonstrates more generation flexibility and controllability

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Adaptive Subspace Projection for Generative Personalization

    cs.CV 2026-05 unverdicted novelty 7.0

    A training-free adaptive subspace projection method mitigates semantic collapsing in generative personalization by isolating and adjusting drift in a low-dimensional subspace using the stable pre-trained embedding as anchor.

  2. Customizing Video Portraits via Identity-ActionDecoupling

    cs.CV 2026-06 unverdicted novelty 6.0

    Proposes IaD framework with Identity Decoupling Loss and Text Alignment Loss for richer, identity-consistent IPT2V without subject-specific fine-tuning.

  3. OmniPrism: Learning Disentangled Visual Concept for Image Generation

    cs.CV 2024-12 unverdicted novelty 6.0

    OmniPrism proposes a disentanglement method using a new paired dataset (PCD-200K), COD contrastive training, and block embeddings to inject separated concepts into diffusion models for multi-aspect image generation.

  4. SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation

    cs.CV 2025-06 unverdicted novelty 5.0

    SynMotion combines disentangled semantic embeddings, parameter-efficient motion adapters, and alternate subject-motion training on a new SPV dataset to improve motion customization in text-to-video and image-to-video ...

  5. SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation

    cs.CV 2024-11 unverdicted novelty 5.0

    SOW uses MLLMs and attention to selectively control unidirectional diffusion for pixel-level fidelity and contextual coherence in text-vision-to-image tasks.

  6. Encoder-Decoder Manifold Alignment for Idempotent Generation

    cs.LG 2026-06 unverdicted novelty 4.0

    Encoder-decoder manifold alignment framework to achieve exact idempotency in generative models.

  7. RIVET: Robust Idempotent Voice Attribute Editing

    cs.SD 2026-06 unverdicted novelty 4.0

    RIVET enforces an idempotency objective during training of voice attribute editing models to improve robustness to noisy labels, outperforming standard training on controlled noise and the GLOBE dataset.