Pith. sign in

REVIEW 1 cited by

InstructBooth: Instruction-following Personalized Text-to-Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03011 v2 pith:OLIQVWRQ submitted 2023-12-04 cs.CV cs.AI

classification cs.CVcs.AI
keywords text-to-imageinstructboothmodelsalignmentimage-textimagespersonalizationpersonalized
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Personalizing text-to-image models using a limited set of images for a specific object has been explored in subject-specific image generation. However, existing methods often face challenges in aligning with text prompts due to overfitting to the limited training images. In this work, we introduce InstructBooth, a novel method designed to enhance image-text alignment in personalized text-to-image models without sacrificing the personalization ability. Our approach first personalizes text-to-image models with a small number of subject-specific images using a unique identifier. After personalization, we fine-tune personalized text-to-image models using reinforcement learning to maximize a reward that quantifies image-text alignment. Additionally, we propose complementary techniques to increase the synergy between these two processes. Our method demonstrates superior image-text alignment compared to existing baselines, while maintaining high personalization ability. In human evaluations, InstructBooth outperforms them when considering all comprehensive factors. Our project page is at https://sites.google.com/view/instructbooth.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Memory-Efficient Personalization of Text-to-Image Diffusion Models via Selective Optimization Strategies

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A hybrid of low-resolution backpropagation and high-resolution zeroth-order optimization, scheduled by a dynamic timestep-dependent probability, matches full-resolution fine-tuning quality while cutting training memory.

Pith tools