Pith. sign in

REVIEW 5 cited by

InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.03411 v1 pith:JFKSGOQ2 submitted 2023-04-06 cs.CV

classification cs.CV
keywords conceptimagetest-timefinetuningimageslearnmodelpre-trained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in personalized image generation allow a pre-trained text-to-image model to learn a new concept from a set of images. However, existing personalization approaches usually require heavy test-time finetuning for each concept, which is time-consuming and difficult to scale. We propose InstantBooth, a novel approach built upon pre-trained text-to-image models that enables instant text-guided image personalization without any test-time finetuning. We achieve this with several major components. First, we learn the general concept of the input images by converting them to a textual token with a learnable image encoder. Second, to keep the fine details of the identity, we learn rich visual feature representation by introducing a few adapter layers to the pre-trained model. We train our components only on text-image pairs without using paired images of the same concept. Compared to test-time finetuning-based methods like DreamBooth and Textual-Inversion, our model can generate competitive results on unseen concepts concerning language-image alignment, image fidelity, and identity preservation while being 100 times faster.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

    cs.CV 2023-07 unverdicted novelty 7.0 of 10

    A single motion module trained on videos adds temporally coherent animation to any personalized text-to-image model derived from the same base without additional tuning.

  2. Equilibrated Diffusion: Frequency-aware Textual Embedding for Equilibrated Image Customization

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Equilibrated Diffusion decomposes concepts in frequency space to independently optimize subject and style embeddings, plus mask-guided diffusion and residual reference attention, for improved subject fidelity and text...

  3. R^2MoE: Redundancy-Removal Mixture of Experts for Lifelong Concept Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    R2MoE adds per-concept LoRA experts with routing distillation and expert pruning, reporting 0.19% forgetting and 15.2M added parameters on CustomConcept101.

  4. VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

    cs.CV 2023-10 unverdicted novelty 6.0 of 10

    Open-source text-to-video and image-to-video diffusion models generate high-quality 1024x576 videos, with the I2V variant claimed as the first to strictly preserve reference image content.

  5. SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation

    cs.CV 2024-11 unverdicted novelty 5.0 of 10

    SOW uses MLLMs and attention to selectively control unidirectional diffusion for pixel-level fidelity and contextual coherence in text-vision-to-image tasks.

Pith tools