REVIEW 5 cited by
HyperDreamBooth: HyperNetworks for Fast Personalization of Text-to-Image Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Personalization has emerged as a prominent aspect within the field of generative AI, enabling the synthesis of individuals in diverse contexts and styles, while retaining high-fidelity to their identities. However, the process of personalization presents inherent challenges in terms of time and memory requirements. Fine-tuning each personalized model needs considerable GPU time investment, and storing a personalized model per subject can be demanding in terms of storage capacity. To overcome these challenges, we propose HyperDreamBooth - a hypernetwork capable of efficiently generating a small set of personalized weights from a single image of a person. By composing these weights into the diffusion model, coupled with fast finetuning, HyperDreamBooth can generate a person's face in various contexts and styles, with high subject details while also preserving the model's crucial knowledge of diverse styles and semantic modifications. Our method achieves personalization on faces in roughly 20 seconds, 25x faster than DreamBooth and 125x faster than Textual Inversion, using as few as one reference image, with the same quality and style diversity as DreamBooth. Also our method yields a model that is 10,000x smaller than a normal DreamBooth model. Project page: https://hyperdreambooth.github.io
Forward citations
Cited by 5 Pith papers
-
Open Models, Open Risks: Measuring Unsafe Generation in Text-to-Image Models In the Wild
On 200+ in-the-wild T2I models, Advanced ASR shows detector-only jailbreak rates overestimate practical unsafe generation, safety is heterogeneous, and high-risk models include both explicit NSFW and seemingly benign ...
-
Stable-Hair v2: Real-World Hair Transfer via Multiple-View Diffusion Model
A multi-view diffusion hair-transfer system transfers a reference hairstyle onto a portrait and renders the edited person from many consistent viewpoints.
-
HyperCLIP: Adapting Vision-Language models with Hypernetworks
HyperCLIP trains a hypernetwork to generate task-specific normalization parameters for a small image encoder, improving zero-shot accuracy over SigLIP baselines on several benchmarks.
-
Omni-ID: Holistic Identity Representation Designed for Generative Tasks
Omni-ID is a fixed-size, multi-view face representation trained with few-to-many reconstruction that reports higher identity preservation than ArcFace and CLIP in face generation and personalized text-to-image tasks.
-
Bridging Rendering and Generative Modeling with Monte Carlo Transport Scheduling
A common variance-time SDE aligns Monte Carlo rendering noise with diffusion-model denoising, enabling low-spp render refinement and stage-ordered material control.
Discussion (0). Continue with ORCID to comment.