REVIEW 13 cited by
PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent advances in text-to-image generation have made remarkable progress in synthesizing realistic human photos conditioned on given text prompts. However, existing personalized generation methods cannot simultaneously satisfy the requirements of high efficiency, promising identity (ID) fidelity, and flexible text controllability. In this work, we introduce PhotoMaker, an efficient personalized text-to-image generation method, which mainly encodes an arbitrary number of input ID images into a stack ID embedding for preserving ID information. Such an embedding, serving as a unified ID representation, can not only encapsulate the characteristics of the same input ID comprehensively, but also accommodate the characteristics of different IDs for subsequent integration. This paves the way for more intriguing and practically valuable applications. Besides, to drive the training of our PhotoMaker, we propose an ID-oriented data construction pipeline to assemble the training data. Under the nourishment of the dataset constructed through the proposed pipeline, our PhotoMaker demonstrates better ID preservation ability than test-time fine-tuning based methods, yet provides significant speed improvements, high-quality generation results, strong generalization capabilities, and a wide range of applications. Our project page is available at https://photo-maker.github.io/
Forward citations
Cited by 13 Pith papers
-
Interact-Custom: Customized Human Object Interaction Image Generation
Interact-Custom generates customized human-object interaction images by first generating a foreground mask from the prompt and then using that mask to guide identity-preserving diffusion generation.
-
Robust ID-Specific Face Restoration via Alignment Learning
RIDFR injects a reference person's identity into diffusion-based face restoration and uses Alignment Learning across multiple same-identity references to suppress pose, expression, and makeup interference.
-
One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
Concatenating all frame prompts into a single prompt, then reweighting singular values and re-anchoring cross-attention, yields training-free identity-consistent text-to-image generation.
-
Arc2Avatar: Generating Expressive 3D Avatars from a Single Image via ID Guidance
Arc2Avatar generates expressive 3D head avatars from a single image by distilling a LoRA-fine-tuned Arc2Face model into 3D Gaussian splats anchored to a FLAME mesh, enabling blendshape expressions.
-
FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles
FaceSpeak generates speech from arbitrary-style portraits by learning separate identity and emotion embeddings from face images, with a new generated multi-style TTS dataset.
-
MagicNaming: Consistent Identity Generation by Finding a "Name Space" in T2I Diffusion Models
An image encoder maps any face to a 'name embedding' that, when prepended to a text prompt, makes an SDXL model generate consistent identities for arbitrary people without fine-tuning.
-
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models
A data-curation engine plus a token-order attention injection module raises spatial accuracy of Stable Diffusion and FLUX models on standard benchmarks.
-
F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and Restoration
FaceQ, a new 12K-image benchmark with multi-dimensional human preference scores, reveals that existing quality metrics poorly match human judgment on AI-generated faces, and F-Eval, an instruction-tuned LMM, outperforms them.
-
Omni-ID: Holistic Identity Representation Designed for Generative Tasks
Omni-ID is a fixed-size, multi-view face representation trained with few-to-many reconstruction that reports higher identity preservation than ArcFace and CLIP in face generation and personalized text-to-image tasks.
-
PersonaCraft: Personalized and Controllable Full-Body Multi-Human Scene Generation Using Occlusion-Aware 3D-Conditioned Diffusion
PersonaCraft adds SMPLx depth and normal conditioning, occlusion boundary enhancement, and occlusion-aware classifier-free guidance to diffusion models, enabling controllable multi-person images that preserve both fac...
-
Try-On-Adapter: A Simple and Flexible Try-On Paradigm
A diffusion-based adapter performs virtual try-on as outpainting from a reference face and garment, reporting FID 5.56 and 7.23 on VITON-HD.
-
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
A tuning-free multi-concept video personalization method that uses anchored prompt tokens and per-reference concept embeddings to prevent identity blending.
-
Face2QR: A Unified Framework for Aesthetic, Face-Preserving, and Scannable QR Code Generation
Face2QR generates scannable, aesthetic QR codes that preserve face identity by integrating face and QR control in a Stable Diffusion pipeline, reshuffling QR modules, and optimizing in latent space.
Discussion (0). Continue with ORCID to comment.