Pith. sign in

REVIEW 13 cited by

PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.04461 v1 pith:MIK252VP submitted 2023-12-07 cs.CV cs.AIcs.LGcs.MM

classification cs.CVcs.AIcs.LGcs.MM
keywords generationphotomakerembeddingapplicationscharacteristicsdatahumaninput
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in text-to-image generation have made remarkable progress in synthesizing realistic human photos conditioned on given text prompts. However, existing personalized generation methods cannot simultaneously satisfy the requirements of high efficiency, promising identity (ID) fidelity, and flexible text controllability. In this work, we introduce PhotoMaker, an efficient personalized text-to-image generation method, which mainly encodes an arbitrary number of input ID images into a stack ID embedding for preserving ID information. Such an embedding, serving as a unified ID representation, can not only encapsulate the characteristics of the same input ID comprehensively, but also accommodate the characteristics of different IDs for subsequent integration. This paves the way for more intriguing and practically valuable applications. Besides, to drive the training of our PhotoMaker, we propose an ID-oriented data construction pipeline to assemble the training data. Under the nourishment of the dataset constructed through the proposed pipeline, our PhotoMaker demonstrates better ID preservation ability than test-time fine-tuning based methods, yet provides significant speed improvements, high-quality generation results, strong generalization capabilities, and a wide range of applications. Our project page is available at https://photo-maker.github.io/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Interact-Custom: Customized Human Object Interaction Image Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Interact-Custom generates customized human-object interaction images by first generating a foreground mask from the prompt and then using that mask to guide identity-preserving diffusion generation.

  2. Robust ID-Specific Face Restoration via Alignment Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RIDFR injects a reference person's identity into diffusion-based face restoration and uses Alignment Learning across multiple same-identity references to suppress pose, expression, and makeup interference.

  3. One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Concatenating all frame prompts into a single prompt, then reweighting singular values and re-anchoring cross-attention, yields training-free identity-consistent text-to-image generation.

  4. Arc2Avatar: Generating Expressive 3D Avatars from a Single Image via ID Guidance

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Arc2Avatar generates expressive 3D head avatars from a single image by distilling a LoRA-fine-tuned Arc2Face model into 3D Gaussian splats anchored to a FLAME mesh, enabling blendshape expressions.

  5. FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles

    cs.SD 2025-01 conditional novelty 6.0 of 10

    FaceSpeak generates speech from arbitrary-style portraits by learning separate identity and emotion embeddings from face images, with a new generated multi-style TTS dataset.

  6. MagicNaming: Consistent Identity Generation by Finding a "Name Space" in T2I Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An image encoder maps any face to a 'name embedding' that, when prepended to a text prompt, makes an SDXL model generate consistent identities for arbitrary people without fine-tuning.

  7. CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A data-curation engine plus a token-order attention injection module raises spatial accuracy of Stable Diffusion and FLUX models on standard benchmarks.

  8. F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and Restoration

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FaceQ, a new 12K-image benchmark with multi-dimensional human preference scores, reveals that existing quality metrics poorly match human judgment on AI-generated faces, and F-Eval, an instruction-tuned LMM, outperforms them.

  9. Omni-ID: Holistic Identity Representation Designed for Generative Tasks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Omni-ID is a fixed-size, multi-view face representation trained with few-to-many reconstruction that reports higher identity preservation than ArcFace and CLIP in face generation and personalized text-to-image tasks.

  10. PersonaCraft: Personalized and Controllable Full-Body Multi-Human Scene Generation Using Occlusion-Aware 3D-Conditioned Diffusion

    cs.CV 2024-11 conditional novelty 6.0 of 10

    PersonaCraft adds SMPLx depth and normal conditioning, occlusion boundary enhancement, and occlusion-aware classifier-free guidance to diffusion models, enabling controllable multi-person images that preserve both fac...

  11. Try-On-Adapter: A Simple and Flexible Try-On Paradigm

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A diffusion-based adapter performs virtual try-on as outpainting from a reference face and garment, reporting FID 5.56 and 7.23 on VITON-HD.

  12. Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A tuning-free multi-concept video personalization method that uses anchored prompt tokens and per-reference concept embeddings to prevent identity blending.

  13. Face2QR: A Unified Framework for Aesthetic, Face-Preserving, and Scannable QR Code Generation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Face2QR generates scannable, aesthetic QR codes that preserve face identity by integrating face and QR control in a Stable Diffusion pipeline, reshuffling QR modules, and optimizing in latent space.

Pith tools