Pith. sign in

REVIEW 11 cited by

PhotoVerse: Tuning-Free Image Customization with Text-to-Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.05793 v1 pith:ANEKZGEG submitted 2023-09-11 cs.CV cs.AI

classification cs.CVcs.AI
keywords identityimageimagesgenerationphotoverseapproacheditabilityfacial
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Personalized text-to-image generation has emerged as a powerful and sought-after tool, empowering users to create customized images based on their specific concepts and prompts. However, existing approaches to personalization encounter multiple challenges, including long tuning times, large storage requirements, the necessity for multiple input images per identity, and limitations in preserving identity and editability. To address these obstacles, we present PhotoVerse, an innovative methodology that incorporates a dual-branch conditioning mechanism in both text and image domains, providing effective control over the image generation process. Furthermore, we introduce facial identity loss as a novel component to enhance the preservation of identity during training. Remarkably, our proposed PhotoVerse eliminates the need for test time tuning and relies solely on a single facial photo of the target identity, significantly reducing the resource cost associated with image generation. After a single training phase, our approach enables generating high-quality images within only a few seconds. Moreover, our method can produce diverse images that encompass various scenes and styles. The extensive evaluation demonstrates the superior performance of our approach, which achieves the dual objectives of preserving identity and facilitating editability. Project page: https://photoverse2d.github.io/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A benchmark and five-million-clip dataset for evaluating and training subject-to-video generation models, with three new metrics for subject consistency, naturalness, and text alignment.

  2. FaceCrafter: Identity-Conditional Diffusion with Disentangled Control over Facial Pose, Expression, and Emotion

    cs.CV 2025-05 conditional novelty 6.0 of 10

    FaceCrafter adds two lightweight cross-attention control modules and an attention disentanglement loss to Arc2Face, achieving more accurate control of facial pose, expression, and emotion with far fewer extra paramete...

  3. BridgeIV: Bridging Customized Image and Video Generation through Test-Time Autoregressive Identity Propagation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    BridgeIV improves subject consistency in customized text-to-video generation by warping attention maps and self-attention values across frames, then refining latents with a CLIP-based reward.

  4. Turn That Frown Upside Down: FaceID Customization via Cross-Training Data

    cs.CV 2025-01 conditional novelty 6.0 of 10

    CrossFaceID, a 40,000-image dataset of controlled same-person facial variations, improves face-identity customization when used to fine-tune IP-Adapter and InstantID models.

  5. ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions

    cs.CV 2025-01 reject novelty 6.0 of 10

    ComposeAnyone generates human images by conditioning a diffusion model on hand-drawn color-block layouts together with decoupled text or reference-image descriptions for each body part.

  6. MagicNaming: Consistent Identity Generation by Finding a "Name Space" in T2I Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An image encoder maps any face to a 'name embedding' that, when prepended to a text prompt, makes an SDXL model generate consistent identities for arbitrary people without fine-tuning.

  7. Identity-Preserving Text-to-Video Generation by Frequency Decomposition

    cs.CV 2024-11 conditional novelty 6.0 of 10

    ConsisID generates identity-preserving videos by injecting low-frequency facial features into shallow layers and high-frequency identity features into attention blocks of a DiT video model.

  8. Vec2Face+ for Face Dataset Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A synthetic face dataset with 4M to 12M images trains a matcher whose average accuracy on five benchmarks is 0.09 to 0.14 points higher than CASIA-WebFace, while twin verification and bias remain unsolved.

  9. XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    XVerse learns token-specific offsets that modify the text-stream modulation of a diffusion transformer, enabling multi-subject identity and attribute control in image generation.

  10. Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A tuning-free multi-concept video personalization method that uses anchored prompt tokens and per-reference concept embeddings to prevent identity blending.

  11. Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era

    cs.CV 2024-11 unverdicted novelty 3.0 of 10

    A survey that organizes over 100 instruction-guided image and multimedia editing papers into a process-based taxonomy, with an emphasis on LLM and MLLM empowered methods.

Pith tools