Pith. sign in

REVIEW 4 cited by

InstantFamily: Masked Attention for Zero-shot Multi-ID Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.19427 v1 pith:HO7S42VT submitted 2024-04-30 cs.CV

classification cs.CV
keywords multi-idgenerationimageimagesinstantfamilymaskedmodeladditionally
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the field of personalized image generation, the ability to create images preserving concepts has significantly improved. Creating an image that naturally integrates multiple concepts in a cohesive and visually appealing composition can indeed be challenging. This paper introduces "InstantFamily," an approach that employs a novel masked cross-attention mechanism and a multimodal embedding stack to achieve zero-shot multi-ID image generation. Our method effectively preserves ID as it utilizes global and local features from a pre-trained face recognition model integrated with text conditions. Additionally, our masked cross-attention mechanism enables the precise control of multi-ID and composition in the generated images. We demonstrate the effectiveness of InstantFamily through experiments showing its dominance in generating images with multi-ID, while resolving well-known multi-ID generation problems. Additionally, our model achieves state-of-the-art performance in both single-ID and multi-ID preservation. Furthermore, our model exhibits remarkable scalability with a greater number of ID preservation than it was originally trained with.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GroupVideo: Multi-Identity Customized Text-to-Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    GroupVideo generates multi-person videos from reference photos plus text, using multimodal identity alignment and ID localization to keep each person's identity consistent.

  2. MultiAnimate: A Unified Framework for Controllable Multi-Character Animation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A diffusion-based framework that animates multiple characters in one scene from separate reference images and pose sequences while preserving each character's identity.

  3. Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Scone unifies subject understanding and generation in a two-stage trained model to improve both composition and distinction in multi-subject image generation, outperforming prior open-source models on new benchmarks.

  4. Retrieval Augmented Comic Image Generation

    cs.CV 2025-06 reject novelty 5.0 of 10

    RaCig combines retrieval-based character assignment with regional IP-Adapter and ControlNet injection to generate comic panels with consistent characters and diverse gestures.

Pith tools