Pith. sign in

REVIEW 7 cited by

MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.11565 v2 pith:URC7RXO3 submitted 2024-04-17 cs.CV cs.AIcs.GR

classification cs.CVcs.AIcs.GR
keywords personalizedbranchpriorgenerationmixture-of-attentionmodelattentioncreation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce a new architecture for personalization of text-to-image diffusion models, coined Mixture-of-Attention (MoA). Inspired by the Mixture-of-Experts mechanism utilized in large language models (LLMs), MoA distributes the generation workload between two attention pathways: a personalized branch and a non-personalized prior branch. MoA is designed to retain the original model's prior by fixing its attention layers in the prior branch, while minimally intervening in the generation process with the personalized branch that learns to embed subjects in the layout and context generated by the prior branch. A novel routing mechanism manages the distribution of pixels in each layer across these branches to optimize the blend of personalized and generic content creation. Once trained, MoA facilitates the creation of high-quality, personalized images featuring multiple subjects with compositions and interactions as diverse as those generated by the original model. Crucially, MoA enhances the distinction between the model's pre-existing capability and the newly augmented personalized intervention, thereby offering a more disentangled subject-context control that was previously unattainable. Project page: https://snap-research.github.io/mixture-of-attention

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MultiAnimate: A Unified Framework for Controllable Multi-Character Animation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A diffusion-based framework that animates multiple characters in one scene from separate reference images and pose sequences while preserving each character's identity.

  2. Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA

    cs.GR 2025-07 conditional novelty 6.0 of 10

    A grid-based LoRA training scheme lets a text-to-video model personalize previously unseen dynamic concepts, subject appearance plus motion, in one feedforward pass with no per-video fine-tuning.

  3. IDProtector: An Adversarial Noise Encoder to Protect Against ID-Preserving Image Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    IDProtector adds imperceptible adversarial noise to a portrait in a single forward pass, disrupting identity-preserving generation by InstantID, IP-Adapter, IP-Adapter-Plus, and PhotoMaker.

  4. Omni-ID: Holistic Identity Representation Designed for Generative Tasks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Omni-ID is a fixed-size, multi-view face representation trained with few-to-many reconstruction that reports higher identity preservation than ArcFace and CLIP in face generation and personalized text-to-image tasks.

  5. DECOR:Decomposition and Projection of Text Embeddings for Text-to-Image Customization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DECOR suppresses undesired word-token semantics in text embeddings via orthogonal projection, reducing prompt misalignment and content leakage in LoRA-customized text-to-image models.

  6. Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A tuning-free multi-concept video personalization method that uses anchored prompt tokens and per-reference concept embeddings to prevent identity blending.

  7. Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook

    cs.CV 2024-11 conditional novelty 4.0 of 10

    The paper introduces BioDeepAV, a benchmark of real and fake talking-face videos, and reports that state-of-the-art deepfake detectors drop sharply when tested on deepfakes from unseen generators.

Pith tools