Pith. sign in

REVIEW 2 cited by

MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.11565 v2 pith:URC7RXO3 submitted 2024-04-17 cs.CV cs.AIcs.GR

classification cs.CVcs.AIcs.GR
keywords personalizedbranchpriorgenerationmixture-of-attentionmodelattentioncreation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce a new architecture for personalization of text-to-image diffusion models, coined Mixture-of-Attention (MoA). Inspired by the Mixture-of-Experts mechanism utilized in large language models (LLMs), MoA distributes the generation workload between two attention pathways: a personalized branch and a non-personalized prior branch. MoA is designed to retain the original model's prior by fixing its attention layers in the prior branch, while minimally intervening in the generation process with the personalized branch that learns to embed subjects in the layout and context generated by the prior branch. A novel routing mechanism manages the distribution of pixels in each layer across these branches to optimize the blend of personalized and generic content creation. Once trained, MoA facilitates the creation of high-quality, personalized images featuring multiple subjects with compositions and interactions as diverse as those generated by the original model. Crucially, MoA enhances the distinction between the model's pre-existing capability and the newly augmented personalized intervention, thereby offering a more disentangled subject-context control that was previously unattainable. Project page: https://snap-research.github.io/mixture-of-attention

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MultiAnimate: A Unified Framework for Controllable Multi-Character Animation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A diffusion-based framework that animates multiple characters in one scene from separate reference images and pose sequences while preserving each character's identity.

  2. Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA

    cs.GR 2025-07 conditional novelty 6.0 of 10

    A grid-based LoRA training scheme lets a text-to-video model personalize previously unseen dynamic concepts, subject appearance plus motion, in one feedforward pass with no per-video fine-tuning.

Pith tools