REVIEW 14 cited by
MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Despite the success in large-scale text-to-image generation and text-conditioned image editing, existing methods still struggle to produce consistent generation and editing results. For example, generation approaches usually fail to synthesize multiple images of the same objects/characters but with different views or poses. Meanwhile, existing editing methods either fail to achieve effective complex non-rigid editing while maintaining the overall textures and identity, or require time-consuming fine-tuning to capture the image-specific appearance. In this paper, we develop MasaCtrl, a tuning-free method to achieve consistent image generation and complex non-rigid image editing simultaneously. Specifically, MasaCtrl converts existing self-attention in diffusion models into mutual self-attention, so that it can query correlated local contents and textures from source images for consistency. To further alleviate the query confusion between foreground and background, we propose a mask-guided mutual self-attention strategy, where the mask can be easily extracted from the cross-attention maps. Extensive experiments show that the proposed MasaCtrl can produce impressive results in both consistent image generation and complex non-rigid real image editing.
Forward citations
Cited by 14 Pith papers
-
D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples
Mask-guided self-attention fusion creates well-aligned target images that stay visually close to poorly-aligned base images, with full denoising trajectories, and DPO on these pairs improves alignment.
-
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.
-
Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional Warping
Drag-based editing becomes pixel-space bidirectional warping plus inpainting, giving real-time previews and 0.3s final edits at 512x512.
-
X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention
X-NeMo trains a 1D identity-agnostic motion descriptor end-to-end with a diffusion model, enabling zero-shot portrait animation with improved identity and expression fidelity.
-
MARBLE: Material Recomposition and Blending in CLIP-Space
MARBLE performs material blending and parametric material-attribute control by manipulating CLIP image embeddings and injecting them into a specific U-Net block of a pre-trained diffusion model.
-
BrushEdit: All-In-One Image Inpainting and Editing
BrushEdit couples a multimodal language model and an object detector with a single arbitrary-mask inpainting model to turn free-form text instructions into interactive, multi-turn image edits.
-
GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction
A geometry-aware, training-free inference framework that refines pretrained video diffusion predictions with projected static history content and view-conditioned routing achieves fifth place on AI City Challenge Track 5.
-
Bridging Rendering and Generative Modeling with Monte Carlo Transport Scheduling
A common variance-time SDE aligns Monte Carlo rendering noise with diffusion-model denoising, enabling low-spp render refinement and stage-ordered material control.
-
FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow Fields
FlowDrag combines 3D mesh deformation with diffusion-based drag editing, using the resulting 2D vector flow to steer the denoising process, and adds a ground-truth benchmark built from video frames.
-
X-Dyna: Expressive Dynamic Human Image Animation
A diffusion-based pipeline that animates a single human image with pose, expression, and dynamic background effects from a driving video, outperforming prior methods on dynamic detail metrics.
-
Z-STAR+: A Zero-shot Style Transfer Method via Adjusting Style Distribution
A training-free style transfer method that fuses content and style latent features in Stable Diffusion via cross-attention reweighting and a scaled adaptive instance normalization.
-
Test-time Conditional Text-to-Image Synthesis Using Diffusion Models
TINTIN conditions Stable Diffusion outputs at test time on color palettes and edge maps by backpropagating losses between decoded images and the condition through the denoising steps.
-
DCI: Dual-Conditional Inversion for Boosting Diffusion-Based Image Editing
DCI combines reference-guided noise correction with fixed-point latent refinement and reports state-of-the-art reconstruction and editing metrics on PIE-Bench.
-
Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion
GE-Adapter combines a temporal smoothness loss, bilateral-filtered DDIM inversion, and shared plus frame-specific prompt tokens to improve text-to-video editing, though the reported evidence is inconsistent.
Discussion (0). Continue with ORCID to comment.