Pith. sign in

REVIEW 14 cited by

MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.08465 v1 pith:R2ZAVY6T submitted 2023-04-17 cs.CV

classification cs.CV
keywords editingimagegenerationconsistentmasactrlself-attentioncomplexexisting
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Despite the success in large-scale text-to-image generation and text-conditioned image editing, existing methods still struggle to produce consistent generation and editing results. For example, generation approaches usually fail to synthesize multiple images of the same objects/characters but with different views or poses. Meanwhile, existing editing methods either fail to achieve effective complex non-rigid editing while maintaining the overall textures and identity, or require time-consuming fine-tuning to capture the image-specific appearance. In this paper, we develop MasaCtrl, a tuning-free method to achieve consistent image generation and complex non-rigid image editing simultaneously. Specifically, MasaCtrl converts existing self-attention in diffusion models into mutual self-attention, so that it can query correlated local contents and textures from source images for consistency. To further alleviate the query confusion between foreground and background, we propose a mask-guided mutual self-attention strategy, where the mask can be easily extracted from the cross-attention maps. Extensive experiments show that the proposed MasaCtrl can produce impressive results in both consistent image generation and complex non-rigid real image editing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples

    cs.CV 2025-05 conditional novelty 7.0 of 10

    Mask-guided self-attention fusion creates well-aligned target images that stay visually close to poorly-aligned base images, with full denoising trajectories, and DPO on these pairs improves alignment.

  2. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  3. Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional Warping

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Drag-based editing becomes pixel-space bidirectional warping plus inpainting, giving real-time previews and 0.3s final edits at 512x512.

  4. X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention

    cs.CV 2025-07 conditional novelty 6.0 of 10

    X-NeMo trains a 1D identity-agnostic motion descriptor end-to-end with a diffusion model, enabling zero-shot portrait animation with improved identity and expression fidelity.

  5. MARBLE: Material Recomposition and Blending in CLIP-Space

    cs.CV 2025-06 conditional novelty 6.0 of 10

    MARBLE performs material blending and parametric material-attribute control by manipulating CLIP image embeddings and injecting them into a specific U-Net block of a pre-trained diffusion model.

  6. BrushEdit: All-In-One Image Inpainting and Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    BrushEdit couples a multimodal language model and an object detector with a single arbitrary-mask inpainting model to turn free-form text instructions into interactive, multi-turn image edits.

  7. GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction

    cs.CV 2026-08 conditional novelty 5.0 of 10

    A geometry-aware, training-free inference framework that refines pretrained video diffusion predictions with projected static history content and view-conditioned routing achieves fifth place on AI City Challenge Track 5.

  8. Bridging Rendering and Generative Modeling with Monte Carlo Transport Scheduling

    cs.CV 2026-02 reject novelty 5.0 of 10

    A common variance-time SDE aligns Monte Carlo rendering noise with diffusion-model denoising, enabling low-spp render refinement and stage-ordered material control.

  9. FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow Fields

    cs.GR 2025-07 conditional novelty 5.0 of 10

    FlowDrag combines 3D mesh deformation with diffusion-based drag editing, using the resulting 2D vector flow to steer the denoising process, and adds a ground-truth benchmark built from video frames.

  10. X-Dyna: Expressive Dynamic Human Image Animation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A diffusion-based pipeline that animates a single human image with pose, expression, and dynamic background effects from a driving video, outperforming prior methods on dynamic detail metrics.

  11. Z-STAR+: A Zero-shot Style Transfer Method via Adjusting Style Distribution

    cs.CV 2024-11 reject novelty 5.0 of 10

    A training-free style transfer method that fuses content and style latent features in Stable Diffusion via cross-attention reweighting and a scaled adaptive instance normalization.

  12. Test-time Conditional Text-to-Image Synthesis Using Diffusion Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    TINTIN conditions Stable Diffusion outputs at test time on color palettes and edge maps by backpropagating losses between decoded images and the condition through the denoising steps.

  13. DCI: Dual-Conditional Inversion for Boosting Diffusion-Based Image Editing

    cs.CV 2025-06 reject novelty 4.0 of 10

    DCI combines reference-guided noise correction with fixed-point latent refinement and reports state-of-the-art reconstruction and editing metrics on PIE-Bench.

  14. Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion

    cs.CV 2025-01 reject novelty 4.0 of 10

    GE-Adapter combines a temporal smoothness loss, bilateral-filtered DDIM inversion, and shared plus frame-specific prompt tokens to improve text-to-video editing, though the reported evidence is inconsistent.

Pith tools