Pith. sign in

REVIEW 10 cited by

MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.08465 v1 pith:R2ZAVY6T submitted 2023-04-17 cs.CV

MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing

classification cs.CV
keywords editingimagegenerationconsistentmasactrlself-attentioncomplexexisting
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Despite the success in large-scale text-to-image generation and text-conditioned image editing, existing methods still struggle to produce consistent generation and editing results. For example, generation approaches usually fail to synthesize multiple images of the same objects/characters but with different views or poses. Meanwhile, existing editing methods either fail to achieve effective complex non-rigid editing while maintaining the overall textures and identity, or require time-consuming fine-tuning to capture the image-specific appearance. In this paper, we develop MasaCtrl, a tuning-free method to achieve consistent image generation and complex non-rigid image editing simultaneously. Specifically, MasaCtrl converts existing self-attention in diffusion models into mutual self-attention, so that it can query correlated local contents and textures from source images for consistency. To further alleviate the query confusion between foreground and background, we propose a mask-guided mutual self-attention strategy, where the mask can be easily extracted from the cross-attention maps. Extensive experiments show that the proposed MasaCtrl can produce impressive results in both consistent image generation and complex non-rigid real image editing.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Keep The Essentials: Efficient Reference Conditioned Generation via Token Dropping

    cs.CV 2026-06 unverdicted novelty 7.0

    Sparse Context achieves 2-4x faster inference in reference-conditioned diffusion models by fine-tuning with random token dropping and applying task-aware selection at inference time, without loss of visual quality.

  2. Semantic Browsing: Controllable Diversity for Image Generation

    cs.CV 2026-06 unverdicted novelty 7.0

    A technique for controllable diversity in text-to-image generation by inducing structured semantic variations at the prompt level via VLM and agentic workflow.

  3. ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

    cs.CV 2024-03 unverdicted novelty 7.0

    ELLA introduces a timestep-aware semantic connector to link LLMs with diffusion models for improved dense prompt following, validated on a new 1K-prompt benchmark.

  4. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  5. Cross-Sensor SAR Data Generation Using Diffusion Models and Feature Migration

    eess.IV 2026-06 unverdicted novelty 6.0

    A diffusion-based data generation framework with attention distillation transfers sensor-specific features to synthesize training data for cross-sensor SAR applications on aircraft targets.

  6. ReAge3D: Re-Aging 3D Faces with View Consistency

    cs.CV 2026-06 unverdicted novelty 6.0

    ReAge3D trains a diffusion re-aging model on synthetic pairs then uses masked propagation from a frontal pivot view to produce consistent multi-view images that supervise 3D face optimization.

  7. InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation

    cs.CV 2026-04 unverdicted novelty 6.0

    InsEdit adapts a video diffusion backbone for text-instruction video editing via Mutual Context Attention, achieving SOTA open-source results with O(100K) data while also supporting image editing.

  8. Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional Warping

    cs.CV 2025-09 conditional novelty 6.0

    Drag-based editing becomes pixel-space bidirectional warping plus inpainting, giving real-time previews and 0.3s final edits at 512x512.

  9. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 5.0

    Hallo4D mitigates 3D/4D generation hallucinations via LMM-based detection, multi-model voting correction, and motion-aware optimization without retraining base generators.

  10. Bridging Rendering and Generative Modeling with Monte Carlo Transport Scheduling

    cs.CV 2026-02 reject novelty 5.0

    A common variance-time SDE aligns Monte Carlo rendering noise with diffusion-model denoising, enabling low-spp render refinement and stage-ordered material control.