Pith. sign in

REVIEW 4 cited by

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.18301 v1 pith:KI5HY24R submitted 2024-11-27 cs.CV

classification cs.CV
keywords lossambiguitysimilargenerationmodelsoverlapsubjectstext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Representing the cutting-edge technique of text-to-image models, the latest Multimodal Diffusion Transformer (MMDiT) largely mitigates many generation issues existing in previous models. However, we discover that it still suffers from subject neglect or mixing when the input text prompt contains multiple subjects of similar semantics or appearance. We identify three possible ambiguities within the MMDiT architecture that cause this problem: Inter-block Ambiguity, Text Encoder Ambiguity, and Semantic Ambiguity. To address these issues, we propose to repair the ambiguous latent on-the-fly by test-time optimization at early denoising steps. In detail, we design three loss functions: Block Alignment Loss, Text Encoder Alignment Loss, and Overlap Loss, each tailored to mitigate these ambiguities. Despite significant improvements, we observe that semantic ambiguity persists when generating multiple similar subjects, as the guidance provided by overlap loss is not explicit enough. Therefore, we further propose Overlap Online Detection and Back-to-Start Sampling Strategy to alleviate the problem. Experimental results on a newly constructed challenging dataset of similar subjects validate the effectiveness of our approach, showing superior generation quality and much higher success rates over existing methods. Our code will be available at https://github.com/wtybest/EnMMDiT.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LaRender: Training-Free Occlusion Control in Image Generation via Latent Rendering

    cs.CV 2025-08 conditional novelty 7.0 of 10

    LaRender replaces cross-attention layers in a pretrained diffusion model with a latent alpha-compositing operation that renders object features in occlusion order, giving training-free occlusion control.

  2. RubricRL: Simple Generalizable Rewards for Text-to-Image Generation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Using an LLM to generate prompt-specific visual rubrics and grade each criterion independently gives a more interpretable reward that improves text-to-image model alignment beyond composite and learned scalar rewards.

  3. Detail++: Training-Free Detail Enhancer for T2I Diffusion Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Detail++ uses progressive multi-branch prompt injection and test-time attention optimization to improve attribute binding in text-to-image generation.

  4. Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SaaS, a self-adaptive attention-scaling method, improves instruction-following fidelity of unified image generation models without training by boosting the cross-attention activation of each sub-instruction in regions...

Pith tools