REVIEW 3 cited by
Training-free Regional Prompting for Diffusion Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Diffusion models have demonstrated excellent capabilities in text-to-image generation. Their semantic understanding (i.e., prompt following) ability has also been greatly improved with large language models (e.g., T5, Llama). However, existing models cannot perfectly handle long and complex text prompts, especially when the text prompts contain various objects with numerous attributes and interrelated spatial relationships. While many regional prompting methods have been proposed for UNet-based models (SD1.5, SDXL), but there are still no implementations based on the recent Diffusion Transformer (DiT) architecture, such as SD3 and FLUX.1.In this report, we propose and implement regional prompting for FLUX.1 based on attention manipulation, which enables DiT with fined-grained compositional text-to-image generation capability in a training-free manner. Code is available at https://github.com/antonioo-c/Regional-Prompting-FLUX.
Forward citations
Cited by 3 Pith papers
-
SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing
A single model performs instruction-guided and line-guided sketch editing by packing sketch, mask, and guidance into RGB channels, trained on a synthetic multi-step edit dataset.
-
RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
RecipeGen is a new benchmark with 26,453 recipes, 196,724 step-aligned images, and 4,491 cooking videos, plus three domain-specific evaluation metrics for recipe generation.
-
RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution
RAGSR combines region-level vision-language captions with regional attention masks to improve fine-grained detail generation in diffusion-based super-resolution.
Discussion (0). Sign in to comment.