Fine-grained prompt decomposition, self-judgment, and localized refinement plus step-level GRPO improve compositional text-to-image generation with unified MLLMs.
Following the generation, the model enters a reflection phase where we extract its internalSelf-Judge(Yes/No) and Self-Feedback(Rationale) logs
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
FiRe: Fine-grained Multimodal Reasoning for Enhanced Image Generation
Fine-grained prompt decomposition, self-judgment, and localized refinement plus step-level GRPO improve compositional text-to-image generation with unified MLLMs.