Biased noise sampling for rectified flows combined with a bidirectional text-image transformer architecture yields state-of-the-art high-resolution text-to-image results that scale predictably with model size.
ArXiv , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
verdicts
CONDITIONAL 2representative citing papers
A learnable control block trained with GRPO to select token generation order in multimodal masked diffusion models improves text-to-image alignment and multimodal understanding over logit-based heuristics.
citing papers explorer
-
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
Biased noise sampling for rectified flows combined with a bidirectional text-image transformer architecture yields state-of-the-art high-resolution text-to-image results that scale predictably with model size.
-
Reinforcing the Generation Order of Multimodal Masked Diffusion Models
A learnable control block trained with GRPO to select token generation order in multimodal masked diffusion models improves text-to-image alignment and multimodal understanding over logit-based heuristics.