Deep fusion of a frozen LLM with a DiT improves text-image alignment over shallow fusion baselines, and a scaled recipe (FuseDiT) achieves competitive results despite limited data and compute.
Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
Deep fusion of a frozen LLM with a DiT improves text-image alignment over shallow fusion baselines, and a scaled recipe (FuseDiT) achieves competitive results despite limited data and compute.