Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:21:22.544386Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2505.10046.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:21:22.544386Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T20:44:20.190064Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-28T20:52:37.899434Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 84166abf-e733-4b08-8131-92ecb78713f0 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38107a9c-d4ce-48ae-80e2-1ac0682d6741 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Improving image generation with bet- ter captions.https://cdn
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 099dd9c6-6957-48f1-8e80-d41b57328739 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis PaliGemma: A versatile 3B VLM for transfer
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6924cf41-b96d-4d04-b941-de9e9bf73817 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1d980851-702e-4225-8b0b-a89b2b9e4e75 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Pixart-σ: Weak-to-strong training of diffu- sion transformer for 4k text-to-image generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5d32a4a-2b8f-4010-8297-357af72623d9 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Pixart-α: Fast training of diffusion trans- former for photorealistic text-to-image synthesis
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d1e53b3-7ec3-4164-a779-abcef857e1f3 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Dreamllm: Synergistic multimodal com- prehension and creation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 18600dbf-4d7b-4c32-9c3e-50e178a4a01c · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis An image is worth 16x16 words: Transformers for image recognition at scale
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea9feaa5-d32c-4a32-97ab-86555854e8b1 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis conceptual-captions-cc12m- llavanext.https://huggingface.co/datasets/ CaptionEmporium / conceptual - captions - cc12m-llavanext, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a2d01e66-b909-4997-8f7d-94bdd26dd061 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Scaling recti- fied flow transformers for high-resolution image synthesis
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5d519a93-288e-437c-85e1-b0880a5f1302 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c1b25db-b937-4edf-a7e1-d762c5b8e6bb · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87813c34-a662-4e5e-935a-201f7f17690f · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9daa17b9-ba47-4df7-8437-507b94b4de11 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cba3a9ee-932d-4840-977a-abba7820078c · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4628444-0840-4eb3-8b39-5cfed74b87b6 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Segment any- thing
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d5f0aabc-8081-4ed7-8bc5-be63e2764acb · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Announcing black forest labs.https: / / blackforestlabs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ed48c7a0-a327-49ce-ad36-75df7b37c629 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b08194e-ee6d-4bc4-b43b-b3e109e38284 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bef1506-ddd8-453f-ae31-edce49ed054c · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d548729c-5054-4608-bfba-5b440b84d94b · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9d11945-1a16-4623-beba-1900f437cd1b · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a13fcdd-d4c0-417e-abf1-64ce1ff25263 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67fc0407-e9f6-4173-a7b4-8f0d5506609f · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Decoupled weight decay regularization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a07767c5-ac48-40b4-a783-d8c948b9a28e · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a79a9207-d8ca-48fc-9664-19cfb0e9705f · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13c44ca0-43f1-49fb-a9ff-afb09bda4419 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Gemma: Open Models Based on Gemini Research and Technology
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74ab8f56-896d-436d-b5a8-245951ea52ae · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Kosmos-g: Generating images in context with multimodal large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 12de4571-f47f-4484-a08e-f83f1b68a468 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Scalable diffusion models with transformers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 41e8791c-670a-47a2-a30a-b81753ade952 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaae57d3-25d5-4324-9ede-01b31b3f3f7f · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Learning transferable visual models from natural language supervision
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fdd8e30c-961c-4114-8a27-3cbf787658ba · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1c29957c-fcc7-458d-97c1-17ac02b26bdf · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 766d0d3c-8087-4f3b-a042-c97d9b6bed46 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Gemma 2: Improving Open Language Models at a Practical Size
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6d5c7ef-ac60-426d-a695-ce10885587f4 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis High-resolution image syn- thesis with latent diffusion models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fce47d88-2f5d-4a7a-a3f7-620ede3bba3a · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis U-net: Convolutional networks for biomedical image segmentation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 770ff9e0-87c8-4cb8-9e05-40f25708c8ea · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04671a1f-5f16-4c92-8f25-e16a2ac660be · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis LMFusion: Adapting Pretrained Language Models for Multimodal Generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32912298-69e7-4213-9abf-079c51bc419b · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Layer by Layer: Uncovering Hidden Representations in Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0726554d-c66b-4672-a230-a6fb5c664a3c · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4e2894ff-197f-47f8-84d4-4fca02de2027 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Journeydb: A benchmark for generative im- age understanding
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3b2636fd-e03b-4e50-8807-cb94f2c15be9 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Is noise conditioning necessary for denoising genera- tive models?arXiv:2502.13129, 2025
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab2af98-7c52-410b-9a57-5542dc66eef5 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c80d04-10c2-482d-b653-eea19c028557 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f92cb59-87ca-4751-bb31-ecba7051fa41 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Self-correcting llm-controlled diffusion models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 349b1ff8-2fd8-4e95-b63e-8894acf49083 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis OmniGen: Unified Image Generation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d3a0de2-db54-4969-b304-7f068e3226fa · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccb77f9a-1191-4d2b-b06b-5e21026ca24c · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ccefb93-3e28-44bd-934f-c88b21e75d7b · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Mastering text-to-image diffu- sion: Recaptioning, planning, and generating with multi- modal llms
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 24a5858c-0b98-4d62-94c8-41a92f6c5dd5 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c23bef9-2e55-45cc-88e1-922db9ce5157 · outbound
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20536b70-f7f7-486f-aba4-ca14b5b728e9 · inbound
Exploring Time Conditioning in Diffusion Generative Models from Disjoint Noisy Data Manifolds Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1453aca1-62fc-4390-8b98-9d0544cdf4a0 · inbound
UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.