Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T13:15:24.299457Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2605.28615.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T13:15:24.299457Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation da92fc0e-3dc2-424e-acd0-5398b05fc4f4 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization VisMin: Visual Minimal-Change Understanding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4b9540b2-3eca-41bf-8f96-0a404812eb55 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 397f5500-d34d-4245-a4b8-d14e052a5854 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Improving image generation with better captions.Computer Science
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cf8ec4d-9e15-4ce6-bd25-9519ae47445a · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models.ACMTransactionson Graphics(TOG), 42:1 – 10, 2023
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c43429ae-55f5-43cb-84d4-0878a77fe81a · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 884d61e2-7a99-4f7f-9b5c-98446b851b3a · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Gentron: Diffusion transformers for image and video generation.2024 IEEE/CVF Conferenceon ComputerVisionand PatternRecognition(CVPR), pages 6441–6451, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9e9757a-3d36-4281-a70d-0ee0cee81a3d · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization DeepSeek-V3 Technical Report
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 22c8f37c-3e86-4cd3-80a9-568ce0c8f71b · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 12e5ab3d-ad4a-4ba9-bcf9-e0935418d10a · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Scaling rectified flow transformers for high-resolution image synthesis
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7939cb22-7437-47a6-b7e2-853ff16cc20d · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Dimba: Transformer-Mamba Diffusion Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9bd29546-d6fa-46d9-9c78-ce15dbaf8ea9 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Ranni: Taming text-to-image diffusion for accurate instruction following.2024IEEE/CVFConferenceonComputerVisionandPatternRecognition(CVPR), pages 4744–4753, 2023
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ce23e08-8cb8-4e1b-a379-7d7e40aea350 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Geneval: An object-focused framework for evaluating text-to-image alignment
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0502b47-7384-41c1-8920-aca5d3635f82 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Contrafusion: Contrastively improving compositional understanding in diffusion models via fine-grained negative images
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df9c6714-1456-44b4-a6fd-4d87ef5e18a3 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c882f8b3-e18b-4f3f-8146-05a01988d4fb · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization arXiv preprint arXiv:2406.06424 , year=
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bcab43c5-1ad2-4ddd-8ff2-61ce94027cd1 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Lora: Low-rank adaptation of large language models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adbcec71-95ac-4dc7-a516-b0955fcb410e · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Ella: Equip diffusion models with llm for enhanced semantic alignment, 2024
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27dbad8e-e2ed-40b6-a4f2-e71b802a6e31 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 284535c7-d273-4ca5-a1f0-286ab423d777 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6e1111f1-bca1-4c11-bf02-1835c8669516 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization yes" or
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e90e67f-754f-46fc-97ec-e4e09aa3de87 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Scalable Ranked Preference Optimization for Text-to-Image Generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 510e5633-8fae-46cd-bcb5-89b3225bd7c7 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Flux.https://github.com/black-forest-labs/flux, 2024
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 303f0d4d-bc8c-4e45-ae45-e447a04a7648 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Aligning Text-to-Image Models using Human Feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cc80049c-57e2-42fb-8447-4cb2862d5099 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46780f46-8b9d-45b9-9b28-593598929161 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Calibrated multi-preference optimization for aligning diffusion models.2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18465–18475, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acdb1d11-57c8-4ee4-982e-01fd1f432468 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Playground v2
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb6d2f2b-4f5b-423e-9a6a-40575aedd698 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6f97e364-a2f8-43ea-bb8e-2bd33348ce8d · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Gligen: Open-set grounded text-to-image generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 557264fc-ffe9-48f0-ab1f-15a838f98369 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization LAION-SG: An Enhanced Large-Scale Dataset for Training Complex Image-Text Models with Structural Annotations
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d7917592-7b31-47fe-b35f-64250911c9eb · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Llm-grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models.Trans
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91ef9bd6-f350-468b-94b6-76255911bfef · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 42e4fffa-f157-4e67-88d8-e82c8036d534 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Aesthetic post-training diffusion models from generic preferences with step-by-step preference optimization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6e39825-1d71-4628-8429-6cbf1de93d74 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 582ff4b1-4d6c-49fb-a319-2788ecc4919d · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Eclipse: A resource-efficient text-to-imagepriorforimagegenerations
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 183dc7a8-7d43-4fb5-b305-1a7a1f153a78 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Enhancing image layout control with loss-guided diffusion models.2025IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 3916–3924, 2024
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1326b9df-2b21-4376-84c5-177c51a1bb88 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Scalable diffusion models with transformers
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ad8a84f-e736-4331-ad3c-f7c5c9d1954b · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4ac09c30-23ee-4d5c-ad04-cf8e721d43b0 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ad5b72b-5caa-46ff-96a8-145544cd971b · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization SAM 2: Segment Anything in Images and Videos
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0dfba7f4-d6ed-44cc-9586-697314419e31 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec046f94-920a-4af2-a4d9-2e965bf65b65 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Joty, and Nikhil Naik
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a0f52fc-1ad2-4af7-9ed1-3e5559e29059 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Emu3: Next-Token Prediction is All You Need
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 835069a5-55c0-4959-9602-71d237bbe30c · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Instancediffusion: Instance- level control for image generation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db404a6b-ccc1-43c2-aba7-98f5806b4c31 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Tokencompose: Text-to-image diffusion with token-level supervision.2024 IEEE/CVF Conferenceon ComputerVisionand PatternRecognition(CVPR), pages 8553–8564, 2023
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation effb093e-0bbe-441b-9d17-e9717133fcc0 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Qwen-Image Technical Report
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 449a95db-e5f9-44f2-89da-6a4985fa6fe7 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e4c0aaf3-2f4d-46d8-92d2-4d07a2b27c62 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 984973a1-1b85-409a-85f0-f32006326c4d · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion.2023IEEE/CVFInternational Conferenceon ComputerVision(ICCV), pages 7418–7427, 2023
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d25c841a-26d7-402b-88e4-7de930f84fb1 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e5da4431-8128-489a-8fe0-2c0fdf5f475a · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 276680a2-f601-4dcf-bb86-2f6e84230380 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b2376fc3-d09e-491b-ab3d-ff9d59f8ace6 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 469978e9-f216-421d-9a3b-551371990134 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Diffusionmodelasanoise-aware latent reward model for step-level preference optimization
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 14a3fe6d-9416-4690-b1b3-fd2691a4c0ad · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b283f70-325c-4771-9399-2aeb653f65b2 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Itercomp: Iterative composition-aware feedback learning from model gallery for text-to-image generation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a0a4529-8acb-4d06-9d30-ebdc5e3679fe · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c74e5544-ffba-43da-b3d8-8d7efb67f563 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fee82bc0-6dc0-45e1-8278-a456b7da04f0 · outbound
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.