Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:54:38.274595Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2507.21391.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:54:38.274595Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 056a1e6a-beb3-4817-b201-201c4fffe76b · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eb2ab8a-9a25-4c24-aadd-539f4e16e3a0 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1414ac53-eb29-4751-a06b-8f3f69844a80 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e9eb8b-c2f4-4e4f-9905-3f8dce5d3266 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation A Note on the Inception Score
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4616089-20a5-497c-8a4a-67aef162a28a · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Attend First, Consolidate Later: On the Importance of Attention in Different LLM Layers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01a57abf-d454-4edd-97d9-20a2f7402118 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Training Diffusion Models with Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c9bb625-c517-49c5-a40f-a8c3e4d0db53 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Rank analysis of incomplete block designs: I
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bb5a2d9f-8017-4cdd-8ae8-61688734c0a3 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed65cea-25c8-4a87-a108-28b39587f9d8 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02c55c24-4378-4579-8a74-f34efdfec3c2 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f92e954-ec8f-447b-a014-6cb2464831dd · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe32db23-e6ac-4c04-a95c-a2c6e99d02ca · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation The socio-moral image database (smid),
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62b1b304-f88a-48e3-a114-a693b7d5c173 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f9e2686-3aa2-4027-9f15-d3df3ff8be4f · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78af6481-780c-46fc-b0f7-791cee9478e1 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Diffusion models beat gans on image synthesis
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1490d63-88ed-496b-aa9e-9014014c8769 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Cogview: Mastering text-to-image generation via transformers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52916587-4e53-4a68-838e-944fbb9d03f4 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Dpok: Reinforcement learning for fine-tuning text-to-image diffu- sion models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 34de7a70-025e-44a6-9eb6-2c9eec8e2305 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Geneval: An object-focused framework for evaluating text- to-image alignment
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60ca7fe7-16c2-4966-bc17-8305b4f4f500 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Generative adversarial networks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8979830-9d36-44d0-be4c-f66261f68e20 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93f16588-f141-461c-beb4-88fdcbfe6f57 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 221cbe1b-c39c-48fa-b1f5-12875ac86735 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 175a9c86-8bea-49fc-b5ff-ec2393d1ae23 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Denoising dif- fusion probabilistic models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6bf3fea-02f6-44eb-904f-15421c7d9897 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation LoRA: Low-Rank Adaptation of Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 848fe526-4132-4bf4-9e22-842a315f02bf · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bc91a6b-bbdc-4591-8105-4bfd5f909bf7 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Pick-a-pic: An open dataset of user preferences for text-to-image generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bf251e1-f89a-4031-8745-a2ac7ded0574 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d72d80f3-ba3e-4c4d-9c51-5d707ddf2eb9 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation What matters when building vision-language models?
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71df43e3-76ce-4791-a801-b71882b18abe · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc9aae33-9a27-401b-b7c8-be632420679e · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94664539-1edf-4155-acf0-6690ccf5666b · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Remov- ing distributional discrepancies in captions improves image- text alignment
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 337ed284-26b4-4f43-b34a-095d129c1003 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Evaluating text-to-visual generation with image-to-text gen- eration
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b495463-d938-487a-836d-29917f2f9c97 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 485cd105-58bf-4858-b115-6c11111da2f0 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Improved baselines with visual instruction tuning, 2023
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0750e59b-d266-4b5c-be87-abf663dab75c · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Visual instruction tuning, 2023
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e18fadd7-3903-415d-8b9c-58d5e539e4c0 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4a9139f-d60a-4b32-a7b7-4b4aaa609763 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af03050d-fe22-4c9c-86de-850fb1ef8d9a · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90215b59-d14e-4456-8c68-0e470e90ac48 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Simpo: Sim- ple preference optimization with a reference-free reward
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d91967a-7d25-4f32-86be-571db2738094 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Training language models to follow instructions with human feedback
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dbeb8a1-6a9a-4e17-bca3-1bd6e508fc90 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19a5864c-3e88-4d4d-8c3a-5565a4ec0ab4 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 486054fb-1c29-405f-a1e9-74cbe132e51b · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 498d6646-c498-490b-95c6-c800bcb74c8e · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Learning transferable visual models from natural language supervi- sion
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95dc7cc0-f7d0-40c6-b4f4-aca3578260c1 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Direct preference optimization: Your language model is secretly a reward model
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 607b8cf6-d085-4e66-8f7e-74768b3fca54 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Direct preference optimization: Your language model is secretly a reward model
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc7536f8-7fcf-463d-90d7-e39f42ce4053 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Zero-shot text-to-image generation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c52fbfde-6f70-4c49-9cd8-36c5ef12c567 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation High-resolution image synthesis with latent diffusion models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d7867f7-2ff2-42e1-8576-59eb0668b5c0 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Photorealistic text-to-image diffusion models with deep language understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0c300b8-ddfb-44bf-b455-1d9f0393f153 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ce9b861-b945-4995-ba66-1d41e41b7056 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Proximal Policy Optimization Algorithms
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a326a054-deaf-4fbf-8f2f-962890284bd8 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation A General Framework for Inference-time Scaling and Steering of Diffusion Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fb5a6c6-185f-4428-8657-6b08bb29b32c · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 04f9475e-a2e7-4fcf-acdb-d64cd9abb94f · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Deep unsupervised learning using nonequilibrium thermodynamics
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation faaa401d-b7f0-4ec2-b1aa-a54f8cd4e031 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee6f6be9-143b-4030-8a6e-1d5888fb3122 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf014844-e467-460e-99a7-964de4cccaef · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d92521d-7af7-4991-839f-b0e355584ce3 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation HelpSteer2-Preference: Complementing Ratings with Preferences
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0649d742-b68a-404e-b33f-82115b5ade71 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation MLLM-as-a-Judge for Image Safety without Human Labeling
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47994f55-4f79-43be-96c8-6f0b8b3df90c · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55c643a1-c5da-4050-9945-d9ba9e50d216 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Human preference score: Better aligning text- to-image models with human preference
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75506a12-6c51-49bb-ab43-87fb31752d09 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Imagere- ward: Learning and evaluating human preferences for text- to-image generation
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3eaeffba-521a-4480-85fb-6f093aab3080 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation A normalized levenshtein distance metric
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f3acc0d-51d8-422a-8f4f-e5647e7a1e6a · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation When and why vision- language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2023
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5e53093-eb1b-4ff2-9218-0079c8c509a9 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b42a0d17-7f87-4a78-8598-5098a9b2a63e · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb6d2b9e-4aab-478f-ba2d-91070b9da082 · outbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Towards language-free training for text-to-image generation
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.