Pith. sign in

Paper Citation Record · LEDGER

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation

As of 7 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2507.21391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21391 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:54:38.274595Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 056a1e6a-beb3-4817-b201-201c4fffe76b · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.526329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.526329Z digest=sha256:d5bcaaaa833552d051db584a79b4fe154e7d1e4260a240deeab8c446ae1eb2d2

Observation 0eb2ab8a-9a25-4c24-aadd-539f4e16e3a0 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.634149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.634149Z digest=sha256:2f8d2886bfa82a1e513d661bd6744e9b16785e092cd2998f9864452a548d3f6a

Observation 1414ac53-eb29-4751-a06b-8f3f69844a80 · outbound

This paper cites Qwen2.5-VL Technical Report.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.751253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.751253Z digest=sha256:aeea51b7122eef6ed14de055b757cbf43e4541bfe67fde051bba3638aeb70652

Observation b0e9eb8b-c2f4-4e4f-9905-3f8dce5d3266 · outbound

This paper cites A Note on the Inception Score.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation A Note on the Inception Score

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.829262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.829262Z digest=sha256:f29b94fa9bf80b145f6e531e1bc282d108c8310f7d4f782e61f0efb75dc131f5

Observation b4616089-20a5-497c-8a4a-67aef162a28a · outbound

This paper cites Attend First, Consolidate Later: On the Importance of Attention in Different LLM Layers.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Attend First, Consolidate Later: On the Importance of Attention in Different LLM Layers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.897882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.897882Z digest=sha256:6261745825655a7a12ae80ec0a44a058f936756df9bbb186a41bf29e377208e4

Observation 01a57abf-d454-4edd-97d9-20a2f7402118 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Training Diffusion Models with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:32.955610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:32.955610Z digest=sha256:da1165ad44e53c0f5f9bc6a1592b40cf8b88dc5fd8ba14234ff47e039d1a4714

Observation 2c9bb625-c517-49c5-a40f-a8c3e4d0db53 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Rank analysis of incomplete block designs: I

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:44.693809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:33.077318Z digest=sha256:e3be20f6d1ea7cf0cde40587eca0d1a5feae1b12a50696077094b56578d4745a

Observation bb5a2d9f-8017-4cdd-8ae8-61688734c0a3 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.162969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.162969Z digest=sha256:e2398baca7caf8d1f9607bc9c5e51e0a8984d0344358f138686ef3e88fecc780

Observation bed65cea-25c8-4a87-a108-28b39587f9d8 · outbound

This paper cites SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.250868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.250868Z digest=sha256:391241b0a1a8678c3055708433cfc7edeab7bc7a41d44bc19e9645d00a048600

Observation 02c55c24-4378-4579-8a74-f34efdfec3c2 · outbound

This paper cites MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.346700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.346700Z digest=sha256:1c52795f70038c2d7941c315f8fda5b205310439d310f7dc9ebe7fbd58b5470e

Observation 8f92e954-ec8f-447b-a014-6cb2464831dd · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.452384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.452384Z digest=sha256:b463a1a78c7fe5ac90ec6a3a2999b32b89d11bc97c375bd868f4bb74333f5271

Observation fe32db23-e6ac-4c04-a95c-a2c6e99d02ca · outbound

This paper cites The socio-moral image database (smid),.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation The socio-moral image database (smid),

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:44.450201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:33.551927Z digest=sha256:d9255f1384b454c9cc9f6850e7729e11c829811170254b1e29bd8877fc279c06

Observation 62b1b304-f88a-48e3-a114-a693b7d5c173 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.643754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.643754Z digest=sha256:ee26b9157b662791166c18e391fdd46999c761118550388de5e89cbb2600ad59

Observation 9f9e2686-3aa2-4027-9f15-d3df3ff8be4f · outbound

This paper cites Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.744321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.744321Z digest=sha256:9322d058d3e9d211aa5feac45c1133e28c6380e5cae785dbb17a39c0acdc9445

Observation 78af6481-780c-46fc-b0f7-791cee9478e1 · outbound

This paper cites Diffusion models beat gans on image synthesis.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Diffusion models beat gans on image synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:33.829967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:33.829967Z digest=sha256:65d20226cb4e331c48a83eb31f13f00238aa0c5347762b33087269290e28b35c

Observation c1490d63-88ed-496b-aa9e-9014014c8769 · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Cogview: Mastering text-to-image generation via transformers

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:44.159261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:33.924185Z digest=sha256:5805abd0aaa1a5d8787dacf52067358b625697eb9de6ef7314495573f773137d

Observation 52916587-4e53-4a68-838e-944fbb9d03f4 · outbound

This paper cites Dpok: Reinforcement learning for fine-tuning text-to-image diffu- sion models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Dpok: Reinforcement learning for fine-tuning text-to-image diffu- sion models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:43.900852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:34.027720Z digest=sha256:633c9a8cd351c56a4d52fc8aab3cbe840690654a93358d4c4d09ddf6cf2bd8f9

Observation 34de7a70-025e-44a6-9eb6-2c9eec8e2305 · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Geneval: An object-focused framework for evaluating text- to-image alignment

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:43.638002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:34.106894Z digest=sha256:51cfc905380fdd135dc51b445f8b5762ea1e792836a8879acc725a482c0bbe64

Observation 60ca7fe7-16c2-4966-bc17-8305b4f4f500 · outbound

This paper cites Generative adversarial networks.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Generative adversarial networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:43.359411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:34.174919Z digest=sha256:81deac3952645ae2af20c72e9f0b7b76af39c1a912890bf4c48efcc0915fc963

Observation d8979830-9d36-44d0-be4c-f66261f68e20 · outbound

This paper cites LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.226626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.226626Z digest=sha256:24917478ab7e92dcc4c66c103cb2a468f23deb5873c9082e98beb272f2d7a1f8

Observation 93f16588-f141-461c-beb4-88fdcbfe6f57 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.274450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.274450Z digest=sha256:a27366b80e39ab798781bdf6d5d0d4a41c5ca8c303376a41b7e1fb43dbecf315

Observation 221cbe1b-c39c-48fa-b1f5-12875ac86735 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.347454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.347454Z digest=sha256:a8c10c8a7d1b428282c1a19fe0f77155b31e557cc526deb9e3bd8cf64508be7b

Observation 175a9c86-8bea-49fc-b5ff-ec2393d1ae23 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Denoising dif- fusion probabilistic models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.455883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.455883Z digest=sha256:9df5d31b4cc0993e5c8a26ae5b3281cbcb591df3df12dff79720b48f4b3c1a1b

Observation c6bf3fea-02f6-44eb-904f-15421c7d9897 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.578707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.578707Z digest=sha256:e789159aea3ac95afc5cdcfd0e43ec8a065bf8139af13c90b203f3f8b989bc24

Observation 848fe526-4132-4bf4-9e22-842a315f02bf · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.709673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.709673Z digest=sha256:c94866c978d17dcbb46d432dcec96dc89efebba9e506c5247e8dc0ae37a13a11

Observation 0bc91a6b-bbdc-4591-8105-4bfd5f909bf7 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:43.109881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:34.861279Z digest=sha256:a10adcd8f67a338a9271537c6f57bcec2239e20e9fb0984dd0e71a7a26d80b00

Observation 4bf251e1-f89a-4031-8745-a2ac7ded0574 · outbound

This paper cites VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.950914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.950914Z digest=sha256:3bd3f274afbe68d48ed4fd91fdb158f8fa000b46a5d2e517c530c873b46d3270

Observation d72d80f3-ba3e-4c4d-9c51-5d707ddf2eb9 · outbound

This paper cites What matters when building vision-language models?.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation What matters when building vision-language models?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.047298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.047298Z digest=sha256:fb50138b2489a945f41bc9f8ac7eeab0fbbc071c953f2713387b29494b6ae6b4

Observation 71df43e3-76ce-4791-a801-b71882b18abe · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:42.807225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:35.145183Z digest=sha256:768b64c6d4a41615de70b8ddda07bfca41752aa207441af7df55b72982baab61

Observation bc9aae33-9a27-401b-b7c8-be632420679e · outbound

This paper cites T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.249132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.249132Z digest=sha256:c02a1bd1064c0afc3408a8f7ae53df7f7383724486052056566715b45b5536ef

Observation 94664539-1edf-4155-acf0-6690ccf5666b · outbound

This paper cites Remov- ing distributional discrepancies in captions improves image- text alignment.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Remov- ing distributional discrepancies in captions improves image- text alignment

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:42.539329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:35.362579Z digest=sha256:62195318388da957e3df9dc916d431a65428332c0a8a5042643201296e2a8331

Observation 337ed284-26b4-4f43-b34a-095d129c1003 · outbound

This paper cites Evaluating text-to-visual generation with image-to-text gen- eration.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Evaluating text-to-visual generation with image-to-text gen- eration

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:42.271746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:35.442644Z digest=sha256:5eefdd8396ce077498e0f4e12ebdbccf07473f07cd3ee5db5cf3842aee995472

Observation 0b495463-d938-487a-836d-29917f2f9c97 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.526604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.526604Z digest=sha256:af831390f1b7bbe641f6c14aaee8f1c1dc362078d3b89f5424b89f83049afaec

Observation 485cd105-58bf-4858-b115-6c11111da2f0 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Improved baselines with visual instruction tuning, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:42.023960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:35.620233Z digest=sha256:aa1351f7c1bb0f4abdf144c16d22e64c1ea05eddab70f50560b1ba26357e818f

Observation 0750e59b-d266-4b5c-be87-abf663dab75c · outbound

This paper cites Visual instruction tuning, 2023.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Visual instruction tuning, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:41.790950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:35.676524Z digest=sha256:f76665202b4abc9320524dd0613e56dff14aec745db07ba67ff3821302e04afb

Observation e18fadd7-3903-415d-8b9c-58d5e539e4c0 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:41.529889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:35.725893Z digest=sha256:7ac79a36f7cb2a2f46799b93f2a13394599f9f6ba426412a1997dbd3f207dda2

Observation d4a9139f-d60a-4b32-a7b7-4b4aaa609763 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.785004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.785004Z digest=sha256:3627140e6f677d46233538e75f574360c901babdc8fb50cb8dfa90927da6a826

Observation af03050d-fe22-4c9c-86de-850fb1ef8d9a · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.880961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.880961Z digest=sha256:d05d8bd3a5d97d141c7c27aafec5371de1c5b6e9f37a50b6932b2b8a6853385b

Observation 90215b59-d14e-4456-8c68-0e470e90ac48 · outbound

This paper cites Simpo: Sim- ple preference optimization with a reference-free reward.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Simpo: Sim- ple preference optimization with a reference-free reward

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:41.318338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:35.929519Z digest=sha256:bcc001459e7de04907d169f472e9e16f244ec947a794a293f5e61302b91e961f

Observation 1d91967a-7d25-4f32-86be-571db2738094 · outbound

This paper cites Training language models to follow instructions with human feedback.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Training language models to follow instructions with human feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:35.985916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:35.985916Z digest=sha256:53e7e9ddc370357fb0cca4afe90c2d943b755de0cd22e3d3c2f19f7b3d5d1c58

Observation 7dbeb8a1-6a9a-4e17-bca3-1bd6e508fc90 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.083635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.083635Z digest=sha256:529b98f876a5b595ae10abe0751ff47ff45b20b1a5d94b187a642bd79213041c

Observation 19a5864c-3e88-4d4d-8c3a-5565a4ec0ab4 · outbound

This paper cites Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:41.061472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:36.177621Z digest=sha256:85aa3cab8d19a376b1014cc08bb632c171e7e7d78597936f305f875eeae96781

Observation 486054fb-1c29-405f-a1e9-74cbe132e51b · outbound

This paper cites UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.257427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.257427Z digest=sha256:4160cb9488050d1107e4ef017fc484aafe644b1079ed352cd887ccecfad7089d

Observation 498d6646-c498-490b-95c6-c800bcb74c8e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Learning transferable visual models from natural language supervi- sion

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.362963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.362963Z digest=sha256:23ae03552f25f5dd798569fa2b22f2e6415c5cab54aa5c956c1fcee0712642cb

Observation 95dc7cc0-f7d0-40c6-b4f4-aca3578260c1 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Direct preference optimization: Your language model is secretly a reward model

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:40.824402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:36.445317Z digest=sha256:769d45c469ce2d70a2146bafd5e15968b662994dbc888e557c9730a19d2ce5b4

Observation 607b8cf6-d085-4e66-8f7e-74768b3fca54 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Direct preference optimization: Your language model is secretly a reward model

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:40.561887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:36.562776Z digest=sha256:1430d3f1314256ac8114fe94f8a6e40d99c9bb54907bc7a83524cccc3f0e9029

Observation dc7536f8-7fcf-463d-90d7-e39f42ce4053 · outbound

This paper cites Zero-shot text-to-image generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Zero-shot text-to-image generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:40.280839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:36.684650Z digest=sha256:0f2ff8c0e8d78903ce26afae31e10ed9b2cb2cfeafff2b54bd3cea6f1850db8d

Observation c52fbfde-6f70-4c49-9cd8-36c5ef12c567 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation High-resolution image synthesis with latent diffusion models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:40.038981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:36.771953Z digest=sha256:a56386ae84394c6d72097afe1f9fd415c66436b12fb0ca38f7c735cf1fc616bc

Observation 4d7867f7-2ff2-42e1-8576-59eb0668b5c0 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.835962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:36.863239Z digest=sha256:438cc8b5ef88ec2ada81cacb9dbb45f668c73428ab67c13a855c4b0611083d28

Observation f0c300b8-ddfb-44bf-b455-1d9f0393f153 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.914296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.914296Z digest=sha256:f87db68a74e8e3bd81d49e57d9d82b1374974037c2eeb0d9d6c3f192db53992f

Observation 3ce9b861-b945-4995-ba66-1d41e41b7056 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Proximal Policy Optimization Algorithms

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:36.988698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:36.988698Z digest=sha256:389e20a0f1e8fa725d1fc26eb7243940ff7a09b6271ce865798983acb9030bdc

Observation a326a054-deaf-4fbf-8f2f-962890284bd8 · outbound

This paper cites A General Framework for Inference-time Scaling and Steering of Diffusion Models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation A General Framework for Inference-time Scaling and Steering of Diffusion Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.089978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.089978Z digest=sha256:ef52c7da24a4792d06237ff82dbcf062660298edcd2e637d70039ffbbce43a32

Observation 0fb5a6c6-185f-4428-8657-6b08bb29b32c · outbound

This paper cites an unresolved cited work.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:54:39.626966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:37.170906Z digest=sha256:8b03b03b5acc4a45609a4e3dfa838a73e1c9552be8a426317318599b7fa2b273

Observation 04f9475e-a2e7-4fcf-acdb-d64cd9abb94f · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Deep unsupervised learning using nonequilibrium thermodynamics

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.484015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:37.288393Z digest=sha256:2c15662ca5209d39c29fbb555c6a51b8762e30a702e8596d7ccd140ec15506c3

Observation faaa401d-b7f0-4ec2-b1aa-a54f8cd4e031 · outbound

This paper cites Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.374599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.374599Z digest=sha256:afbb23eef8e1f2eb60c0e2804e5566bbbff06936e0e66867707db1930fd0bbcf

Observation ee6f6be9-143b-4030-8a6e-1d5888fb3122 · outbound

This paper cites EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.474204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.474204Z digest=sha256:353d962fed940891c4c22420b305bcf302a76057880cda54653c58b0595b003f

Observation cf014844-e467-460e-99a7-964de4cccaef · outbound

This paper cites Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:54:38.472533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:37.533704Z digest=sha256:6cfb05652858d70526e526ff0d3e4fdd2d439a7f2d529f0a065211118dcfcff5

Observation 0d92521d-7af7-4991-839f-b0e355584ce3 · outbound

This paper cites HelpSteer2-Preference: Complementing Ratings with Preferences.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.619519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.619519Z digest=sha256:39906f1c047ad4acaeeb6edc8a23d3305d45e68d75a6ab42db802fe21b5473cc

Observation 0649d742-b68a-404e-b33f-82115b5ade71 · outbound

This paper cites MLLM-as-a-Judge for Image Safety without Human Labeling.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation MLLM-as-a-Judge for Image Safety without Human Labeling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.686456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.686456Z digest=sha256:596c2e52279d62b97e12e30923a89f2270d23503f61f51387daf85478071abd2

Observation 47994f55-4f79-43be-96c8-6f0b8b3df90c · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.787465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.787465Z digest=sha256:a5bbc2edfce8feb30541d34883a11ddc075f11a0e82ef86d42d5cf384d6867d8

Observation 55c643a1-c5da-4050-9945-d9ba9e50d216 · outbound

This paper cites Human preference score: Better aligning text- to-image models with human preference.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Human preference score: Better aligning text- to-image models with human preference

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.344930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:37.859469Z digest=sha256:e4c552bd5ec6ce13caa7a20065357cd5be79f41dfd6411a9fa887b28f7d122ca

Observation 75506a12-6c51-49bb-ab43-87fb31752d09 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Imagere- ward: Learning and evaluating human preferences for text- to-image generation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.261401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:37.927082Z digest=sha256:55c3acc04c9a7e2b2f20260d2e61e8f5a313d6cf2b76a964c573da942ef2284e

Observation 3eaeffba-521a-4480-85fb-6f093aab3080 · outbound

This paper cites A normalized levenshtein distance metric.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation A normalized levenshtein distance metric

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:39.129080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:37.998771Z digest=sha256:0afc035a39480c0e9d87eff8fbfa716b35ebcc015a723791d3fb8950a58203b7

Observation 3f3acc0d-51d8-422a-8f4f-e5647e7a1e6a · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2023.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation When and why vision- language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2023

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:38.984919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:38.090385Z digest=sha256:ec28a31ee70cf969d20fe8aa47517327ed665a5b12455f75f3515e4a92d1cf21

Observation a5e53093-eb1b-4ff2-9218-0079c8c509a9 · outbound

This paper cites Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:38.169486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:38.169486Z digest=sha256:f8d58292df60e08a10be7e10b24ca3c5a6e0f8630323ae5a21a123bc4dc04bb1

Observation b42a0d17-7f87-4a78-8598-5098a9b2a63e · outbound

This paper cites Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:38.229729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:38.229729Z digest=sha256:1d925ca09d97ee06bae5897c5c78cab6d3b894e9ee9c25255e4cd739b1e3e293

Observation cb6d2b9e-4aab-478f-ba2d-91070b9da082 · outbound

This paper cites Towards language-free training for text-to-image generation.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Towards language-free training for text-to-image generation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:54:38.832890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:54:38.274595Z digest=sha256:52813ec1ed328665a6521184dd634a9e37e60cdfacbe3a9187c6dca9fa619242

Pith citing papers

No inbound Pith citation observations are available.