Pith. sign in

Paper Citation Record · LEDGER

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning

As of 7 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 4 inbound Pith citation observations for arXiv:2505.19196.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19196 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:25:26.232105Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:18:33.490433Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:03:13.983158Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a59cfb8-d6c1-4a4a-98d1-7b9a0c718506 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Deep unsupervised learning using nonequilibrium thermodynamics,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.824017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.059212Z digest=sha256:8710c19be5dd82e9f47a78c207e259a13651468fec81bea778475f4f52d7c4df

Observation b9b42b1d-c934-4fe3-a11f-9245bc2f86dc · outbound

This paper cites Denoising diffusion probabilistic models,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Denoising diffusion probabilistic models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.062747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.062747Z digest=sha256:852d828d2958172ffede354d9e799afa2d2437d92fcbe8ed238b6c20c6686275

Observation f9f7afaf-0e9b-44a9-8d97-aa9f714e2f5e · outbound

This paper cites Denoising Diffusion Implicit Models.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Denoising Diffusion Implicit Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.066590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.066590Z digest=sha256:5d39099ae21b5a5a94a6daf5a7c09cf2b6956344179e84e2484c5c4c9e50c5a4

Observation 616f4244-65ae-41b0-af70-5fc41b127016 · outbound

This paper cites Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.069859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.069859Z digest=sha256:7053d5349932bafa91d3cdc8a4157b9ebdfb5c598ae6a9bf514f6dc132dcb0e7

Observation c9ba1b6e-ed7a-4f4a-8446-0f66494ba7c5 · outbound

This paper cites Generative Adversarial Networks.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Generative Adversarial Networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.072893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.072893Z digest=sha256:b0b767baf84091b53efbf80c0d8dfa0226d834a5999c607ccce1e178472872e0

Observation 36fc2aa5-ff92-4347-bb87-46614c818290 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Learning transferable visual models from natural language supervision,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.808598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.076062Z digest=sha256:e20e970672ddadd62b40a6d8c8c52b1a16c7805a9c98bd11d19f98622ad73d3d

Observation a8cf128a-03e1-4295-ba87-01beadc4c0b3 · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.797986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.079004Z digest=sha256:3544bce807eb8b2ca92c8a667594edddb87f2511919afcca0a6ae822b97f409c

Observation c2c730e0-bca1-4d26-9780-15e0f95475fe · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.788543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.081959Z digest=sha256:477199b7f74fd1090bdf0c2afe890b068899aaa003245fe2898e735dcbf53dac

Observation 31e3e112-8969-48c6-af7e-3303037df51c · outbound

This paper cites Microsoft COCO: Common Objects in Context.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Microsoft COCO: Common Objects in Context

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.084718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.084718Z digest=sha256:94aed37902a4d91d9d559216bef093698fca6c512683dceff458565dca64bfa1

Observation 5035b145-6b7c-4329-8388-b38f41b8f4b9 · outbound

This paper cites Imagenet: A large-scale hierarchical image database,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Imagenet: A large-scale hierarchical image database,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.087407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.087407Z digest=sha256:c2868ab291ee4436fbffe89fc2316dfc8cdee128f44335da760b27b25c8ac319

Observation f5969777-9350-40a4-a3bd-dc39466c7730 · outbound

This paper cites High-resolution image synthe- sis with latent diffusion models,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning High-resolution image synthe- sis with latent diffusion models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.773507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.091066Z digest=sha256:5b13e587e6b5764a456812bb51a77b7c97cf95690ef1b874ca94c76c48276e56

Observation 54a42102-2d77-47b5-891a-c44c17bdd6c9 · outbound

This paper cites Improving image generation with better captions,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Improving image generation with better captions,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.764568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.093961Z digest=sha256:c59b8e9d6e245addf8da75ca8a869adca2c79ad2d8fda4c93794bc3ac5a58b0a

Observation f7a42b79-f648-40cd-9bd7-d5abb12dc7c7 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Laion-5b: An open large-scale dataset for training next generation image-text models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.756226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.096690Z digest=sha256:c891b4472b7a77a5f594a3d8bfa0c9532c31a769c616a5b75914dda082e37582

Observation 73f0cfd5-16ce-492f-b66b-8a3922494028 · outbound

This paper cites Training-free structured diffusion guidance for compositional text-to-image synthesis,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Training-free structured diffusion guidance for compositional text-to-image synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.746068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.099271Z digest=sha256:b24119797e4981ebb323569c4b520cf54d8115674dc893580edade473eb93564

Observation ac972a64-6cde-41a2-b6f0-db927337c602 · outbound

This paper cites TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.101775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.101775Z digest=sha256:a661cd53f5b991db7f5bd2eaf75d9117718b257366ecc18e209e7c1f744308e3

Observation 83dfc91e-b25e-4b5a-997d-134f4b12d7c2 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.104634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.104634Z digest=sha256:c1b77c4d0b0acef70e7cef0e1f0c1170e43817461ec1fea44c16311e7c0c2f42

Observation d9ec2f17-2af9-4049-87ef-f4529c2dd380 · outbound

This paper cites Imagereward: learning and evaluating human preferences for text-to-image generation,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Imagereward: learning and evaluating human preferences for text-to-image generation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.737462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.107547Z digest=sha256:68fcd211b74d94f4b67ba20cb341e413506319114d4ea704a71099257ddd5460

Observation b4ae061a-1de1-4a54-b92f-fc9145c7f647 · outbound

This paper cites Pick-a-pic: an open dataset of user preferences for text-to-image generation,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Pick-a-pic: an open dataset of user preferences for text-to-image generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.728358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.110441Z digest=sha256:a5cdf3ffa337e5a3f62c65892b15540b6ab8c5ffea1ce2db996b24db4041a0ec

Observation 9c0f8328-a04a-43ae-9372-84076f1c2057 · outbound

This paper cites OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.112892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.112892Z digest=sha256:44a6f716de6a63468b028779cf3a163d8eb055f977f9bdbc7a46471b08f2d17c

Observation ffc72921-f5e6-4523-9723-d23778d31cd0 · outbound

This paper cites Dpok: reinforcement learning for fine-tuning text-to-image diffusion models,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Dpok: reinforcement learning for fine-tuning text-to-image diffusion models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.719666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.115498Z digest=sha256:60bfa633914f1c5ecf00c0ca2a03cde99cfd830ddb00a0752c4c3f0bfa2f4fbd

Observation a49943ae-083e-4f04-924f-c5f98e89a191 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Training Diffusion Models with Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.118205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.118205Z digest=sha256:1cdd323c42cad8a996f376da42aae6d645e83b1fe2bdcddf3a9af1c0824327ac

Observation 3d82d039-20cc-42ab-9944-ced4ffa73f96 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Proximal Policy Optimization Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.120984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.120984Z digest=sha256:4710f3195a1a7b2f2d196ce68b0a713221335dc3622175d88d3c0f067f2676fb

Observation f5d1d4d8-172b-4707-b6b7-55870db39f66 · outbound

This paper cites Direct prefer- ence optimization: Your language model is secretly a reward model,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Direct prefer- ence optimization: Your language model is secretly a reward model,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.710368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.123805Z digest=sha256:e71d37e6afbaf73607d345cf43c2318497613159cdaf0bb98af4fa0a90878853

Observation 0af8666d-0e7e-4f19-a634-d622111e675f · outbound

This paper cites Stimulating diffusion model for image denoising via adaptive embedding and ensembling,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Stimulating diffusion model for image denoising via adaptive embedding and ensembling,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.701020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.126311Z digest=sha256:170698d511569db305648b3037a1e00805e6e54c847a51f0addb567a51480d1a

Observation baa652e8-00ce-48a7-bf95-b584a9bdb855 · outbound

This paper cites Blue noise for diffusion models,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Blue noise for diffusion models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.691274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.129059Z digest=sha256:6a36cd6f775d9c22bc423ada4511527128766bd7048d9fb9d3b9f8a64b360b97

Observation 7a764e69-28ab-4476-bac1-3702711a3a7c · outbound

This paper cites Boosting Diffusion Models with Moving Average Sampling in Frequency Domain.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Boosting Diffusion Models with Moving Average Sampling in Frequency Domain

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.132556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.132556Z digest=sha256:88dd2b03b55aa8476cd0e44eed7f57af948068322a8e1df59b6f6f5dc73afeab

Observation 80e01d5d-5291-4403-a1f0-853487a63e6a · outbound

This paper cites FreSca: Scaling in Frequency Space Enhances Diffusion Models.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning FreSca: Scaling in Frequency Space Enhances Diffusion Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.136256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.136256Z digest=sha256:a182eeed51ddb77ff313ed1a03c37f7a939ff62e356f348b5002402c576650b4

Observation 293a2962-4691-4cb0-bce0-88f8b86fb19b · outbound

This paper cites A dense reward view on aligning text-to-image diffusion with preference,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning A dense reward view on aligning text-to-image diffusion with preference,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.682210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.139255Z digest=sha256:54dd3793b1c5003ff9e0d9082b63f9751c45766c2dc46da11d2b1597c980be29

Observation 1f77c922-b895-41d0-a9a4-06d761f17ebe · outbound

This paper cites Confronting reward overoptimization for diffusion models: A perspective of inductive and primacy biases,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Confronting reward overoptimization for diffusion models: A perspective of inductive and primacy biases,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.672920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.141926Z digest=sha256:5e7d7c2a9d9df636ebe4843d55fb090040cce33c91bd087358540dce0ab2e186

Observation 7ddc2152-916d-4b1e-b35e-103cb7023a17 · outbound

This paper cites Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.144841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.144841Z digest=sha256:2cf49e8394cdc8904dcdcd1daf05ac51e4b80eb5a3e05702a1b0f64d94584168

Observation 11dccaa8-3582-4b83-b462-f53b6771e70d · outbound

This paper cites Diffusion models beat gans on image synthesis,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Diffusion models beat gans on image synthesis,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.663616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.147939Z digest=sha256:409f464d82e050e78b6b196185bc8c91b9f25ce9470ca686461c1b3bbae872ca

Observation 186e990f-a787-4d59-ab3d-7b884ffe7fd5 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Classifier-Free Diffusion Guidance

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.150649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.150649Z digest=sha256:80208a6e9da99ac15d22b4c45d1b122a07345b87e6a73a3e73e03284912f6060

Observation edd2dd8b-19b6-4ce7-bea8-43db8f0eb1bd · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Aligning Text-to-Image Models using Human Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.154161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.154161Z digest=sha256:9d7b3149c6d6dbcbd25f5e2f715c78e4f01ac1e62d732c866dcd5e4339ca9588

Observation 7795060c-6414-4eb9-821b-c97d09fe8840 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.158191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.158191Z digest=sha256:eb61399b5e39605e0a0af17e67073fa1234d93b96febdd76292493f20714700f

Observation 03783dd0-c41f-4932-9fd6-780499cd33e4 · outbound

This paper cites Optimizing ddpm sampling with shortcut fine-tuning,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Optimizing ddpm sampling with shortcut fine-tuning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.653284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.161300Z digest=sha256:1e0db4bb642a1abef76d202e2a3c6b5b50c84e20f05f2bf56934f9dbdda736c6

Observation 44e5fb01-4654-4a95-ae6e-54baca475030 · outbound

This paper cites Deep reward supervisions for tuning text-to-image diffusion models,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Deep reward supervisions for tuning text-to-image diffusion models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.644330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.164256Z digest=sha256:eb86cd6ee2e44f78bf96e3f000b50c5768dab9828f857745b459c43e234fc35c

Observation 44045a80-0c4d-4cde-b168-b89721392dd0 · outbound

This paper cites Using human feedback to fine-tune diffusion models without any reward model,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Using human feedback to fine-tune diffusion models without any reward model,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.634403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.167113Z digest=sha256:f1b5f5201413e99000503de5df0a2844051d1ce4b09060bc2e938825b357f2de

Observation 4e8fbaae-74bc-41a7-8ce7-76a0602f52ed · outbound

This paper cites Diffusion model alignment using direct preference optimization,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Diffusion model alignment using direct preference optimization,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.549875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.170322Z digest=sha256:6598ec63325477a65de8dfd1862251cbeaba1d473e354f135c0d35c6aace5335

Observation 357d6f4e-48c6-430d-9274-161ff9e8f422 · outbound

This paper cites Steps toward artificial intelligence,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Steps toward artificial intelligence,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.540929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.173109Z digest=sha256:c4ecc13685bafa0c3c76538be8c564b29a08acbb282985ef4540cdd93e906aa3

Observation d0ad225d-b9e1-4768-80e0-10a61a269b37 · outbound

This paper cites Temporal credit assignment in reinforcement learning,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Temporal credit assignment in reinforcement learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.530218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.176690Z digest=sha256:4ec3e9934f413327dd35f0e9945357c553b01a28913f85f947e6db064a2387a7

Observation 65d6c08a-74bd-40f4-9270-b8aac917d6a7 · outbound

This paper cites Learning guidance rewards with trajectory-space smooth- ing,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Learning guidance rewards with trajectory-space smooth- ing,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.520203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.179657Z digest=sha256:68744a4499538f6d5aaedb347ef03a92ab1313d40be3b797fccb9f4e49b58f0b

Observation 33f72736-f8f7-4112-a81b-668036b499ca · outbound

This paper cites Harutyunyan, W.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Harutyunyan, W

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.510627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.182674Z digest=sha256:ff001bcbcbb17c57ed8b8233a4cf0c520f0d0202e7a582a77c740b2fa48c45b0

Observation 3724d282-7afc-42bc-aa97-5da828ff567a · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Policy invariance under reward transformations: Theory and application to reward shaping,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.502040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.185401Z digest=sha256:d03b45f66fe6b32bf816c90abfed80d950c0324c87e9f38d5f82aba9af3eb110

Observation d0a81003-26d7-43ce-ba62-ae59e4f1d1e5 · outbound

This paper cites Text2Reward: Reward Shaping with Language Models for Reinforcement Learning.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.188530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.188530Z digest=sha256:809ee937021b2568bc81ff83cacf1ba6362dac4107d4bd77485432b39535ab6b

Observation c99752e1-3a18-426a-8a2b-ebd7f2dcc812 · outbound

This paper cites DPO meets PPO: Reinforced token optimization for RLHF,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning DPO meets PPO: Reinforced token optimization for RLHF,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.493068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.192604Z digest=sha256:e8d8afeef025b02d3cb4d540d23df8d88457b7faab451e5b009e893bfba4ecba

Observation 7a6199c3-bf2b-4c68-ba3a-3e3b22308795 · outbound

This paper cites R3HF: Reward redistribution for enhancing reinforcement learning from human feedback,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning R3HF: Reward redistribution for enhancing reinforcement learning from human feedback,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.484041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.195874Z digest=sha256:852262d575d16655411f247cc3582e358fea0d6445dada7173bc11a3efc487c1

Observation 9469099a-82cb-40f7-a438-e3935c0b55f9 · outbound

This paper cites Dense reward for free in reinforcement learning from human feedback,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Dense reward for free in reinforcement learning from human feedback,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.474408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.198621Z digest=sha256:2abed70e1484f346d92ab938dc80bbabfb30f2647094646180e3a67c84f6c9a7

Observation 1be93a43-71a6-4aa1-8b09-4246ea680c60 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Simple statistical gradient-following algorithms for connectionist reinforcement learning,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.453992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.204535Z digest=sha256:1f326890093933d422d2b4903fe8efb945e24eba7f754ed5e37d59bd2da2fed5

Observation 69347cfb-4e3b-468c-9ed9-27beddb7a661 · outbound

This paper cites Auto-Encoding Variational Bayes.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Auto-Encoding Variational Bayes

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.207681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.207681Z digest=sha256:df6b53152230a19249679febb56af770a28f1bf84416cad4dad05f046af8d509

Observation f6951719-dd51-470b-898e-5a1b0c4726a2 · outbound

This paper cites Emerging properties in self-supervised vision transformers,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Emerging properties in self-supervised vision transformers,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.444511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.210873Z digest=sha256:2796a9cccff2a34d0ebc07ef1186442b9cfee65c2bf938acbb935f2e87b5c78f

Observation dabf1354-2074-4903-b190-7c30f48f6d6d · outbound

This paper cites DiffSim: Taming Diffusion Models for Evaluating Visual Similarity.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning DiffSim: Taming Diffusion Models for Evaluating Visual Similarity

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.214334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.214334Z digest=sha256:7a2aab0dc29523211d5281ce276f3a65e80fc02f22c03236921e25abc04cc89f

Observation 92df4940-22bb-4894-8c43-65ab8291f10c · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning LoRA: Low-rank adaptation of large language models,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.217670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.217670Z digest=sha256:759dbe9b56880b82c8dafb6a4ba069b20e7f9284dc634d515c4723030f73c0e6

Observation 1762ec38-ef64-4b33-9f98-a3eb686a4f25 · outbound

This paper cites Policy gradient methods for rein- forcement learning with function approximation,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Policy gradient methods for rein- forcement learning with function approximation,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.429414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.220474Z digest=sha256:edc18da89a8b7b1eadcc1e870166cda9f2470b02ab0e5adb61670881e846aff3

Observation f8186217-4635-4ead-9ac7-55d49981bf93 · outbound

This paper cites U-Net: Convolutional Networks for Biomedical Image Segmentation.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning U-Net: Convolutional Networks for Biomedical Image Segmentation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.224156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.224156Z digest=sha256:0f16c786f4f4a6cdf6d9e3d100d0fb3ca92367392a2def0b9fe70abe156971ef

Observation e00ca729-7998-4e5e-9493-2baae7aac08c · outbound

This paper cites Training deep nets with sublinear memory cost,.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Training deep nets with sublinear memory cost,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.419888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.229122Z digest=sha256:03703b8f8609be92b0341213e1eeb07e827c701e445d411a4db9158f477700a6

Observation f5acd2f4-921b-43af-b595-29fdb03e2953 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Training Deep Nets with Sublinear Memory Cost

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:26.232105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:26.232105Z digest=sha256:c932978d869caad8d60041204a871ecf4d350a2def51668a607cfa4ed692f49b

Observation eb825529-10da-4a0b-bf1e-6db2928fdaf5 · outbound

This paper cites Available: https://openreview.net/forum?id=eyxVRMrZ4m.

Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning Available: https://openreview.net/forum?id=eyxVRMrZ4m

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:26.463485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:25:26.201653Z digest=sha256:1db620486344b5e5f928dc2d7416dd566afc1bc24895268d5b87f4c0db343dd6

Pith citing papers

Observation 45e9e2b8-63b4-4484-a920-a61a0bdab5c2 · inbound

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation cites this paper.

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:33.490433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:33.490433Z digest=sha256:cc2babcee479cc0e1d5ecfeb5cbc0bd9d83120a87b4052b66fc44cc9e936bf20

Observation fb275552-9d13-432f-8f4c-003c320e670b · inbound

LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion cites this paper.

LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T21:09:36.040879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:09:36.040879Z digest=sha256:c05ee7cfddd40de035839587d2e0a02aee03dafe86ad66c4680293bb91b5319b

Observation 71e0a791-1444-4142-a587-66be70140e9f · inbound

LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories cites this paper.

LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:25:18.786800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:23:28.424453Z digest=sha256:69dcfc9b5e2b64f9c986f67ac79d7469743efe9ffc158be6208864d91429cbde

Observation e2f44700-b3bf-4fee-b294-a22df8a5e414 · inbound

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation cites this paper.

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.984555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:00:16.005187Z digest=sha256:89436cde4edf13e69da220bfb0be4d4b66b6f0ba135bb3be55df87726b90a898