Pith. sign in

Paper Citation Record · LEDGER

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL

As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2505.15791.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15791 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:18:56.930347Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T11:23:28.424453Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T11:25:18.912960Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa398351-ccc8-4599-a4fd-834ab65439bc · outbound

This paper cites Atom level enzyme active site scaffolding using rfdiffusion2.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Atom level enzyme active site scaffolding using rfdiffusion2

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:59.113875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:18:52.040527Z digest=sha256:48245253d3f1815e83679e4757631b725b657fc8417752bbedcf7901657fe76a

Observation d2c7f045-5afd-4730-9113-14e3e8151be9 · outbound

This paper cites Out of Many, One: Designing and Scaffolding Proteins at the Scale of the Structural Universe with Genie 2.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Out of Many, One: Designing and Scaffolding Proteins at the Scale of the Structural Universe with Genie 2

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:53.460195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:53.460195Z digest=sha256:636a6328555d56b72952e2cf832e329c0211018731b06b17b11d8bf1e6c8f357

Observation 5cf8b26f-86a2-47e3-979f-ef5b74fc1802 · outbound

This paper cites As referenced in the main text, Table 1 includes metrics from the DRaFT [Clark et al., 2023] and PRDP [Deng et al., 2024] papers.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL As referenced in the main text, Table 1 includes metrics from the DRaFT [Clark et al., 2023] and PRDP [Deng et al., 2024] papers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:58.325029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:18:56.835558Z digest=sha256:371f131da62866b1a47dd39ad0526ce37433e0e912dec96fc86bbd1e9efad5c7

Observation 74c7321d-08c8-4c1b-aee3-a557fbd4b929 · outbound

This paper cites DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.512770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.512770Z digest=sha256:59a43b66c1e5f7d1e72e16858af51a0a86965622047588a0e6b6841081d0532f

Observation 3dcc38c7-143c-4d6c-9845-a260871e95c9 · outbound

This paper cites Dealing with Sparse Rewards in Reinforcement Learning.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Dealing with Sparse Rewards in Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.713790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.713790Z digest=sha256:1c891ff59b947567a1226b6d27d7ec77143bb6edaaf8af8b85b7b3625b9b6faa

Observation 7fb24f74-efbe-459a-ac1b-38278052559d · outbound

This paper cites Sequence-Augmented SE(3)-Flow Matching For Conditional Protein Backbone Generation.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Sequence-Augmented SE(3)-Flow Matching For Conditional Protein Backbone Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.921055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.921055Z digest=sha256:481e8e4a443a2c43ebc326e6a1addadde0a39af193fbd2dad74e980fd39ca135

Observation 007c5f7c-0714-4c6f-bc71-fb2220e2149c · outbound

This paper cites Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.995998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.995998Z digest=sha256:3f308a9f5d08f864646aa0267a7711d985868cd6ecab701200714c1c346a7dfb

Observation 4c3004c2-7ed9-498e-b5a0-46ebcdb07874 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Aligning Text-to-Image Models using Human Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:53.080608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:53.080608Z digest=sha256:7d75ed5ec28f014af21c27495cf76137537144e899df8c703898b40d4a3364f8

Observation d3fe9064-24ba-441b-ba91-6e7565fbe53c · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:53.255145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:53.255145Z digest=sha256:a65a168c1efbd878dd53bd5089c3bd76858c30790af620c9b98e92e9321c6ffc

Observation 778a4dbe-3938-45a6-b3fb-84e31e2ff8b8 · outbound

This paper cites Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-Based Decoding.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-Based Decoding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:53.356102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:53.356102Z digest=sha256:038fe350c4969d887a4e0b5b1618629407efa6ad1b3cca7cc58b41087bced337

Observation 60b756a9-3543-4b66-b7a4-8eb2db9f1a9d · outbound

This paper cites Flow Matching for Generative Modeling.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Flow Matching for Generative Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:53.574219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:53.574219Z digest=sha256:517607d349abf93b3c3222f06d541a04fe00b842242c341729cc03e336a5cc5c

Observation 4c18afcc-799a-43ef-be36-8678337ee544 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:53.950021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:53.950021Z digest=sha256:c0ce288a95205f69e573accf27fbbf94c1701c9b650153615a14848aa2897b6c

Observation 51d5e71b-3704-4152-86e4-ff4274939898 · outbound

This paper cites Decoupled Weight Decay Regularization.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Decoupled Weight Decay Regularization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:54.458034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:54.458034Z digest=sha256:90904da769819e3deb8a9a10276dbd920703770d98631a2392b65d7db911ac40

Observation 6a317298-8807-4244-bf64-974ce588f354 · outbound

This paper cites Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:54.669121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:54.669121Z digest=sha256:14b1ecf1f83ecee08a8a8793cdde3a1ecbf7f10d01e5b9804bc2d8da22ee04aa

Observation 927be3ae-e2e6-4ebf-9010-bb91d10dafed · outbound

This paper cites Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:18:57.467666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:18:54.854896Z digest=sha256:58674f372b954afe09672314edeb86026f9d1cb0a444a9b1a4f75c276728032e

Observation b9ee0a31-7f21-4c32-99d2-569451a1693a · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:54.989604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:54.989604Z digest=sha256:eaf619fd78e115ba8582d38e187d3704ace5ed525c89cba72fd1e6e3bb080bd4

Observation 9bd3559c-88e7-4184-bfb8-9c50e89dee06 · outbound

This paper cites Multisample Flow Matching: Straightening Flows with Minibatch Couplings.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Multisample Flow Matching: Straightening Flows with Minibatch Couplings

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.121823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.121823Z digest=sha256:d3fd644d6d8d1e1ecab1e09538281c7075592de9bff7b3079db8ba5267a9ec10

Observation 99439a9f-a924-4e2a-87d4-1a291c60782e · outbound

This paper cites Proximal Policy Optimization Algorithms.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.223435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.223435Z digest=sha256:4683dd74f3d01c11ce277b1c57308a9c0a48f5ec18f54fa47042838d8c2056e5

Observation b04cb345-3e1d-490c-a74b-3109af4dde62 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.607841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.607841Z digest=sha256:b778a01693e52985f7ec20b403f5e5c278eac2a93b40f6ab64e333ddfabcee76

Observation 4bce29fa-00c9-40b7-b738-d00b81ff2a79 · outbound

This paper cites Diffusion Language Models Are Versatile Protein Learners.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Diffusion Language Models Are Versatile Protein Learners

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.707487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.707487Z digest=sha256:60569abd8ec189c1ac261664a614de34c2ebaf84bf8510bcda9c745ded3cb361

Observation 1fcef7d8-965e-4abc-b44a-e5aa6c130d7f · outbound

This paper cites Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.773131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.773131Z digest=sha256:3bbf89fd913db64a6549deae02253120e5db14fa35e2f755d130bd824b7514d0

Observation 0aa61977-60da-48fe-9f63-653bd8dbc57c · outbound

This paper cites VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.855992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.855992Z digest=sha256:cd147fef63551437c8cf880132370691518e677801d919366dbbdea55d30371e

Observation e1e16925-1af0-4627-bc8c-5cfc916a002e · outbound

This paper cites SE(3) diffusion model with application to protein backbone generation.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL SE(3) diffusion model with application to protein backbone generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:56.049830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:56.049830Z digest=sha256:b10da82032e4c096bd88116cd6c0d8a7f79bfb0837d6a93a3f05877380590e30

Observation 4316b84d-e748-4dfb-876f-d098d52ab0f4 · outbound

This paper cites Towards Controllable Diffusion Models via Reward-Guided Exploration.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Towards Controllable Diffusion Models via Reward-Guided Exploration

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:56.162933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:56.162933Z digest=sha256:de48bd1e6fd5f1495b02d9d16d7fc2fca4dbe5c27c80bef9df869843bedba4e8

Observation 0740abd3-2e50-41dd-81ff-a55d42b6508d · outbound

This paper cites Large-scale reinforcement learning for diffusion models.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Large-scale reinforcement learning for diffusion models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:56.295167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:56.295167Z digest=sha256:37a55f72dd33951b572ad389a8f035e162ac0c6ceaf7b987c84701ce57948cdc

Observation 5068b278-5f7a-45ef-94e6-75e76d93dd16 · outbound

This paper cites Behavior Proximal Policy Optimization.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Behavior Proximal Policy Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:56.424024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:56.424024Z digest=sha256:68de992859ba2454c2992487d971f0ee36e7b93723b256fc029c59c543296c1f

Observation ecf48c83-7fc8-4538-8f51-18b38d8930e7 · outbound

This paper cites an unresolved cited work.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:58.916493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:18:56.551973Z digest=sha256:f0a74e4fe93d9b17fdba6eea17d05acc483a6f62e9a3cb1a23d4b938251acb2f

Observation 0a16a69d-6f85-4532-9735-27775c67e570 · outbound

This paper cites Given these components, the flow matching objective for SO(3) can be formulated as: LSO(3)(θ) = Et∼U (0,1),q(R0,R1),Rt∼ρt(Rt|R0,R1) ∥vθ(t, Rt) − ut(Rt|R0, R1)∥2 SO(3).

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Given these components, the flow matching objective for SO(3) can be formulated as: LSO(3)(θ) = Et∼U (0,1),q(R0,R1),Rt∼ρt(Rt|R0,R1) ∥vθ(t, Rt) − ut(Rt|R0, R1)∥2 SO(3)

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:58.731174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:18:56.676501Z digest=sha256:f129a91353be3caf45f36235af038984baf7ad1ab2192a33a517613e1b24a64d

Observation c4c3a304-8b95-48f8-aed2-2517f9ea6ec9 · outbound

This paper cites an unresolved cited work.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:58.520325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:18:56.765412Z digest=sha256:cd8fdff249fdf8432b116404d707748db4142c73981ad86f714f885109d5c951

Observation 2c4718b8-dab6-4b0d-8d7b-384c6c591733 · outbound

This paper cites Color intensity correlates with η magnitude (darker corresponds to higher values).

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Color intensity correlates with η magnitude (darker corresponds to higher values)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:58.060287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:18:56.930347Z digest=sha256:6c696d99b70885a941cc61ae7f7ea674290422877ec86f44fde40c41cdad6aee

Observation 638dc0d6-b238-4d53-9b57-91a7984e2bad · outbound

This paper cites Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.480756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.480756Z digest=sha256:0e5940055efc83dd6afb1292d8fc7026a4606ff70f65ce11553ab6c0bd098ba0

Observation abeef27b-e5a6-4091-b5d6-461eee88ee95 · outbound

This paper cites Denoising Diffusion Implicit Models.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Denoising Diffusion Implicit Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.353442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.353442Z digest=sha256:b1f1cc23b26f65e631296ccddf8592b8e6d00e02dff73db2ffe21c27222c4747

Observation 76dcb53b-f88c-45c0-953d-50265d3b067d · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:54.567472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:54.567472Z digest=sha256:73a90321d9dc751af2a9290a06dbd32ed07de47d2c577ba718b189ebbca742b6

Observation dc7ce98c-5562-40c5-9f55-55aaeec638a8 · outbound

This paper cites Does RLHF Scale? Exploring the Impacts From Data, Model, and Method.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Does RLHF Scale? Exploring the Impacts From Data, Model, and Method

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.823260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.823260Z digest=sha256:bce2acf64d585ad5ad98f48ce3d9dd27aff5ba6a13095ca26e973f9fb8da62fd

Observation a25b8486-174e-4f1d-8625-e5328860af49 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.576998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.576998Z digest=sha256:5e261ddcd6a7e23bdbe57e53f9150a629f7e74b3ba13f2dbe8fa173c372befe5

Observation e4f6683b-e3f5-4300-b6ec-303fcbf2bcc3 · outbound

This paper cites SE(3)-Stochastic Flow Matching for Protein Backbone Generation.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL SE(3)-Stochastic Flow Matching for Protein Backbone Generation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.310087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.310087Z digest=sha256:a364eac317bfcc6b7d67fcbd1665e87d26cadf793af6e65d19192902af3478c7

Observation 78702dee-54bf-4302-adf0-977392a625cb · outbound

This paper cites Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.419212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.419212Z digest=sha256:d384070fd1e9eaa5fcf16d0db0e9e649080db4fb4825d73f33eb36728153a4da

Observation a473a493-0b83-4763-8618-ad6be55cebf8 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Training Diffusion Models with Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.161793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.161793Z digest=sha256:1b69effd31552422ef7d1d033121d9979f79929c41acda13f2aee401bd37bfdc

Pith citing papers

Observation ba09d5f4-eef9-4d98-8718-36e276d1a124 · inbound

LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories cites this paper.

LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:25:18.915785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T11:23:28.424453Z digest=sha256:cf6447eb0cb319bf25748a793bb2df24970ad148256c473876e523d4e8358d03