Pith. sign in

Paper Citation Record · LEDGER

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL

As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2505.15791.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15791 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:18:56.930347Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T11:23:28.424453Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T11:25:18.912960Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa398351-ccc8-4599-a4fd-834ab65439bc · outbound

This paper cites Atom level enzyme active site scaffolding using rfdiffusion2.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Atom level enzyme active site scaffolding using rfdiffusion2

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:59.113875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:18:52.040527Z digest=sha256:f6dc002272b9a2b93435555887fed002ff6275dcb178e91b7b2732ef6f51dbc5

Observation d2c7f045-5afd-4730-9113-14e3e8151be9 · outbound

This paper cites Out of Many, One: Designing and Scaffolding Proteins at the Scale of the Structural Universe with Genie 2.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Out of Many, One: Designing and Scaffolding Proteins at the Scale of the Structural Universe with Genie 2

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:53.460195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:53.460195Z digest=sha256:b31340011a4aef4607bf78da5e58f458d8e20d5d781cf8de09b2e3ef77df4460

Observation 5cf8b26f-86a2-47e3-979f-ef5b74fc1802 · outbound

This paper cites As referenced in the main text, Table 1 includes metrics from the DRaFT [Clark et al., 2023] and PRDP [Deng et al., 2024] papers.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL As referenced in the main text, Table 1 includes metrics from the DRaFT [Clark et al., 2023] and PRDP [Deng et al., 2024] papers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:58.325029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:18:56.835558Z digest=sha256:1c7a4addc014482ea15db5749ad944c86209adab705d99efc28f1df5ba5bebf4

Observation 74c7321d-08c8-4c1b-aee3-a557fbd4b929 · outbound

This paper cites DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.512770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.512770Z digest=sha256:d84530e23f7938bca0e47316efa89a717a0606e98f290e74f5af2b52a3d8918f

Observation 3dcc38c7-143c-4d6c-9845-a260871e95c9 · outbound

This paper cites Dealing with Sparse Rewards in Reinforcement Learning.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Dealing with Sparse Rewards in Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.713790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.713790Z digest=sha256:66ec7154aa8332b997b647d0fe801e77a49bd2659840f66f93e50705f8b50c72

Observation 7fb24f74-efbe-459a-ac1b-38278052559d · outbound

This paper cites Sequence-Augmented SE(3)-Flow Matching For Conditional Protein Backbone Generation.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Sequence-Augmented SE(3)-Flow Matching For Conditional Protein Backbone Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.921055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.921055Z digest=sha256:8ec40a145105d09c6f57a28cd0b0427853b7a09dfd3f850c7e90d3ad209b513a

Observation 007c5f7c-0714-4c6f-bc71-fb2220e2149c · outbound

This paper cites Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.995998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.995998Z digest=sha256:3324561c57c49da71ff787c796c487f913cb48fe2fbc21a8b4da62efc469e415

Observation 4c3004c2-7ed9-498e-b5a0-46ebcdb07874 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Aligning Text-to-Image Models using Human Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:53.080608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:53.080608Z digest=sha256:4198024a1291052ef6dbe860695f3c4fa683f86783927a3305f921e88044dbfb

Observation d3fe9064-24ba-441b-ba91-6e7565fbe53c · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:53.255145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:53.255145Z digest=sha256:05d95a3c10c0540295a8aed6b95eddcb9d93a01df7ed1283fb79a73c07f0ae11

Observation 778a4dbe-3938-45a6-b3fb-84e31e2ff8b8 · outbound

This paper cites Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-Based Decoding.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-Based Decoding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:53.356102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:53.356102Z digest=sha256:88628b8a6d3b39ad5ce08d6f9e02960b34f0e8e2a26b91e180b73831d3ed9f43

Observation 60b756a9-3543-4b66-b7a4-8eb2db9f1a9d · outbound

This paper cites Flow Matching for Generative Modeling.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Flow Matching for Generative Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:53.574219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:53.574219Z digest=sha256:b6715c5bf0cfc860b8c020f2d035aef2a65914fa0613a317d6e71ac6b23c2400

Observation 4c18afcc-799a-43ef-be36-8678337ee544 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:53.950021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:53.950021Z digest=sha256:a029dd1f869c376ea9d0859c8eaa455959b584edd06de3b04e2bbe1fca4ed0b4

Observation 51d5e71b-3704-4152-86e4-ff4274939898 · outbound

This paper cites Decoupled Weight Decay Regularization.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Decoupled Weight Decay Regularization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:54.458034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:54.458034Z digest=sha256:38f306d88342dfb429c98be9b6030a94a3a96857bffbbb7c2b4aa87de4b2c7b5

Observation 6a317298-8807-4244-bf64-974ce588f354 · outbound

This paper cites Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:54.669121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:54.669121Z digest=sha256:1e119899fcd03dded6f2644343fe2cd25dac366695bef1bca77ff69e4b2c9b16

Observation 927be3ae-e2e6-4ebf-9010-bb91d10dafed · outbound

This paper cites Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:18:57.467666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:18:54.854896Z digest=sha256:9fb57b29bc2c9d0307594f45beac7e0192d7df3a4ebb5e1d5387962f577914ed

Observation b9ee0a31-7f21-4c32-99d2-569451a1693a · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:54.989604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:54.989604Z digest=sha256:872005372e08f75b00849c63cf4e09cde34bf0ffad8f5ce7cd0fb7a73110b404

Observation 9bd3559c-88e7-4184-bfb8-9c50e89dee06 · outbound

This paper cites Multisample Flow Matching: Straightening Flows with Minibatch Couplings.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Multisample Flow Matching: Straightening Flows with Minibatch Couplings

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.121823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.121823Z digest=sha256:c3b677d9b63f20512a274809d07b39d03cbbd0935d0b1f409f70f5165e74d285

Observation 99439a9f-a924-4e2a-87d4-1a291c60782e · outbound

This paper cites Proximal Policy Optimization Algorithms.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.223435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.223435Z digest=sha256:9e5226aeb020e5802ad2d7c229ca07234a5f84defc4ba226342c959e661e71da

Observation b04cb345-3e1d-490c-a74b-3109af4dde62 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.607841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.607841Z digest=sha256:82297cb7428c3db1e9382f9d60d5408a1c66e2aaae026f62a9acf18ff0f7b49b

Observation 4bce29fa-00c9-40b7-b738-d00b81ff2a79 · outbound

This paper cites Diffusion Language Models Are Versatile Protein Learners.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Diffusion Language Models Are Versatile Protein Learners

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.707487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.707487Z digest=sha256:9db490bba873d22994ef4a3d9dda88bfa3a68bb726dee99d7ee4e675e50adf9d

Observation 1fcef7d8-965e-4abc-b44a-e5aa6c130d7f · outbound

This paper cites Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.773131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.773131Z digest=sha256:f43158c3f5d4b47881de5caef6346f5d89cdeb19f1bb49cb7c0a3c56588430e2

Observation 0aa61977-60da-48fe-9f63-653bd8dbc57c · outbound

This paper cites VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.855992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.855992Z digest=sha256:127ec4e48b0618d042ab6b3ba30cb7a571cdfb63e341c88c66cb185efb32c760

Observation e1e16925-1af0-4627-bc8c-5cfc916a002e · outbound

This paper cites SE(3) diffusion model with application to protein backbone generation.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL SE(3) diffusion model with application to protein backbone generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:56.049830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:56.049830Z digest=sha256:3909d41a9f01b303b267732205e5f47890b834c54b1d5a273c874736a22fe806

Observation 4316b84d-e748-4dfb-876f-d098d52ab0f4 · outbound

This paper cites Towards Controllable Diffusion Models via Reward-Guided Exploration.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Towards Controllable Diffusion Models via Reward-Guided Exploration

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:56.162933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:56.162933Z digest=sha256:cc1c9298ab689070d54523ad5040afab17cbcfac47f65f01875ef7015ab585b4

Observation 0740abd3-2e50-41dd-81ff-a55d42b6508d · outbound

This paper cites Large-scale reinforcement learning for diffusion models.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Large-scale reinforcement learning for diffusion models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:56.295167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:56.295167Z digest=sha256:17d3343a310985d37a52e71572e9edff7d8d678d767f2311d90d8fb5c333c4dc

Observation 5068b278-5f7a-45ef-94e6-75e76d93dd16 · outbound

This paper cites Behavior Proximal Policy Optimization.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Behavior Proximal Policy Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:56.424024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:56.424024Z digest=sha256:3d278e5df4bd08c13d4b3b53d132e1c91d0f4c7a7f604ab8a222d23f03ae8b6d

Observation ecf48c83-7fc8-4538-8f51-18b38d8930e7 · outbound

This paper cites an unresolved cited work.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:58.916493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:18:56.551973Z digest=sha256:e8add2ed3c8a20c241256b63f136230dac52a430f91f8e20e83acbcd3a9f0d89

Observation 0a16a69d-6f85-4532-9735-27775c67e570 · outbound

This paper cites Given these components, the flow matching objective for SO(3) can be formulated as: LSO(3)(θ) = Et∼U (0,1),q(R0,R1),Rt∼ρt(Rt|R0,R1) ∥vθ(t, Rt) − ut(Rt|R0, R1)∥2 SO(3).

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Given these components, the flow matching objective for SO(3) can be formulated as: LSO(3)(θ) = Et∼U (0,1),q(R0,R1),Rt∼ρt(Rt|R0,R1) ∥vθ(t, Rt) − ut(Rt|R0, R1)∥2 SO(3)

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:58.731174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:18:56.676501Z digest=sha256:0b3528ae5ad806dafa4c2253d638e702f9adc56b8be5cba94965ef4d795ed09c

Observation c4c3a304-8b95-48f8-aed2-2517f9ea6ec9 · outbound

This paper cites an unresolved cited work.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:58.520325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:18:56.765412Z digest=sha256:f32b8fc41dd583439c56ee48b7b681315a9fda4ca180beeb083fd9be862ec925

Observation 2c4718b8-dab6-4b0d-8d7b-384c6c591733 · outbound

This paper cites Color intensity correlates with η magnitude (darker corresponds to higher values).

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Color intensity correlates with η magnitude (darker corresponds to higher values)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:58.060287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:18:56.930347Z digest=sha256:97f2d2097555a6f6690bc944af637184546d024e56ad8ac246176463efa7bed9

Observation 638dc0d6-b238-4d53-9b57-91a7984e2bad · outbound

This paper cites Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.480756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.480756Z digest=sha256:49beaae3afe37d6b9006c5d16b76c94fccfdcf329d1916420e7dfd929cb48734

Observation abeef27b-e5a6-4091-b5d6-461eee88ee95 · outbound

This paper cites Denoising Diffusion Implicit Models.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Denoising Diffusion Implicit Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:55.353442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:55.353442Z digest=sha256:be7e4ea8f9d582872b94515a4404c82b1275b2d4bbcd9fe306fa8affda23993d

Observation 76dcb53b-f88c-45c0-953d-50265d3b067d · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:54.567472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:54.567472Z digest=sha256:90428f2fb62784823a358194bfec041b21d477117daa00f2c7bf3ad81a0615e0

Observation dc7ce98c-5562-40c5-9f55-55aaeec638a8 · outbound

This paper cites Does RLHF Scale? Exploring the Impacts From Data, Model, and Method.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Does RLHF Scale? Exploring the Impacts From Data, Model, and Method

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.823260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.823260Z digest=sha256:0aa3a688f572c04e743f28356c2acabb287f4e938603205d4857943f1f8a56f2

Observation a25b8486-174e-4f1d-8625-e5328860af49 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.576998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.576998Z digest=sha256:31adedfb1e94b174dbaf774269e4c0160b06bbd91ee84dfe9e67cec603ecb662

Observation e4f6683b-e3f5-4300-b6ec-303fcbf2bcc3 · outbound

This paper cites SE(3)-Stochastic Flow Matching for Protein Backbone Generation.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL SE(3)-Stochastic Flow Matching for Protein Backbone Generation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.310087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.310087Z digest=sha256:7de7bc74e377cfd75f309ac33d6d1c116d30c1e7ef8801aa4a84655eeffc6e77

Observation 78702dee-54bf-4302-adf0-977392a625cb · outbound

This paper cites Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.419212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.419212Z digest=sha256:336eb99280e9a3faf01be61af06675468ae315c8c2f8ea10328da59cdd43924c

Observation a473a493-0b83-4763-8618-ad6be55cebf8 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Training Diffusion Models with Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.161793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.161793Z digest=sha256:9cfef0eed3e08e62d97911c3f770fc2f9dcef8d579db9a20bbab95d48b462f8c

Pith citing papers

Observation ba09d5f4-eef9-4d98-8718-36e276d1a124 · inbound

LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories cites this paper.

LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:25:18.915785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:23:28.424453Z digest=sha256:f0dc70df92318fc15b74ea15f9c72b46d9ab96f1d0dba34c455975dcf2fee88d