Pith. sign in

Paper Citation Record · LEDGER

Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2503.11240.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.11240 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:23:03.425660Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T02:25:55.901105Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2c0a9f22-4a7e-4eab-b5c2-0342d3b9d6c4 · inbound

D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples cites this paper.

D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:03.425660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:03.425660Z digest=sha256:351cab17871edd12c5e19ce6b1b037c5db3caa05cea565383194d8558f2ffe95

Observation 2f56465a-12dc-4f28-b8bd-a4ab21f0db97 · inbound

Self-Reflective Reinforcement Learning for Diffusion-based Image Reasoning Generation cites this paper.

Self-Reflective Reinforcement Learning for Diffusion-based Image Reasoning Generation Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:36.702908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:36.702908Z digest=sha256:b3479b4b0b4c72dc5f93403b12597547a81b0bcee92dade8d8fa6ff688f120b2

Observation fa88cfb6-389d-4ca5-9e99-22fc07efb30b · inbound

Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization cites this paper.

Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:04:52.397564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:04:52.397564Z digest=sha256:393531cda5035029933173dd1a6a616dfc901f3ee33029f50ae420cfccad13c2

Observation d2cacf94-a2d2-46e2-886c-be9f94097121 · inbound

Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs cites this paper.

Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:15.559192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T02:10:27.595446Z digest=sha256:c08c4381f1f93a2f0fa57dcc6bda7091f99f6842b6ce4a462f697d93dbe0cee6

Observation 9fd710a6-780b-4245-bb15-a54709896175 · inbound

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF cites this paper.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.902473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:ef783fb340a3f41a192da545267447e99ab77105069e5aa21ca75902348dce04