Pith. sign in

Paper Citation Record · LEDGER

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

As of 19 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 4 inbound Pith citation observations for arXiv:2605.15855.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.15855 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T19:48:17.049547Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:34:44.686328Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T00:23:55.444823Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact17
  • verified fuzzy56
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5bcb7de9-18ad-47d6-a362-0833bbd1c3e4 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Training Diffusion Models with Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:48:57.308257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:b222959955edea1df8686e74cb2b35a17dc300aa06d22a972c77bb29cd0306a9

Observation ccfbb829-fc36-466e-acac-6930ea490396 · outbound

This paper cites A sur- vey on generative diffusion models.IEEE transactions on knowledge and data engineering, 36(7):2814–2830.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? A sur- vey on generative diffusion models.IEEE transactions on knowledge and data engineering, 36(7):2814–2830

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.842876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:24f0c0c88bed1fd33e920ad18b18ac0dd558c9a9f177cd555d64f6a75202fdcb

Observation 1cd3e7a2-3dc1-4e11-b13c-2c4d5a69fbd2 · outbound

This paper cites An Overview of Diffusion Models: Applications, Guided Generation, Statistical Rates and Optimization.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? An Overview of Diffusion Models: Applications, Guided Generation, Statistical Rates and Optimization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.302386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:49b914bc1c40a410e0492d0eac4e1f706e0e273bc240dab8e2002f8a33e20c31

Observation f4680a21-ac6f-434d-bdbf-2c55ae1100dd · outbound

This paper cites Dif- fusiondet: Diffusion model for object detection.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dif- fusiondet: Diffusion model for object detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.841045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:4f38303a11c13e060b6c7b7d4ee327004450ad2187f0c92e71e508e7a10fbad5

Observation bd2739b0-60cc-4a61-b68e-89f31175bda5 · outbound

This paper cites Diffusion models in vision: A survey.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Diffusion models in vision: A survey

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.839134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:f3095d5bc5088776696bb9797c2794a4b18794d814b0b82fdf4f24c3673ca3d5

Observation e39bdd7b-6892-43cb-a10e-85be05c618c2 · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:48:57.332773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:1226896634d4931a9f87d41689f1ebab962b0024c310ede2463ab35b470b3bc7

Observation cc18749b-a09b-4589-be41-c605af67c372 · outbound

This paper cites Dpok: Reinforcement learning for fine-tuning text-to-image diffu- sion models.Advances in Neural Information Processing Systems, 36:79858–79885.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dpok: Reinforcement learning for fine-tuning text-to-image diffu- sion models.Advances in Neural Information Processing Systems, 36:79858–79885

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.837363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:13667de5073296c692074d95104ebac31a740f7add25b9a039f2878c7eef5dec

Observation 135ba8d6-484a-43f4-873e-e5abd472cc61 · outbound

This paper cites Re- inforcement learning for fine-tuning text-to-image diffusion models.Advances in Neural Information Processing Sys- tems, 36.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Re- inforcement learning for fine-tuning text-to-image diffusion models.Advances in Neural Information Processing Sys- tems, 36

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.835544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:a376ebc752d2639799d3155de027525ecac245ff2cce2ec48ec992be6d96603c

Observation 2e0a1158-f96f-4f29-8c63-2cc891fbc0e3 · outbound

This paper cites Reinforcement learning for generative ai: State of the art, opportunities and open research challenges.Journal of Artificial Intelligence Research, 79:417–446.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Reinforcement learning for generative ai: State of the art, opportunities and open research challenges.Journal of Artificial Intelligence Research, 79:417–446

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.833747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:9392709dfd7f1e27ac2635f112612c1658400d135867e35c8ef882ba0ec6a200

Observation 2545ebc2-28a1-4326-b73e-c60f6f9a2b8f · outbound

This paper cites Re- flective policy optimization.International Conference on Machine Learning.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Re- flective policy optimization.International Conference on Machine Learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.831612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:ca532750659bd8c2a423b65b23453b15159eece8a0c9d078ac9f2cb4b07f5841

Observation 3b7a36d4-23a0-49ab-8fc6-b763b125d5b4 · outbound

This paper cites Scaling laws for reward model overoptimization.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Scaling laws for reward model overoptimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.829903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:6f0e6e95dae3e003ffa115c77db7fda78a87d959a5ba0feffafc8bcb2e080cca

Observation 2fdfeb35-f613-4952-8c49-d2c8a44c354d · outbound

This paper cites Integrating Behavior Cloning and Reinforcement Learning for Improved Performance in Dense and Sparse Reward Environments.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Integrating Behavior Cloning and Reinforcement Learning for Improved Performance in Dense and Sparse Reward Environments

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.317233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:f7066a3efef57615d085f6e33279552a50f6389b4653556241b0649e10cd362b

Observation ce1b4e95-c112-4510-8d56-4328a4dbab63 · outbound

This paper cites Dealing with Sparse Rewards in Reinforcement Learning.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dealing with Sparse Rewards in Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.311552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:b9b73ce6e50794d0bfdfc52afdf98cabe08e0b6315db3fb648642766088bd8d9

Observation 2fcf6550-5140-41b0-a1d7-923102585f33 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in Neural Information Processing Systems, 30.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in Neural Information Processing Systems, 30

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.817030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:789fd72a0f9823bc0c46a95de87c246936b189b95c331cfe8b15ab82262d4f04

Observation 1a24358b-e2d5-47ca-a75c-33039563bae6 · outbound

This paper cites Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.821871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:67f8fffbb43862c3faddd3c59fd1f53e7de7245077cdd2050b6b73cad3302711

Observation 03ce1bc0-f3c7-4728-9eac-b4c1c98b4d93 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Imagen Video: High Definition Video Generation with Diffusion Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:48:57.341223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:96c21e4953dcbaba2d459e7edcb02dc3385e7a680afa6da9f57fcc63c2cc031e

Observation 820bd6b8-1e48-4294-b46e-a63293967081 · outbound

This paper cites Video dif- fusion models.Advances in Neural Information Processing Systems, 35:8633–8646.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Video dif- fusion models.Advances in Neural Information Processing Systems, 35:8633–8646

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.828131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:170de3ac2c2d7f3a2ab7bcd1faf11b56d2937b8a10cb493bd960294deda55e70

Observation feb5cc73-2e30-4ccf-8c7a-595a64f29687 · outbound

This paper cites Reward hacking in reinforcement learning and rlhf: A multidisciplinary exami- nation of vulnerabilities, mitigation strategies, and alignment challenges.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Reward hacking in reinforcement learning and rlhf: A multidisciplinary exami- nation of vulnerabilities, mitigation strategies, and alignment challenges

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.798412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:33058eb82d3b15594c9236a501951e9b4a99271f8a2033385532a840caac3619

Observation a3325f06-d978-4d3b-94cc-1951244ba7e8 · outbound

This paper cites Dif- fusion reward: Learning rewards via conditional video dif- fusion.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dif- fusion reward: Learning rewards via conditional video dif- fusion

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.800132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:71e811d3ad1aef24e9fc2db5da038736a9052245b1cd1b782d31bdb69225dfbd

Observation 24a48e48-3fb1-4844-9d57-78672cc942a7 · outbound

This paper cites Diffusion model-based image editing: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Diffusion model-based image editing: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.819867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:6351e5f2abebc5480eca58c031d8b4fbd53c47b4bbd6146635d8864a292dcb9e

Observation e0bfdc6d-64e5-40ec-bf8f-0591e3837307 · outbound

This paper cites Measuring Diversity in Co-creative Image Generation.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Measuring Diversity in Co-creative Image Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.338492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:db6cb999d0b987938f884dd4316a1fd3d11dadffa5d99a68d9668ac8cbe5d062

Observation e9e903b1-bb29-417e-bbd6-b938137d6e03 · outbound

This paper cites Holodiffusion: Training a 3d diffusion model using 2d images.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Holodiffusion: Training a 3d diffusion model using 2d images

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.796486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:32bc1971633c0c56e558208c70d29ac2b35b239dca6ead618cb13ad3d0ce17e7

Observation 659cff2e-f249-4ff5-9109-cf4335076c5f · outbound

This paper cites Test-time Alignment of Diffusion Models without Reward Over-optimization.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Test-time Alignment of Diffusion Models without Reward Over-optimization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.344161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:cd44262608d7f11f00250b5ad76c264a9b0df6cdc72ba00c48bfa5258a1254f1

Observation 8d51bc0f-df24-4786-b634-890e7d98c517 · outbound

This paper cites Variational diffusion models.Advances in neural infor- mation processing systems, 34:21696–21707.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Variational diffusion models.Advances in neural infor- mation processing systems, 34:21696–21707

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.794774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:b66f04c4567c8f75664e694c1c48bde572b01e51cfa2143a7cef177dbb52f4c6

Observation a8720526-3bcf-4794-87e9-993099d1ca9c · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.Ad- vances in Neural Information Processing Systems.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Pick-a-pic: An open dataset of user preferences for text-to-image generation.Ad- vances in Neural Information Processing Systems

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.790557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:c2409af02f4f275337d0241791f8baf2749fcf2fcb9ad8bf4d2097ca451eafdc

Observation 4fd0689c-1a7a-4ba4-8c8a-ff92eae7f2a1 · outbound

This paper cites Improved precision and recall met- ric for assessing generative models.Advances in Neural In- formation Processing Systems, 32.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Improved precision and recall met- ric for assessing generative models.Advances in Neural In- formation Processing Systems, 32

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.824089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:69359406dee1b1670cfd27faf65c08f55acc576beaf54c2d291f4939d1aab4f6

Observation 9335e1c7-c2a3-47d1-b080-5b63029c8805 · outbound

This paper cites Aligning diffusion mod- els by optimizing human utility.Advances in Neural Infor- mation Processing Systems, 37:24897–24925.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Aligning diffusion mod- els by optimizing human utility.Advances in Neural Infor- mation Processing Systems, 37:24897–24925

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.826112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:2454c1b7dac8c459648d9f34fc97fb4c10c3c248615b1984fc648d28b39238d3

Observation c77785a1-63fa-4b32-874d-9484b904a438 · outbound

This paper cites Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.326419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:5c28a9c00d016c1d6bdbb17982abd3adff25539885bc2817e9f29051a85e4ccb

Observation 447ea7be-bcc3-40ac-95bc-2aa795f89557 · outbound

This paper cites No-reference image quality assessment based on spatial and spectral entropies.Signal Processing: Image communica- tion, 29(8):856–863.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? No-reference image quality assessment based on spatial and spectral entropies.Signal Processing: Image communica- tion, 29(8):856–863

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.788614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:1bd4eee86b243e1bebb894c84d9f43d30dd543b6940d432223fb051a414cf5ed

Observation df099ab6-8009-494e-89b6-045612f2282c · outbound

This paper cites Deepcache: Accelerating diffusion models for free.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Deepcache: Accelerating diffusion models for free

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.777784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:40e6b63bfee3ccf97b4438a659b71f4910f50f6952f068473bf204148e565ad2

Observation 0a993b5e-2a98-4313-9795-8538109036db · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.779881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:aca1d5a3be8f03a44a9b9265c7d799c18bbfb44c6b36a04595aa5f6e96562921

Observation 737cf6e0-d5db-45d4-afdd-feaf4fb2bcec · outbound

This paper cites No-reference image quality assessment in the spa- tial domain.IEEE Transactions on Image Processing, 21 (12):4695–4708.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? No-reference image quality assessment in the spa- tial domain.IEEE Transactions on Image Processing, 21 (12):4695–4708

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.775807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:f534431a9f036898785f09c7eb99cbb6b2c4ab59e25a0a190191ecd82ed93b30

Observation 9805fa92-95c3-489a-b8a1-15e769819751 · outbound

This paper cites completely blind.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? completely blind

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.773502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:43c6d6ca86568b805874f1420a0275364f8faa88faf348d75081ac3bbba3fedf

Observation 27547f44-9da5-4373-9bda-b4c76be63de9 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:48:57.335429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:e789077bf10e5f386603d99c1e92f49cf4c0a8aeacd9e76bab9f8b8abc4c14ce

Observation dbf1b86c-b20a-40c2-b057-910954ff97af · outbound

This paper cites Efficient Controllable Diffusion via Optimal Classifier Guidance.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Efficient Controllable Diffusion via Optimal Classifier Guidance

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.286646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:735612890d0e6938910c0cbff553f43d7ccbc781c75fa6c00edc1d2f4e0f869d

Observation e7a12db2-04fd-4761-8fe6-fed9875f37ce · outbound

This paper cites Markov decision processes.Handbooks in Operations Research and Management Science, 2:331– 434.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Markov decision processes.Handbooks in Operations Research and Management Science, 2:331– 434

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.782021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:bfb7e3f49bbfe4f108cb9729218f1f92d8564bcef3a8bd09d6395601dcbcae3b

Observation 407fdf8c-d242-46cf-b305-9554e7a866f5 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Learn- ing transferable visual models from natural language super- vision

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.784099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:4f6be2c38a2879f9f94fbf2e4614c85a3c1cf88049e46b85f7f05dee4c12b59b

Observation 2e3819f6-0f8c-4acb-bf67-aab2c2bd08f7 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:48:57.320043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:4b1afb196b906369e16725aa3b262980dc773f72133fb98ee009e22c809c56da

Observation eb0541e8-4aa8-42d8-8158-e5e48ba79560 · outbound

This paper cites Learning by playing solving sparse reward tasks from scratch.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Learning by playing solving sparse reward tasks from scratch

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.764453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:00f66ae90b54ed4ac00a4f2f65eccc1d070232108e2e7ee662824e7b5f4b680d

Observation 3f73eceb-cbfc-4da3-ace7-5ce04fe79f1c · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? High-resolution image synthesis with latent diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.766741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:8eaa878a8cbcf0dbfe28b8bd17256440d4a3b25d1d819135b4ee8741e7866835

Observation 6b4474c8-117b-4f4d-8e24-e7d418a1d37b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? High-resolution image synthesis with latent diffusion models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.771392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:2993418325f9610883c7c905543ed0a94a56d0528649e00c9d7132bc2e09b817

Observation a67cc0e1-303e-4690-822c-dd5590b4512b · outbound

This paper cites Silhouettes: a graphical aid to the in- terpretation and validation of cluster analysis.Journal of Computational and Applied Mathematics, 20:53–65.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Silhouettes: a graphical aid to the in- terpretation and validation of cluster analysis.Journal of Computational and Applied Mathematics, 20:53–65

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.744326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:d0b76d461128c978db9c57457e3381e51873b7fa2f66aa5175ea30a0040de360

Observation 1128b4b9-5073-4f4d-81ae-1a46f13ba6ed · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.757268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:1ad80ab16e1fabcd33ed69ae75ac7181513d73ecc3795c4068018b7e70b51f50

Observation b15f4320-755c-4cb4-a9d8-9f5089e23cf9 · outbound

This paper cites Improved techniques for training gans.Advances in neural information processing systems, 29.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Improved techniques for training gans.Advances in neural information processing systems, 29

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.759596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:623479846e1076a9a987ccffe2d35ea40ebbf4906695182b01b140ab099a673d

Observation b5158451-e2c6-4c5d-8726-e0b2e8c470b9 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural in- formation processing systems.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural in- formation processing systems

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.754912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:85fcaa070115f4baeaf3112dc341b82bd41fff2e5aa722a79f6c3a095578c0ac

Observation ec9f9c63-5958-4602-85dc-abe14a9b4ca9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Proximal Policy Optimization Algorithms

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:48:57.329746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:8555c852ed5e223c801eee74833a073172d3e4d8b752619edcb77fe0f37c75c1

Observation 77fb2339-f8e7-420c-8b4f-efd13fe239aa · outbound

This paper cites Defining and characterizing reward gam- ing.Advances in Neural Information Processing Systems, 35:9460–9471.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Defining and characterizing reward gam- ing.Advances in Neural Information Processing Systems, 35:9460–9471

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.762057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:2380357b5f3630817408571c25b3a262de56563c0cd09466f65d01897e98b893

Observation cd81aee5-8a30-43df-bf2e-624c424c7ed1 · outbound

This paper cites Defining and characterizing reward gam- ing.Advances in Neural Information Processing Systems, 35:9460–9471.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Defining and characterizing reward gam- ing.Advances in Neural Information Processing Systems, 35:9460–9471

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.768978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:83167ca34b382aac10a39e39495c14e1d0405604489578f461084e4a7a6a4ed1

Observation 0b920364-e15d-46ee-8616-8ee9ec7651ae · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Deep unsupervised learning using nonequilibrium thermodynamics

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.786570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:5ab2b92f3858a80547f3256ee582bf3f151bb2483d4cec3cce6971ebfbd4ea3a

Observation 0444539c-d75f-4c77-85a3-acfdf72f57b5 · outbound

This paper cites Inference-Time Alignment of Diffusion Models with Direct Noise Optimization.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Inference-Time Alignment of Diffusion Models with Direct Noise Optimization

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.291608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:077ab4a7547e347140b39e437c37c6991f3de991eb86a7c377639defce04d960

Observation 63087ef6-f5ab-4105-ad07-bbb4736bc16d · outbound

This paper cites Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.305447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:ffcf23d27808c0954438bbb803aaa6b1d1d08d2e62c36e5d689533687406af0c

Observation 2cac8b13-bfa6-4075-8bd0-de36f0c05fc2 · outbound

This paper cites Visualizing data using t-sne.Journal of machine learning research, 9 (11).

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Visualizing data using t-sne.Journal of machine learning research, 9 (11)

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.752447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:f4dad2df6de6c9eb1eea6980467b06283d1e12c746501182b061ea72b9f4774a

Observation abda76e1-ab2e-4fc0-bd6e-9d5f5107b27c · outbound

This paper cites Diffusion model alignment using direct preference optimization.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Diffusion model alignment using direct preference optimization

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.749288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:404b01b691df96191b5c8722f13afbd2bb02a7ac4c1af1c11843fada0f29c80f

Observation 4c3336c2-a1ee-433d-9678-22bbdf32d63d · outbound

This paper cites Deep-reinforcement-learning-based autonomous uav navi- gation with sparse rewards.IEEE Internet of Things Journal, 7(7):6180–6190.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Deep-reinforcement-learning-based autonomous uav navi- gation with sparse rewards.IEEE Internet of Things Journal, 7(7):6180–6190

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.746932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:7b1419842e8d7b19d58afe325ff55b635d06e9c276df1dc8f1121137369307d2

Observation 257b259f-93f7-4fd9-a2f5-e179c0a14efc · outbound

This paper cites TEAM: Temporal-Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? TEAM: Temporal-Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:01:59.936669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:6b79ee5482d3582cf8d4f935883a6c8889af094c7facd9cc62e368506dfdbd4a

Observation 5eba3af8-bb76-46be-8b49-81d7b20884a7 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:48:57.298828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:3640b4c11f5f538d3cb0fda0048758494ba9a1a0c48ca79940aea579ac998a20

Observation 307ba49e-b2cf-4914-a7f6-4bb9562866c6 · outbound

This paper cites Diffir: Efficient diffusion model for image restoration.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Diffir: Efficient diffusion model for image restoration

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.880287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:2175c615a3e973570f6768d8c0a841b3f37821e9823682e402fcd77b2b9185cb

Observation 466da5dc-a5ed-4ebc-8ac9-0b7cbe1262b9 · outbound

This paper cites Dymo: Training-free diffusion model alignment with dynamic multi-objective scheduling.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dymo: Training-free diffusion model alignment with dynamic multi-objective scheduling

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.882159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:10315eadf640267b5669c81965d82791bb14f70312d5999ec4a79440591fa186

Observation b49e4063-6029-4b37-943f-e0f5cfc219dd · outbound

This paper cites A survey on video dif- fusion models.ACM Computing Surveys, 57(2):1–42.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? A survey on video dif- fusion models.ACM Computing Surveys, 57(2):1–42

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.874259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:ff6b4373e4bdf14e87da4a3f8dc6ad2906882a19892375a2e52405ea59ab96f9

Observation 9c3585fe-e11f-4e7b-95a5-af1d66069cc0 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.876157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:61ee13984e298690b2cf7b0015fdc48fca302558b839a7ab252daddfd88a29dc

Observation 0064e5d0-b387-42e0-a4f7-cf8f0492bd0d · outbound

This paper cites Dream3d: Zero-shot text-to-3d synthesis using 3d shape prior and text-to-image diffusion models.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dream3d: Zero-shot text-to-3d synthesis using 3d shape prior and text-to-image diffusion models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.878023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:79dd9daaebe1331811fb59c74dced2c81ae03e05ea850a33b061a9338667c65e

Observation d451510c-e91f-4077-90e8-442cce488153 · outbound

This paper cites Versatile diffusion: Text, images and variations all in one diffusion model.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Versatile diffusion: Text, images and variations all in one diffusion model

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.883994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:bb8270a96562cd03f40ca79bb42df99a3a36063b64c02d64051148f2684c328a

Observation bb487ac0-3dd8-4dc8-9d17-1ab800ffaca7 · outbound

This paper cites The Exploration-Exploitation Dilemma Revisited: An Entropy Perspective.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? The Exploration-Exploitation Dilemma Revisited: An Entropy Perspective

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.295846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:c21beff33cf90f5baf227c43cdf461f74ed4d1f482afa598036d84aa261edaa5

Observation cd1c5c24-7720-4129-95ea-f650676ffa48 · outbound

This paper cites Entropy-adaptive diffusion policy optimiza- tion with dynamic step alignment.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Entropy-adaptive diffusion policy optimiza- tion with dynamic step alignment

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.865650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:c7f19d2bc87ebcff353f8bfa99410cd53f12f8be0b64aa84cc0215f328c8a4eb

Observation 9c3382a6-86ba-495e-b24b-07294f3eaa40 · outbound

This paper cites Using human feedback to fine-tune diffusion models without any reward model.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Using human feedback to fine-tune diffusion models without any reward model

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.867425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:ab81b505b8931d6e579ae3f566a72d16e200f6c9f1ca66f548fab07584ae4e59

Observation 1477834b-e9f6-49f8-8ff0-89fbe958803b · outbound

This paper cites A novel multi-step reinforcement learning method for solving reward hacking.Applied Intelligence, 49 (8):2874–2888.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? A novel multi-step reinforcement learning method for solving reward hacking.Applied Intelligence, 49 (8):2874–2888

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.869645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:234554506d7ee368844a526f8f6e40ded70ab2f7692ad391a400504a2f988505

Observation 3259f9a6-1279-4ed6-985d-a7fc22732510 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? The unreasonable effectiveness of deep features as a perceptual metric

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.862026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:83edd9ad898ff79473522f093ac3ec97bd7b23a3a98f8b1e227daacee69e2842

Observation 58892d7a-8ed7-445e-9942-f925e4f15021 · outbound

This paper cites Confronting reward overoptimiza- tion for diffusion models: A perspective of inductive and pri- macy biases.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Confronting reward overoptimiza- tion for diffusion models: A perspective of inductive and pri- macy biases

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.858446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:67bd6d9d52895adf62095106363fa0db606fc7948aaca1e14c2b1be37d39e8d1

Observation c186831c-21fe-4e44-a8a6-00d7b8825565 · outbound

This paper cites Alphaholdem: High-performance artificial intelli- gence for heads-up no-limit poker via end-to-end reinforce- ment learning.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Alphaholdem: High-performance artificial intelli- gence for heads-up no-limit poker via end-to-end reinforce- ment learning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.860290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:26aa49f1ad1eba1fcbe20bbd29d2d6f9c85e98d1553f287a5b7dcc328bc98b90

Observation 89ac76cb-2e2d-47d0-aca5-6887a594870f · outbound

This paper cites 3d shape generation and completion through point-voxel diffusion.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? 3d shape generation and completion through point-voxel diffusion

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.863765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:fbe29f6adedde43346a174710b064a6b44d040417441a2ac6bfd985fe960a30b

Observation 1da7ee14-bb16-44f4-9687-cc58c01966fe · outbound

This paper cites Mixture of global and local experts with diffusion transformer for con- trollable face generation.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Mixture of global and local experts with diffusion transformer for con- trollable face generation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.871500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:abc0b879d234aae5c7d9121020d5ab58efecf0d8d8541e2ffca3d8a053475a48

Observation b61469d5-4d07-4b3c-a32c-238dde804236 · outbound

This paper cites RL Fine-Tuning in Diffusion Models Existing diffusion models [4, 5, 24, 59, 62] primarily approximate the data distribution through denoising reconstruction loss.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? RL Fine-Tuning in Diffusion Models Existing diffusion models [4, 5, 24, 59, 62] primarily approximate the data distribution through denoising reconstruction loss

Reference 72

Resolution
malformed identifier
arxiv_id, observed 2026-05-20T19:48:57.314460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:732ee2efaf044adc218f63512004057ba13ab1c21cb67b2ebe955ca81bc351a4

Observation e5304a71-38a6-474e-b7c1-7228977c2d35 · outbound

This paper cites •(1) Visualization Experiments.See Fig.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? •(1) Visualization Experiments.See Fig

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.854117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:62fb367646775e00bf34edae5713eb825ebf713f66481729815ac26efb887d10

Observation 565b7ff3-7f0e-43a3-a3a0-ff88e0c7121f · outbound

This paper cites an unresolved cited work.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:48:57.848394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:3be38cf47c13902f36daad2b9c9f47c724ffb14cc909168263254ac0506260a9

Observation 499adcd8-b8ad-451c-8512-8c13e58f7071 · outbound

This paper cites Proof of Theorem 1.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Proof of Theorem 1

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.850498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:e593e9692aa76c8fb71f3da4deffd7cef55f29e6154484278434ead5d5ee22da

Observation 068f1f5f-4784-4562-83e5-e23d0104f116 · outbound

This paper cites an unresolved cited work.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:48:57.846447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:922517ff34bd794bbc7be012ecbc505fd2fe3d21d21f162f677cf4ec779c13c4

Observation 356c7317-dbc0-47eb-ae78-f52bcb9075aa · outbound

This paper cites Therefore, componentwise, Cov x(i) t , x(j) t+τ = √¯αt ¯αt+τ Σij + r ¯αt+τ ¯αt (1−¯αt)δ ij.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Therefore, componentwise, Cov x(i) t , x(j) t+τ = √¯αt ¯αt+τ Σij + r ¯αt+τ ¯αt (1−¯αt)δ ij

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:48:57.852435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:01d8bc8aa41425e89eb641ee016e7b4a6c40e731b9f74c4874d42e583745089a

Observation a2d4b130-1e26-4dc0-9b74-8d9874997bac · outbound

This paper cites an unresolved cited work.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:48:57.844618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:9b8d840120638ee92ae18ba926f04db776b5c9d236d187999135b9c2a4720029

Observation 499868ac-56b4-44b2-842b-e36e7024ba94 · outbound

This paper cites an unresolved cited work.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:48:57.855751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:d38a350a4b229c940b89502e56ad959973b37cd1effba17a016e750532365f91

Pith citing papers

Observation 800e6608-2c88-48a3-80f0-49efb6a095a1 · inbound

RAGOCR: Optical Compression of Retrieval-Augmented Text via Visual Representation cites this paper.

RAGOCR: Optical Compression of Retrieval-Augmented Text via Visual Representation Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:23:55.517186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T00:23:54.855176Z digest=sha256:5043b1c782ab329ebc48593156b6606df47c2a03f9c4ab52b71d32de74989476

Observation b303e609-b644-420d-8275-3eb5c8c5c215 · inbound

Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models cites this paper.

Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.884241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.884241Z digest=sha256:049099c902e0ed9e14420b149d677971b076802138270810e1ab8d769201c4f5

Observation deabb4af-042c-4d5f-be6f-48b1b3f49ed3 · inbound

PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model cites this paper.

PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T14:34:44.686328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:34:44.686328Z digest=sha256:7ea0eab05422e94e014b194edb1ac8fde4c4f07c7b1d4eee7891d67ef7faeb42

Observation ef9ea5d0-b9ea-4b21-b32f-8d6691654e0b · inbound

Diffusion Image Editing via Asynchronous Token Decoding cites this paper.

Diffusion Image Editing via Asynchronous Token Decoding Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T19:27:32.280009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:27:32.280009Z digest=sha256:9555ff130e8a6945a42634234b130530d09c0ef39c9d867efb42aa6d0a1c32b9