Pith. sign in

Paper Citation Record · LEDGER

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF

As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2607.07693.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07693 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T02:17:20.589485Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact35
  • verified fuzzy4
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch12

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1fe76140-f202-44d1-b07d-16bf1467641a · outbound

This paper cites Hindsight Experience Replay.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Hindsight Experience Replay

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.909974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:563cdf73f68607551694012fded27e92f0f23e3d10d27ae8b7a3596c181042ab

Observation e304a2af-3e64-4c16-9c44-bb570395a8ca · outbound

This paper cites 133011, 3, 7, 8, 12.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF 133011, 3, 7, 8, 12

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T02:25:56.026565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:a1ba5ae9c0b954cedb1805d99b0086faa207c4343acfc19149dcf854c2229d6d

Observation ad5a402c-9aaa-462e-bb2b-1a204b96b0e6 · outbound

This paper cites DSPO: Direct Semantic Preference Optimization for Real-World Image Super-Resolution.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF DSPO: Direct Semantic Preference Optimization for Real-World Image Super-Resolution

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.845525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:cf2e4ded77cc6545c975428bcd1db0ed877dc13ea512a8f2640de3ce2d7e4a0b

Observation faf7892d-89e3-4a22-8b06-16ddceeab3c7 · outbound

This paper cites MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.876895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:51bad89188c82dccbd293ac76145ea0566be7853c9cde6fb3ca550f9530c3e42

Observation 2fc94259-9502-4c79-8de2-e8f5b0dec446 · outbound

This paper cites In: Proceedings of Robotics: Science and Systems (RSS) (2023) 2.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF In: Proceedings of Robotics: Science and Systems (RSS) (2023) 2

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T02:25:56.016587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:f933afb36176c5c5cc09cc6524af53d22ab74b2a78148cb4f8fa3d9503f9972f

Observation 20f0dcde-2523-43e2-83b4-599e538aedfd · outbound

This paper cites Deep reinforcement learning from human preferences.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Deep reinforcement learning from human preferences

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.932859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:33d944debcc60fc03a8d924ba3f7818a6636c151e34a59c6013693d050149f34

Observation 1c41d852-001c-4642-8b42-97bc12a136e5 · outbound

This paper cites Directly Fine-Tuning Diffusion Models on Differentiable Rewards.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Directly Fine-Tuning Diffusion Models on Differentiable Rewards

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.930343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:9cff576734d438535c07786f34cfb8280c4dd688ce426bfcfb902b3f61daecf0

Observation ec6b81b7-a673-4974-8039-af0e933faa18 · outbound

This paper cites DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.947272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:4f1ad0dd9b33b05f1a2bc27ade88848a5d3d8f383b99130274eba1309ceec106

Observation 325b6d17-b4e2-4aa9-89b5-3c411dbdab38 · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.912565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:f8c841c83af1b38011a185b8f3dced76d4b804159a2ff72b85db6c76def9c407

Observation 4bd34cc9-7b99-4330-b53a-b71af152eced · outbound

This paper cites TempFlow-GRPO: When Timing Matters for GRPO in Flow Models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF TempFlow-GRPO: When Timing Matters for GRPO in Flow Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.870955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:faa9d3a5f21fba95269a3b616ab74034d71b2dfadbcc39cc97ffa14f9c3a4e9a

Observation d3bf136e-8151-4f30-b16e-adef5d24e515 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.831004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:7e81eea1c118cc4485e55ea1fa123f1f91c132291c8446a5c4ac5baa91d5165f

Observation c31117ed-4527-44e5-95a5-b99a77ec8b08 · outbound

This paper cites Denoising Diffusion Probabilistic Models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Denoising Diffusion Probabilistic Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.937670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:c24f50f1a109f02bee2dd2f810103c408a9d4d8169d0c12ac2dc00905b808887

Observation f375009b-90de-4073-8196-d8853fcb7f82 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.865485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:09ad133ce7706e7fb50e5af7c264cf3f8e4d3a62a3cf542a44d9f5ba9d802528

Observation 9fd710a6-780b-4245-bb15-a54709896175 · outbound

This paper cites Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.902473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:b4794ad05dde67bf3e8f2e7e12d362bbbd2ac3479918fa009ca1c4c95c7fe52e

Observation 4c63c6ed-0eac-4cfe-9f67-2f0f138a3667 · outbound

This paper cites PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.839812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:7c7b1b6bd418ef0a608ec438ac198130017802dc2923b1d947a2b11a65090e13

Observation 1ce6d9bd-f0e1-424a-8712-342fe507c107 · outbound

This paper cites TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.887521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:3a88065ce195de30222e81d5a07770451a26d45e780de87ad4ef2adda3d328fa

Observation f0d7ae0d-6bd3-4e43-b65f-4847ff0c983d · outbound

This paper cites Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.862799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:f8ba45940fab49e1fc40a072fc3d501608f9bf87ad568ba033b74c3c7931bbf3

Observation c1b1dd6e-26fd-4bf5-aedf-6fd2ee8daacf · outbound

This paper cites Branchgrpo: Stable and efficient grpo with structured branching in diffusion models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Branchgrpo: Stable and efficient grpo with structured branching in diffusion models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T02:25:55.895022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:ad26d36e982b4a0f8b64b7e1539e1445395055cf3be4a19284b7c23ad2257702

Observation 95e0288b-a9ed-4d78-9dfc-95742979b630 · outbound

This paper cites an unresolved cited work.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-07-09T02:25:56.028280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:915006f9ff0c7fb10d8b876e196b1ef9e2ac845205c09d28fe14de6367a120cf

Observation 527924e1-c108-4dd9-a2b5-9f871f963f3b · outbound

This paper cites Continuous control with deep reinforcement learning.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Continuous control with deep reinforcement learning

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.935385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:adb43f7c07754c01545e362c646758051d213dd9b2bc44708b7e3e92dde9ee68

Observation 372b9069-a4e0-49ee-8396-2bfb6cae3fe0 · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Flow-GRPO: Training Flow Matching Models via Online RL

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.899907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:6ff98b845154849209ea8e8412a591a298935642a501e00162bc34dd52fcac59

Observation 35f1aebb-ca82-444f-8f49-795842f1c964 · outbound

This paper cites Synthetic Experience Replay.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Synthetic Experience Replay

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.942485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:64ccdf8461f15a072fc5bb14e4127cd6899ee0369803b2924dac8b052f6bf4a0

Observation 82a6cc54-75ae-400f-aa8f-12b51ae1a4d5 · outbound

This paper cites Dual-Process Image Generation.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Dual-Process Image Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.904789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:d00b1dda64e26ceb06e7515a7c65aee033e44fe38a1dc6bca37c625cc7de4374

Observation db1ce687-aace-4e05-8e87-077bb8278092 · outbound

This paper cites an unresolved cited work.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:25:55.842469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:7f2079ad91924df3db825ac03d8c2e10d73ca6929f365cad082bfedd5d625d21

Observation c9c3d08e-3131-45c3-820c-7cceeb301579 · outbound

This paper cites Do text-free diffusion models learn discriminative visual representations?.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Do text-free diffusion models learn discriminative visual representations?

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.914949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:fee2384ac53938dc99c5cf9b28ca817bed1170e8b1bb8a508f310f18dcbca0f1

Observation c79984dd-eab4-4ca8-8b7e-107df3a0cbd1 · outbound

This paper cites GPT-4 Technical Report.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF GPT-4 Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.857250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:0c8bceccf90c5666d0f994e5d2ede1f54716f17f24c871561d93c811f92899de

Observation 7c6223df-d4d5-4464-85ea-da6e1faa76cd · outbound

This paper cites Training language models to follow instructions with human feedback.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Training language models to follow instructions with human feedback

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.940191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:7e65017dd4e337617d2028d9e7a16a55bb5ef2df08caa0eae244056dd4afe076

Observation f0b870ad-7af4-42ae-b76e-37848e72aaf4 · outbound

This paper cites Reinforcement learning by reward-weighted regression for operational space control.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Reinforcement learning by reward-weighted regression for operational space control

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T02:25:55.802879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:18b9e2f0e6c9b347e6f9162a6465393ab57f3040e4f9a98870a10e25facec200

Observation 6a2bf9f6-84dc-4e57-9ede-0145bc00a5dc · outbound

This paper cites Hard examples are all you need: Maximizing grpo post-training under annotation budgets.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Hard examples are all you need: Maximizing grpo post-training under annotation budgets

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T02:25:55.907595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:5e52ae76a620563c35e19559ee736893ef6c8f1e9e1f87ff1447f81aca3fd309

Observation 2f03f31d-7a68-4176-a776-c54ed1bf771c · outbound

This paper cites Video Diffusion Alignment via Reward Gradients.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Video Diffusion Alignment via Reward Gradients

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.922322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:bc1c143199d96bfc86a45038fd3efbc65e39cc6c48f14dd44558077f568ab50c

Observation eac7afe5-338e-4119-b7d7-9c0af4a5c505 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.882647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:d085aa5dae78086395cd34d654865bf8120ca9a145a29ff91f10e92d36c6a18f

Observation dc409362-ad24-4ba8-be55-e43b05569ebd · outbound

This paper cites an unresolved cited work.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-07-09T02:25:56.024965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:8ce25daa7c16ea36bba4cc34858fce41609ccf1307a0c566d63e526586202baf

Observation bc044944-4a3d-43b5-a8ea-32f03a7eb92a · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR) (2022) 7.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR) (2022) 7

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T02:25:56.023336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:42a31fe49b2e262fb761452801739f9c1e5e132bb63b0836b833c3492b0268ab

Observation 07e2a8de-2d17-4a3e-a807-cd8bff3b5774 · outbound

This paper cites Prioritized Experience Replay.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Prioritized Experience Replay

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.880064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:929030a6aa395701a02906f222fc398c34ec321f91d7cc233dee4e9d9fb7e038

Observation 9147be8d-01af-4b19-ae1d-93fede2baa8f · outbound

This paper cites an unresolved cited work.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-07-09T02:25:56.021586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:3d8297099dac96b1a8dcac7cb729fa2e241f2bd920ead50d4df2b76adcb34c4e

Observation 2b0dea72-f54f-49a0-9830-0a0fdc3f8114 · outbound

This paper cites Trust Region Policy Optimization.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Trust Region Policy Optimization

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.927244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:1ccc0c87c73b04fed52f8c9f56e9fe2cb0fa758ae42cf8fd49f128a3cb957aa1

Observation d64aa1d4-622c-4a35-a9b3-e0c503653fdc · outbound

This paper cites Proximal Policy Optimization Algorithms.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Proximal Policy Optimization Algorithms

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.836731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:2060ea4367c07054801a152aa810bc33e90b28d7074ce38206ee179e294b7f65

Observation 48fe61e6-bdc9-4fd8-892c-99ccbf491efe · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.854539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:8f3b4b3f20544b3b7fba6d7c233920a28d24145b6a923d9d79921ede989f7ac0

Observation 32a72b49-ff3e-48c1-b5b7-a46db07da39d · outbound

This paper cites RL's Razor: Why Online Reinforcement Learning Forgets Less.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF RL's Razor: Why Online Reinforcement Learning Forgets Less

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.889858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:64c6813c2264ad77682857519573073fb6f697905f288818fbc1f00edeb1b7be

Observation a55cffbf-3e07-48f3-be56-bef4b0ec943e · outbound

This paper cites Training Region-based Object Detectors with Online Hard Example Mining.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Training Region-based Object Detectors with Online Hard Example Mining

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.848620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:503e85dc57c4d904a0b84aa0a955e61d4467dfee9ebd436c6c57ccfa432b79a5

Observation eaece74c-41df-4db0-a57e-c698a7874e27 · outbound

This paper cites Deep Unsupervised Learning using Nonequilibrium Thermodynamics.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Deep Unsupervised Learning using Nonequilibrium Thermodynamics

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.834022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:1b6483ea56f36288aafa49b4a51946fe767ba0ed4ee57e232e45f14eeff23380

Observation d7dd12aa-10d5-4484-a49e-bcbdb5baa0e4 · outbound

This paper cites Denoising Diffusion Implicit Models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Denoising Diffusion Implicit Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.885077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:4c0268ea2687fe00f89183853a69e82699bbac31b371321253bfcb18c51e6e19

Observation dd3d5b43-cf21-4969-a33b-b6cf14d54ed2 · outbound

This paper cites MIT press, ??? (2018).

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF MIT press, ??? (2018)

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T02:25:55.796181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:a0a71dfe92ab9e11eb817b96c78feb9951696ea4228bca322ca855ce583f0cc0

Observation 050d4b37-ae3a-4772-964a-9aaec3a18bbc · outbound

This paper cites Diffusion Model Alignment Using Direct Preference Optimization.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Diffusion Model Alignment Using Direct Preference Optimization

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.897454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:1a3ef491b03a1ec21d46f3bbfb668b4aee053f7575771a4bde193afc65220fb4

Observation a8fd6701-d236-4a4f-8fdc-070b8d29c6bc · outbound

This paper cites an unresolved cited work.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-07-09T02:25:56.019952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:5550bc7f56543b501568a36288d64a945baf67ab25e8843dd2c72c59466dba01

Observation c706126f-6539-4203-845b-a6a06bf6d52d · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.892309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:a5302a250b59f391d2098ee1285eb439603fa29d9c870b62261fadee3224bdc4

Observation f6805c19-0029-42c5-a7f3-160d2e405fc9 · outbound

This paper cites Data Retrieval with Importance Weights for Few-Shot Imitation Learning.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Data Retrieval with Importance Weights for Few-Shot Imitation Learning

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.868395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:c50b5e61574fee14658fa4b122e3ca68554b6bdfc648f22d9cef4e9f141f8b00

Observation 574c806b-1fe1-47ff-bf13-81f284638c4b · outbound

This paper cites Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.944838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:d9ea1c0b3a9006a15d67e9d1171b5c6a73167672b4e000bde478c82fc198243b

Observation b28ff67b-39ac-4803-bf5e-ec10c7a3ce4c · outbound

This paper cites Advances in Neural Information Processing Systems 36, 15903–15935 (2023) 7.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Advances in Neural Information Processing Systems 36, 15903–15935 (2023) 7

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T02:25:56.018386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:0dc05ef8bd61f7fa3679bc515e3f7cb7b6f6b19d5796f9a87095d0a619966f43

Observation eabf51b1-b9d6-42ce-81ef-a31b19df1a97 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF DanceGRPO: Unleashing GRPO on Visual Generation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.860040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:b57fa3a473bc4e7a2e6c49a4151087d1b6b8422ec92d08c876cb32d40e415564

Observation a27aef16-94fb-4308-9dee-1966183073fe · outbound

This paper cites Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous Driving and Zero-Shot Instruction Following.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous Driving and Zero-Shot Instruction Following

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.924899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:2f880d61371f02e10b18fd3e13f96032c928d973bb3dcacc8c10a43214c0df81

Observation 3bb1a0b2-44cd-4107-b72b-b8c9908e4275 · outbound

This paper cites Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.919756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:0019104f53417889a7588a26059d441f2af2e5294d19c014ca57e3acef699b8e

Observation b31d0fa8-79a1-408e-aa28-05728b57cc10 · outbound

This paper cites SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.851618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:d557bfb00b29c40be4267aa38f7f327ad80c91d565418a4648ff86f90b04926c

Observation 7cb99871-101a-4543-9a56-60f6cd2d3c6a · outbound

This paper cites Energy-Based Hindsight Experience Prioritization.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Energy-Based Hindsight Experience Prioritization

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.917249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:46cbdb4c720a6c9e6838cb4c39acd419233f4be0352fe0b549452e81127043c7

Observation 6054a268-b56d-4914-810b-d30d2848c609 · outbound

This paper cites arXiv preprint arXiv:2510.01982 (2025) 3.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF arXiv preprint arXiv:2510.01982 (2025) 3

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T02:25:55.873580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:d5077a5a88e4b6297f67e4cab5287b219446ade95aa864863e8a230f60fa533a

Pith citing papers

No inbound Pith citation observations are available.