Pith. sign in

Paper Citation Record · LEDGER

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF

As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2607.07693.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07693 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T02:17:20.589485Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact35
  • verified fuzzy4
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch12

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1fe76140-f202-44d1-b07d-16bf1467641a · outbound

This paper cites Hindsight Experience Replay.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Hindsight Experience Replay

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.909974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:7c421a3831d18072f952764ee851c356ddd5b9cb68a3a8460e52c7c3b562f80e

Observation e304a2af-3e64-4c16-9c44-bb570395a8ca · outbound

This paper cites 133011, 3, 7, 8, 12.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF 133011, 3, 7, 8, 12

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T02:25:56.026565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:bc07685ff414acffe425d5c7924a26b5443f34aed17e3236ac290de843c402ad

Observation ad5a402c-9aaa-462e-bb2b-1a204b96b0e6 · outbound

This paper cites DSPO: Direct Semantic Preference Optimization for Real-World Image Super-Resolution.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF DSPO: Direct Semantic Preference Optimization for Real-World Image Super-Resolution

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.845525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:1e1440c25610415c38d0de2c9611d2078d05f3fe83486b33377943db6695464f

Observation faf7892d-89e3-4a22-8b06-16ddceeab3c7 · outbound

This paper cites MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.876895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:fd4760243a6931861a1751cb87754fa903d3a9496da3f5390470bc2ac47574d6

Observation 2fc94259-9502-4c79-8de2-e8f5b0dec446 · outbound

This paper cites In: Proceedings of Robotics: Science and Systems (RSS) (2023) 2.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF In: Proceedings of Robotics: Science and Systems (RSS) (2023) 2

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T02:25:56.016587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:b7e957c7d7158394d603ffd40fe81933ab0c58ea1070aa1f238aa539ca5010b4

Observation 20f0dcde-2523-43e2-83b4-599e538aedfd · outbound

This paper cites Deep reinforcement learning from human preferences.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Deep reinforcement learning from human preferences

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.932859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:081cbe7d78828385aed0f919a18acd6b4df81a1d1dfe79e4feb7aca84ce6feca

Observation 1c41d852-001c-4642-8b42-97bc12a136e5 · outbound

This paper cites Directly Fine-Tuning Diffusion Models on Differentiable Rewards.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Directly Fine-Tuning Diffusion Models on Differentiable Rewards

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.930343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:2b3244e98324bb2f279318e0f0f29c2dfbd384c4dd300615e512f9d50aabb0d5

Observation ec6b81b7-a673-4974-8039-af0e933faa18 · outbound

This paper cites DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.947272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:f45609072fddbba0b9cf2651fd6179ab95aebe97c4367d0d978f4d65f74156c4

Observation 325b6d17-b4e2-4aa9-89b5-3c411dbdab38 · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.912565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:dce8d61f4f08ccfebd7321446b7db139db588fd9738c7bb2457261ec9937dc41

Observation 4bd34cc9-7b99-4330-b53a-b71af152eced · outbound

This paper cites TempFlow-GRPO: When Timing Matters for GRPO in Flow Models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF TempFlow-GRPO: When Timing Matters for GRPO in Flow Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.870955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:41fcea80af6ce15e67ec374e4c7d034967c3249d823eda6967d71551c2d0bc53

Observation d3bf136e-8151-4f30-b16e-adef5d24e515 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.831004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:3bcc57c20a5c66b42c1d078a7cf80da1f88d42d6fe2a1a64c5d7ec05d96f569f

Observation c31117ed-4527-44e5-95a5-b99a77ec8b08 · outbound

This paper cites Denoising Diffusion Probabilistic Models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Denoising Diffusion Probabilistic Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.937670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:30b21f36de763735a749e9076a0ebd007799a402c2ce9de7933855a6aee9be76

Observation f375009b-90de-4073-8196-d8853fcb7f82 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.865485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:9e270f4aab7b3f0d6fa324e6ac2c384a8b4b2d76d2399790c1558c5ab2ddb7c4

Observation 9fd710a6-780b-4245-bb15-a54709896175 · outbound

This paper cites Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.902473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:c28283070aa02dec8fd0455f3fe39bc1e13ce420fa660dabbac3c8c773e2f7ec

Observation 4c63c6ed-0eac-4cfe-9f67-2f0f138a3667 · outbound

This paper cites PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.839812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:f910518468476c771688f8d7105abd321432c49d905ccf5da9782001e62f6631

Observation 1ce6d9bd-f0e1-424a-8712-342fe507c107 · outbound

This paper cites TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.887521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:09016af87fec4f1b822b9e2e07e251c3445c54fa15319e471e4a296ed93714ab

Observation f0d7ae0d-6bd3-4e43-b65f-4847ff0c983d · outbound

This paper cites Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.862799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:6283c0faf2e10149c3b11319694205d5cb334810a0830d496ef70fa6d0e7c225

Observation c1b1dd6e-26fd-4bf5-aedf-6fd2ee8daacf · outbound

This paper cites Branchgrpo: Stable and efficient grpo with structured branching in diffusion models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Branchgrpo: Stable and efficient grpo with structured branching in diffusion models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T02:25:55.895022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:870352985384dcb42b1ac09dd5d8e96bd4952883e04b12ab234aaee39639f344

Observation 95e0288b-a9ed-4d78-9dfc-95742979b630 · outbound

This paper cites an unresolved cited work.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-07-09T02:25:56.028280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:d87d4f7369f2273857880e539bb1ed01828b212b57ed613176ffb9aaccac3f46

Observation 527924e1-c108-4dd9-a2b5-9f871f963f3b · outbound

This paper cites Continuous control with deep reinforcement learning.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Continuous control with deep reinforcement learning

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.935385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:6679a8acde6ba559ef3436c3ecb5bb202bdf1d6afdd4775fc16cbf69906329b1

Observation 372b9069-a4e0-49ee-8396-2bfb6cae3fe0 · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Flow-GRPO: Training Flow Matching Models via Online RL

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.899907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:ad34202f2f2d183b06bb9381b0cc3d2c7f3b55c9796209aff11f4e800b1380cc

Observation 35f1aebb-ca82-444f-8f49-795842f1c964 · outbound

This paper cites Synthetic Experience Replay.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Synthetic Experience Replay

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.942485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:9556af7586bb77f74e75cd234a583c074350aa1958af87040d92e947ac38d15f

Observation 82a6cc54-75ae-400f-aa8f-12b51ae1a4d5 · outbound

This paper cites Dual-Process Image Generation.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Dual-Process Image Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.904789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:7df904106b05e744215d34281d97a9ebf60907f1b80f7e6149d4c929db46df70

Observation db1ce687-aace-4e05-8e87-077bb8278092 · outbound

This paper cites an unresolved cited work.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:25:55.842469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:d01846210de91c5eed9f1b480e3f4abd72e8ce69d5a9a824d4428daffd6b06d0

Observation c9c3d08e-3131-45c3-820c-7cceeb301579 · outbound

This paper cites Do text-free diffusion models learn discriminative visual representations?.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Do text-free diffusion models learn discriminative visual representations?

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.914949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:876f538d7cc3f19f6d0310b540e23e0c09e43f52194cb0f981ccd136f5a9c653

Observation c79984dd-eab4-4ca8-8b7e-107df3a0cbd1 · outbound

This paper cites GPT-4 Technical Report.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF GPT-4 Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.857250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:be6992dcaea4d7a0051336fd0b6b67c84edb432f3b9f9d5beebf222c41677e78

Observation 7c6223df-d4d5-4464-85ea-da6e1faa76cd · outbound

This paper cites Training language models to follow instructions with human feedback.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Training language models to follow instructions with human feedback

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.940191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:418f308f4f2a8c576e4baf043081fe428f7231eec2106e2930191d2f0b912705

Observation f0b870ad-7af4-42ae-b76e-37848e72aaf4 · outbound

This paper cites Reinforcement learning by reward-weighted regression for operational space control.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Reinforcement learning by reward-weighted regression for operational space control

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T02:25:55.802879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:ecf55dd3b90a0acbd9e65364c708df831a650da2722a88a16a65aacadfb221f2

Observation 6a2bf9f6-84dc-4e57-9ede-0145bc00a5dc · outbound

This paper cites Hard examples are all you need: Maximizing grpo post-training under annotation budgets.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Hard examples are all you need: Maximizing grpo post-training under annotation budgets

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T02:25:55.907595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:5eff59197759ce9a39b972455a0a63b5141cd18d3a7d9747c03849e111f3eaef

Observation 2f03f31d-7a68-4176-a776-c54ed1bf771c · outbound

This paper cites Video Diffusion Alignment via Reward Gradients.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Video Diffusion Alignment via Reward Gradients

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.922322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:473a219ab797fc6a8b5f9733e5a97149764f8746a75c2988066b558027c12d93

Observation eac7afe5-338e-4119-b7d7-9c0af4a5c505 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.882647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:96d945f540b129f16270e7a2b8e7d1eb398d6477d1c79f1b30c5dfc2bc2d1a78

Observation dc409362-ad24-4ba8-be55-e43b05569ebd · outbound

This paper cites an unresolved cited work.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-07-09T02:25:56.024965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:a0dd390e616ee7d4e060f726c979fe9f0dd8f5b6ae4d4e4797879ce503c5eded

Observation bc044944-4a3d-43b5-a8ea-32f03a7eb92a · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR) (2022) 7.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR) (2022) 7

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T02:25:56.023336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:8aaaca22b30518dc71c468dfe1ace117f033ec77c42046f8fba0714bb6afa36b

Observation 07e2a8de-2d17-4a3e-a807-cd8bff3b5774 · outbound

This paper cites Prioritized Experience Replay.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Prioritized Experience Replay

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.880064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:0a33922314dd594ad956a17b02b444628ebd96579c3789020931057c0f568153

Observation 9147be8d-01af-4b19-ae1d-93fede2baa8f · outbound

This paper cites an unresolved cited work.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-07-09T02:25:56.021586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:2a348a04639134d9f9b08d31f43013f1ea30fcfd21b617a9ee61856b77bc8df0

Observation 2b0dea72-f54f-49a0-9830-0a0fdc3f8114 · outbound

This paper cites Trust Region Policy Optimization.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Trust Region Policy Optimization

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.927244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:018b15e0e62b57fdff92abc93dfc8f9f1a71632d788a4de39d03251246641739

Observation d64aa1d4-622c-4a35-a9b3-e0c503653fdc · outbound

This paper cites Proximal Policy Optimization Algorithms.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Proximal Policy Optimization Algorithms

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.836731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:42d697786ebf74f3a2a01bf50630b03c716e7d891b8038893e67187973283478

Observation 48fe61e6-bdc9-4fd8-892c-99ccbf491efe · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.854539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:cc9f9cdaa3551f1c04abc35ed314b95394ebf131eedb8a08daea95724fd7968f

Observation 32a72b49-ff3e-48c1-b5b7-a46db07da39d · outbound

This paper cites RL's Razor: Why Online Reinforcement Learning Forgets Less.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF RL's Razor: Why Online Reinforcement Learning Forgets Less

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.889858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:312dbf9f0e6aab460844daa1b18f12366a70c94ad8674060a0a5c3c7759adcc8

Observation a55cffbf-3e07-48f3-be56-bef4b0ec943e · outbound

This paper cites Training Region-based Object Detectors with Online Hard Example Mining.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Training Region-based Object Detectors with Online Hard Example Mining

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.848620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:4733bcceb15d626bd357fa73410adc037abfe60d2e242a3dfc67f57f08cb2727

Observation eaece74c-41df-4db0-a57e-c698a7874e27 · outbound

This paper cites Deep Unsupervised Learning using Nonequilibrium Thermodynamics.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Deep Unsupervised Learning using Nonequilibrium Thermodynamics

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.834022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:aedb9af167ec54ce278e706b1308982acab1c20fe134e0bd8cd48f499ec3ccd4

Observation d7dd12aa-10d5-4484-a49e-bcbdb5baa0e4 · outbound

This paper cites Denoising Diffusion Implicit Models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Denoising Diffusion Implicit Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.885077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:a8884cb5e3320864f4cc96892c6058d2ea2de0d163daba9c9530a1c71692f873

Observation dd3d5b43-cf21-4969-a33b-b6cf14d54ed2 · outbound

This paper cites MIT press, ??? (2018).

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF MIT press, ??? (2018)

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T02:25:55.796181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:676a32d3991767bd90c5a6e9770a0860a96fb7271cb1356f2862e90d84678c57

Observation 050d4b37-ae3a-4772-964a-9aaec3a18bbc · outbound

This paper cites Diffusion Model Alignment Using Direct Preference Optimization.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Diffusion Model Alignment Using Direct Preference Optimization

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.897454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:fd449aaa5490903a379889601b5c0fa92b5e7d59d2d10f591b7f6b3affe3e983

Observation a8fd6701-d236-4a4f-8fdc-070b8d29c6bc · outbound

This paper cites an unresolved cited work.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-07-09T02:25:56.019952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:19ab99070686dc3008a57c9d95f0a39cf19d3ba0730baacfbe5b9a0f8946d312

Observation c706126f-6539-4203-845b-a6a06bf6d52d · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T02:25:55.892309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:2570401fbdb1f275163a6705eeae5dc0cad8e2dd384d949aed0fbab6b5d0c63d

Observation f6805c19-0029-42c5-a7f3-160d2e405fc9 · outbound

This paper cites Data Retrieval with Importance Weights for Few-Shot Imitation Learning.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Data Retrieval with Importance Weights for Few-Shot Imitation Learning

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.868395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:cda84f42065be79845a5665b633da1b61127236b613f14007ddcb441955b5593

Observation 574c806b-1fe1-47ff-bf13-81f284638c4b · outbound

This paper cites Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.944838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:233fd67c0088392a027781480a1f2e413303c2d33ad8c3ef578bfd79727f392b

Observation b28ff67b-39ac-4803-bf5e-ec10c7a3ce4c · outbound

This paper cites Advances in Neural Information Processing Systems 36, 15903–15935 (2023) 7.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Advances in Neural Information Processing Systems 36, 15903–15935 (2023) 7

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T02:25:56.018386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:e4e0ace3b0f4723b2e4fe92b033890a9526c5906dc8e485c2a0e7589f524a3e5

Observation eabf51b1-b9d6-42ce-81ef-a31b19df1a97 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF DanceGRPO: Unleashing GRPO on Visual Generation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.860040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:49bd4cf7cfeabd13cbbf24563f97622030c5953dfc7101e5d4414d1200fae5fe

Observation a27aef16-94fb-4308-9dee-1966183073fe · outbound

This paper cites Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous Driving and Zero-Shot Instruction Following.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous Driving and Zero-Shot Instruction Following

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.924899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:2452dbe06ad6bd80bdc12a0ed7f3430aebd8ac2c477dd7e828c6c3298209d14c

Observation 3bb1a0b2-44cd-4107-b72b-b8c9908e4275 · outbound

This paper cites Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.919756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:10070fe75e12fb557c04f84f60b0fbf45613f91b394fa5524eb4938bcce7ff40

Observation b31d0fa8-79a1-408e-aa28-05728b57cc10 · outbound

This paper cites SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.851618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:1fbdf808237b519774b6ac7f151d9810624309a4ab63e54f323bcbbe04bd835f

Observation 7cb99871-101a-4543-9a56-60f6cd2d3c6a · outbound

This paper cites Energy-Based Hindsight Experience Prioritization.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Energy-Based Hindsight Experience Prioritization

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:25:55.917249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:0f3b42bf470eb9e7cf8cdb35133a562173073f0f87d383a8581b2e455d82832e

Observation 6054a268-b56d-4914-810b-d30d2848c609 · outbound

This paper cites arXiv preprint arXiv:2510.01982 (2025) 3.

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF arXiv preprint arXiv:2510.01982 (2025) 3

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T02:25:55.873580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T02:17:20.589485Z digest=sha256:fddd54bdd70daa875944acd3341505101d6dae1fc0ccd6a352237dec758758a2

Pith citing papers

No inbound Pith citation observations are available.