Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-09T02:17:20.589485Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2607.07693.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-09T02:17:20.589485Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1fe76140-f202-44d1-b07d-16bf1467641a · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Hindsight Experience Replay
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e304a2af-3e64-4c16-9c44-bb570395a8ca · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF 133011, 3, 7, 8, 12
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad5a402c-9aaa-462e-bb2b-1a204b96b0e6 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF DSPO: Direct Semantic Preference Optimization for Real-World Image Super-Resolution
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation faf7892d-89e3-4a22-8b06-16ddceeab3c7 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2fc94259-9502-4c79-8de2-e8f5b0dec446 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF In: Proceedings of Robotics: Science and Systems (RSS) (2023) 2
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20f0dcde-2523-43e2-83b4-599e538aedfd · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Deep reinforcement learning from human preferences
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c41d852-001c-4642-8b42-97bc12a136e5 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Directly Fine-Tuning Diffusion Models on Differentiable Rewards
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec6b81b7-a673-4974-8039-af0e933faa18 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 325b6d17-b4e2-4aa9-89b5-3c411dbdab38 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bd34cc9-7b99-4330-b53a-b71af152eced · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3bf136e-8151-4f30-b16e-adef5d24e515 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c31117ed-4527-44e5-95a5-b99a77ec8b08 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Denoising Diffusion Probabilistic Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f375009b-90de-4073-8196-d8853fcb7f82 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF LoRA: Low-Rank Adaptation of Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9fd710a6-780b-4245-bb15-a54709896175 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c63c6ed-0eac-4cfe-9f67-2f0f138a3667 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ce6d9bd-f0e1-424a-8712-342fe507c107 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0d7ae0d-6bd3-4e43-b65f-4847ff0c983d · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c1b1dd6e-26fd-4bf5-aedf-6fd2ee8daacf · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Branchgrpo: Stable and efficient grpo with structured branching in diffusion models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95e0288b-a9ed-4d78-9dfc-95742979b630 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 527924e1-c108-4dd9-a2b5-9f871f963f3b · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Continuous control with deep reinforcement learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 372b9069-a4e0-49ee-8396-2bfb6cae3fe0 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Flow-GRPO: Training Flow Matching Models via Online RL
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 35f1aebb-ca82-444f-8f49-795842f1c964 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Synthetic Experience Replay
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82a6cc54-75ae-400f-aa8f-12b51ae1a4d5 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Dual-Process Image Generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation db1ce687-aace-4e05-8e87-077bb8278092 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9c3d08e-3131-45c3-820c-7cceeb301579 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Do text-free diffusion models learn discriminative visual representations?
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c79984dd-eab4-4ca8-8b7e-107df3a0cbd1 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF GPT-4 Technical Report
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c6223df-d4d5-4464-85ea-da6e1faa76cd · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Training language models to follow instructions with human feedback
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0b870ad-7af4-42ae-b76e-37848e72aaf4 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Reinforcement learning by reward-weighted regression for operational space control
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6a2bf9f6-84dc-4e57-9ede-0145bc00a5dc · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Hard examples are all you need: Maximizing grpo post-training under annotation budgets
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f03f31d-7a68-4176-a776-c54ed1bf771c · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Video Diffusion Alignment via Reward Gradients
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eac7afe5-338e-4119-b7d7-9c0af4a5c505 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc409362-ad24-4ba8-be55-e43b05569ebd · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc044944-4a3d-43b5-a8ea-32f03a7eb92a · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR) (2022) 7
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07e2a8de-2d17-4a3e-a807-cd8bff3b5774 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Prioritized Experience Replay
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9147be8d-01af-4b19-ae1d-93fede2baa8f · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b0dea72-f54f-49a0-9830-0a0fdc3f8114 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Trust Region Policy Optimization
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d64aa1d4-622c-4a35-a9b3-e0c503653fdc · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Proximal Policy Optimization Algorithms
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48fe61e6-bdc9-4fd8-892c-99ccbf491efe · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32a72b49-ff3e-48c1-b5b7-a46db07da39d · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF RL's Razor: Why Online Reinforcement Learning Forgets Less
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a55cffbf-3e07-48f3-be56-bef4b0ec943e · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Training Region-based Object Detectors with Online Hard Example Mining
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eaece74c-41df-4db0-a57e-c698a7874e27 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Deep Unsupervised Learning using Nonequilibrium Thermodynamics
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d7dd12aa-10d5-4484-a49e-bcbdb5baa0e4 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Denoising Diffusion Implicit Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd3d5b43-cf21-4969-a33b-b6cf14d54ed2 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF MIT press, ??? (2018)
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 050d4b37-ae3a-4772-964a-9aaec3a18bbc · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Diffusion Model Alignment Using Direct Preference Optimization
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a8fd6701-d236-4a4f-8fdc-070b8d29c6bc · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c706126f-6539-4203-845b-a6a06bf6d52d · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6805c19-0029-42c5-a7f3-160d2e405fc9 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Data Retrieval with Importance Weights for Few-Shot Imitation Learning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 574c806b-1fe1-47ff-bf13-81f284638c4b · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b28ff67b-39ac-4803-bf5e-ec10c7a3ce4c · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Advances in Neural Information Processing Systems 36, 15903–15935 (2023) 7
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eabf51b1-b9d6-42ce-81ef-a31b19df1a97 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF DanceGRPO: Unleashing GRPO on Visual Generation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a27aef16-94fb-4308-9dee-1966183073fe · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous Driving and Zero-Shot Instruction Following
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3bb1a0b2-44cd-4107-b72b-b8c9908e4275 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b31d0fa8-79a1-408e-aa28-05728b57cc10 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7cb99871-101a-4543-9a56-60f6cd2d3c6a · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Energy-Based Hindsight Experience Prioritization
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6054a268-b56d-4914-810b-d30d2848c609 · outbound
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF arXiv preprint arXiv:2510.01982 (2025) 3
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.