Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:07.462011Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 3 inbound Pith citation observations for arXiv:2505.20556.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:07.462011Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T17:11:46.549150Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
31 of 31 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation de6fb29f-26b0-426e-92ab-19cb013c0f0c · outbound
Learning a Pessimistic Reward Model in RLHF [2024b], with probability at least 1 − δ, the last term can be bound by V µ r∗ (π) − V µ r∗ (ˆπ) ≤ CµD (R, π, πref)2 8κ2 · β + 3β N log( Nϵ(R, ∥ · ∥∞) δ )
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36e8ddd2-9c26-4b78-93aa-403379b37e37 · outbound
Learning a Pessimistic Reward Model in RLHF The rejection sampling process is effectively a policy as it takes a prompt as input and stochastically outputs a response
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85ccfc5b-4c63-4600-9e91-6f20604f8e6a · outbound
Learning a Pessimistic Reward Model in RLHF Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f499fe4f-e977-411a-88ef-e517db30b68f · outbound
Learning a Pessimistic Reward Model in RLHF Robust Preference Optimization through Reward Model Distillation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a636902-5b51-4f76-86a8-8078c4290c71 · outbound
Learning a Pessimistic Reward Model in RLHF Mitigating Preference Hacking in Policy Optimization with Pessimism
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c726fd7-51c0-4dcb-858c-88becab87c76 · outbound
Learning a Pessimistic Reward Model in RLHF Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 362722c1-ba46-4730-a3e0-bd411823339a · outbound
Learning a Pessimistic Reward Model in RLHF The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 846f9a8e-399f-4bad-bf6b-ea1336025286 · outbound
Learning a Pessimistic Reward Model in RLHF RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc13c926-2db6-4395-8878-9f430c3aedc4 · outbound
Learning a Pessimistic Reward Model in RLHF Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10740686-15de-458a-8a08-82ef54fc1c73 · outbound
Learning a Pessimistic Reward Model in RLHF Statistical Rejection Sampling Improves Preference Optimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e72c6b5-9edd-4f59-b902-d1674575148e · outbound
Learning a Pessimistic Reward Model in RLHF RRM: Robust Reward Model Training Mitigates Reward Hacking
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2b21913-dd8a-401e-bf9e-3fe44b655998 · outbound
Learning a Pessimistic Reward Model in RLHF WARM: On the Benefits of Weight Averaged Reward Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf583308-af18-4836-b38e-fa24ff026969 · outbound
Learning a Pessimistic Reward Model in RLHF Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd3eca1a-614b-4892-a228-e5e9e4d1e819 · outbound
Learning a Pessimistic Reward Model in RLHF Proximal Policy Optimization Algorithms
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 752f47ed-e83d-4753-81d5-9c3d5d35c94f · outbound
Learning a Pessimistic Reward Model in RLHF Boosting Reward Model with Preference-Conditional Multi-Aspect Synthetic Data Generation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 400d8bdf-8572-4ceb-ba4e-f027564c672b · outbound
Learning a Pessimistic Reward Model in RLHF Generalized Preference Optimization: A Unified Approach to Offline Alignment
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2188078f-3ec4-47cd-8dcd-da45e6a778b8 · outbound
Learning a Pessimistic Reward Model in RLHF Causal Confusion and Reward Misidentification in Preference-Based Reward Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efd341ea-140e-4c02-b775-92ebc86fb584 · outbound
Learning a Pessimistic Reward Model in RLHF LLaMA: Open and Efficient Foundation Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7defe75f-a49b-407f-84a5-9bff0f301c07 · outbound
Learning a Pessimistic Reward Model in RLHF Making RL with Preference-based Feedback Efficient via Randomization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e99903-dfd7-4e45-a6ac-6fd3782b22ad · outbound
Learning a Pessimistic Reward Model in RLHF Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c925f773-402e-4005-aad0-6c67f4ba00c6 · outbound
Learning a Pessimistic Reward Model in RLHF Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3074639-539a-49a7-97b8-496c5ef4f467 · outbound
Learning a Pessimistic Reward Model in RLHF Provable Offline Preference-Based Reinforcement Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a3474c4-bb56-4ffb-89d4-bf52cc09386e · outbound
Learning a Pessimistic Reward Model in RLHF Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43c8e79b-099b-42df-b0df-f7e4b8b3a09b · outbound
Learning a Pessimistic Reward Model in RLHF Logarithmic regret for online kl-regularized reinforcement learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 192b52ba-3723-4214-ab5e-f95fd3bc5dcf · outbound
Learning a Pessimistic Reward Model in RLHF 1 will increase the bound on the performance gap analysis, which is undesired for sampling efficiency
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df8a4b7b-f1ed-41fc-98b0-148f8d11c3fb · outbound
Learning a Pessimistic Reward Model in RLHF Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Reference 1952
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d243c16-4172-4782-96f3-490df15f7716 · outbound
Learning a Pessimistic Reward Model in RLHF Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 301e7a58-e6e9-434e-9935-0ccf2266b1bf · outbound
Learning a Pessimistic Reward Model in RLHF Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44fed6cc-8e4a-4b34-b7fe-75eaa37a3fbd · outbound
Learning a Pessimistic Reward Model in RLHF Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a6d0d9c-1030-42b3-add9-adf75ffc89b3 · outbound
Learning a Pessimistic Reward Model in RLHF Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81bbaf25-9f91-4f8c-a513-b6862e9d21ba · outbound
Learning a Pessimistic Reward Model in RLHF RLHF Workflow: From Reward Modeling to Online RLHF
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4b2d56e-a2c4-49c5-af8c-331bc75722da · inbound
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback Learning a Pessimistic Reward Model in RLHF
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3eb09b5e-b6d2-4617-b84c-8e79bd1ac3ba · inbound
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Learning a Pessimistic Reward Model in RLHF
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bfa36826-0ee0-4a04-925f-e4fd9f11aaef · inbound
A Unifying Lens on Reward Uncertainty in RLHF Learning a Pessimistic Reward Model in RLHF
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.