Pith. sign in

Paper Citation Record · LEDGER

Learning a Pessimistic Reward Model in RLHF

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 3 inbound Pith citation observations for arXiv:2505.20556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20556 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:07.462011Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T17:11:46.549150Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation de6fb29f-26b0-426e-92ab-19cb013c0f0c · outbound

This paper cites [2024b], with probability at least 1 − δ, the last term can be bound by V µ r∗ (π) − V µ r∗ (ˆπ) ≤ CµD (R, π, πref)2 8κ2 · β + 3β N log( Nϵ(R, ∥ · ∥∞) δ ).

Learning a Pessimistic Reward Model in RLHF [2024b], with probability at least 1 − δ, the last term can be bound by V µ r∗ (π) − V µ r∗ (ˆπ) ≤ CµD (R, π, πref)2 8κ2 · β + 3β N log( Nϵ(R, ∥ · ∥∞) δ )

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:09.146820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:01:07.287615Z digest=sha256:4b8cc81a8fdeb10999bf04f9f8b8dc964b8ca54fcb30e9f26284530db4aa9152

Observation 36e8ddd2-9c26-4b78-93aa-403379b37e37 · outbound

This paper cites The rejection sampling process is effectively a policy as it takes a prompt as input and stochastically outputs a response.

Learning a Pessimistic Reward Model in RLHF The rejection sampling process is effectively a policy as it takes a prompt as input and stochastically outputs a response

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:08.680829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:01:07.462011Z digest=sha256:47fc99ea9e2d10369a774ddf2d96dda352ee9cd6747f975b39f0d138b26f1a3d

Observation 85ccfc5b-4c63-4600-9e91-6f20604f8e6a · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Learning a Pessimistic Reward Model in RLHF Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.764200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.764200Z digest=sha256:ae8772aaf6958a69a6cb0329127b94bb1e070717d6e467a591dc6615a71a2a05

Observation f499fe4f-e977-411a-88ef-e517db30b68f · outbound

This paper cites Robust Preference Optimization through Reward Model Distillation.

Learning a Pessimistic Reward Model in RLHF Robust Preference Optimization through Reward Model Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.861147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.861147Z digest=sha256:4af91706054f8643762d4b96943c42f9954e6869fea214015666e9ca95628429

Observation 1a636902-5b51-4f76-86a8-8078c4290c71 · outbound

This paper cites Mitigating Preference Hacking in Policy Optimization with Pessimism.

Learning a Pessimistic Reward Model in RLHF Mitigating Preference Hacking in Policy Optimization with Pessimism

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.976843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.976843Z digest=sha256:c0330652d9e2654f2f14a6daf648beaafb6f777cc158f192553b18328be24606

Observation 6c726fd7-51c0-4dcb-858c-88becab87c76 · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Learning a Pessimistic Reward Model in RLHF Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.069235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.069235Z digest=sha256:c5dac967d76fb42fe89af2b4ec49627c3f935f9f08c0ba85694dfb62a9e77370

Observation 362722c1-ba46-4730-a3e0-bd411823339a · outbound

This paper cites The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization.

Learning a Pessimistic Reward Model in RLHF The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.147815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.147815Z digest=sha256:badd222243e5d24616244b2ac4dd2d26b89e655a8a3a80de77cdee9fed7f0a5a

Observation 846f9a8e-399f-4bad-bf6b-ea1336025286 · outbound

This paper cites RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models.

Learning a Pessimistic Reward Model in RLHF RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.235610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.235610Z digest=sha256:948c7073d2a7caeee7fe7d199164b99bcbde6d8abbd5380e535e9fae795b41e0

Observation cc13c926-2db6-4395-8878-9f430c3aedc4 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Learning a Pessimistic Reward Model in RLHF Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.329676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.329676Z digest=sha256:f6b0b62c81e2e1a138f345d18276d0093a4938afcf0bffa5ee245f9075f9cc43

Observation 10740686-15de-458a-8a08-82ef54fc1c73 · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Learning a Pessimistic Reward Model in RLHF Statistical Rejection Sampling Improves Preference Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.396920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.396920Z digest=sha256:1310bed9b1267f232abe9080a8b32bf83cbed355c1c5cb629fac9c40c148f1e8

Observation 3e72c6b5-9edd-4f59-b902-d1674575148e · outbound

This paper cites RRM: Robust Reward Model Training Mitigates Reward Hacking.

Learning a Pessimistic Reward Model in RLHF RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.490143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.490143Z digest=sha256:21bea7ebbeabc7a493ad0f93da376bcf0e786dfd1a4d8068582d6453babe9df5

Observation a2b21913-dd8a-401e-bf9e-3fe44b655998 · outbound

This paper cites WARM: On the Benefits of Weight Averaged Reward Models.

Learning a Pessimistic Reward Model in RLHF WARM: On the Benefits of Weight Averaged Reward Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.649860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.649860Z digest=sha256:a0c4b3b1c9476daff04d0d399ce35db451da4e1c86cfe19dd7fe01f0b1354677

Observation cf583308-af18-4836-b38e-fa24ff026969 · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Learning a Pessimistic Reward Model in RLHF Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.734257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.734257Z digest=sha256:6b8ab81e704b4ea47feeb238b10f725a9e9ec0108a6ab7f3065e0ba3dfffa98c

Observation dd3eca1a-614b-4892-a228-e5e9e4d1e819 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning a Pessimistic Reward Model in RLHF Proximal Policy Optimization Algorithms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.832817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.832817Z digest=sha256:a6806819b2d046e984d3d32b07503466be1d7f8bb9754bfcd4ae03163c30b4a6

Observation 752f47ed-e83d-4753-81d5-9c3d5d35c94f · outbound

This paper cites Boosting Reward Model with Preference-Conditional Multi-Aspect Synthetic Data Generation.

Learning a Pessimistic Reward Model in RLHF Boosting Reward Model with Preference-Conditional Multi-Aspect Synthetic Data Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.949991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.949991Z digest=sha256:4aff439689e6b6f8091f7eb2b8a9eba02ccbce72f3e95c157cba896e32410edc

Observation 400d8bdf-8572-4ceb-ba4e-f027564c672b · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

Learning a Pessimistic Reward Model in RLHF Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.110662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.110662Z digest=sha256:36fd94ed578dbbc7e302bc6cfa1cd79fd07eea50424c7ce8282f131448a26b21

Observation 2188078f-3ec4-47cd-8dcd-da45e6a778b8 · outbound

This paper cites Causal Confusion and Reward Misidentification in Preference-Based Reward Learning.

Learning a Pessimistic Reward Model in RLHF Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.259098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.259098Z digest=sha256:aa03424fa4f29bfa1c5aa9c922e21bee06cd9fca35e7068a504b920263a3cc19

Observation efd341ea-140e-4c02-b775-92ebc86fb584 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Learning a Pessimistic Reward Model in RLHF LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.403374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.403374Z digest=sha256:6d01746702e28fe8aad4dffc479a6cf189889942259e456bec895b68bec27841

Observation 7defe75f-a49b-407f-84a5-9bff0f301c07 · outbound

This paper cites Making RL with Preference-based Feedback Efficient via Randomization.

Learning a Pessimistic Reward Model in RLHF Making RL with Preference-based Feedback Efficient via Randomization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.516230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.516230Z digest=sha256:a75ad18b2ec40aa605de1f3f314cf1374218f08a324b267f939146326b3139d3

Observation f7e99903-dfd7-4e45-a6ac-6fd3782b22ad · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Learning a Pessimistic Reward Model in RLHF Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.669955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.669955Z digest=sha256:ee8efff1520a3fff543259f1f93cfe2c12dd20f0e42bd0fedbf9819ad9fc3391

Observation c925f773-402e-4005-aad0-6c67f4ba00c6 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Learning a Pessimistic Reward Model in RLHF Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.783481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.783481Z digest=sha256:b73ef0cd961ab043d1ce1943bc7c3107ec860602d348597137a6d2fd6c380f85

Observation a3074639-539a-49a7-97b8-496c5ef4f467 · outbound

This paper cites Provable Offline Preference-Based Reinforcement Learning.

Learning a Pessimistic Reward Model in RLHF Provable Offline Preference-Based Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.936809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.936809Z digest=sha256:da7a6e33abb627f9fb50d16c5dbfcbea30d73a098b538a87593a3eb907498430

Observation 9a3474c4-bb56-4ffb-89d4-bf52cc09386e · outbound

This paper cites Self-Exploring Language Models: Active Preference Elicitation for Online Alignment.

Learning a Pessimistic Reward Model in RLHF Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:07.063483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:07.063483Z digest=sha256:77731a695c7b7ed7a6c811268f4a4bd83e838737a8e504984982179162fe240c

Observation 43c8e79b-099b-42df-b0df-f7e4b8b3a09b · outbound

This paper cites Logarithmic regret for online kl-regularized reinforcement learning.

Learning a Pessimistic Reward Model in RLHF Logarithmic regret for online kl-regularized reinforcement learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:07.182899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:07.182899Z digest=sha256:0744071b144c9aad8b1350ccb3fa3454e42e2ee09a730e5c5cbf181a9515cb3b

Observation 192b52ba-3723-4214-ab5e-f95fd3bc5dcf · outbound

This paper cites 1 will increase the bound on the performance gap analysis, which is undesired for sampling efficiency.

Learning a Pessimistic Reward Model in RLHF 1 will increase the bound on the performance gap analysis, which is undesired for sampling efficiency

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:08.907969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:01:07.361030Z digest=sha256:e8f10d46682c210f40112637237048ff8b97f692d197f2e94cfad8b053397013

Observation df8a4b7b-f1ed-41fc-98b0-148f8d11c3fb · outbound

This paper cites Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF.

Learning a Pessimistic Reward Model in RLHF Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.563716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.563716Z digest=sha256:4cfbcdc03e921a70766ec11192b1059c5664877acbf583c76f3e9f6c2a1a3c4c

Observation 3d243c16-4172-4782-96f3-490df15f7716 · outbound

This paper cites Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization.

Learning a Pessimistic Reward Model in RLHF Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.624213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.624213Z digest=sha256:6a424b9556884b65be3fcb809633118c9fa4eb8ad8bbfbe1617326d63ce35d9c

Observation 301e7a58-e6e9-434e-9935-0ccf2266b1bf · outbound

This paper cites Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling.

Learning a Pessimistic Reward Model in RLHF Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:01:08.191682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:01:05.568274Z digest=sha256:3f8bffbbfc02f480b6ea3c1969cedfa510cefd8eb2ef48c4ade8fd30c1c5231e

Observation 44fed6cc-8e4a-4b34-b7fe-75eaa37a3fbd · outbound

This paper cites Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.

Learning a Pessimistic Reward Model in RLHF Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.472540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.472540Z digest=sha256:43d6e0cefde1a8bb9be2e910e543b44e5010e77dfe491ee077ed4011e2b601ef

Observation 8a6d0d9c-1030-42b3-add9-adf75ffc89b3 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Learning a Pessimistic Reward Model in RLHF Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.407989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.407989Z digest=sha256:9718f71990d3bc41b80468105936a84d2cffc91a2602e05edd9dbaf0f3de272e

Observation 81bbaf25-9f91-4f8c-a513-b6862e9d21ba · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Learning a Pessimistic Reward Model in RLHF RLHF Workflow: From Reward Modeling to Online RLHF

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.698570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.698570Z digest=sha256:38f54c0431dfd69063fba2fd21090977863634ec083ccb1f15329d28854075c2

Pith citing papers

Observation b4b2d56e-a2c4-49c5-af8c-331bc75722da · inbound

Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback cites this paper.

Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback Learning a Pessimistic Reward Model in RLHF

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:46:47.349064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T21:03:53.045304Z digest=sha256:a6d4cc7013d0013024923f2b749bd373417658d7c4a2674cf8f7134413176c5e

Observation 3eb09b5e-b6d2-4617-b84c-8e79bd1ac3ba · inbound

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs cites this paper.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Learning a Pessimistic Reward Model in RLHF

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:41:21.460204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-22T09:38:04.387777Z digest=sha256:e57cf4ffcf3b54d668178ca6e0a5ea74137774649606badaf1411edb297bfa60

Observation bfa36826-0ee0-4a04-925f-e4fd9f11aaef · inbound

A Unifying Lens on Reward Uncertainty in RLHF cites this paper.

A Unifying Lens on Reward Uncertainty in RLHF Learning a Pessimistic Reward Model in RLHF

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.825574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:11:46.549150Z digest=sha256:8b69196621bbc7e6ade69ec62d6148cc11a666b1e2269f2731f5670f0a3a006a