Pith. sign in

Paper Citation Record · LEDGER

Policy-labeled Preference Learning: Is Preference Enough for RLHF?

As of 16 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2505.06273.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.06273 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:58:39.026249Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation df2b8aca-e28b-480d-8132-4f4bb2f184ab · outbound

This paper cites GPT-4 Technical Report.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.547058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.547058Z digest=sha256:3408ae442069fab5905250f4a3cfb2a389536196437e3cae96e107db319d5f99

Observation b019230c-ff8d-4849-aafb-1951c269198e · outbound

This paper cites Models of human preference for learning reward functions.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Models of human preference for learning reward functions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.566308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.566308Z digest=sha256:f8a85d859e701a67c5825379dfa1ab95b890aeb317954bb3785529c105423a2e

Observation 7c08a05f-86c2-4223-a1cd-c779f152a9da · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Offline Reinforcement Learning with Implicit Q-Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.571583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.571583Z digest=sha256:de3e0462d07f102061eab633ce45c22b1ecd39185a2efa0928b43dffd096d0f4

Observation 9e5458fd-c9ec-4ed2-9613-0ef9f29fe47b · outbound

This paper cites Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.580544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.580544Z digest=sha256:5749a16055f362a0025e03aa02f1a110cc39da06b26ddfa182a26325605cb833

Observation e416e5e9-a6b0-4958-b35c-d82664df16a0 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.584826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.584826Z digest=sha256:21fe0279af234c161418e58b959224e8c88b9512afb08aff00861a7d7d2163ba

Observation 1f08dc77-e740-409b-9573-806f67b5233f · outbound

This paper cites Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.653997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.653997Z digest=sha256:da7650dfb0342a70c2be0a29fdcbdcc41f24f0b36f1fedc79e159e0f62ef6ca1

Observation e4bbaf28-01cd-47f8-ac80-a5bb00d5b50c · outbound

This paper cites SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.755725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.755725Z digest=sha256:2cfc27b9ecb0c9b77651b40aded56bb020dc0b0619b83786e971810892f3726d

Observation 1d151406-6d1b-4081-9e4b-4b1e04ecd226 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Proximal Policy Optimization Algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.831743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.831743Z digest=sha256:23b65f2d3fd4cfaba74a36be2d29c89b738c34c33f29d34583a279fefeb09913

Observation fba3397f-af44-4305-abb2-056750b0d8d5 · outbound

This paper cites Aligning Language Models with Demonstrated Feedback.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Aligning Language Models with Demonstrated Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.836609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.836609Z digest=sha256:b924a99fd6c4ff84cf6cb593ab7b446289c84be19a9cbc3ef7c80fb53a487f8c

Observation 435de69b-165f-468f-839a-a7fb790939cf · outbound

This paper cites Forward KL Regularized Preference Optimization for Aligning Diffusion Policies.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Forward KL Regularized Preference Optimization for Aligning Diffusion Policies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.843175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.843175Z digest=sha256:fbd6ed718461a3f15ed99f07eb7404319cc8190108325455dc9410ac8f4b66b6

Observation ce3367d6-9549-4bf5-bd50-2407d0d19eb9 · outbound

This paper cites Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.850516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.850516Z digest=sha256:b7c4e89d815d3050ada36b2f682ed72df6c36c9b1b158b4604b1beacc0837913

Observation b2ecad65-726b-42cc-a8d2-a16489530897 · outbound

This paper cites Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.854695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.854695Z digest=sha256:b9a4efc494e4a2f31bf222c484f035634c248e34fda521105c593a0b57bb56aa

Observation d5af1267-c833-4a0b-b1f6-64ac3b030c33 · outbound

This paper cites Dichotomy of Control: Separating What You Can Control from What You Cannot.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Dichotomy of Control: Separating What You Can Control from What You Cannot

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.858135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.858135Z digest=sha256:17d6d882952cd61cbcbd8db717f033225b8d6e70df792d19578f9bf32cd14ff6

Observation c6c3cc89-936e-48cb-92b4-9bd3348890fb · outbound

This paper cites Token-level Direct Preference Optimization.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Token-level Direct Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.862316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.862316Z digest=sha256:9fa2985c5fbb97993abcaba20ab9cc4646676b3f1fc5195c31e1fdcb175ec580

Observation c2b93ec8-e704-4b73-ba47-f0276b1a35dc · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.865761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.865761Z digest=sha256:3ca641ee1adbf5442cd49baa3724ac5b6fa447f904dbf66ef3d5256675b5b7b1

Observation 900ad02b-2619-4c42-8ec3-433b09c10f34 · outbound

This paper cites an unresolved cited work.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:58:39.709877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:58:38.874698Z digest=sha256:353ae8a5bc3483980e494e16bcb7a24096e3929169b2754a1e794f923f1d9f10

Observation 75a12e4e-98c0-42b7-8eab-b6c62432f22b · outbound

This paper cites an unresolved cited work.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Unresolved cited work

Reference 26

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T23:58:39.557001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:58:38.922235Z digest=sha256:0936eaf819afcd562cb05004c26aee125d3c33fa8441e9b08b4d31cea476f5a7

Observation e42e0a09-fb6a-4394-a405-66e1f2dd15b0 · outbound

This paper cites Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.589423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.589423Z digest=sha256:eb611a03297c96f201eb18e184f6c43f9411677b7cd833695346a1f495aeaae4

Observation 8f2a234c-0723-4e8a-b15d-4a65a4e5fafe · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Fine-Tuning Language Models from Human Preferences

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.869659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.869659Z digest=sha256:93cdfd3c52e04b2b47c5d3c457583230926653d6a17a2a0d28902c51498de022

Observation aa9d8647-3e96-4c3c-bb48-f39ce045d14c · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.827512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.827512Z digest=sha256:d2f4e8d950efac3ac460f9d7e9f8026c19cf999c782c3809de5beaa138e6bf48

Observation e6c00768-6b85-4f49-a442-c968e6367cfe · outbound

This paper cites Quantifying Differences in Reward Functions.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Quantifying Differences in Reward Functions

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.557414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.557414Z digest=sha256:882b830af0b8643c4fdc0f67be07a8953efe9f13d02906f44fc70eed48229b88

Observation bca9145c-5e66-4938-878d-7c77704cfde4 · outbound

This paper cites While following this data generation procedure, we found a step in the reference code where transitions following a success signal were explicitly truncated.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? While following this data generation procedure, we found a step in the reference code where transitions following a success signal were explicitly truncated

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:58:39.524396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:58:39.026249Z digest=sha256:53920ca82885ce26390b504c9b39f960b8b85ed16def283aebd62654d24abff6

Observation dcf98345-91be-4a2b-9828-6b6570c368a3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.846724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.846724Z digest=sha256:7f4c73f87d105fafe42b0261beed6342e927f3dba8f8076b868274e69ba005e2

Observation c92fb850-7e76-4000-90a7-4297541278da · outbound

This paper cites PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.576437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.576437Z digest=sha256:ba909f654cc879b1f7116666762aeb90e459f938935d208a1e5bac1d715fd5b7

Observation f97cbcf5-d39e-43e7-95e7-5ccb7806eade · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.822297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.822297Z digest=sha256:0524f1f0e7e267c0671e13cdc190285736c25c3375ccdcc039c4996284203467

Observation 2e513878-51b6-4341-9784-b1f57a848530 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.552673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.552673Z digest=sha256:40857906890156e11fca5cf8652d97ba2bc5590c3de231d62f94c0f278a56877

Observation b47366cf-38d5-4c56-8816-d23fdc18eec6 · outbound

This paper cites Contrastive Preference Learning: Learning from Human Feedback without RL.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.561818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.561818Z digest=sha256:c6dfb295200075628db57b4838cedf05fe1e2e181c93890249400a9d0e82164f

Pith citing papers

No inbound Pith citation observations are available.