Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:58:39.026249Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2505.06273.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:58:39.026249Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation df2b8aca-e28b-480d-8132-4f4bb2f184ab · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b019230c-ff8d-4849-aafb-1951c269198e · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Models of human preference for learning reward functions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c08a05f-86c2-4223-a1cd-c779f152a9da · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Offline Reinforcement Learning with Implicit Q-Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e5458fd-c9ec-4ed2-9613-0ef9f29fe47b · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e416e5e9-a6b0-4958-b35c-d82664df16a0 · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f08dc77-e740-409b-9573-806f67b5233f · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4bbaf28-01cd-47f8-ac80-a5bb00d5b50c · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d151406-6d1b-4081-9e4b-4b1e04ecd226 · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Proximal Policy Optimization Algorithms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fba3397f-af44-4305-abb2-056750b0d8d5 · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Aligning Language Models with Demonstrated Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 435de69b-165f-468f-839a-a7fb790939cf · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Forward KL Regularized Preference Optimization for Aligning Diffusion Policies
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce3367d6-9549-4bf5-bd50-2407d0d19eb9 · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2ecad65-726b-42cc-a8d2-a16489530897 · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5af1267-c833-4a0b-b1f6-64ac3b030c33 · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Dichotomy of Control: Separating What You Can Control from What You Cannot
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6c3cc89-936e-48cb-92b4-9bd3348890fb · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Token-level Direct Preference Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2b93ec8-e704-4b73-ba47-f0276b1a35dc · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 900ad02b-2619-4c42-8ec3-433b09c10f34 · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 75a12e4e-98c0-42b7-8eab-b6c62432f22b · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e42e0a09-fb6a-4394-a405-66e1f2dd15b0 · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
Reference 1999
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f2a234c-0723-4e8a-b15d-4a65a4e5fafe · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Fine-Tuning Language Models from Human Preferences
Reference 2010
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa9d8647-3e96-4c3c-bb48-f39ce045d14c · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6c00768-6b85-4f49-a442-c968e6367cfe · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Quantifying Differences in Reward Functions
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bca9145c-5e66-4938-878d-7c77704cfde4 · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? While following this data generation procedure, we found a step in the reference code where transitions following a success signal were explicitly truncated
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dcf98345-91be-4a2b-9828-6b6570c368a3 · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c92fb850-7e76-4000-90a7-4297541278da · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f97cbcf5-d39e-43e7-95e7-5ccb7806eade · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e513878-51b6-4341-9784-b1f57a848530 · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b47366cf-38d5-4c56-8816-d23fdc18eec6 · outbound
Policy-labeled Preference Learning: Is Preference Enough for RLHF? Contrastive Preference Learning: Learning from Human Feedback without RL
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.