Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:41:24.133597Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2412.16475.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:41:24.133597Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:46.439486Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T15:15:46.504789Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0697c033-2049-4b76-a444-cb22ee7d507a · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0407d3c7-024b-4ad7-905a-9e86d08624af · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06aa6455-3a21-4490-8e1a-8d7a295660fc · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Accuracy of chatgpt, google bard, and microsoft bing for simplifying radiology reports
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2f75013d-35b0-4a43-9e98-8f815536e98b · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? The evolved radio and its implications for modelling the evolution of novel sensors
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 34dbd208-86fa-4304-ad0c-2627754f1232 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11dbaa5e-9aec-4357-a9b0-a12929b8320d · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Open problems and fundamental limitations of reinforcement learning from human feedback
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 931cb768-d394-44b8-b278-cd9433a34a84 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 908181f5-b5b9-488b-8ebf-c5d191ee899c · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? ODIN : Disentangled reward mitigates hacking in RLHF
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2b7d8406-ef65-4f0e-844e-f51d81e2b508 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Meta discovery: Learning to discover novel classes given very limited data
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ff703073-f67b-47a9-99cb-6265db905465 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Faulty reward functions in the wild, 2016
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6c9fbf7e-f935-4922-90cf-f1ef8190148f · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Reward model ensembles help mitigate overoptimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7c5f317f-021c-437f-83f9-5b2a52b7125f · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? The Expertise Problem: Learning from Specialized Feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d231620d-24d0-4340-9d71-89640e31f792 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Group symmetry in pac learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 38a0f1d1-5cba-438d-8bfe-99a2c9b01b88 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Active teacher selection for reward learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b7e3efa-6565-4eed-8135-f71a844c3077 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Scaling laws for reward model overoptimization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b96e6e82-1da8-4353-88e4-cf4087f4ba94 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Glass floor colleges reject top applicants, accepting only the students likely to enroll, 2001
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d459e898-3910-4869-9dc8-9090ce61848d · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc13a5e3-34c2-4e43-8e03-944febf9a67d · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? AI Alignment: A Comprehensive Survey
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d735d05-947e-4bdd-870b-97b7e625be5c · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Reward (mis) design for autonomous driving
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5ccb960a-0681-47fb-aa64-390ab974365a · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 110d7045-3500-4309-a6a5-1db8d3a2757e · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Mohri, A
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6ab36fc5-0af2-4c6c-a383-e4de8b3f8637 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Overcoming exploration in reinforcement learning with demonstrations
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 577facfb-b71c-4d97-90df-45987cc9d9eb · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c67ff398-8c69-499d-a6a7-05cd5c492973 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? A deep reinforced model for abstractive summarization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 82bbe396-158a-46d0-b37e-a69f0ec74169 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Learning Transferable Visual Models From Natural Language Supervision
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91d54c79-cc00-4caf-b3c8-9d69c3fcdb08 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Direct preference optimization: Your language model is secretly a reward model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c46f08b-c16d-4d5b-b3dc-097a730d8a2d · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e7637ad-d65a-4802-ad75-b54a9395c9da · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85fea4fa-87cb-4d5f-b2e5-0d93ad894e22 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b117e539-d0b6-46b2-91d8-be09761a4c97 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Proximal Policy Optimization Algorithms
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 471b38f0-0d6c-4198-9d2d-de0322efe506 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 699e0a8f-d3cd-489a-b6e3-dfbdf36a16e7 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? A Long Way to Go: Investigating Length Correlations in RLHF
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 088f894f-3ba5-4cf3-b954-7382fe9c58a0 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Defining and characterizing reward gaming
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3177201a-82f7-43eb-9a78-53a765599aa0 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Causal Confusion and Reward Misidentification in Preference-Based Reward Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa9b29dd-165f-484e-b684-85273c1ce7e5 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Relatively rational: Learning utilities and rationalities jointly from pairwise preferences
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 06152cda-8b23-4307-9751-c981a0d929f3 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Bayesian Reward Models for LLM Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a676381-6f24-4375-a3b8-c54775f34c5d · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation da2098a3-63e0-4281-bc6d-d3f3f7282467 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Larger and more instructable language models become less reliable
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4131ff1-1cf1-4919-bf85-d5d5a7d31616 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8c41436-e5c7-48b3-9cc7-d1cf47ff52b8 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Consequences of misaligned ai
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d2848552-7265-461d-a959-0e9bcf4e8739 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? @esa (Ref
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59d96322-98f3-4cfb-92ff-49cb8412f5a5 · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7555979-9685-4b71-b5cc-418ee489e08c · outbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80e08e02-5c8f-49fd-9790-1f82d1c7f942 · inbound
Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? When Can Proxies Improve the Sample Complexity of Preference Learning?
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.