Pith. sign in

Paper Citation Record · LEDGER

When Can Proxies Improve the Sample Complexity of Preference Learning?

As of 24 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2412.16475.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16475 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:41:24.133597Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:46.439486Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:15:46.504789Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0697c033-2049-4b76-a444-cb22ee7d507a · outbound

This paper cites write newline.

When Can Proxies Improve the Sample Complexity of Preference Learning? write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:23.850580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:23.850580Z digest=sha256:36edbb81d2fd27353d3d278479a06e3de8337953cb2cb06fd9649d2dfb50d3b5

Observation 0407d3c7-024b-4ad7-905a-9e86d08624af · outbound

This paper cites GPT-4 Technical Report.

When Can Proxies Improve the Sample Complexity of Preference Learning? GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:23.858731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:23.858731Z digest=sha256:0ea1ad789b319c396f6d49658b64a1c27159380c051965af5e682f8f76f0994e

Observation 06aa6455-3a21-4490-8e1a-8d7a295660fc · outbound

This paper cites Accuracy of chatgpt, google bard, and microsoft bing for simplifying radiology reports.

When Can Proxies Improve the Sample Complexity of Preference Learning? Accuracy of chatgpt, google bard, and microsoft bing for simplifying radiology reports

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:25.132545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:23.866506Z digest=sha256:b9092b9ca51abb2847db653e5efe31e5945d570cf3997170ecef30667316bf9c

Observation 2f75013d-35b0-4a43-9e98-8f815536e98b · outbound

This paper cites The evolved radio and its implications for modelling the evolution of novel sensors.

When Can Proxies Improve the Sample Complexity of Preference Learning? The evolved radio and its implications for modelling the evolution of novel sensors

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:25.116027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:23.876196Z digest=sha256:172a1acc3deb6da16599bcc26fdbcb5831b5482c82540121ccbe34823a8b56a0

Observation 34dbd208-86fa-4304-ad0c-2627754f1232 · outbound

This paper cites an unresolved cited work.

When Can Proxies Improve the Sample Complexity of Preference Learning? Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:23.884929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:23.884929Z digest=sha256:e8604e23a9ce6fe7f7dcdb775d88cf092191187f77bb19c361a919519012fe47

Observation 11dbaa5e-9aec-4357-a9b0-a12929b8320d · outbound

This paper cites Open problems and fundamental limitations of reinforcement learning from human feedback.

When Can Proxies Improve the Sample Complexity of Preference Learning? Open problems and fundamental limitations of reinforcement learning from human feedback

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:25.088000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:23.890654Z digest=sha256:940ea7626fe2596fc3d452eec25ef5465ef607e50f5b4427b7c1cd83e1f0efb3

Observation 931cb768-d394-44b8-b278-cd9433a34a84 · outbound

This paper cites MaxMin-RLHF: Alignment with Diverse Human Preferences.

When Can Proxies Improve the Sample Complexity of Preference Learning? MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:23.896444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:23.896444Z digest=sha256:7cc7900e2e701d7c1d53952d3700ee2293cf18f5bd3e97bc6e52bb4e7b70f96d

Observation 908181f5-b5b9-488b-8ebf-c5d191ee899c · outbound

This paper cites ODIN : Disentangled reward mitigates hacking in RLHF.

When Can Proxies Improve the Sample Complexity of Preference Learning? ODIN : Disentangled reward mitigates hacking in RLHF

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:25.070443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:23.902736Z digest=sha256:6093aa21432314a79b9126357f74ccf208f10a28479386e462fedfa35514ffe7

Observation 2b7d8406-ef65-4f0e-844e-f51d81e2b508 · outbound

This paper cites Meta discovery: Learning to discover novel classes given very limited data.

When Can Proxies Improve the Sample Complexity of Preference Learning? Meta discovery: Learning to discover novel classes given very limited data

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:25.050708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:23.909653Z digest=sha256:b8f4688668076b0aae58f8657bf1a62ad22c2166e09012c4b5a36cedaaf3ce6d

Observation ff703073-f67b-47a9-99cb-6265db905465 · outbound

This paper cites Faulty reward functions in the wild, 2016.

When Can Proxies Improve the Sample Complexity of Preference Learning? Faulty reward functions in the wild, 2016

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:25.030561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:23.915877Z digest=sha256:5b613fc5a558666833411b37ea2621bb99fe3c5d55ef09762984c7a731493920

Observation 6c9fbf7e-f935-4922-90cf-f1ef8190148f · outbound

This paper cites Reward model ensembles help mitigate overoptimization.

When Can Proxies Improve the Sample Complexity of Preference Learning? Reward model ensembles help mitigate overoptimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:25.010653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:23.922221Z digest=sha256:a2a47df74c12be552a269889d1ec92715e9be8a738a11ccfbbaf90a40208593e

Observation 7c5f317f-021c-437f-83f9-5b2a52b7125f · outbound

This paper cites The Expertise Problem: Learning from Specialized Feedback.

When Can Proxies Improve the Sample Complexity of Preference Learning? The Expertise Problem: Learning from Specialized Feedback

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-11T10:41:24.502247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:23.928788Z digest=sha256:36d1d4d6cd0f0346fe8dad8d9d5c5a65f34d56ce441315b4ff945201bd9ca3dc

Observation d231620d-24d0-4340-9d71-89640e31f792 · outbound

This paper cites Group symmetry in pac learning.

When Can Proxies Improve the Sample Complexity of Preference Learning? Group symmetry in pac learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:24.986038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:23.935156Z digest=sha256:61913187934774c270d08a1cfdd6d9089fcfe77a9d738de883a66b9d21b89bd5

Observation 38a0f1d1-5cba-438d-8bfe-99a2c9b01b88 · outbound

This paper cites Active teacher selection for reward learning.

When Can Proxies Improve the Sample Complexity of Preference Learning? Active teacher selection for reward learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:23.943994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:23.943994Z digest=sha256:26189a7b09ca12ef23d2413e0633e64dea35add45ce4481ef83b5e8fd3ae8289

Observation 0b7e3efa-6565-4eed-8135-f71a844c3077 · outbound

This paper cites Scaling laws for reward model overoptimization.

When Can Proxies Improve the Sample Complexity of Preference Learning? Scaling laws for reward model overoptimization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:24.967509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:23.951138Z digest=sha256:84bb5e5dce6b681a0ce9759230227e7fa5e503f89f24557f65d46246aeccc778

Observation b96e6e82-1da8-4353-88e4-cf4087f4ba94 · outbound

This paper cites Glass floor colleges reject top applicants, accepting only the students likely to enroll, 2001.

When Can Proxies Improve the Sample Complexity of Preference Learning? Glass floor colleges reject top applicants, accepting only the students likely to enroll, 2001

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:24.945572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:23.963560Z digest=sha256:5ca2470100f6a2db694651d590f7b3568d31b980d2802e0a6615768de3b908ec

Observation d459e898-3910-4869-9dc8-9090ce61848d · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

When Can Proxies Improve the Sample Complexity of Preference Learning? Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:23.970292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:23.970292Z digest=sha256:3331a0906155629e09f4e3f18bb3c11d5e9b72e64e5fa7f99e0ac65ebd4c44c9

Observation cc13a5e3-34c2-4e43-8e03-944febf9a67d · outbound

This paper cites AI Alignment: A Comprehensive Survey.

When Can Proxies Improve the Sample Complexity of Preference Learning? AI Alignment: A Comprehensive Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:23.977822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:23.977822Z digest=sha256:b99bda7ebd734d6913a5b503120b1392e5eb7a67c504d3260c883d475bba474b

Observation 7d735d05-947e-4bdd-870b-97b7e625be5c · outbound

This paper cites Reward (mis) design for autonomous driving.

When Can Proxies Improve the Sample Complexity of Preference Learning? Reward (mis) design for autonomous driving

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:24.782930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:23.985660Z digest=sha256:0f3ff7bbb13667395494a49748f5adb005d75ec26ab07be64653fffa2e64b409

Observation 5ccb960a-0681-47fb-aa64-390ab974365a · outbound

This paper cites InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling.

When Can Proxies Improve the Sample Complexity of Preference Learning? InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:23.993491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:23.993491Z digest=sha256:95f20b481b656e51e654f037096519662f4be4b42378a78af7487797641a85f9

Observation 110d7045-3500-4309-a6a5-1db8d3a2757e · outbound

This paper cites Mohri, A.

When Can Proxies Improve the Sample Complexity of Preference Learning? Mohri, A

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:24.763858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:23.998936Z digest=sha256:3dafc563dc5f6b7ca5e6a3662e273fbaf868a75cbe6526c9388bc9073096f007

Observation 6ab36fc5-0af2-4c6c-a383-e4de8b3f8637 · outbound

This paper cites Overcoming exploration in reinforcement learning with demonstrations.

When Can Proxies Improve the Sample Complexity of Preference Learning? Overcoming exploration in reinforcement learning with demonstrations

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:24.746697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:24.004198Z digest=sha256:0440c55900923143b3b05c7ee8a12797712508180f9231ba5bb51ced9f7910cd

Observation 577facfb-b71c-4d97-90df-45987cc9d9eb · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

When Can Proxies Improve the Sample Complexity of Preference Learning? The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.008976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.008976Z digest=sha256:a3a177a6ef94109f376f560fbce68978cada1bf0ca93d32fe0b387119ec81b0e

Observation c67ff398-8c69-499d-a6a7-05cd5c492973 · outbound

This paper cites A deep reinforced model for abstractive summarization.

When Can Proxies Improve the Sample Complexity of Preference Learning? A deep reinforced model for abstractive summarization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:24.726845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:24.018198Z digest=sha256:fa37a36a0e7241e6d2069ba4880edc9a79904e04283f9b8c36163dd48ec386d5

Observation 82bbe396-158a-46d0-b37e-a69f0ec74169 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

When Can Proxies Improve the Sample Complexity of Preference Learning? Learning Transferable Visual Models From Natural Language Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.028935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.028935Z digest=sha256:6bcd818c702f5706c0200693f7835d923b24f9b05663c4ee03b85f7c38c42879

Observation 91d54c79-cc00-4caf-b3c8-9d69c3fcdb08 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

When Can Proxies Improve the Sample Complexity of Preference Learning? Direct preference optimization: Your language model is secretly a reward model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.034379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.034379Z digest=sha256:7a9c0f275ae1544372dda4d398b3602e32cc2226fcee8e68f8ecfcd04989fd50

Observation 7c46f08b-c16d-4d5b-b3dc-097a730d8a2d · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

When Can Proxies Improve the Sample Complexity of Preference Learning? Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.040308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.040308Z digest=sha256:21c131bcc374ead01f2b32be2f0049ffe7d1dcbcaa0b3df9703b6a6ad1b47ff9

Observation 6e7637ad-d65a-4802-ad75-b54a9395c9da · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

When Can Proxies Improve the Sample Complexity of Preference Learning? Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.046339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.046339Z digest=sha256:8dc43ceeed8f23f815513abf22f602a4c316c2fb372fc2c55d3319833790e417

Observation 85fea4fa-87cb-4d5f-b2e5-0d93ad894e22 · outbound

This paper cites Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning.

When Can Proxies Improve the Sample Complexity of Preference Learning? Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.051680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.051680Z digest=sha256:6cf80b6a4b05a4bd98f60c7d70d31965405ae867c2b88a6a0062fc2d9a6fce7d

Observation b117e539-d0b6-46b2-91d8-be09761a4c97 · outbound

This paper cites Proximal Policy Optimization Algorithms.

When Can Proxies Improve the Sample Complexity of Preference Learning? Proximal Policy Optimization Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.057383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.057383Z digest=sha256:8ce1e3d523ed65960a706a0d338ed1bfcb1ceb22883c9955d7f5bd8eaabe0f41

Observation 471b38f0-0d6c-4198-9d2d-de0322efe506 · outbound

This paper cites Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback.

When Can Proxies Improve the Sample Complexity of Preference Learning? Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:24.696125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:24.063206Z digest=sha256:78499297944024be010e6f0b838fde6043a8adc12f5e209870c37151236aae17

Observation 699e0a8f-d3cd-489a-b6e3-dfbdf36a16e7 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

When Can Proxies Improve the Sample Complexity of Preference Learning? A Long Way to Go: Investigating Length Correlations in RLHF

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.069209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.069209Z digest=sha256:e10de298b0530b8010d749ae1f4238dd6dda3c5e7d8b87891c74278e9afc8d72

Observation 088f894f-3ba5-4cf3-b954-7382fe9c58a0 · outbound

This paper cites Defining and characterizing reward gaming.

When Can Proxies Improve the Sample Complexity of Preference Learning? Defining and characterizing reward gaming

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.074435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.074435Z digest=sha256:eb2542c9c96d34838c80bc83004eec33d594bfdc0ee10a0a2ae4a61d2e531f5d

Observation 3177201a-82f7-43eb-9a78-53a765599aa0 · outbound

This paper cites Causal Confusion and Reward Misidentification in Preference-Based Reward Learning.

When Can Proxies Improve the Sample Complexity of Preference Learning? Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.080607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.080607Z digest=sha256:1731164563c3eeb7cb2f92fc5c858455c487665a29962fe5cdd99586574bf5aa

Observation aa9b29dd-165f-484e-b684-85273c1ce7e5 · outbound

This paper cites Relatively rational: Learning utilities and rationalities jointly from pairwise preferences.

When Can Proxies Improve the Sample Complexity of Preference Learning? Relatively rational: Learning utilities and rationalities jointly from pairwise preferences

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:24.664012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:24.086756Z digest=sha256:46c78e8c24ba61b85f176ead999e8eaaef02260bbfc4950d4b83ff9eadc3b9f6

Observation 06152cda-8b23-4307-9751-c981a0d929f3 · outbound

This paper cites Bayesian Reward Models for LLM Alignment.

When Can Proxies Improve the Sample Complexity of Preference Learning? Bayesian Reward Models for LLM Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.092589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.092589Z digest=sha256:3443bf04ac9b2a1b38ee9888c0250df7284dd20a900ae63711a931f49e51544e

Observation 3a676381-6f24-4375-a3b8-c54775f34c5d · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

When Can Proxies Improve the Sample Complexity of Preference Learning? Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:24.641039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:24.098121Z digest=sha256:7276065deb85cf9d81a959025b2a326e7fe26c384b76205a4e7d92aed84ca4fb

Observation da2098a3-63e0-4281-bc6d-d3f3f7282467 · outbound

This paper cites Larger and more instructable language models become less reliable.

When Can Proxies Improve the Sample Complexity of Preference Learning? Larger and more instructable language models become less reliable

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.103363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.103363Z digest=sha256:dcd27972c43145f801f0c9fcffdb616186f8c220ee4a77aaecc6a65c71cf3ba7

Observation b4131ff1-1cf1-4919-bf85-d5d5a7d31616 · outbound

This paper cites Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF.

When Can Proxies Improve the Sample Complexity of Preference Learning? Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.109274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.109274Z digest=sha256:0015555dea13fc539634c5d94e014b92a72dc18b30c6639463c04ef1723f5e84

Observation d8c41436-e5c7-48b3-9cc7-d1cf47ff52b8 · outbound

This paper cites Consequences of misaligned ai.

When Can Proxies Improve the Sample Complexity of Preference Learning? Consequences of misaligned ai

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:41:24.618246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T10:41:24.114858Z digest=sha256:6c37e7de2e28ddc46689e3011983e4981feed5660ca4310da34b8361f3ddac11

Observation d2848552-7265-461d-a959-0e9bcf4e8739 · outbound

This paper cites @esa (Ref.

When Can Proxies Improve the Sample Complexity of Preference Learning? @esa (Ref

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.120042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.120042Z digest=sha256:241c90da9283d710d19fcac3bc9b77d3468195f9839c33c429ab9acaf02bc269

Observation 59d96322-98f3-4cfb-92ff-49cb8412f5a5 · outbound

This paper cites an unresolved cited work.

When Can Proxies Improve the Sample Complexity of Preference Learning? Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.126138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.126138Z digest=sha256:98531a7ac4a2be4b3934b836adadbe2353ea2cf2c666b82abf7582165d7c8d0a

Observation d7555979-9685-4b71-b5cc-418ee489e08c · outbound

This paper cites an unresolved cited work.

When Can Proxies Improve the Sample Complexity of Preference Learning? Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.133597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.133597Z digest=sha256:d6aa26c26a9881a560f7ae501273d0547cbfa990ca15f1727ebf5ae84425195d

Pith citing papers

Observation 80e08e02-5c8f-49fd-9790-1f82d1c7f942 · inbound

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? cites this paper.

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? When Can Proxies Improve the Sample Complexity of Preference Learning?

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:15:46.510460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T15:15:46.439486Z digest=sha256:658c66a27156c7e829e541af2d17434e1918593e2d575d97c3bb5be0dcfae999