Pith. sign in

Paper Citation Record · LEDGER

PILAF: Optimal Human Preference Sampling for Reward Modeling

As of 9 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 3 inbound Pith citation observations for arXiv:2502.04270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04270 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:03:53.312441Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:34:25.000468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T01:09:19.294881Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd111811-bc5b-49eb-bf8e-46c0ebfbf20b · outbound

This paper cites write newline.

PILAF: Optimal Human Preference Sampling for Reward Modeling write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:52.987311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:52.987311Z digest=sha256:c6bd6acf8748585d73e6166a10f2084ee107191cbdb232907ea03241196462df

Observation b78b1f42-43ce-4d99-8b4a-a286343de7de · outbound

This paper cites GPT-4 Technical Report.

PILAF: Optimal Human Preference Sampling for Reward Modeling GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:52.994878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:52.994878Z digest=sha256:d2f7d48fd74206791ce1fc06bc623e55902c68ac3dbf7a8a8b75f4a913141ebb

Observation f6bedace-62c6-40fa-b99b-9e3ddcc879ca · outbound

This paper cites G., Guo, Z.

PILAF: Optimal Human Preference Sampling for Reward Modeling G., Guo, Z

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.002290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.002290Z digest=sha256:0c5d53f77469f5159c2df06198b7997027c920dd51543ee9c1df17e3d26454cf

Observation d366c147-6187-445b-a00b-cad225c9df7c · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

PILAF: Optimal Human Preference Sampling for Reward Modeling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.008646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.008646Z digest=sha256:4d7a2965017eb54ee77ccbabb681571d83da7c58c5f9fbf6ba2495713a4e7e41

Observation 838970c1-5395-4dac-8aa1-c3ae94305c32 · outbound

This paper cites Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF.

PILAF: Optimal Human Preference Sampling for Reward Modeling Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.016827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.016827Z digest=sha256:4170e7798700254265f442f8b89002a8ed1c9f6a439e0753fc22a90a140ef1a4

Observation 947a26fa-6a17-409c-b893-4e575b955a2a · outbound

This paper cites an unresolved cited work.

PILAF: Optimal Human Preference Sampling for Reward Modeling Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.022691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.022691Z digest=sha256:89d11207783d28a0e5e4c70b446833cc46cc8ae262c513a1787d14aa7be524c3

Observation 63871527-a189-4f20-b578-52c2574926f1 · outbound

This paper cites Off-Policy Actor-Critic.

PILAF: Optimal Human Preference Sampling for Reward Modeling Off-Policy Actor-Critic

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.028525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.028525Z digest=sha256:e943ae92a0c2224a091edaf4fec4acf5c9ec4c56a110f994377716e7e24ff417

Observation 761777f9-ca75-45c0-b01f-9458df276dd2 · outbound

This paper cites Raft: Reward ranked finetuning for generative foundation model alignment.

PILAF: Optimal Human Preference Sampling for Reward Modeling Raft: Reward ranked finetuning for generative foundation model alignment

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:03:54.430520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:03:53.034733Z digest=sha256:121c70138841ddd0846a54f8dd8220bf6e8f82c3de14ea4e1d9a866b594acb48

Observation 2c0eab95-ced4-4f35-ac75-f71210376015 · outbound

This paper cites Rlhf workflow: From reward modeling to online rlhf, 2024.

PILAF: Optimal Human Preference Sampling for Reward Modeling Rlhf workflow: From reward modeling to online rlhf, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:03:54.404208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:03:53.040539Z digest=sha256:eaecefd5c54e7d2b45cc8d2432ec1d6cdcdefa686a1e473579b746f109c7ecb5

Observation 8ad2621a-5c92-4661-9397-cf82252950d7 · outbound

This paper cites The Llama 3 Herd of Models.

PILAF: Optimal Human Preference Sampling for Reward Modeling The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.045855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.045855Z digest=sha256:83f85ac9b8475808ccc22200aff1ab8ab45898ba2cf5036cc4462fe93c6e5181

Observation 1150a4df-239e-4938-916c-1c35dff9eef5 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

PILAF: Optimal Human Preference Sampling for Reward Modeling KTO: Model Alignment as Prospect Theoretic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.051581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.051581Z digest=sha256:30a5d77a824dedb6901eeb09e3b0770e4b0ac698bc6c73e6771c4dd9d8d380f7

Observation e40fb42b-1869-4cf7-ae12-60c7794bb127 · outbound

This paper cites Scaling laws for reward model overoptimization.

PILAF: Optimal Human Preference Sampling for Reward Modeling Scaling laws for reward model overoptimization

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:03:54.377898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:03:53.059076Z digest=sha256:aa235a3f5469ca5c696b2e86fee2bc72639c4f6fdff875f41cb2405bdaf2fb91

Observation 6cc714cd-169f-4aea-b35e-ed6641c58a34 · outbound

This paper cites L., and Baxter, J.

PILAF: Optimal Human Preference Sampling for Reward Modeling L., and Baxter, J

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:03:54.352936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:03:53.067691Z digest=sha256:859d13d7cdf1ab05d67dee9dcacdb43f6dfca8d4a772f704ec582849f588cf18

Observation 91269a7e-356b-475a-a7ee-5548ece0c4cd · outbound

This paper cites S., Lillicrap, T., Turner, R.

PILAF: Optimal Human Preference Sampling for Reward Modeling S., Lillicrap, T., Turner, R

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:03:54.334407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:03:53.073245Z digest=sha256:a9c1448e16628c2378dc3ca1eff67e86953adbec04850fc5fca38efcf41ffb10

Observation 45ccfa6a-012b-4955-a9d8-ff9a8851a250 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

PILAF: Optimal Human Preference Sampling for Reward Modeling Direct Language Model Alignment from Online AI Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.079636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.079636Z digest=sha256:842dc6e2b989a6c287577a20eea1498b8e40e4dad4ddc4d26e8568fe3b31d4f2

Observation 9915083f-08c0-4946-94d3-e01d97f12f3e · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

PILAF: Optimal Human Preference Sampling for Reward Modeling OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.088095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.088095Z digest=sha256:5de3effa61673a112a41055c0f862668f957190cb2eb31a6db7fd6ac276ae5bf

Observation b963b863-fbed-47c6-a40a-e41b21635caf · outbound

This paper cites Reinforcement Learning from Human Feedback with Active Queries.

PILAF: Optimal Human Preference Sampling for Reward Modeling Reinforcement Learning from Human Feedback with Active Queries

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.095156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.095156Z digest=sha256:0d9f78bf84623b370afa37bd6984e0ade144dc0974cda61475edb9a26362ce58

Observation 9bdf1632-d609-4672-8d08-64d1603b2c39 · outbound

This paper cites and Uehara, M.

PILAF: Optimal Human Preference Sampling for Reward Modeling and Uehara, M

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:03:54.310951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:03:53.101188Z digest=sha256:781278c8ae5d653fed54025868066c2a38545630596136c6023c74e43edd1c21

Observation fd3d0c60-2e79-4346-8888-ecfcda468890 · outbound

This paper cites an unresolved cited work.

PILAF: Optimal Human Preference Sampling for Reward Modeling Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-08T23:03:54.282274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:03:53.117344Z digest=sha256:71578f4aed2753cfc0bb60374e37d6288132a2aa16f8190af048b483b842b044

Observation 84c754f6-4510-434f-9f67-a9ebe0961861 · outbound

This paper cites Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., Smith, N.

PILAF: Optimal Human Preference Sampling for Reward Modeling Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., Smith, N

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:03:54.263739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:03:53.123719Z digest=sha256:407713ae44619a61fb4842ba97d156a57ec950827a2070bbb18ce27a390b73e3

Observation f6dfccbc-8622-4394-a503-5663da3b61b8 · outbound

This paper cites Let's Verify Step by Step.

PILAF: Optimal Human Preference Sampling for Reward Modeling Let's Verify Step by Step

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.130190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.130190Z digest=sha256:a344ec8f2ef4cd53ddf1a911457b5ae37814cba4b6cd424aaf2442506d38a3ca

Observation 0405fa58-2304-4dc4-b7ec-12ea28d1299a · outbound

This paper cites an unresolved cited work.

PILAF: Optimal Human Preference Sampling for Reward Modeling Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-08T23:03:54.234159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:03:53.136921Z digest=sha256:1c54537fc529af1c9a357ffb9b33242075025503610b720a83c7d1ce89bf6801

Observation ed352f6a-cebd-4543-a119-bf15d9a7874b · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

PILAF: Optimal Human Preference Sampling for Reward Modeling Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.142987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.142987Z digest=sha256:b5ab96f16271b68765b616aeb0fafb8750d80aa16f0da14ecb5c31ab7869ec97

Observation 0ae139a1-50fd-426e-833d-3e0c37fbb40f · outbound

This paper cites Decoding-time Realignment of Language Models.

PILAF: Optimal Human Preference Sampling for Reward Modeling Decoding-time Realignment of Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.148096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.148096Z digest=sha256:710a48dbb132ed99d8ad0b28c65fe915ad7c1895b4dd099ec44b97cfbd7c7ff0

Observation 43773dbb-d1c3-4087-a2c6-1ad8d87638cc · outbound

This paper cites J., and Liu, J.

PILAF: Optimal Human Preference Sampling for Reward Modeling J., and Liu, J

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:03:54.194534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:03:53.153316Z digest=sha256:5e56f1f1eb7a46b44ffa7130c113fd2a9a3de93d404cfc53c879054fbf3cff29

Observation d5bd172c-6ebc-44d1-a806-25d0f200b1b2 · outbound

This paper cites Sample-Efficient Alignment for LLMs.

PILAF: Optimal Human Preference Sampling for Reward Modeling Sample-Efficient Alignment for LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.158373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.158373Z digest=sha256:8478804e3d62201eecb0f5172a57da7e5b75edd7c3a45f1f3e2521e0c7fcfb42

Observation 21a221cf-c917-4b45-a6cd-1a665e0ac4c6 · outbound

This paper cites Sample Efficient Preference Alignment in LLMs via Active Exploration.

PILAF: Optimal Human Preference Sampling for Reward Modeling Sample Efficient Preference Alignment in LLMs via Active Exploration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.163898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.163898Z digest=sha256:d0826fb3bb919a992d25fdb0d009f3ddd6a5d68576d9e2f30233fdcb7ae7488a

Observation 391fe1c4-7614-4114-af2f-2c2254b3ff68 · outbound

This paper cites Active preference learning for large language models.

PILAF: Optimal Human Preference Sampling for Reward Modeling Active preference learning for large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:03:54.160584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:03:53.172195Z digest=sha256:5f31bcf451be98a8357f9c14d7231ccf329ffdcbd43e9f43524fe5ebe38512db

Observation 4c2ea069-f21c-4644-a552-e539434b2813 · outbound

This paper cites Training language models to follow instructions with human feedback.

PILAF: Optimal Human Preference Sampling for Reward Modeling Training language models to follow instructions with human feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.179236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.179236Z digest=sha256:d81cc9da9474851b7691f20a87f58708cc22f9caad506fac7637e8ecb987cba7

Observation 9a4e0b24-88bd-4097-b46c-05ca1cf15a58 · outbound

This paper cites D., Ermon, S., and Finn, C.

PILAF: Optimal Human Preference Sampling for Reward Modeling D., Ermon, S., and Finn, C

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.188424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.188424Z digest=sha256:60af65acf7f57eb68deede7e9894d2edfe7e621b063231e4ade4db4c942f48a1

Observation 31a84014-6480-4388-aaaa-8d299a07ea28 · outbound

This paper cites Optimal Design for Reward Modeling in RLHF.

PILAF: Optimal Human Preference Sampling for Reward Modeling Optimal Design for Reward Modeling in RLHF

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.198647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.198647Z digest=sha256:fcd59acd7d02dcf7172006ec830840e58d669f6fdfb90817b6b125d358f6e18a

Observation addf575c-03f8-4356-b91d-4568330cad1a · outbound

This paper cites Proximal Policy Optimization Algorithms.

PILAF: Optimal Human Preference Sampling for Reward Modeling Proximal Policy Optimization Algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.204982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.204982Z digest=sha256:02535f64296a6e73c5e0cc92c744334ea8527d901c4dec98511acc2ff01319c3

Observation 00c51ad8-6c46-4b97-819d-695d06df7778 · outbound

This paper cites The Crucial Role of Samplers in Online Direct Preference Optimization.

PILAF: Optimal Human Preference Sampling for Reward Modeling The Crucial Role of Samplers in Online Direct Preference Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.212190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.212190Z digest=sha256:5bad07c583b7d23af281b27075f22ff6fdac843edbd269b06564a05b9a41293b

Observation 94df5939-c90a-40eb-8662-ed50ecc27064 · outbound

This paper cites Deterministic policy gradient algorithms.

PILAF: Optimal Human Preference Sampling for Reward Modeling Deterministic policy gradient algorithms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.219913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.219913Z digest=sha256:842d08d3596d35fb5a00b3ec9e9ccb6fc5cfa7c00e8d4f7693ac13ba662e1245

Observation 0c0876af-7dad-40b8-a74d-14472e586c96 · outbound

This paper cites S., McAllester, D., Singh, S., and Mansour, Y.

PILAF: Optimal Human Preference Sampling for Reward Modeling S., McAllester, D., Singh, S., and Mansour, Y

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.231744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.231744Z digest=sha256:db91d124f1bf0c9812484011ff0ac86f28a4e21e428653f9d639dc94a460d96f

Observation f14a3650-1869-402e-8243-0d3f6210273e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

PILAF: Optimal Human Preference Sampling for Reward Modeling Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.249843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.249843Z digest=sha256:ad05335b7f27f643141c9d9d29e62ceb11d97bd7764f901b2977a660a28ad614

Observation 271913c1-26bf-46d4-b332-b737d0649152 · outbound

This paper cites an unresolved cited work.

PILAF: Optimal Human Preference Sampling for Reward Modeling Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.257479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.257479Z digest=sha256:a69d756273ab24e5942b7389ce02a18fb3430ebbec98d99dc063eb268fdc2243

Observation 2c180d6a-0f93-4327-9a29-0f37aeb3f691 · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

PILAF: Optimal Human Preference Sampling for Reward Modeling Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.263456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.263456Z digest=sha256:87771a4fa4e2cdc0a26c6df192d5e8c7a378dad33d393845940dd21c2aedd463

Observation 50d3db4a-b21c-490e-9b23-72a1243ea515 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint.

PILAF: Optimal Human Preference Sampling for Reward Modeling Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:03:54.013289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:03:53.269428Z digest=sha256:c0426c7991f4acbd3d1727c64d93c7ec8e21167bcab903d0eb700a9e12c5a24c

Observation 2179c1ad-d998-4e90-9b57-cec8f33112ff · outbound

This paper cites an unresolved cited work.

PILAF: Optimal Human Preference Sampling for Reward Modeling Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-08T23:03:53.986431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:03:53.275113Z digest=sha256:077d9d55b8912e67e52f314595b8ed98051201d86314129eebb9897534f37b7d

Observation 64b1dcd3-0cac-4cc7-a338-22b0006d5f07 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

PILAF: Optimal Human Preference Sampling for Reward Modeling Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.280173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.280173Z digest=sha256:b77cd29a19db8bd4422b641d667842ce8aa0c64290289c82190abef6fd27533b

Observation 564d187d-2518-475f-9d04-2b782da60d06 · outbound

This paper cites Inference Scaling for Long-Context Retrieval Augmented Generation.

PILAF: Optimal Human Preference Sampling for Reward Modeling Inference Scaling for Long-Context Retrieval Augmented Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.286885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.286885Z digest=sha256:e8ee8a5bc165d7c6031e23d81ec19045e306860512772895d91d3e420ceb53f6

Observation c09f1447-2d90-4e96-bf86-e8d9d5c36f80 · outbound

This paper cites Self-Exploring Language Models: Active Preference Elicitation for Online Alignment.

PILAF: Optimal Human Preference Sampling for Reward Modeling Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.293003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.293003Z digest=sha256:4ded47c3d990f87a35f7ae0722ca1e6ceaf6dad39d9729448997951666434391

Observation a929772c-ed3a-491f-bbb6-b01a47469ff6 · outbound

This paper cites @esa (Ref.

PILAF: Optimal Human Preference Sampling for Reward Modeling @esa (Ref

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.299180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.299180Z digest=sha256:e1e063e81665bd51ec10c561097cae947fd57b8d7a24a446757bbb23b0f16d64

Observation 6c99ebb6-acea-4927-83e3-1c108421f5c1 · outbound

This paper cites an unresolved cited work.

PILAF: Optimal Human Preference Sampling for Reward Modeling Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.307149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.307149Z digest=sha256:950a1df4fa0cf26c977035e2eec9fd179a468f543259db363e6ff9425b75fea8

Observation d9d65432-a03b-4fc4-937b-532a80d351f6 · outbound

This paper cites an unresolved cited work.

PILAF: Optimal Human Preference Sampling for Reward Modeling Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.312441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.312441Z digest=sha256:6d5e24577f83f8e294dc441e461ab22285250b1041bec832d146f396dc162700

Pith citing papers

Observation afdb69b1-e454-4c5a-b4ef-a50ba8294804 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities PILAF: Optimal Human Preference Sampling for Reward Modeling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.000468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.000468Z digest=sha256:840a12b7849ad332df5bedf2f50a564d030c3562da83f8bc134fbd05ea7abbb3

Observation 17a7ff4f-ca29-4304-8976-aa512fab2665 · inbound

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses cites this paper.

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses PILAF: Optimal Human Preference Sampling for Reward Modeling

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:58.750065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:13:29.292351Z digest=sha256:e1a17cb1807ae3f738d2a74c2c3517c412e5fb8fa4c830ddf32344cc778b3646

Observation edfe710a-2dec-40be-a4fb-e9c1886a31d9 · inbound

Which Pairs to Compare for LLM Post-Training? cites this paper.

Which Pairs to Compare for LLM Post-Training? PILAF: Optimal Human Preference Sampling for Reward Modeling

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:09:19.296921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T20:37:19.221848Z digest=sha256:a1a3b308c5747085d3037edb34497e55ee74b3942fdc6e9a6163c616dfb8cc86