Pith. sign in

Paper Citation Record · LEDGER

Preference learning made easy: Everything should be understood through win rate

As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2502.10505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10505 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:32:21.331060Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy36
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6180138-a454-4301-9e33-643e150b0bb8 · outbound

This paper cites Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms.

Preference learning made easy: Everything should be understood through win rate Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:23.028818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.903864Z digest=sha256:cb772396b2ef9f4652879428337206b33b93fdee12035d014eed5add57453c20

Observation 1b074722-1adc-408b-b4cb-9b3f3040dc52 · outbound

This paper cites Calibration and consistency of adversarial surrogate losses.

Preference learning made easy: Everything should be understood through win rate Calibration and consistency of adversarial surrogate losses

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:23.007001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.910320Z digest=sha256:0147abb5bf2ac93f8122c0f1561b0cb65b519549d41db13645db9198ece2f0c9

Observation 4c251466-16c3-4b63-b3ac-5d2ad985737c · outbound

This paper cites Multi-class h -consistency bounds.

Preference learning made easy: Everything should be understood through win rate Multi-class h -consistency bounds

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.983272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.916903Z digest=sha256:3bdbf6fcc0cb1638fa24c0ab5d489afd577d8f1cad5da69e71a6c8766845b743

Observation 4c2cc416-e36f-4b17-94ca-93af7a38517c · outbound

This paper cites G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R.

Preference learning made easy: Everything should be understood through win rate G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.965345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.924819Z digest=sha256:019d990b9f8de7f0b07be5b34e90bed73776d1cecdec4f6fcb4ba3bda341830d

Observation d8ab9d49-91d5-49cd-a64c-2dc628851e04 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback.

Preference learning made easy: Everything should be understood through win rate Training a helpful and harmless assistant with reinforcement learning from human feedback

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.938806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.932361Z digest=sha256:43f1a2d4243c0a6a046d16839d89449f01c87e91d63481ee917b7162d9cb16ea

Observation bad98efd-8a4b-48f9-9230-63a8f1eee886 · outbound

This paper cites Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.

Preference learning made easy: Everything should be understood through win rate Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:20.938380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:20.938380Z digest=sha256:58d538ef4dce5d3f3848322914d3216d44c5c4e9c2919b10604034e6eea3419f

Observation 5e77688e-f12f-4247-b719-10fae9e6a0b8 · outbound

This paper cites Quantile Filtered Imitation Learning.

Preference learning made easy: Everything should be understood through win rate Quantile Filtered Imitation Learning

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T18:32:22.012736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.946064Z digest=sha256:6dc790936cc1a139a860dc1e3c1f87a3008a8dd4ae0905a312118d0d6037fa35

Observation 43afe68a-4ec2-412e-85dc-3230471c94a9 · outbound

This paper cites Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition.

Preference learning made easy: Everything should be understood through win rate Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.919649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.952817Z digest=sha256:13363e7b78bcd900e571f613192a9a0f555c05ae84e3659be58b1391f8f7d4a6

Observation 9f07e96b-2159-4740-a1d1-875239dda663 · outbound

This paper cites Human Alignment of Large Language Models through Online Preference Optimisation.

Preference learning made easy: Everything should be understood through win rate Human Alignment of Large Language Models through Online Preference Optimisation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:20.963680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:20.963680Z digest=sha256:ad7ef22dca8128dc3f2b2994a59c2054038a32aa4932cf1a4d54ee0046e78555

Observation 8107c9da-f2c5-40f5-8b2f-330139a41e75 · outbound

This paper cites H., Chen, X., Zhang, Q., Ranganath, R., and Cho, K.

Preference learning made easy: Everything should be understood through win rate H., Chen, X., Zhang, Q., Ranganath, R., and Cho, K

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.898525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.971802Z digest=sha256:d2d493eaccf3df001732cfefd3c2b0e334c6c4a5c3f53bbfc605a32fb086146a

Observation 2902befc-3e72-43b6-acfe-88eb4cef412d · outbound

This paper cites Deep reinforcement learning from human preferences.

Preference learning made easy: Everything should be understood through win rate Deep reinforcement learning from human preferences

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:20.983164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:20.983164Z digest=sha256:cbcb36480596747e0c78f8b1d94dd4c58aac9362741d2e7a5715b89ad07034d1

Observation b619d067-3659-4495-ad93-2921e730fb13 · outbound

This paper cites Raft: Reward ranked finetuning for generative foundation model alignment.

Preference learning made easy: Everything should be understood through win rate Raft: Reward ranked finetuning for generative foundation model alignment

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.864549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.989748Z digest=sha256:0475575a5b80d575942e4732829cdc5e09517b154fa60c68a4150a1db29b3c07

Observation 51dd05ae-79b1-41fc-ba02-c0308782ba0f · outbound

This paper cites Understanding dataset difficulty with v-usable information.

Preference learning made easy: Everything should be understood through win rate Understanding dataset difficulty with v-usable information

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.839675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.001625Z digest=sha256:959028573a26bde299a1ea85f900b0528f4251008c4aca4ea81e569bf6736a42

Observation c5d841ae-b317-4623-9930-c2b9d97c89b4 · outbound

This paper cites Bonbon alignment for large language models and the sweetness of best-of-n sampling.

Preference learning made easy: Everything should be understood through win rate Bonbon alignment for large language models and the sweetness of best-of-n sampling

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.814056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.008913Z digest=sha256:a2caf09e483fa5f48ca1b6a0be27ac77899f1da3850d7ea30ee09f4646b53ea5

Observation b64d9244-8006-498c-96d9-4831e68de4b7 · outbound

This paper cites L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., and de Freitas, N.

Preference learning made easy: Everything should be understood through win rate L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., and de Freitas, N

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.788839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.015499Z digest=sha256:ad48f558d3636202eba629eff219617ff19bdc6c2d2d3a3a4948d5f22006d48d

Observation 5724ed26-706e-4021-a240-b524f4676b17 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Preference learning made easy: Everything should be understood through win rate Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.022389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.022389Z digest=sha256:f950d45ca21b76f2f398230e1951caa8e6ef124f0e6de254a1f9e15540357a6c

Observation 40ae999b-8da0-42a5-b001-6d340fead544 · outbound

This paper cites $f$-PO: Generalizing Preference Optimization with $f$-divergence Minimization.

Preference learning made easy: Everything should be understood through win rate $f$-PO: Generalizing Preference Optimization with $f$-divergence Minimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.028109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.028109Z digest=sha256:0843bd32d5779aa08ac93458167533aa4f18a58b61d9892487c180671ef4ab5a

Observation e3f076e6-6916-4baf-95d1-94ee9fcb0d54 · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Preference learning made easy: Everything should be understood through win rate Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.033460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.033460Z digest=sha256:72a467042999069da9c1f138352e359f6e6062f50d5fb42c85730891bc551956

Observation 52740ac5-432c-4efb-870d-81ff9b8c7ff2 · outbound

This paper cites A., Choi, Y., and Hajishirzi, H.

Preference learning made easy: Everything should be understood through win rate A., Choi, Y., and Hajishirzi, H

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.766300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.038358Z digest=sha256:e037f87e0ef05cfdb8ded71728b7b303a94c9287f91e3f80145609f9e567acd0

Observation 42270018-9ef5-4abf-b241-7186c0528296 · outbound

This paper cites Towards Efficient Exact Optimization of Language Model Alignment.

Preference learning made easy: Everything should be understood through win rate Towards Efficient Exact Optimization of Language Model Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.045898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.045898Z digest=sha256:7e4c4269e49a78b847c43f7726de69165923d6af19a4f32085a1c0a9b2d94a73

Observation eb3b3895-c27c-4a72-9c09-74f8b2d8a384 · outbound

This paper cites A Survey on Human Preference Learning for Large Language Models.

Preference learning made easy: Everything should be understood through win rate A Survey on Human Preference Learning for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.053077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.053077Z digest=sha256:194d1b01ecc63839aed86b6d188d1e6708427301b5928d235c2456ad6c3cab64

Observation f6ba482e-7eed-4cf0-ba69-08f0200da95b · outbound

This paper cites A survey of reinforcement learning from human feedback.

Preference learning made easy: Everything should be understood through win rate A survey of reinforcement learning from human feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.059995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.059995Z digest=sha256:9f9e87fff033eacf956a597e8ebebceb0963d884928fba0c7c2963cdfd795624

Observation 15754356-b360-4f9e-948e-14bdfb38a397 · outbound

This paper cites Understanding the effects of rlhf on llm generalisation and diversity.

Preference learning made easy: Everything should be understood through win rate Understanding the effects of rlhf on llm generalisation and diversity

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.729922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.066417Z digest=sha256:da9ee34801af7ce9e547890bfb1855ecd0352fe6489bcee1e834fa97e0429c23

Observation d73ad913-60ec-4daa-947a-c504d2031a5d · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

Preference learning made easy: Everything should be understood through win rate OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.072721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.072721Z digest=sha256:b24de483c688d88bd14c16d859d93706e2b5b4b01523063259a0f303740fe7c4

Observation 103738c8-633e-451f-be24-7474252aa77b · outbound

This paper cites Huggingface h4 stack exchange preference dataset.

Preference learning made easy: Everything should be understood through win rate Huggingface h4 stack exchange preference dataset

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.710240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.080097Z digest=sha256:59e1a21c444ddd234d833c4cb51e6e08107fe951a041acadb26caa542ba61fdf

Observation 367c276c-8757-48fc-ba82-3f127d3ae281 · outbound

This paper cites Self-Alignment with Instruction Backtranslation.

Preference learning made easy: Everything should be understood through win rate Self-Alignment with Instruction Backtranslation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.086641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.086641Z digest=sha256:8b4a060adff06da882ee1901e3b4d3c06e7428a55a16e11c87db5e62703a98ad

Observation 959bff62-81b5-428d-a8df-db9dabe0082d · outbound

This paper cites an unresolved cited work.

Preference learning made easy: Everything should be understood through win rate Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:32:22.686595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.093625Z digest=sha256:5f1ebbe9d187db7979ad98aaa296701d925871d704eebc10c5231e6eee3a47ff

Observation 14b5bd41-343a-4f2a-b5ca-7ae15d2060c1 · outbound

This paper cites Large Language Models: A Survey.

Preference learning made easy: Everything should be understood through win rate Large Language Models: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.100629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.100629Z digest=sha256:1ccafd43cdbc06a3cade68b0a1b5579842a3fdce2d0971dfcb606b12505d607c

Observation 74293d31-b9a8-4b0c-98c0-135fa23159a2 · outbound

This paper cites Monte carlo gradient estimation in machine learning.

Preference learning made easy: Everything should be understood through win rate Monte carlo gradient estimation in machine learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.663308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.113252Z digest=sha256:97a0f97a906bc6be1821fc481c0fefec7cdaa9601fda02f3fb99647f7be02c63

Observation 39550b1e-bb78-4ec1-85ed-2813f2bada54 · outbound

This paper cites Monte carlo gradient estimation in machine learning.

Preference learning made easy: Everything should be understood through win rate Monte carlo gradient estimation in machine learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.641721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.119438Z digest=sha256:5753d09bed9ae06c6dce31154c22c11eae5b3a07b4142369325283da45d869ab

Observation 465c40fd-57e9-44af-9f9e-309cf5c2cfec · outbound

This paper cites G., Rowland, M., Guo, Z.

Preference learning made easy: Everything should be understood through win rate G., Rowland, M., Guo, Z

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.621373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.125490Z digest=sha256:2ed54ac50265fafe1669bfadcf20f74521806d046be8db07c03774b401d5a390

Observation 20c613a0-c772-4fd1-86f1-899de5cc54a6 · outbound

This paper cites A., Lindsten, F., and Blei, D.

Preference learning made easy: Everything should be understood through win rate A., Lindsten, F., and Blei, D

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.601091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.131839Z digest=sha256:1af85c5edac1a1c292c573673037e6aa16171b21df5bf590ed15ca1199e7c2ba

Observation 5336bf60-2085-4690-9267-135c1e51011f · outbound

This paper cites Gpt-4 technical report.

Preference learning made easy: Everything should be understood through win rate Gpt-4 technical report

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.582335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.139010Z digest=sha256:56298362505c2c35c642094b809d6966c15a14b2fe0edef171227489732bf2a6

Observation 816e2397-35a3-4c92-9c0f-9672f797b1e2 · outbound

This paper cites Training language models to follow instructions with human feedback.

Preference learning made easy: Everything should be understood through win rate Training language models to follow instructions with human feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.145340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.145340Z digest=sha256:e01853338dadd9992fe82c7146d0b89ec0ecc063bbf1ab608652118811b6edc7

Observation fd5fa80e-fc39-44b0-87a1-6a881f593f98 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Preference learning made easy: Everything should be understood through win rate Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.154929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.154929Z digest=sha256:37111e0e22be39b811211e50576c0e6aa396e9fdd8612e02bdbb363020f07b8d

Observation 85b9a4a5-97aa-499d-8585-18d3b2556cd4 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

Preference learning made easy: Everything should be understood through win rate Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.165463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.165463Z digest=sha256:16817298fe416f56c4e4b90082cc45c46e239f38034a9b561db0d4d0c62e2fed

Observation f5b4a645-6411-48d5-a1b2-a2dc0a364c3c · outbound

This paper cites From r to q^* : Your language model is secretly a q-function.

Preference learning made easy: Everything should be understood through win rate From r to q^* : Your language model is secretly a q-function

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.549800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.171588Z digest=sha256:7b3f94a3c9ac9c050d205aa94f23cc833f7a808e6383902b0180e94bc993fd81

Observation 1b439a24-3a82-4014-9314-a15db84d1e18 · outbound

This paper cites D., Ermon, S., and Finn, C.

Preference learning made easy: Everything should be understood through win rate D., Ermon, S., and Finn, C

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.528340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.177265Z digest=sha256:480605bc44e2cb1460fc86de679fc43cf77cfe56024c298a7212c70f06b7b204

Observation 05e21823-4197-4c1b-8316-77d5cd106f45 · outbound

This paper cites Black box variational inference.

Preference learning made easy: Everything should be understood through win rate Black box variational inference

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.507183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.183300Z digest=sha256:82358d5ec6d0ddb7549983e720efc5cf16f0bf7698aa8e146b5b46af097efb56

Observation 2b6830b4-152f-418b-afd5-34dffd1e055f · outbound

This paper cites Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization.

Preference learning made easy: Everything should be understood through win rate Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.189249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.189249Z digest=sha256:9fc2d69084536e27ad7041f50751d37cd88c35abb5e079cb64e57184db2c19cc

Observation a389f205-5390-4018-9842-8d8291d01b0c · outbound

This paper cites Vanishing gradients in reinforcement finetuning of language models.

Preference learning made easy: Everything should be understood through win rate Vanishing gradients in reinforcement finetuning of language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.481360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.195177Z digest=sha256:14a829eaee4d7e8dbafe506ed22480c2d4af0a9a28140bce9f8c09e0c6d9b95d

Observation cacef77f-dca8-49dd-a1ec-72cd0c5e9daf · outbound

This paper cites Direct nash optimization: Teaching language models to self-improve with general preferences, 2024.

Preference learning made easy: Everything should be understood through win rate Direct nash optimization: Teaching language models to self-improve with general preferences, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.452006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.200061Z digest=sha256:00c6d2a83ff8628dc74abb788822e98b3f3a1f2c5d54047b5aec7825aaf81667

Observation 64519d34-4c42-4d2f-a260-b89f02cfed61 · outbound

This paper cites Direct Alignment with Heterogeneous Preferences.

Preference learning made easy: Everything should be understood through win rate Direct Alignment with Heterogeneous Preferences

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.206355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.206355Z digest=sha256:8b1d476d63a381d74faec57b29edde09fd47110f3c6ffd6c0b35f70a46ada8eb

Observation 823f03aa-ab4a-4e29-9d59-69a72139f1aa · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Preference learning made easy: Everything should be understood through win rate A Long Way to Go: Investigating Length Correlations in RLHF

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.215175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.215175Z digest=sha256:6c32a3bdfa425571b3fa3d49c57c8b07ffecca85173b3a830b8ae1ad3a7311e8

Observation fea3bb9f-372c-4b55-89ef-4ac8e716d64b · outbound

This paper cites Distributional preference learning: Understanding and accounting for hidden context in rlhf.

Preference learning made easy: Everything should be understood through win rate Distributional preference learning: Understanding and accounting for hidden context in rlhf

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.424473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.221926Z digest=sha256:1626b141f6c3f870eaccc90ba0a52744f2a8a46bbe37f0f648ad05367ecf91cc

Observation d577a9f7-535e-42c8-8c1a-7fde56c8e413 · outbound

This paper cites How to compare different loss functions and their risks.

Preference learning made easy: Everything should be understood through win rate How to compare different loss functions and their risks

Reference 46

Resolution
verified exact
doi, observed 2026-08-07T18:32:21.406280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.227381Z digest=sha256:a6185c0d0f8e5f14923fb8a175817478b9db6c2579a3c33ec17b1c7e9e6b9799

Observation 3b65fc6e-bc9e-44f5-a051-1c4c21f62af3 · outbound

This paper cites Learning to summarize from human feedback.

Preference learning made easy: Everything should be understood through win rate Learning to summarize from human feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.232480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.232480Z digest=sha256:5937ae5b4912fb483309d9b4ed5154bb31608025595aa5d676ca1b87a20c3a5c

Observation a9667a6d-961a-4104-b483-ad8b2d1fbb5f · outbound

This paper cites S., and Agarwal, A.

Preference learning made easy: Everything should be understood through win rate S., and Agarwal, A

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.394366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.237661Z digest=sha256:42eb1b334e79de744b19831f29d0dc42496e800f64f9e422b8f6f200d8f53f7c

Observation 59644008-f9e1-4566-9f7d-bb638ceb40cd · outbound

This paper cites Preference fine-tuning of llms should leverage suboptimal, on-policy data.

Preference learning made easy: Everything should be understood through win rate Preference fine-tuning of llms should leverage suboptimal, on-policy data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.375130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.243309Z digest=sha256:3e8ca88a1ca72c2586a0bb3f42c4158ee6708010c6447e37fd42c81d8f04ae74

Observation a8b1c143-4b83-49ea-bf47-4462a0fc4edb · outbound

This paper cites D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P.

Preference learning made easy: Everything should be understood through win rate D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.348188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.250115Z digest=sha256:2271fc2fa042b1a8aeebda95c47ff4d0d6b654d9413b299835fe51e4e4ff3e95

Observation 8a4e59a3-d701-4d99-98be-dd44878199ec · outbound

This paper cites Trl: Transformer reinforcement learning.

Preference learning made easy: Everything should be understood through win rate Trl: Transformer reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.322709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.259111Z digest=sha256:d90f17bc05aaf2b84c7def51dd342c0f6638091e9a9dec8cf52e37928f8068d6

Observation 14b72f64-ca56-407f-aee4-3b0792b5404b · outbound

This paper cites Enabling language models to implicitly learn self-improvement.

Preference learning made easy: Everything should be understood through win rate Enabling language models to implicitly learn self-improvement

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.297854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.265704Z digest=sha256:bd5e4e8750d1a1fef1d07db95addaae2237ecf1b65ad8369d8f91a8847992876

Observation 622b1c22-0d0e-4116-9b0c-ca6b54aefe74 · outbound

This paper cites Transforming and combining rewards for aligning large language models.

Preference learning made easy: Everything should be understood through win rate Transforming and combining rewards for aligning large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.261958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.273577Z digest=sha256:679a9e937cf71e5344b944e130be302434441ce37cbe8c424fea83e248c666e7

Observation 14cf1242-2bb1-4135-80d9-aa4b0aefe821 · outbound

This paper cites Policy gradient algorithms.

Preference learning made easy: Everything should be understood through win rate Policy gradient algorithms

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.238493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.282655Z digest=sha256:f9b610016fd016c6eb39178ab848eecc37258415da0e8015b1f85da692b93cea

Observation 74368c74-59ad-45f1-a388-e08c26b885b0 · outbound

This paper cites V., Murray, K., and Kim, Y.

Preference learning made easy: Everything should be understood through win rate V., Murray, K., and Kim, Y

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.201317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.289186Z digest=sha256:7d4519a3e74c4559b69618352ea991d3ae54e2e048b5c2ae17db4e5f18e49eae

Observation 97ad70a4-9085-4d57-ad7e-f3d6738836c3 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Preference learning made easy: Everything should be understood through win rate Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.300233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.300233Z digest=sha256:d5ee4b40a83521492e771fac633ce487892e68bc2864e8423bba5e5cbeba426f

Observation 00f98043-ba0a-4e26-a390-a345ade3e382 · outbound

This paper cites Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J.

Preference learning made easy: Everything should be understood through win rate Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.175500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.307867Z digest=sha256:0ae474489e9be267fe91ad88400e3ded4b840da8fc47bc8825fb5f001bace3d5

Observation 5e8863dd-e6f4-4a5c-8ebe-e6fde31070dd · outbound

This paper cites an unresolved cited work.

Preference learning made easy: Everything should be understood through win rate Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:32:22.143704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.312919Z digest=sha256:61cf1b542c4f5ed291abe5d4cb40fb5865fff2294f9409567eca8343985a9dd8

Observation 862eb67f-a726-4035-9669-b97892d90489 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Preference learning made easy: Everything should be understood through win rate Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.120972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.318779Z digest=sha256:e520faba616f4e072bccfd9392c6b4773ca005295f45d3fd0134e58ae7ede3f0

Observation e9610168-ebbd-4b38-b0c3-8b7fab8f24bc · outbound

This paper cites I., and Jiao, J.

Preference learning made easy: Everything should be understood through win rate I., and Jiao, J

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.100251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.325381Z digest=sha256:137f92ebf5463030765042c694783af3a3aa2de92dfb94fe6b8a78565a871477

Observation 4419939a-a0f5-49a7-b640-ecc4d607b2e2 · outbound

This paper cites write newline.

Preference learning made easy: Everything should be understood through win rate write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.331060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.331060Z digest=sha256:bd6b01ed665df98d9b691a0609194b04cb1dff614b9420c3e3cbc3b8ffa760d3

Pith citing papers

No inbound Pith citation observations are available.