Pith. sign in

Paper Citation Record · LEDGER

Preference learning made easy: Everything should be understood through win rate

As of 19 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2502.10505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10505 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:32:21.331060Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy36
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6180138-a454-4301-9e33-643e150b0bb8 · outbound

This paper cites Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms.

Preference learning made easy: Everything should be understood through win rate Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:23.028818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.903864Z digest=sha256:23b90d1574b251085b927938e49b7b93e8ade2d1dad25d0e19a2269d72dd0785

Observation 1b074722-1adc-408b-b4cb-9b3f3040dc52 · outbound

This paper cites Calibration and consistency of adversarial surrogate losses.

Preference learning made easy: Everything should be understood through win rate Calibration and consistency of adversarial surrogate losses

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:23.007001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.910320Z digest=sha256:23cbd3d84e9eac3aeda8bbafb53e804a0584eb29a1cbaa1304e3c07fcd7e69ce

Observation 4c251466-16c3-4b63-b3ac-5d2ad985737c · outbound

This paper cites Multi-class h -consistency bounds.

Preference learning made easy: Everything should be understood through win rate Multi-class h -consistency bounds

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.983272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.916903Z digest=sha256:06c39054f3265a8ab5d69a28d5bcc5cf3057309489e2b1e862c5fa3d36e196a6

Observation 4c2cc416-e36f-4b17-94ca-93af7a38517c · outbound

This paper cites G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R.

Preference learning made easy: Everything should be understood through win rate G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.965345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.924819Z digest=sha256:c3c21d7c90f8057f3ab4293a13e548bd0422900dfc92d0886bf0db8f36152c4b

Observation d8ab9d49-91d5-49cd-a64c-2dc628851e04 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback.

Preference learning made easy: Everything should be understood through win rate Training a helpful and harmless assistant with reinforcement learning from human feedback

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.938806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.932361Z digest=sha256:8ca38e5f4ecfb2660ea6a46db5e1d091d7720c19047b9a04dd511e5a3f7068ce

Observation bad98efd-8a4b-48f9-9230-63a8f1eee886 · outbound

This paper cites Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.

Preference learning made easy: Everything should be understood through win rate Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:20.938380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:20.938380Z digest=sha256:bb33a35e9fc4e674797ab7eb174c70d360e9cacde2456186d42425610ae12c25

Observation 5e77688e-f12f-4247-b719-10fae9e6a0b8 · outbound

This paper cites Quantile Filtered Imitation Learning.

Preference learning made easy: Everything should be understood through win rate Quantile Filtered Imitation Learning

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T18:32:22.012736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.946064Z digest=sha256:dd43d486e71f36850ffe09ec268e69c66f89587e66436a1f0879e6903a0cf897

Observation 43afe68a-4ec2-412e-85dc-3230471c94a9 · outbound

This paper cites Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition.

Preference learning made easy: Everything should be understood through win rate Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.919649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.952817Z digest=sha256:bc470aba4988f1456645cd9d167aa2cf4d486b3d49b56efd584cc400578833b2

Observation 9f07e96b-2159-4740-a1d1-875239dda663 · outbound

This paper cites Human Alignment of Large Language Models through Online Preference Optimisation.

Preference learning made easy: Everything should be understood through win rate Human Alignment of Large Language Models through Online Preference Optimisation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:20.963680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:20.963680Z digest=sha256:0474f056d5aa69f8a12e0ce0ca25d44dbedf07098a0ccef13fc2b7255e9848a2

Observation 8107c9da-f2c5-40f5-8b2f-330139a41e75 · outbound

This paper cites H., Chen, X., Zhang, Q., Ranganath, R., and Cho, K.

Preference learning made easy: Everything should be understood through win rate H., Chen, X., Zhang, Q., Ranganath, R., and Cho, K

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.898525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.971802Z digest=sha256:9aac28485ddec124a22485848c3d10bdd86d854ef1cf23720a941aef8839b5ea

Observation 2902befc-3e72-43b6-acfe-88eb4cef412d · outbound

This paper cites Deep reinforcement learning from human preferences.

Preference learning made easy: Everything should be understood through win rate Deep reinforcement learning from human preferences

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:20.983164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:20.983164Z digest=sha256:fc55feeb4f3212cd4a494f739773608f5880c252e29468b442840d0e2ffbc40a

Observation b619d067-3659-4495-ad93-2921e730fb13 · outbound

This paper cites Raft: Reward ranked finetuning for generative foundation model alignment.

Preference learning made easy: Everything should be understood through win rate Raft: Reward ranked finetuning for generative foundation model alignment

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.864549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:20.989748Z digest=sha256:5c93a18bf6347af8fa6c4370511e4dba3fe9c2fc31e5a873d54ef625571c2bad

Observation 51dd05ae-79b1-41fc-ba02-c0308782ba0f · outbound

This paper cites Understanding dataset difficulty with v-usable information.

Preference learning made easy: Everything should be understood through win rate Understanding dataset difficulty with v-usable information

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.839675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.001625Z digest=sha256:027ec6d8616a86e25f6d3064ee2b629521410bae656e14438b9e4d1dcdbc18bd

Observation c5d841ae-b317-4623-9930-c2b9d97c89b4 · outbound

This paper cites Bonbon alignment for large language models and the sweetness of best-of-n sampling.

Preference learning made easy: Everything should be understood through win rate Bonbon alignment for large language models and the sweetness of best-of-n sampling

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.814056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.008913Z digest=sha256:fdd678e76c5900b999b799a6774760adbdd7b51444a2703dde5bc2ce2a25d24c

Observation b64d9244-8006-498c-96d9-4831e68de4b7 · outbound

This paper cites L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., and de Freitas, N.

Preference learning made easy: Everything should be understood through win rate L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., and de Freitas, N

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.788839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.015499Z digest=sha256:d7a89bc8a05571c8dac9f4130ce5275fdb99e797ff66c150a1f1bf07d60757e7

Observation 5724ed26-706e-4021-a240-b524f4676b17 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Preference learning made easy: Everything should be understood through win rate Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.022389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.022389Z digest=sha256:b976ab85e27426d35d23b66799034b3784bc13db754d9a694bea58db2c9ee19c

Observation 40ae999b-8da0-42a5-b001-6d340fead544 · outbound

This paper cites $f$-PO: Generalizing Preference Optimization with $f$-divergence Minimization.

Preference learning made easy: Everything should be understood through win rate $f$-PO: Generalizing Preference Optimization with $f$-divergence Minimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.028109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.028109Z digest=sha256:33f152a76a662d9a75bab0b9eb22c79a02758b9d065c558412d79516e6d56ce2

Observation e3f076e6-6916-4baf-95d1-94ee9fcb0d54 · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Preference learning made easy: Everything should be understood through win rate Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.033460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.033460Z digest=sha256:2b83fdb5067b2605793a959cd6a485a293df9ce1429e19c88bf536ca8c7c2f9b

Observation 52740ac5-432c-4efb-870d-81ff9b8c7ff2 · outbound

This paper cites A., Choi, Y., and Hajishirzi, H.

Preference learning made easy: Everything should be understood through win rate A., Choi, Y., and Hajishirzi, H

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.766300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.038358Z digest=sha256:d4968153ffe12cd854edbb5f7f79c65265c31f2bb2be3b9dc128f43543d68b45

Observation 42270018-9ef5-4abf-b241-7186c0528296 · outbound

This paper cites Towards Efficient Exact Optimization of Language Model Alignment.

Preference learning made easy: Everything should be understood through win rate Towards Efficient Exact Optimization of Language Model Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.045898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.045898Z digest=sha256:f4f80b2d4d9c7e66e6271c699d5561228b4fcbe5a5c169bef9f4bb0e19e1beac

Observation eb3b3895-c27c-4a72-9c09-74f8b2d8a384 · outbound

This paper cites A Survey on Human Preference Learning for Large Language Models.

Preference learning made easy: Everything should be understood through win rate A Survey on Human Preference Learning for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.053077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.053077Z digest=sha256:027bfd262b6a6370caf82b431557be314e7d866ca63cf4854fb5065ea0bb617a

Observation f6ba482e-7eed-4cf0-ba69-08f0200da95b · outbound

This paper cites A survey of reinforcement learning from human feedback.

Preference learning made easy: Everything should be understood through win rate A survey of reinforcement learning from human feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.059995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.059995Z digest=sha256:555d00cbc3de15cbc5db5dbf49ac13491ace52664f9f050e118dfe6c96319f56

Observation 15754356-b360-4f9e-948e-14bdfb38a397 · outbound

This paper cites Understanding the effects of rlhf on llm generalisation and diversity.

Preference learning made easy: Everything should be understood through win rate Understanding the effects of rlhf on llm generalisation and diversity

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.729922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.066417Z digest=sha256:afdcf58d7b4b090cb54e1362aeb1855b399460d67f72133112d2aaed8c8fa137

Observation d73ad913-60ec-4daa-947a-c504d2031a5d · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

Preference learning made easy: Everything should be understood through win rate OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.072721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.072721Z digest=sha256:640ef0833640eef07e88de21b7962ff4a37c414835f5295225f4c95e3c7ff288

Observation 103738c8-633e-451f-be24-7474252aa77b · outbound

This paper cites Huggingface h4 stack exchange preference dataset.

Preference learning made easy: Everything should be understood through win rate Huggingface h4 stack exchange preference dataset

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.710240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.080097Z digest=sha256:531ef5bf03c0875092a09ac4a301cc23ca5fbac9e65d982014ff2110cdcf3d27

Observation 367c276c-8757-48fc-ba82-3f127d3ae281 · outbound

This paper cites Self-Alignment with Instruction Backtranslation.

Preference learning made easy: Everything should be understood through win rate Self-Alignment with Instruction Backtranslation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.086641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.086641Z digest=sha256:e21d7333b0624573a009ad23f2df2a4f8f8f10633300c525f619b326dbc6fb8c

Observation 959bff62-81b5-428d-a8df-db9dabe0082d · outbound

This paper cites an unresolved cited work.

Preference learning made easy: Everything should be understood through win rate Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:32:22.686595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.093625Z digest=sha256:4155c23f1bfefd26a166bd5485b4d03ed4e535d17fc2c26cd1e36d7e9b4e9c2e

Observation 14b5bd41-343a-4f2a-b5ca-7ae15d2060c1 · outbound

This paper cites Large Language Models: A Survey.

Preference learning made easy: Everything should be understood through win rate Large Language Models: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.100629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.100629Z digest=sha256:648522c8ff948d53c933949a969fa0a7fd1c144aa2e6bbf6f904f57ec1f389ef

Observation 74293d31-b9a8-4b0c-98c0-135fa23159a2 · outbound

This paper cites Monte carlo gradient estimation in machine learning.

Preference learning made easy: Everything should be understood through win rate Monte carlo gradient estimation in machine learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.663308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.113252Z digest=sha256:3bb401493b02942877c88cf961ca5da9a2d7815b8df83bb4d424516748eae200

Observation 39550b1e-bb78-4ec1-85ed-2813f2bada54 · outbound

This paper cites Monte carlo gradient estimation in machine learning.

Preference learning made easy: Everything should be understood through win rate Monte carlo gradient estimation in machine learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.641721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.119438Z digest=sha256:3e9264a5f77c2246e5caded576049a5831870234d953f28fa2ab7a42a722d04c

Observation 465c40fd-57e9-44af-9f9e-309cf5c2cfec · outbound

This paper cites G., Rowland, M., Guo, Z.

Preference learning made easy: Everything should be understood through win rate G., Rowland, M., Guo, Z

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.621373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.125490Z digest=sha256:d249285390e5cfecb1197c41a204a94a0effd8d03bd401edb550b477242613e7

Observation 20c613a0-c772-4fd1-86f1-899de5cc54a6 · outbound

This paper cites A., Lindsten, F., and Blei, D.

Preference learning made easy: Everything should be understood through win rate A., Lindsten, F., and Blei, D

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.601091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.131839Z digest=sha256:b2e08858fdb431624bc779ae53e076066b89b62c93d4d41dfd0adc25c83d6ba5

Observation 5336bf60-2085-4690-9267-135c1e51011f · outbound

This paper cites Gpt-4 technical report.

Preference learning made easy: Everything should be understood through win rate Gpt-4 technical report

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.582335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.139010Z digest=sha256:fa8e848c2bb2b7cadee7680602b85192e3693d1e43ce1cf048aceeed4d41a4bc

Observation 816e2397-35a3-4c92-9c0f-9672f797b1e2 · outbound

This paper cites Training language models to follow instructions with human feedback.

Preference learning made easy: Everything should be understood through win rate Training language models to follow instructions with human feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.145340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.145340Z digest=sha256:e9fcf8fdf91cb4cb3f6f430d15472152181b8ccbf3e6353ad45afe8325384605

Observation fd5fa80e-fc39-44b0-87a1-6a881f593f98 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Preference learning made easy: Everything should be understood through win rate Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.154929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.154929Z digest=sha256:847994280acfa829825a411579c47bf1886172846b167e9b4fa7763cb38e593c

Observation 85b9a4a5-97aa-499d-8585-18d3b2556cd4 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

Preference learning made easy: Everything should be understood through win rate Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.165463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.165463Z digest=sha256:7b173a620a710929b71c3bc53ec78ce6f625ad2b2ea81b2ca7449e64e0db86a8

Observation f5b4a645-6411-48d5-a1b2-a2dc0a364c3c · outbound

This paper cites From r to q^* : Your language model is secretly a q-function.

Preference learning made easy: Everything should be understood through win rate From r to q^* : Your language model is secretly a q-function

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.549800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.171588Z digest=sha256:cbf6b7e876dfcb4a1b31c9017fc5001992866100f14c5046926ab5f1089afafd

Observation 1b439a24-3a82-4014-9314-a15db84d1e18 · outbound

This paper cites D., Ermon, S., and Finn, C.

Preference learning made easy: Everything should be understood through win rate D., Ermon, S., and Finn, C

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.528340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.177265Z digest=sha256:b45ba3831f569876ce2420896a62bf9f9aa1f5f4f9ef5e1ca98f44d4e3bc03ba

Observation 05e21823-4197-4c1b-8316-77d5cd106f45 · outbound

This paper cites Black box variational inference.

Preference learning made easy: Everything should be understood through win rate Black box variational inference

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.507183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.183300Z digest=sha256:cbfe215f5517c29abc589c9415182ba624551cbfc2652eccabae02ffa6299588

Observation 2b6830b4-152f-418b-afd5-34dffd1e055f · outbound

This paper cites Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization.

Preference learning made easy: Everything should be understood through win rate Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.189249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.189249Z digest=sha256:3f792c6daf29ba9e61ac5e9d1d00f0632ce847dbd20f2b0825db9af9036d3ad4

Observation a389f205-5390-4018-9842-8d8291d01b0c · outbound

This paper cites Vanishing gradients in reinforcement finetuning of language models.

Preference learning made easy: Everything should be understood through win rate Vanishing gradients in reinforcement finetuning of language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.481360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.195177Z digest=sha256:7582393befba646ce955b23890361266a57f3a2c54addc213222316469978db6

Observation cacef77f-dca8-49dd-a1ec-72cd0c5e9daf · outbound

This paper cites Direct nash optimization: Teaching language models to self-improve with general preferences, 2024.

Preference learning made easy: Everything should be understood through win rate Direct nash optimization: Teaching language models to self-improve with general preferences, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.452006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.200061Z digest=sha256:770e00fed2f397b0f80a533d43b8da52b549e8011388784a060fc4467ae38f00

Observation 64519d34-4c42-4d2f-a260-b89f02cfed61 · outbound

This paper cites Direct Alignment with Heterogeneous Preferences.

Preference learning made easy: Everything should be understood through win rate Direct Alignment with Heterogeneous Preferences

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.206355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.206355Z digest=sha256:f6e13b1328aec58f5dd0ae987262286e1ca4bea6a92c8c3055c32c5454fbcd29

Observation 823f03aa-ab4a-4e29-9d59-69a72139f1aa · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Preference learning made easy: Everything should be understood through win rate A Long Way to Go: Investigating Length Correlations in RLHF

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.215175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.215175Z digest=sha256:567032f61483d94a8b1e132e33d3e50c26803376969a9be5f9d6ad30121e4728

Observation fea3bb9f-372c-4b55-89ef-4ac8e716d64b · outbound

This paper cites Distributional preference learning: Understanding and accounting for hidden context in rlhf.

Preference learning made easy: Everything should be understood through win rate Distributional preference learning: Understanding and accounting for hidden context in rlhf

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.424473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.221926Z digest=sha256:3e15a33c6ceec94c73c5eae8c74d9f8f611d5fe4a544fed7a44ec5ef6034f4f8

Observation d577a9f7-535e-42c8-8c1a-7fde56c8e413 · outbound

This paper cites How to compare different loss functions and their risks.

Preference learning made easy: Everything should be understood through win rate How to compare different loss functions and their risks

Reference 46

Resolution
verified exact
doi, observed 2026-08-07T18:32:21.406280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.227381Z digest=sha256:872f9071be8909f5fba8c54533c9685a7c541b6829be2fe177b348345e01139a

Observation 3b65fc6e-bc9e-44f5-a051-1c4c21f62af3 · outbound

This paper cites Learning to summarize from human feedback.

Preference learning made easy: Everything should be understood through win rate Learning to summarize from human feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.232480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.232480Z digest=sha256:5134ab3f18dec9c1607cfd2f77e81776ded9c46faa14a8862c2a3583db3186ce

Observation a9667a6d-961a-4104-b483-ad8b2d1fbb5f · outbound

This paper cites S., and Agarwal, A.

Preference learning made easy: Everything should be understood through win rate S., and Agarwal, A

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.394366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.237661Z digest=sha256:dbaf8d0ccc1805e7858372090fb2de394c6a90bde0485bafdd21a6dfa5756753

Observation 59644008-f9e1-4566-9f7d-bb638ceb40cd · outbound

This paper cites Preference fine-tuning of llms should leverage suboptimal, on-policy data.

Preference learning made easy: Everything should be understood through win rate Preference fine-tuning of llms should leverage suboptimal, on-policy data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.375130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.243309Z digest=sha256:27cbe65ddf80b5c794922904b3041299dcd47fd9b8916f77d02ddf1682b72559

Observation a8b1c143-4b83-49ea-bf47-4462a0fc4edb · outbound

This paper cites D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P.

Preference learning made easy: Everything should be understood through win rate D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.348188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.250115Z digest=sha256:3b9d7d7b9b32016fcc2e9033725b83555b43e0d95bc53981c97173b147e131b1

Observation 8a4e59a3-d701-4d99-98be-dd44878199ec · outbound

This paper cites Trl: Transformer reinforcement learning.

Preference learning made easy: Everything should be understood through win rate Trl: Transformer reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.322709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.259111Z digest=sha256:e5894c37d159c5fece211ea85fc2af62922263ceeccb237f3c795abbd4202e89

Observation 14b72f64-ca56-407f-aee4-3b0792b5404b · outbound

This paper cites Enabling language models to implicitly learn self-improvement.

Preference learning made easy: Everything should be understood through win rate Enabling language models to implicitly learn self-improvement

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.297854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.265704Z digest=sha256:7f266339c8cb87c8b4cb9b1fd56161776c09c55a7604092a983a6af2b7a36bb4

Observation 622b1c22-0d0e-4116-9b0c-ca6b54aefe74 · outbound

This paper cites Transforming and combining rewards for aligning large language models.

Preference learning made easy: Everything should be understood through win rate Transforming and combining rewards for aligning large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.261958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.273577Z digest=sha256:fd930aa331e5c22da1a012af382fb6a5ea482eb8606f892c4a7bea100eb04d55

Observation 14cf1242-2bb1-4135-80d9-aa4b0aefe821 · outbound

This paper cites Policy gradient algorithms.

Preference learning made easy: Everything should be understood through win rate Policy gradient algorithms

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.238493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.282655Z digest=sha256:99e76e7b1a5e957ee716f61fea1ca1755b5d1e47e046604309d77d94245803b9

Observation 74368c74-59ad-45f1-a388-e08c26b885b0 · outbound

This paper cites V., Murray, K., and Kim, Y.

Preference learning made easy: Everything should be understood through win rate V., Murray, K., and Kim, Y

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.201317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.289186Z digest=sha256:42516654d81e0777755c15397b0988846cc729c08723372330c7df4b73d20b99

Observation 97ad70a4-9085-4d57-ad7e-f3d6738836c3 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Preference learning made easy: Everything should be understood through win rate Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.300233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.300233Z digest=sha256:301e7758a8b7eaaf6ae81cd16150e66ba8c14d8a44daec0de50b2a847fd7e7e1

Observation 00f98043-ba0a-4e26-a390-a345ade3e382 · outbound

This paper cites Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J.

Preference learning made easy: Everything should be understood through win rate Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.175500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.307867Z digest=sha256:5c905dfff587edf4e1c0053259f65f1566ccc7df922ff0fa5a31428ba0d769d1

Observation 5e8863dd-e6f4-4a5c-8ebe-e6fde31070dd · outbound

This paper cites an unresolved cited work.

Preference learning made easy: Everything should be understood through win rate Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:32:22.143704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.312919Z digest=sha256:4cb72d7ab6568dca00b14951f57b0965b6f0cd62c52fcb003248d27eca1a87ef

Observation 862eb67f-a726-4035-9669-b97892d90489 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Preference learning made easy: Everything should be understood through win rate Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.120972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.318779Z digest=sha256:7447737932238f69c618a1f2dac2c03cb81a9ca909ea14e3309f5a0f94a8a9d2

Observation e9610168-ebbd-4b38-b0c3-8b7fab8f24bc · outbound

This paper cites I., and Jiao, J.

Preference learning made easy: Everything should be understood through win rate I., and Jiao, J

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:32:22.100251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T18:32:21.325381Z digest=sha256:f2b219a26cbeef3246267bad04d193162ff0a0de8fd6684a4e5ed036b07d6f99

Observation 4419939a-a0f5-49a7-b640-ecc4d607b2e2 · outbound

This paper cites write newline.

Preference learning made easy: Everything should be understood through win rate write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:21.331060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:21.331060Z digest=sha256:f5d67ea7181c31553c7f2a1642fb6aa7e263cdaad1d6be37a5144bbc12f9d8d5

Pith citing papers

No inbound Pith citation observations are available.